Abstract
Speech technology applications have emerged as a promising method for assessing speech-language abilities and at-home therapy, including prosody. Many applications assume that observed prosody errors are due to an underlying disorder; however, they may be instead due to atypical representations of prosody such as immature and developing speech motor control, or compensatory adaptations by those with congenital neuromotor disorders. The result is the same – vocal productions may not be a reliable measure of prosody knowledge. Therefore, in this study we examine the usability of a new technology application to express prosody knowledge without relying on vocalizations using the Prosodic Marionette (PM) graphical user interface for artificial resynthesis of speech prosody. We tested the ability of neurotypical participants to use the PM interface to control prosody through 2D movements of word-icon blocks vertically (fundamental frequency), horizontally (pause length), and by stretching (word duration) to correctly mark target prosodic contrasts. Nearly all participants used vertical movements to correctly mark fundamental frequency changes where appropriate (e.g., raised second word for pitch accent on second word). A smaller percentage of participants used the stretching feature to mark duration changes; when used, participants correctly lengthened the appropriate word (e.g., stretch the second item to accent the second word). Our results suggest the PM interface can be used reliably to correctly signal speech prosody, which validates future use of the interface to assess prosody in clinical and developmental populations with atypical speech motor control.
Keywords: prosody, speech motor control, speech production, acoustics
1. Introduction
Linguistic prosody conveys many important aspects of speech including stress, rhythm, intonation, and phrase structure, as speakers alter fundamental frequency, pause length, word duration, and speech intensity (Bolinger, 1989, Lehiste, 1970, 1976, Shattuck-Hufnagel and Turk, 1996). Impairments to speech production can alter not just segmental speech, but prosody as well, and lead to loss of intelligibility and naturalness. In fact, disordered prosody has been considered a key component of dysarthria and used for clinical assessment and diagnosis (Darley et al., 1975, Duffy, 2007). A major confound of prosody assessments is the inability to dissociate receptive from expressive prosody (Wells and Peppé, 2003, Peppé and McCann, 2003). In particular, it is difficult to separate these two processes using assessment techniques that are based on spoken responses (either imitations or in conversation). While speech technology has been used in the general assessment of speech and language it often includes evaluation of prosody through acoustic analysis of lexical stress (Shahin et al., 2015), intensity, (Rodríguez et al., 2012), and tone (Rodríguez et al., 2012). The utility of these applications can not be understated. Together they provide an opportunity to significantly increase client-clinician interactions by standardizing objective assessment measurements (Shahin et al., 2015), automatically computing baseline clinical measures and for tracking intervention progress (Shahin et al., 2015), and for increasing client motivation / compliance through game-like exercises that target speech skills (e.g., voice-, intensity-, and pitch-controlled computer games; Rodríguez et al., 2012). However, these approaches still require vocal production, and are therefore unable to separate the effects of impaired motor control from deficits in prosody knowledge. Despite this limitation, the general framework of approaches used in current speech technologies holds great promise for developing new methods for providing information to users about their speech production, which may be helpful in assessing prosody production. In this paper, we extend these past technologies for speech assessment and training to a new application called the Prosodic Marionette, that is capable of examining prosody production knowledge using a non-vocal task.
Our new speech technology application is based on previous work on multimodal feedback of speech information with emphasis on graphical feedback of visual representations of speech features, including visualization of pitch, intensity, and duration (Ferguson et al., 2012, Hailpern et al., 2010, Patel and Furr, 2011, Patel and McNab, 2011, Pervaiz and Patel, 2014). Specifically, pitch and intensity have been used in multi-modal game-like interfaces designed to enhance clinical speech interventions (Patel and Salata, 2006, Ferguson et al., 2012, Rodríguez et al., 2012), and others have included duration cues and syllable detection to target deficits often observed in children with autism spectrum disorder and/or other speech disorders like childhood apraxia of speech (Hailpern et al., 2010, Shahin et al., 2015). Pitch, intensity, and duration are easily represented on a 2D graphical display. For instance, pitch and intensity can be intuitively represented along the vertical dimension of a computer screen, with changes in pitch and/or intensity used to move objects up and down (Hailpern et al., 2010, Ferguson et al., 2012, Rodríguez et al., 2012). Likewise, duration cues can be represented both in the vertical dimension, with increases in duration leading to upward movements of a game avatar (Rodríguez et al., 2012) or along the horizontal domain with time increasing left to right (Hailpern et al., 2010). Another innovative approach incorporated visual feedback of vocal intensity to a Google Glass display using similar graphical representations intended to improve speech clarity for individuals with multiple sclerosis and Parkinson’s disease (Pervaiz and Patel, 2014). Other studies have focused on improving expressive reading by including information about intensity, pitch, and duration that use both typical graphical representations of prosodic features (e.g., curves and meters representing fundamental frequency and sound pressure level) and augmented text (e.g., text that follows a curved path that represents the fundamental frequency) to provide visual feedback about prosody production (Patel and Furr, 2011, Patel and McNab, 2011). Similar to previous speech technologies, these expressive reading applications have used changes in the vertical screen position to represent changes in pitch, vertical height (e.g., font size) for intensity, and horizontal spacing and width for pause lengths and word durations (Patel and Furr, 2011, Patel et al., 2014).
The Prosodic Marionette draws from these established conventions for graphical display of prosody information, but reverses the process by permitting users to visually re-arrange linguistic items within a phrasal context on a graphical user interface in order to alter three major prosodic dimensions: fundamental frequency, word duration, and inter-word pause length for synthesized auditory output. Therefore, rather than displaying information from acoustic analyses of self-produced utterances, the Prosodic Marionette places movable widgets on a graphical interface that, when adjusted, can be synthesized to produce novel prosody. With the Prosodic Marionette visual-spatial interface, users are able to demonstrate prosodic knowledge through manual adjustments rather than spoken production, which may be more robust or reliable than an impaired or immature vocal motor system as may be found in individuals with neuromotor disorders (e.g., cerebral palsy) or for children who are in the typical developmental process. By eliminating speech motor control as a dependency for assessing prosody, the Prosodic Marionette may help to directly assess prosodic knowledge in a production task. The comparison of productions from both the Prosodic Marionette and spoken vocalizations have the potential to provide a more precise description of errors in prosody observed during development and for individuals with neuromotor disorders.
The following sections describe in detail the underlying structure and framework of the Prosodic Marionette interface, and present the results of a usability study to assess interface performance and functionality for marking contrastive stress patterns for pitch accent placements and boundary tones in yes/no question contexts (Patel and Campellone, 2009, Patel, 2004). The purpose of our usability study was to test two major hypotheses in a cohort of healthy adults without any neurological or cognitive impairments: (1) participant vocal imitations of synthesized target stimuli are faithful to the underlying prosody, and (2) participants can use the Prosodic Marionette interface to mark appropriate prosodic patterns (e.g., word-level focus and final boundary tones). Specifically, the first hypothesis will examine whether participant vocal imitations use fundamental frequency, word duration, and pause lengths matching those of the target stimuli that vary in their prosodic focus and phrase patterns. We expect that participants will be able to match all prosodic information in the target stimuli, which will serve as a baseline for evaluating the usability of the Prosodic Marionette interface. The second hypothesis will examine the same prosody measurements to determine whether participants are able to match the target prosodic contours using the Prosodic Marionette interface. Here we expect that participants will be able to manipulate the graphical elements of the interface to change the prosody output; however, our research question is whether some contrasts are more reliable than others (e.g., word focus versus phrase contours) and if participants preferentially rely on certain visual-spatial manipulations of the graphical interface that result in modification of specific prosody information (fundamental frequency, word duration, and pause length). Successful verification of both hypotheses will enable future research to test the utility of the Prosodic Marionette as an assessment tool that compares visual-spatial reconstruction of prosody with vocal productions that may be affected by speech disorder.
2. Method
2.1. Material and Stimuli
We designed the Prosodic Marionette software interface to test individuals’ knowledge of a range of prosodic contrasts that differ in both intonation (fundamental frequency) and timing (word duration and pause length). The Prosodic Marionette interface has the advantage of examining a variety of prosodic contrasts appropriate for a given language, independent of vocal motor control. It specifically examines one’s ability to reconstruct prosodic contrasts in a visual-spatial domain rather than via speaking in order to assess prosodic knowledge. The goal of Prosodic Marionette tasks is to match a target production through spatial manipulation of word-icon blocks on a graphical user interface bearing the intonational components of an utterance. Specifically, we represent pitch on the vertical axis and time on the horizontal axis with space between widgets indicating pauses and the width of widgets for duration. We did not initially include intensity in order to reduce interface complexity and given our past results that changes in pitch, pause length, and word duration were affected the most through the use of augmented displays of speech prosody (Patel et al., 2014), since the Prosodic Marionette resembles the interface of Patel et al. (2014) though used for prosody synthesis control rather than augmenting text for expressive reading. Further, the software interface was designed to be configured for a range of possible participants across ages (children through adults) and vocal motor function (with and without neuromotor disorders and dysarthria), and retains the same set of capabilities independent from the specific participant group. Figure 1 (bottom) illustrates the graphical interface and the dimensions available for manipulating prosody.
Figure 1:

Prosodic Marionette interface configured for adult participants. (Top) the imitation mode, IM, with no graphical cues for prosody features. (Bottom) the manual adjustment mode, PM, with example movements provided to highlight the three major prosody features able to be manipulated: fundamental frequency (both pitch accents and boundary tones), word duration, and inter-word pause length.
The Prosodic Marionette interface has two primary modes, one is used to record vocal imitations of target prosodic contours (Figure 1, top) and a second displays word-icon blocks connected with solid lines that are movable to indicate pitch and temporal changes (Figure 1, bottom). In this mode, moving the word-icon blocks up raises fundamental frequency and down lowers fundamental frequency. Moving a block left removes silence between the previous block to the left and increases the silence between the next block to the right, and vice versa. Participants can also drag a “hook” (half-circle graphical widget) to visually stretch the word-icon block to change the vowel duration of the associated word. Finally, participants can move a “tail” up or down to alter the fundamental frequency of the boundary tone; upward movements are translated into rising boundary tones and downward movements (below the level of the final word) are translated into falling boundary tones.
In a typical experimental session, participants first see the target utterance and hear it being produced by a speech resynthesizer (additional detail in Section 2.1.1). The target utterance is both written in traditional orthography and accompanied by icons to accommodate younger children and/or adults with developing and/or impaired literacy skills. In the imitation mode, or IM, participants listen to the target utterance, then press the record button to vocally reproduce the prosodic target. The graphical layout for IM places each word-icon block equally spaced on the horizontal plane, which does not convey any prosodic information and can not be manipulated (cf. Read N’ Karaoke; Patel and Furr, 2011). Participants press the button a second time to end recording. Vocally imitated productions can be replayed acoustically via the play button and compared to the target utterance. In the manual adjustment mode, or PM, participants again listen to the target utterance, then move and stretch the word-icon blocks on the graphical display to mimic the target prosodic contour. The default layout in this mode places all blocks equidistant on a horizontal plane to convey a neutral prosodic contour, i.e. no rises/falls in fundamental frequency or temporal variations. The PM experimental task then is to move each block to reflect the target prosody, e.g., raise the appropriate block location to indicate an increase in fundamental frequency. A prosodic contour representing fundamental frequency (stretched appropriately for inter-word pause lengths) is then interpolated between word blocks to provide visual feedback (blue line, Figure 1), and the resulting contour combined with any enlarged word-blocks (signifying word duration) is used to generate a resynthesized utterance for comparison of the target and reconstruction.
2.1.1. Auditory resynthesis
Both target utterances and Prosodic Marionette reconstructions are provided as feedback to participants through resynthesis of a neutral utterance with new prosodic information, depicted in Figure 2, using a custom script written for the Praat speech analysis software (Boersma and Weenink, 2011). The resynthesis method requires high-quality template audio samples and a segmental-level transcription for manipulation of fundamental frequency, pause lengths, and word durations, which results in natural-sounding audio stimuli and reconstructions. Ideally, the neutral prosody stimuli should be extracted from carrier sentences such as, “Bob said, ‘May needs one lime,”’ in order to minimize phrase onset effects. Reconstructions are generated by first modifying the appropriate neutral sound sample for either the target stimulus prosody or based on converted word-icon block screen coordinates (2D position) and widths, then resynthesized using a custom Praat script (Boersma and Weenink, 2011).1 The vertical screen coordinates ranged from a minimum of 50 Hz on the lower border to a maximum of 180 Hz on the upper border with initial graphical display at 102 Hz (40% from the bottom border). The horizontal screen coordinates represented a maximum utterance length of 2 seconds; graphical display of each word-box was dependent on the orthography only, and the initial spacing between words was distributed equally between all boxes. The initial duration represented by each word-box was equal to the natural production duration in the neutral sample. Therefore, the effects of all horizontal and box-stretching movements on pause length and word duration were in proportion to the amount of time remaining after subtracting the sum of all neutral word durations from the maximum utterance length (2 seconds). The limits on maximum utterance length and fundamental frequency range can be changed for longer stimuli and vocal ranges appropriate for children, and male and female adults. A demonstration is available online in Supplementary Video 1.
Figure 2:

Flowchart for Prosodic Marionette processing. Vertical and horizontal screen coordinates, and word-block widths are translated into fundamental frequency, pause length and word duration. The resulting prosody specification is used to resynthesize the target, neutral sample according to participant manipulations using a Praat script.
2.1.2. Graphical user interface
The Prosodic Marionette graphical user interface (GUI), shown in Figure 1, was implemented in Python using the Kivy cross-platform framework for natural user interface development (kivy.org) with PyAudio (Pham, 2006) for audio recording and playback. The Kivy framework provides native support for both touch screen and mouse-based interaction with GUI elements, which allows input access flexibility for use with multiple user populations with varying degrees of manual dexterity. For instance, touch interfaces can be more intuitive for children while adults with neuromotor disorders may require specialized computer mice, joysticks, or stylus interfaces to interact with computer systems.
2.1.3. Stimuli
The stimuli used for this usability study focused on minimal-pair-like differences in prosodic contrasts from age-appropriate stimuli for adults (similar to Patel and Campellone, 2009). For instance, Prosodic Marionette reconstructions of statement versus question sentences should primarily differ in the height of the final word and/or the height of the “tail” GUI element, which are used to mark the phrase boundary tone. Similarly, pitch accents placed on the second versus the fourth word of a four word sentence should differ primarily in the height (fundamental frequency) and width (word duration) of the second or fourth word-icon block, as appropriate. In addition, increases in word duration for the final word of all stimuli are expected due to the phrase-final lengthening effect (Klatt, 1976, Wightman et al., 1992). The stimuli used in this study are shown in Figure 3 and represent the following prosodic contrasts:
High pitch accent on the second [PA2] versus fourth [PA4] words, and
Rising [PBQ] versus falling [PBS] boundary tone to indicate yes/no questions versus statements.2
Figure 3:

Fundamental frequency for target sentence stimuli and prosodic contrasts. Statement and Question stimuli produced with broad focus. Words 1–4 represented by W1–W4, Boundary Tone by BT.
2.2. Participants
We recruited 15 adults from the University of Kansas and the surrounding community to participate in a study to validate the usability and performance of the Prosodic Marionette graphical user interface (age range: 19 – 30, mean 23.7 years; 5 male, 10 female). All participants were native speakers of American English, self-reported normal speech, language and hearing, and received monetary compensation for their time. This study investigates baseline usability from a normative population of healthy adults without any neurological or cognitive impairment. All study procedures were approved by the Institutional Review Board of the University of Kansas and all participants provided their informed consent prior to engaging in study activities.
2.3. Procedure
Experimental sessions were conducted in a sound-isolated booth. All participants sat in front of a touchscreen tablet computer (Dell XPS 18) and wore a pair of headphones (Sennheiser HD280) in addition to a head-mounted microphone (AKG C520) positioned lateral to the corner of the mouth. All participants completed tasks for both the IM and PM modes of the interface and the order of the two activities was counterbalanced across participants. Vocal imitations of the target stimuli obtained from the IM mode were used to establish a normative baseline of production based on the target stimuli for comparison with PM mode reconstructions.
The interface displayed target stimuli as word-icon blocks and was used to record participant responses in both the IM and PM tasks. The experimenter explained the task and engaged each participant in two demonstration trials for each task (IM and PM) prior to proceeding to the experimental runs. Each run consisted of 16 trials (two repetitions of each target stimulus listed in Figure 3) for each of the production modes, IM and PM. The participants were offered breaks between the two activities and were allowed breaks at any time they requested.
During the IM trials, participants were instructed to vocally imitate each target phrase exactly as they heard it produced by the computer. During the PM trials, participants were instructed to visuo-spatially manipulate the word-icon blocks to reproduce the target prosody. Participants had the option of moving the word-icon blocks with their finger or a mouse, though in this study all participants used the touch interface. Participants could listen to their IM imitations/PM reconstructions and could refine their responses up to two times prior to moving onto the next trial.
2.4. Data Measurement
In order to streamline the analysis process, we designed a semiautomated acoustic analysis software program that takes advantage of the known target stimuli in order to reduce manual effort and possible human error during transcription. For each audio file, our procedure first uses the Snack Sound Toolkit (KTH Sweden) to estimate fundamental frequency using the average magnitude difference function (AMDF) method3 then computes a forced-alignment via the Hidden Markov Model Toolkit (http://htk.eng.cam.ac.uk/) and Prosodylab-Aligner (Gorman et al., 2011) to obtain estimates of word onset and termination, and word transcriptions. Analysis of Prosodic Marionette data is ideally suited to semi-automatic transcription using forced-alignment; all of the utterances are completely known, are simple enough such that errors in production are unlikely, and have clear acoustics without noise. Next, word boundaries are adjusted using a custom user interface based on visual cues (amplitude waveform and spectrogram plots) and auditory cues (playback of sound samples between pairs of start–end markers). Last, the optimal fundamental frequency is chosen as either (1) the value at the inflection point within each word (minima or maxima), or (2) the mean word fundamental frequency if there are no inflections. A linear regression line is found for the last 2/3 of each word to determine whether fundamental frequency was rising or falling, and used to estimate the boundary tone fundamental frequency for the final word. Final manual adjustments could be made to correct for any idiosyncratic errors in the automated procedure. Once completed, all transcription and prosody values are stored in log files for later analysis.
2.5. Statistical analysis
Prior to statistical analysis, all fundamental frequency values were transformed according to the mel scale (m = 1127 log (1 + f0/700)), where m is the mel-scaled frequency and f0 is fundamental frequency in Hertz. The mel-scale is a psychoacoustic measure that equalizes perception of pitch across frequencies. Identical procedures were used for analysis of the IM and PM data. A general linear model of mel-scaled fundamental frequency (participant ID as a random factor) was first used to examine whether there were any differences due to participant sex (differences are expected for IM, but none are expected for PM). If present, a third factor of sex was added to the mixed effects general linear models described below. Next, we examined each sentence stimulus on its own for differences between conditions (either PA2 vs PA4, or PBQ vs PBS); the sentence stimulus in one condition is a natural control for its paired condition (e.g., May NEEDS one lime [PA2] vs May needs one LIME [PA4]) in order to account for any word- and sentence-specific differences in fundamental frequency, word duration, and pause length. Each sentence stimulus was examined using a mixed-effects general linear model to determine differences between within-subject factors of Word (Word 1, Word 2, Word 3, Word 4, Boundary Tone) and Condition (Pitch Accent 2 vs Pitch Accent 4 & Question vs Statement), participant ID was used as a random factor. An analysis of deviance was used for all hypothesis testing via Type II Wald χ2 test and significance assessed at p < 0.05. Statistically significant interactions between factors of Condition and Word (and Sex if applicable) were examined for simple effects (mixed-effects linear model with post-hoc comparisons per level of Condition across all Word levels and Mann-Whitney U Tests per level of Word across all levels of Condition). All statistical tests were assessed for significance after correction for multiple comparisons. These steps were repeated for word duration and pause length in both the IM and PM tasks.
3. Results
In this section we report on participant performance in the IM and PM tasks. Specifically, we evaluate the main hypothesis that participants are able to assign proper prosodic contours (fundamental frequency, pause length, and word duration) to match target stimuli using the Prosodic Marionette interface. Of particular interest are the simple interactions between Condition (pitch accent placement vs boundary tone placement) and Word (position in the sentence), which will help us to determine whether certain prosodic contrasts are more reliably imitated (IM) or reproduced (PM) than others. In addition, we repeat all statistical analyses for the three main prosody measurements (fundamental frequency, word duration, and pause length) to determine whether there are preferences or biases toward one adjustments of one measure over another according to task (IM or PM) and contrast. For instance, a statistically significant difference in word 2 fundamental frequency between stimuli with a pitch accent placement on the second and fourth words would indicate participants marked the correct prosody. Similarly, a statistically significant difference in boundary tone fundamental frequency between the question and statement stimuli also indicates participants were able to mark the correct prosody. We therefore include an analysis of all main effects (Condition and Word), their interactions (Condition X Word) and simple effects for any statistically significant findings to fully examine participants’ ability to mark prosody using the Prosodic Marionette interface. In addition, we examine the usage of each interface control for modifying prosodic features as a measure of interface effectiveness. The IM task results are presented first to verify participants were able to vocally produce the target prosody, and to provide a baseline from which to interpret prosody results in the PM task.
3.1. IM task
The IM task serves as a baseline for interpreting the PM task results by asking participants to directly imitate the synthesized stimuli. The acoustic analyses of IM task productions also confirms the expected pattern of prosodic contours used in this study (see Figure 3). In this analysis, we first identified whether there were any differences in fundamental frequency present due to sex to help account for individual variance in subsequent statistical analysis of the factors Condition and Word. This was done in order to determine how participants differentially marked each of the paired prosodic contrasts. We expected and found statistically significant differences in fundamental frequency between male and female participants (linear mixed-effects analysis of deviance, χ2(1) = 221.34, p < 0.001); therefore, a third factor of Sex was included in all subsequent analyses of fundamental frequency and simple effects were evaluated for each sex separately. The median observed fundamental frequencies and interquartile ranges for the IM task are shown in Figure 4 for all sentences and conditions. Analyses of word duration and pause length were completed using all participants independent of sex.
Figure 4:

(left) Medians and interquartile ranges for fundamental frequency from IM task productions. Responses from female participants are in black, and male participants in gray. The dashed and solid lines represent the paired conditions PA2-PA4 or PBQ-PBS (see legends). Words 1–4 represented by W1–4 and boundary tone by BT. (right) Results of statistical tests of fundamental frequency using the same tests and levels of significance used in Figure 5.
There were statistically significant differences in fundamental frequency for the three-way interaction effect of Sex X Condition X Word for all sentences (“Lane mows the lawn,” χ2(4) = 12.58, p = 0.014; “Mel has a dime,” χ2(4) = 43.73, p < 0.001; “May needs one lime,” χ2(4) = 22.76, p < 0.001; “Noam won a broom,” χ2(4) = 40.79, p < 0.001) as well as for the two-way interaction of Condition X Word for each level of Sex (Figure 4, right, under the heading “Condition X Word”). There were also statistically significant differences in word duration for the interaction effect of Condition X Word for sentences with pitch accent contrasts only (Lane mows the lawn: χ2(3) = 35.65, p < 0.001, May needs one lime: χ2(3) = 62.83, p < 0.001). Although there were no interaction effects for pause length, statistically significant main effects of Word were found for all sentences after Bonferroni correction for the number of sentences (“Lane mows the lawn,” χ2(2) = 73.83, p < 0.001; “Mel has a dime,” χ2(2) = 40.92, p < 0.001; “May needs one lime,” χ2(2) = 27.7 = 06, p < 0.001; “Noam won a broom,” χ2(2) = 56.24, p < 0.001). In this case, a statistically significant main effect without interactions indicates a similar approach to pause length was used within the broad contrast categories of pitch accent placement and phrase boundary. Pauses, in our stimuli, should be minimized for connected speech, unless there are specific changes as a result of prosodic marking.
The simple effects of Condition and Word were further examined to determine specific differences in fundamental frequency and word duration for sentences with statistically significant interaction effects, and a post-hoc multiple comparisons test was used to examine the statistically significant main effect of pause length. For the effect of Condition we used a one-way analysis of deviance of fundamental frequency using a mixed effects general linear model of the within subjects factor Word (5 levels) for each Condition per Sentence. For the effect of Word we examined differences in both fundamental frequency and word duration using a Mann-Whitney U Test of Condition (2 levels) for each Word and Boundary Tone per Sentence.
3.1.1. Simple effects: Condition
All results for the analysis of the simple effects of Condition for fundamental frequency are shown in Figure 4, right (under the heading “Condition”).
In the pitch accent placement contrast, participants raised their fundamental frequency on the second word in the PA2 condition compared to the PA4 condition for both sentence stimuli (female: 88.9 mel, 108.6 mel; male: 52.1 mel, 77.6 mel; shown in Figure 4, right). Similarly, participants raised their fundamental frequency on the fourth word in the PA4 condition compared to the PA2 condition for both pitch accent contrast stimuli (female: 81.8 mel, 96.3 mel; male: 37.4 mel, 64.3 mel; shown in Figure 4, right). In addition, participants increased duration of the second word in the PA2 condition compared to the PA4 condition (42.5 ms and 87.4 ms, shown in Table 1) and on the fourth word in the PA4 condition compared to the PA2 condition (74.9 ms and 125.8 ms for each stimulus in Table 1). This effect was eliminated though for the sentence “Lane mows the lawn” in the PA2 condition after correcting for multiple comparisons.
Table 1:
Duration IM results, by sentence and word, Mann-Whitney U Test. Phrase boundary stimuli had no significant differences in word durations. ∆ is the difference in median word duration for significant comparisons. Only word 2 and word 4 were significant after Bonferroni correction for words (N=4) and sentence (N=4), p<0.05/16, except for word 2 (“Lane mows the lawn”), which is no longer statistically significant after Bonferroni correction.
| Lane mows the lawn | May needs one lime | |
|---|---|---|
| Word 2 | PA2 > PA4 (p=.131/16) | PA2 > PA4 |
| U=687, ∆=42.5 ms | U=1234, ∆=87.4 ms | |
| Word 4 | PA4 > PA2 | PA4 > PA2 |
| U=212.5, ∆=74.9 ms | U=226.5, ∆=125.8 ms |
In the phrase boundary contrasts, participants raised the boundary tone in the PBQ condition and lowered it in the PBS condition for both sentences (female: 204.4 mel and 209.2 mel for “Mel has a dime” and “Noam won a broom,” respectively; male: 128.8 mel and 144.8 mel; full statistics shown in Figure 4, right). There were no changes in word duration or pause length between phrase boundary contrasts.
3.1.2. Simple effects: Word
All results for the simple effects of Word analysis for fundamental frequency can be found in Figure 4, right, with the heading “Word.”
In the pitch accent placement contrasts, pairwise comparisons between all levels of Word revealed similar, and expected, patterns of fundamental frequency usage among all participants with some variations by specific stimuli. The two major patterns were:
The highest fundamental frequency was found for word 2 in both stimuli in the PA2 condition; the lowest on the boundary tone.
The highest fundamental frequency was found for word 4 in the PA4 condition; the lowest on the boundary tone.
In the PA2 condition, there was a general decrease in fundamental frequency through the boundary tone (with an exception on word 2), while in the PA4 condition fundamental frequency gradually increased until reaching a maximum at word 4, after which a sharp decline was observed on the boundary tone. In addition, pairwise comparisons for the main effect of Word for pause length indicated statistically significantly longer intervals between words 2 and 3 compared to all others for both stimuli with pitch accent contrasts independent of condition (Tukey’s HSD test, p < 0.001).
In the phrase boundary contrasts, statistically significant effects of Word were found for the simple effects analyses per level of Condition for both sentences. Follow-up examination of the pairwise comparisons between levels of Word indicated that:
Participants increased the boundary tone fundamental frequency to levels greater than all other words in the PBQ condition, and
The boundary tone had the lowest fundamental frequency for the PBS condition stimuli.
The major differences from these patterns occurred for the PBS stimuli with participants using greater fundamental frequency on word 3 for the stimulus “Mel has a dime” and on word 4 for the stimulus “Noam won a broom,” though both patterns match those given in the target stimuli (see Figure 3). We also found that pairwise comparisons of the statistically significant main effect of Word for pause length indicated longer pause durations between words 3 and 4 compared to all other pause intervals for both stimuli (“Mel has a dime:” 21 ms relative to the pause preceding word 2, Z = 2.788, p < 0.001; 13.8 ms relative to the pause preceding word 3, Z = 2.467, p < 0.05; “Noam won a broom:” 21 ms relative to the pause preceding word 2, Z = 4.928, p < 0.001; 21 ms relative to the pause preceding word 3, Z = 4.935, p < 0.001).
3.2. PM task
The median fundamental frequency and interquartile ranges per word in the PM task are shown in Figure 5 for all sentences and conditions. We found statistically significant differences in fundamental frequency for the interaction Condition X Word for all four sentence stimuli, corrected for multiple comparisons (Figure 5, right). However, there was only one statistically significant interaction for word duration (May needs one lime: χ2(3) = 19.14, p < 0.001), and no interaction effects for pause length reached statistical significance. Despite no significant interaction effects, a statistically significant main effect of Word was found for pause length for the sentence “Lane mows the lawn” (χ2(2) = 27.23, p < 0.001). Unlike word duration, differences in pause length are a function of prosodic contrast only (e.g., all things equal, pauses should be eliminated for connected speech versus variation in word duration as a function of word length and prosody). Main effects of fundamental frequency were not examined due to a statistically significant interaction effect, and main effects of word duration were not examined since differences in word duration are only informative as a function of prosody.
Figure 5:

(left) Medians and interquartile ranges for fundamental frequency from PM reconstructions. The dashed and solid lines represent the paired conditions PA2-PA4 or PBQ-PBS (see legends). Words 1–4 represented by W1–4 and the boundary tone by BT. (right) Results of statistical tests of fundamental frequency. Only statistically significant results are shown. Interaction effects are assessed with Bonferroni correction for each sentence (α = 0.05/4); simple effects of Condition assessed with Bonferroni correction α = 0.05/20 for 5 words and 4 sentences; simple effects of Word assessed with Bonferroni correction α = 0.05/8 for 4 sentences and 2 conditions. Equality in the simple effects of Word results indicates no significant difference (e.g., Word 4 ≥ Word 3 ≥ Word 2 indicates Word 4 is not different than Word 3, but statistically significantly greater than Word 2). Commas indicate no statistically significant differences.
The simple effects of Condition and Word were further examined to determine specific differences in fundamental frequency and word duration for sentences with statistically significant interaction effects, and a post-hoc multiple comparisons test was used to examine the statistically significant main effect of pause length for the sentence “Lane mows the lawn.” For the effect of Condition we used a one-way analysis of deviance of fundamental frequency using a mixed effects general linear model of the within subjects factor Word (5 levels) for each Condition per Sentence (an identical procedure was used to examine word duration for the sentence “May needs one lime”). For the effect of Word we examined differences in both fundamental frequency and word duration using a Mann-Whitney U Test of Condition (2 levels) for each Word and Boundary Tone per Sentence.
3.2.1. Simple effects: Condition
All results for the simple effects of Condition for fundamental frequency are shown for each sentence in Figure 5 (right) after the heading “Condition.” The pitch accent placement stimuli are in the first two rows, and phrase boundary stimuli in the bottom two rows.
In the pitch accent placement contrast, both sentences had higher fundamental frequency on word 2 in the PA2 condition than the PA4 condition (48.2 and 46.3 mel, respectively), and only raised fundamental frequency on word 4 in the PA4 condition versus the PA2 condition for the stimulus “May needs one lime” (25.6 mel). Participants also raised fundamental frequency on word 3 for this sentence (23.3 mel). In addition to fundamental frequency, participants increased the length of word 2 in the PA2 condition for the sentence “May needs one lime” compared to the PA4 condition (i.e., NEEDS) by 55.7 ms, though this simple effect was not statistically significant after correcting for multiple comparisons (U = 796, p = 0.006; α = 0.05/20). Pairwise comparisons of the levels of Word for the statistically significant main effect of Word for pause length (“Lane mows the lawn”) found participants increased the interval between words 2 and 3 by 37 ms on average compared to all other pause lengths independent of pitch accent conditions PA2 or PA4 (39.3 ms relative to the pause preceding word 2, Tukey’s HSD Z = 4.477, p < 0.001; 35.6 ms relative to the pause preceding word 4, Tukey’s HSD Z = 4.046, p < 0.001).
In the phrase boundary contrast, we found participants generally raised the boundary tone fundamental frequency in the PBQ condition, and lowered in the PBS condition (difference between conditions: 98.5 and 100.8 mel, respectively for each sentence). In addition, participants raised fundamental frequency on word 3 in the PBS condition compared to the PBQ condition for the stimulus “Mel has a dime” (16.0 mel), but no others, which is similar to the target prosody in Figure 3 and the IM responses in Figure 4. There were no differences in word durations or pause length for either stimulus in the phrase boundary contrast.
3.2.2. Simple effects: Word
All results for the simple effects of Word for fundamental frequency are shown for each sentence in Figure 5 (right) after the heading “Word.” There was a statistically significant effect of word fundamental frequency for all sentences and conditions.
In the pitch accent placement contrasts, post-hoc testing (Tukey’s HSD) revealed participants used expected fundamental frequency patterns for all stimuli with minor variations. In both sentences in the PA2 condition, fundamental frequency was greatest for word 2, followed by equal fundamental frequency in words 1, 3 and 4, and lowest fundamental frequency in the boundary tone. In the PA4 condition there were no statistically significant differences in fundamental frequency for the sentence “Lane mows the lawn” between words 1–4, and a statistically significantly lower fundamental frequency on the boundary tone. For the stimulus “May needs one lime,” statistically significant differences were found with greatest fundamental frequency on words 3 and 4 compared to words 1 and 2 and the boundary tone.
In the phrase boundary contrasts, statistically significant post-hoc comparisons (Tukey’s HSD) were found for both sentences in the PBQ condition, with highest fundamental frequency on the boundary tone compared to all other words. In the PBS condition, fundamental frequency was generally lowest for the boundary tone than all other words. In the sentence “Noam won a broom,” the boundary tone fundamental frequency was lower than words 2–4, and equal to the level of word 1.
3.2.3. Follow-up analysis of temporal features
Our statistical analysis revealed that changes to temporal features were not as robust as those made to fundamental frequency, despite an expectation that both words with pitch accents and phrase boundaries should lead to increases in duration. Therefore, we completed a follow-up analysis to determine how many participants altered word duration in any amount from their initial, baseline values and how many participants placed a space of any size between words. This analysis provides a sense for whether participants used the temporal features, or whether they prioritized matching to fundamental frequency. A summary of these results can be seen in Table 2. Overall, we found a maximum of 51% of participants increased duration (initial values were minimum durations, so any modification was an increase) in a single Sentence x Condition pair (“May needs one lime,” PA2), and that when used, word duration was increased on the contrast-appropriate word (e.g., word 2 in the PA2 condition). For example, participants increased word duration most often on the final word for all conditions (except for the PA2 condition), which agrees with prior observations of phrase final lengthening (Klatt, 1976, Wightman et al., 1992) and requisite increases for applying pitch accents to the final word (e.g., PA4). For the PA2 conditions, participants most often increased word duration on the second word. Even fewer participants made any changes to pause intervals, with a maximum of just 18% in any one Sentence X Condition pair (“Mel has a dime,” PBQ).
Table 2:
Proportion trials with increased duration over baseline values (left four columns) and increased pause duration above zero (last 3 columns) in the adult-configured interface. Values in bold are the maximum proportions over the four words (duration) and three pause intervals (pause) per stimulus.
| Word Duration | Preceding Pause | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Sentence | Condition | Word 1 | Word 2 | Word 3 | Word 4 | Word 2 | Word 3 | Word 4 | |
| Mel has a dime. | PBS | 0.054 | 0.054 | 0.081 | 0.297 | 0.108 | 0.135 | 0.162 | |
| Mel has a dime? | PBQ | 0.059 | 0.000 | 0.029 | 0.294 | 0.118 | 0.147 | 0.176 | |
| Noam won a broom. | PBS | 0.171 | 0.143 | 0.114 | 0.429 | 0.086 | 0.086 | 0.086 | |
| Noam won a broom? | PBQ | 0.189 | 0.027 | 0.135 | 0.486 | 0.108 | 0.135 | 0.162 | |
| Lane MOWS the lawn. | PA2 | 0.147 | 0.353 | 0.176 | 0.147 | 0.059 | 0.088 | 0.059 | |
| Lane mows the LAWN. | PA4 | 0.147 | 0.176 | 0.088 | 0.294 | 0.118 | 0.059 | 0.088 | |
| May NEEDS one lime. | PA2 | 0.077 | 0.513 | 0.026 | 0.231 | 0.108 | 0.077 | 0.103 | |
| May needs one LIME. | PA4 | 0.121 | 0.212 | 0.242 | 0.424 | 0.085 | 0.091 | 0.121 | |
4. Discussion
We conducted an experiment to examine the usability and validity of the Prosodic Marionette, a tool for studying prosody using a visual-spatial computer interface for resynthesizing prosody. The major benchmarks for success include the following:
Hypothesis 1: Participants can accurately vocally imitate prosodic contrasts based on auditory perception of artificially resynthesized speech segments (IM mode).
Hypothesis 2: Participants can accurately reproduce prosodic contrasts using the Prosodic Marionette interface (PM mode).
In the first benchmark, we confirmed vocal imitation of the prosodic patterns in our target stimuli (IM). The same analyses then specifically tested the hypotheses that participants used the Prosodic Marionette interface to mark appropriate, prosodic patterns for each pair of stimuli (PM).
4.1. Accurate imitation
To successfully imitate the target prosody, participants must be able to perceive all prosodic information from synthetic stimuli and reproduce them vocally. We examined performance according to fundamental frequency, word duration, and pause length, for each sentence stimulus. Overall, participants fully imitated the appropriate changes to fundamental frequency for all sentences in all prosodic conditions (pitch accent placements and phrase boundaries). The major features were all found statistically significant:
Raised word 2 fundamental frequency for the PA2 condition,
Raised word 4 fundamental frequency for the PA4 condition,
Raised the boundary tone fundamental frequency in the PBQ condition, and
Lowered boundary tone fundamental frequency for the PBS condition.
Participants were even faithful to small differences between the two PBS stimuli; the stimulus “Mel has a dime” included an increase on word 3 as well as on word 4 while “Noam won a broom” had an increase on word 4 only (see Figure 3). This was reflected in Figure 4 with greatest (and equal) fundamental frequency for word 3 and word 4 in “Mel has a dime” compared to all other words.
Participants were also able to accurately mark the target prosody with appropriate changes to word duration and pause length. Specifically, participants increased the duration of word 2 and word 4 in the pitch accent contrasts relative to each other. Although there were no differences between phrase boundary conditions, participants did show the highest duration on word 4. Due to the design of our experiment, it is not possible in the IM task to determine whether these increases in duration are due to the word itself, or the prosody; however, past studies on phrase-final word lengthening associated with both statement and question stimuli (Klatt, 1976, Wightman et al., 1992) corroborate these findings. There were also no condition-specific differences in pause length, rather we found participants consistently increased the pause between words 2 and 3 for all pitch accent placement stimuli, and between words 3 and 4 for phrase boundary stimuli. These results taken together confirm the benchmark for accurate vocal imitation of target resynthesized prosodic stimuli.
4.2. Accurate reproduction via PM
We examined fundamental frequency, word duration, and pause length measurements of prosody to assess the accuracy of participants’ recreation of prosodic contrasts using the Prosodic Marionette. The main results for participants’ use of fundamental frequency in Figure 5 qualitatively matches the targets specified in Figure 3, with changes in fundamental frequency placed where appropriate. Specifically, participants used the following major prosodic markers, modified through vertical movements of word-icon blocks in the Prosodic Marionette interface:
High fundamental frequency on word 2 to mark a high pitch accent placed on word 2,
High fundamental frequency on word 4 to mark a high pitch accent placed on word 4,
Low boundary tone fundamental frequency to mark a statement,
High boundary tone fundamental frequency to mark a yes-no question, and
Low boundary tone fundamental frequency for all non-question stimuli to generally mark statement prosody.
These user responses, focusing on fundamental frequency, indicate general success on the benchmark for basic use of the Prosodic Marionette interface by raising a word-icon block to mark high pitch accent placements, and using the “tail” GUI feature to correctly mark rising and falling boundary tones. For all other words there was some inter-participant variation, though generally all responses aligned well with the target utterances.
In the temporal domain, a majority of participants did not use the Prosodic Marionette’s ability to alter temporal features of word duration and pause length, which limits the ability to interpret the temporal feature results. However, when they did, the majority of participants altered only the duration of prosody-bearing words (see left columns, Table 2). For example, participants increased the word length above initial levels for the second and fourth words in the pitch accent placement contrasts as appropriate. A further examination of the ratios of participants who made any change in duration found between 29 – 48% who increased the duration of the final word for stimuli representing statement, question, and pitch accent on the final word. Approximately 35% and 51% of participants increased the duration of the second word for the stimuli “Lane MOWS the lawn” and “May NEEDS one lime,” respectively, to mark a pitch accent on the second word. These results are supported by the statistical analysis of word duration for pitch accent contrasts (e.g., an increase in duration on word 2 of “May needs one lime” by 56 ms in the PA2 condition). In addition, participants also removed all pauses between words (see Table 2, right three columns). No more than 17% of participants inserted a preceding pause on any given word, which was fairly consistently placed before the final word for boundary tones, but inconsistently applied to pitch accents.
These results demonstrate general usability and proficiency for marking prosodic contrasts using the Prosodic Marionette interface, especially for the fundamental frequency feature. Participants were able to successfully use pitch cues through vertical movements of the word-icon blocks and duration cues, when used, by stretching the word-icon blocks. None of the prosodic contrasts investigated in this study required a pause interval preceding any word, and in fact, the appropriate action should be to close the gap between words for a fluent sentence. A majority of participants fully closed the gaps between words (between 82 – 94% on any given word). Future work should focus on participants’ abilities to use the temporal features of the Prosodic Marionette to mark word duration and pause length cues.
5. Future directions and limitations
5.1. Assessing prosody using the Prosodic Marionette
A future goal of the Prosodic Marionette interface is as a new research and clinical tool for assessing prosody, for instance in typically and atypically developing children, and children and adults with neuromotor disorders leading to dysarthria. Assessing prosodic knowledge using the Prosodic Marionette interface consists of evaluating prosody performance within each of the production modes (IM and PM), then comparing performance across production modes. Typically developing adults without any neurological or neuromotor impairments are expected to accurately produce target prosodic contours both vocally and through the PM interface. That is, individuals are able to perceive and reproduce prosodic contours equally, and accurately using either their vocal-motor systems or the PM visual-spatial interface. Equal performance implies the vocal motor system (IM) and perceptual system (PM) are intact without impairment.
Young children may have immature motor control of the speech production mechanism that may prevent accurate production of prosodic contours, and obscure vocal demonstration of prosodic knowledge (Patel and Grigos, 2006). A non-vocal assessment, like the Prosodic Marionette, may shed light on the developmental process of acquiring prosody. If children have adult-like knowledge of prosody, and have sensitive enough perceptual systems to perceive prosody, then they should be able to reproduce target prosodic contours visuo-spatially via limb-motor control. Children may alternatively perform similarly to adults for some prosodic contrasts leading to IM accuracy that may either equal or exceed PM reconstructions. In this case, differences between IM and PM may identify a minimum level of executive function needed to complete the PM tasks. Nonetheless, the specific differences between IM and PM productions over a range of prosodic contrasts can provide insight into the trajectory of prosodic development in children (Cruttenden, 1985). The assessment of prosody using the Prosodic Marionette interface, however, must be adapted for access considerations and age-appropriateness (e.g., sentence stimuli).
In adults with acquired dysarthria, past research identifies deficits in prosody (namely monopitch, monoloudness, and equal stress) as hallmarks of the disorder (Darley et al., 1975). Using the Prosodic Marionette interface, we expect vocal imitations of prosodic contrasts by individuals with acquired dysarthria to be monopitch and monoloud. However, if they still have intact prosodic knowledge, they should be able to produce target prosody in the PM production mode. Individuals with limb neuromotor impairments without any deficits to the vocal motor system are hypothesized to be able to accurately reproduce prosodic contours using their vocal-motor systems, but perform more poorly in the PM task as a result of their impairment (subject to access modification effectiveness). Accurate IM performance with poor-performing or inaccurate PM reproductions implies the vocal motor system is intact, but that either the limb-motor system is impaired, immature, or otherwise unreliable – or the PM task is too complex and difficult to be completed. Care must be taken to ensure participants have sufficient cognitive capacity to fully understand task instructions. For instance, individuals who are unable to follow complex, multistep instructions may perform poorly in the PM task, though demonstrating prosodic knowledge via vocal imitation. Individuals with speech-motor deficits or immaturities, are hypothesized to be able to accurately reproduce prosodic contours using the PM interface only, but unable match target contours using vocal-motor imitation as a result of their speech-motor status. This asymmetry in performance implies perception and prosodic knowledge are intact, and only the vocal-motor system is impaired, immature, or otherwise unreliable.
Focusing on the goal of research and clinical applications of the Prosodic Marionette interface reveals a limitation of the current software. In the present study, participants generally performed equally between Prosodic Marionette reconstructions and vocal imitations as expected for neurotypical, post-lingual adults, but they slightly differed in their responses between production modalities on word duration and pause length (matched expectations for IM but not PM). These results indicate that participants valued changes in fundamental frequency the most when using the Prosodic Marionette, relied on word durations to a much lower extent, and for the most part ignored pause length. These differences serve as a normative baseline using the present stimuli for expected maximum performance in future studies using the Prosodic Marionette with individuals with speech motor impairments. In addition, these differences demonstrate that an imbalance between melodic (i.e., fundamental frequency) and temporal features may be observed in the PM task, which implies that the choice of stimuli is critically important for answering specific research questions about users’ skill with melodic versus temporal features of prosody. Furthermore, future experiments may uncover that some participants, with or without neuromotor disorders, rely on certain features more than others as a function of stimulus. An advantage of the Prosodic Marionette is that such asymmetries are quantifiable using the PM task interface though future study is needed to define and validate a clinically useful objective measure for making comparisons between the IM and PM tasks. The overreliance on fundamental frequency in the current study is likely specific to the stimuli used in this experiment; none included an explicit change in temporal parameters (e.g., none had exaggerated word durations, or extended pauses as might be found between phrases within a sentence). Rather, changes in duration and pause length in this experiment were expected as implicit markers of relatively simple prosodic contrasts, while the major prosodic feature was fundamental frequency. Future study is needed to fully explore participants’ responses to stimuli with explicit changes to temporal features.
5.2. Exploring complex graphical interface control of prosody
One of the advantages of the Prosodic Marionette interface is providing participants an ability to modify prosody through multiple dimensions (melodic and temporal). As a tool, access to these dimensions is advantageous for specifying the desired prosody as specifically as possible. In practice, the unconstrained availability of these options may prove challenging for some participants, including young children and adults with sensory, motor, or cognitive impairments. To counter these challenges, we recommend restricting access to certain dimensions of prosody manipulation according to experimental hypotheses in a controlled manner (e.g., separately test participants’ ability to move word-icon blocks up and down, then left to right for testing fundamental frequency and temporal dimensions). In addition, future work is needed on constructing target stimuli that maximize Prosodic Marionette responses along the specifically tested dimensions (e.g., stimuli that evoke measurable changes in pause lengths versus those for altering word duration). Another limitation of the present version of the interface is an inability to manipulate word intensity along with the other acoustic and temporal features. A beta-version of the interface is currently being developed for integration of intensity control and future study will be needed to examine the usability of graphical interface elements for controlling intensity and its interaction with the other interface manipulations of fundamental frequency, pause length, and word duration. Finally, the Prosodic Marionette interface was initially designed for experimental investigations of prosodic knowledge in a variety of participant populations. Future uses of the interface may be focused on functional applications such as second language learning or clinical intervention for disorders in prosody production.
Supplementary Material
6. Acknowledgments
We would like to thank Seth Polsley for his assistance implementing the PM software, Jonathan Barnes for his help with the recordings, and Nanette Veilleux and Stefanie Shattuck-Hufnagel for their discussions to improve the interface design. This research was supported by the National Institutes of Health (NIDCD: R21DC013095; NICHD: R03HD064787, U54HD090216).
Footnotes
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final citable form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
Praat data structures are manipulated using the TextGridTools Python package (Buschmeier and Vlodardzak, 2013).
“WH” questions are produced using a falling boundary tone and can also be examined with appropriate prosodic contours.
The following parameters were used for fundamental frequency estimation: window length 7.5 ms, frame length 10 ms, maximum F0 250 Hz, minimum F0 51 Hz.
References
- Boersma P and Weenink D (2011). PRAAT: doing phonetics by computer (Version 5.2.25) [Computer Software; ]. [Google Scholar]
- Bolinger D (1989). Intonation and its uses: melody in grammar and discourse Stanford University Press, Stanford. [Google Scholar]
- Buschmeier H and Vlodardzak M (2013). TextGridTools: A TextGrid processing and analysis toolkit for Python. In Proceedings der 27. Konferenz zur Elektronischen Sprachsignalverarbeitung, pages 152–157, Bielefeld, Germany. [Google Scholar]
- Cruttenden A (1985). Intonation comprehension in ten-year-olds. Journal of Child Language, 12(3):643–661. [Google Scholar]
- Darley FL, Aronson AE, and Brown JR (1975). Motor speech disorders W. B. Saunders, Philadelphia, PA. [Google Scholar]
- Duffy JR (2007). Motor speech disorders: history, current practice, future trends and goals. In Weismer G, editor, Motor Speech Disorders, pages 7–57. Plural Publishing, San Diego, CA. [Google Scholar]
- Ferguson S, Johnston A, Ballard K, Tan CT, and Perera-Schulz D (2012). Visual feedback of acoustic data for speech therapy. In Proceedings of the 7th Audio Mostly Conference on A Conference on Interaction with Sound - AM ‘12, pages 135–140, Corfu, Greece: ACM. [Google Scholar]
- Gorman KR, Howell J, and Wagner M (2011). Prosodylab-Aligner: A tool for forced alignment of laboratory speech. Canadian Acoustics, 39(3):192–193. [Google Scholar]
- Hailpern J, Karahalios K, DeThorne L, and Halle J (2010). Vocsyl: Visualizing Syllable Production for Children with ASD and Speech Delays. In Proceedings of the 12th international ACM SIGACCESS conference on Computers and accessibility - ASSETS ‘10, page 297, Orlando, FL, USA: ACM. [Google Scholar]
- Klatt DH (1976). Linguistic uses of segmental duration in English: acoustic and perceptual evidence. Journal of the Acoustical Society of America, 59(5):1208–1221. [DOI] [PubMed] [Google Scholar]
- Lehiste I (1970). Suprasegmentals MIT Press, Cambridge. [Google Scholar]
- Lehiste I (1976). Suprasegmental features of speech. In Lass NJ, editor, Contemporary Issues in Experimental Phonetics, pages 225–239. Academic Press, New York, NY. [Google Scholar]
- Patel R (2004). The acoustics of contrastive prosody in adults with cerebral palsy. Journal of Medical Speech-Language Pathology, 12:189–193. [Google Scholar]
- Patel R and Campellone P (2009). Acoustic and perceptual cues to contrastive stress in dysarthria. Journal of Speech, Language, and Hearing Research, 52(1):206–222. [DOI] [PubMed] [Google Scholar]
- Patel R and Furr W (2011). ReadN’Karaoke: Visualizing Prosody in Children’s Books for Expressive Oral Reading. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, pages 3203–3206, Vancouver, BC: ACM. [Google Scholar]
- Patel R and Grigos MI (2006). Acoustic characterization of the question-statement contrast in 4, 7 and 11 year-old children. Speech Communication, 48(10):1308–1318. [Google Scholar]
- Patel R, Kember H, and Natale S (2014). Feasibility of augmenting text with visual prosodic cues to enhance oral reading. Speech Communication, 65:109–118. [Google Scholar]
- Patel R and McNab C (2011). Displaying prosodic text to enhance expressive oral reading. Speech Communication, 53(3):431–441. [Google Scholar]
- Patel R and Salata A (2006). Using computer games to mediate caregiver-child communication for children with severe dysarthria. Journal of Medical Speech Language Pathology, 14(4):279–284. [Google Scholar]
- Peppé S and McCann J (2003). Assessing intonation and prosody in children with atypical language development: the PEPS-C test and the revised version. Clinical Linguistics & Phonetics, 17(4–5):345– 354. [DOI] [PubMed] [Google Scholar]
- Pervaiz M and Patel R (2014). SpeechOmeter: Heads-up Monitoring to Improve Speech Clarity. In Proceedings of the 16th International ACM SIGACCESS Conference on Computers & Accessibility, pages 319–320, Rochester, NY, USA: ACM. [Google Scholar]
- Pham H (2006). Pyaudio [computer software; ]. [Google Scholar]
- Rodriguez WR, Saz O, and Lleida E (2012). A prelingual tool for the education of altered voices. Speech Communication, 54(5):583–600. [Google Scholar]
- Shahin M, Ahmed B, Parnandi A, Karappa V, McKechnie J, Ballard KJ, and Gutierrez-Osuna R (2015). Tabby Talks: An automated tool for the assessment of childhood apraxia of speech. Speech Communication, 70:49–64. [Google Scholar]
- Shattuck-Hufnagel S and Turk AE (1996). A prosody tutorial for investigators of auditory sentence processing. Journal of Psycholinguistic Research, 25(2):193–247. [DOI] [PubMed] [Google Scholar]
- Wells B and Peppe S (2003). Intonation abilities of children with speech and language impairment. Journal of Speech, Language, and Hearing Research, 46(1):5–20. [DOI] [PubMed] [Google Scholar]
- Wightman CW, Shattuck-Hufnagel S, Ostendorf M, and Price PJ (1992). Segmental durations in the vicinity of prosodic phrase boundaries. Journal of the Acoustical Society of America, 91(3):1707–1717. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
