Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Jun 5.
Published before final editing as: J Voice. 2024 Dec 5:S0892-1997(24)00373-4. doi: 10.1016/j.jvoice.2024.10.024

Prosodic Preferences of Surface Electromyography-based Subvocal Speech for People with Laryngectomy

Laura Raiff a,b, Dea Turashvili a,b,c, James T Heaton d, Gianluca De Luca a,b, Joshua C Kline a,b, Jennifer M Vojtech a,b,e
PMCID: PMC12137685  NIHMSID: NIHMS2032185  PMID: 39643558

Abstract

Introduction:

People who undergo a total laryngectomy lose their natural voice and depend on alaryngeal technologies for communication. However, these technologies are often difficult to use and lack prosody. Surface electromyographic-based silent speech interfaces are novel communication systems that overcome many of the shortcomings of traditional alaryngeal speech and have the potential to seamlessly incorporate individualized prosody. The purpose of this study was to (1) validate the ability of alaryngeal silent speech to effectively incorporate pitch modulations—a key prosodic element in natural speech—into synthesized speech assessed through listening experiments, and (2) determine the key features of these communication devices according to core users.

Methodology:

People with laryngectomy (PWL, n=15) and their primary communication partners (n=5) listened to synthesized sentences with differing prosodic content generated from deep regression neural networks developed in our prior work. Specifically, the fundamental frequency fo contour of each sentence was manipulated in 4 ways: (1) flattened to the average fo, (2) altered to discrete sentence-level classification of muscle activity, (3) altered to continuous mapping of muscle activity, and (4) filtered to emulate speech from an electrolarynx (EL). Listeners ranked the fo contours of each sentence in terms of speech naturalness and the importance of various speech aid features.

Results:

Continuous contours rated higher than all other types of contours, and monotonic EL contours rated the lowest. Speech aid features were rated highest to lowest in the following order: sound quality, intelligibility, pitch, delay, volume, handsfree, maintenance, cost, wearability, training, and visibility.

Conclusion:

These results will help inform future development of silent speech interfaces and shape priorities of communication devices toward the preferences of their users.

Keywords: electromyography, pitch, alaryngeal speech, silent speech interface, laryngectomy, speech naturalness

Introduction

Voice is one of the most unique and ubiquitous methods of transmitting identity. A short utterance provides effortless self-expression, informing a speaker’s gender, personality, mood, desires, and beyond. Consequently, people who lose the ability to speak naturally suffer negatively altered self-perception, social withdrawal, and lower quality of life, among other severe negative psychosocial concomitants.1,2

Total laryngectomy involves complete removal of the larynx and, thus, an individual’s ability to speak naturally. This surgical procedure is the primary treatment for advanced-stage laryngeal cancer and the standard of care when nonsurgical management of laryngeal cancer fails.3 Without their vocal folds to produce voice, people with laryngectomy (PWL) must rely on alternative vibration sources to produce speech; however, these technologies remain inferior to natural laryngeal speech. Two common methods involve vibration of the pharyngoesophageal (PE) segment: tracheoesophageal speech (TE, requires intensive maintenance and additional surgery4) and esophageal speech (ES, requires extensive training, and few master the technique5,6). An electrolarynx (EL) is a common, non-invasive method that offers a rapid learning curve. As such, it often serves as a primary tool for communication or as a backup option for patients using tracheoesophageal (TE) or esophageal speech (ES). When placed against the neck, the handheld device supplies electromechanical vibration to the vocal tract. Yet the constant use of a patient’s hand can be excessively cumbersome and especially challenging for patients with impaired manual dexterity.4,7

Physical limitations aside, alaryngeal options provide most PWLs with a serviceable, albeit unnatural, voice. Still, many do not hear their voice: naïve listeners and PWLs consistently rate all alaryngeal options as sounding less intelligible and more artificial compared to pre-surgical voices.2 The lack of naturalness can be partially attributed to these technologies’ lack of effective dynamic pitch modulation capabilities. Modulating pitch is an essential prosodic aspect of verbal communication that facilitates meaning, attitude, emotion, and more. With TE and ES, excitation of the PE segment has unstable oscillation patterns and lower periodicity compared to natural vocal folds, all leading to a much lower fundamental frequency fo, inconsistent pitch contours, unnatural spectral shaping, and an overall rough-sounding voice.8 EL speech is often perceived as inferior to TE and ES in terms of naturalness as it has a very mechanical and monotone sound quality.5 Attempts to incorporate dynamic pitch modulation to ELs have had mixed outcomes. Studies have shown intelligibility was slightly higher for speakers using a pitch-controlled EL (e.g., a TrueTone EL) compared to those with flattened frequencies.9,10 However, commercially available pitch-controlled ELs use buttons and/or switches that are arguably impractical for real-time modulation of pitch due to system complexity and high cognitive loads.7 Additionally, despite the increase in fo control, speakers do not think that their pitch-modulated EL speech represents how they sounded before total laryngectomy.11 Other efforts for improving EL naturalness provide hands-free control via biosignals associated with vocal f0. (1216) Despite success within the research community, these devices are not yet commercially available since they still fundamentally sound like an EL (e.g., mechanical or robot-like).

Constraints of existing alaryngeal options have inspired the emerging field of silent speech interfaces (SSIs). SSIs enable oral communication in the absence of vocalization by decoding physiological signals associated with speech production.17 Without requiring new vocalization or articulation techniques, SSIs are ideal for PWLs. Various sensory modalities have been used to capture these signals,1827 but surface electromyography (sEMG)-based SSIs are the most promising due to their non-invasiveness, portability, real-time applicability, and high efficacy.11,28,29

Surface EMG-based SSIs capture articulatory muscle activity of the face and voice-related activity of the neck during subvocal speech (inaudibly mouthing words) via non-invasive electrodes affixed to the skin surface. Previous work has demonstrated these systems’ ability to accurately translate sEMG signals to continuous speech with low latency, and thus show promise as a viable alternative to conventional alaryngeal speech.30 Yet incorporating prosody into synthetic speech is a developing effort for sEMG-based SSIs. While a natural relationship between extrinsic laryngeal muscles and pitch modulation has been established,1214,3133 effectively translating this relationship into synthetic voice synthesis remains an area of exploration.

Among various signal-processing approaches applied to date, machine learning techniques have had the most success in extracting pitch from sEMG signals. Nakamura et al.34 first demonstrated the potential of using sEMG signals from the neck to predict fo using Gaussian mixture models, achieving a moderate correlation between observed and predicted fo of r=.49. Diener et al.35 took a different approach by using a feed-forward neural network to predict fo contours from a quantization approach rather than a continuous sequence, but did not observe good performance (r=.27). Building off these previous studies, our recent work achieved the highest correlation between observed and predicted fo(r=.92) using deep regression neural networks.36 Although these recent developments show improvements in the objective accuracy of fo prediction from an sEMG signal, evaluating whether prosody is effectively incorporated into synthetic speech requires subjective assessments from listeners. Few SSI studies have explored these outcomes,11,35 and it remains unclear the extent to which different types of fo cues (e.g., fine-tuned or generic) affect the reception of synthetic speech. Systems pursuing a fine-tuned approach of fo estimation developed user-specific models that at least initially rely on acoustic information.3537 Such systems are incompatible with most PWL as they typically do not seek speech aids until after laryngectomy surgery. Furthermore, no study has investigated whether these user-specific systems generate more natural-sounding prosodic cues compared to generic systems. For instance, it is possible that a PWL would prefer generically generated prosodic information instead of a fine-tuned fo contour, as they do not require acoustic input and resulting computational overhead (e.g. increased delay and design constraints). It is additionally unclear how the preferences of PWL compare to their communication partners, who typically play a central role in facilitating effective communication. Yet, to our knowledge, previous investigations have only involved perceptual preferences from naïve listeners rather than the intended target population.11,35 As such, one of the primary goals of this work was to elucidate the preferences of both PWLs and their primary communication partners regarding methods of fo generation in synthetic speech. We utilized the synthetic speech outputs from sEMG-based SSIs described in our prior work11,36 to understand how manipulations to—or the absence of—a time-varying fo contour affect the perceived naturalness of sEMG-based synthetic speech. For this investigation, we compared listener preferences across three types of fo contours:

  1. A flat fo contour that does not vary with time,

  2. A generic fo contour that is based on discrete classifications of phrase-level muscle activity, and

  3. A fine-tuned fo contour that is generated by continuous mapping of time-varying muscle activity to fo.

We additionally included a synthetic version of EL speech as a reference in our analysis, given its popularity as an alaryngeal speech option for PWLs. We hypothesized that the fo contour of the more nuanced approach of continuous mapping of muscle activity would be preferred by PWLs and their primary communication partners over the flat contour, discrete mapping, and synthesized EL speech.

The second goal of this work was to determine the characteristics of speech aids (e.g., SSI or EL) that PWLs deemed most important for effective communication. Factors such as intelligibility,38 voice quality,39 cost,4042 and ease of use,42 among others, are known to affect speech rehabilitation following total laryngectomy; however, recent empirical evidence as to how these characteristics are prioritized by PWLs and their primary communication partners is limited. This lack of information raises concerns that current research into high-tech speech aids, such as SSIs, might be prioritizing aspects that are less important to their intended users. For instance, a speech aid might offer excellent voice quality but be challenging to use, leading to low adoption rates among PWLs. Conversely, a less sophisticated device might be favored due to its relative cost and ease of use, even if it provides lower voice quality. It therefore remains unclear which characteristics should be prioritized. Thus, a secondary goal of our work was to address these gaps by surveying the importance of various speech aid characteristics to these core users.

The development of SSIs or other high-tech speech aids that incorporate the most preferred prosodic qualities of synthetic speech may be an important step toward improving the quality of life of PWLs. Beyond discerning the preferred prosodic qualities of synthetic speech for PWLs and their primary communication partners, we provide a list of essential considerations to guide future research and development in communication strategies and technologies for PWLs.

Methods

Development of Speech Stimuli

Collection of sEMG and Acoustic Signals

Speech stimuli were sourced from a retrospective analysis, the methods of which are described in Vojtech et al.36 In brief, speakers with typical voices were seated in a quiet room where they were instructed to perform a series of speech tasks as sEMG and acoustic data were concurrently recorded. The speakers produced sustained vowels, legatos, syllables, phrases, reading passages, questions, and monologues to induce a range of vocal behaviors and articulatory patterns. The sEMG signals were collected via two wireless Trigno Quattro sensors (Delsys, Natick, MA, USA), which provided eight single-differential electrodes. Each sensor was placed over a distinct region of the face and neck as described in Meltzner et al.43,44 Acoustic signals were recorded using a headset omnidirectional microphone (Movo LV-6C XLR) that was positioned at a 45-degree angle from the midline and 4–7 cm from the lips. Microphone signals were pre-amplified (ART Tube MP Project Series) and digitized at 44.1 kHz (National Instruments USB NI-6251). Time-aligned sEMG and acoustic signal acquisition was performed using a triggering setup in EMGworks software (Delsys, Natick, MA, USA) with a custom trigger module to connect the data acquisition board and sEMG base station trigger port.

For the purposes of the current study, we focused on the sEMG and acoustic data from four participants who were native English speakers (2 female, 2 male; M = 29.8 years, SD = 9.6 years). None of the speaker participants had a history of speech, language, or hearing disorders. Three speech tokens of varying phrasal stress patterns—initial word stress, final word stress, and unstressed—were chosen from the retrospectively collected speech corpus for our analysis. These tokens included: (1) “THAT won’t make a difference,” (2) “Are you SERIOUS?,” and (3) “There is nothing I can do about it.” These selections were based on empirical assessments conducted by the authors to confirm phrasal stress on the intended word, 45 indicated by capitalized words.

Manipulating fo Contours

Using the concurrently collected acoustic and sEMG data from each speech token, we generated a series of synthetic speech stimuli for subsequent perceptual evaluation. Three types of fo contours were synthesized (flat, discrete, and continuous) and overlaid each acoustic signal (see Fig. 1 for example). Flat fo contours had uniform pitch variation across the speech signal and served as a baseline condition. Discrete and continuous fo contours were synthesized to examine the effects of pitch variation type—either distinct pitch variations associated with specific stress patterns or natural pitch variations that mirror typical speech inflections and prosodic contours — on perceived speech naturalness. Generation of these contours is described as follows:

Figure 1.

Figure 1.

Illustration of discrete (purple), continuous (gold), and flat (green) fo contours generated for one speaker from the phrase “Are you SERIOUS?”

  1. Flat Contours: Flat speech stimuli were generated using the “monotonize” function of the Praat Vocal Toolkit,46,47 which automatically analyzes the fo values across an acoustic recording input and then sets all fo values to a consistent median value, which was set per stimulus as the mean fo of the original signal.

  2. Discrete Contours: Speech stimuli with discrete fo contours were synthesized using a modified version of the sEMG-based SSI method described in Vojtech et al.11 As phrasal stress categorization and lexical recognition is outside the scope of the current study, we developed a streamlined text-to-speech algorithm inspired by the referenced methodology to generate an fo contour via prototypical representations of phrasal stress.45 Stress on the first word (“THAT”) or last word (“SERIOUS”) was achieved through empirically derived gain modifications to increase the fo of the intended word.4850 Pitch manipulations were implemented to overlay the generated fo contours onto flat speech stimuli using routines based on the Pitch Synchronous Overlap and Add (PSOLA) algorithm.51,52

  3. Continuous Contours: Continuous speech stimuli were generated using the deep regression neural networks described in Vojtech et al.,36 which were already trained on the concurrent sEMG and acoustic data of the four speakers. In brief, the fo contour was extracted from each speech token using the autocorrelation-based Praat method of the Parselmouth46,53 package in Python (v.3.8). Minimum and maximum fo values were set to 65 Hz and 475 Hz, respectively, 5456 and the time step for this algorithm was set to default (0.75/minimum fo). The acoustic and sEMG data were time-aligned using a dynamic time-warping algorithm from the linmdtw package57 in Python with a hop value of 0.010 s. All acoustic and sEMG signals were then windowed at a frame size of 40 ms with a 20-ms step shift; for each sEMG signal window, a set of 20 features was calculated using procedures detailed in Vojtech et al.36 Deep regression neural networks were trained for each speaker using the set of sEMG features and observed fo from their signals, then predicted fo for each window. Using these trained models for each of the four speakers, we used existing routines from the Parselmouth package in Python to replace the existing fo contour of each speech token with the predicted fo.

  4. EL Speech: In addition to the three types of fo contours that were overlaid onto existing speech tokens, we synthesized speech stimuli to mimic the acoustic characteristics of EL speech. The EL speech stimuli were derived using the vocoder and EQ 10 bands functions of the Praat Vocal Toolkit.46,47 The mean fo value of the original speech token was first estimated using the autocorrelation-based Praat method with minimum and maximum fo values set to 65 Hz and 475 Hz; all other parameters were kept at their default settings. The speech token was then filtered via source-filter synthesis with a pulse train carrier waveform at a frequency matching that of the identified mean fo. The spectral characteristics of the resulting speech were additionally adjusted using the automated equalizer function from the Praat Vocal Toolkit, with an emphasis on 1) boosting the lower frequencies to enhance the fo prominence and 2) attenuating the higher frequencies to replicate the acoustic profile of EL speech. To mimic the recording conditions of original acoustic samples (and hence, the discrete and continuous synthetic samples) that were collected in a low-noise environment, and to make the synthetic EL speech more like true EL speech,5860 white noise was added to the processed EL speech stimuli at a signal-to-noise ratio of 25 dB.

Evaluation of Speech Stimuli

Listeners

Fifteen PWLs (10 male, 5 female; M = 54 years, SD = 4.13 years) and five primary communication partners (1 male, 4 female; M = 59 years, SD = 9.56 years) were recruited for the study through in-person outreach at community-based support groups by author J.T.H. All participants spoke English and were naïve to the purpose of the study. Refer to Table 1 for demographic information of PWLs and their primary communication partners. Prior to participation, all listeners provided informed, written consent in compliance with the Western Institutional Review Board (Protocol #20182089).

Table 1.

Demographic information of people with laryngectomy (PWLs) and primary communication partners (“PCP”).

Time Since Laryngectomy

ID Sex Age Years Months Primary Method of Communication Other Methods Used (if applicable)

PWL01 Male 58 6 0 TE EL
PWL02 Female 59 2 8 ES
PWL03 Male 62 0 10 TE
PWL04 Male 51 15 6 TE EL
PWL05 Female 52 1 6 EL, TE
PWL06 Male 51 0 5 EL
PWL07 Female 57 4 7 TE EL, ES
PWL08 Male 57 0 4 TE
PWL09 Female 50 3 2 EL, TE
PWL10 Male 52 4 8 TE EL, Dry-erase board, Text-to-speech apps
PWL11 Male 56 0 9 EL EL, Text-to-speech apps
PWL12 Female 57 2 3 TE
PWL13 Male 51 2 9 EL, TE ES
PWL14 Male 48 1 10 EL
PWL15 Male 37 23 11 EL ES
PCP01 Female 52 N/A N/A N/A N/A
PCP02 Female 67 N/A N/A N/A N/A
PCP03 Female 69 N/A N/A N/A N/A
PCP04 Female 62 N/A N/A N/A N/A
PCP05 Male 57 N/A N/A N/A N/A

Note. TE=tracheoesophageal speech, ES=esophageal speech, EL=electrolaryngeal speech.

Online Survey

The listening procedure was completed remotely by all participants using the online behavioral research platform Gorilla Experiment Builder (www.gorilla.sc). Listeners were instructed to complete both study tasks in a single session in a quiet environment. Headphones were recommended to provide optimal acoustic fidelity but not required.. Each listening session began with a volume adjustment from the Gorilla open materials repository to allow listeners to adjust their device volume to a comfortable level during the presentation of a sample speech token. Once they set the volume, they were instructed to keep the volume constant throughout the entire experiment.

Listeners were presented with a visual sort-and-rate (VSR) task, in which they were instructed to rate the overall naturalness of speech stimuli relative to one another. VSR has been reported with high reliability for ratings of naturalness in speakers with voice disorders61 and for synthesized speech tokens.62 A definition of speech naturalness was provided for the entire duration of the task as “Speech naturalness is not judged by how well you understand the words that are being said, but rather your preference for how it sounds in terms of rate, rhythm, intonation and voice quality”.63 This definition remained visible throughout the naturalness rating tasks.

In the VSR task, listeners could play each audio file as many times as needed and were not instructed as to the order in which to rate the naturalness of the synthetic speech tokens. A progress bar displayed at the bottom of the screen guided listeners through each task page, where they rated four synthetic speech tokens at a time. Each token featured the same phrase (“THAT won’t make a difference,” “Are you SERIOUS?,” or “There is nothing I can do about it.”) spoken by the same speaker, but synthesized using one of four manipulation methods (flat, discrete, continuous, EL). Listeners assessed the naturalness of each audio file by adjusting a continuous scale graded from 0 to 100, anchored between “least natural” (0) and “most natural” (100). After rating the naturalness of all four tokens, listeners were instructed to play each audio file again and adjust their ratings relative to each other, ensuring that no two tokens received the same score. Once they were comfortable with their responses, they were instructed to press “Next” to move onto the next page of the task. The VSR task involved evaluating a series of 48 unique synthetic speech tokens (1 phrase × 3 stress types × 4 manipulation methods × 4 speakers). Additionally, 12 tokens were repeated in a pseudorandom order (1 phrase × 3 stress types × 4 manipulation methods) to assess intralistener reliability, resulting in a total of 60 synthetic speech token judgements. The order of each task page was randomized to minimize biased responses from anticipating speakers’ voices.

After completing the VSR task, listeners were asked to rate the degree of importance of various features of speech aids or alaryngeal speaking techniques on a continuous scale graded from 0 to 100, anchored between “least important” (0) and “most important” (100). Listeners were asked to consider all features prior to rating them. As in the VSR task, no two features were allowed to have the same rating. Definitions of each feature were provided to participants and are described in Table 2.

Table 2.

Definitions of each speech aid feature.

Speech Aid Feature Definition

Device Cost Some devices might not be fully covered by insurance (e.g., Medicare). How important is it to avoid out-of-pocket costs (for example, $1000 initial cost)
No Wearable Sensors Some devices might require you to wear small sensors near your mouth or under your chin. How important is it to avoid having skin-placed sensors?
Device Visibility Some devices are visible to people around you. How important is it to have a device not visible or obvious to others?
Device Volume Some devices or alaryngeal speech techniques are louder than others. How important is it for you to speak loudly when you want?
Pitch Modulation Some devices or alaryngeal speech techniques are monotone (flat pitch) while others allow pitch control. How important is it for you to control your vocal pitch?
Speech Intelligibility Some devices or alaryngeal speech techniques are easier to understand than others. How important is it to have everything you say understood?
Sound Quality Some devices or alaryngeal speech techniques sound better than others. How important is it to have a normal, natural-sounding voice and speech quality?
Hands-free Device Some devices or alaryngeal speech techniques require the use of one hand. How important is it to be able to speak hands-free?
No Speech Delay Some devices have a 1- to 3-second delay between when you start to speak and when a synthetic version of your speech is generated. How important is it that your speech has no delay?
No Training Required Some devices or alaryngeal speech techniques require training and practice before performing well. How important is it not to need training or practice?
Ease of Device Maintenance Some devices or alaryngeal speech techniques require routine maintenance (e.g., battery changes, valve replacement, etc.). How important is it to avoid maintenance?

Statistical Analysis

Naturalness ratings (0–100) and speech aid feature importance scores (0–100) were extracted directly from the VAS scores. First, intrarater reliability was calculated for the naturalness rating task using two-way random effects intraclass correlation coefficients (ICCs) to assess consistency of values. Listeners with poor reliability scores (ICC < 0.50) were removed from further analysis.64 Due to the unbalanced sample sizes (15 PWLs and 5 primary communication partners), Levene’s test for homogeneity of variance was calculated for both analyses.

A two-way mixed effects analysis of variance (ANOVA) was performed to evaluate how mean naturalness ratings were impacted among the four types of fo contours (flat, discrete, continuous, EL) and group (PWL, primary communication partner) across participants. Type of contour, group, and the interaction between contour type and group were fixed factors. An alpha level of .05 was used for significance tests. A squared partial curvilinear correlation ηp2 was used to calculate effect sizes. Fixed effects tests used the Satterthwaite approximation and variance estimation was calculated via maximum likelihood. Tukey simultaneous post hoc tests were used to evaluate pairwise comparisons of mean naturalness ratings across the types of fo contours using 95% confidence intervals.

A two-way ANOVA was performed to evaluate differences in mean importance score between speech aid features and listener group. An alpha level of .05 was used for significance tests. A squared partial curvilinear correlation was used to calculate effect sizes. Tukey simultaneous post hoc tests were then performed to evaluate pairwise comparisons of the mean importance score between speech aid features using 95% confidence intervals. Effect sizes were calculated via Cohen’s d. All analyses were completed using Minitab version 20.

Results

Naturalness Ratings

One participant (PWL15) was found to have low intrarater reliability, and their data was not included in the naturalness rating analysis. For the remaining 19 participants, the average intrarater reliability was ICC = .87 (SD = .13, range: .45–.99). There were no significant differences in the variance of naturalness ratings between PWLs or primary communication partners (p = .27). Table 3 shows the results of the two-way mixed effects ANOVA. The model showed a significant, small interaction effect of contour×groupp<.001,np2=.02 and a significant effect of contour on the naturalness ratings (p < .001). A Tukey post-hoc test revealed that PWLs rated continuous (M = 62.63, 95% C.I. = [59.50, 65.76]), discrete (M = 58.69, 95% C.I. = [55.56, 61.82]), and flat (M = 58.83, 95% C.I. = [55.70, 61.96]) contours significantly higher than EL speech (M = 6.61, 95% C.I. = [3.48, 9.74], all p < .001). Primary communication partner also rated EL speech (M = 2.74, 95% C.I. = [−1.71, 7.01]) significantly lower than all other contours (p < .001). However, Tukey post-hoc tests found primary communication partners rated continuous contours (M = 74.73, 95% C.I. = [70.37, 79.10]) significantly higher than flat contours (M = 62.53, 95% C.I. = [58.17, 66.90], p = .004), but discrete contours (M = 69, 95% C.I. = [64.64, 73.36]) did not rate significantly different from continuous (p = .634) or flat contours (p = .478; Fig. 2).

Table 3.

Results of analysis of variance test performed for speech naturalness ratings.

Factor df np2 F p

Participant 18 0.74 0.25 < .001
Contour 3 0.21 513.14 < .001
Group 1 1.35 .260
Contour x Group 3 0.86 7.54 < .001

Figure 2.

Figure 2.

Mean ratings of speech naturalness for each type of fo contour for PWL (maroon) and communication partners (peach).Error bars represent 95% confidence intervals. Post-hoc pairwise comparisons revealed statistical significance as indicated by the brackets (*p < .05).

Speech Aid Feature Preferences

The Levene’s test showed that the variance of the feature importance scores between PWL and primary communication partners was not significantly different (p = .50). The results from the two-way ANOVA showed that only speech aid feature demonstrated a significant, large effect in the model (p < .001, ηp2=0.32; Table 4). The Tukey post-hoc analysis showed that sound quality was the most important speech aid feature, with a significantly higher preference than maintenance, cost, wearability, training, and visibility (Fig. 3a). However, sound quality was not statistically significantly higher than intelligibility, pitch, delay, volume, handsfree, or maintenance (which received the subsequent greatest preference scores in the listed order; Fig. 3b). Visibility demonstrated the lowest score for importance. All statistically significant comparisons showed a strong effect size (d > 0.8) except for training, which showed a medium effect with handsfree, intelligibility, and pitch (d > 0.6) and small effect with volume and delay (d > 0.2).

Table 4.

Results of analysis of variance test performed for speech aid feature ratings

Factor df np2 F p

Feature 10 0.32 9.32 < .001
Group 1 0.12 .732
Feature x Group 10 0.99 .456

Figure 3.

Figure 3.

Figure 3.

a) Mean of scores for each speech aid feature. Error bars represent 95% confidence intervals.

b) Statistical significance matrix of speech aid features. Orange squares=p < .05, teal squares=p > .05, gray square indicates no comparison was performed.

Discussion

This study aimed to elucidate how a core group of potential SSI users perceive different modulations of prosodic cues in synthesized speech as well as their preferred features in speech aids. Three sentences were synthesized from a state-of-the-art, sEMG-based SSI. Each synthesized sentence had four variations of fo contours: flat, discrete, continuous, and EL. Twenty listeners, consisting of PWLs and primary communication partners, rated their perceived naturalness of each variation of each sentence. Additionally, listeners provided ratings of 11 common speech aid features in terms of which they believed to be the most or least important. The results of this study may serve as a guide for future development of SSIs to prioritize features that meet PWL needs regarding speech prosody and device characteristics.

Perceived Speech Naturalness

We hypothesized that the continuous pitch contours would be rated the highest among listeners. Our results supported this conclusion as continuous pitch contours were rated higher than all other types. This result was unsurprising as the continuous contours were personalized to the speaker’s muscle activity rather than generically generated based on the grammatical content of the sentence.

Interestingly, we found no difference between discrete and flat contours. This finding contradicts our previous work which found that synthesized speech was perceived as more natural with sentence-level variations in fo than with a flat fo.62 These differences may be explained by the listener group of the current study comprising experienced users of alaryngeal speech while our previous study involved naïve listeners. Since PWLs and their primary communication partners may be more acquainted with the prosodic limitations of alaryngeal speech, the flat fo contour may sound more familiar thus leading it to have a similar naturalness rating to the discrete contours. This phenomenon has been observed in several other types of disordered speech where naïve listeners rate atypical speech more harshly than patients or clinicians.6567 While it is reasonable to speculate this to be an underlying cause, more work is needed to investigate this notion, specifically for synthesized speech from PWLs. Our contradictory results may additionally be explained by different methodologies to produce synthetic speech: a generic voice bank as used in our previous work instead of signal mapping based on the individual’s own acoustic characteristics as used here. Future work should aim to more comprehensively examine the reception of these different methodologies by both SSI users and listeners.

There was a significant interaction effect between listener group and the different types of contours on the mean naturalness rating. Both PWLs and their primary communication partners agreed that continuous contours were the most natural, EL speech was the least natural, and that there was no significant difference between discrete and flat contours. Yet overall naturalness rating of the continuous contours differed between listener groups. For PWLs, continuous contours were rated the highest, but not significantly different from the other two fo contours. In contrast, primary communication partners found continuous contours significantly preferable than flat contours but did not perceive them as different from discrete contours. The discrepancy in statistical significance may be explained by the imbalance in group size (14 PWLs, 5 partners). Additionally, the difference may arise from the varying perspectives on speech naturalness. PWLs may have rated speech naturalness relative to their own monotonic alaryngeal voice, whereas primary communication partners likely assessed naturalness of the synthesized sentences relative to typical speech. This effect parallels findings of our previous work, where EL users preferred using a monopitch EL despite having access to an EL with pitch control; in this case, users reported that ELs with pitch control were cognitively complex and not worth the additional effort to master.11 Over time, PWLs may identify with their relatively monotonic voices and no longer associate modulation with naturalness, thereby resulting in a weak differentiation between flat and dynamic fo contours. Likewise, their communication partners might become accustomed to diminished prosody and thereby discount nuanced differences between continuous and discrete fo patterns, despite appreciating flat contours as different from continuous. These subtle shifts in fo contour perception could have been accentuated by the marked difference between our synthesized EL versus laryngeal speech samples, causing both listener groups to perceptually lump the non-EL samples together. Given that EL speech is consistently perceived as highly unnatural (as verified in our results), future work might better distinguish among fo modulation strategies for synthesized SSI speech by eliminating comparisons to EL samples and instead using natural speech as a point of comparison.

Speech Aid Feature Preferences

Most studies on speech aids—particularly those focusing on speech-generating devices—prioritize speech quality and intelligibility as primary outcomes. However, there is a notable gap in research concerning other characteristics that may be critical to end users, especially from the perspective of current needs and technological advancements. To our knowledge, no recent studies have systematically identified and rated the relative importance of various features of speech aids according to users aside from sparse user-specific case studies. Addressing this gap is essential for identifying common trends in preferred features and understanding which features vary in importance among different users. To initiate this discourse, we surveyed a subset of individuals who use speech aids (i.e., PWLs) and their primary communication partners to rate the importance of several different speech aid features. Across participants, features were rated in the following order: sound quality, intelligibility, pitch, delay, volume, handsfree, maintenance, cost, wearability, training, and visibility.

The 11 features can be categorized into two distinct groups: communication reception and physical operation. The results of our study indicate that users of speech aids prefer effective communication (intelligibility, sound quality, pitch, volume, and delay) over operational considerations (training, handsfree functionality, maintenance, cost, wearability, and visibility). Specifically, all five features associated with communication reception were rated as more important than features associated with the physical operation of the speech aid. These results underscore the need for alaryngeal speech aid development to focus on enhancing communication effectiveness, such as incorporating prosodic cues in the speech output of these devices. Broadly, these findings resonate with the perspectives observed among augmentative and alternative communication (AAC) users, in that the benefits of “being understood” outweigh the challenges of using a communication aid.68 One focus group discussing the research priorities of AAC users and their communication partners similarly concluded that such systems should be responsive to individual needs and allow users to communicate successfully in specific situations.69 The results of our study further reinforces these values while also providing new quantitative evidence of what PWLs prioritize in their speech aids.

Limitations and Future Work

Our current study demonstrates how a core set of SSI users who rely on alaryngeal speech prefer synthesized versions of typical speech relative to synthesized EL speech. Since PwL and their primary communication partners are experienced listeners, we acknowledge the potential for bias in their preferences compared to naïve listeners. Future research should include both experienced and naïve listeners to better understand how exposure to alaryngeal speech influences speaker and listener preferences. Furthermore, it is important to note that speech and voice have a wealth of characteristics that were outside the scope of this study. For one, prosody includes elements beyond fo modulation—including intensity, rate, and other spectral characteristics—which future work should consider. Although speech naturalness includes various elements of prosody, other metrics are also relevant for assessing speech function. In this study, we selected speech naturalness as a primary measure to balance the depth of information obtained with the overall study duration. Yet other measures, such as acceptability and communication efficiency, may provide insights into additional aspects of speech health.62,70

While the purpose of this study was to assess a key aspect of synthetic speech perception (i.e., naturalness), future studies should consider other metrics and a range of social contexts to capture a comprehensive assessment of the speech. Beyond prosodic elements, other variables affect the rating of naturalness. Dall et al.71 found that spontaneous conversational speech was consistently rated more natural-sounding than speech read from a prompt (as was done in our study). Additionally, the ratings were significantly different even when listeners were instructed to perceive the speech “as if they were having a conversation” or “as if somebody was reading aloud.”71 Using conversational speech was outside the scope of our current study given our focus on the perceptual impact of pitch contours, but broader speaking contexts should be studied. Future work should also examine a larger range of fo contours and synthesis approaches. For example, although our discrete fo contours were designed to best represent natural and appropriate pitch patterns, other patterns with wider or narrower fo ranges may have been perceived differently. Likewise, the continuous fo contours derived directly from speaker participants’ sEMG recordings could have mapped more widely or narrowly onto fo ranges instead of strictly relating to their acoustic recordings. Ultimately, sEMG-based prosodic control will likely be individualized for PWLs working with their healthcare team during their initial SSI fitting and instruction.

Some patients might do best with pitch contours selected through sEMG-based pattern recognition—particularly if their treatment includes radical neck dissection and radiation therapy, thereby limiting pitch-related sEMG content. In contrast, other patients might prefer pitch control continuously mapped onto their sEMG signals. In either case, dynamic fo range could be easily adjusted to best serve a user’s language, dialect, speaking context, and mood all afforded by the powerful flexibility of speech synthesis. The speakers in our study were pseudorandomly selected from a larger sample (n = 10) to ensure an equal male-to-female ratio while balancing statistical power and the duration of the perceptual study. However, it is important to note that these speakers had normal neck anatomy, thereby generating sEMG signals that may not accurately reflect the physiological variations present in individuals with altered neck anatomy (e.g., PWL). The selected speakers were also able to produce fo contours for model training, which may not be attainable for PWL who rely on alternative speech production methods. Consequently, while our study focused on perceived outcomes rather than the specific inputs of SSIs, our findings may not be broadly applicable to the PWL population. The limited sample size of speakers (n = 4) further compounds this issue, as it restricts the diversity of speaker characteristics and experiences represented in the results. As a result, caution should be exercised when extrapolating these findings to the wider PWL community. Future research with a more diverse and representative sample of speakers is essential to enhance the generalizability of the findings and ensure they are applicable to individuals with varying speech production challenges.

Our study gained some valuable initial insight into what speech aid features are most important to their users. Although PWLs and primary communication partners did not exhibit clear differences in their preferences, it is possible that other factors influence these ratings. For example, in the development of hearing aids—another assistive technology with extensive consumer research due to its broader market—studies have found that user preferences vary based on demographic factors and the severity of hearing loss.72 As such, future work should adopt similar considerations to investigate what factors influence consumers’ preferences in both SSIs and their application in speech aids.

Conclusion

As SSIs continue to advance, obtaining feedback from target users is essential to guiding the effective development of these technologies. Incorporating the preferences of both users and listeners is key for the clinical viability and acceptance of sEMG-based SSIs for alaryngeal speech. Our study shows that PWLs and their primary communication partners strongly preferred the synthetic speech with prosody generated by our models over monotone EL speech. Although it is well documented that prosodic cues involving pitch are an important factor in perceived speech naturalness, the method in which these cues are introduced into synthetic speech had not yet been investigated; our findings indicate that continuous mapping of fo to time-varying muscle activity was preferred over discrete, sentence-level mapping. However, future studies should investigate how these preferences hold across different types of speech, such as conversational speech. Finally, we also obtained a list of speech aid features that PWLs and their primary communication partners valued most. Having quality communication was deemed more important than reducing physical burdens as quality of speech, intelligibility, pitch control, little delay, and volume control rated as the most important features. Taken together, these listener preferences and speech aid design feature rankings provide foundational insight for SSI development.

Footnotes

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

References

  • 1.Danker H, Wollbrück D, Singer S, Fuchs M, Brähler E, Meyer A. Social withdrawal after laryngectomy. Eur Arch Otorhinolaryngol. 2010;267(4):593–600. doi: 10.1007/s00405-009-1087-4 [DOI] [PubMed] [Google Scholar]
  • 2.Sharpe G, Camoes Costa V, Doubé W, Sita J, McCarthy C, Carding P. Communication changes with laryngectomy and impact on quality of life: a review. Qual Life Res. 2019;28(4):863–877. doi: 10.1007/s11136-018-2033-y [DOI] [PubMed] [Google Scholar]
  • 3.Andaloro C, Widrich J. Total Laryngectomy. In: StatPearls. StatPearls Publishing; 2023. Accessed December 4, 2023. http://www.ncbi.nlm.nih.gov/books/NBK556041/ [PubMed] [Google Scholar]
  • 4.Kaye R, Tang CG, Sinclair CF. The electrolarynx: voice restoration after total laryngectomy. Med Devices (Auckl). 2017;10:133–140. doi: 10.2147/MDER.S133225 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Cox SR. Review of the Electrolarynx: The Past and Present. Perspectives of the ASHA Special Interest Groups. 2019;4(1):118–129. doi: 10.1044/2018_PERS-SIG3-2018-0013 [DOI] [Google Scholar]
  • 6.Eibling DE. Voice Restoration after Total Laryngectomy. In: Myers EN, Carrau RL, Eibling DE, et al. , eds. Operative Otolaryngology: Head and Neck Surgery. Vol 1. 2nd ed. W.B. Saunders; 2008:431–437. doi: 10.1016/B978-1-4160-2445-3.50054-6 [DOI] [Google Scholar]
  • 7.Meltzner GS, Hillman RE, Heaton JT, Houston KM, Kobler J, Qi Y. Electrolaryngeal speech : the state of the art and future directions for development. In: Contemporary Considerations in the Treatment and Rehabilitation of Head and Neck Cancer : Voice, Speech, and Swallowing. ProEd; 2005:571–590. [Google Scholar]
  • 8.Drugman T, Rijckaert M, Janssens C, Remacle M. Tracheoesophageal speech: A dedicated objective acoustic assessment. Computer Speech & Language. 2015;30(1):16–31. doi: 10.1016/j.csl.2014.07.003 [DOI] [Google Scholar]
  • 9.Watson PJ, Schlauch RS. The Effect of Fundamental Frequency on the Intelligibility of Speech With Flattened Intonation Contours. American Journal of Speech-Language Pathology. 2008;17(4):348–355. doi: 10.1044/1058-0360(2008/07-0048) [DOI] [PubMed] [Google Scholar]
  • 10.Watson PJ, Schlauch RS. Fundamental frequency variation with an electrolarynx improves speech understanding: a case study. Am J Speech Lang Pathol. 2009;18(2):162–167. doi: 10.1044/1058-0360(2008/08-0025) [DOI] [PubMed] [Google Scholar]
  • 11.Vojtech JM, Chan MD, Shiwani B, et al. Surface Electromyography–Based Recognition, Synthesis, and Perception of Prosodic Subvocal Speech. Journal of Speech, Language, and Hearing Research. 2021;64(6S):2134–2153. doi: 10.1044/2021_JSLHR-20-00257 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Goldstein EA, Heaton JT, Kobler JB, Stanley GB, Hillman RE. Design and implementation of a hands-free electrolarynx device controlled by neck strap muscle electromyographic activity. IEEE Transactions on Biomedical Engineering. 2004;51(2):325–332. doi: 10.1109/TBME.2003.820373 [DOI] [PubMed] [Google Scholar]
  • 13.Goldstein EA, Heaton JT, Stepp CE, Hillman RE. Training Effects on Speech Production Using a Hands-Free Electromyographically Controlled Electrolarynx. Journal of Speech, Language, and Hearing Research. 2007;50(2):335–351. doi: 10.1044/1092-4388(2007/024) [DOI] [PubMed] [Google Scholar]
  • 14.Kubert HL, Stepp CE, Zeitels SM, et al. Electromyographic control of a hands-free electrolarynx using neck strap muscles. Journal of Communication Disorders. 2009;42(3):211–225. doi: 10.1016/j.jcomdis.2008.12.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Stepp CE, Heaton JT, Rolland RG, Hillman RE. Neck and Face Surface Electromyography for Prosthetic Voice Control After Total Laryngectomy. IEEE Transactions on Neural Systems and Rehabilitation Engineering. 2009;17(2):146–155. doi: 10.1109/TNSRE.2009.2017805 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Heaton JT, Robertson M, Griffin C. Development of a wireless electromyographically controlled electrolarynx voice prosthesis. In: 2011 Annual International Conference of the IEEE Engineering in Medicine and Biology Society. ; 2011:5352–5355. doi: 10.1109/IEMBS.2011.6091324 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Denby B, Schultz T, Honda K, Hueber T, Gilbert JM, Brumberg JS. Silent speech interfaces. Speech Communication. 2010;52(4):270–287. doi: 10.1016/j.specom.2009.08.002 [DOI] [Google Scholar]
  • 18.Angrick M, Herff C, Mugler E, et al. Speech synthesis from ECoG using densely connected 3D convolutional neural networks. J Neural Eng. 2019;16(3):036019. doi: 10.1088/1741-2552/ab0c59 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Brumberg JS, Guenther FH, Kennedy PR. An Auditory Output Brain–Computer Interface for Speech Communication. In: Guger C, Allison BZ, Edlinger G, eds. Brain-Computer Interface Research: A State-of-the-Art Summary. SpringerBriefs in Electrical and Computer Engineering. Springer; 2013:7–14. doi: 10.1007/978-3-642-36083-1_2 [DOI] [Google Scholar]
  • 20.Crevier-Buchman L, Gendrot C, Denby B, et al. Articulatory strategies for lip and tongue movements in silent versus vocalized speech. In: ; 2011:1. Accessed February 29, 2024. https://shs.hal.science/halshs-00610870 [Google Scholar]
  • 21.Fabre D, Hueber T, Girin L, Alameda-Pineda X, Badin P. Automatic animation of an articulatory tongue model from ultrasound images of the vocal tract. Speech Communication. 2017;93:63–75. doi: 10.1016/j.specom.2017.08.002 [DOI] [Google Scholar]
  • 22.Fagan MJ, Ell SR, Gilbert JM, Sarrazin E, Chapman PM. Development of a (silent) speech recognition system for patients following laryngectomy. Medical Engineering & Physics. 2008;30(4):419–425. doi: 10.1016/j.medengphy.2007.05.003 [DOI] [PubMed] [Google Scholar]
  • 23.Hirahara T, Otani M, Shimizu S, et al. Silent-speech enhancement using body-conducted vocal-tract resonance signals. Speech Communication. 2010;52(4):301–313. doi: 10.1016/j.specom.2009.12.001 [DOI] [Google Scholar]
  • 24.Hueber T, Benaroya EL, Chollet G, Denby B, Dreyfus G, Stone M. Development of a silent speech interface driven by ultrasound and optical images of the tongue and lips. Speech Communication. 2010;52(4):288–300. doi: 10.1016/j.specom.2009.11.004 [DOI] [Google Scholar]
  • 25.Kimura N, Gemicioglu T, Womack J, et al. SilentSpeller: Towards mobile, hands-free, silent speech text entry using electropalatography. In: Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems. CHI ‘22. Association for Computing Machinery; 2022:1–19. doi: 10.1145/3491102.3502015 [DOI] [Google Scholar]
  • 26.Nakajima Y, Kashioka H, Shikano K, Campbell N. Non-audible murmur recognition input interface using stethoscopic microphone attached to the skin. In: 2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003. Proceedings. (ICASSP ‘03). Vol 5. ; 2003:V–708. doi: 10.1109/ICASSP.2003.1200069 [DOI] [Google Scholar]
  • 27.Porbadnigk A, Wester M, Calliess J p., Schultz T EEG-BASED SPEECH RECOGNITION - Impact of Temporal Effects: In: Proceedings of the International Conference on Bio-Inspired Systems and Signal Processing. SciTePress - Science and and Technology Publications; 2009:376–381. doi: 10.5220/0001554303760381 [DOI] [Google Scholar]
  • 28.Diener L, Herff C, Janke M, Schultz T. An initial investigation into the real-time conversion of facial surface EMG signals to audible speech. In: 2016 38th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). ; 2016:888–891. doi: 10.1109/EMBC.2016.7590843 [DOI] [PubMed] [Google Scholar]
  • 29.Janke M, Diener L. EMG-to-Speech: Direct Generation of Speech From Facial Electromyographic Signals. IEEE/ACM Transactions on Audio, Speech, and Language Processing. 2017;25(12):2375–2385. doi: 10.1109/TASLP.2017.2738568 [DOI] [Google Scholar]
  • 30.Jou SC, Schultz T, Walliczek M, Kraft F, Waibel A. Towards continuous speech recognition using surface electromyography. In: Interspeech 2006. ISCA; 2006:paper 1592-Mon3WeS.3–0. doi: 10.21437/Interspeech.2006-212 [DOI] [Google Scholar]
  • 31.Atkinson JE. Correlation analysis of the physiological factors controlling fundamental voice frequency. The Journal of the Acoustical Society of America. 1978;63(1):211–222. doi: 10.1121/1.381716 [DOI] [PubMed] [Google Scholar]
  • 32.Heaton JT, Goldstein EA, Kobler JB, et al. Surface Electromyographic Activity in Total Laryngectomy Patients following Laryngeal Nerve Transfer to Neck Strap Muscles. Ann Otol Rhinol Laryngol. 2004;113(9):754–764. doi: 10.1177/000348940411300915 [DOI] [PubMed] [Google Scholar]
  • 33.Stepp CE, Hillman RE, Heaton JT. Use of Neck Strap Muscle Intermuscular Coherence as an Indicator of Vocal Hyperfunction. IEEE Transactions on Neural Systems and Rehabilitation Engineering. 2010;18(3):329–335. doi: 10.1109/TNSRE.2009.2039605 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Nakamura K, Janke M, Wand M, Schultz T. Estimation of fundamental frequency from surface electromyographic data: EMG-to-F0. In: 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE; 2011:573–576. doi: 10.1109/ICASSP.2011.5946468 [DOI] [Google Scholar]
  • 35.Diener L, Umesh T, Schultz T. Improving Fundamental Frequency Generation in EMG-to-Speech Conversion Using a Quantization Approach. In: 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE; 2019:682–689. doi: 10.1109/ASRU46091.2019.9003804 [DOI] [Google Scholar]
  • 36.Vojtech JM, Mitchell CL, Raiff L, Kline JC, De Luca G. Prediction of Voice Fundamental Frequency and Intensity from Surface Electromyographic Signals of the Face and Neck. Vibration. 2022;5(4):692–710. doi: 10.3390/vibration5040041 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Ahmadi F, Araujo Ribeiro M, Halaki M. Surface electromyography of neck strap muscles for estimating the intended pitch of a bionic voice source. In: 2014 IEEE Biomedical Circuits and Systems Conference (BioCAS) Proceedings. IEEE; 2014:37–40. doi: 10.1109/BioCAS.2014.6981639 [DOI] [Google Scholar]
  • 38.Clements KS, Rassekh CH, Seikaly H, Hokanson JA, Calhoun KH. Communication after laryngectomy. An assessment of patient satisfaction. Arch Otolaryngol Head Neck Surg. 1997;123(5):493–496. doi: 10.1001/archotol.1997.01900050039004 [DOI] [PubMed] [Google Scholar]
  • 39.Hilgers FJM, Ackerstaff AH, Aaronson NK, Schouwenburg PF, Van Zandwijk N. Physical and psychosocial consequences of total laryngectomy. Clinical Otolaryngology & Allied Sciences. 1990;15(5):421–425. doi: 10.1111/j.1365-2273.1990.tb00494.x [DOI] [PubMed] [Google Scholar]
  • 40.Staffieri A, Mostafea BE, Varghese BT, et al. Cost of tracheoesophageal prostheses in developing countries. Facing the problem from an internal perspective. Acta Otolaryngol. 2006;126(1):4–9. doi: 10.1080/00016480500265935 [DOI] [PubMed] [Google Scholar]
  • 41.Xi S Effectiveness of voice rehabilitation on vocalisation in postlaryngectomy patients: a systematic review. Int J Evid Based Healthc. 2010;8(4):256–258. doi: 10.1111/j.1744-1609.2010.00177.x [DOI] [PubMed] [Google Scholar]
  • 42.Kapila M, Deore N, Palav RS, Kazi RA, Shah RP, Jagade MV. A brief review of voice restoration following total laryngectomy. Indian Journal of Cancer. 2011;48(1):99. doi: 10.4103/0019-509X.75841 [DOI] [PubMed] [Google Scholar]
  • 43.Meltzner GS, Heaton JT, Deng Y, De Luca G, Roy SH, Kline JC. Silent Speech Recognition as an Alternative Communication Device for Persons With Laryngectomy. IEEE/ACM Trans Audio Speech Lang Process. 2017;25(12):2386–2398. doi: 10.1109/TASLP.2017.2740000 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Meltzner GS, Heaton JT, Deng Y, Luca GD, Roy SH, Kline JC. Development of sEMG sensors and algorithms for silent speech recognition. J Neural Eng. 2018;15(4):046031. doi: 10.1088/1741-2552/aac965 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Patel R, Campellone P. Acoustic and perceptual cues to contrastive stress in dysarthria. J Speech Lang Hear Res. 2009;52(1):206–222. doi: 10.1044/1092-4388(2008/07-0078) [DOI] [PubMed] [Google Scholar]
  • 46.Boersma P, Weenink D. Praat: doing phonetics by computer. Published online 2021. http://www.praat.org/
  • 47.Corretge R Praat Vocal Toolkit. Published online 2012. https://www.praatvocaltoolkit.com
  • 48.Bolinger DL. A Theory of Pitch Accent in English. WORD. 1958;14(2–3):109–149. doi: 10.1080/00437956.1958.11659660 [DOI] [Google Scholar]
  • 49.Fry DB. Duration and Intensity as Physical Correlates of Linguistic Stress. The Journal of the Acoustical Society of America. 1955;27(4):765–768. doi: 10.1121/1.1908022 [DOI] [Google Scholar]
  • 50.Fry DB. Experiments in the Perception of Stress. Lang Speech. 1958;1(2):126–152. doi: 10.1177/002383095800100207 [DOI] [Google Scholar]
  • 51.Charpentier F, Stella M. Diphone synthesis using an overlap-add technique for speech waveforms concatenation. In: ICASSP ‘86. IEEE International Conference on Acoustics, Speech, and Signal Processing. Vol 11. Institute of Electrical and Electronics Engineers; 1986:2015–2018. doi: 10.1109/ICASSP.1986.1168657 [DOI] [Google Scholar]
  • 52.Moulines E, Charpentier F. Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones. Speech Communication. 1990;9(5):453–467. doi: 10.1016/0167-6393(90)90021-Z [DOI] [Google Scholar]
  • 53.Jadoul Y, Thompson B, de Boer B. Introducing Parselmouth: A Python interface to Praat. Journal of Phonetics. 2018;71:1–15. doi: 10.1016/j.wocn.2018.07.001 [DOI] [Google Scholar]
  • 54.Awan SN, Mueller PB. Speaking fundamental frequency characteristics of centenarian females. Clinical Linguistics & Phonetics. 1992;6(3):249–254. doi: 10.3109/02699209208985533 [DOI] [PubMed] [Google Scholar]
  • 55.Baken R Clinical Measurement of Speech and Voice. College-Hill Press; 1987. [Google Scholar]
  • 56.Coleman RF, Markham IW. Normal variations in habitual pitch. Journal of Voice. 1991;5(2):173–177. doi: 10.1016/S0892-1997(05)80181-X [DOI] [Google Scholar]
  • 57.Christopher J, Tralie Dempsey E. Exact, parallelizable dynamic time warping alignment with linear memory. In: ; 2020:462–469. doi: 10.5281/zenodo.4245470 [DOI] [Google Scholar]
  • 58.Koul RK, Allen GD. Segmental Intelligibility and Speech Interference Thresholds of High-Quality Synthetic Speech in Presence of Noise. Journal of Speech, Language, and Hearing Research. 1993;36(4):790–798. doi: 10.1044/jshr.3604.790 [DOI] [PubMed] [Google Scholar]
  • 59.Dreher JJ, O’Neill JJ. Effects of ambient noise on speaker intelligibility of words and phrases. The Laryngoscope. 1958;68(3):539–548. doi: 10.1002/lary.5540680335 [DOI] [PubMed] [Google Scholar]
  • 60.Loizou PC. Speech Enhancement: Theory and Practice. CRC Press; 2007. doi: 10.1201/9781420015836 [DOI] [Google Scholar]
  • 61.Anand S, Stepp CE. Listener Perception of Monopitch, Naturalness, and Intelligibility for Speakers With Parkinson’s Disease. Journal of Speech, Language, and Hearing Research. 2015;58(4):1134–1144. doi: 10.1044/2015_JSLHR-S-14-0243 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Vojtech JM, Noordzij JP, Cler GJ, Stepp CE. The Effects of Modulating Fundamental Frequency and Speech Rate on the Intelligibility, Communication Efficiency, and Perceived Naturalness of Synthetic Speech. American Journal of Speech-Language Pathology. 2019;28(2S):875–886. doi: 10.1044/2019_AJSLP-MSC18-18-0052 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Yorkston KM, Beukelman DR, Strand EA, Hakel M. Management of Motor Speech Disorders in Children and Adults. 3rd ed. ProEd; 2010. [Google Scholar]
  • 64.Koo TK, Li MY. A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. Journal of Chiropractic Medicine. 2016;15(2):155–163. doi: 10.1016/j.jcm.2016.02.012 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Franken MC, Bezooijen R van, Boves L. Stuttering and Communicative Suitability of Speech. Journal of Speech, Language, and Hearing Research. 1997;40(1):83–94. doi: 10.1044/jslhr.4001.83 [DOI] [PubMed] [Google Scholar]
  • 66.Kubitskey KM. Experience and Perception: How Experience Affects Perception of Naturalness Change in Speakers with Dysarthria. The Ohio State University; 2015. Accessed April 1, 2024. https://etd.ohiolink.edu/acprod/odb_etd/etd/r/1501/10?clear=10&p10_accession_num=osu1438186598 [Google Scholar]
  • 67.Ward EC, MacBean NA. Perceptual judgements of tracheoesophageal speech: the issue of listener bias. Asia Pacific Journal of Speech, Language and Hearing. 2003;8(1):24–35. doi: 10.1179/136132803805576336 [DOI] [Google Scholar]
  • 68.Childes JM, Palmer AD, Fried -Oken Melanie, Graville DJ. The Use of Technology for Phone and Face-to-Face Communication After Total Laryngectomy. American Journal of Speech-Language Pathology. 2017;26(1):99–112. doi: 10.1044/2016_AJSLP-14-0106 [DOI] [PubMed] [Google Scholar]
  • 69.O’Keefe BM, Kozak NB, O’Keefe BM, et al. Research priorities in augmentative and alternative communication as identified by people who use AAC and their facilitators. Augmentative and Alternative Communication. 2007;23(1):89–96. doi: 10.1080/07434610601116517 [DOI] [PubMed] [Google Scholar]
  • 70.Klopfenstein M Speech naturalness ratings and perceptual correlates of highly natural and unnatural speech in hypokinetic dysarthria secondary to Parkinson’s disease. JIRCD. 2016;7(1):123–146. doi: 10.1558/jircd.v7i1.27932 [DOI] [Google Scholar]
  • 71.Dall R, Yamagishi J, King S. Rating Naturalness in Speech Synthesis: The Effect of Style and Expectation. In: Speech Prosody 2014. ISCA; 2014:1012–1016. doi: 10.21437/SpeechProsody.2014-192 [DOI] [Google Scholar]
  • 72.Manchaiah V, Picou EM, Bailey A, Rodrigo H. Consumer Ratings of the Most Desirable Hearing Aid Attributes. J Am Acad Audiol. 2021;32(8):537–546. doi: 10.1055/s-0041-1732442 [DOI] [PubMed] [Google Scholar]

RESOURCES