Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2020 May 1.
Published in final edited form as: J Voice. 2017 Dec 23;33(3):269–276. doi: 10.1016/j.jvoice.2017.11.012

Estimation of source-filter interaction regions based on Electroglottography

Anil Palaparthi 1,2, Lynn Maxfield 3, Ingo R Titze 4,5,6
PMCID: PMC6014870  NIHMSID: NIHMS922908  PMID: 29277351

Abstract

Source-filter interaction (SFI) is a phenomenon in which acoustic airway pressures influence the glottal airflow at the source (Level 1) and the vibration pattern of the vocal folds (Level 2). This interaction is most significant when dominant source harmonics are near airway resonances. The influence of acoustic airway pressures on vocal fold vibration (Level 2) was studied systematically by changing the supraglottal vocal tract length in human subjects with tube extensions. The subjects were asked to perform fundamental frequency (fo) glides while phonating through tubes of various lengths. An algorithm was developed using the quasi-open quotient extracted from the electroglottograph (EGG). Regions of sudden vocal fold vibration pattern change due to SFI were inferred from contact area changes. The algorithm correctly identified 89% of male and 84.8% of female quantal changes in contact pattern associated with interactions between source harmonics and formants during ascending glides. During the descending glides, the algorithm correctly identified 84% of male and 81.1% of female quantal changes in contact pattern. These results are in comparison with those obtained from the fo based algorithm (Maxfield et al., 2017).

Keywords: Source-Filter interaction, electroglottography, step detection, quasi-open quotient, voice production

I. INTRODUCTION

Source-filter interaction in voice production is described as the influence of the vocal tract on the glottal source. Voice source analysis (Rothenberg, 1981, 1988; Fant, 1986), excised larynx experiments (Alipour et al., 2001; Smith et al., 2013), physical model experiments (Chan and Titze, 2006; Zhang et al., 2006), human subject experiments with non-singers (Titze et al., 2008; Zanartu et al., 2011) and singers (Titze and Worley, 2009) have all shown evidence of source-filter interaction. As long as the dominant source frequencies lie well below the resonant frequencies of the vocal tract, the filter influences the source mainly in terms of glottal airflow skewing and waveform ripple (Rothenberg, 1981; Fant, 1986). While these interactions with glottal airflow occur in speech across all ages and genders, greater interactions between the source and filter are often found in speech of females and children (Story and Bunton, 2013). The complexity of interactions increases when lower partials of the source frequently cross the resonances; this is most often the case in singing, where the fundamental frequency range spans several octaves (Titze, 2008; Maxfield et al., 2017).

Interaction between the vocal source and the filter occurs at two levels (Titze, 2008). In Level 1 interaction, acoustic vocal tract pressures affect the transglottal pressure, and with it the glottal airflow (Fig. 1, inner loop). At this level, vocal fold vibration can remain relatively undisturbed (Koc and Ciloglu, 2016). In the glottal airflow, however, frequencies creating harmonic distortion are produced that contribute to the source spectrum. This interaction is present in all phonation (speaking, shouting, and singing), male, female, and child. It contributes to the spectral slope and the spectral ripple in the glottal airflow (Rothenberg, 1981; Fant, 1986; Fant and Lin, 1987), even when the spectrum of the glottal area is harmonic and no disruptions of vocal fold vibration occur (Titze, 2008).

Fig. 1.

Fig. 1

Level 1 (inner loop) and level 2 (outer loop) source-filter interaction.

Level 2 interaction, defined as any source interaction with the vocal tract that disturbs the vibration pattern of the vocal folds, is conceptualized in the outer loop in Fig. 1. This interaction occurs primarily when dominant modes of vibration of the vocal folds are affected by intraglottal pressure, which can be profoundly different from transglottal pressure. In a highly simplified approximation, the two pressures can be written as

Ptg=PsPitransglottal pressure (1)
Pig=(Ps+Pi)/2intraglottal pressure (2)

where Ps is the subglottal pressure and Pi is the supraglottal pressure. Note that the supraglottal pressure (Pi) affects the driving pressure for flow (Ptg) in the opposite way it affects the driving pressure for the tissue (Pig). Transglottal pressure decreases with Pi while intraglottal pressure increases with Pi. The supraglottal pressure, Pi can change suddenly with acoustic reactance of the vocal tract near an airway resonance (Lucero et al., 2012; Hatzikirou et al., 2006). In these instances, frequency jumps, subharmonic frequencies, chaos, and other bifurcations or instabilities can occur (Titze, 2008; Giovanni et al., 1999).

Electroglottography (EGG), which is an approximation of the contact area between the vocal folds in vibration, can be used to isolate Level 2 interaction from Level 1 interaction because it doesn’t directly involve glottal airflow. In this article, the changes in the vibratory pattern of the vocal folds due to Level 2 interaction are inferred using changes in the vocal fold contact area in human subjects performing fo glides. The more conventional way to predict the changes in vocal fold vibration is to use the microphone signal and extract an fo contour (Maxfield et al., 2017), but not every Level 2 interaction is necessarily reflected in a sizeable fo change. Two different modes of vibration may have similar frequencies, but their vibration patterns may differ. A more robust way to determine the vocal fold vibration changes is to use the EGG signal. Thus, the motivation for this study was to determine the relative robustness of ΔEGG versus Δfo for Level 2 interaction. The current study was also extended to include both ascending and descending fo glides in comparison to the Maxfield et al., (2017) study which considered only ascending fo glides.

For general speech analysis and synthesis applications, it is important to know which features of the sound source are intended by the speaker (controlled) and which are incidental (uncontrolled) due to interaction. In an analysis-synthesis loop, for example, if fo is extracted and the incidental source-filter variations become part of the programmed fo contour in synthesis, their effect may be doubled, or at least incorrectly represented. Thus, knowing where unplanned source variations are likely to occur is an important part of signal analysis and processing, especially in high fo speech or singing where harmonics with considerable energy pass through vocal tract resonances.

A. Brief Review of former fo based algorithm

In the fo based method of Maxfield et al., (2017), a rate of fo change, f˙o was computed for each glide using the equation

f˙o(n)=fo(n)fo(n1)fo(n)Δt (3)

where Δt is the sampling period in s, and is the sample number. Average rate of fo change, f˙p, was computed by fitting a third order polynomial to f˙o. Significant deviations of f˙o from f˙p at each sample were computed by setting f˙o below a threshold to zero (Fig. 2). The threshold τ was defined as the standard deviation of f˙o from the average, f˙p over the entire glide.

τ=1Nn1N(f˙o(n)f˙p(n))2 (4)

Resonance frequencies and bandwidths of the vocal tract were not measured directly in this study. Instead, formants identified from spectrograms were considered as an approximation of the resonances of the vocal tract such that the source-filter interaction corresponds to the source harmonic-formant interaction. In our notation, fn is the n-th source harmonic frequency and Fn is the n-th resonance (formant) frequency (Titze et al., 2015).

Fig. 2.

Fig. 2

(Color Online) Rate of change of fo versus fo, indicating source-filter interaction regions during an ascending fo glide (Maxfield et al., 2017).

This analysis indicated that approximately 85% of all deviations were associated with at least one of the resonance/harmonic crossing regions (Maxfield et al., 2017) during ascending glides. This result appeared to make fo change a robust predictor of Level 2 interaction. However, there was some evidence that subjects tended to mask (or prevent) fo perturbations to obtain a smooth fo glide (Davis, 1979). Could this fo smoothing be concomitant with a change in the vibration pattern? This became the primary question for the current study: Is fo change by itself as good a predictor of Level 2 interaction as the quasi-open quotient of the EGG signal? If so, the result would be important for future signal processing of speech.

II. METHODS

A. Subject selection and data collection

Subject selection and data collection procedures were the same as those reported in Maxfield et al., (2017). Here, only a brief summary of the procedure is presented. Four males and three females were asked to repeat fo glides, beginning in vocal fry and extending to 500 Hz for males and 700 Hz for females; they then descended back to vocal fry. Between repetitions, the vocal tract was extended in length using 6 tubes of varying length, effectively altering the resonance frequencies and ensuring that several harmonics, including the fundamental, would be forced to pass through at least one resonance region. The tubes were fitted with a mouthpiece to ensure uniform jaw and lip placement and held between the subjects’ teeth with a tight lip seal around the tube throughout the phonation. This arrangement would also limit the variation in resonances during the glide. The distal end of the tubes was attached to a handle, which was also a mount for a microphone (Countryman Isomax B3), maintaining a constant distance of 1.5 cm from the distal end of the tube. Subjects were fitted with an EGG collar (Kay Elemetrics Model 6103). The microphone signal was amplified using an FMC pre-amp (RNP model) before being captured, along with the EGG signal, using a 16-channel analog-digital converter box (ADInstruments, Powerlab 16/30). All signals were recorded using Labchart7 data collection software. Table I lists the approximate first and second formant frequencies estimated for both male and female subjects by artificially elongating the vocal tract with tubes of various lengths. The no-tube condition was deliberately not included because subjects were familiar with the resonances of their own vocal tract and the corresponding interactions. The simple formula for a quarter-wave resonance tube

Fn=(2n1)c4L (5)

was used for resonance estimation (not an exact calculation with end effects), where n is an integer, c is the sound velocity, and L is the length of the tube combined with the length of the vocal tract (estimated at 17 cm for males and 15 cm for females). It can be observed that the tubes are likely to vary the first resonance frequency up to about 150 Hz, and second resonance frequency up to about 450 Hz, although the actual resonance frequencies may vary due to a nonuniform vocal tract shape behind the tube.

TABLE I.

Anticipated F1 and F2 Frequencies Calculated for Artificially Elongated Vocal Tracts

Tube length (cm) 5 7 9 11 15 19
Male F1 (Hz) 398 365 337 313 273 243
F2 (Hz) 1159 1062 981 911 797 708
Female F1 (Hz) 437 398 337 313 291 257
F2 (Hz) 1275 1159 1062 981 850 750

B. Data Analysis

In the current study, both ascending and descending portions of an up-down fo glide were used to study the influence of SFI on vocal fold vibration. Improved estimates of formant frequencies were obtained with linear predictive coding (LPC) analysis of vocal fry, determined for each recording using the PRAAT voice and speech analysis software. Vocal fry was used to view the formant frequencies without the uncertainties caused by widely-spread harmonics at higher frequencies. F1, F2, and F3 were extracted during a 500 ms segment of stable vocal fry, both before and after the up-down glides and were averaged to obtain a single estimate of the formants.

1. Open Quotient and Quasi-Open Quotient

An assumption was made that a specific measure on the EGG waveform, the quasi-open-quotient (QOQ), is sensitive to most changes in the vibration pattern of the vocal folds. The DECOM (DEgg Correlation-based Open Quotient Measurement) method described by Henrich et al. (2004) was used to compute the QOQ. It is important to clarify here that an EGG-based measure of glottal contact area change does not perfectly correlate with a visually-determined measure of glottal closure. Indeed, it has been rightly suggested that the terms “open quotient” and “closed quotient” should be used only when referring to metrics derived from methods with which a definite opening and closing of the glottis can be determined (endoscopic videokymography, for example) (Herbst et al., 2017). Recognizing this disparity while also maintaining accuracy to the DECOM method, the resulting DEGG-determined metric has been labeled QOQ (Hacki, 1996) hereafter. The DECOM method uses the derivative of the electroglottograph (DEGG) signal to determine QOQ. The peaks in the DEGG signal can be related to the glottal opening and closing instants. The duration between the glottal opening event and the consecutive glottal closing event is the open time. The duration between two consecutive glottal closing events is the fundamental period. The QOQ can then be computed as the ratio of open time to the fundamental period. The QOQ was computed for every 10 ms over a window length of 30 ms.

2. Level-2 SFI regions determined from QOQ

The QOQ typically increases in the amount of 0.1 to 0.3 at modal-falsetto transitions during an fo glide (Henrich et al., 2005). In the absence of laryngeal register transitions, QOQ changes only slightly around a mean value of 0.5 in a gradual fo change. However, in the current study, small changes in the QOQ on the order of 0.02–0.1 were also observed to be a result of SFI. These changes in QOQ were attributed to changes in vibratory pattern of the vocal folds due to differences in acoustic pressures as harmonics passed through resonance regions (Titze, 2008). These changes could be visualized as steps in the QOQ contour. Hence, an algorithm based on step detection was developed to determine the SFI regions during an fo glide. The region between successive ascending and descending steps consists of an SFI region (Fig. 3e, red portions). The algorithm was designed based on the assumption that QOQ increases during SFI. There were a few instances where QOQ decreased, especially during the descending glides, but these instances were very few in comparison to those where QOQ increased.

Fig. 3.

Fig. 3

(Color Online) Determining Level 2 SFI regions from QOQ. (a) Raw Quasi-Open Quotient contour signal. (b) Band pass filtered QOQ signal. (c) Median filtered QOQ signal.(d) Steps in the QOQ signal detected by Potts algorithm. Intermediate regions were represented in green. (e) Merging of intermediate regions with the extreme regions. The estimated SFI regions were represented in red in the QOQ step signal.

The QOQ contour extracted using DECOM exhibited both drift and high frequency noise (Fig. 3a). These were removed by passing the QOQ signal through a band pass filter with lower cutoff frequency of 0.1 Hz and higher cutoff frequency of 5 Hz. The resultant filtered QOQ signal can be seen in Fig. 3b. The filtered QOQ signal was then passed through a median filter with cutoff frequency of 9 Hz to enhance the steps in the QOQ signal (Fig. 3c). These cutoff frequencies were chosen based on the strength of the signal. There is not much energy in the signal after 9 Hz as the magnitude is reduced by 40 dB at 9 Hz. A median filter replaces each sample with the median of neighboring samples thus enhancing the steps. Finally, the standard Potts model (Storath et al., 2014) was used for step detection in the median filtered QOQ signal (Fig. 3d, blue and green dots). The Potts model solves the non-convex optimization problem given by

Pγ(x)=argminxNγχ0+Aχbpp (6)

The term

χ0=#{i:χiχi+1} (7)

penalizes the number of steps and the term

Aχbpp=i=1N|Aχibi|p (8)

measures fidelity of steps χ to data b. The parameter γ controls the tradeoff between jump sparsity and data fidelity and was set to 0.1 for male and 0.05 for female data. Matrix A was considered to be an identity matrix in our study and the parameter p was set to 2. The intermediate regions obtained from the Potts model (Fig. 3d, green portions) were merged with one of the adjacent extreme regions (Fig. 3d, blue portions) based on their proximity to those extreme regions (Fig. 3e). However, if the length of the intermediate region was larger than the adjacent lower extreme region, it was merged with the adjacent upper extreme region irrespective of the proximity. The different values of γ for males and females were chosen to reduce the number of these intermediate regions during the analysis. It should be noted that the steps in the QOQ signal start to appear after median filtering itself (Fig. 3c). The additional analysis of step detection using Potts model was performed only to highlight the regions of SFI from non-SFI regions. These parameters were maintained the same for both the ascending and the descending glides.

III. RESULTS

A typical spectrogram of an fo glide of a male subject is shown in Fig. 4. Formants can be seen from the fry portion of the glide. It can be observed that the first three formants interact the most with the first four harmonics at various locations during the glide as fo begins considerably lower than the first formant. However, the second and third formants interacted the most with the first four harmonics in the case of a female subject.

Fig. 4.

Fig. 4

Spectrogram of a male subject performing fo glide with approximate formant locations marked.

A. Ascending glide

Fig. 5 shows an example raw quasi-open quotient signal and its step signal from an ascending glide of a male subject similar to Fig. 3e, along with the source-filter interaction regions determined using spectrogram (vertical bars). The interaction regions determined using the QOQ step detection algorithm were shown in red in the QOQ step signal. An interaction region in time is defined as the duration for a source harmonic to cross a 50 Hz bandwidth of a formant. Bandwidth of 50 Hz was chosen based on the nominal bandwidths of/a/vowel formants whose minimum bandwidth was found to be 43 Hz (Titze et al. 2014). Because of the use of a mouthpiece for formant stability, and the selection of 50 Hz bandwidth, an exact estimation of the formant frequency is not necessary for this study. For male subjects, interactions where 2fo≈F1, fo≈F1, 2fo≈F2, 3fo≈F2, and 4fo≈F2 were of primary interest. These interactions occupied about 65% of the frequency range of a fo glide and 54% of the glide time. As can be observed from Fig. 5, the steps coincided with the beginning and end of the interaction regions. It can also be observed that some of the interactions overlap (ex: fo≈F1 and 3fo≈F2).

Fig. 5.

Fig. 5

(Color Online) (Top) Raw Quasi-Open Quotient waveform. (Bottom) Quasi-Open Quotient step signal along with source-filter interaction regions determined using spectrogram (vertical bars) for a male subject during ascending glide. The SFI regions determined using QOQ step detection algorithm were shown in red in the QOQ step signal.

Across the four male subjects, a total of 280 interaction regions, corresponding to 2fo≈F1, fo≈F1, 2fo≈F2, 3fo≈F2, and 4fo≈F2 interactions, were identified using spectrograms. From the EGG signals, a total of 286 regions were identified using the algorithm described above. Of these regions, 250 (89%) reliably aligned with the source-filter interaction regions. Thus, the algorithm resulted in 11% false negatives (interaction present but not identified by the QOQ algorithm) and 12% false positives (QOQ regions not associated with any interaction region). The number of correctly identified interactions are listed with respect to individual interactions in Table II. The first row shows the percentages for male subject data during an ascending glide. Interactions in the 2fo≈F1 and 3fo≈F2 regions were identified more reliably than the other interactions. The 2fo≈F2 interaction was the least identifiable for males as it generally occurred during the end of the ascending glide.

TABLE II.

Number of Correctly Identified Interactions by QOQ Step Detection Algorithm over Total Interactions for Male and Female Subjects. Values in the Parenthesis Indicate Percentage of Correctly Identified Interactions

Glide Subjects 2fo≈F1 fo≈F1 2fo≈F2 3fo≈F2 4fo≈F2 3fo≈F3 4fo≈F3
Ascending Male 55/56 (98.2) 51/58 (87.9) 44/53 (83.0) 52/57 (91.2) 49/57 (85.9)
Female 44/52 (84.6) 38/51 (74.5) 48/51 (94.1) 27/33 (81.8) 33/38 (86.8) 39/45 (86.6)
Descending Male 53/63 (84.1) 54/64 (84.3) 54/62 (87.1) 50/62 (80.6) 51/62 (82.2)
Female 38/43 (88.3) 40/49 (81.6) 38/47 (80.8) 15/26 (57.7) 30/36 (83.3) 37/43 (86.1)

For female subjects, the important 2fo≈F1 interaction was not well identified because the first resonance frequency with tube extensions was very low (257–437 Hz, Table I). Females quickly transitioned from vocal fry to fo values in their normal speech range (around 200 Hz), which put 2fo near or above 400 Hz, providing few points for 2fo≈F1 analysis. Interactions with the second and third formants were easier to determine due to this higher starting fundamental frequency. Most of the harmonics other than fo were above the first formant. Because of the higher starting fo, all 2fo≈F1 interactions were omitted from the study, leaving fo≈F1, 2fo≈F2, 3fo≈F2, 4fo≈F2 interactions, as well as 3fo≈F3, and 4fo≈F3 interactions, to be considered for female subjects. In the case of one subject, there were no 4fo≈F2 interactions because her fundamental frequency started higher than F2/4.

Figure 6 shows a quasi-open quotient step signal from a female subject overlapped with the source-filter interaction regions determined using spectrogram (vertical bar) during an ascending glide. It can be observed that the female subjects have greater source-filter interaction compared to the male subjects. On an elapsed-time scale, average interaction time for female subjects during a fo glide is about 63%, a 9% increase compared to the male subjects. Across the three female subjects, 270 interaction regions were studied. The QOQ step detection algorithm identified 249 regions. Among them, a total of 229 regions were identified reliably with the interaction regions leading to 84.8% of correct estimations, 15.2% false negatives, and 8% false positives. The second row of Table II provides the percentages of correctly identified interactions by the QOQ algorithm for the female subjects during ascending glides. It can be observed that the two extreme interactions (i.e. 4fo≈F2 and 2fo≈F2) were the hardest to identify because the interactions occurred either shortly after the onset of the glide or near the end of the glide and were weak.

Fig. 6.

Fig. 6

(Color Online) (Top) Raw Quasi-Open Quotient waveform. (Bottom) Quasi-Open Quotient step signal along with source-filter interaction regions determined using spectrogram (vertical bars) for a female subject during ascending glide. SFI regions determined using QOQ step detection algorithm were represented in red in the QOQ step signal.

B. Descending glides

We report results from descending glides separately because, on visual inspection, they were not produced as consistently as the ascending glides by all the subjects. Also, Level-2 SFI regions during descending glides were determined from both fo and QOQ for completeness with the former Maxfield et al., (2017) study. The procedure to determine the level-2 SFI regions from fo during descending glides was the same to the method used for ascending glides (Maxfield et al., 2017). Fig. 7 shows the rate of change of fo versus fo for a female, indicating source-filter interaction regions during a descending fo glide. The descending glides yielded 584 fo instabilities across the 4 males and 3 females. Of these 584 instabilities, 508 (87%) occurred at frequencies at which one of the first four harmonics was crossing one of the first three formants. fo instabilities from male subjects associated with 88% and female subjects with 86% of crossings. These percentages are similar to the male (88%) and female (80%) subjects reported for the ascending glides and again significantly higher than chance.

Fig 7.

Fig 7

(Color Online) Rate of change of fo versus fo, indicating source-filter interaction regions for a female subject during a descending fo glide.

Fig. 8 shows a QOQ step signal from a male subject overlapped with the source-filter interaction regions. It can be observed that the steps during descending glide were similar to those during the ascending glide. The algorithm was able to identify the start and end of the interaction regions during the descending glide. Across the four male subjects, a total of 313 interaction regions were identified using spectrograms, corresponding to 2fo≈F1, fo≈F1, 2fo≈F2, 3fo≈F2, and 4fo≈F2 interactions. The algorithm described above was able to reliably align with 262 (84%) source-filter interaction regions. Thus, the algorithm resulted in 16% false negatives. The number of correctly identified interactions for descending glides were also listed with respect to individual interactions in Table II. The third row shows the percentages for male subject data. All the interaction regions were estimated with similar reliability unlike the ascending glides where 2fo≈F1 and 3fo≈F2 regions were identified more accurately compared to the other regions.

Fig. 8.

Fig. 8

(Color Online) (Top) Raw Quasi-Open Quotient waveform. (Bottom) Quasi-Open Quotient step signal along with source-filter interaction regions determined using spectrogram (vertical bars) for a male subject during descending glide. SFI regions determined using QOQ step detection algorithm were represented in red in the QOQ step signal.

For female subjects, the 2fo≈F1 interaction was not well identified, even for the descending glides. The interaction regions fo≈F1, 2fo≈F2, 3fo≈F2, 4fo≈F2, 3fo≈F3, and 4fo≈F3 were considered similar to the ascending glides. Across the three female subjects, 244 interaction regions were studied. Among them, a total of 198 regions were reliably identified with the interaction regions determined using spectrogram leading to 81.1% of correct estimations, and 18.9% false negatives. The fourth row of Table II provides the percentages of correctly identified interactions during descending glide for the female subjects. It can be observed that the 4fo≈F2 interaction was the least identified among the other interactions as it occurred mostly during the end of the glide. The other interactions were identified with greater accuracy, similar to the ascending glides.

IV. DISCUSSION AND CONCLUSIONS

Results from the current study demonstrate that electroglottographic signals can be used to identify quantal changes in vocal fold contact, and therewith presumably changes in the vibratory patterns of the vocal folds, resulting from source-filter interaction. Treating the changes in the EGG signal as quantal steps allowed prediction of the start and end of the interaction regions. The finding of earlier studies that source-filter interaction is stronger in females (Bozeman, 2013) was also confirmed. It was further validated that females encounter more interaction regions (six versus five) compared to males due to their wider fo ranges. It was also found that higher order harmonic-resonance crossings may have a significant impact on fo stability. However, the current study was performed only on normal healthy subjects and thus doesn’t provide any insight into how the results change in the presence of voice disorders. In particular, when the modes of vibration of the left vocal fold are not synchronized with those of the right vocal fold, more interaction may occur.

In order to robustly predict quantal changes in contact pattern of the vocal folds from an EGG signal, the fidelity of the EGG signal is important. Obtaining an EGG signal during fo glides was a challenge as the subjects often tended to raise their larynx during the glide and the signal became weak. We tried to instruct the subjects to consciously be aware of the position of their larynx and to perform repetitions until a strong EGG signal was obtained. Beyond recording fidelity, the next step in accurately predicting quantal changes in contact pattern is obtaining a smooth quasi-open quotient signal from the EGG without many discontinuities. In this study, we used the DECOM method given in (Henrich et al., 2004) which resulted in a smooth quasi-open quotient signal in most of the cases. When the QOQ signal appeared to be inaccurate, such test cases were removed from further analysis. With a good QOQ signal, the QOQ step detection algorithm was able to estimate the regions of quantal change in the contact pattern in the vocal folds. Predicting the vibratory mode details may also be possible with the QOQ step detection algorithm. This approach was not studied here, but should be included in a future study.

Maxfield et al., (2017) and the current study both used data from the same subjects and the percentages of correctly identified interaction regions were slightly higher from the QOQ step detection algorithm for ascending glides (89% versus 88% for males and 84.8% versus 80%) than for descending glides (84% versus 88% for males and 81.1% versus 86% for females). Overall, the difference in results between the two algorithms is not significant. Hence, it can be concluded that both fo based and QOQ based algorithms perform similarly in estimating the regions of source-filter interaction. We also tried to develop an algorithm based on the slope of EGG signal during opening and closing instances to identify SFI regions. The algorithm turned out to be as complex as the QOQ based algorithm and the results were also not much different. There are both advantages and disadvantages with fo based and QOQ based algorithms. The QOQ based algorithm is more complex compared to the fo based algorithm in terms of computation and analysis. On the other hand, accurately determining fo in each time window using automated algorithms is difficult. The results would change significantly if fo were to be determined incorrectly even in one or two time windows. In this and the previous Maxfield et al., (2017) studies, we used semi-automated algorithms where we manually corrected the incorrect fo estimations after running through automated algorithms. However, the results from QOQ based algorithm were not that strongly dependent on few errors in QOQ estimation as filtering was used.

The regions of quantal changes in vocal fold contact due to source-filter interaction were more reliably identified during the ascending glides as compared to the descending glides (89% versus 84% for males, and 84.8% versus 81.1% for females) using the QOQ based method. The major reason for this degraded performance was the reduced energy in the descending glide signals compared to the ascending glide signals. Subjects may have depleted their usable lung pressure at the end of the descending glide, potentially creating weak signals. It became hard to identify the SFI regions as well as estimate the QOQ accurately during descending glides compared to the ascending glides. It was possible to identify fo accurately especially using semi-automated techniques even from weak signals. Hence, the results from fo based method were more comparable between the ascending and descending glides.

This study could have been done using different tube lengths for males than for females, thus spanning a wider and more appropriate range of F1–F2 for both genders. In particular, a tube shorter than 5 cm could have been used for females to increase F1. This use of same tube lengths for both sexes may have contributed to the discrepancy in results between males and females. The fixed mouthpiece gave good formant control and less opportunity for the subject to predict where interactions would take place compared to if natural vowels had been selected. A future study might address the adaptation or compensation achievable when these interactions are predictable and unwanted. Future studies might also address the perceptual significance of voice quality due to sudden changes in source-filter interaction (Samlan and Kreiman, 2014).

Acknowledgments

This work was supported by National Institute on Deafness and Other Communication Disorders Grant R01 DC012045, awarded to Ingo R. Titze, Principal Investigator.

Footnotes

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final citable form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

Contributor Information

Anil Palaparthi, National Center for Voice and Speech, The University of Utah, Salt Lake City, UT, 84101 USA; Department of Bioengineering, The University of Utah, Salt Lake City, UT, 84112 USA.

Lynn Maxfield, National Center for Voice and Speech, The University of Utah, Salt Lake City, UT, 84101 USA.

Ingo R. Titze, National Center for Voice and Speech, The University of Utah, Salt Lake City, UT, 84101 USA Department of Bioengineering, The University of Utah, Salt Lake City, UT, 84112 USA; Department of Communication Sciences and Disorders, The University of Iowa, Iowa City, IA, 52242 USA.

References

  1. Alipour F, Montequin D, Tayama N. Aerodynamic profiles of a hemilarynx with a vocal tract. Ann Otol Rhinol Laryngol. 2001;110:550–555. doi: 10.1177/000348940111000609. [DOI] [PubMed] [Google Scholar]
  2. Bozeman KW. Practical Vocal Acoustics-Pedagogic Applications for Teachers and Singers. Pendragon: Press Hilsdale; 2013. pp. 32–35. (Vox Musicae: The Voice, Vocal Pedagogy, and Song, No. 9). [Google Scholar]
  3. Chan RW, Titze IR. Dependence of phonation threshold pressure on vocal tract acoustics and vocal fold tissue mechanics. J Acoust Soc Am. 2006;119:2351–2362. doi: 10.1121/1.2173516. [DOI] [PubMed] [Google Scholar]
  4. Davis S. Acoustic characteristics of normal and pathological voices. In: Lass NK, editor. Speech and Language: Advances in Basic Research and Practice. Vol. 1. Academic Press; New York: 1979. pp. 271–335. [Google Scholar]
  5. Fant G. Glottal flow: Models and interaction. J Phonetics. 1986;14:393–399. [Google Scholar]
  6. Fant G, Lin Q. STL Q Prog Status Rep. Royal Institute of Technology; Stockholm: 1987. Glottal voice source-vocal tract acoustic interaction; pp. 1–13. [Google Scholar]
  7. Giovanni A, Ouaknine M, Guelfucci B, et al. Nonlinear behavior of vocal fold vibration. J Voice. 1999;27:261–266. doi: 10.1016/s0892-1997(99)80002-2. [DOI] [PubMed] [Google Scholar]
  8. Hacki T. Electroglottographic quasi-open quotient and amplitude in crescendo phonation. J Voice. 1996;10:342–347. doi: 10.1016/s0892-1997(96)80025-7. [DOI] [PubMed] [Google Scholar]
  9. Hatzikirou H, Fitch W, Herzer H. Voice instabilities due to source-tract interaction. Acta Acust United Acust. 2006;92:468–475. [Google Scholar]
  10. Herbst CT, Schutte HK, Bowling DL, Svec JG. Comparing chalk with cheese-The EGG contact quotient is only a limited surrogate of the closed quotient. 2017;31(4):401–409. doi: 10.1016/j.jvoice.2016.11.007. [DOI] [PubMed] [Google Scholar]
  11. Henrich N, D’Alessandro C, Doval B, Castellengo M. On the use of the derivative of electroglottographic signals for characterization of nonpathological phonation. J Acoust Soc Am. 2004;115(3):1321–1332. doi: 10.1121/1.1646401. [DOI] [PubMed] [Google Scholar]
  12. Henrich N, D’Alessandro C, Doval B, Castellengo M. Glottal open quotient in singing: Measurements and correlation with laryngeal mechanisms, vocal intensity, and fundamental frequency. J Acous Soc Am. 2005;117(3):1417–1430. doi: 10.1121/1.1850031. [DOI] [PubMed] [Google Scholar]
  13. Koc T, Ciloglu T. Nonlinear interactive source-filter models for speech. Computer Speech and Language. 2016;36:365–394. [Google Scholar]
  14. Lucero JC, Louremco C, Hermant N, Hirtum AV, Pelorson X. Effect of source-tract acoustic coupling on the oscillation onset of the vocal folds. J Acoust Soc Am. 2012;132(1):403–411. doi: 10.1121/1.4728170. [DOI] [PubMed] [Google Scholar]
  15. Maxfield LM, Palaparthi A, Titze IR. New evidence that nonlinear source-filter coupling affects harmonic intensity and fo stability during instances of harmonics crossing formants. J Voice. 2017;31(2):149–156. doi: 10.1016/j.jvoice.2016.04.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Rothenberg M. Acoustic interaction between the glottal source and the vocal tract. In: Stevens KN, Hirano M, editors. Vocal Fold Physiology. University of Tokyo Press; Tokyo: 1981. pp. 305–328. [Google Scholar]
  17. Rothenberg M. Acoustic reinforcement of vocal fold vibratory behavior in singling. In: Fujimura O, editor. Vocal Physiology: Voice Production, Mechanisms and Functions. Raven Press; New York: 1988. pp. 379–389. [Google Scholar]
  18. Samlan RA, Kreiman J. Perceptual consequences of changes in epilaryngeal area and shape. J Acoust Soc Am. 2014;136:2798–2806. doi: 10.1121/1.4896459. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Smith BL, Nemcek SP, Swinarski KA, Jiang JJ. Nonlinear source-filter coupling due to the addition of a simplified vocal tract model for excised larynx experiments. J Voice. 2013;27(3):261–266. doi: 10.1016/j.jvoice.2012.12.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Storath M, Weinmann A, Demaret L. Jump-Sparse and sparse recovery using Potts functionals. IEEE Transactions on Signal Processing. 2014;62:3654–3666. [Google Scholar]
  21. Story BH, Bunton K. Production of child-like vowels with nonlinear interaction of glottal flow and vocal tract resonance. Proc Mtgs Acoust. 2013;19:060303. [Google Scholar]
  22. Titze IR. Nonlinear source-filter coupling in phonation: Theory. J Acoust Soc Am. 2008;123(5):2733–2749. doi: 10.1121/1.2832337. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Titze IR, Baken RJ, Bozeman KW, et al. Toward a consensus on symbolic notation of harmonics, resonances, and formants in vocalization. J Acoust Soc Am. 2015;137(5):3005–3007. doi: 10.1121/1.4919349. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Titze IR, Riede T, Popolo P. Nonlinear source-filter coupling in phonation: Vocal exercises. J Acoust Soc Am. 2008;123(4):1902–1915. doi: 10.1121/1.2832339. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Titze IR, Worley AS. Modeling source-filter interaction in belting and high-pitched operatic male singing. J Acoust Soc Am. 2009;126(3):1530–1540. doi: 10.1121/1.3160296. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Titze IR, Palaparthi A, Smith SL. Benchmarks for time-domain simulation of sound propagation in soft-walled airways: Steady configurations. J Acoust Soc Am. 2014;136(6):3249–3261. doi: 10.1121/1.4900563. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Zanartu M, Mehta DD, Ho JC, Wodicka GR, Hillman RE. Observation and analysis of in vivo vocal fold tissue instabilities produced by nonlinear source-filter coupling: A case study. J Acoust Soc Am. 2011;129(1):326–339. doi: 10.1121/1.3514536. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Zhang Z, Neubauer J, Berry DA. The influence of subglottal acoustics on laboratory models of phonation. J Acoust Soc Am. 2006;120:1558–1569. doi: 10.1121/1.2225682. [DOI] [PubMed] [Google Scholar]

RESOURCES