Abstract
Most cues to speech intelligibility are within a narrow frequency range, with its upper limit not exceeding 4 kHz. It is still unclear whether speaker-related (indexical) information is available past this limit or how speaker characteristics are distributed at frequencies within and outside the intelligibility range. Using low-pass and high-pass filtering, we examined the perceptual salience of dialect and gender cues in both intelligible and unintelligible speech. Setting the upper frequency limit at 11 kHz, spontaneously produced unique utterances (n = 400) from 40 speakers were high-pass filtered with frequency cutoffs from 0.7 to 5.56 kHz and presented to listeners for dialect and gender identification and intelligibility evaluation. The same material and experimental procedures were used to probe perception of low-pass filtered and unmodified speech with cutoffs from 0.5 to 1.1 kHz. Applying statistical signal detection theory analyses, we found that cues to gender were well preserved at low and high frequencies and did not depend on intelligibility, and the redundancy of gender cues at higher frequencies reduced response bias. Cues to dialect were relatively strong at low and high frequencies; however, most were in intelligible speech, modulated by a differential intelligibility advantage of male and female speakers at low and high frequencies.
I. INTRODUCTION
Low-pass (LP) and high-pass (HP) filtering has been utilized in speech perception research since the early work of Harvey Fletcher and his collaborators at Bell Telephone Laboratories in the 1920s, whose objective was to evaluate speech quality in the telephone network. LP-filtered speech retains information in lower frequencies, and higher frequencies are eliminated above a selected cutoff. Conversely, HP-filtered speech retains high-frequency information, and the spectral content below a selected frequency cutoff is removed.
In a series of early studies [see Allen (1996) for a review of Fletcher's numerous contributions to the development of the speech intelligibility index], it was shown that identification of nonsense syllables consisting of various combinations of vowels and consonants was very high at LP filter cutoff frequency of 3 kHz or greater or HP filter cutoff at 1 kHz or lower. This is because LP-filtered speech at 3 kHz provides most of the frequency information needed for identification of vowels and consonants (e.g., Kewley-Port et al., 1983; Lehiste and Peterson, 1959), and eliminating information below 1 kHz in HP filtering has little effect on the perception of consonants that supply about 90% of the acoustic information important for speech comprehension (Fletcher and Galt, 1950; French and Steinberg, 1947; Licklider and Miller, 1951). For nonsense syllables, the effects of LP and HP filtering were similar for frequency cutoff of 1.9 kHz, resulting in an equal amount of intelligibility (68%) for each filtering type (French and Steinberg, 1947). Together, LP and HP filtering studies have established that the acoustic energy between 250 Hz and 4 kHz is most important to speech intelligibility, which led to the wide acceptance of the “telephone bandwidth” in analog telephony that, limited to 300–3400 Hz, did not considerably reduce intelligibility when compared with full-spectrum speech.
As a result, the high-frequency region above 4 kHz has not attracted much attention in speech research as most studies focused on defining efficiency in the processing of intelligible speech, and it was assumed that information in higher frequencies made only minor contributions to comprehension. Specifically, the upper frequency limit with respect to its maximum possible contribution to intelligibility was estimated to be 7 kHz (French and Steinberg, 1947). From the perspective of human hearing, there is a physiological basis for this intelligibility limit related to the morphology and function of human cochlea. The filtering properties of the cochlea are nonlinear, and auditory filters broaden with increasing frequency, which indicates that high-frequency filters pass a wider range of frequencies than low-frequency filters. French and Steinberg's model showed that the wide frequency band from 5.7 to 7 kHz still contained some cues to intelligibility, but its independent contribution (5%) was equal to that of each lower filter. Nevertheless, the upper 7-kHz limit is more effective considering the perception of fricatives as it can disambiguate words such as “sick” and “thick,” and modern technologies using high-definition audio—including cellular telephony—have adopted the frequency band extended to 7 kHz.
However, while focusing on intelligibility of filtered speech, research has not yet seriously considered the utility of those speech cues that, distributed across speech spectrum with various strengths, do not directly contribute to comprehension. Over the last few decades, speech perception studies have increasingly recognized the role of speaker characteristics that are interwoven into the acoustic signal and are automatically processed in listening to natural speech. These “indexical” characteristics include a range of social, physical, and psychological information, such as social status, gender, age, personality, linguistic background, regional dialect, foreign accent, race, etc., which is conveyed to listeners through a speaker's voice (see the following volumes for reviews, perspectives, and discussions: Labov, 2010; Kreiman and Sidtis, 2011; Pardo et al., 2022).
The goal of the current investigation is to gain a better understanding of spectral distribution of socio-indexical cues that provide information about a speaker's identity, considering both the low- and high-frequency ranges. Our indexical speaker variables include regional dialect (situating speakers in their geographically defined regional sound patterns) and gender (cueing biological sex, either male or female). Using LP and HP filtering, we examine the perceptual salience of these two types of cues in relation to the intelligibility of filtered speech. We seek to determine whether these indexical cues are equally strong in LP- and HP-filtered speech when frequency information important for speech comprehension is reduced or eliminated. Our research questions and predictions are elucidated after considering the known effects of LP and HP filtering on the perception of indexical cues in speech.
A. Indexical information in LP-filtered speech
LP filtering has been used primarily in the context of speech intelligibility. However, its potential to study those aspects of speech and speaker characteristics that are conveyed by variations in fundamental frequency (F0) and lower formant frequencies has prompted numerous investigations into prosody, lexical tones, vocal characteristics of race, vocal emotions, and emotional disorders (e.g., Kitayama and Ishii, 2002; McNally et al., 2001; Scherer, 2003). For example, LP filtering at various cutoff frequencies was used in research on affectual infant-directed speech (Burnham et al., 2002; Kitamura and Burnham, 2003; Knoll et al., 2009) and on perception of race (Lass et al., 1980; Thomas and Reaser, 2004).
Only a handful of studies have used LP filtering to investigate the contribution of prosodic cues to identification of regional dialects or different varieties of the same language spoken in geographically distant locations (e.g., Portuguese spoken in Portugal and Brazil). Traditionally, the research examining regional variation in pronunciation patterns has focused on segments (i.e., vowels and consonants) as the primary sources of variability, and prosodic cues (i.e., intonation, rhythm, and perceptible features such as pitch range, pitch level, or speaking rate) have been considered secondary (e.g., Labov et al., 2006). Since prosodic information is most prominently carried out by F0 variation, LP filtering has been a logical choice to “delexicalize” speech by removing most of the spectral content while preserving intonation and the temporal organization of syllables. Frequency cutoffs between 300 and 400 Hz have been the primary choice in experiments because voice F0 does not typically rise higher (both in male and female speech; e.g., Pépiot, 2014), and retained rhythm information is robust, as evidenced in language discrimination tasks with newborn infants (e.g., Nazzi et al., 1998).
Indeed, the frequency band up to 400 Hz contained sufficient information for perceiving rhythmic differences between Brazilian Portuguese and European Portuguese (Frota et al., 2002). In another study, Van Bezooijen and Gooskens (1999) found that several Dutch and English dialects could be distinguished above chance by prosodic features in utterances LP-filtered at 350 Hz. However, both unaltered and monotonized speech (removing intonation but retaining verbal information) provided many more dialect cues, suggesting that contribution of prosodic features to dialect identification is comparatively weaker. The role of prosodic cues (intonation) was also evaluated in a study investigating prosodic differences between the closely related dialects of Orkney and Shetland in Scotland (Van Leyden and Van Heuven, 2006). Listeners' judgments were well above chance when speech was LP-filtered at 300 Hz, indicating that prosodic information differentiating dialects is available in unintelligible speech. In yet another study, perceivable differences between several German dialects were found—with various strengths—when utterances were bandpass-filtered from 70 to 270 Hz (Schaeffler and Summers, 1999). Together, investigations using LP filtering suggest that dialect information is available in unintelligible speech; however, prosodic features alone provide listeners with only a subset of cues necessary to successfully differentiate dialects.
Gender of an adult speaker (here, defined as biological sex and not as a social or cultural variable), reflecting physical differences between males and females, is one of the most salient indexical features conveyed by human voice. On average, the vocal tract of an adult female is 15% shorter than of an adult male (Fitch and Giedd, 1999), and physiological differences in the pharynx and the larynx give rise to differences in F0, perceived as lower pitch in males and higher pitch in females. Differences in vocal tract length affect the resonances of the vocal tract such that formant frequencies in vowels are about 20% higher in females than in males (Hillenbrand et al., 1995). Of course, less extreme differences between male and female voices can create perceptual ambiguities, and listeners may use additional cues (e.g., pitch dynamics, intonation, speaking rate) to guide their identification decisions.
Using LP filtering, Lass et al. (1976) reported high gender identification accuracy (91%) from isolated vowels LP-filtered at 255 Hz, underscoring the salience of F0 information in speaker gender recognition. In another study, gender identification was near ceiling for sentences LP-filtered at 255 Hz, which supports the importance of F0 (without contribution of vocal tract resonances) in making gender identification decisions (Lass et al., 1980). From this work, it can be concluded that F0 alone is the main predictor of the perceived gender differences, and full-spectrum speech may provide additional cues to disambiguate male and female voices when F0 falls within a gender ambiguous range.
B. Speech and speaker information in the extended high-frequency (EHF) spectrum
HP filtering has not been used in speech perception research as often as LP filtering, primarily because frequencies beyond 7 kHz have not been considered important for speech intelligibility. However, recent research in human auditory perception has increasingly recognized that EHF hearing above the intelligibility limit has ecological utility [review in Hunter et al. (2020)]. Although the upper limit of human hearing is around 20 kHz, still little is known how EHF hearing may support speech perception because the primary investigations of human hearing sensitivity (and audiological assessment of hearing) have been limited to pure tone sensitivity up to 8 kHz.
The utility of speech information in the EHF range was manifested in perception experiments that minimized the contribution of lower frequencies. For example, using bandstop filtering (removing the mid-frequency band), listeners successfully combined information from the high- and low-frequency speech spectrum in the perception of nonsense syllables (Lippmann, 1996). In another study, vowel and consonant categorization was above chance when syllables were HP-filtered above 5.7 kHz and lower-frequency energy was masked by noise (Vitela et al., 2015). It was also shown that information in EHF spectrum contributes to speech recognition in the presence of other competing talkers (Monson et al., 2019) or in noise (Deshpande and Holambe, 2011; Motlagh Zadeh et al., 2019). This occurs because, typically, signal degradations adversely affect the lower- rather than higher-frequency cues, forcing the listeners to utilize EHF hearing to compensate for missing information in the speech signal.
Acoustic information available at higher frequencies is most evident in high-frequency sounds, including fricatives. Explorations and modeling of the acoustic characteristics of fricatives have increased over the last three decades, and new measures have been proposed largely because speech could be routinely recorded at higher sampling rates, 44.1 kHz or even 48 kHz. This opened up the possibility of considering higher-frequency regions in quantifying fricative spectra, which led to defining new parameters specifically for use with sampling rates of 32 kHz or higher (Shadle, 2023). One of the most reliable acoustic measures that differentiates fricatives in speech production is the location of their spectral energy peaks, typically measured as the highest-amplitude peak of the fast Fourier transform (FFT) spectrum. Spectral peaks are higher in women than in men (Jongman et al., 2000; Tabain, 1998); however, considering the upper-shoulder frequencies of the peak, their means (measured at points 10 dB down from the peak) for female speakers can be very high. For /s/, the mean upper-shoulder frequencies were found to range from 12.2 to 15.4 kHz or even 18 kHz, thus, approaching the upper limit of human hearing (Alexander, 2019).
Little is known regarding what speaker characteristics are available in the high-frequency range. Using HP filtering, several perception studies have shown that listeners can identify speaker gender (male, female) when low-frequency spectral energy is removed. Gender was identified with 82% accuracy from 250-ms segments of naturally produced vowels HP-filtered at 3.5 kHz (Donai and Lass, 2015) and significantly above chance from vowels and sentences HP-filtered up to 8.5 and 12 kHz, respectively (Donai and Halbritter, 2017). In another study, Monson et al. (2014) reported high accuracy (92%) for gender identification from phrases HP-filtered at 5.7 kHz. Together, these experiments demonstrated that cues to speaker gender can be strong in EHF spectrum and that the presence of low-frequency spectral energy in the speech signal is not necessary for accurate gender identification.
Whether dialect-related cues can be found in the high-frequency region is still unknown. Evidence exists that, in unmodified speech, perceptual cues to dialect variation are present not only in vowels and consonants but also in prosodic aspects of speech, such as rhythm, intonation, lexical stress, pitch range, speaking rate, and the use of pauses (e.g., Clopper and Smiljanic, 2011, 2015; Jacewicz et al., 2010). However, there is no experimental evidence that these cues are perceptually salient at extended high frequencies. Using HP filtering, our empirical study will contribute new findings regarding the presence and strength of perceptual cues to dialect identification in high-frequency ranges.
C. The current study
The current study seeks to determine the availability and strength of indexical cues to a speaker's regional dialect and gender in low- and high-frequency spectral energy when frequency information important for speech comprehension is reduced or unavailable. Studies using LP filtering typically assessed the importance of suprasegmental (prosodic) cues for identification of regional and foreign accents by comparing the effects of a single LP filter cutoff frequency with the original (unmodified) speech (e.g., Van Bezooijen and Gooskens, 1999; Frota et al., 2002; Kolly et al., 2014; Schaeffler and Summers, 1999; Van Leyden and Van Heuven, 2006). A consistent finding was that a LP filter cutoff at or below 400 Hz (350 or 300 Hz) was sufficient to elicit listeners' identification performance above chance level. Speaker gender can be accurately identified with near-ceiling accuracy at a 255-Hz cutoff (Lass et al., 1976; Lass et al., 1980). Based on this evidence, our first research question is
-
(1)
Does a progressively higher LP filter cutoff, providing increasingly more intelligibility cues, supply increasingly more indexical information about a speaker's regional dialect and gender?
Addressing this question, we used LP filters with cutoff frequencies of 500, 700, 900, and 1100 Hz (experiment 1), predicting that listeners' dialect identification scores will significantly improve with each higher filter level as more verbal information becomes available. Based on the reviewed literature, we expected an above-chance performance at LP-500 Hz; however, whether the 1.1 kHz cutoff could provide sufficient verbal content to yield dialect recognition scores approaching unmodified speech was less clear. Given that a LP-1.2 kHz cutoff has been considered a bridge between unintelligible LP-filtered and unmodified speech for female speakers (Knoll et al., 2009), it can be expected that dialect identification at LP-1.1 kHz will not yet reach the levels of unmodified speech. Further, more dialect cues will be available in male than female speech because lower vowel formants in males will increase segmental (and semantic) content, giving rise to a comparatively greater intelligibility of male speech. While this outcome is predictable based on the speech production mechanism, we point out that the current study is the first to assess statistically the anticipated effects of the gender-related differences in LP-filtered speech on dialect identification because previous investigations either used male speakers only (Van Bezooijen and Gooskens, 1999; Van Leyden and Van Heuven, 2006) or female speakers only (Frota et al., 2002) or used both male and female speakers but did not inquire into the differences between the two (Kolly et al., 2014; Schaeffler and Summers, 1999).
Although gender identification scores may not significantly improve with each higher LP filter level, we nevertheless sought to verify findings of the earlier studies that were not without limitations. Specifically, gender was identified from isolated vowels in Lass et al. (1976) and from four experimental sentences in Lass et al. (1980), produced by 20 speakers (ten male) in either study. High accuracy (in percentage correct) could be related to the limited number of both stimuli and speakers, and it is unclear whether scores near ceiling can be obtained when listeners are presented with more extensive speaker and item variability. In the current study, we doubled the number of speakers (n = 40; 20 male) and increased the number of stimuli (n = 400; ten unique phrases from each speaker).
We used the signal detection theory (SDT) framework (Green and Swets, 1966; Macmillan and Creelman, 2005) to analyze both gender and dialect identification scores. Unlike accuracy in percentage correct, SDT is a preferred statistical approach in analyzing listener behavior under various degrees of stimulus uncertainty because it allows for an independent account of perceptual sensitivity to stimulus characteristics (i.e., the selection process and the associated processing efficiency that increases with stimulus strength) and response bias (i.e., decision threshold or criterion). The second aspect of SDT, the decision behavior, has been often ignored in speech perception research, and interpretation of response bias still allows for a fair amount of flexibility [see Gorea and Sagi (2005) for a discussion of decision behavior]. In essence, the decision criterion reflects the strategy underlying response of the listeners in terms of their willingness to say Male or Female (in the context of the current study). Studying the decision processes (i.e., response bias) associated with perceptual processes (i.e., perceptual sensitivity) is particularly important when listeners cope with extensive stimulus variability in multiple experimental conditions. Attending to multiple speakers and items may encourage nonoptimal behavior (i.e., increase response bias), such as the tendency to become more cautious (i.e., more conservative) over the course of a session (See et al., 1997). It is possible that listeners can become more uncertain and cautiously biased toward responding Male or Female at selected LP filter levels, which experiment 1 seeks to determine.
Our second research question is analogous to the first.
-
(2)
Does a progressively higher HP filter cutoff, providing increasingly fewer intelligibility cues, supply increasingly less indexical information about a speaker's regional dialect and gender?
In experiment 2, we presented the same speech material to a different group of listeners, and the utterances were HP-filtered at 700, 1175, 1973, 3312, and 5560 Hz (the upper limit of the high-frequency region was 11 kHz). Our selection of HP filter cutoffs was informed by both intelligibility and gender identification studies that used HP filtering, and our predictions for dialect and gender identification in relation to intelligibility are based on the following evidence. Using intelligibility data from French and Steinberg (1947) as a general guideline, HP-700 Hz cutoff was expected to yield high intelligibility scores (near ceiling). Each higher-frequency cutoff was determined logarithmically, by adding the log step size of 0.225 [log10(0.225)] to the log of each frequency beginning with 700 Hz. In this way, HP-1175 Hz corresponded to a cutoff of about 1.2 kHz in French and Steinberg, which was expected to yield slightly lower intelligibility scores, about 90%. The expectations for each higher cutoff were as follows: HP-1973 Hz corresponded to 1.9 kHz (intelligibility of about 70%), HP-3312 Hz corresponded to about 3 kHz (intelligibility of 30% or lower), and HP-5560 Hz corresponded to 5 kHz (intelligibility of 0%).
Although French and Steinberg's intelligibility results are widely accepted, it is important to remember that these data are based on the recognition of nonsense syllables averaged across male and female speakers. We therefore expected some discrepancies between our intelligibility results and theirs, particularly as they relate to male and female productions. Presumably, gender-based differences may be best reflected at HP-1175 Hz and HP-1973 Hz, with female speech yielding higher intelligibility scores than male speech. This is because these two cutoffs will negatively affect identification of consonants in males, and a reduction in consonant information is expected to decrease speech intelligibility (e.g., Licklider and Miller, 1951). As summarized by Kent and Read (1992, p. 157), “as a rule of thumb, the frequency value for a particular acoustic feature will be on the order of 20% higher for a woman than for a man,” and these two cutoffs will eliminate some of the lower vowel formant transitions (contributing to the perception of stops) in men, along with the low spectral regions important for the perception of nasal and liquid consonants.
Given that most phonological cues to dialects are in intelligible speech, we can predict a linear relationship between dialect identification and intelligibility so that each higher HP filter cutoff will yield increasingly lower dialect identification scores, modulated by speaker gender as explicated above. However, if prosodic cues to dialects are present at higher frequencies, listeners will still be able to distinguish between the dialects—at least above chance—even if intelligibility cues at the two highest cutoffs were severely reduced (at HP-3312 Hz) or no longer available (at HP-5560 Hz).
Based on the recent evidence about the availability of perceptual cues to speaker gender at high frequencies (Donai and Halbritter, 2017; Donai and Lass, 2015; Monson et al., 2014), we predicted that gender identification would remain high at either HP-3312 Hz (Donai and Lass, 2015) or HP-5560 Hz (Monson et al., 2014). However, consistent with experiment 1, we sought to verify the strength of gender cues using a larger speech sample and an SDT analytical approach.
II. METHODS
A. Listeners
1. Experiment 1: LP filtering
Twenty participants (ten female, ten male) ranging in age from 19 to 24 years [mean (M) = 21.9; standard deviation = ±2.1 years] served as listeners in experiment 1. In terms of gender identity, all participants self-identified as cisgender (i.e., conforming to their biological sex at birth). All participants were undergraduate students at The Ohio State University, were native speakers of American English, and had lived in Columbus, Ohio (OH), at least 4 years prior to the study. Nineteen participants grew up in central OH and spoke the Midland variety of American English typical of the region (which includes Columbus). One participant was from Oregon but has lived in central OH for 5 years. All participants reported normal hearing and no history of speech-language disorders. All gave written informed consent to participate and completed a demographic and language background questionnaire. The participants were asked to participate in two separate listening tasks in two sessions spaced 2 weeks apart. They were compensated with a monetary fee ($15 per session) for their time. The study was approved by the Institutional Review Board for human research at The Ohio State University.
2. Experiment 2: HP filtering
Twenty-two participants (16 female, 6 male; cisgender) between the ages of 19 and 24 years (21.2 ± 1.1) served as listeners in experiment 2. None of the participants in experiment 2 participated in experiment 1. Nineteen participants were undergraduate students at The Ohio State University, and three participants were not associated with the university. Inclusion criteria for participation and all procedural steps preceding data collection were as in experiment 1.
B. Stimulus material
1. Speech samples
The speech stimuli were selected from a large database of previously recorded speech samples including spontaneous conversations of several generations of speakers [n = 536, age range: 8–96 years old; more information about the corpus can be found in Fox and Jacewicz (2012), Jacewicz et al. (2010), and Jacewicz et al. (2011a)]. The speakers were local residents from tightly defined (by three adjacent counties) speech communities situated in three major dialect regions in the United States, the North (southeastern Wisconsin), the Midland (central OH), and the South [western North Carolina (NC)]. For the current study, we selected 40 speakers, 20 from central OH and 20 from western NC. The speakers were matched for age and gender as closely as possible, based on availability of participants in the database who ranged in age from 51 to 65 years. This age range was selected because it represented a specific stage of sound change in cross-generational speech patterns in each community that was abundant in traditional features of each dialect [see Jacewicz et al. (2011a) for further details]. We note that the term regional dialect (rather than regional accent) has historically been used in American sociolinguistics because, considering phonetics and phonology, the major dialect regions were defined on the basis of vowel production and systemic vowel rotations (chain shifts) within each geographic region (Labov et al., 2006). Relevant to the current study, the North is defined by the Northern Cities Shift that includes northern OH, while the South is defined by the Southern Shift that extends to southern OH, and Midland has a distinct set of vowel changes that can be found in central OH. Compared to the North and the South, the Midland lacks salient regional features and is sometimes termed General American English.
The speech community in western NC was situated in the Inland South (a subregion of the South), representing a concentration of the most salient features of Southern American English, including monophtongization of /a/, extensive /u/-fronting, and diphthongization (i.e., breaking) of front vowels, which is perceived (and commonly referred to) as Southern drawl. These vowel features are widespread—with various degrees—in the Southern states and are more prevalent among older generations of speakers. In addition to the spectral diphthong-like changes, the perceptual salience of the Southern drawl can possibly be enhanced by a distinctive type of pitch contour in monosyllabic stressed vowels. Specifically, measurements of F0 shape show a robust F0 drop from peak to offset, and the fall time is longer because the peak is located early in the vowel. In contrast, vowels in the North and central OH have a comparatively later peak alignment, and the F0 contour appears truncated [see Jacewicz and Fox (2018) for detailed acoustic analyses of these patterns]. It is possible that the salience of the Southern drawl and its perceived “slowness” comes from a combination of formant and F0 dynamics that are executed over a comparatively greater vowel duration in this dialect (Jacewicz et al., 2007, 2011b).
There were ten male and ten female speakers in each dialect group, self-identified as cisgender. For central OH, the mean ages were 57.7 ± 3.4 for males and 60.8 ± 5.9 for females; for western NC, they were 58.8 ± 5.9 for males and 59.4 ± 2.8 for females.
In choosing the speech material for the study, care was taken to ensure that the selected utterances fell between the conversational pauses of the speaker and did not contain any lexical information that might suggest the regional background of that speaker or their gender (e.g., sentences such as “I have lived in Columbus all my life” or “My wife went to the store” were not included). Ten unique utterances were chosen from each speaker for a total of 400 for use in the experiments. The selected utterances were operationally defined as being shorter (containing ≤8 syllables) or longer (>8 syllables). The utterances were separated into ten different sets of 40 unique sentences/phrases, ensuring each set contained one sentence/phrase from each speaker, and each set was balanced for dialect and gender, as well as the number of shorter and longer utterances. Within each of the ten sets, the utterances were randomized, separately for each set.
Care was taken to have a similar number of syllables in each set for each dialect (OH mean: 8.42 syllables/sentence ± 2.31; NC mean: 8.88 ± 2.34). Mean durations of utterances were 1792 ± 542 ms for OH and 2063 ± 716 ms for NC. These duration differences reflected dialect-inherent speech tempo variations. Specifically, articulation rate in the current data set (measured in number of syllables per second) was significantly greater in OH (4.93 syllables/s) than NC (4.53 syllables/s) [F(1, 36) = 4.72, p = 0.036]. Significant differences were also found for the effects of gender, with males having greater articulation rate than females (4.96 vs 4.53 syllables/s, respectively) [F(1, 36) = 6.09, p = 0.018]; there was no significant interaction. These results indicate that OH speech is faster than NC speech and that males speak faster than females, which is consistent with previously reported differences between NC and Wisconsin speakers (Jacewicz et al., 2010). A complete stimulus list and associated details for each stimulus set can be found in supplementary material1 available online.
The speakers were matched for speaking F0 as closely as possible within each male and female gender and dialect group. Acoustic analysis of F0 was conducted on the entire stimulus set (N = 400 utterances) using autocorrelation method, and the measures of F0 (mean, minimum, maximum, and range) were obtained for each utterance. As can be seen in Table I, females had higher F0 than males, and their F0 ranges were also comparatively greater. The differences between males and females were significant for each F0 measure, indicating that the higher F0 in females is evident in both higher F0 minima [F(1, 36) = 91.42, p < 0.001] and F0 maxima [F(1, 36) = 78.09, p < 0.001] and that female speech is comparatively more “expressive,” having greater F0 ranges [F(1, 36) = 25.40, p < 0.001]. The differences between the two dialects were not significant for any F0 measure, suggesting that, in the current sample, the two dialects did not differ in their overall speaking F0. However, we note that the average F0 range was slightly greater in NC (M = 93.54 Hz) than in OH (M = 82.72 Hz), which may have some effects on the perception of the Southern speech as being comparatively more “dynamic.” The interaction between dialect and gender was not significant for any F0 measure.
TABLE I.
Summary of F0 measurements averaged over each dialect and gender group.
| Speaker group | Mean F0 ± standard deviation | Minimum F0 | Maximum F0 | F0 range |
|---|---|---|---|---|
| OH male | 115.54 ± 15.08 | 89.32 | 149.77 | 60.45 |
| NC male | 121.68 ± 19.39 | 88.55 | 163.61 | 75.06 |
| OH female | 186.83 ± 25.71 | 140.77 | 245.76 | 104.99 |
| NC female | 183.45 ± 28.10 | 130.29 | 242.30 | 112.01 |
2. Signal processing and stimulus preparation
Prior to digital filtering, all utterances were root mean square (rms) equalized using matlab (The MathWorks Inc., Natick, MA). LP and HP filtering was done in matlab using the Signal Processing Toolbox. In experiment 1, there were four conditions with stimulus sentences LP-filtered at 500, 700, 900, and 1100 Hz. The control condition was with the unmodified (original) sentences recorded at a 44.1 kHz sampling rate with a 16-bit resolution. The individual sentences were LP-filtered with a digital Parks–McClellan equiripple finite impulse response filter using the FIRPM function. For each passband cutoff frequency, the stop band frequency was 50 Hz, which provided sharp attenuation slopes. The filtered signals were visually inspected. Following LP filtering, all sentences were rms equalized in matlab to reduce intensity differences between stimuli that could be used as a perceptual cue for listeners in the experiment. For each individual listener, two of the ten stimulus sentence sets were randomly assigned to each of these five conditions to ensure that each listener responded to 80 unique utterances (40 OH, 40 NC) in each of the five conditions.
In experiment 2, five experimental conditions were created with the same sentences HP-filtered at 700, 1175, 1973, 3312, and 5560 Hz. Each sentence was HP-filtered with a digital Butterworth filter using the FDESIGN.HIGHPASS function in matlab, with passband attenuation of 3 dB. The upper limit of the high-frequency region was at 11.025 kHz because the original recordings at 44.1 kHz were downsampled to 22.05 kHz prior to HP processing. Following HP filtering, the sentences were rms equalized in matlab to reduce intensity differences. The unmodified utterances were not used in experiment 2. As in experiment 1, each listener responded to 80 unique utterances (40 OH, 40 NC) in each of the five conditions.
C. Experimental procedure
Each experiment consisted of two listening tasks completed in two separate 1-h sessions: identification and intelligibility. Participants were asked to come back 2 weeks later to complete the second task. The order of the tasks was counterbalanced. A common protocol was followed in experiments 1 and 2. Testing was completed in a sound attenuating booth. Each listener was tested separately while sitting in front of a computer monitor. Signals were presented diotically at 70 dB sound pressure level (SPL) via Sennheiser (Wedemark, Germany) 640 circumaural headphones. The experiment was under computer control, using custom programs written in matlab.
In the identification task, listeners were asked to identify the gender and dialect of each speaker. After hearing each utterance, the participant indicated with a mouse click whether the speaker was male or female, from OH or NC. There were four response boxes displayed on the monitor: Central Ohio Male Speaker, North Carolina Male Speaker, Central Ohio Female Speaker, and North Carolina Female Speaker. Stimulus presentation was self-paced. No repetitions were possible (the custom matlab routines did not allow any stimulus to be played more than once).
The five conditions in each experiment were tested in separate blocks, two for each condition. Together, there were ten blocks (40 utterances each) in experiment 1 (LP-filtered and unmodified stimuli) and experiment 2 (HP-filtered stimuli). The presentation order of the ten blocks was pseudorandomized to ensure that listeners first heard sentences in each of the five filtering conditions (five blocks) before these conditions were repeated. Again, these randomizations were done for each individual listener such that no listener received the same sentence sets under the same conditions in the same order. A practice block was administered prior to each experiment, consisting of 20 utterances (two for each filtering condition) to ensure the participant could perform the task. The practice block did not include any of the experimental stimulus sentences but did include experimental variations in gender and dialect. Breaks were allowed as needed.
In the intelligibility task, stimulus presentation was exactly as in the identification task. The only difference was that after playing each sentence, a text box appeared on the screen, and the participants typed into the box what they heard following the prompt: “Now type into the box the sentence that you heard.” Participants clicked “OK” once they were satisfied with the response, which saved the response and activated the next stimulus presentation. The experiment was self-paced. No repetitions were allowed.
D. Statistical analysis
A common statistical approach was followed in analyzing data in experiments 1 and 2. For the identification task, listener responses were analyzed using SDT (Green and Swets, 1966; Macmillan and Creelman, 2005) that allows for the separation of perceptual sensitivity to stimulus and response bias (Lynn and Barrett, 2014). The SDT analyses were conducted separately for the effects of dialect and gender. In the analysis of dialect effects, the correct categorization of an OH talker was a hit, and the categorization of a NC talker as an OH talker was a false alarm. In analyzing the gender effects, the correct categorization of a male talker was a hit, and the categorization of a female talker as male was a false alarm. We used nonparametric measures of sensitivity (A′) (Snodgrass and Corwin, 1988) and bias (B′′D) (Donaldson, 1992) that do not make assumptions about underlying distributions. The sensitivity and bias scores were subsequently modeled using linear mixed-effects regression; the analyses were conducted in IBM SPSS Statistics version 28 (2022). Details about individual models are provided in Sec. III.
For the intelligibility task, the digitally recorded responses were scored based on keywords, two or three per utterance, following the scoring system developed for this task (all keywords used in scoring are marked in the utterance list in the supplementary material1 available online). Words with added or deleted morphemes were counted as incorrect, and those containing obvious spelling errors were counted as correct (e.g., if context appropriate, “bean” spelled as “been” was counted as correct). Care was taken to have all sets balanced for difficulty such that each utterance contained commonly known words; the topics of conversations from which the utterances were selected included family, hobbies, friends, daily life stories, etc. Proportion correct scores were modeled using linear mixed-effects regression as in the identification task.
III. RESULTS
Figure 1 summarizes identification and intelligibility results for both the LP and HP filtering experiments. As expected, identification decisions about a speaker's gender were near ceiling for all LP filter frequency cutoffs; however, HP filtering also retained most cues to gender, particularly for the first three filters. Dialect identification scores were comparatively lower, even when all perceptual cues were available in unmodified speech. The two types of filtering produced two different response patterns: Decisions about a speaker's dialect improved with each consecutive LP filter and worsened with each higher HP filter. Below, we present the statistical analyses of sensitivity in decision-making, response bias, and stimulus intelligibility separately for each filtering type.
FIG. 1.
(Color online) Average dialect (by gender) and gender (by dialect) identification (A′) and intelligibility accuracy by gender (in proportion correct) for low-pass filtered, unmodified, and high-pass filtered speech. Error bars indicate one standard error (SE).
A. Experiment 1: LP filtering
Sensitivity (A′) and response bias (B"D) scores were modeled using linear mixed-effects regression. The models were fitted using restricted maximum-likelihood estimation. Significance was evaluated by applying the Satterthwaite approximation of degrees of freedom for F-tests and t-tests. Decisions about significance were based on type III tests of fixed effects (the F-tests). The estimates of fixed effects were based on one-sample t-tests for the beta regression coefficients for each participant. Estimates of all models in experiment 1 are listed in Table II (see Appendix A).
1. Dialect identification
Sensitivity (A′) in making decisions about dialect was modeled with filter level (500, 700, 900, 1100, unmodified), speaker gender (male, female), and their interaction as fixed effects and a listener as a random effect. The model for dialect identification included the main effects of filter [F(4, 175) = 22.99, p < 0.001] and gender [F(1, 175) = 29.99, p < 0.001]. Unmodified speech provided significantly more cues to dialect when compared with each LP filter (p ≤ 0.003, Bonferroni-adjusted), and the lowest LP-500 Hz cutoff provided the fewest cues (p ≤ 0.002); the small differences between the LP-700 Hz, LP-900 Hz, and LP-1100 Hz cutoffs were not significant. The main effect of gender indicated that more cues to dialect were available in male speech (M = 0.806) than in female speech (M = 0.713). This was particularly true for the two lowest cutoffs, suggesting that the lower F0 in males could supply more of the spectral cues to dialect when compared with female speech. However, this interpretation should be considered with caution because the interaction between filter and gender was not significant (p = 0.066), and its inclusion did not improve the model fit.
2. Dialect response bias
To better understand listeners' decisions under uncertainty, the responses were analyzed for possible effects of bias. In SDT, a listener can have a conservative bias (being more cautious and less willing to guess a hit) or a liberal bias (being less cautious and more willing to guess a hit). In the current study, an OH speaker was a hit, and a NC speaker was a false alarm. Thus, when uncertain, a conservative listener will less likely respond Ohio (and more likely respond North Carolina), and a liberal listener will more likely respond Ohio (and less likely respond North Carolina). In the B"D measure of perceptual sensitivity used here (derived from each listener's proportions of hits and false alarms), the values lie between +1 (conservative bias) and –1 (liberal bias), and 0 indicates no bias. Therefore, a conservative listener will have a bias toward responding North Carolina (positive B"D values), a liberal listener will be biased toward Ohio (negative B"D values), and responses of an unbiased listener will not be different from 0.
The model for dialect bias included a significant filter by gender interaction [F(4, 171) = 5.94, p < 0.001]. The interaction arose because listeners were more conservative and biased toward North Carolina (positive values) in responding to female speakers in unmodified speech [one-sample t-test, M = 0.33, t(19) = 3.99, p < 0.001, d = 0.41], but to male speakers at two cutoffs, LP-700 Hz [M = 0.16, t(19) = 1.93, p = 0.034, d = 0.37] and LP-1100 Hz [M = 0.27, t(19) = 3.42, p = 0.001, d = 0.36)]. As can be seen in Fig. 2, listeners were unbiased (i.e., not significantly different from 0) for the remaining conditions. Overall, the results for bias indicate that listeners were significantly more likely to respond North Carolina when in doubt about a speaker's dialect, but their decisions were differentially influenced by a speaker's gender.
FIG. 2.
(Color online) Average response bias (B′′D) for dialect (by gender) and gender (by dialect) identification for low-pass filtered, unmodified, and high-pass filtered speech. Error bars indicate one SE.
3. Gender identification
To examine sensitivity (A′) in making decisions about gender, a linear mixed-effects regression model was fitted with filter level (500, 700, 900, 1100, unmodified), speaker dialect (OH, NC), and their interaction as fixed effects and a listener as a random effect. The model for gender identification included the fixed effects of filter [F(4, 175) = 10.35, p < 0.001] and dialect [F(1, 175) = 6.55, p = 0.011], with no interaction. As can be seen in Fig. 1, the differences between each consecutive filter were small, but some reached significance at the 0.05 level as determined by pairwise comparisons with Bonferroni correction. Specifically, unmodified speech provided significantly more cues to gender than either the LP-500 Hz, LP-700 Hz, or LP-900 Hz cutoff, and LP-500 Hz provided fewer cues than LP-1100 Hz. The difference between LP-1100 Hz and unmodified was not significant (p = 0.082), indicating that LP-filtered speech at 1100 Hz supplied enough cues to gender, yielding identification rates comparable with those for unmodified speech. The main effect of dialect indicated that gender cues were slightly easier to detect in OH speech (M = 0.980) than in NC (M = 0.976); however, this could be related to individual speakers' variation rather than representing a dialect-inherent feature.
4. Gender response bias
In calculations of gender bias, a male speaker was a hit, and a female speaker was a false alarm. Thus, when uncertain, a conservative listener will less likely respond Male (and more likely respond Female), and a liberal listener will more likely respond Male (and less likely respond Female). A conservative listener will have a bias toward responding Female (positive B"D values), a liberal listener will be biased toward Male (negative B"D values), and responses of an unbiased listener will not be different from 0.
The model for gender bias included the fixed effects of filter [F(4, 171) = 3.59, p = 0.008] and dialect [F(1, 171) = 18.50, p < 0.001]. Pairwise comparisons showed that listeners were liberal in their judgments at the three lowest LP filters (responding Male) and for NC speakers. To explore these effects further, one-sample t-tests were used to determine the significance of bias for each separate LP filter and dialect when compared with the reference level 0 (no bias). As shown in Fig. 2, listeners were significantly more likely to choose Male (negative values) in response to NC speakers for the three lowest cutoffs [LP-500 Hz: M = –0.23, t(19) = –2.13, p = 0.023, Cohen's d = 0.52; LP-700 Hz: M = –0.30, t(19) = –3.18, p = 0.002, d = 0.43; and LP-900 Hz: M = –0.24, t(19) = –2.40, p = 0.013, d = 0.46]. Listeners were significantly biased toward responding Female (positive values) to OH speakers at LP-1110 Hz [M = 0.20, t(19) = 2.71, p = 0.007, d = 0.33]. They were unbiased (i.e., not significantly different from 0) for the remaining conditions by dialect combinations.
In summary, we found that listeners were significantly biased at all four filter levels. We also note a relation between sensitivity (A′) and bias for the three lowest cutoffs: When cues to gender were only available in low-frequency energy up to 900 Hz, listeners were liberal in their decisions about gender and tended to choose Male when hearing NC speakers (but not OH speakers). However, when more speech cues became available at 1100 Hz, listeners were conservative and chose Female when responding to OH speakers (but not NC speakers), although their gender sensitivity scores (A′) were as high as for unmodified speech. However, they showed no bias in responding to unmodified speech. These results suggest that only the unprocessed utterances provided enough information to minimize listeners' uncertainties (and the associated guessing) in making decisions about gender and that LP-filtered speech—although rich in gender cues—still did not supply all necessary cues.
5. Intelligibility
Raw scores for each participant (based on keywords) were converted to proportion correct scores and modeled using linear mixed-effects regression with filter level (500, 700, 900, 1100, unmodified), speaker gender (male, female), speaker dialect (OH, NC), and their interactions as fixed effects and a listener as a random effect.
The optimal model included the fixed effects of filter [F(4, 7971) = 1467.05, p < 0.001] and gender [F(1, 7971) = 20.50, p < 0.001] and their interaction [F(4, 7971) = 3.98, p = 0.003]. Dialect was not significant (p = 0.132), and the two- and three-way interactions with dialect were not significant; they were removed from the model. As expected, intelligibility was mostly lost at the lowest frequency cutoff (500 Hz) and progressively increased with each consecutive filter, reaching the ceiling for unmodified speech (p < 0.001 for all pairwise comparisons). A significant filter by gender interaction arose because at the two lowest filter levels, male speakers were slightly more intelligible than females [paired t-test; LP-500 Hz: t(19) = 6.53, p < 0.001; LP-700 Hz: t(19) = 3.64, p < 0.001]. As shown in Fig. 1, the highest LP-1100 Hz cutoff still did not contain enough semantic content to match the intelligibility of unmodified speech.
B. Experiment 2: HP filtering
Listeners' sensitivity (A′) and bias (B"D) scores in experiment 2 were modeled exactly as in experiment 1, the HP filter levels (700, 1175, 1973, 3312, and 5560 Hz) being the only difference. The estimates of fixed effects from linear mixed-effects regression can be found in Table III (see Appendix B).
1. Dialect identification
Overall, dialect identification scores for HP-filtered speech were high (Fig. 1). The main effect of filter [F(4, 189) = 109.63, p < 0.001] showed that each higher filter provided significantly fewer cues to dialect [p < 0.001 for all pairwise comparisons except for the HP-700 Hz (M = 0.90) and HP-1175 Hz (M = 0.85) pair, p = 0.290, where A′ scores were high]; listeners performed at chance at HP-5560 Hz (M = 0.49). The model included a filter by gender interaction [F(4, 189) = 4.78, p = 0.001]. The interaction arose because listeners perceived significantly more dialect cues in female speakers than in males at the first three cutoffs [paired t-test, HP-700 Hz: M = 0.91 vs M = 0.89, t(21) = 2.06, p = 0.026; HP-1175 Hz: M = 0.87 vs M = 0.82, t(21) = 2.97, p = 0.004; HP-1973 Hz: M = 0.74 vs M = 0.64, t(21) = 4.78, p < 0.001], but no significant gender-related differences were found at the two highest cutoffs. We note that for the two highest filters, means were higher for males than for females and approached significance at the HP-3312 Hz cutoff [M = 0.64 vs M = 0.57, t(21) = –1.66, p = 0.056].
2. Dialect response bias
The model for dialect bias included a filter by gender interaction [F(4, 171) = 4.53, p = 0.002]. Listeners were conservative and significantly biased toward North Carolina (significant compared to 0 using one-sample t-test) when responding to female speakers at the first three filters [HP-700 Hz: M = 0.21, t(21) = 2.38, p = 0.013, d = 0.41; HP-1175 Hz: M = 0.25, t(21) = 3.40, p = 0.001, d = 0.35; and HP-1973 Hz: M = 0.12, t(21) = 1.75, p = 0.047, d = 0.31] and to male speakers at HP-1973 Hz [M = 0.27, t(21) = 3.01, p = 0.003, d = 0.42]. We note that the conservative bias for female speakers corresponds to their higher dialect sensitivity scores. Listeners were also biased at the highest HP-5560 Hz cutoff [liberal bias toward Ohio for female speakers, M = –0.20, t(21) = –2.63, p = 0.008, d = 0.35]; however, this finding is of limited importance given that dialect identification scores indicated chance performance.
3. Gender identification
Overall, gender identification was high. Sensitivity scores (A′) to gender were near ceiling at the first three filters (see Fig. 1) and significantly declined only at the two highest filters, HP-3312 Hz (M = 0.91) and HP-5560 Hz (M = 0.85); all pairwise comparisons with these two filters were significant at p < 0.001. The model included a filter by dialect interaction [F(4, 189) = 5.06, p < 0.001] that arose because of the mixed patterns for the two highest filters. At HP-5560 Hz, listeners detected more gender cues in NC speakers (M = 0.87) than in OH speakers (M = 0.82) [paired t-test, t(21) = 4.07, p < 0.001], while this order was reversed at HP-3312 Hz [M = 0.92 for OH and M = 0.89 for NC, t(21) = 2.13, p = 0.023]. Dialect also significantly influenced listeners' decisions about gender at the “easiest” HP-700 Hz cutoff; however, the difference was very small [M = 0.99 for NC and M = 0.98 for OH, t(21) = 3.17, p = 0.002].
4. Gender response bias
The model for gender bias also included a filter by dialect interaction [F(4, 189) = 3.45, p = 0.009]. Exploring this interaction, a one-sample t-test showed that listeners were significantly biased only at the highest HP-5560 Hz cutoff (see Fig. 2). They were conservative (responding Female, positive B"D values) when hearing OH speakers [M = 0.22, t(21) = 2.17, p = 0.021, d = 0.47] but liberal (responding Male, negative B"D values) when hearing NC speakers [M = –0.18, t(21) = –1.95, p = 0.033, d = 0.43]. We note here a correspondence between the sensitivity (A′) results and response bias. There was no such correspondence at HP-3312 Hz as listeners were not significantly biased [p = 0.481 for OH and p = 0.084 for NC], suggesting that greater amounts of cues to gender at a comparatively lower filter level reduced listeners' uncertainty (and the associated guessing). The listeners were also unbiased for the remaining listening conditions.
5. Intelligibility
Raw scores for each participant were converted to proportion correct scores and modeled using linear mixed-effects regression with filter level (700, 1175, 1973, 3312, and 5560 Hz), speaker gender (male, female), speaker dialect (OH, NC), and their interactions as fixed effects and a listener as a random effect. The model for intelligibility included three two-way interactions; a three-way interaction did not improve the model fit and was removed. The results showed that intelligibility dropped sharply when speech was HP-filtered at 1973 Hz (from M = 0.80 at HP-1175 Hz to M = 0.28 at HP-1973 Hz), and utterances were almost impossible to understand at the two highest filters (all pairwise comparisons for the main effect of filter were significant at p < 0.001). The main effect of gender indicated higher accuracy scores for female speakers than for males. However, a significant gender by filter interaction [F(4, 8762) = 51.31, p < 0.001] arose because the differences between females and males were small and not significantly different from each other for pairwise comparisons at the “easiest” (HP-700 Hz) and the most difficult (HP-3312 Hz and HP-5560 Hz) cutoffs, but the female advantage was apparent at the two remaining filters (see Fig. 1). Specifically, the difference between females and males was the largest at HP-1973 Hz and second largest at HP-1175 Hz, and both differed significantly from one another [t(21) = –6.02, p < 0.001]. This interaction indicates that, when compared with males, female speech provided appreciably more intelligibility cues at HP-1973 Hz (by M = 0.23) and, to a lesser extent, at HP-1175 Hz (by M = 0.08).
The second significant interaction between dialect and filter [F(48 762) = 11.88, p < 0.001] arose because accuracy was higher for OH speakers than for NC speakers for the first three cutoffs [HP-700 Hz: t(21) = 7.27, p < 0.001; HP-1175 Hz: t(21) = 6.90, p < 0.001; and HP-1973 Hz: t(21) = 5.81, p < 0.001], and there were no significant dialect-related differences for the two highest filters. This interaction indicates that under challenging conditions, listeners from OH understood OH speakers better than NC speakers, and, thus, familiarity with speaker dialect enhanced accuracy scores. This finding is consistent with research showing that perceptual attunement to phonetic features of a particular variety of language emerges early in life [see Werker (2018) for a review]. The third significant interaction between dialect and gender [F(4, 8762) = 6.72, p = 0.010] added that accuracy was higher for OH females when compared to NC females [t(4397) = 2.56, p = 0.005].
IV. DISCUSSION
This study aimed at exploring the availability of spectral information in the low- and high-frequency (up to 11 kHz) regions to the perception of a speaker's dialect and gender. Using LP and HP filtering, we showed that cues to gender were well preserved in both low- and high-frequency ranges and did not depend on intelligibility, and the redundancy of gender cues at higher frequencies reduced response bias. Unlike for gender, the most perceptual cues to dialect were in intelligible speech, although dialect could still be identified above chance even when semantic information was no longer available. We discuss these results in more detail below. When discussing the intelligibility results, we refer to the accuracy data in percentages rather than proportions to prevent potential confusion with the measurement scales for sensitivity and bias.
A. LP filtering
The intelligibility results for the LP filtering experiment demonstrated that LP-filtered speech up to 900 Hz was mostly unintelligible. Intelligibility improved at LP-1100 Hz, reaching an overall accuracy of 56%, which indicates that the available cues were still far from approximating intelligibility of unmodified speech. Recall that French and Steinberg (1947) reported 68% accuracy for nonsense syllables at LP-1.9 kHz. Although we have not probed effects of higher LP cutoffs, our results appear consistent with their study because we could expect further intelligibility improvements as more acoustic cues are provided within the 1.1–1.9 kHz band.
Our prediction of greater intelligibility of male than female speech was confirmed for the two lowest filter levels; however, no gender-related differences were found for LP-900 Hz cutoff and above. These results indicate that intelligibility predictions based solely on information in F0 and lower vowel formants in male speech need to be revised regarding the contribution of the acoustics of consonant production. That is, intelligibility of female speech was not different from males when more cues to consonantal sonorants (nasals and liquids) and dynamic cues in first formant transitions became available at LP-900 Hz, but information in higher formant transitions and fricatives was still lacking for both genders.
The results for dialect identification showed that the most cues to dialect were in unmodified speech, indicating that perceptual sensitivity to dialect features is at its best when all semantic, segmental, and prosodic information remains intact. As predicted, listeners were able to identify dialects above the chance level at LP-500 Hz. This finding is consistent with the reports that prosodic cues differentiating dialects in LP-filtered speech at 300–400 Hz are strong enough to yield above-chance performance (Van Bezooijen and Gooskens, 1999; Frota et al., 2002; Van Leyden and Van Heuven, 2006). However, the differences between LP-700 Hz, LP-900 Hz, and LP-1100 Hz were not significant, and dialect identification for these cutoffs was high, more so in male than in female speech. The lack of improvement with each higher-frequency cutoff does not correspond to the significant increase in intelligibility at each filter level. This outcome answers our first research question in that a progressively higher LP filter cutoff (above LP-700 Hz) did not supply increasingly more information about a speaker's dialect, suggesting that identification decisions were most likely based on prosodic cues, dialect-specific articulation rates, and familiarity with those dialect features that could be detected in mostly unintelligible speech.
However, speaker gender did contribute to listeners' dialect identification decisions. Overall, listeners found more dialect cues in males (M = 0.78) than in females (M = 0.68). Although the greatest differences were at the two lowest frequency cutoffs, the unmodified speech still supplied slightly more cues to dialect in males than in females (M = 0. 90 and M = 0.86, respectively). The results for LP-filtered speech indicate that comparatively more of the segmental content along with dynamic information in vowel formants (i.e., Southern breaking) in male speakers could contribute some additional cues, possibly increasing the salience of the Southern drawl relative to female speakers of this dialect.
Overall, listeners showed some gender-related bias in making decisions about dialect. They were significantly biased toward North Carolina only in responding to male speakers at 700 and 1100 Hz cutoffs and to female speakers in unmodified speech. The lack of a systematic pattern complicates a convincing interpretation of which dialect characteristics could have contributed to this variable outcome. It is surprising that listeners were biased when all dialect cues were available in unmodified speech. It could be that, when in doubt, they shifted their criterion toward North Carolina when hearing slower speaking rate and greater variation in pitch, but it is less clear why they would do that for females in one condition and for males in another, especially when there was no obvious relation between intelligibility and dialect identification scores. That is, it is unclear why listeners were biased at 700 Hz cutoff and not at 500 Hz and why uncertainty increased as more semantic content was provided at the highest filter level and in unmodified speech, where dialect identification was at its best. Clearly, more work is needed to better understand the causes of response bias, also considering listeners' characteristics that underlie their cognitive decisions.
Predictably, gender identification did not depend on intelligibility and remained near ceiling across all filter levels, although the 1100 Hz cutoff still supplied some additional cues. These results are consistent with previous studies (Lass et al., 1976; Lass et al., 1980). However, irrespective of their high sensitivity scores, listeners were significantly biased in making decisions about gender at all LP filter levels, including LP-1100 Hz. They were only unbiased in responding to unmodified speech. The bias results contribute a new finding, suggesting that LP filtering—although preserving F0 and at least some formant frequency, prosodic, and verbal information—nevertheless increased uncertainty in listeners' decision-making, and only the redundancy of unmodified speech provided all needed cues for accurate gender identification.
We found that for all three filters up to 900 Hz, listeners were biased toward responding Male when hearing NC speakers, suggesting that certain dialect cues in the voices of NC speakers (whether male or female), prompted the Male response. Based on the acoustic measurements, the bias could be related to a slower articulation rate of NC speakers and greater F0 ranges in their utterances when compared with OH speakers (see Table I). The combination of these two properties could have contributed to listeners' adjusting their criterion for NC male speech as being comparatively slower and more “dynamic” and using this criterion even when hearing NC female speakers. This seems plausible because, although little is known about a potential contribution of speech rate to perceptually differentiating between males and females, there is some evidence that speech tempo is not a strong perceptual cue to gender (Leung et al., 2018), and thus, shifting the criterion likely involved listeners' decisions based on dialect differences and not gender differences. The presence of the first formant (and even the second formant in some utterances) could have contributed additional dynamic cues signaling the Southern breaking in vowels (e.g., Fox and Jacewicz, 2009). We note that the stimulus utterances did not contain pauses and were rms equalized to reduce intensity differences, and thus, no perceptual cues were provided in either pausing patterns or intensity variations. Information in the sentence intonation contours was likely unreliable because the utterances were extracted from spontaneous narratives, and no systematic patterns were produced by either NC or OH speakers. Therefore, the differences in articulation rates and speaking F0 ranges along with selected formant frequency cues seem to be the main contributors to how listeners shifted their criterion based on their dialect-specific expectations.
On the other hand, at LP-1100 Hz, listeners were biased toward responding Female when hearing OH speech. It is unclear why they became more conservative (i.e., more cautious in their decisions about gender) when intelligibility increased at the highest LP filter. At present, we have no obvious explanation for why listeners shifted their criterion so that a comparatively faster and less prosodically “expressive” OH speech could likely be associated with female speakers. Analogous perceptual data from NC listeners could verify whether greater familiarity with the ambient dialect contributed to the response bias observed here and, in general, increase our understanding of bias in SDT.
B. HP filtering
We found that intelligibility remained high when speech was HP-filtered at 700 Hz (93%) and even 1175 Hz (80%). This result is consistent with the early findings for nonsense syllables that speech comprehension was not markedly compromised by removing the spectral content below 1 kHz (Fletcher, 1953; French and Steinberg, 1947). However, there was a sharp intelligibility drop at HP-1973 Hz, more so in males than in females, which is inconsistent with the above studies. Specifically, intelligibility for nonsense syllables at 1.9 kHz in French and Steinberg (1947) was well above chance (68%), whereas it was well below chance in our results for spontaneous speech (40% for females, 16% for males). A further discrepancy was at HP-3312 Hz (30% versus 10% for females and 6% for males in our study). These conflicting results indicate that HP filtering, distorting segmental cues, is more detrimental to the intelligibility of a string of polysyllabic words in spontaneous (and often semantically unpredictable) utterances than to intelligibility of isolated nonsense syllables.
The gender-based differences at HP-1175 Hz and HP-1973 Hz confirmed our predictions that female speech would provide comparatively more intelligibility cues due to the reduction of consonantal information in male speech. Another relevant finding was that familiarity with a dialect enhanced intelligibility of HP-filtered speech, so that listeners from OH found significantly more verbal cues in speakers from OH than from NC.
Dialect identification was high at both HP-700 Hz (M = 0.90) and HP-1175 Hz (M = 0.85), which was comparable with unmodified speech (M = 0.88). Predictably, most cues to dialect were in intelligible speech. Addressing our second research question, the predicted linear relationship between dialect identification and intelligibility was confirmed in that each higher cutoff frequency decreased the amount of available dialect cues, and greater intelligibility of female speech at HP-1175 Hz and HP-1973 Hz resulted in higher dialect identification rates compared with males. However, we underscore that semantic content was not the only carrier of dialect information. We found that dialect sensitivity scores were still relatively high even when intelligibility was lost. Specifically, they were above chance at HP-3312 Hz (M = 0.60) and at chance at HP-5560 Hz (M = 0.49), when listeners only heard unintelligible speech of high-pitched chirping quality. Possibly, some dialect cues guiding listeners' decisions could still be found in articulation rate and speaking F0 ranges. We have not explored a potential contribution of rhythm to dialect identification, which needs to be addressed in future research using the most effective rhythm metrics.
The results for dialect bias in HP-filtered speech showed again that listeners were significantly biased toward responding North Carolina (except for the highest filter). This finding is consistent with that for the LP-filtered speech and adds an incremental insight into our understanding of dialect bias. In particular, listeners were biased when both intelligibility and dialect sensitivity scores were high, including HP-700 Hz, HP-1175 Hz, and HP-1973 Hz, unmodified speech, and LP-1100 Hz. The contribution of a speaker's gender is still unclear, but we can infer that listeners' uncertainty about dialect increased when more rather than less dialect information was available in speech. At present, we still do not have a clear understanding of which speech characteristics might contribute to variations in listener confidence, especially when analyzing response bias for socially based variables that may differentially affect decision processes of individual listeners. A similar conclusion was reached by Owren et al. (2007), who analyzed their data for gender bias. Analyses of perceptual responses to speakers' characteristics within a signal detection framework are still rare, and further research is needed to establish consistencies and common trends among studies.
In contrast to the intelligibility and dialect identification results, sensitivity to speaker gender remained high, indicating that cues to a speaker's gender are distributed over a wide frequency range and do not critically depend on either verbal information or F0 in the low-frequency region. Our results are, thus, consistent with several previous studies showing the utility of the high-frequency energy in gender perception (Donai and Halbritter, 2017; Donai and Lass, 2015; Monson et al., 2014).
The results for gender bias showed that listeners were only significantly biased at HP-5560 Hz, where sensitivity to a speaker's gender was reduced when compared with the other filter levels. We can conclude that the comparatively fewer cues at the highest frequency cutoff increased listeners' uncertainty and introduced a significant response bias. The direction of the bias was consistent with that in the LP filtering experiment. When in doubt, listeners were more likely to respond Male when hearing NC speakers, and Female when hearing OH speech. Recall that listeners were not biased in responding to unmodified speech in experiment 1, and they were not biased when hearing HP-filtered speech (except for the highest filter). This indicates that cues to gender in the high-frequency region reduced uncertainty, whereas those available solely in the low-frequency range reduced confidence. This finding is significant because it suggests that the redundancy of gender information in the high-frequency range can provide additional cues that perceptually disambiguate a male or a female voice. Further research is needed to address this issue in a more focused design.
Our results for HP-filtered speech can be considered foundational, breaking ground for further explorations of physical, psychological, and social characteristics of speakers based on those properties of speech that are out of the range of intelligibility. We found that aspects of voice quality are still preserved in unintelligible speech, and the high-frequency energy can still be useful in reducing listener uncertainties about a speaker's gender or providing supplementary information about dialect. Because gender cues are well preserved in the high-frequency region, it is possible that other speaker characteristics conveyed by voice, such as age, race, health status, emotions, mood, or arousal, could also be derived from high-frequency information [see Kreiman and Sidtis (2011) for a range of possibilities].
Our results for LP and HP filtering have implications for the development of speech intelligibility models. Comparing the results of both types of filtering, French and Steinberg (1947) reported that the effects of LP and HP filtering are similar at around 2 kHz for nonsense syllables, resulting in intelligibility of about 68% for either filtering type. Since that time, the relative importance of various frequencies to speech intelligibility has been studied, revised, and standardized in the Speech Intelligibility Index Standard [ANSI S3.5–1997; see Amlani et al. (2002) for review], and new intelligibility models and measures have been proposed (e.g., Chen, 2011; Müsch and Buus, 2001a,b). Our results for spontaneous speech show that, when speech was HP-filtered at around 2 kHz, intelligibility dropped to about 30%, which is inconsistent with these established norms. Considering the effects of both LP and HP filtering, we found that LP-filtered speech at 1.1 kHz approached intelligibility, and HP-filtered speech close to 2 kHz approached incomprehensibility. Therefore, the frequency band between 1.1 and 2 kHz can be considered critical to speech intelligibility. It is possible that this relatively narrow band can encompass variable pronunciations of phonetic segments across multiple speakers and provide sufficient lexical information to categorize ambiguous consonants during top-down decision processes underlying perception of spontaneous speech (Norris et al., 2003). Systematic research is needed to explore the contribution of this important region in various speaking contexts.
V. CONCLUSIONS AND FUTURE DIRECTIONS
This study demonstrates that cues to a speaker's gender are distributed over a wide frequency range in speech and are well preserved even in the high-frequency band of 5.5–11 kHz. The redundancy of gender cues at high frequencies can reduce response bias, whereas listeners can be significantly biased when only low-frequency information is available. Cues to dialect are still available in low- and high-frequency regions; however, most cues to dialect are in intelligible speech. The results suggest that the frequency band between 1.1 and 2 kHz can be critical to intelligibility of spontaneous speech. Further studies need to investigate whether other speaker characteristics conveyed by voice, such as age, race, or health status, can be cued outside of the intelligibility range, including the high-frequency region. Future development of speech intelligibility models needs to be informed by talker variability in spontaneous speech. More research is needed to investigate the extent to which accuracy of listeners' decisions should be adjusted for their response bias.
ACKNOWLEDGMENTS
We thank Zane Smith and Magan McClurg for their assistance with data collection and processing. This research was supported, in part, by the National Institute of Deafness and Other Communication Disorders of the National Institutes of Health under Award No. R01DC006871.
APPENDIX A
Table II shows linear mixed-effects modeling results for experiment 1 (low-pass filtering).
TABLE II.
Linear mixed-effects modeling results for experiment 1 (low-pass filtering). Shown are summaries of best-fit model estimates (β) with their corresponding SEs, test statistics, and significance values.
| Variable | Fixed effect | β | SE | t | p-value |
|---|---|---|---|---|---|
| Gender identification | (Intercept) | 0.985 | 0.002 | 479.248 | <0.001 |
| Dialect (OH)a | 0.004 | 0.002 | 2.559 | 0.011 | |
| LP filter level (500 Hz)b | −0.016 | 0.003 | −6.104 | <0.001 | |
| LP filter level (700 Hz) | −0.012 | 0.003 | −4.477 | <0.001 | |
| LP filter level (900 Hz) | −0.010 | 0.003 | −3.758 | <0.001 | |
| LP filter level (1100 Hz) | −0.007 | 0.003 | −2.661 | 0.009 | |
| Gender response bias | (Intercept) | −0.115 | 0.065 | −1.778 | 0.078 |
| Dialect (OH) | 0.204 | 0.048 | 4.298 | <0.001 | |
| LP filter level (500 Hz) | −0.142 | 0.075 | −1.889 | 0.061 | |
| LP filter level (700 Hz) | −0.175 | 0.075 | −2.335 | 0.021 | |
| LP filter level (900 Hz) | −0.083 | 0.075 | −1.109 | 0.269 | |
| LP filter level (1100 Hz) | 0.069 | 0.075 | 0.913 | 0.362 | |
| Dialect identification | (Intercept) | 0.835 | 0.024 | 35.175 | <0.001 |
| Gender (male)c | 0.093 | 0.017 | 5.476 | <0.001 | |
| LP filter level (500 Hz) | −0.252 | 0.027 | −9.382 | <0.001 | |
| LP filter level (700 Hz) | −0.149 | 0.027 | −5.543 | <0.001 | |
| LP filter level (900 Hz) | −0.098 | 0.027 | −3.670 | <0.001 | |
| LP filter level (1100 Hz) | −0.111 | 0.027 | −4.142 | <0.001 | |
| Dialect response bias | (Intercept) | 0.325 | 0.086 | 3.791 | <0.001 |
| Gender (male) | −0.379 | 0.106 | −3.574 | <0.001 | |
| LP filter level (500 Hz) | −0.372 | 0.106 | −3.508 | <0.001 | |
| LP filter level (700 Hz) | −0.253 | 0.106 | −2.385 | 0.018 | |
| LP filter level (900 Hz) | −0.315 | 0.106 | −2.974 | 0.003 | |
| LP filter level (1100 Hz) | −0.385 | 0.106 | −3.636 | <0.001 | |
| Gender (male) × LP filter level (500 Hz) | 0.440 | 0.150 | 2.934 | 0.004 | |
| Gender (male) × LP filter level (700 Hz) | 0.464 | 0.150 | 3.096 | 0.002 | |
| Gender (male) × LP filter level (900 Hz) | 0.471 | 0.150 | 3.144 | 0.002 | |
| Gender (male) × LP filter level (1100 Hz) | 0.713 | 0.150 | 4.755 | <0.001 | |
| Intelligibility | (Intercept) | 0.978 | 0.020 | 49.454 | <0.001 |
| Gender (male) | 0.004 | 0.017 | 0.210 | 0.834 | |
| LP filter level (500 Hz) | −0.912 | 0.017 | −54.001 | <0.001 | |
| LP filter level (700 Hz) | −0.663 | 0.017 | −39.215 | <0.001 | |
| LP filter level (900 Hz) | −0.481 | 0.017 | −28.449 | <0.001 | |
| LP filter level (1100 Hz) | −0.433 | 0.017 | −25.613 | <0.001 | |
| Gender (male) × LP filter level (500 Hz) | 0.057 | 0.024 | 2.380 | 0.017 | |
| Gender (male) × LP filter level (700 Hz) | 0.072 | 0.024 | 3.017 | 0.003 | |
| Gender (male) × LP filter level (900 Hz) | −0.004 | 0.024 | −0.166 | 0.868 | |
| Gender (male) × LP filter level (1100 Hz) | 0.028 | 0.024 | 1.186 | 0.236 |
Reference level: NC.
Reference level: unmodified.
Reference level: female.
APPENDIX B
Table III shows linear mixed-effects modeling results for experiment 2 (high-pass filtering).
TABLE III.
Linear mixed-effects modeling results for experiment 2 (high-pass filtering). Shown are summaries of best-fit model estimates (β) with their corresponding SEs, test statistics, and significance values.
| Variable | Fixed effect | β | SE | t | p-value |
|---|---|---|---|---|---|
| Gender identification | (Intercept) | 0.872 | 0.010 | 84.493 | <0.001 |
| Dialect (OH)a | −0.049 | 0.012 | −3.936 | <0.001 | |
| HP filter level (700 Hz)b | 0.114 | 0.012 | 9.217 | <0.001 | |
| HP filter level (1175 Hz) | 0.104 | 0.012 | 8.362 | <0.001 | |
| HP filter level (1973 Hz) | 0.081 | 0.012 | 6.555 | <0.001 | |
| HP filter level (3312 Hz) | 0.020 | 0.012 | 1.639 | 0.103 | |
| Dialect (OH) × HP filter level (700 Hz) | 0.043 | 0.018 | 2.429 | 0.016 | |
| Dialect (OH) × HP filter level (1175 Hz) | 0.045 | 0.018 | 2.584 | 0.011 | |
| Dialect (OH) × HP filter level (1973 Hz) | 0.056 | 0.018 | 3.198 | 0.002 | |
| Dialect (OH) × HP filter level (3312 Hz) | 0.076 | 0.018 | 4.334 | <0.001 | |
| Gender response bias | (Intercept) | −0.177 | 0.091 | −1.945 | 0.054 |
| Dialect (OH) | 0.393 | 0.112 | 3.504 | <0.001 | |
| HP filter level (700 Hz) | 0.154 | 0.112 | 1.371 | 0.172 | |
| HP filter level (1175 Hz) | 0.153 | 0.112 | 1.367 | 0.173 | |
| HP filter level (1973 Hz) | −0.026 | 0.112 | −0.232 | 0.817 | |
| HP filter level (3312 Hz) | 0.034 | 0.112 | 0.300 | 0.765 | |
| Dialect (OH) × HP filter level (700 Hz) | −0.511 | 0.159 | −3.220 | 0.002 | |
| Dialect (OH) × HP filter level (1175 Hz) | −0.492 | 0.159 | −3.100 | 0.002 | |
| Dialect (OH) × HP filter level (1973 Hz) | −0.299 | 0.159 | −1.883 | 0.061 | |
| Dialect (OH) × HP filter level (3312 Hz) | −0.245 | 0.159 | −1.546 | 0.124 | |
| Dialect identification | (Intercept) | 0.460 | 0.024 | 18.877 | <0.001 |
| Gender (male)c | 0.063 | 0.032 | 1.945 | 0.053 | |
| HP filter level (700 Hz) | 0.448 | 0.032 | 13.871 | <0.001 | |
| HP filter level (1175 Hz) | 0.413 | 0.032 | 12.804 | <0.001 | |
| HP filter level (1973 Hz) | 0.281 | 0.032 | 8.690 | <0.001 | |
| HP filter level (3312 Hz) | 0.111 | 0.032 | 3.425 | <0.001 | |
| Dialect (OH) × HP filter level (700 Hz) | −0.081 | 0.046 | −1.776 | 0.077 | |
| Dialect (OH) × HP filter level (1175 Hz) | −0.113 | 0.046 | −2.467 | 0.015 | |
| Dialect (OH) × HP filter level (1973 Hz) | −0.159 | 0.046 | −3.478 | <0.001 | |
| Dialect (OH) × HP filter level (3312 Hz) | 0.002 | 0.046 | 0.044 | 0.965 | |
| Dialect response bias | (Intercept) | −0.199 | 0.083 | −2.392 | 0.018 |
| Gender (male) | 0.206 | 0.106 | 1.943 | 0.053 | |
| HP filter level (700 Hz) | 0.408 | 0.106 | 3.853 | <0.001 | |
| HP filter level (1175 Hz) | 0.450 | 0.106 | 4.248 | <0.001 | |
| HP filter level (1973 Hz) | 0.316 | 0.106 | 2.982 | 0.003 | |
| HP filter level (3312 Hz) | 0.262 | 0.106 | 2.478 | 0.014 | |
| Gender (male) × HP filter level (700 Hz) | −0.501 | 0.150 | −3.347 | <0.001 | |
| Gender (male) × HP filter level (1175 Hz) | −0.441 | 0.150 | −2.949 | 0.004 | |
| Gender (male) × HP filter level (1973 Hz) | −0.054 | 0.150 | −0.361 | 0.719 | |
| Gender (male) × HP filter level (3312 Hz) | −0.192 | 0.150 | −1.284 | 0.201 | |
| Intelligibility | (Intercept) | 0.034 | 0.014 | 2.454 | 0.015 |
| Gender (male) | −0.034 | 0.014 | −2.486 | 0.013 | |
| Dialect (OH) | −0.005 | 0.014 | 0.700 | −0.385 | |
| HP filter level (700 Hz) | 0.895 | 0.015 | 59.098 | <0.001 | |
| HP filter level (1175 Hz) | 0.759 | 0.015 | 50.062 | <0.001 | |
| HP filter level (1973 Hz) | 0.328 | 0.015 | 21.650 | <0.001 | |
| HP filter level (3312 Hz) | 0.070 | 0.015 | 4.637 | <0.001 | |
| Gender (male) × Dialect (OH) | 0.029 | 0.011 | 2.592 | 0.010 | |
| Gender (male) × HP filter level (700 Hz) | −0.001 | 0.018 | −0.056 | 0.956 | |
| Gender (male) × HP filter level (1175 Hz) | −0.059 | 0.018 | −3.381 | <0.001 | |
| Gender (male) × HP filter level (1973 Hz) | −0.213 | 0.018 | −12.145 | <0.001 | |
| Gender (male) × HP filter level (3312 Hz) | −0.027 | 0.018 | −1.555 | 0.120 | |
| Dialect (OH) × HP filter level (700 Hz) | 0.036 | 0.018 | 2.078 | 0.038 | |
| Dialect (OH) × HP filter level (1175 Hz) | 0.095 | 0.018 | 5.433 | <0.001 | |
| Dialect (OH) × HP filter level (1973 Hz) | 0.074 | 0.018 | 4.228 | <0.001 | |
| Dialect (OH) × HP filter level (3312 Hz) | 0.001 | 0.018 | 0.084 | 0.933 |
Reference level: NC.
Reference level: unmodified.
Reference level: female.
This paper is part of a special issue on Perception and Production of Sounds in the High-Frequency Range of Human Speech.
Footnotes
See supplementary material at https://doi.org/10.1121/10.0020906 for a complete list of utterances used in the experiments, marked key words, and details associated with each stimulus set.
References
- 1. Alexander, J. M. (2019). “ The S-SH Confusion Test and the effects of frequency lowering,” J. Speech Lang. Hear. Res. 62, 1486–1505. 10.1044/2018_JSLHR-H-18-0267 [DOI] [PubMed] [Google Scholar]
- 2. Allen, J. B. (1996). “ Harvey Fletcher's role in the creation of communication acoustics,” J. Acoust. Soc. Am. 99, 1825–1839. 10.1121/1.415364 [DOI] [PubMed] [Google Scholar]
- 3. Amlani, A. M. , Punch, J. L. , and Ching, T. Y. C. (2002). “ Methods and applications of the audibility index in hearing aid selection and fitting,” Trends Amplif. 6, 81–129. 10.1177/108471380200600302 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Burnham, D. , Kitamura, C. , and Vollmer-Conna, U. (2002). “ What's new pussycat? On talking to babies and animals,” Science 296(5572), 1435. 10.1126/science.1069587 [DOI] [PubMed] [Google Scholar]
- 5. Chen, F. (2011). “ The relative importance of temporal envelope information for intelligibility prediction: A study on cochlear-implant vocoded speech,” Med. Eng. Phys. 33, 1033–1038. 10.1016/j.medengphy.2011.04.004 [DOI] [PubMed] [Google Scholar]
- 6. Clopper, C. , and Smiljanic, R. (2011). “ Effects of gender and regional dialect on prosodic patterns in American English,” J. Phon. 39, 237–245. 10.1016/j.wocn.2011.02.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Clopper, C. , and Smiljanic, R. (2015). “ Regional variation in temporal organization in American English,” J. Phon. 49, 1–15. 10.1016/j.wocn.2014.10.002 [DOI] [Google Scholar]
- 8. Deshpande, M. S. , and Holambe, R. S. (2011). “ Robust speaker identification in the presence of car noise,” Int. J. Biom. 3, 189–205. 10.1504/IJBM.2011.040815 [DOI] [Google Scholar]
- 9. Donai, J. J. , and Halbritter, R. M. (2017). “ Gender identification using high-frequency speech energy: Effects of increasing the low-frequency limit,” Ear Hear. 38, 65–73. 10.1097/AUD.0000000000000353 [DOI] [PubMed] [Google Scholar]
- 10. Donai, J. J. , and Lass, N. J. (2015). “ Gender identification from high-pass filtered vowel segments: The use of high-frequency energy,” Atten. Percept. Psychophys. 77, 2452–2462. 10.3758/s13414-015-0945-y [DOI] [PubMed] [Google Scholar]
- 11. Donaldson, W. (1992). “ Measuring recognition memory,” J. Exp. Psychol. Gen. 121, 275–277. 10.1037/0096-3445.121.3.275 [DOI] [PubMed] [Google Scholar]
- 12. Fitch, W. T. , and Giedd, J. (1999). “ Morphology and development of the human vocal tract: A study using magnetic resonance imaging,” J. Acoust. Soc. Am. 106, 1511–1522. 10.1121/1.427148 [DOI] [PubMed] [Google Scholar]
- 13. Fletcher, H. (1953). Speech and Hearing in Communication ( Van Nostrand, New York: ) (reprinted by the Acoustical Society of America, 1995). [Google Scholar]
- 14. Fletcher, H. , and Galt, R. H. (1950). “ The perception of speech and its relation to telephony,” J. Acoust. Soc. Am. 22, 89–151. 10.1121/1.1906605 [DOI] [Google Scholar]
- 15. Fox, R. A. , and Jacewicz, E. (2009). “ Cross-dialectal variation in formant dynamics of American English vowels,” J. Acoust. Soc. Am. 126, 2603–2618. 10.1121/1.3212921 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Fox, R. A. , and Jacewicz, E. (2012). “ Dialectal and generational variations in vowels in spontaneous speech,” in Proceedings of Interspeech 2012, September 9–13, Portland, OR (International Speech Communication Association, Baixas, France), pp. 1404–1407. [Google Scholar]
- 17. French, N. R. , and Steinberg, J. C. (1947). “ Factors governing the intelligibility of speech sounds,” J. Acoust. Soc. Am. 19, 90–119. 10.1121/1.1916407 [DOI] [Google Scholar]
- 18. Frota, S. , Vigario, M. , and Martins, F. (2002). “ Language discrimination and rhythm classes: Evidence from Portuguese,” in Proceedings of Speech Prosody 2002, April 11–13, Aix-en-Provence, France (International Speech Communication Association, Baixas, France), pp. 319–322. [Google Scholar]
- 19. Gorea, A. , and Sagi, D. (2005). “ Decision and attention,” in Neurobiology of Attention, edited by Itti L., Rees G., and Tsotsos J. K. ( Elsevier Academic, Amsterdam: ), pp. 152–159. [Google Scholar]
- 20. Green, D. W. , and Swets, J. A. (1966). Signal Detection Theory and Psychophysics ( Wiley, New York: ). [Google Scholar]
- 21. Hillenbrand, J. , Getty, L. A. , Clark, M. J. , and Wheeler, K. (1995). “ Acoustic characteristics of American English vowels,” J. Acoust. Soc. Am. 97, 3099–3111. 10.1121/1.411872 [DOI] [PubMed] [Google Scholar]
- 22. Hunter, L. L. , Monson, B. B. , Moore, D. R. , Dhar, S. , Wright, B. A. , Munro, K. J. , Zadeh, L. M. , Blankenship, C. M. , Stiepan, S. M. , and Siegel, K. H. (2020). “ Extended high frequency hearing and speech perception implications in adults and children,” Hear. Res. 397, 107922. 10.1016/j.heares.2020.107922 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Jacewicz, E. , and Fox, R. A. (2018). “ Regional variation in fundamental frequency of American English vowels,” Phonetica 75, 273–309. 10.1159/000484610 [DOI] [PubMed] [Google Scholar]
- 24. Jacewicz, E. , Fox, R. A. , and Salmons, J. (2007). “ Vowel duration in three American English dialects,” Am. Speech 82(4), 367–385. 10.1215/00031283-2007-024 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Jacewicz, E. , Fox, R. A. , and Salmons, J. (2011a). “ Cross-generational vowel change in American English,” Lang. Var. Change 23(1), 45–86. 10.1017/S0954394510000219 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Jacewicz, E. , Fox, R. A. , and Salmons, J. (2011b). “ Vowel change across three age groups of speakers in three regional varieties of American English,” J. Phon. 39, 683–693. 10.1016/j.wocn.2011.07.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Jacewicz, E. , Fox, R. A. , and Wei, L. (2010). “ Between-speaker and within-speaker variation in speech tempo of American English,” J. Acoust. Soc. Am. 128(2), 839–850. 10.1121/1.3459842 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Jongman, A. , Wayland, R. , and Wong, S. (2000). “ Acoustic characteristics of English fricatives,” J. Acoust. Soc. Am. 108, 1252–1263. 10.1121/1.1288413 [DOI] [PubMed] [Google Scholar]
- 29. Kent, R. D. , and Read, C. (1992). The Acoustic Analysis of Speech ( Singular, San Diego, CA: ). [Google Scholar]
- 30. Kewley-Port, D. , Pisoni, D. B. , and Studdert-Kennedy, M. (1983). “ Perception of static and dynamic acoustic cues to place of articulation in initial stop consonants,” J. Acoust. Soc. Am. 73, 1779–1793. 10.1121/1.389402 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Kitamura, C. , and Burnham, D. (2003). “ Pitch and communicative intent in mother's speech: Adjustments for age and sex in the first year,” Infancy 4, 85–110. 10.1207/S15327078IN0401_5 [DOI] [Google Scholar]
- 32. Kitayama, S. , and Ishii, K. (2002). “ Word and voice: Spontaneous attention to emotional utterances in two languages,” Cogn. Emot. 16, 29–59. 10.1080/0269993943000121 [DOI] [Google Scholar]
- 33. Knoll, M. A. , Uther, M. , and Costall, A. (2009). “ Effects of low-pass filtering on the judgment of vocal affect in speech directed to infants, adults and foreigners,” Speech Commun. 51, 210–216. 10.1016/j.specom.2008.08.001 [DOI] [Google Scholar]
- 34. Kolly, M.-J. , Leemann, A. , and Dellwo, V. (2014). “ Foreign accent recognition based on temporal information contained in lowpass-filtered speech,” in Proceedings of Interspeech 2014, September 14–18, Singapore, pp. 2175–2179. [Google Scholar]
- 35. Kreiman, J. , and Sidtis, D. (2011). Foundations of Voice Studies: An Interdisciplinary Approach to Voice Production and Perception ( Wiley-Blackwell, Malden, MA: ). [Google Scholar]
- 36. Labov, W. (2010). Principles of Linguistic Change: Cognitive and Cultural Factors ( Wiley-Blackwell, Malden, MA: ). [Google Scholar]
- 37. Labov, W. , Ash, S. , and Boberg, C. (2006). The Atlas of North American English: Phonetics, Phonology and Sound Change ( De Gruyter Mouton, Berlin: ). [Google Scholar]
- 38. Lass, N. J. , Almerino, C. A. , Jordan, L. F. , and Walsh, J. M. (1980). “ The effect of filtered speech on speaker race and sex identifications,” J. Phon. 8, 101–112. 10.1016/S0095-4470(19)31445-7 [DOI] [Google Scholar]
- 39. Lass, N. J. , Hughes, K. R. , Bowyer, M. D. , Waters, L. T. , and Bourne, V. T. (1976). “ Speaker sex identification from voiced, whispered and filtered isolated vowels,” J. Acoust. Soc. Am. 59, 675–678. 10.1121/1.380917 [DOI] [PubMed] [Google Scholar]
- 40. Lehiste, I. , and Peterson, G. E. (1959). “ The identification of filtered vowels,” Phonetica 4, 161–177. 10.1159/000258001 [DOI] [Google Scholar]
- 41. Leung, Y. , Oates, J. , and Chan, S. P. (2018). “ Voice, articulation and prosody contribute to listener perceptions of speaker gender: A systematic review and meta-analysis,” J. Speech. Lang. Hear. Res. 61, 266–297. 10.1044/2017_JSLHR-S-17-0067 [DOI] [PubMed] [Google Scholar]
- 42. Licklider, J. C. R. , and Miller, G. A. (1951). “ The perception of speech,” in Handbook of Experimental Psychology, edited by Stevens S. S. ( Wiley, New York: ), pp. 1040–1074. [Google Scholar]
- 43. Lippmann, R. P. (1996). “ Accurate consonant perception without mid-frequency speech energy,” IEEE Trans. Speech Audio Process. 4, 66–69. 10.1109/TSA.1996.481454 [DOI] [Google Scholar]
- 44. Lynn, S. K. , and Barrett, L. F. (2014). “ ‘ Utilizing’ signal detection theory,” Psychol. Sci. 25, 1663–1673. 10.1177/0956797614541991 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45. Macmillan, N. A. , and Creelman, C. D. (2005). Detection Theory: A User's Guide ( Lawrence Erlbaum, Mahwah, NJ: ). [Google Scholar]
- 46. McNally, R. J. , Otto, M. W. , and Hornig, C. D. (2001). “ The voice of emotional memory: Content-filtered speech in panic disorder, social phobia, and major depressive disorder,” Behav. Res. Ther. 39(11), 1329–1337. 10.1016/S0005-7967(00)00100-5 [DOI] [PubMed] [Google Scholar]
- 47. Monson, B. B. , Lotto, A. J. , and Story, B. H. (2014). “ Detection of high-frequency energy level changes in speech and singing,” J. Acoust. Soc. Am. 135, 400–406. 10.1121/1.4829525 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48. Monson, B. B. , Rock, J. , Schulz, A. , Hoffman, E. , and Buss, E. (2019). “ Ecological cocktail party listening reveals the utility of extended high-frequency hearing,” Hear. Res. 381, 107773. 10.1016/j.heares.2019.107773 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49. Motlagh Zadeh, L. , Silbert, H. N. , Sternasty, K. , Swanepoel, D. W. , Hunter, L. L. , and Moore, R. D. (2019). “ Extended high frequency hearing enhances speech perception in noise,” Proc. Natl. Acad. Sci. U.S.A. 116(47), 23753–23759. 10.1073/pnas.1903315116 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50. Müsch, H. , and Buus, S. (2001a). “ Using statistical decision theory to predict speech intelligibility. I. Model structure,” J. Acoust. Soc. Am. 109, 2896–2909. 10.1121/1.1371971 [DOI] [PubMed] [Google Scholar]
- 51. Müsch, H. , and Buus, S. (2001b). “ Using statistical decision theory to predict speech intelligibility. II. Measurement and prediction of consonant-discrimination performance,” J. Acoust. Soc. Am. 109, 2910–2920. 10.1121/1.1371972 [DOI] [PubMed] [Google Scholar]
- 52. Nazzi, T. , Bertoncini, J. , and Mehler, J. (1998). “ Language discrimination by newborns: Toward an understanding of the role of rhythm,” J. Exp. Psychol. Hum. Percept. Perform. 24, 756–766. 10.1037/0096-1523.24.3.756 [DOI] [PubMed] [Google Scholar]
- 53. Norris, D. , McQueen, J. M. , and Cutler, A. (2003). “ Perceptual learning in speech.” Cogn. Psychol. 47, 204–238. 10.1016/S0010-0285(03)00006-9 [DOI] [PubMed] [Google Scholar]
- 54. Owren, M. J. , Berkowitz, M. , and Bachorowski, J. A. (2007). “ Listeners judge talker sex more efficiently from male than from female vowels,” Percept. Psychophys. 69, 930–941. 10.3758/BF03193930 [DOI] [PubMed] [Google Scholar]
- 55. Pardo, J. S. , Pellegrino, E. , Dellwo, V. , and Möbius, B. (2022). “ Special issue: Vocal accommodation in speech communication,” J. Phon. 95, 101196. 10.1016/j.wocn.2022.101196 [DOI] [Google Scholar]
- 56. Pépiot, E. (2014). “ Male and female speech: A study of mean f0, f0 range, phonation type and speech rate in Parisian French and American English speakers,” in Proceedings of Speech Prosody 2014, May 20–23, Dublin, Ireland, pp. 305–309. [Google Scholar]
- 57. Schaeffler, F. , and Summers, R. (1999). “ Recognizing German dialects by prosodic features alone,” in Proceedings of the 14th International Congress of Phonetic Sciences, August 1–7, San Francisco, CA, pp. 2311–2314. [Google Scholar]
- 58. Scherer, K. R. (2003). “ Vocal communication of emotion: A review of research paradigms,” Speech Commun. 40, 227–256. 10.1016/S0167-6393(02)00084-5 [DOI] [Google Scholar]
- 59. See, J. D. , Warm, J. S. , Dember, W. N. , and Howe, S. R. (1997). “ Vigilance and signal detection theory: An empirical evaluation of five measures of response bias,” Hum. Factors 39, 14–29. 10.1518/001872097778940704 [DOI] [Google Scholar]
- 60. Shadle, C. H. (2023). “ Alternatives to moments for characterizing fricatives: Reconsidering Forrest et al. (1988),” J. Acoust. Soc. Am. 153(2), 1412–1426. 10.1121/10.0017231 [DOI] [PubMed] [Google Scholar]
- 61. Snodgrass, J. G. , and Corwin, J. (1988). “ Pragmatics of measuring recognition memory: Applications to dementia and amnesia,” J. Exp. Psychol. Gen. 117, 34–50. 10.1037/0096-3445.117.1.34 [DOI] [PubMed] [Google Scholar]
- 62. Tabain, M. (1998). “ Non-sibilant fricatives in English: Spectral information above 10 kHz,” Phonetica 55(3), 107–130. 10.1159/000028427 [DOI] [PubMed] [Google Scholar]
- 63. Thomas, E. R. , and Reaser, J. (2004). “ Delimiting perceptual cues used for the ethnic labeling of African American and European American voices,” J. Socioling. 8(1), 54–87. 10.1111/j.1467-9841.2004.00251.x [DOI] [Google Scholar]
- 64. Van Bezooijen, R. , and Gooskens, C. (1999). “ Identification of language varieties: The contribution of different linguistic levels,” J. Lang. Soc. Psychol. 18, 31–48. 10.1177/0261927X99018001003 [DOI] [Google Scholar]
- 65. Van Leyden, K. , and Van Heuven, V. (2006). “ On the prosody of Orkney and Shetland dialects,” Phonetica 63, 149–174. 10.1159/000095306 [DOI] [PubMed] [Google Scholar]
- 66. Vitela, A. D. , Monson, B. B. , and Lotto, A. J. (2015). “ Phoneme categorization relying solely on high-frequency energy,” J. Acoust. Soc. Am. 137, EL65–EL70. 10.1121/1.4903917 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67. Werker, J. F. (2018). “ Perceptual beginnings to language acquisition,” Appl. Psycholinguist. 39, 703–728. 10.1017/S0142716418000152 [DOI] [Google Scholar]


