Abstract
Purpose:
We evaluated whether naive listeners' ratings of the gender typicality of the speech of children assigned male at birth (AMAB) and children assigned female at birth (AFAB) were different at two time points: one at which children were 2.5–3.5 years old and one when they were 4.5–5.5 years old. We also examined whether measures of speech, language, and inhibitory control predicted developmental changes in these ratings.
Method:
A group of adults (N = 80) rated single-word productions of 55 AMAB and 55 AFAB children on a continuous scale from “definitely a boy” to “definitely a girl.” Children's productions were taken from previous longitudinal study of phonological development and vocabulary growth. As part of that study, children completed a battery of standardized and nonstandardized tests at both time points.
Results:
Listener ratings for AMAB and AFAB children were significantly different at both time points. The difference was larger at the later time point, and this was due entirely to changes in the ratings of AMAB children's speech. A measure of language production and a measure of inhibitory control predicted developmental changes in these ratings, albeit only weakly, and not in a consistent direction.
Conclusions:
The gender typicality of AMAB and AFAB children's speech is perceptibly different for children as young as 2.5 years old. Developmental changes in perceived gender typicality are driven by changes in the speech of AMAB children. The learning of gendered speech is not constrained or facilitated by overall speech and language skill.
Speech is a remarkably efficient signal. A single spoken word can convey multiple types of information simultaneously. Munson et al. (2011) describe the different types of information that might be conveyed by a single production of the word [kʰæt]. This includes information about, for example, the regular semantic meaning (that the talker is referring to an object of the class felis catus and not of the class canis lupis familiaris), and information about the speakers themselves. This study is concerned with this second type of information, the social information about a talker that speech conveys. This can include macrosociological characteristics such as age, gender, social class, and race and microsociological characteristics such as membership in different local social groups (Babel & Munson, 2014).
The specific focus of the current investigation is on children's acquisition of socially meaningful phonetic variation that conveys their nascent gender and how this interacts with other aspects of language development. Munson and Babel (2019) and Tripp and Munson (2021) show that phonetic differences between men and women extend far beyond the acoustic differences between cisgender men and cisgender women that are expected to occur because of sex dimorphism in the speech production mechanism. Cisgender individuals are ones whose gender identity accords with broad social expectations of the identity should follow from the sex that was assigned at birth. As reviewed by Lieberman (1986), cisgender adult men, as a group, have longer vocal tracts and larger, more-massive vocal folds than cisgender adult women. Longer vocal tracts lead to lower resonant frequencies for sounds produced with a relatively open articulatory posture, such as vowels. Lighter, less-massive vocal folds lead to a higher rate of vibration and, hence, a higher frequency sound source (i.e., a higher f0) for speech. This sex dimorphism is most robustly present after puberty, when cisgender men's larynxes descend (Vorperian et al., 2005) and their vocal folds thicken relative to those of cisgender women (Harries et al., 1998).
Given these anatomical and physiological differences, it is not surprising that cisgender men and cisgender women differ, as a group, in their overall formant-frequency scaling and f0 (e.g., Lee et al., 1999). However, as described below, men and women also differ in many phonetic parameters that are not grounded in the sex dimorphism in the speech production mechanism that occurs in cisgender individuals. Because these differences are not grounded in sex dimorphism, they are inclusive of all men and women and not merely the subset who are cisgender.
Moreover, even canonically sex-dimorphic acoustic characteristics such as overall formant-frequency scaling and fundamental frequency (f0) are under talkers' conscious control: Although anatomy and physiology constrain absolute formant-frequency and f0 ranges, individuals can decrease and increase these with articulatory maneuvers. This has been shown in a series of studies by Cartei and colleagues (e.g., Cartei et al., 2012). In those studies, individuals were asked to emulate masculine or feminine speech. Individuals changed formant-frequency scaling and f0 to achieve different degrees of masculinity and femininity in speech.
Phonetic differences between men and women can be argued to reflect not sex dimorphism but gender. Gender is a complex social, cultural, and political construct; the full description of which is outside of the scope of this article. A few critical points about gender are relevant to the current investigation. First, gender reflects social agency. That is, gendered personae are identities that individuals claim and assert. Second, gender reflects performance (Butler, 2011). This performance involves both conscious and unconscious choices about behavior and self-presentation that convey the claimed and asserted gender identity. Third, these performances are part of participation in different communities of practice (Eckert & McConnell-Ginet, 1992).
Gendered Speech Development
If gender reflects the agentive performance of identity, and performances that are situated in communities of practice, then it is reasonable to hypothesize that development should involve learning to perform gender in ways that align with children's nascent gender. This learning should be reflected in speech and language production. Examining the speech of children provides a particularly useful test of the hypothesis that gendered speech reflects learned attributes. A long-standing assumption has been that prior to puberty; there are no anatomical differences in the vocal tracts of children assigned male at birth (AMAB) and children assigned female at birth (AFAB). Through the remainder of this article, we use the AMAB and AFAB labels to describe studies of children's speech, including this study. We reserve the terms boy/male and girl/female only for studies in which children were asked their gender and were given a variety of choices beyond just male/boy and female/girl, including an open-ended option. This reflects the contemporary understanding of gender as something that is claimed and asserted by individuals and not imposed by others.
Vorperian et al. (2011) examined vocal tract anatomy in children in a wide age range. Vorperian et al.) found differences in oral cavity horizontal length for a cohort AMAB and AFAB children aged 5–10 years old, but not in a younger cohort, or an older cohort aged 10–14 years or a younger cohort. This finding runs contrary to the assertion that AMAB and AFAB individuals' vocal tract sizes and shapes should not differ prior to puberty. However, the status of sex differences in vocal tract anatomy prior to puberty is one of active debate. Barbier et al. (2016) found no evidence of prepubertal sexual dimorphism in a different corpus and argued that the findings of Vorperian et al. were a statistical artifact due, in part, to the broad age range they studied. Barbier et al. did find sex dimorphism in pharyngeal cavity size at age 12 and in vocal tract length at age 15 years.
Even if the findings of Vorperian et al. (2011) on sex dimorphism in 5- to 10-year-old children are not a statistical artifact, any differences in the speech of AMAB and AFAB children younger than 5 years would reflect children's learning of different speech styles. This could involve learning the articulatory maneuvers needed to approximate the sex dimorphism in adults' speech, learning gendered variants that are not grounded in sex dimorphism or both. Indeed, there is evidence that AMAB and AFAB children's speech does differ. Perry et al. (2001) examined adult listeners' perception of the speech of four groups of AMAB and AFAB children, aged 4 years old, 8 years old, 12 years old, and 16 years old. Adults listened to carrier phrases containing the set of /h/−vowel−/d/ words and rated the children on a 6-point scale (1 = positively a female; 2 = appeared to be a female; 3 = unsure, may have been a female; 4 = unsure, may have been a male; 5 = appeared to be a male; 6 = positively a male). All four groups of AMAB and AFAB groups elicited different ratings. For the 4-, 8-, and 12-year-old children, mean ratings were in the 3–4 range of the scale. For the 16-year-old children, the ratings were closer to the endpoints. Presumably, this reflects the 16-year-old children having undergone puberty. Consistent with this, acoustic analyses found robust differences in f0 and formant-frequency measures for the 16-year-olds. For the three younger groups, there were no differences in f0 but reliable sex differences in the frequencies of the first three formants. Across the four age groups, formant frequencies were correlated with a variety of measures of body size, such as height, weight, and neck circumference. However, differences between AMAB and AFAB children remained even when the influence of body size was controlled statistically. This suggests that the formant differences reflect active articulatory maneuvers rather than the passive influence of vocal tract morphology. This hypothesis is supported by the findings of Cartei et al. (2014, 2019). Cartei et al. showed that children, like adults, are able to control f0 and overall formant-frequency scaling when asked to emulate the speech of men or women.
The findings of Perry et al. (2001) are consistent with findings in a variety of studies (e.g., Amir et al., 2012; Weinberg & Bennett, 1971). Still, other studies have demonstrated that AMAB and AFAB children's speech is acoustically different. Fox and Nissen (2005) showed that AMAB and AFAB children produced the fricatives /s/ and /ʃ/ differently from one another in ways that mirror the differences seen between adult men and women: AFAB children had higher spectral mean frequencies than AMAB children. This finding is particularly interesting given that it has been argued persuasively that differences between adult men and women's /s/ are unlikely to be due to sex dimorphism in the speech-production mechanism (Calder, 2019; Fuchs & Toda, 2010; Stuart-Smith, 2007; Zimman, 2017b).
Individual Differences in Gendered Speech in Children
Together, findings from studies on both the perception of gender in children's speech, on differences between AMAB and AFAB children's speech, and on children's ability to produce speech that resembles that of adult men and women, show that children learn gendered speech prior to the sex dimorphism that occurs at puberty. The findings of Perry et al. (2001) show that this learning occurs prior to the point in development at which Vorperian et al. (2011) found transient differences in oral cavity size between AMAB and AFAB children.
The mechanism by which gendered speech is learned has been speculated in previous research and is discussed in more detail in the Discussion section of this study. One of the mechanisms is socially selective learning, in which children learn aspects of language from a subset of the people they encounter during language acquisition. Socially selective learning has been demonstrated experimentally in a variety of studies (e.g., Koenig & Harris, 2005). The learning of gendered speech represents a potentially universal case of socially selective learning, which may affect both typical language acquisition and language acquisition in unusual circumstances, such as during speech-language therapy. Hence, it is vitally important to understand the factors that affect the learning of gendered speech and how it interacts with other aspects of language acquisition.
One way to examine factors that affect the learning of gendered speech is to examine whether individual differences in the development of gender are associated with differences in the learning of gendered speech. Two studies examined this. Munson et al. (2015) used the rating scale of Perry et al. (2001) to examine ratings of the voices of two groups of prepubertal AMAB children. One of these groups was identified by a psychologist as having gender identity disorder (GID), a no longer used diagnostic label that is defined by demonstrating a variety of behaviors that are stereotypically not associated with cisgender boys. The other group comprised cisgender boys matched in age. The children in the gender-nonconforming GID group were rated as sounding less prototypically boy-like than the cisgender boys. Li et al. (2016) examined the acoustics of /s/ in a wide age range of AMAB and AFAB children aged 4–16 years. Li et al. found that AMAB children's /s/ acoustics varied as a function of their performance on a questionnaire designed to measure gender identity. Children whose performance was more consistent with those of cisgender boys had /s/ productions that resembled those of adult men. This was true even when acoustic estimates of vocal tract length were controlled statistically. Together, Munson et al. (2015) and Li et al. show that individual differences in gender identity predict individual differences in gendered speech in children.
Another potential source of individual differences in gendered speech is individual differences children's overall linguistic skill: children with stronger language skills might be better able to learn sociolinguistic variation than those whose are less proficient at language. Evidence for this hypothesis comes from cross-sectional studies of the acquisition of sociolinguistic variation. Li (2017) examined the acquisition of a gendered variant of /ɕ/ in varieties of Mandarin spoken in Northern Mainland China. The gender marking in /ɕ/ serves to enhance the difference between /ɕ/ and the similar fricative /ʂ/. Li found that 2-year-olds' productions of /ɕ/, /ʂ/, and the third sibilant fricative /s/ were acoustically undifferentiated from one another. For 3-year-olds, there was evidence of differentiation among the three fricatives but not of gender marking. Evidence of gender marking was found for 4- and 5-year-old AFAB children. Critically, this gender marking happened only in age cohorts who had already acquired differences between the sounds that are subject to gender marking. Li's finding suggests that gender marking should lag behind other aspects of language acquisition. We hypothesize that individual differences in language acquisition might lead to individual differences in the acquisition of gender marking. Children whose early language performance is substantially above their peers should also acquire socially meaningful linguistic variation earlier than those peers with lower language scores.
This Study
This study further explores the acquisition of gendered speech by examining listener ratings of the gender typicality of children's speech. This study makes many unique contributions to this literature. One is that we examine this question longitudinally at two time points: one where children's average age is 3 years and one where it is 5 years. Nearly all of the previous studies on this topic have examined cross-sectional data. Only one recent study, Fung et al. (2021), examined the acquisition of gendered speech longitudinally. Fung et al. examined listeners' perception of the speech of six AMAB and six AFAB children at age 2.5, 4, and 5.5 years, using productions from a publicly available longitudinal corpus of children's speech. Listeners were asked to classify children as male or female. They found that listeners rated the two groups differently even at the earliest time point, when children were younger than in any published study, other than a conference presentation of the data in this article. This study looks at listener ratings in a larger group of children.
Moreover, the children in the current investigation participated in a longitudinal study of phonological development and vocabulary growth. As part of their participation in that study, they completed a number of standardized measures of vocabulary size, speech-production accuracy, speech perception, and executive function, at each of the time points. This allows us to evaluate our hypothesis that early speech and language performance predicts ratings of gendered speech at a later time point.
Given the relatively exploratory nature of this analysis, we examine a broad set of speech and language skills. These are not intended to represent the range of skills that contribute to the learning of gendered speech but instead to represent the range of measures that are conventionally made in speech and language assessments and research studies. They are limited to measures that were given in a study whose focus was on the relationship between phonological development and vocabulary growth, rather than on the development of gendered speech. Despite this weakness, each one of these measures is plausibly related to gendered speech development. We examine speech perception based on the assumption that children with stronger speech perception would be better able to detect gendered speech variants. Indeed, there is evidence that children are able to perceive the alignment between visual and phonetic cues to gender early in life (Patterson & Werker, 2002). We examine speech production accuracy based on the assumption that children with better articulation would be more able to manipulate phonetic detail to convey gender, as was shown by Li (2017). We examine vocabulary size based on the observation that vocabulary size predicts a wide range of phonological behaviors seemingly unrelated to the lexicon (e.g., Beckman et al., 2007). We examine the volume of expressive communication from day-long recordings because of the finding that greater practice is associated with better fine phonetic control (Cychosz et al., 2021), which might lead to greater fluency in producing socially meaningful phonetic variation. We examine inhibitory control under the assumption that superior skill in this area might be associated with better switching among different linguistic codes, as has been found previously in studies of bilingual individuals (e.g., Borragan et al., 2018). Again, we acknowledge that the relationship between each of these measures and gendered speech development is likely to be complex. For example, inhibitory control is only one component of the broader skill of cognitive flexibility (Deák, 2003), and other components of cognitive flexibility might be more relevant for learning sociolinguistic variation than inhibitory control per se.
This study allows us to examine a number of other important questions about the acquisition of gendered speech that have not yet been addressed. First, it allows us to explore the different acoustic characteristics of speech that are associated with differences in ratings. Second, it allows us to examine whether any developmental changes are driven more by the behavior of AMAB children or AFAB children. One possible scenario is that development involves AFAB children sounding progressively more girl-like and AMAB children sounding progressively more boy-like. It is also possible that development could be driven primarily by one of the two groups. Given the sex dimorphism in adult vocal tracts, AFAB children's vocal tracts are more similar to those of adult cisgender women than AMAB children's vocal tracts are to adult cisgender men. Hence, the phonetic maneuvers required for a child to sound like an adult man are likely to be more extreme than those to sound like an adult woman. This may mean that the learning of male-like phonetic variation lags that of learning female-like phonetic variation, and that developmental changes would be driven primarily by the behavior of AFAB children.
Finally, this study allows us to examine how the perception of gender in children's voices interacts with the perception of their age. The growth of the vocal tract described by Vorperian et al. (2005) has predictable effects on acoustics. As shown by Lee et al. (1999), formant frequencies and f0 decrease over development. Hence, f0 and formant frequencies are potentially useful cues to a child's age. Sex dimorphism in the speech production mechanism also leads to differences in f0 and formant frequencies. Part of learning gendered speech may involve emulating these differences, and adult raters may estimate children's gender, in part, based on these acoustic parameters. In short, the perception of age and gender in children's speech might be tightly interwoven.
There has been relatively little research on the interaction between perceived age and perceived gender of children's voices. Plummer et al. (2013) showed a trade-off between these parameters in the perception of a synthetic voice designed to reflect the average vocal tract size and f0 of 10-year-old children: Some listeners perceived the voice as adult and female, and others perceived it as adolescent and male. If the findings of Plummer et al. are replicated in this study, then it is possible that a difference in perceived gender between AFAB and AMAB children is due to a difference in perceived age, with the AFAB children being perceived as older than the AMAB children. Barreda and Assmann (2018) examined the perception of age as a function of whether participants were told that they were listening to an AMAB or AFAB child. Their study modeled predictions of actual and perceived age of speech samples from a wide age range of AMAB and AFAB children, in tasks where raters were told whether they were listening to an AMAB or AFAB child, and ones in which it was not. In contrast to the findings of Plummer et al., the models predicting perceived age were qualitatively similar regardless of whether listeners were told whether a child was AMAB or AFAB. This finding suggests that raters were implicitly perceiving gender and accounting for it in their age perception. By collecting both perceived gender and perceived age in this study, we are able to test whether perceived gender ratings are robust even when the perceived age of the same children's voices is controlled statistically.
In sum, our research questions are as follows:
Do naive listeners provide different ratings for AMAB and AFAB children's speech on a scale from “definitely a boy” to “definitely a girl” for children at age 3 years and age 5 years? What characteristics of the stimuli being rated (both acoustic characteristics and transcribed accuracy) predict listener ratings?
Does the difference in ratings of AMAB and AFAB children's speech on this gender scale change between age 3 years and 5 years? If so, do the changes apply equally to the two groups?
Do individual children's language and speech abilities at age 3 years predict the ratings at age 5 years?
Are ratings of the perceived gender of children's speech robust when the same listeners' perception of children's age is controlled statistically?
Method
Participants
There were two groups of participants in this study, each of which is described separately. The Child Talkers refers to the 110 children whose productions served as stimuli in the perception study. The Adult Listeners refers to the 80 adults who rated the children's productions in the gender-perception task. The procedures used to collect data from the children were determined by the institutional review boards (IRBs) of the University of Wisconsin (Protocol Number SE-2011-0030) and the University of Minnesota (Study Number 1103S97246) to conform to U.S. regulations regarding the protection of human participants in research. The procedures used to collect ratings of the gender of children's speech were determined by the University of Minnesota IRB to conform to the relevant regulations (Study Number 0708S1506). All participants provided informed consent prior to participating.
Child Talkers
The first group is composed of 110 children (55 AFAB and 55 AMAB) who were recruited through word of mouth, online advertisements, newspaper bulletins, flyers placed in day care centers, schools, and other local gathering locations. These children were originally involved in a longitudinal study investigating speech perception, speech production, and word learning by children with varying socioeconomic statuses (SESs), stages of speech and language development (including typically developing and late talkers), and hearing statuses (cochlear implant users). Summaries of some of the primary research questions of the original study can be found in Munson et al. (2021), Cychosz et al. (2021), and Erskine et al. (2020). While 164 children were included in the longitudinal study, 110 were selected for this study, with each child AMAB being matched to a child AFAB within 3 months of age. Only children who were typically developing, without hearing impairments, and who could be matched in age in AMAB–AFAB pairs were selected from the larger sample of 164. The children included in this study represent a range of SES, and all children were native speakers of a variety of English. Through parent–child observations and parent interviews, dialect was determined for each child. The children's genders were provided by their primary caregivers, who reported the child as male, female, or other. An explicit distinction between sex and gender was not made to these parents, as gender was not a primary focus of the original study. We imagine that the male/female designations reflect largely the parents' expectations given their child's sex assigned at birth. Hence, we refer to this variable as sex assigned at birth throughout the analysis section.
At each of the three time points in this longitudinal study, each child completed a battery of standardized and experimental measures of speech production, speech perception, phonological processing, vocabulary knowledge, phonological awareness, and hearing ability, across two to three sessions. The standardized and nonstandardized measures that the children completed at the first time point included estimates of vocabulary size, expressive language, speech production, speech perception, and inhibitory control. Vocabulary size was measured by the Peabody Picture Vocabulary Test–Fourth Edition (Dunn & Dunn, 2007, which measures receptive vocabulary knowledge) and the Expressive Vocabulary Test, Second Edition (EVT-2, Williams, 2007, which measures expressive vocabulary knowledge). Speech production accuracy was measured by the Goldman-Fristoe Test of Articulation-2 (Goldman & Fristoe, 2000). Speech perception was measured by a novel minimal pair identification task described in Erskine et al. (2020). Briefly, in that task, participants were presented with two pictures representing words that differ in a single phoneme (i.e., moon and man). Each picture was first presented individually along with a spoken prompt of its name. The pictures were then presented together, and a different prompt of one of the words was presented. Children responded by pointing to the correct picture. By naming the pictures in advance, the task minimized the effect of lexical knowledge on performance. Despite the seeming simplicity of this task, performance on it varied widely. Inhibitory control was measured by a developmentally appropriate Stroop task (Zelazo et al., 2013), in which children were asked to point to a series of pictures alternating between larger and smaller fruits. Successful performance on this task required inhibition of the previous response. Expressive language was measured using the Language ENvironment Analysis (LENA) system's estimates of child vocalizations and conversational turns (Cristia et al., 2020; Xu et al., 2012). At the last time point, vocabulary size was estimated by the EVT-2. Speech production was measured using the GFTA-2. Speech perception was measured using the Speech Analysis and Interactive Learning System (Rvachew, 1994). Nonverbal processing ability was measured using the matrices section of the Kaufman Brief Intelligence Test-2 (Kaufman & Kaufman, 2014).
Each session was 90- to 120-min long. The first time point took place when the children were between the ages of 28 and 39 months old. The second time point took place when the children were 40–52 months old. The final time point took place when the children were between the ages of 53 and 66 months old. The wide age range at each time point was part of the design of the original study.
For this study, we focus on data collected at the first and third time point. In hopes of eliminating confusion due to the absence of the second time point in our analyses, the first time point will be referred to as the first time point (FTP), and the third time point will be referred to as the last time point (LTP). The latter term was also used by Munson et al. (2021), who analyzed data from the second and third time points. Mean age for the children at FTP was 32.5 months (SD = 3.6). Mean age for LTP was 56.7 months (SD = 3.9). The difference in age between AMAB and AFAB was not statistically significant at either FTP (t[108] < 1, p = .75) or LTP (t[108] < 1, p = .85). In order to limit the length of the gender-rating experiments, the 110 children were divided into four groups of approximately equal size, each of whom was rated in a separate experiment. These groups did not differ significantly in their age at either time point, as assessed by a single-factor, between-subjects analysis of variance, F(3, 306) = 1.286, p = .28, for FTP, F(3, 306) = 1.046, p = 0.38, for LTP. The groups also did not differ in the proportion of AMAB and AFAB children, as assessed by a chi-square difference test (χ2 [df = 3] = .279, p = .964).
Adult Listeners
Eighty adults served as listeners in the gender-perception task, 20 in each of the four experiments. They were between the ages of 18 and 40 years. All of the listeners self-identified as native English speakers (defined as having learned English from birth from at least one parent who spoke English natively and who communicated with them in English from birth) with no history of hearing impairment, or language/learning disorder. They passed a pure-tone hearing screening bilaterally at 25 dB HL.
The Adult Listeners were recruited through posters that were dispersed around the University of Minnesota, Twin Cities campus, from personal contacts, and by word of mouth. On average, the task lasted approximately 45 min.
Stimuli
The stimuli were eight-word productions by each of the 110 child talkers, four each from FTP and LTP of the longitudinal study, as described above. The production data, in this study, were taken from a Real Word Repetition (RWR) task in the longitudinal study. The RWR task was designed to elicit a large number of productions of target sounds to assess growth in articulatory control in the preschool years. In the RWR task, children saw pictures of familiar words while hearing an auditory prompt naming that picture. Children were instructed to repeat the familiar word. At FTP, multiple productions of target words were collected, as only a small set of words met the criteria for being included in this study. These included that the word has an age of acquisition of 2 years or earlier and be picturable. Moreover, the word list was designed to measure children's acquisition of different phoneme contrasts (i.e., /t/ vs. /k/ and /s/ vs. /ʃ/), and the contrasts also had to be elicited in balanced vowel contexts, which limited the choice of words. At LTP, more different words were chosen, as the age of acquisition criterion was for 4 years old, which allowed for a larger set of candidate words.
Two sets of production prompts were used. One was produced by a White woman who was born in Minnesota and who spoke the local mainstream dialect, abbreviated MAE (Mainstream American English). The other was produced in African American English (AAE) by an African American woman who was born in Wisconsin and who regularly code-switches between the local regional mainstream spoken in her birth city and the variety of African American English spoken there. The use of different dialect prompts was motivated by goals of the original study. The production prompts were played in the home dialect, as assessed by the experiment coordinator. Of the 110 children in this study, seven received the words in AAE and 103 in MAE.
Words used at FTP were cookie, scissors, shovel, and garbage. Words used at LTP include cookie, summer, shovel, and rabbit. Cookie and shovel were chosen for both time points. The other words were unique to either FTP or LTP. These specific words were chosen from the longer word list, because they contain a variety of vowels and word-initial consonants (including /s/, which has been showing previously to cue listener judgments of gender), and because they do not have a strong association with cisnormative notions of gender.
A research assistant chose one exemplar of each word at each time point, after listening to all possible recordings of a word at that time point for each child. The recordings were selected based on the researcher's judgment that the words were produced as accurately as possible (i.e., if multiple trials of a word were available, the one that was perceived to be more accurate was chosen) and that they were representative of the child's normal mode of production (i.e., words that were produced in “silly” voices were not included). Furthermore, recordings where the child was speaking over the experimenter or shouting the word were not selected. The recording was then edited to contain only the child's production of the word, absent of the eliciting prompts.
Each stimulus word was transcribed and rated for articulatory accuracy. Measuring accuracy allows us to examine whether the stimuli followed the pattern observed in many developmental studies (such as the normative data for the GFTA-2), in which AFAB children produce words more accurately than age-matched AMAB children. Two undergraduate students, both of whom had both completed coursework in phonetics, served as transcribers. They were asked to listen to each sound in the word and rate it as accurate, inaccurate, or intermediate. The use of the “intermediate” category is based on a recommendation by Stoel-Gammon (2001) as a solution for coding children's productions that sound mildly distorted but which are not clearly substitution errors. The use of intermediate productions is regularly used in the phonetic transcription protocol for the laboratory in which this research took place, including other transcriptions of these children's speech (e.g., Munson et al., 2021).
Three acoustic measures of the stimuli were taken, so that we could evaluate which acoustic measures predict gender ratings. These were not meant to be exhaustive. Instead, we picked two measures that have been shown to mark gender in adults and have no basis in sex dimorphism (two acoustic measures of /s/) and one measure that is strongly sex dimorphic in adults (f0). These were made using Praat (Boersma, 2001). The first of these were acoustic measures of the word-initial /s/ in scissors at FTP and summer at LTP. We examined the first moment, spectral centroid, and the second moment, spectral variance, for the middle 40-ms interval of frication. These two measures have been shown previously to predict judgments of the gender typicality of AMAB children's speech (Munson et al., 2015). We examined the f0 of the midpoint of the vowel in the first (or, for monosyllables, sole) syllable of the word, which was the stressed syllable. In cases where the automatic pitch-tracking algorithm failed because of nonmodal voice quality, the f0 was estimated using the average of three hand-measured pitch periods in the approximate center of the vowel.
Procedure
The rating task took place during a single session. The experiment comprised a gender-rating block and an age-rating block. The gender-rating block always preceded the age-rating block. There were two sections to the gender-rating block. As in Perry et al. (2001), the experiment was blocked by talker age. In each of these sections, all of the stimuli from either FTP or LTP were presented. Listeners were told that they were listening to children of the same approximate age during each of the sections. The order of two sections was randomized across participants. On each trial, the listeners were told the word that the child was saying (“listen to the child say the word [X]”). Listeners rated by clicking on a double-headed arrow on the computer screen. Click location was logged automatically. The line was 1,000 pixels long, so there were 1,000 potential unique responses to each token. The text “definitely a boy” appeared above the left edge and the text “definitely a girl” appeared above the right edge. Participants were instructed to use the entire scale. The use of a continuous rating scale allowed us to elicit finer grained data than a binary judgment of “male” or “female.” The raw ratings in pixels were transformed to the proportion of the length of the line, and these were used for analysis.
In the age-rating section of the experiment, participants were presented with a single block with both FTP and LTP stimuli in it. Only the words shovel and cookie were included in this block. Only two words were used in this block of the experiment to keep the overall length manageable. After hearing each word, participants judged the age on a 9-point scale (1 = less than 2.5 years, 2 = 2.5–3 years old, 3 = 3–3.5 years old, 4 = 3.5–4 years old, 5 = 4–4.5 years old, 6 = 4.5–5 years old, 7 = 5–5.5 years old, 8 = 5.5–6 years old, and 9 = more than 6 years old).
Results
Stimulus Characteristics
The Appendix describes in detail the statistical analyses of the measures we took of the stimuli. In sum, there were ample cues to time point (and hence to age) in the accuracy and acoustic characteristics of the stimuli. There were considerably fewer cues to sex assigned at birth: AMAB and AFAB children produced words with similar accuracy and with similar f0s. Only /s/ showed differences between the two groups.
Standardized Test Performance
Children's performance on the measures of vocabulary, expressive language, speech production, speech perception, and inhibitory control are shown in Table 1. The standard deviations in this table show that there was a wide range of performance on these measures. Mean performance on the measures of vocabulary size were approximately 1 SD above the mean. Although the LTP speech and language measures are not used in any analyses in this article, we use them here to examine whether the groups' speech and language abilities are comparable at that time point.
Table 1.
Standardized test performance.
| Time point | Test | AMAB mean (SD) | AFAB mean (SD) |
|---|---|---|---|
| First time point | Expressive Vocabulary Test, Second Edition a Standard Score | 116.4 (14.6) 3 | 118.3 (16.9) 1 |
| Peabody Picture Vocabulary Test–Fourth Edition b Standard Score | 114.5 (17.1) 3 | 113.0 (16.3) 1 | |
| LENA Conversational Turn Count c | 55.3 (24.6) 3 | 53.5 (23.9) 3 | |
| LENA Child Vocalization Count c | 222.8 (89.7) 3 | 222.6 (88.5) 3 | |
| Goldman-Fristoe Test of Articulation-2 d Standard Score* | 96.5 (10.6) 6 | 90.2 (13.2) 2 | |
| Inhibitory Control e Raw Score | 2.2 (0.7) 1 | 2.1 (0.7) 2 | |
| Minimal Pair Perception f Percent Correct | 73.3% (14.9%) 5 | 67.8% (18.8%) 3 | |
| Last time point | Expressive Vocabulary Test, Second Edition a Standard Score | 120.1 (15) 1 | 117.4 (13.5) 1 |
| Goldman-Fristoe Test of Articulation-2 d Standard Score | 92.2 (11) 3 | 91.9 (12.9) 3 | |
| Kaufman Brief Intelligence Test-2 g Matrices Subtest Standard Score | 106.7 (14.2) 2 | 107.7 (10.9) 2 | |
| SAILS h Percent Correct | 71.3% (9.8%) 3 | 73.6% (10.9%) 4 |
Note. AMAB = assigned male at birth; AFAB = assigned female at birth; LENA = Language ENvironment Analysis; SAILS = Speech Analysis and Interactive Learning System.
N = 55.
N = 54.
N = 53.
N = 52.
N = 50.
N = 49.
Difference between AMAB and AFAB children significant in an independent-samples t test, t(99.6) = 2.66, p = .009.
A series of independent-samples t tests examined whether the AMAB and AFAB children differed on the six FTP and four LTP measures. Of the 10 comparisons in Table 1, only one achieved statistical significance: The AFAB children's standard scores on the GFTA-2 at FTP were significantly poorer than those of the AMAB children (p = .009). This was true when the uncorrected α level of 0.05 was used but not when the level was corrected for multiple comparisons. Interestingly, the raw scores on the GFTA-2 did not differ significantly. The difference occurred only because different norms are used for AFAB and AMAB children in the age range at FTP. The standard scores on the GFTA-2 at LTP did not differ between AMAB and AFAB children.
Perceived Gender
Linear mixed-effects models were used to examine the perceived gender ratings, using the lme4 package in R (Bates et al., 2015). Significance tests were calculated using the LMERTest package (Kuznetsova et al., 2017). Degrees of freedom were estimated using Satterthwaite's approximation. The dependent measure was click location in the gender-rating task, transformed to the length of the rating scale, where 0 was the end of the line labeled “definitely a boy” and 1 was the end of the line labeled “definitely a girl.” Models were built progressively, starting with a base model (model g0) that had random intercepts for listener and item (i.e., unique combinations of talker, time point, and word). The next model, Model g1, added a fixed effect of time point (contrast-coded) and a random slope for the effect of time point on listener. Model g2 had fixed effects of time point and sex assigned at birth (contrast coded) and random slopes for the effects of time point and sex assigned at birth on listener. Model g3 had fixed effects of time point, sex assigned at birth, and their interaction, and random slopes for the effect of these on listeners. The significance of each model was determined both by whether it resulted in a better fit than the next-simplest model, and whether the coefficients for the fixed effects in the more complex model were significant.
Models g1, g2, and g3 all fit the data better than the next-simplest model (Model g0–g1: χ2 [df = 2] = 46.2, p < .001; Model g1–g2: χ2 [df = 2] = 161.6, p < .001; Model g2–g3: χ2 [df = 2] = 10.1, p = .006). The coefficients for Model g3 are shown in Table 2. As this table shows, the coefficients for sex assigned at birth and for the interaction between sex assigned at birth and time point were robustly significant.
Table 2.
Model predicting the full set of gender ratings from fixed effects of sex assigned at birth (SAB), time point (TP), and their interaction. The model also had random intercepts for participants and items and a random slope for the effects of SAB, TP, and their interaction on individual listeners' ratings.
| Estimate | Estimate | Standard error | Degrees of freedom | t value | Pr(> |t|) |
|---|---|---|---|---|---|
| (Intercept) | .530 | .007 | 317.2 | 80.611 | < .001 |
| SAB (AMAB = −1, AFAB = 1) | .054 | .006 | 629.2 | 9.68 | < .001 |
| TP (FTP = −1, LTP = 1) | −.010 | .006 | 692.3 | −1.762 | .070 |
| SAB × TP | .014 | .006 | 772.9 | 2.73 | .006 |
Note. Pr = probability; AMAB = assigned male at birth; AFAB = assigned female at birth; FTP = first time point; LTP = last time point.
Figure 1 shows boxplots of the mean ratings for each child, separated by sex assigned at birth and time point. The median ratings of the full set of AFAB children were higher than those for the AMAB children at both time points, indicating that they were perceived to be more girl-like. The difference was larger at LTP than FTP.
Figure 1.

Mean gender ratings for each child, separated by sex assigned at birth and time point. FTP = first time point; LTP = last time point; AFAB = assigned female at birth; AMAB = assigned male at birth.
Variants of Models g0 and g1 were run separately for AFAB and AMAB children only to examine whether ratings for both groups changed from FTP to LTP. For AFAB children, Model g1 did not fit the data better than Model g0 (χ2 [df = 1] = 3625, p = .547). In contrast, for AMAB children, Model g1 fits the data better than Model g0 (χ2 [df = 1] = 11.177, p < .001), and the coefficient for time point was significant in the AMAB model g1. Together, this shows that the interaction between sex assigned at birth and time point occurred because ratings for AMAB children became progressively more boy-like from FTP to LTP, but ratings of AFAB children did not change.
The next analysis examined whether the effects of sex assigned at birth and time point were robust when we controlled for the effect of stimulus accuracy, f0, and /s/ acoustics on ratings. By comparing the effect, if any, of f0 and /s/ acoustics on listener ratings, we could assess whether listeners were attending to the cue that was most robustly present in the stimuli (/s/ acoustics, especially /s/ spectral centroid). We could also assess whether listeners were relying on cues to gender in children's speech that were not present in these stimuli (accuracy), or cues to gender in adults' speech, which were not present in these stimuli (f0). In order to compare across predictors whose scales were different, these were z-transformed prior to being added to models.
For the analysis of the influence of f0 on ratings, we first fit a baseline model (Model f 0 _0) with only random intercepts for listener and item. Model f 0 _1 also included a fixed effect for f0 and random slopes for the effect of f0 on individual listeners. Model f 0 _2 also included an interaction between sex assigned at birth and time point and random slopes for the effect of these on individual listeners. Model assessment proceeds as it did with Models g0 through g2. This allowed us to examine whether time point and sex assigned at birth predicted ratings beyond what was already predicted by f0. Each model fits the data significantly better than the next simplest model (Model f0_0–Model f0_1: χ2 [df = 2] = 241.08, p < .001; Model f0_1–Model f0_2: χ2 [df = 6] = 187.95, p < .001). The coefficients for sex assigned at birth, time point, and their interaction in Model f0_1 were similar in significance, size, and direction to those in Model g2. The coefficient for f0 was also significant (β = 0.017, standard error of the mean [SEM] = 0.007, t[435.8] = 2.664, p = .008). The positive coefficient showed that higher f0 was associated with judgments closer to the “definitely a girl” end of the rating scale.
The analysis of the effect of stimulus accuracy on ratings paralleled that for f0, in that three models (model accuracy_0, model accuracy_1, model accuracy_2) were fit and compared with one another. Each model fit the data better than the next-simplest model (model accuracy_0–model accuracy_1: χ2 [df = 2] = 34.886, p < .001; model accuracy_1–model accuracy_2: χ2 [df = 6] = 218.73, p < .001). In model accuracy_2, the coefficients for sex assigned at birth, time point, and for the interaction between sex assigned at birth and time point were all significant and in the same direction and approximate magnitude as in Model g3. The coefficient for accuracy was also significant (β = 0.020, SEM = 0.006, t[739.2] = 3.337, p = .001). The positive coefficient showed that more accurate productions were associated with judgments closer to the “definitely a girl” end of the rating scale.
The analysis of the effect of /s/, although based on the smaller set of tokens and conducted separately for spectral centroid and spectral variance, was otherwise parallel to that for f0 and accuracy, in that three models (model {centroid,variance}_0, model {centroid,variance}_1, model {centroid,variance}_2) were fit and compared with one another. For spectral variance, each model fits the data better than the next-simplest model (model variance_0–model variance_1: χ2 [df = 1] = 16.602, p < .001, model variance_1–model variance_2: χ2 [df = 6] = 442.17, p < .001). In model variance_2, the coefficients for sex assigned at birth, time point, and their interaction were all significant and in the same direction and of the same approximate size as in Model g2. However, the coefficient for spectral variance was not significant. For spectral centroid, each model fit the data better than the next-simplest model (model centroid_0–model centroid_1: χ2 [df = 1] = 157.76, p < .001, model centroid_1–model centroid_2: χ2 [df = 6] = 379.98, p < .001). In model centroid_2, the coefficients for sex assigned at birth, time point, and their interaction were all significant and in the same direction and of the same approximate size as in Model g2. The coefficient for centroid was also significant (β = 0.038, SEM = 0.004, t[430.9] = 9.019, p < .001). The positive coefficient showed that more-accurate productions were associated with judgments closer to the “definitely a girl” end of the rating scale.
In sum, the analysis of the effect of stimulus characteristics on ratings showed that stimulus characteristics did not explain the ratings entirely: Even when accuracy, /s/ acoustics, and f0 were included in models predicting gender ratings, there were still effects of sex assigned at birth, time point, and their interaction. Stimulus characteristics predicted ratings regardless of whether the AFAB and AMAB children differed in that characteristic. However, the strongest characteristic, as indicated by the size of the β coefficient, was the characteristic that did differ between the groups, /s/ centroid.
Perceived Age
LMER was used to examine whether age ratings differed as a function of sex assigned at birth, time point, and their interaction. This was assessed using methods parallel to those used to examine gender ratings. A series of four models (Models a0–a3) were fit. Models a1 and a2 fit the data better than the next-simplest model (Model a0–a1: χ2 [df = 2] = 693.4, p < .001; Model a1–a2: χ2 [df = 2] = 6.33, p = .04). Model a3 did not converge. When the random slope estimating the influence of the interaction between sex assigned at birth and time point on individual listeners was removed, the model did converge. That model did not fit the data better than Model a2 (χ2 [df = 1] = 1.49, p = .22). The coefficients for Model a2 are shown in Table 3.
Table 3.
Model predicting age ratings from fixed effects of sex assigned at birth (SAB) and time point (TP) and their interaction.
| Estimate | Estimate | Standard error | Degrees of freedom | t value | Pr(> |t|) |
|---|---|---|---|---|---|
| (Intercept) | 4.441 | .100 | 118.2 | 44.431 | < .001 |
| SAB (AMAB = −1, AFAB = 1) | −0.120 | .048 | 408.0 | −2.479 | .014 |
| TP (FTP = −1, LTP = 1) | 0.945 | .065 | 247.6 | 14.637 | < .001 |
Note. The model also had random intercepts for participants and items and random slopes for the effect of SAB and TP on individual listeners’ ratings. Pr = probability; AMAB = assigned male at birth; AFAB = assigned female at birth; FTP = first Time Point; LTP = last time point.
As this table shows, the coefficients for sex assigned at birth and for the interaction between sex assigned at birth and time point were robustly significant. Figure 2 shows boxplots of the mean ratings for each child, separated by sex assigned at birth and time point.
Figure 2.

Mean age ratings for each child, separated by sex assigned at birth and time point. FTP = first time point; LTP = last time point; AFAB = assigned female at birth; AMAB = assigned male at birth.
As this figure shows, the AMAB children were rated to sound older than the AFAB children at both time points, and both groups were rated to sound older at the second time point. The age range of the children is associated with ratings in the 1–3 range of the scale at FTP and the 7–9 range at LTP. Relatively few of the children's mean ratings fell within these ranges. Most notably, the effect of sex assigned at birth on perceived age does not suggest that the perceived gender findings in the previous section are a consequence of group differences in perceived age. If that were the case, the AFAB children would have been rated as older than the AMAB children.
Predicting Perceived Gender From Perceived Age
To further test the (in)dependence of perceived gender and perceived age, we conducted an analysis of whether sex assigned at birth and time point predicted gender ratings when effects of perceived age were controlled statistically. This analysis used the gender ratings for the cookie and shovel stimuli, as these were the only stimuli for which there were both gender and age ratings at both time points. The dependent measure was gender ratings. Three models were built. Model ag0 had only random intercepts for item and listener. Model ag1 also included a fixed effect of perceived age, and a random slope for the effect of perceived age on individual listeners. Model ag2 also included sex assigned at birth, time point, and their interaction and random slopes for the effect of time point and sex assigned at birth on listeners. A model that also included a random slope for the influence of the interaction between time point and sex assigned at birth on individual listeners did not converge. Model ag1 did not fit the data better than the next-simplest model at the α = 0.05 level but did achieve significance at the less stringent α = 0.06 level (Model ag0–ag1: χ2 [df = 1] = 3.596, p = .058). The coefficient associated with perceived age in Model ag1 was also significant only at the less stringent α = 0.06 level (β = −0.003, SEM = 0.001, t[711.5] = −1.906, p = .057). The negative value shows that the older children were the more boy-like they were rated. Model ag2 fit the data better than Model ag1 (χ2 [df = 6] = 78.03, p < .001). In Model ag2, the coefficients for sex assigned at birth, time point, and their interaction were significant, and the direction was the same as in Model g3, shown in Table 3. The coefficient for perceived age was not significant (p = .14). This analysis shows that the effects of sex assigned at birth on perceived gender were not confounded by perceived age.
Predicting LTP Gender Ratings
The final analysis examined whether the FTP speech and language measures summarized in Table 1 predicted gender ratings at LTP. The first of these examined whether children's speech and language abilities at FTP predicted gender ratings of their speech at LTP. In order to compare ratings for the AMAB and AFAB children in a single analysis, we subtracted the ratings for the AMAB children from 1, so that we could compare all 110 children's ratings relative to the cisgender standard for their sex assigned at birth. This derived rating tracks children's adherence to cisgender norms for speech. The assumption that all of the children in this study are cisgender is likely incorrect, given population studies of gender development. We return to this inherent weakness of this study in the discussion. It does, however, maximize statistical power, as all 110 children are examined in a single analysis.
These derived gender ratings at LTP were the dependent measure in a series of models. Our goal was to assess whether FTP speech and language abilities predicted these ratings beyond derived gender ratings at FTP. Hence, we calculated the average derived gender ratings for the FTP measures for each listener–talker combination. For example, for each of the four derived ratings of a given talker by a given listener at LTP, the average derived rating of that same talker by that same listener at FTP served as a fixed effect. This allowed us to control for the effect of perceived gender of that talker by that listener at FTP on the LTP ratings. In a base model (model prediction_0), there were random effects for listener and item. In the next model (prediction_1), there was a fixed effect of FTP ratings, as described above. In the final model (prediction_2), there were additional fixed effects for the seven FTP speech and language measures shown in Table 1.
Prior to constructing model prediction_2, missing data from the seven predictor variables were imputed using the R package MICE (Multiple Imputation by Chained Equations; Van Buuren & Groothuis-Oudshoorn, 2010). Data imputation was supported by the finding that Little's Missing Completely at Random test (Little, 1988) determined that the missing data were randomly missing, χ2 [df = 34] = 42, p = .84.
Each of these models fits the data better than the next-simplest model (model prediction_0–model prediction_1: χ2 [df = 1] = 152.86, p < .001, model prediction_1-model prediction_2: χ2 [df = 7] = 26.76, p < .001). The coefficients for model prediction_2 are shown in Table 4.
Table 4.
Model predicting the derived last time point (LTP) ratings from each listener's average rating for each talker at first time point (FTP) and from FTP speech and language scores.
| Estimate | Estimate | Standard error | Degrees of freedom | t value | Pr(> |t|) |
|---|---|---|---|---|---|
| Intercept | .570 | .008 | 414.2 | 75.911 | < .001 |
| FTP gender ratings | .037 | .003 | 8767.0 | 12.402 | < .001 |
| LENA Child Vocalization Count | .037 | .012 | 429.5 | 3.190 | .002 |
| LENA Conversational Turn Count | −.048 | .012 | 429.5 | −4.119 | < .001 |
| Inhibitory Control Raw Score | −.020 | .008 | 429.6 | −2.391 | .017 |
| Minimal Pair Perception Percent Correct | −.006 | .008 | 430.8 | −0.731 | .465 |
| Goldman-Fristoe Test of Articulation-2 Standard Score | −.005 | .009 | 428.1 | −0.635 | .526 |
| Expressive Vocabulary Test, Second Edition Standard Score | −.002 | .011 | 428.0 | −0.223 | .824 |
| Peabody Picture Vocabulary Test–Fourth Edition Standard Score | .017 | .010 | 427.7 | 1.653 | .099 |
Note. The model also had random intercepts for participants and items. Pr = probability; LENA = Language ENvironment Analysis.
As Table 4 shows, the coefficients for four of the predictors were statistically significantly larger than 0: FTP gender ratings, both LENA measures, and the measure of inhibitory control. The effect of the FTP gender ratings is not surprising. The positive coefficient for the LENA Child Vocalization Count can be interpreted as evidence that children who engage in more vocal rehearsal at FTP have speech that meets cisgender norms more closely at LTP than children who engage in less vocal rehearsal. Previous analyses of these children have shown that the LENA Child Vocalization Count predicts a variety of other measures for these children (i.e., the extent of coarticulation in word-initial consonant–vowel sequences, Cychosz et al., 2021). Together, these findings are evidence that child vocalization count is a useful predictor of other speech and language behaviors in this cohort.
The two other FTP measures that predicted LTP gender ratings, the measure of inhibitory control, and LENA Conversational Turn Count, had negative coefficients, meaning that children with poorer performance on these measures at FTP have speech that meets cisgender norms more closely at LTP than children with better performance on these measures. The reason for this negative relationship is not clear: It is not obvious why engaging in more conversational turns and having better inhibitory control would negatively affect the development of gendered speech. One speculation is that, for some children, there is a trade-off between learning socially meaningful variation and learning other aspects of language. If that were the case, the lower derived gender ratings at FTP might reflect children selectively prioritizing learning aspects of language other than socially meaningful variation. If that were true, there might be a reciprocal relationship between some language scores and the acquisition of gendered language because the children with high scores at FTP prioritize the learning of other aspects of language. Another possibility is that the children with higher scores on the two FTP measures were disproportionately not cisgender and that the negative relationship reflects the children's successful learning of noncisgender speech norms.
Summary and Discussion
Research Questions Revisited
This section evaluates the research questions. We found that naive listeners do indeed provide different ratings for AMAB and AFAB children's speech on a scale from “definitely a boy” to “definitely a girl,” both for children at FTP (when children were 2.5–3.5 years old) and LTP (when they were 4.5–5.5 years old). The finding that there is a difference at age 3 years is consistent with the findings of Fung et al. (2021), who examined listeners' perception of the gender of six AMAB and six AFAB children. Together, this study and Fung et al. provide powerful evidence that gendered speech is learned early. The fact that gendered speech can be detected before the sex dimorphism that occurs at puberty—and even before the transient sex differences documented by Vorperian et al. (2011)—is evidence that gendered speech is indeed learned, and not a consequence of sex dimorphism. This study found that the accuracy of the stimuli, and three acoustics characteristics of the stimuli being rated, f0, and two measures of /s/ acoustics, predicted ratings, but that differences between AMAB and AFAB children remained even when these were taken into account. We also found that listeners perceived the AMAB children to sound older than the AFAB children but that the differences in perceived gender between the two groups was robust even when perceived age was controlled statistically.
We found that the difference in ratings was larger at LTP than FTP and that the bigger differences between the two groups were attributable entirely to changes in the ratings of AMAB children. This study is the first to demonstrate that developmental changes in gendered speech are driven by the performance of just one group of children. Children's vocal tracts more closely resemble those of adult cisgender women than adult cisgender men. If the learning of gendered speech involves emulating the sex dimorphic characteristics of adult speech, then the onus to learn gendered speech falls disproportionately on children who are emulating adult men. The changes in ratings of AMAB children could reflect that group being more actively engaged in learning gendered speech than AFAB children.
Finally, we found that some language abilities at FTP predicted gendered speech at LTP. Specifically, the volume of child vocalizations from a daylong recording, conversational turns from a daylong recording, and a measure of inhibitory control at FTP predicted gender ratings at FPT. However, the relationship for the latter two variables was the opposite of what we predicted: higher scores (i.e., more conversational turns and better inhibitory control) at FTP was associated with less prototypically gendered speech at LTP. The reason for this inverse relationship is unclear. Equally important, although, was the finding that measures of vocabulary size and speech production accuracy at FTP did not predict gender ratings at LTP. That finding indicates that children with a wide range of vocabulary sizes and degrees of speech production accuracy are capable of learning gendered speech variation. It is important to emphasize again the exploratory nature of this investigation. The measures that were examined in this retrospective study were those available to us from a study examining a very different topic. While each of the measures was plausibly related to gendered speech development, none of them was tailored to the research question. For example, the measure of inhibition that we examined might only be indirectly related to the development of gendered speech. A measure of, for example, children's ability to switch between attention to different sources of information during speech perception (which itself might be related to inhibition) might predict gendered speech acquisition more strongly.
Limitations
This study had a number of limitations. One of the most notable of these was that we had no measure of gender identity for these children. Our analyses were based on cisnormative assumptions, that is, that the target for AMAB children was a male-sounding voice and a female-sounding voice was a target for AFAB children. Kidd et al. (2021) found that 1.8% of high school–age age students in an urban school district were transgender. If the same statistic held of this cohort, we can conclude that the cisnormative assumption is invalid for at least two of our participants. Future, prospective work on this topic should use measures of gender identity that capture the continuously varying, multidimensional nature of gender, as Li et al. (2016) did.
The use of listener ratings in this study is both a benefit and a limitation. It is a benefit in that the ratings reflect a gestalt perception of gender. As described in the introduction, gender is endemic in the speech signal. Gender is reflected in variation in sounds whose articulation requires fine motor control and which are later acquired (like sibilant fricatives), in features that are less articulatorily complex and which are early acquired (like the long voice-onset times of voiceless stops), and in characteristics of sequences of sounds (like the overall f0 of an utterance). Moreover, as summarized by Leung et al. (2018), listeners incorporate these variables when making judgments of gender. It is possible that children at different developmental levels will convey gender through different sets of features, depending on their ability to control the relevant articulations. Gestalt perceptual ratings allow for the possibility that two children might be equally good at performing gender but might do so with very different speech cues. Although this is undoubtedly a strength, perceptual ratings also have a weakness, in that ratings might reflect stereotypes about a gender rather than authentic gender differences. In this study, for example, we found that the accuracy of a token predicted listener ratings, despite that variable not differing between the AMAB and AFAB children in this study. This might reflect stereotypes about the relationship between speech accuracy and gender in adults, like those described by Heffernan (2010). Future work should systematically examine the perception tasks under which stereotypes are least likely to be activated, so that ratings reflect authentic differences between groups.
Related to this, the use of a continuous scale to measure listener perception is also both a benefit and a limitation. In a strictly statistical sense, a continuous scale is advantageous because it provides data that are more granular than is possible from a binary male/female judgment. Indeed, as Figure 1 shows, few of the children in this study would likely be perceived as vividly male or female, as most of the gender ratings were in the middle tertile of the scale. In a conceptual sense, a continuous scale is good in that it acknowledges some of the bedrock principles of gender studies: that gender exists beyond the binary and that variation in gender is continuous. The limitation to this scale, however, is that it does not fully capture the multidimensional nature of gender. As argued by Butler, gender is a multidimensional construct that goes beyond simple masculinity and femininity. The children in this study who were rated toward the ends of the scale might have been perceived to be boy-like or girl-like for very different reasons. Future research should incorporate insights from gender studies to devise ways of measuring gender that go beyond simple appraisals of masculinity and femininity.
Learning Gendered Speech: Hypotheses and Implications
If children do indeed learn gendered variation in speech, the next logical step in this program of research is to examine the mechanism of this learning. Ladegaard and Bleses (2003) explored two hypotheses of how children acquire gender specific speech styles: the frequency hypothesis and the role-model hypothesis. The frequency hypothesis claims that children acquire the phonetic variants of whoever they hear more frequently, that is, their primary caretakers and from observing the ways that adults change their speech style depending on if they are talking to a boy or a girl. The frequency hypothesis is supported by the work of Foulkes et al. (2005), who examined gendered variants of word-medial /t/ in the variety of English spoken in and around Newcastle, England. Foulkes et al. found that caregivers used more of the female-typed variants (in this case, a preaspirated /t/) when speaking to girls, and more male-typed variants (glottalized /t/) when speaking to boys. In that study, children used the cisnormative gender-appropriate variant of /t/ by age 4 years.
One piece of evidence contrary to the idea that gendered speech is learned solely through child-directed speech comes from the findings of Munson et al. (2015). Munson et al. (2015) examined the speech of cisgender boys and AMAB children whose behaviors were deemed to be not cisnormative. Children in the latter group were diagnosed with the now-obsolete diagnostic label GID. Many of these children received an evaluation for GID, because their parents were concerned about their noncisnormative behavior. It seems unlikely that parents who were troubled by children's gender expression would model linguistic variants that were not cisnormative. Parents' reactions to their children's gender conformity or nonconformity vary individually and culturally. Kane (2006) presents the results of interviews of a cohort of parents in New England about their reactions and feelings about gender nonconformity in their preschool children. These parents tended to welcome nonconformity among their AFAB children but less among their AMAB children. Particularly, heterosexual fathers were the most likely to be motivated to endorse their concept of masculinity with their AMAB children. Kane's findings imply that parents might actively correct or counter noncisnormative language variants, particularly in AMAB children.
The role model hypothesis claims that children try to adopt the speech style variations of those they feel more comfortable with, more closely identify with, or both. The role model hypothesis is consistent with a large literature showing that children's language learning in laboratory tasks is affected by their epistemic trust of different adults. Koenig and Harris (2005) showed that children preferred to learn language labels for new objects from individuals whom they observed labeling known objects accurately and not from people who labeled known objects inaccurately or who professed not to know a familiar object's name. Tripp et al. (2021) argue that epistemic trust underlies performance on a variety of different speech and language processing tasks, even ones that do not examine it directly.
Gendered speech learning may follow from a similar principle: Children may prefer to learn speech variants from people whom they trust or who they see as embodying characteristics of themselves. This hypothesis is broadly consistent with a large literature showing that the quality of interaction between infants and their caregivers influences their likeliness to imitate others. Parents are able to direct their children's attention to stimuli in their environment that are significant to them, thus influencing their children to attend more to gender specific stimuli. As children develop their gender identities, they evaluate a group of people positively as soon as he or she identifies with that group. Experimental studies using novel toys demonstrated that children show more interest and remember more details about toys that they believed to be intended for their gender (Martin & Ruble, 2004). Gendered speech learning is likely to involve the learning of specific gendered features from different individuals. This assertion follows from work on the acquisition and evolution of gendered speech styles over the lifespan, including in trans* folks. Zimman (2017a) argued that the acquisition of gendered speech variants in transmasculine men reflects their learning of specific variants (distinctive productions of /s/ and distinctive use of f0) that reflect different components of gendered and different gendered meanings. The same process may be at play in first-language acquisition of gendered speech.
One especially useful framework for understanding the development of gendered speech is Plaut (2010) diversity science framework. Plaut emphasizes that studies of differences in behavior should take into account not only the differences themselves but also the differences in the perception of differences. Tripp and Munson (2021) applied Plaut's framework to understanding gender differences in speech. They propose that individual differences in the perception of gender rely on a variety of factors, including individual differences in the number of gender categories that individuals believe to exist. Variation in children's acquisition of gender might reflect individual differences in the number and type of gender categories that children believe exist. That same principle may affect the ratings of gender that were collected in this experiment. The raters themselves were not the topic of this investigation. We reasoned that if we had enough ratings, we could estimate how a child's gender would be perceived by the population of adults that they interact with regularly. This tactic of averaging over a population of raters has proven very fruitful in studies of other aspects of children's speech development, such as their development of phonetic contrasts (Schellinger et al., 2017). However, it is possible that the raters in this study themselves differed in their perception of gender categories. Examining this systematically in future studies is an important component of building a broader understanding of how gendered behaviors emerge in different individuals and communities.
Understanding the nature of learning gendered speech is important because of the universality of this process. It is also important because of the unique learning situation that faces children with speech and language impairments. In some therapeutic settings, children learn new sounds and words from a single trainer, such as a single speech-language pathologist/speech and language therapist. Individual children's drive to learn a particular speech style might affect their willingness to learn from a trainer whose speech does not embody that style. Understanding these social drives could explain some of the variation in clinical outcomes that cannot be explained by other factors.
Acknowledgments
This research was funded by NIH grant R01 DC02932 to Jan Edwards (lead PI), Munson (MPI), and Mary E. Beckman (MPI) and by a University of Minnesota Undergraduate Opportunities Research Program grant to Kiana Koeppe. We thank Alayo Tripp, Erin Durban, Kerry Ebert, Robert Schlauch, and Sheri Stronach for helpful feedback on this work and Katherine Dougherty and Gisela Smith for help testing adult participants. We thank the entire Learning to Talk research team, and especially Jan Edwards and Mary E. Beckman, without whose work we would not have had child productions to use in this study.
Appendix
Stimulus Characteristics
This appendix describes in detail the acoustics and transcribed accuracy of the stimuli used in this study. The mean number of phonemes correct per word was averaged across the two raters. The use of average accuracy ratings is motivated by the findings of Schellinger et al. (2017), which disagreements between transcribers often show that the sound being transcribed is not a clear member of a particular category. Fully accurate sounds were coded as 1, deleted and substituted sounds coded as 0, and intermediate sounds coded as 0.5. The average phoneme accuracy scores at first time point (FTP) were 0.71 for girls (SD = 0.22) and 0.72 for boys (SD = 0.21) and at last time point (LTP) were 0.89 for girls (SD = 0.14) and 0.90 for boys (SD = 0.13). A linear mixed-effects model examined whether there was a significant effect of time point and talker gender on the phoneme accuracy scores. This model was fitted using the lme4 package in R (Bates et al., 2015). The significance of the model coefficients were determined with the R package LMERTest (Kuznetsova et al., 2017), using Satterthwaite's approximation for degrees of freedom. The fixed effects in this model were time point and gender, and the random effects were child and word. Models were built progressively, starting with a base model with only random intercepts for child and word. The next model added a fixed effect for time point (coded using contrast coding) and a random slope for the effect of time point on individual children (modeled as being uncorrelated with the random intercept for children). This model fits the data significantly better than the base model, χ2 [df = 3] = 110.96, p < .001, indicating that the effect of time point on accuracy was significant. A model including a fixed effect for gender and a random effect of gender on word did not improve model fit, either alone or in an interaction with time point. This indicates that there was not a statistically significant difference in accuracy between assigned male at birth (AMAB) and assigned female at birth (AFAB) children at either time point.
A linear mixed-effects model was used to predict the f0 variables from time point and sex assigned at birth. The model fitting was similar to that for the models predicting accuracy. The model that predicted f0 from time point fits the data better than a baseline model with only random effects of participant ID and word (χ2 [df = 2] = 103.52, p < .001). However, adding sex assigned at birth, either alone or in interaction with time point, did not improve model fit. These data are shown in Figure A1. As this figure shows, the f0 was lower at LTP than at FTP. The AFAB children had higher mean f0 than AMAB children did at FTP, although this difference was not statistically significant. The mean f0s at LTP was nearly identical.
Figure A1.

Mean f0 of the tonic vowel at midpoint for each child, separated by sex assigned at birth (SAB) and time point. FTP = first time point; LTP = last time point; AFAB = assigned female at birth; AMAB = assigned male at birth.
Given that there was only one token of /s/ at each time point, simple linear models were used to predict spectral centroid and spectral variance from sex assigned at birth, time point, and their interaction. The model predicting spectral centroid was significant overall, F[3,202] = 5.407, p = .001. The coefficients for sex assigned at birth and time point were significant (sex assigned at birth: β = 231, t = 1.991, p = .047; for time point: β = 409, t = 3.521, p = .005), but their interaction was not significant. Figure A2 shows that the spectral centroid was higher for assigned female at birth (AFAB) children than assigned male at birth (AMAB) children at both time points. The effect of sex assigned at birth on spectral centroid mirrors the difference seen in adult men and women (Jongman et al., 2000). FTP = first time point; LTP = last time point.
Figure A2.

Spectral centroid of word-initial /s/ for each child, separated by sex assigned at birth (SAB) and time point. FTP = first time point; LTP = last time point; AFAB = assigned female at birth; AMAB = assigned male at birth.
The model predicting spectral variance was significant overall, F(3, 202) = 4.666, p = .005. For the model predicting spectral variance, the coefficient for time point was significant, as was the coefficient for the interaction between time point and sex assigned at birth (time point: β = −97, t = −2.271, p = .024; for Time Point × Sex Assigned at Birth: β = −122, t = −2.857, p = .005). Figure A2 shows that the assigned female at birth (AFAB) children had higher spectral variance than assigned male at birth (AMAB) children at first time point (FTP) but lower spectral variance at last time point (LTP). Munson et al. (2015) found that lower spectral variance was associated with judgments of more masculine speech. Post hoc models showed that the difference at FTP was not statistically significant, while the difference at LTP was. These differences are shown in Figure A3.
Figure A3.

Spectral variance of word-initial /s/ for each child, separated by sex assigned at birth (SAB) and time point. FTP = first time point; LTP = last time point; AFAB = assigned female at birth; AMAB = assigned male at birth.
In sum, the cues to sex assigned at birth were most robustly present in acoustics of /s/. While f0 and stimulus accuracy did differ across time points, they did not differ between the assigned male at birth (AMAB) and assigned female at birth (AFAB) children. FTP = first time point; LTP = last time point.
Funding Statement
This research was funded by NIH grant R01 DC02932 to Jan Edwards (lead PI), Munson (MPI), and Mary E. Beckman (MPI) and by a University of Minnesota Undergraduate Opportunities Research Program grant to Kiana Koeppe.
References
- Amir, O. , Engel, M. , Shabtai, E. , & Amir, N. (2012). Identification of children's gender and age by listeners. Journal of Voice, 26, 313–321. https://doi.org/10.1016/j.jvoice.2011.06.001 [DOI] [PubMed] [Google Scholar]
- Babel, M. , & Munson, B. (2014). Producing socially meaningful linguistic variation. In Goldrick M., Ferreira V., & Miozzo M. (Eds.), The Oxford handbook of language production (pp. 308–325). Oxford University Press. https://doi.org/10.1093/oxfordhb/9780199735471.013.022 [Google Scholar]
- Barbier, G. , Boё, J. , Captier, G. , & Laboissiѐre R. (2016). Human vocal tract growth: A longitudinal study of the development of various anatomical structures. In 16th Annual Conference of the International Speech Communication Association (Interspeech 2015), International Speech Communication Association, Sep 2015, Dresden, Germany. https://doi.org/10.21437/Interspeech.2015 [Google Scholar]
- Barreda, S. , & Assmann, P. F. (2018). Modeling the perception of children's age from speech acoustics. The Journal of the Acoustical Society of America, 143, EL361–EL366. https://doi.org/10.1121/1.5037614 [DOI] [PubMed] [Google Scholar]
- Bates, D. , Maechler, M. , Bolker, B. , & Walker, S. (2015). lme4: Linear mixed-effects models using Eigen and S4. R package version 1.1–9. https://CRAN.R-project.org/package=lme4
- Beckman, M. E. , Munson, B. , & Edwards, J. (2007). Vocabulary growth and the developmental expansion of types of phonological knowledge. In Laboratory Phonology 9. Mouton de Gruyter. [Google Scholar]
- Boersma, P. (2001). Praat, A system for doing phonetics by computer. Glot International, 5(9), 341–345. [Google Scholar]
- Borragan, M. , Martin, C. D. , De Bruin, A. , & Duñabeitia, J. A. (2018). Exploring different types of inhibition during bilingual language production. Frontiers in Psychology, 9, 2256. https://doi.org/10.3389/fpsyg.2018.02256 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Butler, J. (2011). Gender trouble: Feminism and the subversion of identity. Routledge. https://doi.org/10.4324/9780203824979 [Google Scholar]
- Calder, J. (2019). The fierceness of fronted /s/: Linguistic rhematization through visual transformation. Language in Society, 48(1), 31–64. https://doi.org/10.1017/S004740451800115X [Google Scholar]
- Cartei, V. , Cowles, W. , Banerjee, R. , & Reby, D. (2014). Control of voice gender in pre-pubertal children. British Journal of Developmental Psychology, 32(1), 100–106. https://doi.org/10.1111/bjdp.12027 [DOI] [PubMed] [Google Scholar]
- Cartei, V. , Cowles, H. W. , & Reby, D. (2012). Spontaneous voice gender imitation abilities in adult speakers. PLOS ONE, 7, Article e31353. https://doi.org/10.1371/journal.pone.0031353 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cartei, V. , Garnham, A. , Oakhill, J. , Banerjee, R. , Roberts, L. , & Reby, D. (2019). Children can control the expression of masculinity and femininity through the voice. Royal Society Open Science, 6, 190656. https://doi.org/10.1098/rsos.190656 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cristia, A. , Bulgarelli, F. , & Bergelson, E. (2020). Accuracy of the language environment analysis system segmentation and metrics: A systematic review. Journal of Speech, Language, and Hearing Research, 63(4), 1093–1105. https://doi.org/10.1044/2020_JSLHR-19-00017 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cychosz, M. , Munson, B. , & Edwards, J. R. (2021). Practice and experience predict coarticulation in child speech. Language Learning and Development, 17(4), 366–396. https://doi.org/10.1080/15475441.2021.1890080 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Deák, G. O. (2003). The development of cognitive flexibility and language abilities. Advances in Child Development and Behavior, 31, 271–327. https://doi.org/10.1016/S0065-2407(03)31007-9 [DOI] [PubMed] [Google Scholar]
- Dunn, L. , & Dunn, D. (2007). Peabody Picture Vocabulary Test–Fourth Edition (PPVT-4). Pearson. [Google Scholar]
- Eckert, P. , & McConnell-Ginet, S. (1992). Think practically and look locally: Language and gender as community-based practice. Annual Review of Anthropology, 21, 461–488. https://doi.org/10.1146/annurev.an.21.100192.002333 [Google Scholar]
- Erskine, M. E. , Munson, B. , & Edwards, J. R. (2020). Relationship between early phonological processing and later phonological awareness: Evidence from nonword repetition. Applied Psycholinguistics, 41(2), 319–346. https://doi.org/10.1017/S0142716419000547 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Foulkes, P. , Docherty, G. , & Watt, D. (2005). Phonological variation in child-directed speech. Language, 81(1), 177–206. https://doi.org/10.1353/lan.2005.0018 [Google Scholar]
- Fox, R. A. , & Nissen, S. L. (2005). Sex-related acoustic changes in voiceless English fricatives. Journal of Speech, Language, and Hearing Research, 48(4), 753–765. https://doi.org/10.1044/1092-4388(2005/052) [DOI] [PubMed] [Google Scholar]
- Fuchs, S. , & Toda, M . (2010). Do differences in male versus female /s/ reflect biological or sociophonetic factors? In Turbulent sounds: An interdisciplinary guide (pp. 281–302). Mouton de Gruyter. https://doi.org/10.1515/9783110226584 [Google Scholar]
- Fung, P. , Schertz, J. , & Johnson, E. K. (2021). The development of gendered speech in children: Insights from adult L1 and L2 perceptions. Journal of the Acoustical Society of America Express Letters, 1(1), Article 014407. https://doi.org/10.1121/10.0003322 [DOI] [PubMed] [Google Scholar]
- Goldman, M. , & Fristoe, R. (2000). Goldman-Fristoe Test of Articulation–Second Edition (GFTA-2). Pearson. https://doi.org/10.1037/t15098-000 [Google Scholar]
- Harries, M. , Hawkins, S. , Hacking, J. , & Hughes, I. (1998). Changes in the male voice at puberty: Vocal fold length and its relationship to the fundamental frequency of the voice. The Journal of Laryngology & Otology, 112(5), 451–454. http://doi.org/10.1136/adc.77.5.445 [DOI] [PubMed] [Google Scholar]
- Heffernan, K. (2010). Mumbling is macho: Phonetic distinctiveness in the speech of American radio DJs. American Speech, 85(1), 67–90. https://doi.org/10.1215/00031283-2010-003 [Google Scholar]
- Jongman, A. , Wayland, R. , & Wong, S. (2000). Acoustic characteristics of English fricatives. The Journal of the Acoustical Society of America, 108(3), 1252–1263. https://doi.org/10.1121/1.1288413 [DOI] [PubMed] [Google Scholar]
- Kane, E. W. (2006). “No way my boys are going to be like that!” parents' responses to children's gender nonconformity. Gender and Society, 20(2), 149–176. https://doi.org/10.1177/0891243205284276 [Google Scholar]
- Kaufman, A. S. , & Kaufman, N. L. (2014). Kaufman Brief Intelligence Test (K-BIT). Pearson. https://doi.org/10.1002/9781118660584.ese1325 [Google Scholar]
- Kidd, K. M. , Sequeira, G. M. , Douglas, C. , Paglisotti, T. , Inwards-Breland, D. J. , Miller, E. , & Coulter, R. W. (2021). Prevalence of gender-diverse youth in an urban school district. Pediatrics, 147(6), Article e2020049823. https://doi.org/10.1542/peds.2020-049823 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Koenig, M. , & Harris, P. (2005). Preschoolers mistrust ignorant and inaccurate speakers. Child Development, 76(6), 1261–1277. https://doi.org/10.1111/j.1467-8624.2005.00849.x [DOI] [PubMed] [Google Scholar]
- Kuznetsova, A. , Brockhoff, P. B. , & Christensen, R. H. B. (2017). lmerTestPackage: Tests in linear mixed effects models. Journal of Statistical Software, 82(13), 1–26. https://doi.org/10.18637/jss.v082.i13 [Google Scholar]
- Ladegaard, H. J. , & Bleses, D. (2003). Gender differences in young children's speech: The acquisition of sociolinguistic competence. International Journal of Applied Linguistics, 13(2), 222–233. https://doi.org/10.1111/1473-4192.00045 [Google Scholar]
- Lee, S. , Potamianos, A. , & Narayanan, S. (1999). Acoustics of children's speech: Developmental changes of temporal and spectral parameters. The Journal of the Acoustical Society of America, 105, 1455–1468. https://doi.org/10.1121/1.426686 [DOI] [PubMed] [Google Scholar]
- Leung, Y. , Oates, J. , & Pang Chan, S. (2018). Voice, articulation, and prosody contribute to listener perceptions of speaker gender: A systematic review and meta-analysis. Journal of Speech, Language, and Hearing Research, 61(2), 266–297. https://doi.org/10.1044/2017_JSLHR-S-17-0067 [DOI] [PubMed] [Google Scholar]
- Li, F. (2017). The development of gender-specific patterns in the production of voiceless sibilant fricatives in Mandarin Chinese. Linguistics, 55(5), 1021–1044. https://doi.org/10.1515/ling-2017-0019 [Google Scholar]
- Li, F. , Rendall, D. , Vasey, P. L. , Kinsman, M. , Ward-Sutherland, A. , & Diano, G. (2016). The development of sex/gender-specific /s/ and its relationship to gender identity in children and adolescents. Journal of Phonetics, 57, 59–70. https://doi.org/10.1016/j.wocn.2016.05.004 [Google Scholar]
- Lieberman, P. (1986). Some aspects of dimorphism and human speech. Human Evolution, 1, 67–75. https://doi.org/10.1007/BF02437286 [Google Scholar]
- Little, R. (1988). A test of missing completely at random for multivariate data with missing values. Journal of the American Statistical Association, 83(404), 1198–1202. https://doi.org/10.1080/01621459.1988.10478722 [Google Scholar]
- Martin, C. L. , & Ruble, D. (2004). Children's search for gender cues: Cognitive perspectives on gender development. Current Directions in Psychological Science, 13(2), 67–70. https://doi.org/10.1111/j.0963-7214.2004.00276.x [Google Scholar]
- Munson, B. , & Babel, M. (2019). The phonetics of sex and gender. In Katz W. & Assmann P. (Eds.), Routledge handbook of phonetics (pp. 499–525). Routledge. https://doi.org/10.4324/9780429056253 [Google Scholar]
- Munson, B. , Beckman, M. E. , & Edwards, J. (2011). Phonological representations in language acquisition: Climbing the ladder of abstraction. In Cohn A., Fougeron C., & Huffman M. (Eds.), Oxford handbook in laboratory phonology (pp. 288–309). Oxford University Press. https://doi.org/10.1093/oxfordhb/9780199575039.001.0001 [Google Scholar]
- Munson, B. , Crocker, L. , Pierrehumbert, J. B. , Owen-Anderson, A. , & Zucker, K. J. (2015). Gender typicality in children's speech: A comparison of boys with and without gender identity disorder. The Journal of the Acoustical Society of America, 137, 1995–2003. https://doi.org/10.1121/1.4916202 [DOI] [PubMed] [Google Scholar]
- Munson, B. , Logerquist, M. , Kim, H. , Martell, A. , & Edwards, J. (2021). Does early phonetic differentiation predict later phonetic development? Evidence from a longitudinal study of /ɹ/ development in preschool children. Journal of Speech, Language, and Hearing Research, 64(7), 2417–2437. https://doi.org/10.1044/2021_JSLHR-20-00555 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Patterson, M. L. , & Werker, J. F. (2002). Infants' ability to match dynamic phonetic and gender information in the face and voice. Journal of Experimental Child Psychology, 81, 93–115. https://doi.org/10.1006/jecp.2001.2644 [DOI] [PubMed] [Google Scholar]
- Perry, T. L. , Ohde, R. N. , & Ashmead, D. H. (2001). The acoustic bases for gender identification from children's voices. The Journal of the Acoustical Society of America, 109, 2988–2998. https://doi.org/10.1121/1.1370525 [DOI] [PubMed] [Google Scholar]
- Plaut, V. C. (2010). Diversity science: Why and how difference makes a difference. Psychological Inquiry, 21(2), 77–99. https://doi.org/10.1080/10478401003676501 [Google Scholar]
- Plummer, A. R. , Ménard, L. , Munson, B. , & Beckman, M. E. (2013). Comparing vowel category response surfaces over age-varying maximal vowel spaces within and across language communities. Interspeech, 2013, 421–425. https://doi.org/10.21437/Interspeech.2013 [Google Scholar]
- Rvachew, S. (1994). Speech perception training can facilitate sound production learning. Journal of Speech and Hearing Research, 37(2), 347–357. https://doi.org/10.1044/jshr.3702.347 [DOI] [PubMed] [Google Scholar]
- Schellinger, S. K. , Munson, B. , & Edwards, J. (2017). Gradient perception of children's productions of /s/ and /θ/: A comparative study of rating methods. Clinical Linguistics & Phonetics, 31(1), 80–103. https://doi.org/10.1080/02699206.2016.1205665 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Stoel-Gammon, C. (2001). Transcibing the speech of young children. Topics in Language Disorders, 21, 12–21. https://doi.org/10.1097/00011363-200121040-00004 [Google Scholar]
- Stuart-Smith, J. (2007). Empirical evidence for gendered speech production: /s/ in Glaswegian. In Laboratory Phonology 9 (pp. 65–86). Mouton de Gruyter. [Google Scholar]
- Tripp, A. , Feldman, N. H. , & Idsardi, W. J. (2021). Social inference may guide early lexical learning. Frontiers in Psychology, 12, 645247. https://doi.org/10.3389/fpsyg.2021.645247 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tripp, A. , & Munson, B. (2021). Perceiving gender while perceiving language: Integrating psycholinguistics and gender theory. Wiley Interdisciplinary Reviews: Cognitive Science, Article e1583. https://doi.org/10.1002/wcs.1583 [DOI] [PubMed] [Google Scholar]
- Van Buuren, S. , & Groothuis-Oudshoorn, K. (2010). MICE: Multivariate imputation by chained equations in R. Journal of Statistical Software, 45(3), 1–68. https://doi.org/10.18637/jss.v045.i03 [Google Scholar]
- Vorperian, H. K. , Kent, R. D. , Lindstrom, M. J. , Kalina, C. M. , Gentry, L. R. , & Yandell, B. S. (2005). Development of vocal tract length during early childhood: A magnetic resonance imaging study. The Journal of the Acoustical Society of America, 117, 338–350. https://doi.org/10.1121/1.1835958 [DOI] [PubMed] [Google Scholar]
- Vorperian, H. K. , Wang, S. , Schimek, E. M. , Durtschi, R. B. , Kent, R. D. , Gentry, L. R. , & Chung, M. K. (2011). Developmental sexual dimorphism of the oral and pharyngeal portions of the vocal tract: An imaging study. Journal of Speech, Language, and Hearing Research, 54(4), 995–1010. https://doi.org/10.1044/1092-4388(2010/10-0097) [DOI] [PMC free article] [PubMed] [Google Scholar]
- Weinberg, B. , & Bennett, S. (1971). Speaker sex recognition of 5- and 6-year-old children's voices. The Journal of the Acoustical Society of America, 50, 1210–1213. https://doi.org/10.1121/1.1912757 [DOI] [PubMed] [Google Scholar]
- Williams, K. (2007). Expressive Vocabulary Test, Second Edition. Pearson. https://doi.org/10.1037/t15094-000 [Google Scholar]
- Xu, D. , Gilkerson, J. , & Richards, J. A. (2012). Objective child vocal development measurement with naturalistic daylong audio recording. Interspeech, 2012, 1123–1126. https://doi.org/10.21437/Interspeech.2012 [DOI] [PubMed] [Google Scholar]
- Zelazo, P. , Anderson, J. , Richler, J. , Wallner-Allen, K. , Beaumont, J. , & Weintraub, S. (2013). NIH Toolbox Cognitive Function Battery (CFB): Measuring executive function and attention. In Zelazo P. & Bauer P. (Eds.), National Institutes of Health Toolbox—Cognitive Function Battery (NIH Toolbox CFB): Validation for children between 3 and 15 years Monographs of the Society for Research in Child Development. 4. Vol. 78 (pp. 16–33). Society for Research in Child Development. https://doi.org/10.1111/mono.12032 [DOI] [PubMed] [Google Scholar]
- Zimman, L. (2017a). Gender as stylistic bricolage: Transmasculine voices and the relationship between fundamental frequency and /s/. Language in Society, 46(3), 339–370. https://doi.org/10.1017/S0047404517000070 [Google Scholar]
- Zimman, L. (2017b). Variability in /s/ among transgender speakers: Evidence for a socially grounded account of gender and sibilants. Linguistics, 55(5), 993–1019. https://doi.org/10.1515/ling-2017-0018 [Google Scholar]
