Abstract
The primary vocal registers of modal, falsetto, and fry have been studied in adults but not per se in infancy. The vocal ligament is thought to play a critical role in the modal-falsetto contrast, yet it is still developing during infancy ([1] Tateya & Tateya, 2015). Cover tissues are also implicated in the modal-fry contrast but the low fo cutoff of 70 Hz, shared between genders, suggests a psychoacoustic basis for the contrast. Buder, Chorna, Oller, and Robinson ([2], 2009) used the labels of “loft,” “modal,” and “pulse” for distinct vibratory regimes that appear to be identifiable based on spectrographic inspection of harmonic structure and auditory judgments in infants, but this work did not supply acoustic measurements to verify which of these nominally labeled regimes resembled adult registers. In this report, we identify clear transitions between registers within infant vocalizations, and measure these registers and their transitions for fo and relative harmonic amplitudes (H1 – H2). By selectively sampling first-year vocalizations, this manuscript quantifies acoustic patterns that correspond to vocal fold vibration types not previously cataloged in infancy. Results support a developmental basis for vocal registers, revealing that a well-developed ligament is not needed for loft-modal quality shifts as seen in harmonic amplitude measures. Results also reveal that a distinctively pulsatile register can occur in infants at much higher fos than expected on psychoacoustic grounds. Overall results are consistent with cover tissues in infancy that are, for vibratory purposes, highly compliant and readily detached.
Keywords: infant vocal development, registers, regimes, falsetto, modal, vocal fry, acoustics
1. Introduction
At least three primary registers are agreed to occur in adult phonation, known most widely as modal, falsetto, and fry [3]. Modal, which is most typical, is the default register for speaking and singing in habitual fundamental frequency (fo) range, corresponding to “chest” voice in singers. Falsetto occurs in higher fo ranges, corresponding to “head” voice in singers. Vocal fry occurs in the lowest fo ranges, also being called “creaky,” “laryngealized,” or “glottalized” voice by linguists who have observed this phonatory contrast to be phonemic in some languages [4, 5]. In these contexts, registers are controlled adjustments of phonation and may be assumed as such to be well controlled behaviors. However, harmonic and waveform patterns auditorily resembling falsetto and fry, along with other non-modal patterns such as subharmonics, have been observed in infancy [6, 7], and in this context they can be regarded as the naturally and spontaneously occurring regimes of an inherently non-linear dynamic system [2, 8, 9].
Svec et al. [10] contrast the spontaneity of natural ‘bifurcations’ that are to be expected in highly non-linear dynamic systems such as phonation with what they characterized as the “prevailing opinion” (of that time) that a muscular adjustment is required for register shifts. In a non-linear dynamic bifurcation, sudden changes in vibratory regime may occur abruptly as a control parameter (such as longitudinal vocal fold tension or subglottal pressure) is smoothly varied. The richness of vibratory regimes studied by Buder et al. [2] is consistent with the view that the bifurcation principle is operative in infant phonation (though underlying mechanisms and parameters remain unresolved).
Explication of levels of analysis and associated terminology can help maintain clear conceptual distinctions between vibratory mechanisms, acoustic outputs, and auditory qualities. Phonation can be characterized as resulting from interactions between airflow and vocal fold vibrations. These interactions typically correspond to clear acoustic effects. Moreover, as voice scientists also classify vocal outputs based on auditory qualities, the adjective ‘phonatory’ is often applied to that level as well. Scientific study of adult registers has been conducted at all three levels, but seems to have been rooted primarily in auditory impressions in the context of acoustic and mechanistic investigations. In the present article, we will also use the term register as an auditory-acoustic construct intended to inform understanding of the underlying phonatory mechanisms, and thus the term “phonation” here is intended to encompass both vibratory and acoustic phenomena. Along similar lines in our usage, the term ‘vocal quality’ originates in the auditory domain with concomitant acoustic markers, pointing towards potentially distinct phonatory mechanisms. “Regimes” refers to the broadest set of vocal fold vibration possibilities as distinguished primarily via inspection of the harmonic structure of phonatory output.
In the present paper, we are most interested in register transitions as presumably spontaneous phenomena in the phonatory exploration that characterizes much of vocal development in the first year of life [11]. While falsetto and fry are the most widely used terms for non-modal registers in contemporary literature, in the context of infant phonatory types it seems preferable to return to the terms originally coined by Hollien [12] — “loft” and “pulse.” These terms also connote acoustic bases more directly and as such are relatively agnostic with respect to underlying mechanisms. They also leave open the possibility that occurrence of similar phonation types in infancy are distinct from those used intentionally for performance and/or linguistic purposes.
While not entirely conclusive, a well-developed literature on registers in adult phonation has established some consensus on the underlying mechanisms that distinguish the primary non-modal registers. Loft is generally thought to involve lateralization of the muscular body of the vocal folds and a decoupling of the cover tissues from the body, yielding a relatively small vertical contact area to participate in vibration [3]. The resulting voice quality is therefore not only higher in fo, but relatively thin sounding, with a conspicuously higher amplitude of the first harmonic relative to the second and overall greater spectral slope. Pulse, lower in fo than modal but also with a distinctive ‘rasping’ quality, is generally thought to involve greater activation of the body with a lax cover [3], associated with reduced airflow and a longer closed time [13].
The study of registers in infancy is of particular interest because of differences between adult and infant vocal fold composition. Developmental studies of the vocal mechanisms have reinforced as well as expanded the many ways in which we know that the infant is not just a smaller version of an adult vocalizer [14]. The vocal fold tissues themselves differ substantially in the first year of life from those of the mature adult, the lamina propria being relatively undifferentiated and macula flavae mostly undeveloped [15]. While this topic has been investigated fairly systematically in the anatomical and medical literature, revealing a protracted schedule of tissue differentiation that begins fetally but extends into adolescence [16, 17, 18, 19], there are very few studies that have analyzed non-cry phonation with these tissue mechanics in mind—see however Fuamenya et al. [20] for recent work in the context of crying.
The cover-tissue difference in infants is particularly relevant to questions regarding registers—in adults, both loft and pulse purportedly involve cover-body configurations that are quite different than that used to produce modal. Specifically, cover-body decoupling in loft phonation is thought to involve the ligamental nature of the intermediate and deep layers of the lamina propria. However, the distinctions between these connective layers are essentially lacking in the cover tissues of infant vocal folds. Pulse is thought to be associated with a lax cover in adults, yet the cover is particularly compliant in infancy. A basic theoretical question therefore arises in the search for registers in infant phonation. Given the presumed roles of cover tissues in loft and pulse register production, and the known differences in infant vocal fold cover tissues, do registers occur in infancy, and if they do, how do they differ from adult registers? This question requires acoustic criteria for the non-modal registers.
Fo is a direct acoustic correlate of the two non-modal registers when produced by mature speakers, but does not by itself distinguish those registers from modal phonation. Infants do not produce registers on demand, so criteria other than fo are needed. Buder et al. [2] identified harmonic spacing patterns consistent with pulse and loft fo ranges, but did not provide independent acoustic markers. Physiological register shifts between modal and falsetto may occur at different fo values depending on the individual, often with rather large pitch jumps [10]. As pitch may shift without a concomitant register shift, distinctive vocal quality measurements are needed.
Compared to loft, pulse in adult phonation is more clearly distinguished from modal by fo alone, transitioning remarkably at the same value of c. 70 Hz for both men and women [3]. While investigators tend to agree that pulse is produced by a distinct vibratory mechanism [3], psychoacoustics alone could explain why 70 Hz is the value at which individual glottal pulses become audibly distinct [3]. For infancy, in which fo values at transitions to pulse have not previously been determined, and because salient pitch shifts to pulse have not been reported as they have for modal-loft transitions, it is again important to consider additional acoustic measures in conjunction with fo. A defining feature of pulse phonation is the critical damping of each glottal pulse excitation before the next pulse occurs, causing a distinctive waveform shape marked by the appearance of temporal gaps between glottal pulses [3]. This waveform-shape criterion operates independently of harmonic amplitudes and thus can serve as an independent feature that may signal pulse register in infancy.
Relative amplitudes of the first two harmonics have been considered as indicating glottal status in registers, most definitively for loft [21, 22]. Relative harmonic amplitude measures, including the difference between the first and second harmonics, were introduced in the empirical literature by Hanson [23] as a reflection of the relatively open glottal configuration found in female voices: First harmonics were higher amplitude relative to the second in women [23]. More recent studies indicate that the amplitudes of the first two harmonics are sensitive not only to open quotient and breathiness but to a variety of phonation types, including register [24]. Examination of the literature in general supports use of both fo and H1-H2 to distinguish both between modal and loft and between modal and pulse registers.
Table 1 summarizes previous studies that reported either transitional fo or representative H1-H2 values for registers in non-disordered speakers. The literature employing these measures for studies of register is surprisingly spotty, with no adult studies having characterized both registers together in terms of both measures. While scientists developing theories of voice production per se are clearly interested in all possible manners of voice production [3, 25], the phenomena of loft and pulse registers are treated in applied studies as distinct domains: Loft (i.e. falsetto) has been of primary interest in singing, and normative pulse (i.e., ‘creak,’ ‘laryngealization,’ or ‘glotttalization’) has been investigated as signaling phonemic contrasts, or marking phrasal boundaries [26], in speech. Speaking and singing are distinct manners of voice production so it has also made sense to study them distinctively.
Table 1.
References reporting fo and/or H1-H2 values for modal vs. non-modal registers
| Citation | Population | fo (Hz) | H1-H2 (dB) |
|---|---|---|---|
| Modal vs. Loft | |||
| Neiman et al. [21] | Adult males | Loft ≥ 250, Modal lower | Loft H1 > H2 (in 23 of 25) Modal H1 < H2 (in 18 of 25) |
| Salamao & Sundberg [22] | Adult males singing | Various, with regions of overlap | H1 > H2, but with 14.2 dB greater difference in loft than in modal.a |
| Švec et al. [10] | Excised male larynx (aerodynamically excited) | Modal to Loftb: 168↗332 Loft to Modalb: 234↘146 |
|
| Modal vs. Pulse | |||
| Gordon & Ladefoged [4] | Male speaker of Zapotec | Pulse H1<H2, Modal H1>H2 | |
| McGlone [28] | Adult males Adult females |
Pulse: 34–51 Pulse: 28–49 |
|
| McGlone & Shipp [29] | Adult males | Pulse: 18–65, Modal: 87–117 | |
| Blomgren et al.[30] | Adult males Adult females |
Mean Pulse = 49, Mean Modal = 117 Mean Pulse = 48, Mean Modal = 211 |
|
| Avelino [31] | 3 Adult males 3 Adult females |
Modal H1-H2 > Pulse H1-H2 (in 2 of 3)c Modal H1-H2 > Pulse (in 3 of 3) |
|
Measured from spectrum of inverse filtered flow glottogram.
Note that loft value after upward jump is higher than loft value before downward jump, and vice versa for modal values, indicating ‘bistable’ region.
Female values all positive, male values all but one (pulse) negative.
While acoustic measures may corroborate distinct registers in opportunistically observed vocalizations, identification of registers in such materials may begin with auditory judgments. Abrupt pitch shifts have been widely observed in infant phonation [6, 7], but not all of them are necessarily associated with register transitions, so auditory and acoustic classification of distinctive qualities across such shifts remains necessary. Similarly, “growl” vocalizations are often marked by low frequency pulsing [2, 27], but precise comparisons in acoustic terms across transitions are needed to demonstrate whether and in what ways these qualities and apparent registers are categorically distinct. Our approach emphasizes that within-vocalization register transitions specifically satisfy two scientific concerns raised by difficulties in observing infant phonatory patterns: (1) listening across a transition optimizes auditory identification, and (2) some of the possible physiological variables affecting register shifts may be held constant across the transition. Under these conditions, significant changes in acoustic quantities optimally support the premise that registers are driven by distinctive phonatory mechanisms.
In summary, the main goal of this study is to document and quantify distinctive registers in infancy by investigating the nature of these registers as produced by immature vocal folds. Again, to our knowledge the investigation of both loft and pulse using basic acoustic metrics in one study has not previously been published even with adult subjects, much less with children or infants. The theory of adult phonation suggests that due to the absence of a vocal ligament in infants clear breaks between modal and loft are unlikely to occur at this stage of development, and there is also little basis on which to predict that infants will exhibit a clearly distinct pulse register at any specific frequencies. Secondarily, acoustic quantification of the observed registers will enhance metrics used to classify infant phonation into distinct types, helping to associate them with regimes and distinct vocal qualities. This latter goal will support yet unfulfilled objectives in studies of infant vocal development, specifically identifying the physiological and acoustic bases for the protophone categories purported to form a key basis (that of systematic contrastive sound production) for subsequent language development [33, 34, 35].
2. Material and Methods
2.1 Participants and Recordings
Two 20-minute recording sessions for each of three typically developing female infants were selected for coding at each of three ages: ‘early’ (3–4 months), ‘mid’ (5–7 months), and ‘late’ (9–11 months) for a total of 18 sessions (360 minutes). Infants were both video and audio recorded in a laboratory equipped with 4 cameras and set up as a child’s playroom. For representativeness, recordings were paired, one from a session in which the mother and infant were freely interacting and the other during a mother-experimenter interview session in which infants were otherwise often vocally active ‘separately.’
The infants were fitted with custom-built vests that housed a wireless microphone system (Samson Airline UHF AL1 transmitter, equipped with a Countryman Associates low-profile low-friction flat frequency response MEMWF0WNC capsule, sending to a Samson UHF AM1 receiver). The vest configuration followed an original design developed by Buder and Stoel-Gammon [36], with the microphone capsule housed within a velcro patch and oriented to maintain the mouth-to-microphone distance at approximately 10–15 cm. TF32 software [37] operating a DT321 acquisition card (Data Translation, Inc.) was used to digitize the infant signals at 48 kHz after low-pass filtering at 20 kHz via an AAF-3 anti-aliasing filter board.
2.2 Materials
Recording sessions were analyzed in the AACT (Action Analysis, Coding, and Training) environment [38] which presents synchronized video and audio and allows users to demarcate and label intervals on spectrographic displays during coding. All non-cry, non-laugh, non-vegetative protophone vocalizations [27, 35] meeting minimal audibility and duration criteria (> 50 ms) were coded in breath group units [39]. These vocalizations were subsequently coded into intervals (referred to below as “segments’) representing the following mutually exclusive and exhaustive phonatory regimes [2]: Modal (clear, parallel, and moderately spaced harmonics), HiModal (more widely-spaced harmonics or an audible falsetto quality), Pulse (very closely spaced harmonics, widely spaced glottal pulses, and a ‘zipper-like’ sound), Subharmonics (lower-amplitude harmonics appearing in between main harmonics), Biphonation (two different sets of harmonics moving in non-parallel directions), C-Stops (within vocalization adduction-caused gaps in phonation), O-Stops (within vocalization abduction-caused gaps in phonation), and Chaos (very unclear harmonic structure with aperiodicity in glottal pulses): See [2] for more detailed definitions and examples.
Of special interest in the current investigation, the HiModal code was a stand-in for possible loft register: Coders were trained only to mark very high-pitched intervals or intervals in which the thin and weak quality of a non-modal ‘falsetto’ voice was salient. As the existence of a true loft register in infancy had not been verified in prior literature, no training was provided to certify that regime coders could reliably distinguish loft from modal (but we note that the adult literature generally seems to lack that certification as well). Hence, the HiModal codes in our dataset included cases that might have been judged to involve ‘loft” if such a code had been included in the coding protocol, but the codes also included many segments which were high-pitched but still perceived to have been produced by the same mechanism as lower-pitched modal segments. Pulse was deemed to be a more salient quality in infancy, reliably distinguished on the basis of close harmonics, critically-damped glottal pulses, and a specifically raspy “zipper” or “frog” quality [2]. Pulse regimes mark “growl” protophones, but so can other low-fo or even mid-fo ‘rough’ regimes such as subharmonics, biphonation, and chaos [27].
In all, 2,445 non-cry, non-vegetative, non-laugh, protophone vocalizations were identified and coded for regimes. From this corpus, an exhaustive search was made, via inspection of spectrograms and listening, for vocalizations in which quality transitions between modal and ‘high modal’ codes or between modal and pulse codes were audible within the utterance. Transitions could be in either order, but with no intervening regimes. The following paragraphs detail how we inspected for markers of loft and pulse, independently of the fo and H1-H2 measurements that would then be used to characterize the registers, to corroborate whether the previous regime coding boundaries represented a register transition and where that transition occurred.
For modal↔ loft transitions, only those in which there was a perceptually very salient difference in vocal quality and spectral tilt across the transitions were retained for analysis. Transitions were marked accordingly. For modal↔ pulse transitions clear temporal gaps between glottal pulses should be observed characterizing the pulse segments [3]. While waveform shape was one consideration in previous ‘pulse’ regime coding, a screening criterion for such temporal gaps had not been utilized.
For this purpose, RMS amplitude contours were extracted with a very narrow 0.5 ms window to inspect for “temporal gaps” between glottal pulses—see Figure 1. Inspection of these contours was a primary consideration but not always sufficient: Variations in overall vocal intensity and noise background had to be taken into account when judging where the RMS amplitude had reached the ambient floor, and in such cases waveform morphology could also be a consideration. Furthermore, as is especially conspicuous in Figure 1 panel (b), an interim ‘transitional’ segment was sometimes observed in which the temporal gaps were lower in amplitude than in the adjacent modal segment but did not actually reach the amplitude. Most often such gaps between modal and pulse segments were well under 50 ms.
Figure 1.
Illustrations of RMS procedure for modal-pulse discrimination. Each panel, extracted from TF32, displays the waveform on top (c. 0.5 s), a narrowband spectrogram (c. 0–4kHz), and RMS amplitude with a 0.5 ms window. Label displays beneath the amplitude trace include “(--------)” codes indicating the segments of interest, also indicated by arrows. Panel (a) present a particularly clear example with only a brief transition gap, while Panel (b) presents the most difficult case encountered in this corpus, with an extended transition due to brief, intervening regimes and an unclear amplitude floor due to noise overlay. The samples are from different infants but both at the youngest age represented in the data.
Figure 2 provides a schematic to illustrate where acoustic measures were extracted relative to the identified transitions. Measurement locations differed somewhat between the two register types because of the transitional gaps between modal and pulse that were not observed in modal↔ loft transitions. A primary question regarding fo is the size of the shift between registers; for this question it was important to make fo measurements without variable gaps between registers. On the other hand, as H1-H2 measures were applied here to characterize the registers as such, it was important to accommodate transitional gaps between modal and pulse (especially in cases such as that illustrate in Figure 1b). In summary, all measures were taken as closely “adjacent” to the transitions as possible except for H1-H2 in pulse, which was taken as “near” to transitions as possible while still ensuring that the measures clearly represented pulse phonation.
Figure 2.
Schematic of fo and H1-H2 measurement locations (‘waveform’ shapes in this schematic do not depict actual glottal waveforms and are merely intended to be evocative of register distinctions). Measurements were taken at ‘Mid’ locations for all three registers, and ‘Adjacent’ to loft-modal transitions. In pulse-modal transitions, transitional segments such as illustrated in Figure 1 were typically observed: For this reason, fo measures were still taken ‘adjacent’ to transitions, but H1-H2 measures were taken for pulse only ‘Near’ the transition. This latter criterion meant that for pulse a .5 ms windowed RMS contour had to reveal a continuous sequence of glottal pulses that were critically damped exhibiting temporal gaps for H1-H2 measurement.
Employing the criteria identified above for modal↔ non-modal transitions, analysts endeavored to find five clear examples of each transition type from each of the three infants at each of the three ages. However, sometimes fewer than five, or even just one good example, could be found to represent an infant/age. The resulting sample is detailed in Table 2 of the Results section below.
Table 2.
Observations of within-utterance register transitions. The total number of transitions = 151. M-L = modal to loft, L-M = loft to modal, P-M = pulse to modal, M-P = modal to pulse.
| Infant | Age (months) | Vocs w/Loft | Transitions | Vocs w/Pulse | Transitions | Modal Segments |
|---|---|---|---|---|---|---|
| AD | 3 | 6 | 6 M-L, 3 L-M | 10 | 7 M-P, 3 P-M | 18 |
| 6 | 3 | 1 M-L, 3 L-M | 8 | 5 M-P, 3 P-M | 12 | |
| 9 | 4 | 5 M-L, 3 L-M | 6 | 3 M-P, 3 PM | 16 | |
| EA | 3 | 7 | 5 M-L, 5 L-M | 5 | 2 M-P, 1 P-M | 13 |
| 5 | 8 | 6 M-L, 5 L-M | 9 | 2 M-P, 7 P-M | 19 | |
| 10 | 5 | 3 M-L, 2 L-M | 5 | 1 M-P, 4 P-M | 10 | |
| SM | 4 | 12 | 8 M-L, 9 L-M | 8 | 5 M-P, 3 P-M | 25 |
| 6 | 10 | 8 M-L, 5 L-M | 7 | 1 M-P, 6 P-M | 20 | |
| 11 | 10 | 9 M-L, 4 L-M | 5 | 1 M-P, 4 P-M | 16 | |
| Totals: | 65 | 51 M-L, 39 L-M | 63 | 27 M-P, 34 P-M | 149 |
Note: The tallies of transitions for each infant and age represent the number of clear examples that could be readily found among the 40 minutes of recorded examined but not necessarily the maximum number to be found. Smaller numbers (5 or fewer) reflect a paucity of examples (e.g. infant EA at the youngest age for pulse), but in other cases more than the listed number were readily identified (e.g. infant SM for loft, or infant AD at younger ages for pulse), but only the clearest exemplars were retained to provide representation across infants and ages.
The Action Analysis, Coding and Training (AACT) software was used for extracting harmonic amplitude values. AACT implements the TF32 acoustic analysis library for extracting parameters such as harmonic frequencies and amplitudes, given an initial fo cursor placement in a (optionally frequency-zoomed) narrowband spectrum. AACT/TF32 software automatically generates the following seven values for harmonic amplitude extraction: the time associated with the cursor placement, the fundamental frequency (regional fo as determined by frequency cursor placement), regional-peak frequency and amplitude for the first harmonic, H1 (taken as fo for this study), and regional-peak frequency and amplitude close to double H1, which was taken to be H2, and finally H1-H2 as the amplitude of H2 subtracted from the amplitude of H1. The harmonic spectrum was visually inspected to confirm the accuracy of each extraction. Materials with ambiguous harmonic structure due to overlapping vocalizations, noise, or unclear harmonic structure were measured in consultation with co-authors until consensus was reached. In many cases, problems with identification of harmonics occurred due to very low intensity and/or very low frequency, causing either or both of the first harmonics to fall below the spectral amplitude floor. Such cases were excluded from the final data. These problems account for somewhat reduced numbers of observations in some categories, especially for pulse H1-H2 measures. H1-H2 measures were not transformed according to vowel identity as in the H1*-H2* measures proposed by Hanson [18] primarily because formant measurements are notoriously difficult in infant materials due to sparse harmonic sampling, nasality and a host of other issues [40]. We note, however, that prior literature on registers in adult phonation has also presented raw H1-H2 measures (see Table 1).
During training of two analysts for H1-H2 measures from the entire corpus of regime segments identified as loft or modal, a large sample for assessing reliability was obtained, demonstrating good inter-coder correlation for H1-H2 (Pearson r(353) = .90). A third analyst placed location markers for the pulse measures; reanalysis of 20% of the pulse-associated measurement locations by a reliability coder revealed a strong yet somewhat lower correlation (Pearson r(84) = .84). Intercoder reliability was also obtained on 20% of the data for the location of pulse↔ modal transition locations using the RMS-based technique; coder differences on average differed by less than 3 ms (standard deviation 27 ms).
As a first approach to the topic with no specific hypothesis-driven framework in terms of such measures in infant register samples, the data are summarized ‘descriptively,’ by univariate statistics (t-test, ANOVAs) on the measures, examined within-groups as appropriate, e.g. across transitions. Nonetheless, even with a modestly sized dataset and large overall variability, basic results of interest were significant in the third order alpha level—p < .001, so we have reported all effects with alpha level .05 or less and no ‘number of tests’ adjustments have been applied.
3. Results
3.1 Occurrence of Registers
Table 3 lists vocalizations selected from the initial pool of 2445 according to the criteria listed in Method. This subset does not represent an exhaustive inventory of all transition occurrences, but does include the utterances submitted to analysis and demonstrates that all contiguous and clearly distinct registers occurred within vocalization at least once at each age for each infant. Sample sound files described in the Appendix below are included as supplemental materials for this article
Table 3.
Fo values of registers in Hz at midpoint of regime segment.
| Minimum | Mean | Maximum | n | |
|---|---|---|---|---|
| Loftmid | 399 | 791 | 2322 | 90 |
| Modalmid | 228 | 384 | 779 | 148 |
| Pulsemid | 35 | 122 | 211 | 58 |
3.2 Acoustic Characteristics of Registers and Register Transitions
3.2.1 Register differences
An ANOVA confirmed that the three registers were clearly distinct by fo in Hz (F(2, 292) = 238, p < .001) and also in each pairwise comparison (Tukey HSD, p < .001). These effects were stronger yet in semitones (F(2,292) = 681), where the distributions were clearly more well normalized as seen in Figure 3b below by the symmetrical whiskers, the distribution of outliers both above and below the median, and the absence of extreme outliers above the median. See Table 3 for values and numbers of fo observations. Note that modal and loft showed overlap, while modal and pulse registers were non-overlapping. It is notable that pulse in the infant data was on average at frequencies that would be in a modal range for adults—122 Hz—and was observed to be as high as 211 Hz even at regime midpoint.
Figure 3.
Box and whisker plots depicting fo in (a) Hz and (b) semitones, measured mid-register.
ANOVA also indicated an overall difference between registers in these data by the H1-H2 measure (F[2, 276) = 93.4, p < .001), but pairwise comparisons by Tukey HSD were significant at p < .001 only for loft versus the other two registers. Both modal and pulse showed a high second harmonic relative to the first in infant phonation, while loft phonation in the infants yielded the expected higher first harmonic relative to the second. There was still much variation within registers and overlap across registers, but ‘loft’ was distinct in both measures.
3.2.2 Register transitions
We now turn to inspection of the measures taken adjacent to transitions. Recall that for these comparisons, as schematized in Figure 2, the fo data were taken as close to the transition as possible, ‘Adjacent’ to it, in both types of register shifts. Loft transitional H1-H2 measures were also taken ‘Adjacent’ to shifts, but pulse transitional H1–H2 measures were taken as ‘Near’ to the transitions as possible, so long as pulse-by-pulse temporal gaps had been clearly observed (again, typically no more than 10–20 ms separated from clearly modal regimes). Table 5 lists fo measure statistics by transition type for each register separately, and also includes the average changes for each transition type in both Hz and semitones, and regions of overlap between the upward and downward going transitions. Figure 5 depicts the transition differences in fo.
Table 5.
Fo values in Hz of non-modal and modal registers adjacent to transition points by transition types. Bolded entries indicate last adjacent values in the direction of transitions. Italics indicate average last adjacency. ST = semitones.
| Transition | Register | Minimum | Mean | Maximum | Mean Transition | Shifts |
|---|---|---|---|---|---|---|
| Modal→Loft | Modal | 289 | 428 | 603 | 313 Hz, (9.1 ST) | 428 ↗ 741 |
|
|
||||||
| n = 51 | Loft | 440 | 741 | 1219 | ||
|
| ||||||
| Loft→Modal | Loft | 396 | 687 | 1254 | −272 Hz, (−8.4 ST) | 687 ↘415 |
|
|
||||||
| n = 39 | Modal | 290 | 415 | 647 | ||
|
| ||||||
| Modal→Pulse | Modal | 136 | 267 | 377 | −66 Hz, (−5.3 ST) | 267 ↘201 |
|
|
||||||
| n = 29 | Pulse | 70 | 201 | 293.7 | ||
|
|
||||||
| Pulse→Modal | Pulse | 124 | 206 | 275 | 96 Hz (6.5 ST) | 206 ↗302 |
|
|
||||||
| n = 32 | Modal | 193 | 302 | 560 | ||
Figure 5.
Fo transitions by type and direction. Error bars are standard errors.
3.2.3 Loft/Modal fo transitions
Paired t-tests confirmed that both upward and downward shifts traversed significantly distinct endpoints: modal→loft: t(50) = 13.9. p < .001, loft→modal: t(38) = −11.2, p < .001. These effects were even larger when assessed on the semitone scale: modal→loft: t(50) = 22.0. p < .001, loft→modal: t(38) = −15.9, p < .001. The average last Hz value of modal prior to a transition to loft (428) was about the same as the average first Hz value of modal after a transition from loft (415), and a t-test affirmed that these values were statistically indistinguishable (t(88) = −0.77, p = .44). The first Hz value of loft after a transition from modal (741) was also comparable to the last value of loft before a transition to modal (687), and these were also statistically indistinct (t(88) = −1.21, p = .23). The upward and downward shifts appeared comparable in Hz and their magnitudes were statistically indistinct (t(88) = 1.2, p = .22).
3.2.4 Pulse/Modal fo transitions
Paired t-tests confirmed that both upward and downward shifts traversed significantly distinct endpoints: modal→pulse: t(31)=−5.46, p < .001, pulse→modal: t(28)= −5.94, p < .001. These effects are comparable when assessed on the semitone scale: modal→pulse: t(31) = −4.76. p < .001, pulse→modal: t(28) = 6.41, p < .001. The average last Hz value of modal prior to a transition to pulse (267) was smaller than the average first Hz value of modal after a transition from pulse (302), and this difference did reach statistical significance (t(59) = −2.05, p < .05), but this level of significance might be viewed with caution given the number of tests applied to these data (and when one high outlier value was removed from the post-pulse modal data the p value was .08). The first Hz value of pulse after transition from modal (201) was comparable to the last Hz value of pulse before transition to modal (206), and these were not statistically distinct (t(59) = −0.4, p = .68). The upward shift of 96 was larger than the downward of −66, but, apparently due to variability these magnitudes were not statistically distinct (t = −1.51, p = .14).
Table 6 lays out the H1-H2 measures in the same format as fo measures in Table 4, and Figure 6 depicts transition line charts for this measure. Recall that for mid-register, depicted in Figure 2 and quantified in Table 3, large H1-H2 differences were observed in relative harmonic amplitudes for loft versus modal and loft versus pulse, but not for modal versus pulse.
Table 6.
H1-H2 values in dB of non-modal and modal registers close to transition points by transition types. Bolded entries indicate last adjacent values in the direction of transitions. Italics indicate average ‘last’ adjacency.
| Transition | Register | Minimum | Mean | Maximum | Mean Transition |
|---|---|---|---|---|---|
| Modal→Loft | Modal | −26.3 | −7.4 | 9.9 | +14.5 |
|
|
|||||
| n = 51 | Loft | −16.4 | 7.1 | 37.8 | |
|
| |||||
| Loft→Modal | Loft | −17.9 | 8.0 | 31.5 | −15.2 |
|
|
|||||
| n = 39 | Modal | −21.5 | −7.2 | 8.9 | |
|
| |||||
| Modal→Pulse | Modal | −19 | −9.2 | 6.9 | +0.8 |
|
|
|||||
| n = 29 | Pulse | −20.3 | −8.4 | 4.2 | |
|
| |||||
| Pulse→Modal | Pulse | −21.3 | −8.0 | 6.2 | −4.3 |
|
|
|||||
| n = 32 | Modal | −39.1 | −12.3 | 9.8 | |
Table 4.
H1 – H2 values in dB of registers at midpoint of regime segment.
| Minimum | Mean | Maximum | n | |
|---|---|---|---|---|
|
|
||||
| Loftmid | −12 | 9.3 | 33 | 89 |
|
| ||||
| Modalmid | −22 | −6.0 | 17.2 | 148 |
|
| ||||
| Pulsemid | −17 | −6.0 | 6.4 | 42a |
Figure 6.
Bar charts depicting H1-H2 transitions by type and direction
3.2.5 Loft/Modal H1-H2 transitions
Large shifts in average H1-H2 values were observed in both transition directions, reflecting an abruptly occurring change between a relatively strong second harmonic in modal to a relatively strong first harmonic in loft. Paired t-tests reveal that these harmonic amplitude changes were statistically significant across the modal to loft transition (t(50) = −7.61, p < .001) and across the loft to modal transition (t(38) = 7.13, p < .001). Note in Table 6, however, that a very large range of relative harmonic amplitudes was observed in both modal and loft—this range is examined in more detail below. The shifts appear to be comparable in both directions, and their magnitudes were statistically indistinct (t(88) = .02, p = .99)
3.2.6 Pulse/Modal H1–H2 transitions
Because H1-H2 differences were not found between pulse and modal mid register, it is not surprising that they were also not observed to be very different across the transitions. More surprising, is the observation somewhat less negative H1-H2 differences in pulse as compared to nearby modal values, indicating comparatively somewhat less amplitude in the second harmonic, as in loft. After modal to pulse transitions, the second harmonic was higher very slightly relative to the first, but not with statistical significance according to matched pair t-test (t(31) = −.66, p = .51). Across pulse to modal transitions, a small overall increase in the second harmonic relative to the first was significantly distinct at the p < .01 level (t(29) = 2.61, p < .05). Again, this small effect must be viewed with caution given the small numbers of observations and larger number of tests conducted on this dataset overall, but it remains of interest as tending oppositely to expectation based on adult models.
3.3 Variation within Registers
As seen in Figure 4 for mid-register variation and the transitional value ranges listed in Table 6, there was considerable variation of relative harmonic amplitudes in these data both within and across registers. Some of this variation may be related to fundamental frequency: The overall Pearson correlation taken between fo and H1-H2 for all the available measures in this corpus was significantly positive (r = .50, p < .001), possibly reflecting a general covariation of these acoustic variables with pressure, amplitude, glottal resistance, and/or other parameters likely to cause H1-H2 changes. However, the scatter plot depicted in Figure 7 suggests that this overall correlation was just as much if not entirely due to variation across the register clusters in this space. Furthermore, if the registers are produced by different vibratory mechanisms [3], patterns of co-variation might also be expected to differ within register.
Figure 4.

Box and whisker plots depicting H1-H2, measured mid-register.
Figure 7.
Scatterplot of all H1-H2 measures over fo (inclusive of mid-register, loft/modal adjacencies, pulse-adjacent modal, and modal-near pulse).
Examining solely within registers, a small but significant positive correlation was obtained for the modal values (r = .18, p < .01) while the others were non-significant at zero or negative values (loft r = .01, p = .89; pulse r = −.15, p = .13). Furthermore, when only mid-register values were examined, all correlations became non-significant, and the negative correlation for pulse register became more so at r = −.23. In conclusion, fo and H1-H2 measures were not observed to co-vary except when all three registers were combined, and so this variation may be epiphenomenal to the distribution across clusters, and not characteristic of the mechanisms except possibly within modal. For the purpose of using fo and H1-H2 as acoustic classifiers distinguishing registers, we therefore proceed to analyze modal↔ non-modal distinctions for the two non-modal registers separately.
3.4 Acoustic Classification of Registers
Table 7 presents logistic regression modeling results using fo and H1-H2 as predictors and each of the two pairwise modal/non-modal register contrasts as outcomes. Note that pulse was fully predicted by fo, with no further predictive value when H1-H2 was added. Loft was best predicted by fo, but H1-H2 added a good bit of final variance-explained, increasing prediction from .86 to .90. Overall, at .90 prediction of loft with only two measures, and .96 for pulse with only one, the measurement set contributes a great deal of explanation for the categories.
Table 7.
Logistic regression results for discriminating registers across transitions
| H1-H2 | fo | H1-H2 and fo | ||||||
|---|---|---|---|---|---|---|---|---|
|
|
|
|
||||||
| Prediction | AreaROC | Cutoff | Prediction | AreaROC | Cutoff | Prediction | AreaROC | |
| Loft-Modal | .74 | .88 | 2.8 | .86 | .97 | 523 | .90 | .98 |
| Pulse-Modal | .62a | .48a | NA | .96 | .996 | 209 | ||
Only fo significant.
The fo threshold value for modal versus pulse in this dataset was 208.5 Hz, conforming nicely with the general rule of thumb of 200 Hz that ‘regime’ coders had used when identifying pulse [2]. Note again, however, that this would be a remarkably high fo at which to hear individual pulses, yet pulse was the percept for these infants: a generally rasping ‘washboard’/’zipper’/’frog’ sound, occurred despite the high fo. The result is surprising because adult pulse is reportedly perceived at the much lower fo where individual pulses become detectable.
With both variables predicting modal versus loft, the combination threshold values were fo = 516 Hz and H1-H2 = 1.6 dB. This fo cutoff was slightly lower than the threshold for fo alone, and this H1-H2 threshold was slightly more negative than in the harmonics alone model. In combination the measures increased prediction but at negligibly small improvement in comparison to either measure alone. However, given the large gaps across loft↔ modal transitions and the large variation of both measures within registers, it is almost certainly of value to utilize both measures when discriminating these registers. Examining correctly classified instances, it could be seen that sometimes one measure ‘rescued’ the other. For example, the segment with the lowest fo that was still identified as loft—fo = 396 Hz—had an H1-H2 measure of 9.8 dB, well above the cutoffs. The highest fo still identified as modal—fo = 779 Hz—had an H1-H2 measure of −12.6 dB. These observations validate perceptual impressions that perceived registers in these infants were not simply a matter of pitch.
4. Discussion
Each of the 3 infants in this study demonstrated within-vocalization register transitions at least once in relatively brief recorded samples at each of the ages examined during their first year of vocal development. The loft-modal transitions were particularly salient, corresponding in the acoustic data to large jumps in fo and large changes in H1-H2, typically positive for perceived loft and negative for perceived modal. Pulse-modal transitions were equally evident however, and were also signaled by clear acoustic markers. While these transitions were distinguished better by fo than by H1-H2, pulse surprisingly trended towards more positive H1-H2 values: Given the theoretical bases for this measure, such values would presumably result from a less-pressed configuration of the glottis in comparison to the immediately adjacent modal phonation.
The harmonic amplitude measures alone did not clearly support for the infants the distinct mechanism for pulse phonation posited in the adult literature. Moreover, perceived register transitions occurred around 200 Hz, casting doubt upon the applicability for infants of the psychoacoustic argument from adult literature suggesting that pulse is identifiable because individual pulses are only audibly distinct when fo is lower than 70 Hz. By the same token, the simple fact that infant vocal folds will naturally vibrate at higher frequencies does not explain the perceived register transition at the higher fo. Furthermore, the lack of clearly distinct H1-H2 values in perceived pulse register seems to indicate that the infants were not producing pulse by ‘constricting’ the glottis as is generally thought to occur in adults [7].
Most broadly in comparison to adult phonation, these results indicate that the loft register operates very similarly in infants although it is signaled by somewhat higher fo, while pulse may operate quite differently in infancy than in adults. The findings should be interpreted in the context of known differences between adult and infant vocal fold composition, specifically regarding the relatively undifferentiated and compliant lamina propria in infants. In formulating contemporary myoelastic-aerodynamic theory, van den Berg himself [41] identified the ligament as most likely needed for loft phonation. The stiffness of adult ligamental layers in conjunction with macula flavae contribute measurable differences to natural vocal fold vibration frequencies [42]. As it is established that these two aspects of vocal fold histology—ligament and macula flavae—are underdeveloped during possibly the entire first year of life [1], it is most interesting to observe that infants produce a distinctive loft register quite readily, apparently contradicting the received wisdom that a ligamental layer is critically involved.
A propensity for infant vocal folds to produce sound signaling a pulse register, even at relatively high fo values but without distinctive H1-H2 values, may be understood in terms of infant tissue characteristics. ‘De-ligamented’ lamina propria are likely to be highly compliant and also more readily detachable from the body. Such loose and detachable vocal fold cover tissues could provide the lax cover and marginalized body tissues that fry phonation is thought to involve, potentially explaining the lack of distinctive H1-H2 characteristics even for glottal pulses that exhibit critical damping as evidenced in temporal gaps between pulses. In other words, a reason that this perceived register did not correspond to an anticipated ‘pressed’ marker, signaled by a high second harmonic (relative to nearby modal register), might be attributable to the very compliance of these cover tissues. Furthermore, this loose detachable aspect might also explain the occurrence of ‘falsetto’ style vibration without the involvement of a ligament. The natural protophone categories that result (often called vowel-like sounds, squeals, and growls) are routinely recognized by caregivers and form anchors for vocal interaction while infants are still pre-linguistic [11].
Turning to the observed threshold values obtained in the logistic regression results, these values will be useful in classification work underway towards understanding the acoustic dimensions of infant protophones [11, 27]. While previously our regime codes included only an agnostic ‘High Modal’ category, the fo and harmonic amplitude measures reported here suggest that a loft register is indeed present in infants. The register distinction may provide the basis for distinguishing two types of “squeals” in infant vocalizations, one as corresponding to loft, and the other, high-pitched modal. Register distinctions made on an acoustic basis can now be added to other acoustic dimensions, such as fo variability and other non-modal regimes, and the combination should help in modeling of human coding of infant protophone types [43, 44]. Perceptual coding of phonation by young children with autism suggests that the register distinction may be important for understanding audible signs characterizing this population [45].
4.1 Summary
The results presented here demonstrated that within-vocalization transitions between modal and the two primary non-modal registers of loft and pulse occur at each of three different ages during the first year of life in each of three infants sampled. These transitions resemble those observed in adults even though the cover tissues of infant vocal folds are quite different than found in adults. In addition to describing phonatory phenomena not previously documented in infant phonation, the results present specific challenges to voice science by demonstrating that (a) loft phonation does not require a vocal ligament, and (b) pulse phonation can occur audibly at frequencies above 200 Hz.
4.2 Limitations, Future Research
A major limitation of this study is the small number of speakers involved, only three female infants. As with so many other aspects of speech analysis, all the measures of this study required prior determination of fo, and early infant phonation is notorious for confounding this determination (due e.g., to subharmonics, biphonation, and chaos). With ongoing developments in automated recognition of harmonic structures in spectrograms, however, it can be hoped that more rapid and ultimately automatic detection of registers in infant phonation might be possible.
With more rapid data collection on a larger number of subjects, applications may extend beyond certification of protophone dimensions and categories in a variety of directions. Age effects, while not found in this limited dataset, ought to be observed beyond the first year of life, especially in a lowering of the pulse fo threshold. It is possible that infants’ use of registers could be imitative, or vary by situation as an indicator of emotion or interactional intent. Quantified discrimination of pre-linguistic phonation and distress-call (‘crying’) phonation has yet to be performed, and knowledge of register-specific modes of infant phonation will provide an essential backdrop for this effort. Early detection of voice-related signs and symptoms of speech disorders is another longer-term direction, but meanwhile, clearly much basic voice science is yet to be done tracing the ontogeny of register range measures with the schedule of ligament and macula flavae developments.
Supplementary Material
Acknowledgments
Eugene H. Buder and D. Kimbrough Oller received support for this research from Grants R01 DC006099 and DC011027 from the National Institute on Deafness and Other Communication Disorders. Oller’s role in this research was also founded by an endowment from the Plough Foundation, which supports his Chair of Excellence at the University of Memphis.
Appendix – Sound File Descriptions
Sound File 1: The vocalization contains two measured modal-loft transitions produced from the middle to the end of the vocalization. This vocalization was produced by 3-month-old infant EA.
Sound File 2: The vocalization contains a loft-modal transition at the end of the clip (but also contains some environmental noise from a rattle). The vocalization was produced by infant SM at 6 months.
Sound File 3: The vocalization contains another example of a modal-loft transition, with another upward pitch break after that. This vocalization was produced by infant AD at 9 months.
Sound File 4: This vocalization is complex but contains an example of a pulse-modal transition being produced by 3-month-old infant AD. The transition occurs near the middle of the vocalization, following the loft-modal transition at the beginning. The pulse fit the criteria for visual and auditory identification but is an example of pulse with an fo above 200 Hz, measuring at 223.
Sound File 5: The vocalization consists of a pulse-modal transition with pulse fo measured at 266. This example was produced by infant EA at 5 months.
Sound File 6: The low-intensity vocalization includes two pulse segments and a modal segment in the middle and was produced by infant SM at 11 months. Analysis focused on the modal-pulse transition that occurs at the end of vocalization with fo measured at 223 Hz.
Footnotes
Conflicts of interest: none.
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final citable form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
References
- 1.Tateya T, Tateya I. Regenerative Medicine in Otolaryngology. Springer; New York: 2015. Vocal fold development; pp. 161–169. [Google Scholar]
- 2.Buder EH, Chorna L, Oller DK, Robinson R. Vibratory regime classification of infant phonation. J Voice. 2008;22:553–564. doi: 10.1016/j.jvoice.2006.12.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Titze IR. Principles of Voice Production. Prentice Hall; Englewood Cliffs, NJ: 1994. [Google Scholar]
- 4.Gordon M, Ladefoged P. Phonation types: A cross-linguistic overview. J Phon. 2001;29:383–406. [Google Scholar]
- 5.Keating PA, Garellek M, Kreiman J. Acoustic properties of different kinds of creaky voice. Proceedings of the 18th International Congress of Phonetic Sciences; Glasgow. 2015. [Google Scholar]
- 6.Kent RD, Murray AD. Acoustic features of infant vocalic utterances at 3, 6, and 9 months. J Acoust Soc Am. 1982;72:353–365. doi: 10.1121/1.388089. [DOI] [PubMed] [Google Scholar]
- 7.Robb MP, Saxman JH. Acoustic observations in young children’s non-cry vocalizations. J Acoust Soc Am. 1988;83:1876–1882. doi: 10.1121/1.396523. [DOI] [PubMed] [Google Scholar]
- 8.Mende W, Herzel H, Wermke K. Bifurcations and chaos in newborn infant cries. Phys Lett A. 1990;145:418–424. [Google Scholar]
- 9.Robb MP. Bifurcations and chaos in the cries of full-term and preterm infants. Folia Phoniatr Logop. 2003;55:233–240. doi: 10.1159/000072154. [DOI] [PubMed] [Google Scholar]
- 10.Svec JG, Schutte HK, Miller DG. On pitch jumps between chest and falsetto registers in voice: Data from living and excised human larynges. J Acoust Soc Am. 1999;106:1523–1531. doi: 10.1121/1.427149. [DOI] [PubMed] [Google Scholar]
- 11.Oller DK. The Emergence of the Speech Capacity. Lawrence Erlbaum Associates; Mahwah, N.J: 2000. [Google Scholar]
- 12.Hollien H. On vocal registers. J Phon. 1972;2:124–143. [Google Scholar]
- 13.Hollien H, Girard GT, Coleman RF. Vocal fold vibratory patterns of pulse register phonation. Folia Phoniatr. 1977;29:200–205. doi: 10.1159/000264089. [DOI] [PubMed] [Google Scholar]
- 14.Vorperian HK, Wang S, Chung MK, Schimek EM, Durtschi RB, Kent RD, Gentry LR. Anatomic development of the oral and pharyngeal portions of the vocal tract: An imaging study a. J Acoust Soc Am. 2009;125:1666–1678. doi: 10.1121/1.3075589. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Sato K, Hirano M, Nakashima T. Fine structure of the human newborn and infant vocal fold mucosae. Ann Otolaryngol Rhinol Laryngol. 2001;110:417–424. doi: 10.1177/000348940111000505. [DOI] [PubMed] [Google Scholar]
- 16.Boseley ME, Hartnick CJ. Development of the human true vocal fold: depth of cell layers and quantifying cell types within the lamina propria. Ann of Otology, Rhinology, & Laryng. 2006;115:784–788. doi: 10.1177/000348940611501012. [DOI] [PubMed] [Google Scholar]
- 17.Hartnick CJ, Rehbar R, Prasad V. Development and maturation of the pediatric human vocal fold lamina propria. The Laryngoscope. 2005;115:4–15. doi: 10.1097/01.mlg.0000150685.54893.e9. [DOI] [PubMed] [Google Scholar]
- 18.Nita LM, Battlehner CN, Ferreira MA, Imamura R, Sennes LU, Caldini EG, Tsuji DH. The presence of a vocal ligament in fetuses: a histochemical and ultrastructural study. J Anat. 2009;215:692–697. doi: 10.1111/j.1469-7580.2009.01146.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Rosenberg TL, Schweinfurth JM. Cell density of the lamina propria of neonatal vocal folds. Ann of Otology, Rhinology, & Laryng. 2009;118:87–90. doi: 10.1177/000348940911800202. [DOI] [PubMed] [Google Scholar]
- 20.Fuamenya NA, Robb MP, Wermke K. Noisy but effective: Crying across the first 3 months of life. J Voice. 2015;29:281–286. doi: 10.1016/j.jvoice.2014.07.014. [DOI] [PubMed] [Google Scholar]
- 21.Neiman M, Robb M, Lerman J, Duffy R. Acoustic examination of naturalistic modal and falsetto voice registers. Logop Phoniatr Vocology. 1997;22:135–138. [Google Scholar]
- 22.Salomao GL, Sundberg J. What do male singers mean by modal and falsetto register? An investigation of the glottal voice source. Logop Phoniatr Vocology. 2009;34:73–83. doi: 10.1080/14015430902879918. [DOI] [PubMed] [Google Scholar]
- 23.Hanson HM. Glottal characteristics of female speakers: Acoustic correlates. J Acoust Soc Am. 1997;101:466–481. doi: 10.1121/1.417991. [DOI] [PubMed] [Google Scholar]
- 24.Kreiman J, Shue YL, Chen G, Iseli M, Gerratt BR, Neubauer J, Alwan A. Variability in the relationships among voice quality, harmonic amplitudes, open quotient, and glottal area waveform shape in sustained phonation a. J Acoust Soc Am. 2012;132:2625–2632. doi: 10.1121/1.4747007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Roubeau B, Henrich N, Castellengo M. Laryngeal vibratory mechanisms: The notion of vocal register revisited. J Voice. 2009;23:425–438. doi: 10.1016/j.jvoice.2007.10.014. [DOI] [PubMed] [Google Scholar]
- 26.Dilley L, Shattuck-Hufnagel S, Ostendorf M. Glottalization of word-initial vowels as a function of prosodic structure. J Phon. 1996;24:423–444. [Google Scholar]
- 27.Buder EH, Warlaumont AS, Oller DK. An acoustic phonetic catalog of prespeech vocalizations from a developmental perspective. In: Peter B, MacLeod AAN, editors. Comprehensive Perspectives on Speech Sound Development and Disorders: Pathways from Linguistic Theory to Clinical Practice. Nova Science Publishers Inc; Hauppage, NY: 2013. pp. 103–134. [Google Scholar]
- 28.McGlone RE. Air flow during vocal fry phonation. J Speech Lang Hear Res. 1967;10:299–304. doi: 10.1044/jshr.1002.299. [DOI] [PubMed] [Google Scholar]
- 29.McGlone RE, Shipp T. Some physiologic correlates of vocal-fry phonation. J Speech Lang, Hear Res. 1971;14:769–775. doi: 10.1044/jshr.1404.769. [DOI] [PubMed] [Google Scholar]
- 30.Blomgren M, Chen Y, Ng ML, Gilbert HR. Acoustic, aerodynamic, physiologic, and perceptual properties of modal and vocal fry registers. J Acoust Soc Am. 1998;103:2649–2658. doi: 10.1121/1.422785. [DOI] [PubMed] [Google Scholar]
- 31.Avelino H. Acoustic and electroglottographic analyses of nonpathological, nonmodal phonation. J Voice. 2010;24:270–280. doi: 10.1016/j.jvoice.2008.10.002. [DOI] [PubMed] [Google Scholar]
- 33.Koopmans-van Beinum FJ, van der Stelt JM. Early stages in the development of speech movements. In: Zetterstrom R, editor. Precursors of Early Speech. Stockton Press; New York: 1986. pp. 37–50. [Google Scholar]
- 34.Nathani S, Ertmer D, Stark RE. Assessing vocal development in infants and toddlers. Clin Linguist Phon. 2006;20:351–369. doi: 10.1080/02699200500211451. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Oller DK. The emergence of the sounds of speech in infancy. In: Yeni-Komshian GH, Kavanagh JF, Ferguson CA, editors. Child Phonology (Vol 1: Production) Academic Press; New York: 1980. pp. 93–112. [Google Scholar]
- 36.Buder EH, Stoel-Gammon C. American and Swedish children’s acquisition of vowel duration: Effects of vowel identity and final stop voicing. J Acoust Soc Am. 2002;111:1854–1864. doi: 10.1121/1.1463448. [DOI] [PubMed] [Google Scholar]
- 37.Milenkovic P. TF32. University of Wisconsin-Madison; Madison, WI: 2001. computer software. [Google Scholar]
- 38.Delgado RE, Milenkovic P. AACT- Action analysis coding and training software. Intelligent Hearing Systems Corp; Miami: 2017. computer software. [Google Scholar]
- 39.Lynch MP, Oller DK, Steffens ML, Buder EH. Phrasing in prelinguistic vocalizations. Dev Psychobiol. 1995;1:3–25. doi: 10.1002/dev.420280103. [DOI] [PubMed] [Google Scholar]
- 40.Robb MP, Cacace AT. Estimation of formant frequencies in infant cry. Int J Ped Oto. 1995;33:57–67. doi: 10.1016/0165-5876(94)01112-b. [DOI] [PubMed] [Google Scholar]
- 41.van den Berg J. Myoelastic-aerodynamic theory of voice production. In: Kent RD, Atal BS, Miller JL, editors. Papers in Speech Communication: Speech Production. Acoustical Society of America; Woodbury, NY: 1991. pp. 121–138. [Google Scholar]
- 42.Titze IR, Hunter EJ. Normal vibration frequencies of the vocal ligament. J Acoust Soc Am. 2004;115:2264–2269. doi: 10.1121/1.1698832. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Buder EH. Topography of human vocal development. Symposium on Learning about the Vocal World: Deciphering the Statistics of Communication; Atlanta, GA. 2015. [Google Scholar]
- 44.Buder EH, Oller DK. Acoustic structure of common early protophones. Annual Meeting of the American Speech-Language-Hearing Association; Denver, CO. 2015. [Google Scholar]
- 45.Wallace MC, Cleary JE, Buder EH, Oller DK, Sheinkopf SJ, Mundy P. An acoustical analysis of vocalizations of children with autism. Annual Meeting of the American Speech-Language-Hearing Association; Chicago. 2008. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.






