Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Jun 1.
Published in final edited form as: Infant Behav Dev. 2025 Apr 25;79:102059. doi: 10.1016/j.infbeh.2025.102059

Evaluating Canonical Babbling Ratios Extracted from Day-Long Audio Recordings in Infants Later Diagnosed with Autism Spectrum Disorder

Shoba S Meera a,b, Divya Swaminathan a, Rahul Pawar c, Lisa Yankowitz d, Kevin Donovan b, Khavi Khuu b, Julia Parish-Morris d, Steven F Warren e, Annette Estes f,g, Lonnie Zwaigenbaum h, Mark Clements i, David V Anderson i, Robert T Schultz d, Heather C Hazlett b,j, Tanya St John f,g, Juhi Pandey d, Natasha Marrus k, Kelly Botteron l, Stephen R Dager f,m, Meghan R Swanson n,*, Linda R Watson o,*, Joseph Piven, for The IBIS Networkb,j,*
PMCID: PMC12148686  NIHMSID: NIHMS2077365  PMID: 40286507

Abstract

Canonical babbling (CB) is a critical developmental milestone that typically occurs in the second half of the first year of life. Studies focusing on CB in infants at elevated familial likelihood for autism spectrum disorder (ASD) or who later receive an ASD diagnosis are limited and the evidence is mixed. CB comprises a series of canonical syllables (CS) which are defined as the rapid transitions between consonant-like sounds paired with vowel-like sounds (e.g., [gugugu]). One way of measuring CB is by computing canonical babbling ratio (CBR) i.e. total number of CS divided by the total number of syllables. If the child has reached the criterion of 0.15 CBR it is said that they have achieved the CB stage. For several years now, CB has been measured using short lab based or home-based video recordings which may not represent a child’s natural vocalization pattern since child vocalizations fluctuate throughout the day. Day long audio recordings, that capture a child’s vocalizations throughout the day, has the potential to overcome this limitation. Therefore, the current study aimed to answer whether CBRs computed from day-long audio recordings using the language environment analysis (LENA®) were different among infants at elevated familial likelihood for ASD who receive an ASD diagnosis (EL-ASD; n=11), who did not receive an ASD diagnosis (EL-Neg; n=32) and infants at low likelihood for ASD (LL-Neg; n=25) at 9 and 15 months. The study also aimed to evaluate if there are group differences in reaching the canonical babbling stage at 9 and 15 months and are CBRs at 9 and 15 months associated with later language abilities at 24 months. Findings indicated no group differences in mean CBRs at 9 and 15 months and no association with later language abilities. However, we found that children in the EL-ASD group were less likely to reach the 0.15 CBR threshold for being in the canonical babbling stage by 9 months of age. Thus, suggesting that a diagnosis of ASD is associated with delays in CB for some children. Future work in this area must include a larger sample and more standardized annotation protocols to harmonize results across studies and ensure replication.

Keywords: Canonical babbling (CB), Canonical babbling ratio (CBR), Day-Long Audio Recording, Autism Spectrum Disorder (ASD), Elevated familial Likelihood for autism

1.1. Canonical Babbling

The first year of a child’s life is dominated by rapid changes in development including speech-language development. During the first six months of life, infants produce a wide variety of speech-like vocalizations collectively called protophones (Oller, 2000). Protophones include quasi-vowels, full vowels, squeals, growls, yells, whispers, and raspberries. By the end of the first year, typically between 6 to 10 months, infants make a significant developmental leap, producing well-formed canonical syllables and thereby make an entry into the canonical babbling stage (Oller, 1980). Canonical syllables are well-formed consonant-like sounds paired with vowel-like sounds, characterized by rapid transitions between them (Oller, 1980, 2000). Canonical babbling consists of series of canonical syllables such as [baba], [tita], or [gugugu]. Canonical babbling forms the basis for the first spoken words, which typically emerge around the first birthday.

In neurotypical infants, canonical babbling typically occurs between 6 and 10 months of age (e.g., Cychosz et al., 2021, Oller, 1980; Oller & Eilers, 1988; Stark, 1980) and is associated with later language outcomes (e.g., Delays in attaining canonical babbling have an impact on later language abilities (e.g., Fasolo et al., 2008; Ha & Oller, 2022; Iverson & Wozniak, 2007; Lang et al., 2019; McDaniel et al., 2018; Oller et al., 1998; Oller & Eilers, 1988; Plumb & Wetherby, 2013; Yankowitz et al., 2022). Deviances in canonical babbling, including later onset of canonical babbling, reductions in babbling rates, or lack of complexity of canonical babbling, have been reported in various developmental disorders, such as in infants with Williams syndrome (e.g., Masataka, 2001), autism spectrum disorder (e.g., Patten et al., 2014; Yankowitz et al., 2022), childhood apraxia of speech (e.g., Overby & Caspari, 2015), Fragile X syndrome (e.g., Belardi et al., 2017), Rett syndrome (e.g., Bartl-Pokorny et al., 2013), Angelman syndrome (e.g., Hamrick et al., 2023), tuberous sclerosis complex (e.g., Gipson et al., 2023) – suggestive of canonical babbling serving as a potential early behavioral marker for developmental disability diagnoses. Given its pivotal role in early speech-language development as the milestone that precedes first words and its potential as an early behavioral marker for various developmental disorders, it is no wonder that canonical babbling has been the subject of extensive research. Yet, there is variability in the way canonical babbling has been studied and reported. While an easy way of measuring canonical babbling is to ask parents if they have heard their child using such vocalizations (e.g., see the MacArthur Bates Communicative Development Inventory (Fenson, 2007), and the Communication and Symbolic Behavior Scale Caregiver Questionnaire (Wetherby & Prizant, 2002), the more commonly utilized approach, although labor-intensive, is collecting speech samples and hand-annotating them for canonical babbling.

There is wide methodological variance in how canonical babbling has been measured and reported from speech samples, which can limit reproducibility and comparability. For example, some studies report canonical babbling onset, i.e., the age when canonical babbles first appear (e.g., Oller et al., 1998). Other studies report canonical babbling ratio, the ratio of canonical syllables to total syllables (e.g., Lynch et al., 1995; Oller & Eilers, 1988). Canonical babbling ratios have been presented as mean differences between groups (e.g., Chapman et al., 2001; Ha & Oller, 2022; Long et al., 2024; Patten et al., 2014; Rvachew et al., 2005; Yankowitz et al., 2022) or as percentage of children above a certain cut off value/threshold (e.g., Belardi et al., 2017; Cychosz et al., 2021; Lynch et al., 1995; Patten et al., 2014; Yankowitz et al., 2022). When infants cross a canonical babbling ratio of 0.15 standard criterion they are considered to be in the canonical babbling stage (Lynch et al., 1995). Studies have also reported total canonical syllables or total number of consonants (Chenausky et al., 2017), percentage of canonical syllables (DeVeney et al., 2021; Paul et al., 2011), the number of directed and non-directed canonical syllables (Garrido et al., 2017) and canonical proportions (Cychosz et al., 2021).

1.2. Canonical Babbling in infants at an elevated likelihood for autism

Prospective longitudinal studies of infants at elevated familial likelihood for ASD have demonstrated that infants who go on to receive a diagnosis of ASD often show atypical behaviors in the first year of life, well before the diagnostic features typically consolidate into the full syndrome and well before typical age of clinical diagnosis (Zwaigenbaum et al., 2019). Behavioral markers such as eye contact, imitation, shared affect, orienting to name, response to caregiver language input, joint attention, gestures, repetitive patterns of behaviour and sensory modulation difficulties in the first year of life have strong associations with a later diagnosis of ASD (Riboldi et al., 2024; Tanner & Dounavi, 2021). However, as presented in Table-1, studies that have examined canonical babbling in infants later diagnosed with ASD or infants at elevated likelihood for autism are limited and findings are mixed (Chenausky et al., 2017; Chericoni et al., 2016; DeVeney et al., 2021; Garrido et al., 2017; Iverson & Wozniak, 2007; Long et al., 2024; Patten et al., 2014; Paul et al., 2011; Talbott et al., 2016; Yankowitz et al., 2022). One reason for these mixed findings could be the age ranges studied (see Table-1). Other reasons could be methodological variance, such as the variable under study (e.g., canonical babbling onset versus canonical babbling ratios), length of the sample studied, comparison groups employed to make conclusions (e.g., ASD compared with neurotypical versus elevated likelihood compared with low likelihood and so on), and type of the speech sample recording (e.g., free play versus clinician directed play sample versus day- long audio recordings.

Table-1:

Summary of studies evaluating canonical babbling in infant-toddlers less than 2 years of age at-an elevated likelihood for ASD

Authors Sample Sample size Approach Duration of sample annotated/Type of sample collected Data collection (Time points) Key variable^ Key variable findings
Iverson & Wozniak, 2007 Elevated familial likelihood for ASD Elevated familial likelihood for ASD=21 (n=19 for babble onset) No risk comparis on group=18 Prospective 45-minute home video Monthly between 5–14 mo Age of onset of reduplicated babble. More infants in Elevated familial likelihood for ASD (29%) than no risk comparison group (0%) were significantly delayed in onset of reduplicated babble
Paul et al., 2011 Elevated familial likelihood for ASD Elevated familial likelihood for ASD=28–38
Low familial likelihood for ASD = 20–31
Prospective First 50 speech like vocalizatio ns from home videos Thrice (6mo, 9mo, 12mo) Percentage canonical syllables Significant difference at 9 mo only (Elevated familial likelihood for ASD < Low familial likelihood for ASD)
Patten et al., 2014 Community ASD=23 Typically developing =14 Retrospective 5-minute home videos * 2 Twice (9-12mo, 15-18mo) Canonical Babbling Ratio Significant differences at both time points (mod to large effect sizes) ASD<typically developing ASD infants were less likely to reach canonical babbling stage between 9–12 months and produced lower canonical babbling ratio compared to TD infants
Talbott et al., 2016 Elevated familial likelihood for ASD Elevated familial likelihood for ASD=30 Low familial likelihood for ASD (LRC)=30 Prospective 20-minute lab and home videos Once (9mo) Canonical Syllables/min No differences
Chericoni et al., 2016 Community ASD=10 Typically developing (TD)=10 Retrospective 30 seconds to 2-minute home videos * 2 different sequences Thrice (0–6mo, 6–12mo, 12-18mo) Long reduplicative babbling and, 2-syllable babbling No differences
Chenausky et al., 2017 Elevated familial likelihood for ASD Elevated familial likelihood for ASD with confirmed ASD diagnosis=10 Elevated familial likelihood for ASD who did not receive ASD diagnosis=18 Low familial likelihood for ASD (LRC)=18 Prospective 30-minute lab videos Thrice (12mo, 18mo, 24mo) Number of consonants Number of consonants were lower in Elevated familial likelihood for ASD with confirmed ASD diagnosis than in the Low familial likelihood for ASD group (LRC) at all three time points
Garrido et al., 2017 Community ASD (fail on screening tool)=34 No ASD (pass on screening tool)=23 Retrospective 30-minute lab videos Once (13mo - 15mo) Number of directed and non-directed canonical syllables Significant differences in both number of directed canonical syllables (Autism<Non-ASD) and number of non-directed canonical syllables (Autism>non-ASD)
DeVeney et al., 2021 Elevated familial likelihood for ASD and premature + low birth weight Elevated familial likelihood for ASD and premature + low birth weight=13 Low familial likelihood for ASD (Low risk)=31 Prospective 20-min play observation done at home Twice (6mo, 12mo) Number of speech-like vocalizations, Percentage canonical syllables (CS) No differences at 6 mo & 12 mo; But significant differences in change between 612 mo in Percentage CS only (Elevated familial likelihood for ASD and pre-mature + low birth weight<ASD - Low risk)
Yankowitz et al., 2022 Elevated familial likelihood for ASD Elevated familial likelihood for ASD with confirmed ASD diagnosis=44 Elevated familial likelihood for ASD who did not receive ASD diagnosis=141 Low familial likelihood for ASD = 82 Prospective 10–30- minute lab videos Twice (6mo, 12mo) Canonical babbling ratio Type of babbling (reduplicativ e, variegated or neither) Significant differences in canonical babbling ratio at 12 mo only (Elevated familial likelihood for ASD with confirmed ASD diagnosis < Elevated familial likelihood for ASD who did not receive ASD diagnosis=Low familial likelihood for ASD) No differences between groups in reaching canonical babbling stage Reduplicated babbles - Elevated familial likelihood for ASD with confirmed ASD diagnosis < Elevated familial likelihood for ASD who did not receive ASD diagnosis < Low familial likelihood for ASD. No significant differences in rates of producing variegated babbles between the three groups.
Long et al., 2024 Elevated familial likelihood for ASD ASD= 44 Typically developin g (TD)=127 Prospective 8*5 segments with the highest number of infant utterances - Day long audio recordings Monthly between 1–13 mo Canonical Babbling Ratio No differences in Canonical Babbling Ratio at all time points between groups. Canonical babbling increases with age. Male ASD infants showed significant differences in canonical babbling ratio trajectories compared with female ASD infants, and to both male and female TD infants.
^

Key findings related to canonical babbling

Rows shaded in grey indicate studies conducted on elevated familial likelihood for ASD sample (i.e. infant siblings of autistic children)

1.3. Methods used to study and analyze canonical babbling in infants at elevated likelihood for autism

Two main approaches have been adopted by researchers to study canonical babbling in infants at elevated likelihood for autism or autistic children; (1) prospective - following infants from birth until they receive a diagnosis (e.g., Chenausky et al., 2017; DeVeney et al., 2021; Iverson & Wozniak, 2007; Long et al., 2024; Paul et al., 2011; Talbott et al., 2016; Yankowitz et al., 2022) (2) retrospective – analyzing already available recordings in infants diagnosed with ASD (e.g., Chericoni et al., 2016; Garrido et al., 2017; Patten et al., 2014). In both study designs, most studies have used short home videos (e.g., Patten et al., 2014) or short videos from lab assessments (e.g., Garrido et al., 2017) which do not necessarily capture an infant’s naturalistic home language environment. The length of the video recording and the time of day when the video is collected can also be limiting factors, as infant vocalizations fluctuate throughout the day (Warren et al., 2010). Finally, in the case of short home video recordings, there is a high chance of variability in the sample being analyzed. For instance, some videos may include various event types (e.g., birthday party, picnic, baby class), others may have varied ambient noise, audio/video recording quality, and recording duration.

Another way to collect data and analyze canonical babbling in by utilizing samples drawn from day-long audio recordings (e.g., the LENA®). Annotating samples drawn from day-long audio recordings offer several advantages over short home-videos or lab-based methods for assessing canonical babbling. These recordings are likely to provide a more naturalistic understanding of an infant’s linguistic abilities and home language environment, minimizing the influence of situational factors like special events (e.g., birthday parties) and capturing the fluctuations in vocalizations across the day (Warren et al., 2010). However, to best of our knowledge only one previous study has utilized day-long audio recordings in infants at elevated likelihood for autism (Long et al., 2024). This study evaluated canonical babbling ratios between infants at elevated familial likelihood for autism who received a diagnosis of autism and those with low familial likelihood who were neurotypical. The study did not report differences in canonical babbling ratios between groups (see Table-1), but male autistic infants showed significant differences in canonical babbling ratio trajectories compared with female autistic infants, and to both male and female neurotypical infants.

1.4. The present study

Given the importance of canonical babbling in infants at elevated likelihood for autism but mixed findings in the existing literature, the present study was undertaken to leverage potential advantages of using day long audio recordings to study canonical babbling. We used day-long audio recordings to compute canonical babbling ratios in 68 infants at elevated and low familial likelihood for autism. This elevated familial likelihood sample included n=11 infants who went on to receive a diagnosis of ASD (EL-ASD). We posed three questions – (1) Are there differences in canonical babbling ratios between infants at elevated likelihood for ASD who receive a diagnosis (EL-ASD), infants at elevated likelihood for ASD who do not receive a diagnosis (EL-ASD), and infants at low likelihood for ASD (LL-Neg)? (2) Are there differences in infants meeting the conventional criterion of reaching the canonical babbling stage (i.e., a 0.15 canonical babbling ratio) at 9 and 15 months, between EL-ASD, EL-Neg, and LL-Neg groups? (3) Are canonical babbling ratios at 9 and 15 months associated with later language abilities?

2. Methods

2.1. Participants

This study included 68 infants from the Infant Brain Imaging Study (IBIS)a longitudinal study of infants at elevated or low familial likelihood for ASD. Data were collected at the IBIS network’s clinical sites in the USA: University of North Carolina at Chapel Hill; Children’s Hospital of Philadelphia; Washington University in St. Louis; and University of Washington. Research protocols were approved by institutional review boards from respective clinical sites. All parents provided written, informed consent. Management of data collection, curation, and archiving was accomplished using the Longitudinal Online Research and Imaging System (LORIS) platform (Das et al., 2016).

2.1.1. Inclusion and Exclusion

Infants within the elevated likelihood group (EL, n=43) had an older sibling with a community diagnosis of ASD who also met criteria for ASD on the Social Communication Questionnaire (SCQ) (Rutter, M et al., 2003) and Autism Diagnostic Interview (ADI-R) (Lord et al., 1994). Infants in the low likelihood group (LL; n=25) had an older, neurotypical sibling, determined by parental interview on the Family Interview for Genetic Studies (FIGS) (Maxwell, 1992), and no first-degree relatives with a neurodevelopmental disorder. Infants were excluded from the study based on the following criteria: evidence of a genetic condition or syndrome, significant medical or neurological condition affecting development, significant vision or hearing impairment, birth weight <2000 grams or gestational age <36 weeks, perinatal brain injury secondary to birth complications or exposure to specific medication or neurotoxins during gestation, contraindication for MRI, predominant home language other than English, children who were adopted or half siblings, had a first degree relative with psychosis, schizophrenia, or bipolar disorder, or children who were twins.

2.2. Infant speech recordings

Infant speech samples were extracted from day long audio recordings using Language ENvironment Analysis (LENA®) digital language processors (DLP) (Xu et al., 2009). Each DLP weighed 2 ounces and was worn by the infant all day, using a vest designed to hold the DLP in place and provide optimal acoustic properties for recording (Ford et al., 2008). Families were mailed packets containing the DLP and vest or were provided packets during a research visit. Families were instructed to complete two days of recording (~16 hours per day; minimum of 8 waking hours), starting the recording when the infant woke up for the first time in the morning and letting the recording run uninterrupted throughout the day and into the night. Data collection methods have been described in detail previously (Swanson et al., 2019). LENA® Pro software suite V3.3.4 (Xu et al., 2009) was used to automatically process infant vocalizations; infant vocalizations are speech-like segments that are at least 0.6 seconds in duration and do not include non-speech sounds, e.g., crying, laughing, burping (Gilkerson & Richards, 2020). We used the LENA® generated child labels (CHN) as the first step to inform the selection of the sub-sample of each recording for hand annotation (i.e., a 5-minute non-contiguous high child vocalization sample – described in detail in a later section). Current software for generating automatic labels and counts is not as accurate as human listeners (Cristia et al., 2020, 2021), it provides an estimate of the number of child vocalizations that occur with 76% sensitivity (Xu et al., 2009) although the sensitivity was not determined for samples with autistic children.

2.3. Measures

The following developmental and diagnostic assessments were used to characterize the groups: The Mullen Scales of Early Learning (MSEL) (Mullen, 1995) is a standardized developmental assessment for children aged 0–68 months. It provides an Early Learning Composite (ELC) standard score which indexes overall development, and five subscale t-scores (fine motor, gross motor, visual reception, expressive and receptive language). The Autism Diagnostic Observation Schedule (ADOS) (Lord et al., 2000) is a semi-structured, observational play assessment of social interaction, communication, and repetitive behaviors for diagnosis and classification of ASD. Conventional scoring algorithms were applied to create a total calibrated severity score (CSS) (Gotham et al., 2009).

2.4. Procedure

Participants in the current study included all infants who met the following criteria: (1) at least one day of LENA® recording at age 9 months, and (2) cognitive and/or diagnostic assessments at 24 months (n=65) or 36 months (n=3). The sample size for this study was not predetermined. All available data from the multi-site longitudinal IBIS study that met the above specified criteria were included. Data collection occurred between 2015 and 2017, with annotation of speech samples and analysis conducted between 2017 and 2019. At the time, the sample size was comparable to other studies in the field (e.g., Chenausky et al., 2017; Iverson & Wozniak, 2007; Paul et al., 2011; Talbott et al., 2016; see Table 1). Participants were classified into the three groups based on cognitive and diagnostic assessments carried out at 24 or 36 months. The first group (EL-ASD infants) comprised 11 EL subjects (22.9%) who received a diagnosis of either DSM-IV-TR autistic disorder or pervasive developmental disorder not otherwise specified (referred to here as ASD), based on a clinical best estimate by licensed clinicians, supported by all available assessment data including the ADOS and MSEL. The remaining participants within the EL group were classified as EL-negative (EL-Neg; n=32). The third group was the LL-Negative group (LL-Neg: n=25). Of the remaining participants who did not meet criteria for ASD or signs of LD, they were classified into two groups namely, EL-negative (EL-Neg; n=32) and LL-negative (LL-Neg; n=25). All LL-Neg infants scored within normative ranges on the MSEL (> 85 Early Learning Composite Score). ADOS data was incomplete for 4 participants; 3 EL-Neg infants screened negative for ASD on the ADI-R; 1 LL screened negative on the SCQ. Demographic information appears in the results section. See Estes et al., 2015 for a detailed description of the diagnostic procedures.

Five infants had siblings in the study who also contributed home language recordings (4 EL sets of siblings, 1 LL set of siblings). A priori, one infant from each family was selected for inclusion in the current analyses based on the following criteria: (1) availability of diagnostic outcome data, or (2) if both infants had diagnostic outcome data, the infant with 9-month home language recording, or (3) if both infants had diagnostic outcome data, and both had 9-month home language recording, then the infant with 15-month home language recording was included in the analysis. Further, seven infants, five from the EL-Neg group and two from the LL-Neg group, were classified as having ‘signs of language delay (LD)’ based on a t-score below 35 (1.5 SDs below the mean) on the MSEL receptive language and/or expressive language subscales. Although these infants contributed home language recordings, they were excluded from the analysis a priori. More details for this classification are presented in (Swanson et al., 2017). Infants with signs of LD were not included in the subsequent analyses since, (i) we were interested to understand canonical babbling ratios in infants with a later diagnosis of ASD in comparison to infants with no concerns/signs of language delay at 24/36 months in the EL and LL groups, and (ii) the numbers within the LD groups were very small to run any statistical analysis.

2.4.1. Generating 5-minute samples from day long audio recordings for annotation

Once LENA® recordings were collected and processed, we used the LENA® Advanced Data EXtractor (ADEX) (Xu et al., 2008) program to identify the highest numbers of child vocalizations. Although the longitudinal study collected two days of recordings, we annotated only one day of the recording due to limited availability of annotators First, we divided the 16-hour day long audio recording to 5-second contiguous segments using study-specific code on Matlab (The MathWorks Inc. MATLAB Version: 9.13.0, 2022). Speech boundaries/segment boundaries determined by the LENA® system were not violated. Hence, some segments were longer than 5 seconds. Next, these contiguous 5-second segments were rank ordered from the highest to the lowest number of estimated child vocalization count. The number of child vocalization counts per 5-second segments were based on the automatic counts generated by LENA®. The first 60 5-second segments with the highest child vocalization count were concatenated to form a single ~5 min non-contiguous speech sample (i.e., 5sec*60segments = 5 minutes) for hand annotation. The method was chosen to ensure that the annotation sample included the target variable, child vocalizations. A similar approach of selecting periods with high vocalization activity has been utilized in a recent study (Meera et al., 2025)

2.5.2. Annotators

Trained annotators were blind to infant group status (i.e., EL or LL) and diagnostic outcome. Annotators were undergraduate students who volunteered their time in the study lab. All the segments within a recording were annotated by two individual annotators. In addition, a third annotator annotated the segments that the two annotators disagreed on and arrived at a final decision. Annotators were reliable with the first author and lead trainer, a speech-language pathologist with expertise in annotating child vocalizations, and with a co-author (KK), an undergraduate student who took part in annotation efforts for the entire study period (1.5 years). All annotators were native English speakers except for one, whose first language was Turkish, but they were fluent in English as well. Annotators reported no history of hearing loss or difficulty perceiving speech.

2.5.2. Annotator training

All annotators underwent a training period that lasted for 2 weeks (4 sessions of 1 hour each). In session one, the lead trainer provided definitions and audio examples of different annotation categories (e.g., canonical babbling). Definitions were based upon work done by Garrido and colleagues (Garrido, et al., 2017). A code book/annotator guide was developed and was available to the annotators during the training period. Sessions two to four included different aspects of annotation including the decision-making process when annotating each segment (see Figure-1). Annotators were trained to use the acoustic display of the waveform to identify various segments. Three recordings were designated as training recordings. The annotators and lead trainer completed 2 recordings together and discussed difficult segments in detail. Later the annotators completed one recording independently and the lead trainer evaluated the annotations. If the annotator and lead trainer agreed 80% of the time, the annotator was considered ready for annotation.

Figure 1:

Figure 1:

Annotation tree for identifying canonical syllables and other child vocalizations

2.5.4. Annotation procedure

A graphical user interface (GUI) based labeling toolkit was developed to facilitate annotation efforts (Pawar et al., 2017). The labels that could be assigned to each segment were: primary label (i.e., target child vocalization, CHD); secondary labels, i.e., child speech (CSP) and child non-speech (CNSP); and tertiary labels, i.e., canonical syllables and non-canonical syllables. Options under each label appeared as a drop-down menu and the annotator chose the best option. If the annotator could not decide on a certain label, they could leave it blank. Label fields were followed by a textbox that allowed the annotator to transcribe the utterance and note any difficulties they faced when making a judgment. When the annotator did not label a given segment, they were required to explain why a certain segment was not classified within the text box. See Figure-2 for an example of the GUI developed for annotating child vocalizations. This software is available from the authors upon request.

Figure 2:

Figure 2:

Graphical user interface developed for hand annotation of canonical syllables and other child vocalizations

A naturalistic listening approach was used to annotate infant speech and allocate labels to a given speech sample. Previous studies (Belardi et al., 2017; Patten et al., 2014) have found this approach to be reliable. The naturalistic listening approach is designed to have annotators listen in a manner similar to the way a parent or caregiver would listen to their child, hearing each utterance just once. We allowed for annotators to go back to the utterance a second time only if (a) it was longer than 2–3 seconds (less than 1% of utterances fell into this category), (b) if competing speakers were present (e.g., another child), and (c) if the utterance sounded muffled. If, after listening to a segment for a second time, the annotator continued to be unsure, they were instructed to fill the text box with their comment and move on to the next segment. 5-minute recordings were assigned to the annotators in a random order.

Two annotators individually annotated all vocalization segments by assigning them to the above-mentioned labels (CHD, CSP, canonical syllables, etc. – see fig-1). Sounds that were both speech and non-speech (e.g., talking while crying) were not categorized under any of the labels (this comprised less than 1% of vocalizations). Inter-annotator agreement, i.e., inter-rater reliability (pre-consensus), based on Cohen’s Kappa was 92.27%, k = 0.68 (moderate agreement).

2.6. Analytical plan

All analyses were performed using R software, version 4.0.2 (R Core Team. R: A Language and Environment for Statistical Computing - Version 4.0.2, 2020). Group differences in demographics were examined using ANOVA and Chi square tests. To address our first question, are there differences in canonical babbling ratios between groups at 9 and 15 months, linear mixed effect models were fit to the data. Fixed effects included (a) group i.e., EL-ASD, EL-Neg, and LL-Neg (LL-Neg was set as the reference group), (b) timepoint i.e., T1 and T2 (T2 was set as the reference group), (c) an interaction term between group and timepoint and (d) age was added as a co-variate. Random effects included a random intercept for each participant. Mixed-effects models were estimated using the ‘lmer’ function from the lme4 package. To address the next question, the likelihood of canonical babbling ratios meeting 0.15 criterion threshold, we used Fisher’s exact test for both time points (i.e., 9 and 15 months). Fisher’s exact test was used due to small sample size. Our third question was to examine if canonical babbling ratios were associated with later language skills (MSEL Receptive Language t-score, MSEL Expressive Language t-score) for which we used Pearson’s correlations. We evaluated these associations at both time points (9 and 15 months). Further, we did not consider commonly used control variables such as maternal education and sex of the infant since we were underpowered to run such analyses.

3. Results

3.1. Demographics

Analyses were performed to confirm that groups were matched on key demographic variables; no significant differences were present. Participant demographic information, number of participants at each visit, MSEL and ADOS scores at 24 months of age are provided in Table-2.

Table-2:

Descriptive and demographic data by group

Variable HL-ASDa HL-Negb LL-Negc Test Statistic, p-value
9m (n) 11 32 25
15m (n) 10 22 20
9 & 15m visit (n) 10 22 20
Mean (SD)
Age 9-month visit 10.23 (1.08) 9.83 (0.76) 9.88 (0.66) F = 1.07, p = 0.34
Age 15-month visit 15.64 (0.69) 15.84 (0.80) 15.68 (0.79) F = 0.33, p = 0.71
Months between 9 and 15m 5.38 (1.25) 6.04 (0.77) 5.98 (0.91) F = 1.93, p = 0.155
Male: n (%) 8 (72.73) 21 (65.63) 14 (56) χ2 = 1.075, p = 0.584
Female: n (%) 3 (27.27) 11 (34.37) 11 (44)
Maternal Education n (%)
High school diploma 5 (45.45) 8 (25.00) 3 (25.00) χ2 = 8.04, p = 0.089
College degree 3 (27.27) 14 (43.75) 7 (43.75)
Graduate degree 3 (27.27) 10 (31.25) 15 (31.25)
Paternal Education# n (%)
High school diploma 5 (45.45) 13 (40.63) 4 (16.67) χ2 = 4.83, p = 0.304
College degree 3 (27.27) 10 (31.25) 10 (41.67)
Graduate degree 3 (27.27) 9 (28.13) 10 (41.67)
Race and ethnicity n (%)
African American 0 1 (3.12) 2 (8)
White 8 (72.27) 26 (81.25) 18 (72)
Asian 1 (9.09) 0 0
More than one race 2 (18.18) 5 (15.62) 5 (20)
24 mo clinic visit^ Mean (SD)
Age at visit 25.65 (2.31) 24.68 (1) 24.92 (1.54) F = 1.69, p = 0.192
MSEL early learning composite 87.4 (21.35) 104.6 (14.9) 114.46 (11.31) F =11.82, p = 0.001 a<b<c
MSEL receptive language t-score 43.91 (18.32) 54.73 (7.26) 58.13 (7.05) F = 7.93, p = 0.001 a<b, a<c
MSEL expressive language t-score 47.36 (14) 50.6 (11.43) 54 (7.86) F = 1.55, p = 0.219
MSEL fine motor t-score 40.7 (8.82) 50.17 (9.84) 57.83 (8.68) F = 12.69, p = 0.001 a<b<c
MSEL visual reception t-score 44.6 (1181) 53.77 (9.67) 58.79 (9.62) F = 7.17, p = 0.001 a<b, a<c
ADOS calibrated severity score 7.1 (1.91) 2.76 (2.24) 2 (1.9) F = 22.18, p = 0.001 a<b, a<c

EL-ASD: Elevated likelihood for autism spectrum disorder; EL-Neg: Elevated likelihood but not diagnosed with autism spectrum disorder; LL-Neg: Low likelihood for autism spectrum disorder MSEL: The Mullen Scales of Early Learning; ADOS: The Autism Diagnostic Observation Schedule

#

data missing for 1 participant (EL-Neg group)

^

excluding 3 infants who were seen at 36 months (EL-Neg=2, LL-Neg=1)

3.2. Group differences in canonical babbling ratios and other child vocalization measures

Raw means, standard deviations and range of canonical babbling ratios and other child vocalization measures by group and for both timepoints are presented in Table-S1. Results from the linear mixed effects models (see Table-3) indicated that there were no significant main effects of group, timepoint, on canonical babbling ratios and other child vocalization measures. Likewise, there were no significant interactions between group and timepoint on canonical babbling ratios and other child vocalization measures.

Table-3:

Linear Mixed Effects Models estimating Canonical Babbling Ratios and other child vocalization measures

Canonical Babbling Ratio Canonical syllables Non-Canonical syllables Total vocalizations
Estimate (95% CI) SE DF p-value Estimate (95% CI) SE DF p-value Estimate (95% CI) SE DF p-value Estimate (95% CI) SE DF p-value
Intercept −0.11 (−0.81, 0.58) 0.35 110.73 0.746 −21.50 (−187.4, 144.48) 83.77 112.25 0.798 112.36 (−1.00, 225.72) 57.19 108.60 0.052 65.17 (−148.48, 278.81) 107.82 111.06 0.547
Time T1 0.10 (−0.17, 0.37) 0.14 112.95 0.467 12.60 (−52.49, 77.70) 32.85 110.25 0.384 −2.41 (−47.53, 42.72) 22.78 112.65 0.916 19.53 (−64.08, 103.13) 42.18 108.66 0.644
Group – EL-ASD −0.09 (−0.23, 0.04) 0.07 112.22 0.183 −13.42 (−47.72, 20.88) 17.30 105.93 0.440 6.62 (−16.42, 29.66) 11.63 112.72 0.571 −6.49 (−51.16, 38.18) 22.53 104.15 0.774
Group – EL-Neg −0.22 (−0.13, 0.09) 0.05 112.81 0.699 8.34 (−18.76, 35.44) 13.68 110.25 0.543 10.93 (−7.48, 29.33) 9.29 112.95 0.242 20.57 (−14.62, 55.75) 17.75 109.27 0.249
Group EL-ASD*Time T1 −0.01 (−0.19, 0.17) 0.09 55.33 0.932 3.08 (−35.54, 41.703) 19.27 54.63 0.874 −2.63 (−33.12, 10.77) 15.51 51.42 0.866 −0.43 (−48.76, 47.89) 24.12 55.44 0.986
Group EL-Neg*Time T1 −0.01 (−0.14, 0.13) 0.07 58.56 0.935 −14.88 (−44.82, 15.06) 14.95 56.46 0.324 13.17 (−37.12, 10.77) 11.95 54.90 0.275 29.27 (−66.77, 8.25) 18.73 56.96 0.124
Age 0.05 (0.01, 0.09) 0.02 111.41 0.046 6.72 (−3.74, 17.18) 5.28 111.73 0.206 −4.04 (−11.21, 3.13) 3.62 109.60 0.267 4.33 (−9.12, 17.79) 6.79 110.28 0.525

EL-ASD – Elevated likelihood for autism, EL-Neg – Elevated likelihood but not diagnosed with autism; LL-Neg – Low likelihood for autism (reference Group); Time T1 – Timepoint 1 (9-month visit), Time T2 – Timepoint 2 (15-month visit; reference group)

3.3. Meeting the conventional criterion of 0.15 canonical babbling ratio for the canonical babbling stage

At 9 months of age, 72.7% (n=8) of EL-ASD infants, 93.8% (n=30) of EL-Neg and all LL-Neg infants (100%, n= 25), met criteria for being in the conventional canonical babbling stage (see Figure-3 panel A). At 9 months there was a significant difference among groups in number of infants reaching the canonical babbling stage (p = 0.029). At 15 months of age nearly all infants met criteria for being in the canonical babbling stage. Only 1 infant each in the EL-ASD group (10%) and EL-Neg group (4.5%) did not meet criteria for being in the canonical babbling stage (see Figure-3 panel B). This difference did not meet statistical significance.

Figure 3:

Figure 3:

Canonical babbling ratios at 9 months (Panel A) and 15 months (Panel B)

3.4. Association between canonical babbling ratios at 9 and 15 months and later language skills at 24 months

Pearson’s correlations revealed no significant associations between canonical babbling ratio at 9 months and MSEL Receptive Language t-scores (r = 0.38, p = 0.69,) or MSEL Expressive Language t-scores (r = −0.07, p =0.54) at 24 months. Similarly, there were no significant associations between canonical babbling ratio at 15 months and MSEL Receptive Language t-scores (r = 0.88, p=0.38) or MSEL Expressive Language t-scores (r = 0.22, p= 0.11) at 24 months of age. The correlations by diagnostic group for canonical babbling ratios and later language abilities, and for other child vocalization (e.g., total syllables) measures and later language skills too failed to reach statistical significance and are presented in Tables S1 and S2.

4. Discussion

We examined canonical babbling ratios in infants at elevated familial likelihood for ASD who later met criteria for ASD (EL-ASD), comparing them to elevated likelihood infants who did not meet criteria for autism (EL-Neg), as well as typically developing infants not at elevated familial likelihood for ASD (LL-Neg). Our study contributes to a literature with mixed findings on this important early speech-language milestone. To the best of our knowledge, this is one of the few studies that has analyzed canonical babbling ratios in infants later diagnosed with ASD using day-long recordings, aside from a recent study by Long et al., 2024.

Although our sample is smaller than that of Long et al., (2024), like them, we did not find statistically significant differences in mean canonical babbling ratios between groups. This finding contradicts our hypothesis and differs from two previous studies that used audio-video data, where canonical babbling ratios were significantly different between EL-ASD, EL-Neg, and LL-Neg in a prospective study at 12 months (Yankowitz et al., 2022), a study design similar to ours, and between ASD and neurotypical at both 9–12mo and 15–18mo in a retrospective study (Patten et al., 2014). Both of these studies, which reported significant mean differences in canonical babbling ratios, were based on short lab or home recordings using audio-video data. In contrast, the current study and Long et al., (2024) employed day-long audio recordings and did not report significant differences.

A natural question is whether short audio-video data from lab/home recordings provide a better source or whether these recordings have overrepresented differences in canonical babbling ratios in infants at elevated familial likelihood for autism or those diagnosed with autism. A study by Bergelson et al., (2019) correctly highlights the distinct advantages and disadvantages of video and audio recordings, although not specifically for canonical babbling ratios but for noun usage, suggesting that sampling methods are crucial when designing a study. Addressing these questions will require comparisons across the same sample using both audio-video lab and home recordings, as well as day-long audio recordings, which represent an important future direction for the field. If lab-based audio-video recordings proved to be at least as informative, if not more, the time and effort invested in home or day-long audio recordings may not be necessary. Additionally, as discussed previously, it is important to acknowledge that directly comparing previous studies may be premature due to differences in methodology. For example, like Yankowitz et al., (2022) and Patten et al., (2014) our study too excluded non-speech-like utterances like squeals, growls etc., leading to a different pattern of arriving at what constitutes total syllables. In contrast, Long et al., (2024) included non-speech utterances and thus differed in calculation of canonical babbling ratios. Additionally, differences in the groups under study and the length of the recordings, could also have contributed to the discrepancies in findings.

In contrast to Long et al., (2024) and Yankowitz et al., (2022), our study found a significant difference in the number of infants reaching the canonical babbling stage (canonical babbling ratio > 0.15) across the three groups at 9 months, a finding similar to Patten et al., (2014). Specifically, infants in the HL-ASD group were less likely to have reached the canonical babbling stage by 9 months compared to infants in the other two groups. Further, while Long et al., (2024) reported that this criterion of 0.15 was not typically reached until 11–13 months, in both typically developing and ASD groups, our study demonstrated that 72% of infants in the EL-ASD group and 100% infants in the LL-Neg group had reached the 0.15 stage by 9-month visit (~9–10 months of age). This finding is likely to be a function of differing methodology used in both studies. The only way to determine whether there is a delay in reaching the canonical babbling stage when annotating using day-long audio recordings, or the reverse, based on the current criterion of 0.15, is by replicating the findings in an independent sample with similar annotating efforts.

Canonical babbling in the first year of life is often associated with later language abilities based on studies demonstrating this phenomenon in neurotypical children, ASD, and other developmental disorders (e.g., Fasolo et al., 2008; Ha & Oller, 2022; Lang et al., 2019; McDaniel et al., 2018; Oller et al., 1998; Yankowitz et al., 2022). Yankowitz et al., (2022) is one of the only studies that has examined the association between canonical babbling ratio and later language abilities in children at increased likelihood for autism. Their findings suggested that canonical babbling ratios at 12 months were positively correlated with expressive language abilities, as measured by the Mullen Scales of Early Learning (MSEL) at 24 months, as well as the number of words produced on the MacArthur-Bates Communicative Development Inventories (M-CDI) at 24 months. However, our study did not find these associations, which is a deviation from what we expected.

4.1. Limitations

There are several limitations to this study. The sample of autistic children is small, so the results must be interpreted with caution and require replication. Additionally, the lack of canonical babbling ratio differences between groups, while consistent with the only other study using day-long audio recordings (Long et al., (2024) may be due to the methodology we used for annotating 5-minute samples. It is possible that we captured the ‘chatty’ moments for all infants since we used a 5-minute sample with high child vocalization. By focusing on 5-minute segments with high child vocalization activity, we may have inadvertently captured “chatty” moments across all infants, regardless of diagnostic group. This could mask potential differences in canonical babbling ratio that might emerge in more representative or randomly selected samples. High-vocalization periods are likely to reflect optimal conditions for babbling, potentially inflating canonical babbling ratio values and reducing variability between groups. However, this possibility must be empirically tested by comparing it with the annotation of unselected vocalization samples, which is a future direction for this study. Another future direction for this work, and the field in general, is to develop automatic methods for counting canonical babbling, allowing the use of entire day-long audio recordings to estimate canonical babbling counts (e.g., Futaisi et al., 2019). An additional factor affecting the interpretation of our results is that all EL infants in this study were identified as being at elevated familial risk for ASD. Therefore, these findings cannot be generalized to a community-ascertained sample or to other elevated ASD likelihood samples (e.g., low birth weight). The sample predominantly comprises white families and thus does not reflect ethnic and racial diversity. Although we employed a third independent consensus annotator for all disagreed segments, we acknowledge that pre-consensus interrater reliability was lower than desired. In addition to canonical babbling ratios, information on the onset of babbling, rate of babbling, and the directedness of canonical babbling would have strengthened the findings. To understand the impact of methodological differences on canonical babbling ratio estimates, future research should compare day-long audio recordings with lab measures or home video samples. Analyzing the directedness of babbling and describing the consonant repertoire in canonical babbling are the next steps in this line of research. Although recent work by Oller and colleagues has demonstrated differences in canonical babbling between boys and girls in both neurotypical and ASD populations (Long et al., 2024; Oller et al., 2020, 2021, 2023), our study did not evaluate sex differences in canonical babbling among groups due to limited number if girls in the study especially in the EL-ASD group.

4.2. Conclusion and future directions

Despite several limitations, this study represents a well-characterized sample where all infants had diagnostic outcome data at 24 or 36 months. Although the findings indicate no group differences in mean canonical babbling ratios at 9 and 15 months and no association with later language abilities, we found that children in the EL-ASD group were less likely to reach the 0.15 canonical babbling ratio threshold for being in the canonical babbling stage by 9 months of age. This suggests that a diagnosis of ASD is associated with delays in canonical babbling for some children. To ensure robust findings on canonical babbling, particularly canonical babbling ratio in ASD, it is essential to establish standardized protocols for annotation and dissemination of results. Otherwise, as discussed in our study, comparisons between studies will remain challenging. In fact, this problem of methodological differences is not unique to ASD but also prevalent in the literature on early vocal patterns, including canonical babbling, across other developmental disorders, which presents a challenge in synthesizing findings across studies (Long et al., 2023).

Supplementary Material

1

Highlights:

  • Canonical Babbling Ratios (CBRs) is the ratio obtained by diving total number of canonical syllables (i.e., rapid transitions between consonant-like sounds with vowel-like sounds e.g., [gugugu]) by the total number of syllables.

  • CBRs were computed from day long audio recordings in infants at elevated familial likelihood for ASD who received an ASD diagnosis (EL-ASD), who did not receive an ASD diagnosis (EL-Neg) and infants at low likelihood for ASD (LL-neg) at 9 and 15 months of age.

  • There were no group differences in mean CBRs at 9 and 15 months and no association with later language abilities at 24 months.

  • Children in the EL-ASD group were less likely to reach the 0.15 CBR threshold for being in the canonical babbling stage by 9 months of age.

Acknowledgements

The authors thank the children and their families for their ongoing participation in this longitudinal study, as well as all research assistants who annotated the data.

Funding:

This work was supported by grants through the National Institutes of Health (R01-HD055741 PI Piven, R01-HD055741-S1 PI Piven, P50-HD10357305 PI Piven, U54-EB005149 PI: Kikinis) and the Simons Foundation (SFARI Grant 140209). Dr. Meera was supported by the Fulbright-Nehru Post-Doctoral Fellowship grant USIEF-2264/FNDPR/2017. The funders had no role in study design, data collection, analysis, data interpretation, or writing of the report.

Footnotes

Declaration of Competing Interest

The authors have no conflicts of interest to declare.

Author contributions: CRediT

1.Conceptualization: Shoba Meera, Meghan Swanson, Linda Watson, Joseph Piven

2.Data curation: Shoba S. Meera, Khavi Khuu, Meghan R. Swanson

3.Formal analysis: Shoba S. Meera, Divya Swaminthan, Kevin Donavan, Rahul Pawar,

4.Funding acquisition: Joseph Piven, Heather Hazlett, Annette Estes, Robert T. Schultz, Kelly Botteron, Stephen R. Dager

5.Investigation and project administration: Annette Estes, Lonnie Zwaigenbaum, Robert T. Schultz, Heather C. Hazlett, Tanya St. John, Juhi Pandey, Kelly Botteron, Stephen R. Dager, Joseph Piven

6.Software: Rahul Pawar, Mark Clements, David Anderson

7.Supervision: Meghan Swanson, Linda Watson, Joseph Piven

8. Writing – original draft: Shoba S. Meera

9. Writing – review & editing: Shoba S. Meera, Divya Swaminathan, Rahul Pawar, Lisa Yankowitz, Kevin Donovan, Khavi Khuu, Julia Parish-Morris, Steven F. Warren, Annette Estes, Lonnie Zwaigenbaum, Mark Clements, David V. Anderson, Robert T. Schultz, Heather C. Hazlett, Tanya St. John, Juhi Pandey, Natasha Marrus, Kelly Botteron, Stephen R. Dager, Meghan R. Swanson, Linda R. Watson, and Joseph Piven

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

Data Availability Statement

De-identified data will be made available upon reasonable request to the authors

References:

  1. Bartl-Pokorny KD, Marschik PB, Sigafoos J, Tager-Flusberg H, Kaufmann WE, Grossmann T, & Einspieler C (2013). Early socio-communicative forms and functions in typical Rett syndrome. Research in Developmental Disabilities, 34(10), 3133–3138. 10.1016/j.ridd.2013.06.040 [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Belardi K, Watson LR, Faldowski RA, Hazlett H, Crais E, Baranek GT, McComish C, Patten E, & Oller DK (2017). A retrospective video analysis of canonical babbling and volubility in infants with fragile X syndrome at 9–12 months of age. Journal of Autism and Developmental Disorders, 47(4), 1193–1206. 10.1007/s10803-017-3033-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Bergelson E, Amatuni A, Dailey S, Koorathota S, & Tor S (2019). Day by day, hour by hour: Naturalistic language input to infants. Developmental Science, 22(1). 10.1111/desc.12715 [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Chapman KL, Hardin-Jones M, Schulte J, & Halter KA (2001). Vocal development of 9-month-old babies with cleft palate. Journal of Speech, Language, and Hearing Research: JSLHR, 44(6), 1268–1283. 10.1044/1092-4388(2001/099) [DOI] [PubMed] [Google Scholar]
  5. Chenausky K, Nelson C, & Tager-Flusberg H (2017). Vocalization Rate and Consonant Production in Toddlers at High and Low Risk for Autism. Journal of Speech, Language, and Hearing Research : JSLHR, 60(4), 865–876. 10.1044/2016_JSLHR-S-15-0400 [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Chericoni N, de Brito Wanderley D, Costanzo V, Diniz-Gonçalves A, Leitgel Gille M, Parlato E, Cohen D, Apicella F, Calderoni S, & Muratori F (2016). Pre-linguistic Vocal Trajectories at 6–18 Months of Age As Early Markers of Autism. Frontiers in Psychology, 7, 1595. 10.3389/fpsyg.2016.01595 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Cristia A, Bulgarelli F, & Bergelson E (2020). Accuracy of the Language Environment Analysis System Segmentation and Metrics: A Systematic Review. Journal of Speech, Language, and Hearing Research, 63(4), 1093–1105. 10.1044/2020_JSLHR-19-00017 [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Cristia A, Lavechin M, Scaff C, Soderstrom M, Rowland C, Räsänen O, Bunce J, & Bergelson E (2021). A thorough evaluation of the Language Environment Analysis (LENA) system. Behavior Research Methods, 53(2), 467–486. 10.3758/s13428-020-01393-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Cychosz M, Cristia A, Bergelson E, Casillas M, Baudet G, Warlaumont AS, Scaff C, Yankowitz L, & Seidl A (2021). Vocal development in a large-scale crosslinguistic corpus. Developmental Science, 24(5), e13090. 10.1111/desc.13090 [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Das S, Glatard T, MacIntyre LC, Madjar C, Rogers C, Rousseau M-E, Rioux P, MacFarlane D, Mohades Z, Gnanasekaran R, Makowski C, Kostopoulos P, Adalat R, Khalili-Mahani N, Niso G, Moreau JT, & Evans AC (2016). The MNI data-sharing and processing ecosystem. NeuroImage, 124(Pt B), 1188–1195. 10.1016/j.neuroimage.2015.08.076 [DOI] [PubMed] [Google Scholar]
  11. DeVeney SL, Kyvelidou A, & Mather P (2021). A home-based longitudinal study of vocalization behaviors across infants at low and elevated risk of autism: Autism & Developmental Language Impairments. Autism & Developmental Language Impairments, 6, 1–18. 10.1177/23969415211057658 [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Estes A, Zwaigenbaum L, Gu H, St John T, Paterson S, Elison JT, Hazlett H, Botteron K, Dager SR, Schultz RT, Kostopoulos P, Evans A, Dawson G, Eliason J, Alvarez S, Piven J, & IBIS network. (2015). Behavioral, cognitive, and adaptive development in infants with autism spectrum disorder in the first 2 years of life. Journal of Neurodevelopmental Disorders, 7(1), 24. 10.1186/s11689-015-9117-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Fasolo M, Majorano M, & D’Odorico L (2008). Babbling and first words in children with slow expressive development. Clinical Linguistics & Phonetics, 22(2), 83–94. 10.1080/02699200701600015 [DOI] [PubMed] [Google Scholar]
  14. Fenson L (2007). MacArthur-Bates communicative development inventories. Paul H. Brookes Publishing Company; Baltimore, MD. [Google Scholar]
  15. Ford M, Baer CT, Xu D, Yapanel U, & Gray S (2008). The lenatm language environment analysis system. https://www.lena.org/wp-content/uploads/2016/07/LTR-03-2_Audio_Specifications.pdf
  16. Futaisi NA, Zhang Z, Cristia A, Warlaumont A, & Schuller B (2019). VCMNet: Weakly Supervised Learning for Automatic Infant Vocalisation Maturity Analysis. 2019 International Conference on Multimodal Interaction. https://www.academia.edu/96837027/VCMNet_Weakly_Supervised_Learning_for_Automatic_Infant_Vocalisation_Maturity_Analysis [Google Scholar]
  17. Garrido D, Watson LR, Carballo G, Garcia‐Retamero R, & Crais ER (2017). Infants at‐risk for autism spectrum disorder: Patterns of vocalizations at 14 months. Autism Research, 10(8), 1372–1383. 10.1002/aur.1788 [DOI] [PubMed] [Google Scholar]
  18. Gilkerson J, & Richards JA (2020). A Guide to Understanding the Design and Purpose of the LENA® System. LENA Foundation: Boulder, CO. https://www.lena.org/wp-content/uploads/2020/07/LTR-12_How_LENA_Works.pdf [Google Scholar]
  19. Gipson TT, Oller DK, Messinger DS, & Perry LK (2023). Understanding speech and language in tuberous sclerosis complex. Frontiers in Human Neuroscience, 17, 1149071. 10.3389/fnhum.2023.1149071 [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Gotham K, Pickles A, & Lord C (2009). Standardizing ADOS scores for a measure of severity in autism spectrum disorders. Journal of Autism and Developmental Disorders, 39(5), 693–705. 10.1007/s10803-008-0674-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Ha S, & Oller KD (2022). Longitudinal Study of Vocal Development and Language Environments in Infants With Cleft Palate. The Cleft Palate Craniofacial Journal, 59(10), 1286–1298. 10.1177/10556656211042513 [DOI] [PubMed] [Google Scholar]
  22. Hamrick LR, Seidl A, & Kelleher BL (2023). Semi-Automatic Assessment of Vocalization Quality for Children With and Without Angelman Syndrome. American Journal on Intellectual and Developmental Disabilities, 128(6), 425–448. 10.1352/1944-7558-128.6.425 [DOI] [PubMed] [Google Scholar]
  23. Iverson JM, & Wozniak RH (2007). Variation in vocal-motor development in infant siblings of children with autism. Journal of Autism and Developmental Disorders, 37(1), 158–170. 10.1007/s10803-006-0339-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Lang S, Bartl-Pokorny KD, Pokorny FB, Garrido D, Mani N, Fox-Boyer AV, Zhang D, & Marschik PB (2019). Canonical Babbling: A Marker for Earlier Identification of Late Detected Developmental Disorders? Current Developmental Disorders Reports, 6(3), 111–118. 10.1007/s40474-019-00166-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Long HL, Christensen L, Hayes S, & Hustad KC (2023). Vocal Characteristics of Infants at Risk for Speech Motor Involvement: A Scoping Review. Journal of Speech, Language, and Hearing Research : JSLHR, 66(11), 4432–4460. 10.1044/2023_JSLHR-23-00336 [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Long HL, Ramsay G, Bene ER, Su PL, Yoo H, Klaiman C, Pulver SL, Richardson S, Pileggi ML, Brane N, & Oller DK (2024). Canonical babbling trajectories across the first year of life in autism and typical development. Autism: The International Journal of Research and Practice, 13623613241253908. 10.1177/13623613241253908 [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Lord C, Risi S, Lambrecht L, Cook EH, Leventhal BL, DiLavore PC, Pickles A, & Rutter M (2000). The autism diagnostic observation schedule-generic: A standard measure of social and communication deficits associated with the spectrum of autism. Journal of Autism and Developmental Disorders, 30(3), 205–223. [PubMed] [Google Scholar]
  28. Lord C, Rutter M, & Le Couteur A (1994). Autism Diagnostic Interview-Revised: A revised version of a diagnostic interview for caregivers of individuals with possible pervasive developmental disorders. Journal of Autism and Developmental Disorders, 24(5), 659–685. 10.1007/BF02172145 [DOI] [PubMed] [Google Scholar]
  29. Lynch MP, Oller DK, Steffens ML, Levine SL, Basinger DL, & Umbel V (1995). Onset of speech-like vocalizations in infants with Down syndrome. American Journal of Mental Retardation: AJMR, 100(1), 68–86. [PubMed] [Google Scholar]
  30. Masataka N (2001). Why early linguistic milestones are delayed in children with Williams syndrome: Late onset of hand banging as a possible rate–limiting constraint on the emergence of canonical babbling. https://onlinelibrary.wiley.com/doi/abs/10.1111/1467-7687.00161
  31. Maxwell EM (1992). Family Interview for Genetic Studies (FIGS): A Manual for FIGS. Genome-Wide Association Study of Schizophrenia. [Google Scholar]
  32. McDaniel J, Slaboch KD, & Yoder P (2018). A Meta-Analysis of the Association Between Vocalizations and Expressive Language in Children with Autism Spectrum Disorder. Research in Developmental Disabilities, 72, 202–213. 10.1016/j.ridd.2017.11.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Meera SS, Swaminathan D, Venkata Murali SR, Raju R, Srikar M, Shyam Sundar S, Amudhan S, Cristia A, Pawar R, Rao A, Vasuki PP, Volme S, & Mysore A (2025). Validation of the Language ENvironment Analysis (LENA) Automated Speech Processing Algorithm Labels for Adult and Child Segments in a Sample of Families From India. Journal of Speech, Language, and Hearing Research, 68(1), 40–53. 10.1044/2024_JSLHR-24-00099 [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Mullen M (1995). Mullen Scales of Early Learning. In Pp. 58–64. Circle Pines, MN: AGS. 10.1007/978-1-4419-1698-3_596 [DOI] [Google Scholar]
  35. Oller DK (1980). The Emergence of the Sounds of Speech in Infancy. In Yeni-komshian GH, Kavanagh JF, & Ferguson CA (Eds.), Child Phonology (pp. 93–112). Academic Press. 10.1016/B978-0-12-770601-6.50011-5 [DOI] [Google Scholar]
  36. Oller DK (2000). The Emergence of the Speech Capacity. Psychology Press. 10.4324/9781410602565 [DOI] [Google Scholar]
  37. Oller DK, & Eilers RE (1988). The Role of Audition in Infant Babbling. Child Development, 59(2), 441–449. 10.2307/1130323 [DOI] [PubMed] [Google Scholar]
  38. Oller DK, Eilers R, Neal-Beevers A, & Cobo-Lewis A (1998). Late Onset Canonical Babbling: A Possible Early Marker of Abnormal Development. American Journal of Mental Retardation : AJMR, 103, 249–263. 10.1352/0895-8017(1998)103<0249:LOCBAP>2.0.CO;2 [DOI] [PubMed] [Google Scholar]
  39. Oller DK, Gilkerson J, Richards JA, Hannon S, Griebel U, Bowman DD, Brown JA, Yoo H, & Warren SF (2023). Sex differences in infant vocalization and the origin of language. iScience, 26(6), 106884. 10.1016/j.isci.2023.106884 [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Oller DK, Griebel U, Bowman DD, Bene E, Long HL, Yoo H, & Ramsay G (2020). Infant boys are more vocal than infant girls. Current Biology, 30(10), R426–R427. 10.1016/j.cub.2020.03.049 [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Oller DK, Ramsay G, Bene E, Long HL, & Griebel U (2021). Protophones, the precursors to speech, dominate the human infant vocal landscape. Philosophical Transactions of the Royal Society B: Biological Sciences, 376(1836), 20200255. 10.1098/rstb.2020.0255 [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Overby M, & Caspari SS (2015). Volubility, consonant, and syllable characteristics in infants and toddlers later diagnosed with childhood apraxia of speech: A pilot study. Journal of Communication Disorders, 55, 44–62. 10.1016/j.jcomdis.2015.04.001 [DOI] [PubMed] [Google Scholar]
  43. Patten E, Belardi K, Baranek GT, Watson LR, Labban JD, & Oller DK (2014). Vocal patterns in infants with autism spectrum disorder: Canonical babbling status and vocalization frequency. Journal of Autism and Developmental Disorders, 44(10), 2413–2428. 10.1007/s10803-014-2047-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Paul R, Fuerst Y, Ramsay G, Chawarska K, & Klin A (2011). Out of the mouths of babes: Vocal production in infant siblings of children with ASD. Journal of Child Psychology and Psychiatry, and Allied Disciplines, 52(5), 588–598. 10.1111/j.1469-7610.2010.02332.x [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Pawar R, Albin A, Gupta U, Rao H, Carberry C, Hamo A, Jones RM, Lord C, & Clements MA (2017). Automatic analysis of LENA recordings for language assessment in children aged five to fourteen years with application to individuals with autism. 2017 IEEE EMBS International Conference on Biomedical & Health Informatics (BHI), 245–248. 10.1109/BHI.2017.7897251 [DOI] [Google Scholar]
  46. Plumb AM, & Wetherby AM (2013). Vocalization Development in Toddlers With Autism Spectrum Disorder. Journal of Speech, Language, and Hearing Research, 56(2), 721–734. 10.1044/1092-4388(2012/11-0104) [DOI] [PubMed] [Google Scholar]
  47. R Core Team. R: A language and environment for statistical computing (Version 4.0.2). (2020). [Computer software]. R Foundation for Statistical Computing. https://www.R-project.org/ [Google Scholar]
  48. Riboldi EM, Capelli E, Cantiani C, Beretta C, Molteni M, & Riva V (2024). Differentiating early sensory profiles in toddlers at elevated likelihood of autism and association with later clinical outcome and diagnosis. Autism, 28(7), 1654–1666. 10.1177/13623613231200081 [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Rutter M, Bailey A, & Lord C,M. (2003). Social communication questionnaire. Los Angeles, CA: Western Psychological Services. https://cir.nii.ac.jp/crid/1370285712560901636 [Google Scholar]
  50. Rvachew S, Creighton D, Feldman N, & Sauve R (2005). Vocal development of infants with very low birth weight. Clinical Linguistics & Phonetics, 19(4), 275–294. 10.1080/02699200410001703457 [DOI] [PubMed] [Google Scholar]
  51. Stark RE (1980). Stages of Speech Development in the First Year of Life. In Child Phonology (pp. 73–92). Stark RE “Stages of speech development in the first year of life. Child Phonology, Vol. 1, Production; (edited by Yeni-Komshian G et al. ). 10.1016/B978-0-12-770601-6.50010-3 [DOI] [Google Scholar]
  52. Swanson MR, Donovan K, Paterson S, Wolff JJ, Parish-Morris J, Meera SS, Watson LR, Estes AM, Marrus N, Elison JT, Shen MD, McNeilly HB, MacIntyre L, Zwaigenbaum L, John T. St., Botteron K, Dager S, & Piven J. (2019). Early Language Exposure Supports Later Language Skills in Infants With and Without Autism. Autism Research : Official Journal of the International Society for Autism Research, 12(12), 1784–1795. 10.1002/aur.2163 [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Swanson MR, Shen MD, Wolff JJ, Elison JT, Emerson RW, Styner MA, Hazlett HC, Truong K, Watson LR, Paterson S, Marrus N, Botteron KN, Pandey J, Schultz RT, Dager SR, Zwaigenbaum L, Estes AM, Piven J, & IBIS Network. (2017). Subcortical Brain and Behavior Phenotypes Differentiate Infants With Autism Versus Language Delay. Biological Psychiatry. Cognitive Neuroscience and Neuroimaging, 2(8), 664–672. 10.1016/j.bpsc.2017.07.007 [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Talbott MR, Nelson CA, & Tager‐Flusberg H (2016). Maternal Vocal Feedback to 9‐Month‐Old Infant Siblings of Children with ASD. Autism Research, 9(4), 460–470. 10.1002/aur.1521 [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Tanner A, & Dounavi K (2021). The Emergence of Autism Symptoms Prior to 18 Months of Age: A Systematic Literature Review. Journal of Autism and Developmental Disorders, 51(3), 973–993. 10.1007/s10803-020-04618-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. The MathWorks Inc. MATLAB version: 9.13.0. (2022). [Computer software; ]. https://www.mathworks.com [Google Scholar]
  57. Warren SF, Gilkerson J, Richards JA, Oller DK, Xu D, Yapanel U, & Gray S (2010). What automated vocal analysis reveals about the vocal production and language learning environment of young children with autism. Journal of Autism and Developmental Disorders, 40(5), 555–569. 10.1007/s10803-009-0902-5 [DOI] [PubMed] [Google Scholar]
  58. Wetherby AM, & Prizant BM (2002). Communication and Symbolic Behavior Scales: Developmental Profile, 1st normed ed (pp. x, 177). Paul H Brookes Publishing Co. [Google Scholar]
  59. Xu D, Yapanel U, & Gray S (2009). Reliability of the LENA Language Environment Analysis System in young children’s natural home environment. Boulder, CO: Lena Foundation, 1–16. [Google Scholar]
  60. Xu D, Yapanel U, Gray S, & Baer CT (2008). The LENA language environment analysis system: The interpreted time segments (ITS) file. Infoture Inc.: Boulder, CO, USA. [Google Scholar]
  61. Yankowitz LD, Petrulla V, Plate S, Tunc B, Guthrie W, Meera SS, Tena K, Pandey J, Swanson MR, Pruett JR, Cola M, Russell A, Marrus N, Hazlett HC, Botteron K, Constantino JN, Dager SR, Estes A, Zwaigenbaum L, … IBIS Network. (2022). Infants later diagnosed with autism have lower canonical babbling ratios in the first year of life. Molecular Autism, 13(1), 28. 10.1186/s13229-022-00503-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  62. Zwaigenbaum L, Brian JA, & Ip A (2019). Early detection for autism spectrum disorder in young children. Paediatrics & Child Health, 24(7), 424–432. 10.1093/pch/pxz119 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1

Data Availability Statement

De-identified data will be made available upon reasonable request to the authors

RESOURCES