Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Aug 7.
Published before final editing as: Infant Behav Dev. 2026 Jun 30;84:102215. doi: 10.1016/j.infbeh.2026.102215

Body position classification using wearable sensors in infants with cerebral palsy

Kari S Kretch a,*, Florencia A Enriques a, John M Franchak b, Drew H Abney c, Christopher A Bell c, Christian M Jerry c, Katherine Lindig c, Grace Steffen c
PMCID: PMC13446499  NIHMSID: NIHMS2199839  PMID: 42379128

Abstract

Infants learn through everyday interactions with the physical and social environment. For infants with cerebral palsy (CP), motor impairments may disrupt everyday learning opportunities. How can we measure real-world motor behavior and learning opportunities in infants with CP? Machine learning models have been developed to quantify body position throughout a day in infants with typical development (TD) using wearable sensor data. However, these models have not been validated in infants with motor impairments. This study assessed the validity of body position classification using machine learning and sensors in infants with CP. Ten infants with CP (7–18 months; one session each) and 19 infants with TD (4–12 months; 45 sessions) wore four inertial sensors on their legs throughout a day. Ninety minutes were video recorded and manually scored for body position in five categories: supine, prone, sitting, standing, and held. Random forest classifiers were trained to predict body position from sensor-derived motion features. Models trained on datasets that varied in size (9, 45, 54 sessions) and group composition (CP, TD, CP and TD) were compared to determine the most accurate model for infants with CP. Larger training sets and training sets that included data from infants with CP were the most accurate; the final models achieved similar performance in CP and TD (86% and 89% accuracy), captured meaningful individual differences (ICCs = 0.682–0.999), and generated predictions that were correlated with motor skill assessments. Findings demonstrate that wearable sensors and machine learning can accurately classify real-world body position in infants with CP.

Keywords: motor development, posture, cerebral palsy, wearable sensors, machine learning

1. Introduction

Behavioral science is replete with laboratory studies. Standardized environments, controlled experiments, and clever manipulations have produced insights on plausible mechanisms of human behavior, cognition, and development. Additionally, clinical research in rehabilitation—recovery after an injury or medical condition or improvement of functional skills in individuals with disabilities—tends to measure intervention effectiveness using standardized assessments conducted in a controlled lab or clinic setting. However, actual behavior, cognition, development, and recovery take place in the real world. Thus, fully understanding these processes requires studying them outside the lab. For infant researchers, whose methods have always relied heavily on observation of naturalistic behavior due to the limited capacity of our participants to follow experimental instructions or provide verbal reports of their mental activities, it is a natural and necessary shift to focus on actual natural behavior “in the wild.” In exchange for experimental control, researchers studying infants in the real world gain ecological validity, insights into the vital roles of context and environment, and a wealth of theory-generating data on how complex developmental processes actually unfold (de Barbaro & Fausey, 2022). Similarly, the historical shift in rehabilitation research toward focusing on function and participation requires measuring function and participation in real-world settings (Rimmer, 2006). Therefore, the development of methods for measuring real-world behavior is essential for both developmental and rehabilitation science.

1.1. Importance of real-world body position in infant development

One meaningful aspect of everyday behavior in early development is body position: the physical configuration of the body in relation to gravity and environmental supports. A common approach to measuring body position is by classification into recognizable categories, such as standing, sitting, prone, supine, and being held (Franchak, 2019; Franchak et al., 2021, 2023). Body position can be measured for brief observation periods using manual video coding (Franchak et al., 2018; Kretch et al., 2022; Schneider et al., 2022; Thurman & Corbetta, 2017), and for longer periods of time using retrospective (Dudek-Shriber & Zelazny, 2007; Graciosa et al., 2024; Kuo et al., 2008; Majnemer & Barr, 2007) or real-time (Franchak, 2019; Franchak et al., 2024; Kretch, 2026; Kretch, Luna, et al., 2025) parent report; newer methods using wearable sensors will be described in the section below.

Body position is important for development because it constrains how infants interact with the world; these interactions, in turn, provide learning and interaction opportunities that shape developmental outcomes (Franchak, 2020; Iverson, 2022). For example, being held in arms or worn in a sling or carrier decreases crying, increases parent speech and responsiveness, and facilitates attachment (Anisfeld et al., 1990; Hunziker & Barr, 1986; Little et al., 2019; Mireault et al., 2018). Sitting—compared to lying prone or supine—improves infants’ ability to explore objects, expands their view of the environment, and enriches social interaction with caregivers (Franchak et al., 2018; Kretch et al., 2014, 2022; Kretch, Marcinowski, et al., 2025; Marcinowski et al., 2019; Schneider et al., 2022; Soska & Adolph, 2014; Soska et al., 2010). Standing further expands visual, social, and object-related interactions (Calabretta et al., 2022; Clearfield, 2011; Clearfield et al., 2008; Franchak et al., 2018; Karasik et al., 2012, 2011, 2014; Kretch et al., 2014; West & Iverson, 2021; Yamamoto et al., 2019). Crucially, the more time infants spend in particular positions throughout their everyday lives, the more exposure they have to the corresponding learning opportunities.

How infants distribute their time among different body positions changes with development. The amount of the day spent in held and supine positions decreases dramatically over the first year, while time in upright sitting and standing positions increases; prone time increases until about 9 months of age and then decreases again (Franchak, 2019). Body position is also closely linked to motor skill acquisition. For example, the amount of time spent sitting is associated with sitting skill (Franchak, 2019; Kretch, Luna, et al., 2025), the amount of time standing is associated with walking skill (Franchak, 2019; Franchak et al., 2024, 2026), and the peak in prone time aligns with the development of crawling (Franchak, 2019).

To summarize, motor skill development changes how infants are positioned during their everyday lives, and positioning influences developmentally-relevant learning and interaction opportunities. This cascade of events is believed to underpin documented links between the development of particular motor skills, like sitting and walking, and developmental improvements in other areas, like language and cognition (Oudgenoeg-Paz et al., 2015, 2012; Soska et al., 2010; Walle & Campos, 2014). In order to substantiate these mechanisms, measurements of developmental changes and inter-individual variability in real-world body position are essential.

1.2. Measuring natural infant behavior using inertial sensors

Although we have learned a great deal from the aggregated rough estimates of time per day in different positions through video observations and parent reports, these methods cannot provide precise time series of positioning episodes throughout an entire day. Recent years have seen a surge in the popularity of wearable technologies that passively measure infant behavior and health metrics over long periods time in natural environments (Zhu et al., 2015). These technologies produce high-resolution, continuous data that can address questions like the temporal structure or variability of activity over days or weeks or the co-occurrence of different types of behaviors, like body position and speech input (Rousey et al., 2026).

One common type of wearable technology is an inertial sensor, a small device that measures movement relative to an internal reference frame. These sensors include accelerometers to measure linear acceleration (changes in speed and/or direction), gyroscopes to measure angular velocity (speed and/or direction of rotation), and inertial movement units (IMUs) that contain both types of sensors. Rechargeable IMUs have the capacity to measure and store days’ worth of data on a single battery charge, and capture six synchronized, high resolution time series of linear (X/Y/Z) and rotational (roll/pitch/yaw) movement.

To detect meaningful movement patterns from raw IMU time series data, researchers can develop algorithms or select threshold values based on a priori knowledge about mechanical forces during movement. For example, Smith and colleagues determined acceleration cutoffs for detecting infant kicks (Smith et al., 2015); Hewitt and colleagues validated inclination cutoffs for detecting prone positioning (Hewitt et al., 2019). Applying these human-defined biomechanical algorithms to time series of accelerometer or IMU data has enabled the measurement of movement outcomes such as properties of upper and lower limb movements (Abrishami et al., 2019; Jeong et al., 2021; Shida-Tokeshi et al., 2018; Smith et al., 2015; Trujillo-Priego et al., 2017), active vs. sedentary states (Bruijns et al., 2020), and trunk orientation (Greenspan et al., 2021; Hewitt et al., 2019).

Inertial sensors can also be used to measure more complex, full body activities that are not easily detectable from simple biomechanical algorithms or cut points. An alternative method to classify raw IMU data into semantic descriptions of behavior is to employ supervised machine learning (Pargent et al., 2023) to perform activity recognition (Lara & Labrador, 2013). Raw data are segmented into dozens or hundreds of predictors, or features, paired with ground truth labels of the target activities. A machine learning algorithm uses these training data to “learn” which combination and weightings of features best predict the target activity classes, analogously to how a regression models an outcome variable from a set of predictors. Notably, machine learning algorithms create models purely based on statistical regularities in the data, rather than on biomechanical reasoning. The resulting model can then be used to predict the target activity from an unlabeled set of features. By applying these models to labeled data sets and comparing the model predictions to the ground truth labels, researchers can quantify the accuracy of the models and determine whether they can be generalized to novel, unlabeled data.

Machine learning has produced models that successfully classify activities like sitting, standing, and walking in adults (Narayanan et al., 2020) and children (Hendry et al., 2023; Nam & Park, 2013); these methods have recently been adapted to measure body position in infants (Airaksinen et al., 2022, 2025, 2020; Duda-Goławska et al., 2024; Franchak et al., 2021, 2023). Despite the irregularity and variability of infant movement, this recent work has demonstrated that infant body position can be accurately classified. Laboratory validation studies with infants and toddlers found approximately 90% accuracy in differentiating between multiple position categories, such as upright, supine, prone, sitting, and being held (Airaksinen et al., 2020; Franchak et al., 2021). Importantly, these methods have been validated in the natural environment over long recording periods to obtain estimates of the amount of time infants spend in various positions throughout a typical day (Franchak et al., 2023). Although accuracy (~85%) and Cohen’s kappa (~.75) scores were lower in natural settings compared to laboratory settings, correlations between ground truth and model predicted position durations were sufficiently strong to capture individual differences.

1.3. Measuring natural behavior in cerebral palsy

Importantly, body position data generated from these methods is associated with clinical assessments of infant motor skill (Airaksinen et al., 2022, 2025). Therefore, body position classification using wearable sensors may be useful for understanding development in infants with neuromotor disorders such as cerebral palsy (CP). CP is the most common cause of motor disability in childhood and affects about 1 in 345 children in the US (Durkin et al., 2016). CP is a disorder of posture and movement that arises due to a disturbance in early brain development (Rosenbaum et al., 2007). While CP by definition involves motor impairment, presentation is highly variable in severity (locomotor skills can range from walking unassisted to power wheelchair use), movement patterns (primary subtypes display spasticity/muscle stiffness, dystonia/excessive involuntary movement, or ataxia/poorly coordinated movements), and topography (some individuals are impacted only on one side of the body or only in the lower extremities) (Bax et al., 2005; Rosenbaum et al., 2007).

While these impairments—and their effects on motor skills across development in individuals with CP—are relatively well-understood, we know less about the everyday use of motor skills in the everyday lives of people with CP. A recent study using ecological momentary assessment (repeated parent reports of real-time body position throughout the day) found substantial differences in estimates of everyday body position duration in infants with or at risk for CP, compared to their peers with typical development (Kretch, 2026). In particular, infants with CP spent dramatically more time supine and less time standing than age-matched infants with typical development. These differences in everyday body position may limit opportunities for interaction with the physical and social world and compound developmental delays in other areas. Body position throughout everyday life is thus an important developmental variable; accurate measurement of body position over long periods of time in the home environment can provide insights into processes of atypical development and the effects of motor delays on developmental outcomes. Additionally, these measurements can be used in the future as meaningful real-world outcomes to assess the effects of interventions that seek to improve motor skills and everyday participation, either in clinical trials or on an individual basis in clinical practice.

A variety of clinically important behaviors have been measured in adults and children with CP using wearable sensors, including gait mechanics (Carcreff et al., 2020, 2018, 2017; Paraschiv-Ionescu et al., 2019), upper extremity use (Beani et al., 2025; Braito et al., 2018; Srinivasan et al., 2024), dystonic movements (den Hartog et al., 2022; Vanmechelen et al., 2023), and sedentary vs active time (Clanchy et al., 2011; Oftedal et al., 2015; Orlando et al., 2025; Xiong et al., 2022). Activity recognition has also been employed to measure body positions and activities like walking, cycling, and wheelchair propulsion in children and adults with CP, either using biomechanical algorithms like inclination cut-points (Claridge et al., 2019; Sato & Hirai, 2011; Sato et al., 2014) or using machine learning models (Ahmadi et al., 2018; Goodlich et al., 2020; Tørring et al., 2024).

The first two years of life comprise a critical developmental window for shaping trajectories in children with CP (Morgan et al., 2021). Understanding the everyday experiences (like body position) that influence development through this period is a high priority for research and early intervention practice. However, measuring body position in infants with CP presents challenges. Infants with CP may present with less stereotypical movement patterns than either typically developing infants or older children with CP. Additionally, impairments can vary widely between infants with CP and across the lifespan (Bax et al., 2005; Rosenbaum et al., 2007). To our knowledge, no studies to date have examined the accuracy of activity recognition, including body position classification, using wearable sensors in infants with CP.

1.4. Machine learning models: Generalizability to new populations

A key concern for machine learning models is their ability to generalize to new data sets. As in biological learning, properties of the data encountered during training affect how well the models perform when applied to novel data. Generalizability issues are often encountered when attempting to use activity recognition models trained on data from healthy, neurotypical populations to classify the behavior of individuals with mobility impairments (Rast & Labruyère, 2020). Put simply, models can only correctly classify activities to the extent that the relations between motion features and activities are similar to those learned in the training process. For example, a model may learn that motion features A, B, and C tend to have values around X, Y, and Z while neurotypical people are sitting. If another population sits in a different way that produces motion feature values outside of the range expected by the model, the model will perform poorly at recognizing sitting in the new population. This issue is encountered frequently in rehabilitation research with populations with neurological conditions: For example, activity recognition models trained from neurotypical participants do not perform well on individuals post-stroke (Capela et al., 2016; O’Brien et al., 2017; Oh et al., 2023) or with Parkinson’s Disease (Albert et al., 2012).

One solution is to train condition-specific models, using individuals with the target condition in the training set; this approach has been found to improve prediction accuracy in people post-stroke (Albert et al., 2012; O’Brien et al., 2017; Oh et al., 2023). However, this solution has limitations: Clinical populations tend to be heterogeneous and, for some conditions, difficult to recruit in large enough numbers to train a generalizable model. Another approach is to include both neurotypical and atypical individuals in the training data, which maximizes variability and sample size. Indeed, in a study of individuals post-stroke, the best accuracy (for people post-stroke) was seen with a model trained on a combined data set of healthy and stroke participants (Oh et al., 2023).

An understanding of the important characteristics of training sets is crucial for researchers attempting to apply existing machine learning models and to create new models. For those seeking to employ existing models with novel data, knowing which aspects of a training set support generalization can help guide selection of current models and identify when generating a new model is necessary. This is an important issue in infant research, because few off-the-shelf models are trained on infant data (de Barbaro et al., 2026), and even fewer on infants with atypical development. As new computational tools make the creation of machine learning models accessible to an increasing number of developmental scientists, more researchers will face decisions about the size and composition of the training sets. However, there is very little published literature to indicate which decisions might influence model generalizability.

1.5. Current study: Assessing body position classifiers in infants with CP

The aim of the current study was to assess the validity of body position classification using wearable sensors and machine learning in infants with CP. We focused on body position as a crucial activity classification taxonomy for infants because body position influences visual, manual, locomotor, and interpersonal interactions with the physical and social world (Calabretta et al., 2022; Clearfield, 2011; Clearfield et al., 2008; Franchak et al., 2018; Karasik et al., 2011, 2014; Kretch et al., 2014, 2022, 2023; Kretch, Marcinowski, et al., 2025; Long et al., 2021; Marcinowski et al., 2019; Rousey et al., 2026; Soska & Adolph, 2014; West & Iverson, 2021; Yamamoto et al., 2019) in ways that may have downstream effects on a variety of developmental outcomes (Eppler, 1995; Needham, 2000; Oudgenoeg-Paz et al., 2012; Soska et al., 2010; Walle & Campos, 2014). In brief, how infants are positioned during daily activities matters for development. A better understanding of everyday positioning in infants with CP could provide crucial insight on sequelae of motor impairments and opportunities for early intervention.

This study had three primary objectives. The first objective was to compare the accuracy of machine learning models trained on different data sets to classify body position in infants with CP. We used the same data collection and model training methodology that has previously demonstrated high accuracy in typically developing infants (Franchak et al., 2023), and incorporated a new sample of infants with typical development as a control test set to provide a baseline of accuracy for comparison. We trained classification models on training sets varying in size (9 or 45 infants) and composition (infants with CP, infants with typical development, or a combination) to determine the most accurate model for infants with CP. Once we determined the most accurate model for infants with CP, the second objective was to evaluate the ability of the model to predict individual body position categories and to capture individual differences in body position frequency. Third, we used the model to generate predictions for the frequency of the different body positions across the full day of measurement, and assessed associations with a clinical measure of infants’ motor skill in those positions as an assessment of convergent validity.

A key feature of this study is that the models were tested on data from unprompted, spontaneous activity in the natural environment. Prior work in activity recognition, especially in clinical populations (including CP), has largely relied on controlled laboratory observations of scripted activities for training and testing classification accuracy. However, accuracy is typically lower under the variable conditions that comprise real-world activity (Franchak et al., 2021). If the ultimate goal is to measure real-world activity, then high classification accuracy must be demonstrated in real-world environments.

2. Method

2.1. Participants

Two groups of participants—infants with cerebral palsy (CP) and infants with typical development (TD)—were tested using similar procedures at two different sites. All procedural differences between the sites were minor and are documented below. Demographic information for all participants is provided in Table 1.

Table 1.

Sample Demographics

Cerebral Palsy
(n=10)
Typical Development
(n=19)
Child sex
 Female 5 (50.0%) 11 (57.9%)
 Male 5 (50.0%) 8 (42.1%)
Child race
 White 5 (50.0%) 17 (89.5%)
 Black or African American 2 (20.0%) 0 (00.0%)
 Asian 1 (10.0%) 0 (00.0%)
 More Than One Race 1 (10.0%) 2 (10.5%)
 Unknown or Not Reported 1 (10.0%) 0 (00.0%)
Child ethnicity
 Not Hispanic 5 (50.0%) 19 (100%)
 Hispanic 5 (50.0%) 0 (00.0%)
Child gestational age
 Full term (≥37 wk) 5 (50.0%) 17 (89.5%)
 Moderate-late preterm (32–37 wk) 2 (20.0%) 2 (10.5%)
 Very preterm (28–31 wk) 2 (20.0%) 0 (00.0%)
 Extremely preterm (<28 wk) 1 (10.0%) 0 (00.0%)
Motor skills Visit (months)
4 (n=13) 6 (n=12) 8 (n=13) 12 (n=7)
 Roll (prone to supine) 6 (60.0%) 9 (69.2%) 9 (75.0%) 13 (100%) 7 (100%)
 Roll (supine to prone) 5 (50.0%) 5 (38.5%) 10 (83.3%) 13 (100%) 7 (100%)
 Sit1 7 (70.0%) 0 (00.0%) 7 (58.3%) 12 (92.3%) 7 (100%)
 Crawl2 3 (30.0%) 0 (00.0%) 0 (00.0%) 9 (69.2%) 6 (85.7%)
 Pull to stand3 3 (30.0%) 0 (00.0%) 0 (00.0%) 5 (38.5%) 4 (57.1%)
 Cruise4 3 (30.0%) - - - -
 Walk5 0 (00.0%) 0 (00.0%) 0 (00.0%) 0 (00.0%) 2 (28.6%)
Estimated GMFCS level 6
 I 3 (30.0%)
 II 3 (30.0%)
 III 1 (10.0%)
 IV 0 (00.0%)
 V 3 (30.0%)

Note. Motor skills were determined via semi-structured interview for the CP group and via specific items from the Early Motor Questionnaire (EMQ) in the TD group, resulting in slightly different definitions of some skills.

1

For the CP group, sitting was defined as “sit without support for 30 seconds”; for the TD group, we used EMQ item 6.1: “When placed into a sitting position on the floor, the child can sit independently without support (hands lifted).”

2

For the CP group, crawling was defined as “crawling across a room”; for the TD group, we used EMQ item 5.3: “When placed into a crawling position resting on hands and knees, the child will crawl forward for a few steps.”

3

For the CP group, pulling to stand was defined as “pulling to stand on furniture”; for the TD group, we used EMQ item 6.4: “When placed into a sitting position on the floor, the child can hold on to some furniture and pull into a standing position.”

4

For the CP group, cruising was defined as “walking sideways using furniture for support”; this skill was not assessed for the TD group.

5

For the CP group, walking was defined as “walking across a room”; for the TD group, we used EMQ item 7.4: “When placed into a standing position, the child will walk alone 4 or 5 steps independently with arms raised.”

6

Gross Motor Function Classification System levels determined using the “Before 2nd Birthday” section of the GMFCS-Expanded and Revised. However, note that these descriptors did not neatly classify the infants in this sample because each level specified skills that did not always co-occur. Therefore, we assigned infants to the level that matches their sitting skill, even though this may be inconsistent with their rolling or crawling skill. We note that GMFCS is not expected to be stable before the age of 2, so the assigned GMFCS levels are only estimates at this point in development.

2.1.1. Infants with cerebral palsy

Ten infants with a parent-reported CP diagnosis (one 8-month-old, one 10-month-old, two 12-month-olds, one 13-month-old, two 15-month-olds, two 17-month-olds, and one 18-month-old) participated in the study for a single session at their homes in the greater Los Angeles area. Participants were primarily recruited from Children’s Hospital Los Angeles, social media ads, and referrals from local clinicians or researchers. Families received a $50 Amazon gift card after the study completion.

2.1.2. Infants with typical development

We leveraged an available subset of data from an ongoing longitudinal study focused on movement and sleep behaviors in infants with TD. The sample included 19 unique infants with home visits conducted at 4 (n = 13), 6 (n = 12), 8 (n = 13), and/or 12 (n = 7) months of age, resulting in a total of N = 45 sessions. Four infants contributed data to one session only, six infants had two sessions, seven infants had three sessions, and two infants had all four sessions. Participants were primarily recruited from the Athens, Georgia area using flyers posted in the community, community events, and participant referrals. Families received $100 cash for their first and fourth visit, and $200 cash for their second and third visits, for a total of $600 across four visits. Families also received an additional $25 for every referral provided who enrolled in the study and completed at least one visit.

Note that the age ranges differed between the groups (4–12 months for the infants with TD and 7–18 months for the infants with CP). However, because the motor skills of the infants with CP were delayed, the distribution of motor milestone achievement was similar between the groups (see bottom rows of Table 1). For example, most (all infants with CP and all but two TD sessions) infants were non-ambulatory; a bit over half (70% of infants with CP and 58% of TD sessions) were able to sit independently. Therefore, we would expect the infants in both groups to demonstrate a similar range of motor behaviors during natural activity.

2.2. Wearable sensors and garment

We constructed a set of custom leggings to secure four IMUs (Axivity AX6; 2.3×3.25×0.89 cm, 0.11 kg) in place throughout the day of recording (Figure 1). The leggings contained four custom-sewn tightly fitted pockets to prevent the IMUs from rotation or displacement: Two were placed on the lateral left and right hip, and the other two were placed on the lateral left and right ankle, using a standardized orientation. Multiple sizing options ensured the leggings fit each infant properly and were not too long or too loose. The sensors had enough battery capacity to last for about 12 hours collecting at 50 Hz.

Figure 1.

Figure 1

Infant wearing leggings with embedded sensors. Arrows represent sensor locations.

We chose this configuration of sensors on the legs only because prior work has demonstrated that leg sensors are sufficient to classify the five body positions in infants with TD (Franchak et al., 2023). Limiting our sensor placement to the legs allows us to use a single, easy to use garment that minimizes complexity for parents and does not affect infant behavior.

2.3. Procedure

At each session, a trained researcher visited the participants’ homes in the morning for about one hour. During the visit, the researcher explained the purpose of the study to the parents, confirmed the infant’s date of birth and diagnosis to verify eligibility, and obtained written informed consent. To obtain further information on motor milestone achievement and motor skill repertoire in the CP sample, the researcher interviewed the parent to determine whether the infant was able to roll, sit, pull to stand, crawl, cruise, and walk. For the TD sample, similar parent report motor skill data were obtained at each visit using the Early Motor Questionnaire (Libertus & Landa, 2013).

After consent was obtained, the researcher began recording with a tripod-mounted video camera. At the beginning of the recording, the researcher dropped the leggings on the floor in view of the camera to create a synchronization point in the sensor and video data. The researcher then asked the caregiver to place the leggings on the infant.

For the CP sample, the researcher administered the Alberta Infant Motor Scale (AIMS), a standardized assessment of gross motor skills in four postures—supine, prone, sitting, and standing—for infants from birth to independent walking (Piper & Darrah, 2022). We chose to use the AIMS as our quantitative measure of motor skill for convergent validity because it is designed for the infancy period (0–18 months or until skilled independent walking), provides specific measures of skill in each of the measured body positions, and contains enough items in each position to be sensitive to individual differences. The video-recorded AIMS assessment served two purposes: It allowed us to measure infants’ posture-specific motor skill, and it provided video of infants in each position for training and testing the classification models. The AIMS contains 58 items across the four postures, ordered by difficulty/maturity, to be marked as observed or not observed. Infants were video recorded in the four AIMS postures for at least two minutes each. If the infant was unable to maintain any of the postures independently, the researcher or parent provided support. For skills on the AIMS list that were not observed spontaneously, the researcher prompted the behaviors through toy placement and verbal cues without providing physical assistance. In addition to the AIMS postures, the researcher also recorded the parent holding and carrying the infant and the infant placed in any positioning devices the family used throughout the day (e.g., highchairs, exersaucers), for two minutes each. For the TD sample, the researcher did not perform the full AIMS assessment but did prompt the parent to place the infant in each of the positions, to hold and carry the infant, and to place them in positioning devices.

After the AIMS assessment and prompted activities, infants were recorded in natural activity. For the CP sample, the researcher left the home, and parents were instructed to go about their natural activities while keeping the camera recording for one hour, moving or adjusting the camera as needed to ensure a full view of the infant. For the TD sample, the researcher remained in the home and recorded infants’ natural activity: Parents were instructed to “play how you normally might” for 15 minutes of recording. Including the AIMS/prompted activities and natural activity recording, videos were approximately 90 minutes in duration for the CP sample and 40 minutes for the TD sample.

Parents were instructed to keep the leggings on the infant for the rest of the day until bedtime (for the CP sample) or for at least five hours (for the TD sample), and to only remove them for bathing, diaper changes, and trips out of the house. Parents were asked to record the timing of these legging removals as well as naptimes so that those time periods could be removed from analysis of the full-day data.

2.4. AIMS scoring

The AIMS assessment was scored from video following the session. A trained rater (licensed physical therapist with expertise in early intervention) identified each target skill observed in each position. Per the AIMS administration guidelines, a final score for each position was obtained by summing all the observed skills as well as the skills that precede the least mature observed skill. Thus, supine scores can range from 0–9, prone scores can range from 0–21, sitting scores can range from 0–12, and standing scores can range from 0–16.

2.5. Body position annotation

To obtain a time series of ground-truth body position, trained research assistants coded the videos frame by frame using Datavyu software (datavyu.org). Body position had five different categories. Supine was defined as lying on the back or side. Prone was defined as any face down position including lying on the stomach or on hands and knees, including crawling or stationary prone positioning. Sitting included instances where the infant was sitting on the ground or other surface (furniture or caregiver’s lap) with the bottom resting on the surface of support. Sitting also included times when the infant was sitting upright in a device or restraint such as a highchair or car seat. Standing was coded when the infant was standing on the ground or other surface on one or two feet, or tall kneeling, including cruising/walking or stationary standing positioning. Held included times when the infant was held by a caregiver off the ground surface. Categories were mutually exclusive and exhaustive; in other words, every moment was classified into one of the five categories. Coders scored the onset and offset of each instance of each body position, producing a time series where each video frame had an associated position code. Time when the baby was off camera or the leggings were removed or interfered with were noted as missing and removed from analysis.

For the CP sample, a primary coder annotated the full videos, and a second coder completed annotation for 25% of each video to determine inter-rater reliability. For the TD sample, a primary coder annotated all videos, and a second coder completed annotation for a random third of the videos. Percent agreement ranged from 60–99% for the CP sample (M = 93%) and 78–98% for the TD sample (M = 90%); kappa ranged from 0.49-.99 for the CP sample (M = 0.91) and 0.75–0.97 for the TD sample (M = 0.88). (Note that for the CP sample, nine of the ten infants had excellent interrater reliability—92–99% agreement and 0.90–0.99 kappa. The tenth infant spent large amounts of time in a reclined position that was difficult to classify between supine and sitting, resulting in low reliability. The coders and a tiebreaker viewed the disagreements together and determined the final codes.)

2.6. Machine learning classification

We fit a series of machine learning models to predict ground truth body position from the IMU motion data following previously described modeling procedures (Franchak et al., 2021, 2023). Each session’s accelerometer and gyroscope time series were aligned to the body position code time series based on the synchronization point. A data set combining all sessions from infants with CP and with TD was pre-processed and partitioned into different training and testing sets, as described below and in Supplemental Material 1.

Motion features and ground truth labels were calculated from the synchronized time series data in 4-s sliding windows starting every second (e.g., 9:41:01–9:41:05, 9:41:02–9:41:06, 9:41:03–9:41:07). As in past activity recognition work (Franchak et al., 2021, 2023), using 4-s windows allowed us to calculate time-varying motion features from the 200 samples of 50 Hz raw acceleration and gyroscope data. There were 436 motion features that summarized the raw movement signals in each window. Of those, 240 features were summary statistics (minimum, maximum, sum, mean, median, SD, kurtosis, skew, 25th percentile, 75th percentile) of a single sensor’s data in a single orientation (e.g., left hip acceleration in the X direction, right ankle gyroscope pitch). In addition, 196 cross-sensor and cross-orientation motion features were derived by calculating the magnitude, difference, and correlation between different signals (e.g., mean acceleration in the X direction across all 4 sensors, correlation between left and right ankle roll). Each window was assigned a position code only if the infant spent at least 75% of the window (> 3 s) in a single position; ambiguous windows in which infants were in two different positions, such as 2 s sitting and 2 s upright, would add noise to model training and evaluation so they were excluded.

The resulting windowed dataset contained 436 motion features that served as predictors and a single ground truth label that could be used to train supervised machine learning models. We fit models using the Random Forest algorithm (Breiman, 2001) implemented in the DecisionTree package (Sadeghi et al., 2022) in the Julia programming language (Bezanson et al., 2017) version v1.11. Initial testing indicated that hyperparameter tuning did not lead to appreciable differences in model performance (~1% overall accuracy), thus, we used default values for model hyperparameters except for the number of trees in each model. We used 150 trees in the model evaluation fits to reduce computation time and 250 trees for the final model to prioritize accuracy.

We partitioned the combined dataset containing 10 sessions from infants with CP and 45 sessions from infants with TD into training and testing datasets to evaluate models that varied in composition and training data set size (see Figure 2). Model composition allowed us to test the generalizability of models that contained only infants with CP (CP-trained), only infants with TD (TD-trained), and models with both groups of infants (TDCP-trained). We varied training data set size by comparing small models that had 9 or 10 sessions, medium models that had 44 or 45 sessions, and full models that contained 54 sessions. Crucially, every training/testing split partitioned entire participant sessions into either the training or testing set, because including a session in both training and testing would inflate accuracy and undermine claims about model generalization. By ensuring that tested session’s data were held out completely, our model evaluation statistics describe how well the model should perform on a newly collected participant or session. Full details on the partitioning procedures can be found in Supplemental Material 1.

Figure 2.

Figure 2

Training/testing data set partition examples for (A) small models with 9 sessions in the training set and (B) medium models with 45 participants in the training set. In both examples, infants with CP are colored blue and infants with TD are colored orange. Test participants are highlighted by a magenta rectangle, and training datasets are bounded by dashed rectangles. Panel A shows a 9-ppt CP-trained model (blue rectangle) and 5 different 9-participant TD-trained models used to predict the test one infant with CP. Panel B shows a 45-participant TD-CP trained model used to test one infant with CP and 9 infants with TD.

2.7. Model evaluation

We calculated multiple standard performance metrics to evaluate model performance. A time series of model-predicted labels was obtained by applying the predict function in R to the motion features of each held-out testing dataset. Predicted labels for each window were aligned with ground truth labels to allow for evaluation—did the predicted label for a given 4-s window match the ground truth label. We report three overall performance statistics—accuracy, kappa, and Macro F1-score—for every training/testing partition. Accuracy was the overall percentage of windows in which the model-predicted and ground truth labels matched. The Cohen’s kappa statistic is a measure of agreement that adjusts for the base rate of each category. Finally, the Macro F1-score was calculated as the unweighted average of the F1-score for each body position category—the harmonic mean of the sensitivity and positive predictive value for that category. The Macro F1-score helps to account for imbalance in base rates and bias in detecting each class. To assess whether performance differed between model types (CP-trained vs. TD-trained for small models, TD-trained vs. TDCP-trained for medium models) and between groups (TD vs. CP), we performed linear mixed models using the lmer function from the lmerTest package in R. Models contained the following predictors: group (0 = CP, 1 = TD), model type (for small models: 0 = CP-trained, 1 = TD-trained; for medium models, 0 = TDCP-trained, 1 = TD-trained), and the group*model type interaction. All predictors were centered at the mean so that coefficients represent main effects in the presence of interactions. Because sessions had data for multiple model types, a random intercept was included at the session level. We used t-tests to assess whether performance of the full models differed between CP and TD groups.

In addition to macro performance, we report by-class performance metrics and prevalence for the full model for the LOOCV prediction for each participant with CP. Each by-class metric describes different biases in the types of errors that models make when classifying events. Sensitivity was calculated based on the percentage of true outcomes for a given class were correctly identified (e.g., percentage of correctly detected sitting events out of true sitting events). Positive predictive value (PPV) was the percentage of predicted events that truly belonged to that category (e.g., percentage of times the model predicted sitting that the infant was sitting). Finally, specificity was calculated as the percentage of negative events for a class that were correctly classified (e.g., the percentage of not-sitting times that the infant was classified in a position other than sitting).

Finally, we compared the model predicted prevalence versus the ground truth prevalence to identify how well models could represent individual differences in the rate of each position during the video-recorded portion of the study. Prevalence was calculated as the percentage of windows the infant was predicted to be in a certain position (predicted prevalence) and actually coded to be in that position (ground truth prevalence). We calculated Intraclass Correlation Coefficients (ICCs) for each position to quantify the agreement between predicted prevalence and ground truth prevalence.

2.8. Full day predictions

Full day predictions were created for each of the 10 infants with CP based on the final 55-participant TDCP-trained model. Caregiver-logged nap times and other legging removal times were used to exclude predictions from the full day calculations based on clock time. The timelines in Supplemental Material 3 show each infant’s full day predictions; blank times represent naps and other times that data were not recorded. We examined correlations between the proportion of windows classified as prone, sitting, and standing in the full-day data and the corresponding AIMS score (prone, sitting, and standing). Because supine placement does not require motor skill on the part of the infant, we examined associations for AIMS prone, sitting, and standing scores only.

2.9. Data sharing

All deidentified data and analysis code is publicly available at [URL to be added].

3. Results

3.1. Description of body positions in the training data

Before describing the model performance results, we examine the distribution of body positions in the training data to ensure that each position was sufficiently represented. Figure 3 shows the ground truth (human-coded) proportion of time spent in each position category during the video recordings. Because the data collection procedure included prompted activities, every infant in both groups spent some time in each position. However, the amount of time in each position was unbalanced, with more time spent in supine and sitting than prone, standing, or held. Accordingly, we report kappa and F1 statistics in our subsequent model comparison results since overall accuracy does not account for class imbalance.

Figure 3.

Figure 3

Proportion of the video-recorded observation (comprising the training or testing data) that infants were in each body position. Each symbol represents one infant’s data, and the horizontal lines represent the group mean.

It is important to note that these distributions include some mixture of prompted activities and natural activity, and the natural activity portion was considerably longer (i.e., a larger proportion of the video recording) in the CP sample. Thus, these distributions should not be interpreted as representative of infants’ everyday positioning routines or differences in everyday positioning activity between CP and TD.

3.2. Comparison of models varying in training set size and composition

3.2.1. Small models

The first set of models aimed to evaluate the performance of models trained on data from infants with CP to predict body position in infants with CP. As a baseline for comparison, we also trained and tested models on infants with TD. To determine the impact of the novel (CP) training set composition, we held the training set size constant and compared the performance of these models to the performance of models trained on similarly sized samples of infants with TD (Figure 2a). To determine the impact of the novel (CP) test set composition, all small models were also tested on infants with TD.

Figure 4a depicts the three performance metrics—accuracy, kappa, and F1-score—for each combination of training and testing sets. Overall, the CP-trained models performed well at predicting body position in infants with CP, with accuracy M = 84.2%, kappa M = 0.753, and F1 M = 0.788. Performance varied widely between individuals, with a minimum accuracy of 62.8%, minimum kappa of 0.508, and minimum F1 of 0.634. In comparison, the five folds of small TD-trained models performed poorly at predicting body position in infants with CP, with M accuracy of the five folds ranging from 64.1–79.4%, M kappa ranging from 0.519–0.690, and M F1 ranging from 0.639–0.763. However, all small models—the CP-trained model and the five folds of TD-trained models—performed well at predicting body position in infants with TD, and their performance was similar to that of the CP-trained models tested on infants with CP. The linear mixed model confirmed significant effects of group and the group*model type interaction for all three metrics (Table 2).

Figure 4.

Figure 4

Three overall performance metrics—accuracy, kappa, and F1 score—for (A) small models and (B) medium models. Each symbol represents one infant’s metrics, and the horizontal lines represent the group mean. Blue symbols and lines represent results when testing models on infants with CP, and orange symbols/lines represent results when testing models on infants with TD. For the small models, CP model type represents the models trained on data from 9 or 10 infants with CP, and TD1–5 model types represent the five folds of models trained on data from 9 or 10 infants with TD. For the medium models, TD model type represents the models trained on 44 or 45 infants with TD, and TDCP1–5 model types represent the models trained on 9 or 10 infants with CP plus 35 or 36 infants with TD.

Table 2.

Linear mixed model results for performance metrics

Small Models Medium Models
Accuracy
Intercept 0.834
[0.808, 0.861]
0.884
[0.859, 0.908]
Group 0.108*
[0.040, 0.176]
0.043
[−0.020, 0.107]
Model Type −0.006
[−0.031, 0.019]
−0.013**
[−0.020, −0.007]
Group*Model Type 0.138**
[0.075, 0.202]
0.071**
[0.0534, 0.088]
Kappa
Intercept 0.757
[0.720, 0.795]
0.823
[0.786, 0.860]
Group 0.157*
[0.061, 0.254]
0.080
[−0.016, 0.177]
Model Type −0.006
[−0.037, 0.026]
−0.018**
[−0.027, −0.009]
Group*Model Type 0.181**
[0.100, 0.263]
0.096**
[0.072, 0.119]
F1
Intercept .803
[0.779, 0.827]
0.857
[0.832, 0.881]
Group .099*
[0.038, 0.160]
0.050
[−0.014, 0.114]
Model Type .006
[−0.014, 0.026]
−0.009*
[−0.014, −0.003]
Group*Model Type .1058**
[0.054, 0.158]
0.060**
[0.046, 0.075]

Note. Coefficient estimates and 95% confidence intervals from linear mixed models. Group is coded as CP = 0 and TD = 1. For the small models, Model Type is coded as CP-trained = 0, TD-trained = 1. For the medium models, Model Type is coded as TDCP-trained = 0, TD-trained = 1. All predictors are centered at the mean. All models include a random intercept for session.

*

p<.01.

**

p<.001.

3.2.2. Medium models

The second set of models aimed to evaluate the performance of models trained on a larger data set to predict body position in infants with CP. Because recruiting a large number of infants with CP is challenging, we increased the size of the training sets by incorporating data from children with TD (Figure 2b). To determine whether including training data from infants with CP is necessary, we held the training set size constant and compared the performance of five folds of models trained on data from infants with CP and TD with a model trained on data from infants with TD only. As a baseline for comparison, all medium models were also tested on infants with TD.

Figure 4b depicts the performance metrics for each combination of training and testing sets for the medium models. The medium TDCP-trained models performed well at predicting body position in infants with CP, with M accuracy of the five folds ranging from 84.8–86.7%, M kappa ranging from 0.758–0.784, and M F1 ranging from 0.818–0.835. The medium TD-trained model performed less well at predicting body position in infants with CP, with accuracy M = 78.8%, kappa M = 0.678, and F1 M = 0.767. In contrast, all medium models performed well at predicting body position in infants with TD, and their performance was similar to that of the TDCP-trained models tested on infants with CP. The linear mixed model confirmed significant effects of model type and the group*model type interaction for all three metrics (Table 2).

3.3. Evaluation of full model

The previous results suggest that predictive models for identifying body position in infants with CP benefited from 1) using a larger training set and 2) including infants with CP in the training set. Therefore, to maximize performance, we trained a full model containing all available data, combining infants with CP and TD. To evaluate the performance of this model, we performed a LOOCV procedure by training a series of models on all sessions except the one being left out to test from the combined sample of infants with CP and TD. This included testing on the infants with TD to provide a baseline of the “best” performance for the model.

3.3.1. Overall metrics

Figure 5 depicts the performance metrics for the full models tested on infants with CP and infants with TD, which were similar to those for the medium TDCP-trained models. Note that compared to the first set of models—the small CP-trained models—the average performance of the full models is slightly better for predicting body position in infants with CP. Additionally, the inter-individual variability is reduced when using the full TDCP-trained models: minimum accuracy of 72.9%, kappa of 0.548, and F1 of 0.675. Therefore, the larger, combined training set improved model performance for the most difficult-to-predict participants. The overall performance of the full models for infants with CP was not significantly different from their performance for infants with TD; accuracy (CP M = .86, TD M = .89), t(13.4) = −1.19, p = 0.254; kappa (CP M = .77, TD M = .84), t(13.2) = −1.49, p = 0.161; F1 (CP M = .82, TD M = .87), t(15.2) = −1.49, p = 0.156.

Figure 5.

Figure 5

Three overall performance metrics—accuracy, kappa, and F1 score—for full models. Each symbol represents one infant’s metrics, and the horizontal lines represent the group mean. Blue symbols and lines represent results when testing models on infants with CP, and orange symbols/lines represent results when testing models on infants with TD.

3.3.2. By-class metrics

Figure 6 presents a confusion matrix showing the aggregated counts and percentages of ground truth versus prediction for every 4-s window for all infants with CP. In this type of visualization, each 4-s window is classified into exactly one square of the matrix so that the total of all squares equals 100%. Good classification accuracy is represented in a confusion matrix by high percentages in the squares along the main diagonal (held episodes being correctly classified as held, prone episodes being correctly classified as prone, and so on) and poor classification accuracy is represented by high percentages in the squares outside of the main diagonal (held episodes being incorrectly classified as standing, prone episodes being incorrectly classified as supine, and so on).

Figure 6.

Figure 6

Confusion matrix showing the proportion (larger text) and number (smaller text) of 4-sec windows that were classified in each category by model predictions and ground truth video coding, pooled over the 10 infants with CP. The darkness of the cell color scales to the frequency of that combination of prediction and ground truth: lower values close to zero—e.g., ground truth of prone classified as held which represented only 8 windows—in the lightest blue, and higher values—e.g., ground truth of sitting classified as sitting which represented 20,340 windows—in the darkest blue). The darkest cells are seen along the center diagonal, indicating that most windows were classified correctly. The only notable confusions outside of the diagonal are misclassifying sitting as supine (6.4% of windows) and misclassifying supine as sitting (4.4% of windows).

The matrix shows excellent alignment between the predicted and actual positions, and demonstrates that the most commonly confused classes were sitting and supine. In other words, true sitting events were sometimes incorrectly labeled as supine and vice versa. Table 3 lists the by-class metrics—sensitivity, specificity, and PPV—for each position. Sensitivity was lowest in sitting (M = 83.9%) and highest in prone (M = 95.2%); in other words, the models were excellent at identifying prone when it actually occurred, but missed some true instances of sitting. Specificity was high for all positions but was lowest in supine (M = 91.4%) and highest in prone (M = 99.7%); in other words, the model hardly ever labeled a non-prone window as prone, but occasionally labeled non-supine windows as supine. Finally, PPV was lowest in held (M = 71.7%) and highest in prone (M = 96.8%); in other words, a window labeled prone was very likely to truly be prone, but a window labeled held was slightly less likely to truly be held.

Table 3.

By-class metrics for each position

Mean Median Minimum Maximum
Supine
Sensitivity 0.87 0.96 0.521 1.00
Specificity 0.914 0.976 0.714 0.998
PPV 0.756 0.802 0.256 0.997
Prone
Sensitivity 0.952 0.991 0.723 1.00
Specificity 0.997 0.999 0.984 1
PPV 0.968 0.983 0.88 0.999
Sitting
Sensitivity 0.839 0.92 0.575 0.994
Specificity 0.927 0.984 0.695 1.00
PPV 0.855 0.984 0.353 1.00
Standing
Sensitivity 0.864 0.907 0.667 1.00
Specificity 0.992 0.995 0.981 1.00
PPV 0.785 0.909 0.174 0.998
Held
Sensitivity 0.898 0.926 0.759 0.993
Specificity 0.978 0.988 0.935 0.999
PPV 0.717 0.785 0.255 0.988

3.3.3. Predicted vs. ground truth prevalence

Finally, to assess the ability of the model predictions to capture individual differences in time spent in the different positions, we compared the prevalence predicted by the model—the percentage of windows the model classified as each position—to the ground truth prevalence from the video coding—the percentage of windows the infant was actually in the position according to the human coder. Figure 7 depicts the ground truth and predicted prevalence for each position. ICCs were excellent for prone (0.999), standing (0.984), and held (0.943), good for sitting (0.756), and moderate for supine (0.682).

Figure 7.

Figure 7

Predicted prevalence of each position class (total proportion of 4-sec windows that the full model predicted the child was in the position) compared to the ground truth prevalence (total proportion of 4-sec windows that the human coder indicated the child was in the position) for all 10 infants with CP.

3.4. Exploratory analysis: Between-group differences in raw features

The accuracy of the TDCP-trained model for both infants with CP and infants with TD suggests that the random forest procedure identified generalizable features that could best predict body positions across both groups. To explore whether the groups differed in the key features, we first identified the 5 features with the highest feature importance scores from the final model. Then, for each of the 5 features for each of the 5 position categories, we calculated a mean feature value for each infant and examined the distributions of these values by group. Supplemental Material 2 depicts these distributions for the top 5 features. Overall, noticeable differences were observed between positions; for example, the highest importance feature, 25th percentile of right hip Y acceleration, had high values while infants were sitting and supine, moderate values when upright and held, and low values when prone. However, there were no apparent differences in distributions between the CP and TD groups, suggesting that the important features related similarly to body positions in both groups.

3.5. Full day positioning experience

Not surprisingly, given the variability in age and motor skill level, the amount of time in the five positions varied widely between participants (see Supplemental Material 3). Our final analysis examined whether this inter-individual variability in postures used during real-world activity was associated with our clinical measure of skill in each posture. Prone positioning was moderately associated with prone AIMS scores (r(8) = 0.582, p = 0.078), but this correlation did not reach significance in this small sample; sitting positioning was not at all associated with sitting AIMS scores (r(8) = 0.034, p = 0.926); and standing positioning was strongly associated with standing AIMS scores (r(8) = 0.845, p = .002; Figure 8).

Figure 8.

Figure 8

AIMS prone, sitting, and standing scores compared to the proportion of the full unrecorded day that each child with CP was in prone, sitting, and standing positions.

4. Discussion

This study assessed the validity of body position classification using wearable sensors and machine learning in infants with CP. Previous work has demonstrated that a machine learning model trained on data from infants with TD could accurately predict body position in infants with TD during natural behavior in the home environment (Franchak et al., 2021, 2023). Here, we found that a combined model trained on data from infants with and without CP accurately predicted body position in infants with CP nearly as well as in infants with TD. This model showed good overall performance (with variation in accuracy between positions), and its predictions for a full day of activity aligned with clinical measures of posture-specific motor skill.

4.1. Size and composition of the training set

We independently varied the size and composition of the training sets to determine which training data produced the most accurate model for infants with CP. We first examined the performance of models trained on small (n = 9) datasets, comparing models trained exclusively on data from infants with CP to models trained on same sized samples of infants with TD. Models trained on CP data showed moderate overall accuracy when predicting body position in infants with CP. Performance of these models varied considerably across individual infants, with the highest reaching 98% accuracy and the lowest at only 63% accuracy. Note that this was similar to the performance distributions seen in infants with TD when trained on 9-session data sets and may represent ceiling performance for models trained on this small training set size. In contrast, TD-trained models performed poorly when applied to infants with CP. This mirrors the poor performance of models trained on neurotypical adults for predicting activity in people with stroke or Parkinson’s Disease (Albert et al., 2012; Capela et al., 2016; O’Brien et al., 2017; Oh et al., 2023), and suggests that the correspondence between motion features and body positions in infants with CP are different than those in infants with TD. Poor performance is not likely to be due to the different levels of motor development, as both CP and TD samples included infants who could and could not sit, crawl, and cruise. More likely, this reflects atypical movement patterns characteristic of CP, including abnormal muscle tone, spasticity, poor motor control and coordination, and excessive or involuntary movements. For example, one infant with CP was noted to display unusual excessive movement of the legs while floor sitting. This was mistakenly labeled by the TD-trained models as supine, perhaps reflecting the common pattern of infants kicking while in a supine position. Our findings suggest caution is warranted for researchers seeking to generalize machine learning models trained on neurotypical individuals to clinical populations.

Interestingly, all small models—including the model trained only on infants with CP—performed moderately well when tested on infants with TD. This finding suggests that although the variability and atypical movement patterns present in infants with CP pose a challenge for models trained on TD data, models trained on CP data may capture a broader range of movement patterns that generalize well to infants with TD.

The second set of analyses evaluated models trained on larger datasets, with and without infants with CP. Merely increasing the size of the training set was not sufficient to improve performance for infants with CP: The medium models trained on data from infants with TD still achieved limited accuracy. However, when infants with CP were included in the medium model training set, performance was strong. Analogously to findings from the small models, infants with TD were well-predicted by medium models trained on either infants with TD alone or the combination. These findings together suggest that infants with TD are more robust to the composition of the training set; however, for infants with CP, training set composition matters. To maximize performance, we trained full models using all available data (i.e., all infants with TD and all infants with CP, holding only the tested session out of the training set). Like the medium TDCP-trained models, the full models performed well both for infants with CP and TD.

Compared to the first attempt—the small model trained on only 9 infants with CP—the full model produced only a modest increase in average performance (M = 84% vs. M = 86% accuracy). However, full models demonstrated considerably increased performance at the floor of the distribution (min = 63% vs. min = 73% accuracy). This improvement suggests that increasing the sample size may help the model deal with individual variability in movement patterns, which is particularly important in heterogeneous clinical populations such as CP.

Overall, results suggest that for clinical populations, predictive accuracy benefits both from increased training sample size and from inclusion of the target clinical population in the training data. When large clinical populations are difficult to recruit (as in infants with CP), combining data from the clinical and typical groups is a good solution (see also Oh et al., 2023).

4.2. Evaluating the best model

After confirming good overall accuracy (comparable to that seen in infants with TD), we further examined the performance of this model to predict each activity class—i.e., each body position—in infants with CP. This allows us to understand the limits of the model’s performance in measuring specific outcomes important for future inquiry. Overall, we saw that the model had considerable strengths and moderate limitations in measuring different positions.

The model was excellent at identifying prone, standing, and held categories, with good sensitivity and specificity values and few false positives/false negatives (as seen in the confusion matrix). Note that the lower average PPV values for standing and held reflect their low prevalence, as these were the least frequent positions. Additionally, there was strong agreement between model predictions and ground truth in the overall prevalence of prone, standing, and held; in other words, the model correctly identified infants who spent more vs. less time in these three positions.

Classifying sitting and supine positioning was more challenging. These positions had slightly lower sensitivity and specificity values, and the confusion matrix shows that they were occasionally mistaken for each other, with the model incorrectly labeling some supine windows as sitting (4% of windows overall) and some sitting windows as supine (6% of windows overall). Additionally, agreement between model-prediction prevalence and ground truth prevalence of supine and sitting was lower. Why were sitting and supine confused? One reason may relate to similarities in the orientation of the lower limbs and the hip and ankle sensor signals during supine and long sitting. Additionally, there may be inherent ambiguity between the categories. In particular, infants—especially younger infants or those with less mature motor skills—spend time in reclined positions that are partway between sitting upright and lying flat (Franchak, 2019). For example, one infant with CP spent significant amounts of time in a reclined car seat and later in an adult-supported reclined position while being fed a bottle (note that this was the infant with poor inter-rater reliability due to one coder selecting supine and the other sitting in these instances). Some limitations in prediction accuracy likely reflect the inherent challenge of parsing continuous and complex phenomena like body position into categorical classifications. Notably, the lower accuracy in these positions was not due to insufficient training data, as sitting and supine were the most frequent positions during the video recordings. Finally, we note that although the diagnostic accuracy was lower on average for supine and sitting, the median values are high (96% sensitivity and 98% specificity for supine, 92% sensitivity and 98% sensitivity for sitting). Therefore, accuracy even for these positions was excellent in a majority of participants, with a few individuals being more difficult to categorize. Importantly, this confusion is not unique to CP; supine and sitting were also the least accurately classified positions in previous work with infants with TD (Franchak et al., 2023).

Although most analyses in this manuscript focused on the 90-minute observed (video recorded) portion, we also examined the longer, unobserved portion of the session to examine how much time infants spent in different positions over the course of a full day. Time spent in the five body positions varied widely across participants, underscoring the substantial inter-individual variability in daily motor behavior among infants with CP in this age range. The sample also encompassed a wide range of motor skill levels, as seen in AIMS scores (prone scores ranged from 1–17, sitting scores from 0–12, and standing scores from 0–10). Our final analysis explored whether variability in real-world body position was associated with the clinical measures of position-specific motor skill. Overall, we would expect that infants with more skill in a particular position would spend more time in that position in daily life (Franchak, 2019; Franchak et al., 2024; Kretch et al., 2024; Kretch, Luna, et al., 2025). Our results on this analysis were mixed. We found that infants with greater standing skill spent more time standing (a large, statistically significant correlation), providing preliminary evidence of convergent validity. Additionally, infants with greater prone skill spent marginally more time prone (a moderate, non-significant correlation in this sample of 10 infants), suggesting that real-world prone experience may be related to prone skill. However, there was no association between sitting skill and sitting frequency. This may be because infants with low sitting skill can still spend time in a sitting position supported by caregivers or seating devices (Callahan & Sisler, 1997; Karasik et al., 2015; Kretch, 2025; Kretch et al., 2024, 2023), and infants with the most advanced motor skills replace sitting time with standing (Franchak et al., 2024, 2018; Thurman & Corbetta, 2017). Overall, positioning in daily life is influenced by many factors other than motor skill, making real-world positioning itself an important independent outcome measure for basic and clinical research.

4.3. Future directions: Methodological and scientific considerations

Although overall performance was strong even in infants with CP, there are some steps that could be taken if even higher accuracy was needed. First, we chose a four-sensor protocol as a compromise between accuracy and ease of data collection. It is likely that adding additional sensors—on the trunk, head, or upper extremities—could improve accuracy over using lower extremity sensors alone. However, dealing with multiple garments would likely be a greater burden for caregivers and more obtrusive for infants. Future work could determine whether the benefit of added sensors is large enough to merit the added data collection complexity. Additionally, personalized models could be created using a smaller portion of each infant’s data to predict the rest of that infant’s daily activity (Ahmadi et al., 2020; Franchak et al., 2023); this could be beneficial in a highly heterogeneous diagnostic group like CP. However, this approach requires collection of video recordings and laborious manual annotation for every participant, which may be prohibitively time-consuming for large-scale studies. If resources allow, personalized models would likely generate the most accurate predictions because they take individual differences in infants’ movement patterns into account. Finally, adding more infants with CP to the training data may improve model performance, as our findings suggest that predictive performance improves with both larger training datasets and inclusion of infants with CP.

The development of a validated body position prediction model opens possibilities for researchers to ask a multitude of important questions about everyday positioning experience in the home environment in infants with CP. For example, we can investigate how everyday positioning changes over development, or how frequencies, temporal distribution, and variability of positioning events differ from those seen in typically developing infants. This type of data has significant potential to provide new insights into processes of divergent early development in CP. Additionally, clinical trials examining early motor interventions for infants with CP typically rely on standardized assessments that reflect motor capacity (Boyd et al., 2017; Dusing et al., 2022; Morgan et al., 2023; Prosser et al., 2018); the addition of sensor-based measurements of everyday positioning could support the assessment of real-life motor performance (Bjornson et al., 2019; Halma et al., 2020; Michielsen et al., 2009). Additionally, investigations of effects of infant and family factors—such as home layout, family composition, or parenting styles—on infant body position can be extended to children with CP, providing valuable data to inform early intervention delivery. Finally, the promising results found here with children with CP suggest that this approach can likely be implemented in other populations of infants with motor impairments and unique movement patterns.

At the moment, we recommend this method as a research tool for understanding everyday behavior (body position) and its contributions to development; our model expands the utility of this method beyond children with typical development to encompass a larger range of neurodiversity and reveal potential opportunities for clinical intervention. The success of classifying body positions suggests that similar models could be generated to classify other relevant behaviors, even in infants with CP, given the proper model training conditions. In the future, activity recognition models could be developed into home-based monitoring tools that can provide valuable information for clinicians on the everyday performance of motor skills or developmentally important activities.

4.4. Limitations and conclusions

Limitations should be considered when interpreting these findings. First, the small sample size for the CP group limited our ability to detect statistical differences between the samples. It is possible that the overall performance metrics for the full model are in fact higher for infants with TD than those with CP (M = 86% vs. M = 89% accuracy), but our group comparisons were underpowered. Additionally, the small sample precludes a detailed exploration of factors influencing accuracy in this population, including examining whether model performance differs across clinically meaningful CP subgroups—for example, at different levels of motor development or with bilateral vs. unilateral presentations. A second limitation is that the CP and TD groups were collected at two different sites; differences in location, demographics, and minor differences in the procedures used at the two sites may have contributed to variability between groups.

In conclusion, this study demonstrated that wearable sensors combined with machine learning can accurately classify body position in infants with CP in real-world contexts. Model performance was strongest when training datasets were larger and included infants with CP, highlighting the importance of incorporating representative movement patterns during model training. Although classification accuracy varied across individuals and positions, the final models achieved comparable performance for infants with CP and TD and were able to capture meaningful inter-individual differences in body position. We recommend the model developed here for use in studies of everyday, real-world body position in infants with CP. The use of this tool to examine divergent developmental trajectories of everyday body position statistics and influential contextual factors can contribute to moving the fields of developmental science and pediatric rehabilitation toward a more comprehensive understanding of typical and atypical infant development outside the confines of the lab environment.

Supplementary Material

Supplementary Material 1
Supplementary Material 3
Supplementary Material 2

Acknowledgments:

This work was supported by a pilot grant from the Center for Smart Use of Technologies to Assess Real World Outcomes/C-STAR (National Institutes of Health P2CHD101899) to KSK, a James S. McDonnell Foundation Opportunity Grant to DHA, and a grant from the National Science Foundation (BCS 2521429/2521430/2521431) to JMF, KSK, and DHA. We gratefully acknowledge members of the USC Learning, Development, and Rehabilitation Lab and the UGA Developmental Dynamics Lab for their assistance with behavioral coding and data collection, and the families who contributed their time to participate in our studies.

References

  1. Abrishami M, Nocera L, Mert M, Trujillo-Priego IA, Purushotham S, Shahabi C, & Smith BA (2019). Identification of developmental delay in infants using wearable sensors: Full-day leg movement statistical feature analysis. IEEE Journal of Translational Engineering in Health and Medicine, 7, 1–7. 10.1109/JTEHM.2019.2893223 [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Ahmadi MN, O’Neil ME, Baque E, Boyd RN, & Trost SG (2020). Machine learning to quantify physical activity in children with cerebral palsy: Comparison of group, group-personalized, and fully-personalized activity classification models. Sensors (Basel, Switzerland), 20(14), 3976. 10.3390/s20143976 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Ahmadi MN, O’Neil M, Fragala-Pinkham M, Lennon N, & Trost S (2018). Machine learning algorithms for activity recognition in ambulant children and adolescents with cerebral palsy. Journal of Neuroengineering and Rehabilitation, 15(1), 105. 10.1186/s12984-018-0456-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Airaksinen M, Gallen A, Kivi A, Vijayakrishnan P, Häyrinen T, Ilén E, Räsänen O, Haataja LM, & Vanhatalo S (2022). Intelligent wearable allows out-of-the-lab tracking of developing motor abilities in infants. Communication & Medicine, 2, 69. 10.1038/s43856-022-00131-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Airaksinen M, Gallen A, Taylor E, de Sena S, Palsa T, Haataja L, & Vanhatalo S (2025). Assessing infant gross motor performance with an at-home wearable. Pediatrics, e2024068647. 10.1542/peds.2024-068647 [DOI] [PubMed] [Google Scholar]
  6. Airaksinen M, Räsänen O, Ilén E, Häyrinen T, Kivi A, Marchi V, Gallen A, Blom S, Varhe A, Kaartinen N, Haataja L, & Vanhatalo S (2020). Automatic Posture and Movement Tracking of Infants with Wearable Movement Sensors. Scientific Reports, 10(1), 169. 10.1038/s41598-019-56862-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Albert MV, Toledo S, Shapiro M, & Kording K (2012). Using mobile phones for activity recognition in Parkinson’s patients. Frontiers in Neurology, 3, 158. 10.3389/fneur.2012.00158 [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Anisfeld E, Casper V, Nozyce M, & Cunningham N (1990). Does infant carrying promote attachment? An experimental study of the effects of increased physical contact on the development of attachment. Child Development, 61(5), 1617. 10.2307/1130769 [DOI] [PubMed] [Google Scholar]
  9. Bax M, Goldstein M, Rosenbaum P, Leviton A, Paneth N, Dan B, Jacobsson B, Damiano D, & Executive Committee for the Definition of Cerebral Palsy. (2005). Proposed definition and classification of cerebral palsy, April 2005. Developmental Medicine and Child Neurology, 47(8), 571–576. 10.1017/s001216220500112x [DOI] [PubMed] [Google Scholar]
  10. Beani E, de ‘Cavalieri MF, Filogna S, Barzacchi V, Cianchetti M, Maselli M, Martini G, Menici V, Prencipe G, Sicola E, Cioni G, & Sgandurra G (2025). Wearable sensors for measuring spontaneous upper limb use in children with unilateral cerebral palsy and typical development. Journal of Neuroengineering and Rehabilitation, 22(1), 71. 10.1186/s12984-025-01601-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Bezanson J, Edelman A, Karpinski S, & Shah VB (2017). Julia: A fresh approach to numerical computing. SIAM Review. Society for Industrial and Applied Mathematics, 59(1), 65–98. 10.1137/141000671 [DOI] [Google Scholar]
  12. Bjornson KF, Moreau N, & Bodkin AW (2019). Short-burst interval treadmill training walking capacity and performance in cerebral palsy: a pilot study. Developmental Neurorehabilitation, 22(2), 126–133. 10.1080/17518423.2018.1462270 [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Boyd RN, Ziviani J, Sakzewski L, Novak I, Badawi N, Pannek K, Elliott C, Greaves S, Guzzetta A, Whittingham K, Valentine J, Morgan C, Wallen M, Eliasson A-C, Findlay L, Ware R, Fiori S, & Rose S (2017). REACH: study protocol of a randomised trial of rehabilitation very early in congenital hemiplegia. BMJ Open, 7(9), e017204. 10.1136/bmjopen-2017-017204 [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Braito I, Maselli M, Sgandurra G, Inguaggiato E, Beani E, Cecchi F, Cioni G, & Boyd R (2018). Assessment of upper limb use in children with typical development and neurodevelopmental disorders by inertial sensors: a systematic review. Journal of Neuroengineering and Rehabilitation, 15(1), 94. 10.1186/s12984-018-0447-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Breiman L (2001). Random forests. Machine Learning, 45(1), 5–32. 10.1023/a:1010933404324 [DOI] [Google Scholar]
  16. Bruijns BA, Truelove S, Johnson AM, Gilliland J, & Tucker P (2020). Infants’ and toddlers’ physical activity and sedentary time as measured by accelerometry: a systematic review and meta-analysis. The International Journal of Behavioral Nutrition and Physical Activity, 17(1), 14. 10.1186/s12966-020-0912-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Calabretta BT, Schneider JL, & Iverson JM (2022). Bidding on the go: Links between walking, social actions, and caregiver responses in infant siblings of children with autism spectrum disorder. Autism Research: Official Journal of the International Society for Autism Research. 10.1002/aur.2830 [DOI] [PubMed] [Google Scholar]
  18. Callahan CW, & Sisler C (1997). Use of seating devices in infants too young to sit. Archives of Pediatrics & Adolescent Medicine, 151(3), 233–235. 10.1001/archpedi.1997.02170400019004 [DOI] [PubMed] [Google Scholar]
  19. Capela NA, Lemaire ED, Baddour N, Rudolf M, Goljar N, & Burger H (2016). Evaluation of a smartphone human activity recognition application with able-bodied and stroke participants. Journal of Neuroengineering and Rehabilitation, 13(1), 5. 10.1186/s12984-016-0114-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Carcreff L, Gerber CN, Paraschiv-Ionescu A, De Coulon G, Newman CJ, Aminian K, & Armand S (2020). Comparison of gait characteristics between clinical and daily life settings in children with cerebral palsy. Scientific Reports, 10(1), 2091. 10.1038/s41598-020-59002-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Carcreff L, Gerber CN, Paraschiv-Ionescu A, De Coulon G, Newman CJ, Armand S, & Aminian K (2018). What is the best configuration of wearable sensors to measure spatiotemporal gait parameters in children with cerebral palsy? Sensors (Basel, Switzerland), 18(2), 394. 10.3390/s18020394 [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Carcreff L, Ionescu A, Gerber C, De Coulon G, Aminian K, Newman C, & Armand S (2017). Assessment of the spatiotemporal gait parameters of children with cerebral palsy in daily-life settings: comparison between wearable systems using different sensor location. Gait & Posture, 57, 237–238. 10.1016/j.gaitpost.2017.06.392 [DOI] [Google Scholar]
  23. Clanchy KM, Tweedy SM, & Boyd R (2011). Measurement of habitual physical activity performance in adolescents with cerebral palsy: a systematic review. Developmental Medicine and Child Neurology, 53(6), 499–505. 10.1111/j.1469-8749.2010.03910.x [DOI] [PubMed] [Google Scholar]
  24. Claridge EA, van den Berg-Emons RJG, Horemans HLD, van der Slot WMA, van der Stam N, Tang A, Timmons BW, Gorter JW, & Bussmann JBJ (2019). Detection of body postures and movements in ambulatory adults with cerebral palsy: a novel and valid measure of physical behaviour. Journal of Neuroengineering and Rehabilitation, 16(1), 125. 10.1186/s12984-019-0594-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Clearfield MW (2011). Learning to walk changes infants’ social interactions. Infant Behavior & Development, 34(1), 15–25. 10.1016/j.infbeh.2010.04.008 [DOI] [PubMed] [Google Scholar]
  26. Clearfield MW, Osborne CN, & Mullen M (2008). Learning by looking: Infants’ social looking behavior across the transition from crawling to walking. Journal of Experimental Child Psychology, 100(4), 297–307. 10.1016/j.jecp.2008.03.005 [DOI] [PubMed] [Google Scholar]
  27. de Barbaro K, & Fausey CM (2022). Ten Lessons About Infants’ Everyday Experiences. Current Directions in Psychological Science, 31(1), 28–33. 10.1177/09637214211059536 [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. de Barbaro K, Madden-Rusnak A, & Timmons A (2026). Thinking critically about algorithms for automated detection of behavior: 11 guidelines for social and behavioral scientists. Developmental Science, 29(3), e70144. 10.1111/desc.70144 [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. den Hartog D, van der Krogt MM, van der Burg S, Aleo I, Gijsbers J, Bonouvrié LA, Harlaar J, Buizer AI, & Haberfehlner H (2022). Home-Based Measurements of Dystonia in Cerebral Palsy Using Smartphone-Coupled Inertial Sensor Technology and Machine Learning: A Proof-of-Concept Study. Sensors, 22(12). 10.3390/s22124386 [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Duda-Goławska J, Rogowski A, Laudańska Z, Żygierewicz J, & Tomalski P (2024). Identifying infant body position from inertial sensors with machine learning: Which parameters matter? Sensors (Basel, Switzerland), 24(23), 7809. 10.3390/s24237809 [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Dudek-Shriber L, & Zelazny S (2007). The effects of prone positioning on the quality and acquisition of developmental milestones in four-month-old infants. Pediatric Physical Therapy: The Official Publication of the Section on Pediatrics of the American Physical Therapy Association, 19(1), 48–55. 10.1097/01.pep.0000234963.72945.b1 [DOI] [PubMed] [Google Scholar]
  32. Durkin MS, Benedict RE, Christensen D, Dubois LA, Fitzgerald RT, Kirby RS, Maenner MJ, Van Naarden Braun K, Wingate MS, & Yeargin-Allsopp M (2016). Prevalence of Cerebral Palsy among 8-Year-Old Children in 2010 and Preliminary Evidence of Trends in Its Relationship to Low Birthweight. Paediatric and Perinatal Epidemiology, 30(5), 496–510. 10.1111/ppe.12299 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Dusing SC, Harbourne RT, Hsu L-Y, Koziol NA, Kretch K, Sargent B, Jensen-Willett S, McCoy SW, & Vanderbilt DL (2022). The SIT-PT Trial Protocol: a Dose-Matched Randomized Clinical Trial Comparing 2 Physical Therapist Interventions for Infants and Toddlers with Cerebral Palsy. Physical Therapy. 10.1093/ptj/pzac039 [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Eppler MA (1995). Development of manipulatory skills and the deployment of attention. Infant Behavior & Development, 18(4), 391–405. 10.1016/0163-6383(95)90029-2 [DOI] [Google Scholar]
  35. Franchak JM (2019). Changing Opportunities for Learning in Everyday Life: Infant Body Position Over the First Year. Infancy: The Official Journal of the International Society on Infant Studies, 24(2), 187–209. 10.1111/infa.12272 [DOI] [PubMed] [Google Scholar]
  36. Franchak JM (2020). The ecology of infants’ perceptual-motor exploration. Current Opinion in Psychology, 32, 110–114. 10.1016/j.copsyc.2019.06.035 [DOI] [PubMed] [Google Scholar]
  37. Franchak JM, Kadooka K, & Fausey CM (2024). Longitudinal relations between independent walking, body position, and object experiences in home life. Developmental Psychology, 60(2), 228–242. 10.1037/dev0001678 [DOI] [PubMed] [Google Scholar]
  38. Franchak JM, Kretch KS, & Adolph KE (2018). See and be seen: Infant–caregiver social looking during locomotor free play. Developmental Science, 21(4), e12626. 10.1111/desc.12626 [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Franchak JM, Rousey HN, & Wang H (2026). Natural statistics of infants’ everyday motor experiences relate to sitting and walking development. Developmental Psychology. 10.1037/dev0002185 [DOI] [PubMed] [Google Scholar]
  40. Franchak JM, Scott V, & Luo C (2021). A Contactless Method for Measuring Full-Day, Naturalistic Motor Behavior Using Wearable Inertial Sensors. Frontiers in Psychology, 12, 701343. 10.3389/fpsyg.2021.701343 [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Franchak JM, Tang M, Rousey H, & Luo C (2023). Long-form recording of infant body position in the home using wearable inertial sensors. Behavior Research Methods. 10.3758/s13428-023-02236-9 [DOI] [PubMed] [Google Scholar]
  42. Goodlich BI, Armstrong EL, Horan SA, Baque E, Carty CP, Ahmadi MN, & Trost SG (2020). Machine learning to quantify habitual physical activity in children with cerebral palsy. Developmental Medicine and Child Neurology, 62(9), 1054–1060. 10.1111/dmcn.14560 [DOI] [PubMed] [Google Scholar]
  43. Graciosa MD, Ferronato PAM, Drezner R, & de Jesus Manoel E (2024). Emergence of locomotor behaviors: Associations with infant characteristics, developmental status, parental beliefs, and practices in typically developing Brazilian infants aged 5 to 15 months. Infant Behavior & Development, 76, 101965. 10.1016/j.infbeh.2024.101965 [DOI] [PubMed] [Google Scholar]
  44. Greenspan B, Cunha AB, & Lobo MA (2021). Design and validation of a smart garment to measure positioning practices of parents with young infants. Infant Behavior & Development, 62, 101530. 10.1016/j.infbeh.2021.101530 [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Halma E, Bussmann JBJ, van den Berg-Emons HJG, Sneekes EM, Pangalila R, Schasfoort FC, & SPACE BOP study group. (2020). Relationship between changes in motor capacity and objectively measured motor performance in ambulatory children with spastic cerebral palsy. Child: Care, Health and Development, 46(1), 66–73. 10.1111/cch.12719 [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Hendry D, Rohl AL, Rasmussen CL, Zabatiero J, Cliff DP, Smith SS, Mackenzie J, Pattinson CL, Straker L, & Campbell A (2023). Objective measurement of posture and movement in young children using wearable sensors and customised mathematical approaches: A systematic review. Sensors (Basel, Switzerland), 23(24), 9661. 10.3390/s23249661 [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Hewitt L, Stanley RM, Cliff D, & Okely AD (2019). Objective measurement of tummy time in infants (0–6 months): A validation study. PloS One, 14(2), e0210977. 10.1371/journal.pone.0210977 [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Hunziker UA, & Barr RG (1986). Increased carrying reduces infant crying: a randomized controlled trial. Pediatrics, 77(5), 641–648. 10.1542/peds.77.5.641 [DOI] [PubMed] [Google Scholar]
  49. Iverson JM (2022). Developing language in a developing body, revisited: The cascading effects of motor development on the acquisition of language. Wiley Interdisciplinary Reviews. Cognitive Science, e1626. 10.1002/wcs.1626 [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Jeong H, Kwak SS, Sohn S, Lee JY, Lee YJ, O’Brien MK, Park Y, Avila R, Kim J-T, Yoo J-Y, Irie M, Jang H, Ouyang W, Shawen N, Kang YJ, Kim SS, Tzavelis A, Lee K, Andersen RA, … Rogers JA (2021). Miniaturized wireless, skin-integrated sensor networks for quantifying full-body movement behaviors and vital signs in infants. Proceedings of the National Academy of Sciences of the United States of America, 118(43). 10.1073/pnas.2104925118 [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Karasik LB, Adolph KE, Tamis-LeMonda CS, & Zuckerman AL (2012). Carry on: Spontaneous object carrying in 13-month-old crawling and walking infants. Developmental Psychology, 48(2), 389–397. 10.1037/a0026040 [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Karasik LB, Tamis-LeMonda CS, & Adolph KE (2011). Transition from crawling to walking and infants’ actions with objects and people. Child Development, 82(4), 1199–1209. 10.1111/j.1467-8624.2011.01595.x [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Karasik LB, Tamis-Lemonda CS, & Adolph KE (2014). Crawling and walking infants elicit different verbal responses from mothers. Developmental Science, 17(3), 388–395. 10.1111/desc.12129 [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Karasik LB, Tamis-LeMonda CS, Adolph KE, & Bornstein MH (2015). Places and postures: A cross-cultural comparison of sitting in 5-month-olds. Journal of Cross-Cultural Psychology, 46(8), 1023–1038. 10.1177/0022022115593803 [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Kretch KS (2025). Effects of sitting support and positioning on infant-parent coordinated attention. Developmental Psychology. 10.1037/dev0002067 [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. Kretch KS (2026). Everyday positioning experience in typically developing infants and infants with or at risk for cerebral palsy. Research in Developmental Disabilities, 172(105279), 105279. 10.1016/j.ridd.2026.105279 [DOI] [PMC free article] [PubMed] [Google Scholar]
  57. Kretch KS, Franchak JM, & Adolph KE (2014). Crawling and walking infants see the world differently. Child Development, 85(4), 1503–1518. 10.1111/cdev.12206 [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Kretch KS, Koziol NA, Marcinowski EC, Hsu L-Y, Harbourne RT, Lobo MA, McCoy SW, Willett SL, & Dusing SC (2024). Sitting capacity and performance in infants with typical development and infants with motor delay. Physical & Occupational Therapy in Pediatrics, 44(2), 164–179. 10.1080/01942638.2023.2241537 [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Kretch KS, Koziol NA, Marcinowski EC, Kane AE, Inamdar K, Brown ED, Bovaird JA, Harbourne RT, Hsu L-Y, Lobo MA, & Dusing SC (2022). Infant posture and caregiver-provided cognitive opportunities in typically developing infants and infants with motor delay. Developmental Psychobiology, 64(1). 10.1002/dev.22233 [DOI] [PubMed] [Google Scholar]
  60. Kretch KS, Luna A, Fausey CM, & Franchak JM (2025). Infant sitting status, sitting age, and everyday positioning experience across the transition to independent sitting. Infancy: The Official Journal of the International Society on Infant Studies, 30(6). 10.1111/infa.70059 [DOI] [PubMed] [Google Scholar]
  61. Kretch KS, Marcinowski EC, Hsu L-Y, Koziol NA, Harbourne RT, Lobo MA, & Dusing SC (2023). Opportunities for learning and social interaction in infant sitting: Effects of sitting support, sitting skill, and gross motor delay. Developmental Science, 26(3), e13318. 10.1111/desc.13318 [DOI] [PMC free article] [PubMed] [Google Scholar]
  62. Kretch KS, Marcinowski EC, Koziol NA, Harbourne RT, Hsu L-Y, Lobo MA, Willett SL, & Dusing SC (2025). Sitting and caregiver speech input in typically developing infants and infants with cerebral palsy. PloS One, 20(5), e0324106. 10.1371/journal.pone.0324106 [DOI] [PMC free article] [PubMed] [Google Scholar]
  63. Kuo Y-L, Liao H-F, Chen P-C, Hsieh W-S, & Hwang A-W (2008). The influence of wakeful prone positioning on motor development during the early life. Journal of Developmental and Behavioral Pediatrics: JDBP, 29(5), 367–376. 10.1097/DBP.0b013e3181856d54 [DOI] [PubMed] [Google Scholar]
  64. Lara OD, & Labrador MA (2013). A survey on human activity recognition using wearable sensors. IEEE Communications Surveys & Tutorials, 15(3), 1192–1209. 10.1109/surv.2012.110112.00192 [DOI] [Google Scholar]
  65. Libertus K, & Landa RJ (2013). The Early Motor Questionnaire (EMQ): a parental report measure of early motor development. Infant Behavior & Development, 36(4), 833–842. 10.1016/j.infbeh.2013.09.007 [DOI] [PMC free article] [PubMed] [Google Scholar]
  66. Little EE, Legare CH, & Carver LJ (2019). Culture, carrying, and communication: Beliefs and behavior associated with babywearing. Infant Behavior & Development, 57(101320), 101320. 10.1016/j.infbeh.2019.04.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  67. Long BL, Sanchez A, Kraus AM, Agrawal K, & Frank MC (2021). Automated detections reveal the social information in the changing infant view. Child Development. 10.1111/cdev.13648 [DOI] [PubMed] [Google Scholar]
  68. Majnemer A, & Barr RG (2007). Influence of supine sleep positioning on early motor milestone acquisition. Developmental Medicine and Child Neurology, 47(6), 370–376. 10.1111/j.1469-8749.2005.tb01156.x [DOI] [PubMed] [Google Scholar]
  69. Marcinowski EC, Tripathi T, Hsu L-Y, Westcott McCoy S, & Dusing SC (2019). Sitting skill and the emergence of arms-free sitting affects the frequency of object looking and exploration. Developmental Psychobiology, 61(7), 1035–1047. 10.1002/dev.21854 [DOI] [PubMed] [Google Scholar]
  70. Michielsen ME, de Niet M, Ribbers GM, Stam HJ, & Bussmann JB (2009). Evidence of a logarithmic relationship between motor capacity and actual performance in daily life of the paretic arm following stroke. Journal of Rehabilitation Medicine: Official Journal of the UEMS European Board of Physical and Rehabilitation Medicine, 41(5), 327–331. 10.2340/16501977-0351 [DOI] [PubMed] [Google Scholar]
  71. Mireault GC, Rainville BS, & Laughlin B (2018). Push or carry? Pragmatic opportunities for language development in strollers vs. Backpacks. Infancy: The Official Journal of the International Society on Infant Studies, 23(4), 616–624. 10.1111/infa.12238 [DOI] [PMC free article] [PubMed] [Google Scholar]
  72. Morgan C, Badawi N, Boyd RN, Spittle AJ, Dale RC, Kirby A, Hunt RW, Whittingham K, Pannek K, Morton RL, Tarnow-Mordi W, Fahey MC, Walker K, Prelog K, Elliott C, Valentine J, Guzzetta A, Olivey S, GAME study team, & Novak I, (2023). Harnessing neuroplasticity to improve motor performance in infants with cerebral palsy: a study protocol for the GAME randomised controlled trial. BMJ Open, 13(3), e070649. 10.1136/bmjopen-2022-070649 [DOI] [PMC free article] [PubMed] [Google Scholar]
  73. Morgan C, Fetters L, Adde L, Badawi N, Bancale A, Boyd RN, Chorna O, Cioni G, Damiano DL, Darrah J, de Vries LS, Dusing S, Einspieler C, Eliasson A-C, Ferriero D, Fehlings D, Forssberg H, Gordon AM, Greaves S, … Novak I (2021). Early Intervention for Children Aged 0 to 2 Years With or at High Risk of Cerebral Palsy: International Clinical Practice Guideline Based on Systematic Reviews. JAMA Pediatrics. 10.1001/jamapediatrics.2021.0878 [DOI] [PMC free article] [PubMed] [Google Scholar]
  74. Nam Y, & Park JW (2013). Child activity recognition based on cooperative fusion model of a triaxial accelerometer and a barometric pressure sensor. IEEE Journal of Biomedical and Health Informatics, 17(2), 420–426. 10.1109/JBHI.2012.2235075 [DOI] [PubMed] [Google Scholar]
  75. Narayanan A, Desai F, Stewart T, Duncan S, & Mackay L (2020). Application of raw accelerometer data and machine-learning techniques to characterize human movement behavior: A systematic scoping review. Journal of Physical Activity & Health, 17(3), 360–383. 10.1123/jpah.2019-0088 [DOI] [PubMed] [Google Scholar]
  76. Needham A (2000). Improvements in Object Exploration Skills May Facilitate the Development of Object Segregation in Early Infancy. Journal of Cognition and Development: Official Journal of the Cognitive Development Society, 1(2), 131–156. 10.1207/S15327647JCD010201 [DOI] [Google Scholar]
  77. O’Brien MK, Shawen N, Mummidisetty CK, Kaur S, Bo X, Poellabauer C, Kording K, & Jayaraman A (2017). Activity Recognition for Persons With Stroke Using Mobile Phone Technology: Toward Improved Performance in a Home Setting. Journal of Medical Internet Research, 19(5), e184. 10.2196/jmir.7385 [DOI] [PMC free article] [PubMed] [Google Scholar]
  78. Oftedal S, Bell KL, Davies PSW, Ware RS, & Boyd RN (2015). Sedentary and active time in toddlers with and without cerebral palsy. Medicine and Science in Sports and Exercise, 47(10), 2076–2083. 10.1249/MSS.0000000000000653 [DOI] [PubMed] [Google Scholar]
  79. Oh Y, Choi S-A, Shin Y, Jeong Y, Lim J, & Kim S (2023). Investigating activity recognition for hemiparetic stroke patients using wearable sensors: A deep learning approach with data augmentation. Sensors (Basel, Switzerland), 24(1), 210. 10.3390/s24010210 [DOI] [PMC free article] [PubMed] [Google Scholar]
  80. Orlando JM, Smith BA, Hafer JF, Paremski A, Amodeo M, Lobo MA, & Prosser LA (2025). Physical activity in pre-ambulatory children with cerebral palsy: An exploratory validation study to distinguish active vs. Sedentary time using wearable sensors. Sensors (Basel, Switzerland), 25(4), 1261. 10.3390/s25041261 [DOI] [PMC free article] [PubMed] [Google Scholar]
  81. Oudgenoeg-Paz O, Leseman PPM, & Volman M. (chiel) J. M. (2015). Exploration as a mediator of the relation between the attainment of motor milestones and the development of spatial cognition and spatial language. Developmental Psychology, 51(9), 1241–1253. 10.1037/a0039572 [DOI] [PubMed] [Google Scholar]
  82. Oudgenoeg-Paz O, Volman M. (chiel) J. M., & Leseman PPM (2012). Attainment of sitting and walking predicts development of productive vocabulary between ages 16 and 28 months. Infant Behavior & Development, 35(4), 733–736. 10.1016/j.infbeh.2012.07.010 [DOI] [PubMed] [Google Scholar]
  83. Paraschiv-Ionescu A, Newman CJ, Carcreff L, Gerber CN, Armand S, & Aminian K (2019). Locomotion and cadence detection using a single trunk-fixed accelerometer: validity for children with cerebral palsy in daily life-like conditions. Journal of Neuroengineering and Rehabilitation, 16(1), 24. 10.1186/s12984-019-0494-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  84. Pargent F, Schoedel R, & Stachl C (2023). Best practices in supervised machine learning: A tutorial for psychologists. Advances in Methods and Practices in Psychological Science, 6(3). 10.1177/25152459231162559 [DOI] [Google Scholar]
  85. Piper MC, & Darrah J (2022). Motor Assessment of the Developing Infant, 2nd Edition. Elsevier. [Google Scholar]
  86. Prosser LA, Pierce SR, Dillingham TR, Bernbaum JC, & Jawad AF (2018). iMOVE: Intensive Mobility training with Variability and Error compared to conventional rehabilitation for young children with cerebral palsy: the protocol for a single blind randomized controlled trial. BMC Pediatrics, 18(1), 329. 10.1186/s12887-018-1303-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  87. Rast FM, & Labruyère R (2020). Systematic review on the application of wearable inertial sensors to quantify everyday life motor activity in people with mobility impairments. Journal of Neuroengineering and Rehabilitation, 17(1), 148. 10.1186/s12984-020-00779-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  88. Rimmer JH (2006). Use of the ICF in identifying factors that impact participation in physical activity/rehabilitation among people with disabilities. Disability and Rehabilitation, 28(17), 1087–1095. 10.1080/09638280500493860 [DOI] [PubMed] [Google Scholar]
  89. Rosenbaum PL, Paneth N, Leviton A, Goldstein M, Bax M, Damiano D, Dan B, & Jacobsson B (2007). A report: the definition and classification of cerebral palsy April 2006. Developmental Medicine and Child Neurology. Supplement, 109, 8–14. 10.1111/j.1469-8749.2007.tb12610.x [DOI] [PubMed] [Google Scholar]
  90. Rousey HN, Tang M, Garcia S, & Franchak JM (2026). Within-day variations in infant body position predict caregiver speech input. Developmental Science, 29(2), e70120. 10.1111/desc.70120 [DOI] [PubMed] [Google Scholar]
  91. Sadeghi B, Chiarawongse P, Squire K, Jones DC, Noack A, St-Jean C, Huijzer R, Schätzle R, Butterworth I, Peng Y-F, & Blaom A (2022). DecisionTree.jl - A Julia implementation of the CART Decision Tree and Random Forest algorithms. Zenodo. 10.5281/ZENODO.7359268 [DOI] [Google Scholar]
  92. Sato H, & Hirai T (2011). A preliminary study describing body position in daily life in children with severe cerebral palsy using a wearable device. Disability and Rehabilitation, 33(25–26), 2529–2534. 10.3109/09638288.2011.579221 [DOI] [PubMed] [Google Scholar]
  93. Sato H, Iwasaki T, Yokoyama M, & Inoue T (2014). Monitoring of body position and motion in children with severe cerebral palsy for 24 hours. Disability and Rehabilitation, 36(14), 1156–1160. 10.3109/09638288.2013.833308 [DOI] [PubMed] [Google Scholar]
  94. Schneider JL, Roemer EJ, Northrup JB, & Iverson JM (2022). Dynamics of the dyad: How mothers and infants co-construct interaction spaces during object play. Developmental Science, e13281. 10.1111/desc.13281 [DOI] [PMC free article] [PubMed] [Google Scholar]
  95. Shida-Tokeshi J, Lane CJ, Trujillo-Priego IA, Deng W, Vanderbilt DL, Loeb GE, & Smith BA (2018). Relationships between full-day arm movement characteristics and developmental status in infants with typical development as they learn to reach: An observational study. Gates Open Research, 2, 17. 10.12688/gatesopenres.12813.2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  96. Smith BA, Trujillo-Priego IA, Lane CJ, Finley JM, & Horak FB (2015). Daily Quantity of Infant Leg Movement: Wearable Sensor Algorithm and Relationship to Walking Onset. Sensors, 15(8), 19006–19020. 10.3390/s150819006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  97. Soska KC, & Adolph KE (2014). Postural Position Constrains Multimodal Object Exploration in Infants. Infancy: The Official Journal of the International Society on Infant Studies, 19(2), 138–161. 10.1111/infa.12039 [DOI] [PMC free article] [PubMed] [Google Scholar]
  98. Soska KC, Adolph KE, & Johnson SP (2010). Systems in development: motor skill acquisition facilitates three-dimensional object completion. Developmental Psychology, 46(1), 129–138. 10.1037/a0014618 [DOI] [PMC free article] [PubMed] [Google Scholar]
  99. Srinivasan S, Amonkar N, Kumavor PD, & Bubela D (2024). Measuring upper extremity activity of children with unilateral cerebral palsy using wrist-worn accelerometers: A pilot study. The American Journal of Occupational Therapy: Official Publication of the American Occupational Therapy Association, 78(2), 7802180050. 10.5014/ajot.2024.050443 [DOI] [PubMed] [Google Scholar]
  100. Thurman SL, & Corbetta D (2017). Spatial exploration and changes in infant–mother dyads around transitions in infant locomotion. Developmental Psychology, 53(7), 1207–1221. 10.1037/dev0000328 [DOI] [PubMed] [Google Scholar]
  101. Tørring MF, Logacjov A, Brændvik SM, Ustad A, Roeleveld K, & Bardal EM (2024). Validation of two novel human activity recognition models for typically developing children and children with Cerebral Palsy. PloS One, 19(9), e0308853. 10.1371/journal.pone.0308853 [DOI] [PMC free article] [PubMed] [Google Scholar]
  102. Trujillo-Priego IA, Lane CJ, Vanderbilt DL, Deng W, Loeb GE, Shida J, & Smith BA (2017). Development of a Wearable Sensor Algorithm to Detect the Quantity and Kinematic Characteristics of Infant Arm Movement Bouts Produced across a Full Day in the Natural Environment. Technologies, 5(3). 10.3390/technologies5030039 [DOI] [PMC free article] [PubMed] [Google Scholar]
  103. Vanmechelen I, Bekteshi S, Haberfehlner H, Feys H, Desloovere K, Aerts J-M, & Monbaliu E (2023). Reliability and discriminative validity of wearable sensors for the quantification of upper limb Movement disorders in individuals with dyskinetic cerebral palsy. Sensors (Basel, Switzerland), 23(3), 1574. 10.3390/s23031574 [DOI] [PMC free article] [PubMed] [Google Scholar]
  104. Walle EA, & Campos JJ (2014). Infant language development is related to the acquisition of walking. Developmental Psychology, 50(2), 336–348. 10.1037/a0033238 [DOI] [PubMed] [Google Scholar]
  105. West KL, & Iverson JM (2021). Communication changes when infants begin to walk. Developmental Science, e13102. 10.1111/desc.13102 [DOI] [PMC free article] [PubMed] [Google Scholar]
  106. Xiong JS-P, Reedman SE, Kho ME, Timmons BW, Verschuren O, & Gorter JW (2022). Operationalization, measurement, and health indicators of sedentary behavior in individuals with cerebral palsy: a scoping review. Disability and Rehabilitation, 44(20), 6070–6081. 10.1080/09638288.2021.1949050 [DOI] [PubMed] [Google Scholar]
  107. Yamamoto H, Sato A, & Itakura S (2019). Eye tracking in an everyday environment reveals the interpersonal distance that affords infant-parent gaze communication. Scientific Reports, 9(1), 10352. 10.1038/s41598-019-46650-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  108. Zhu Z, Liu T, Li G, Li T, & Inoue Y (2015). Wearable sensor systems for infants. Sensors (Basel, Switzerland), 15(2), 3721–3749. 10.3390/s150203721 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1
Supplementary Material 3
Supplementary Material 2

RESOURCES