Abstract
Abstract
Background
Early childhood development (ECD) lays the foundation for lifelong health, academic success and social well-being, yet over 250 million children in low- and middle-income countries are at risk of not reaching their developmental potential. Traditional measures fail to fully capture the risks associated with a child’s development outcomes. Artificial intelligence techniques, particularly machine learning (ML), offer an innovative approach by analysing complex datasets to detect subtle developmental patterns.
Objective
To map the existing literature on the use of ML in ECD research, including its geographical distribution, to identify research gaps and inform future directions. The review focuses on applied ML techniques, data types, feature sets, outcomes, data splitting and validation strategies, model performance, model explainability, key themes, clinical relevance and reported limitations.
Design
Scoping review using the Arksey and O‘Malley framework with enhancements by Levac et al.
Data sources
A systematic search was conducted on 16 June 2024 across PubMed, Web of Science, IEEE Xplore and PsycINFO, supplemented by grey literature (OpenGrey) and reference hand-searching. No publication date limits were applied.
Eligibility criteria
Included studies applied ML or its variants (eg, deep learning (DL), natural language processing) to developmental outcomes in children aged 0–8 years. Studies were in English and addressed cognitive, language, motor or social-emotional development. Excluded were studies focusing on robotics; neurodevelopmental disorders such as autism spectrum disorder, attention-deficit/hyperactivity disorder and communication disorders; disease or medical conditions; and review articles.
Data extraction and charting
Three reviewers independently extracted data using a structured MS Excel template, covering study ML techniques, data types, feature sets, outcomes, outcome measures, data splitting and validation strategies, model performance, model explainability, key themes, clinical relevance and limitations. A narrative synthesis was conducted, supported by descriptive statistics and visualisations.
Results
Of the 759 articles retrieved, 27 met the inclusion criteria. Most studies (78%) originated from high-income countries, with none from sub-Saharan Africa. Supervised ML classifiers (40.7%) and DL techniques (22.2%) were the most used approaches. Cognitive development was the most frequently targeted outcome (33.3%), often measured using the Bayley Scales of Infant and Toddler Development-III (33.3%). Data types varied, with image, video and sensor-based data being most prevalent. Key predictive features were grouped into six categories: brain features; anthropometric and clinical/biological markers; socio-demographic and environmental factors; medical history and nutritional indicators; linguistic and expressive features; and motor indicators. Most studies (74.1%) focused solely on prediction, with the majority conducting predictions at age 2 years and above. Only 41% of studies employed explainability methods, and validation strategies varied widely. Few studies (7.4%) conducted external validation, and only one had progressed to a clinical trial. Common limitations included small sample sizes, lack of external validation and imbalanced datasets.
Conclusion
There is growing interest in using ML for ECD research, but current research lacks geographical diversity, external validation, explainability and practical implementation. Future work should focus on developing inclusive, interpretable and externally validated models that are integrated into real-world implementation.
Keywords: Child, Machine Learning, Cognition
STRENGTHS AND LIMITATIONS OF THIS STUDY.
A comprehensive and systematic search was conducted across four major databases and grey literature source.
The study followed the Preferred Reporting Items for Systematic Review and Meta-Analyses extension for Scoping Reviews guidelines to ensure transparency and methodological rigour.
A wide range of machine learning techniques and variants were captured, along with key model aspects such as output, features, performance, data types, data splitting and validation strategies, clinical relevance and explainability methods.
The review included only English-language articles, potentially excluding relevant studies published in other languages.
Studies focusing on neurodevelopmental disorders such as autism spectrum disorder, attention-deficit/hyperactivity disorder and communication disorders were excluded, which may limit the scope of findings.
Introduction
Early childhood development (ECD) lays the foundation for a person’s lifelong health, academic achievement, economic performance and societal impact.1,3 This is a process that involves the development of physical or motor, cognitive, language and social-emotional skills in children from birth to around 8 years of age.4 Providing a supportive and stimulating environment for children, especially during their early years (from birth to 3 years), is crucial for optimal development.5 During this period, a child’s brain undergoes rapid development and is highly sensitive to environmental stimuli, both positive and negative, making it a pivotal stage for growth and development.1
Despite substantial evidence underscoring the importance of early intervention programmes for optimal childhood development,2 over 250 million children in low- and middle-income countries (LMICs) are at risk of not reaching their full developmental potential.6 This issue is particularly prevalent in sub-Saharan Africa (SSA), where factors such as poverty, malnutrition and a lack of stimulating home environments contribute to these developmental risks.7 8 However, effective intervention can only occur if these children are identified early. Without early identification, opportunities to support their growth and development are missed, leading to long-term consequences for their academic, economic, behavioural and socio-emotional success. Thus, early identification is not only the first step, it is a crucial step to ensure that every child at risk receives the support they need to thrive.
Traditional population-level assessments rely on proxy measures like stunting or poverty levels, which may not fully capture a child’s holistic developmental progress.5 In contrast, direct assessment tools offer precise evaluations but are resource-intensive and time-consuming.5 This gap in ECD assessments is especially evident in low-resource settings, where scalable and accurate tools are urgently needed. Artificial intelligence techniques such as machine learning (ML) can bridge the gap between large-scale assessments and individual evaluations by analysing complex datasets to identify subtle patterns in child development.9 By leveraging existing data like caregiver reports and clinical records, ML can predict developmental outcomes at lower costs, making it possible to identify at-risk children early on. This enables timely and targeted interventions that are crucial for supporting optimal development. Integrating ML into developmental assessments has the potential to transform child development monitoring globally, providing a scalable and cost-effective solution that enhances early identification and addresses the limitations of traditional methods.
Despite ML’s potential for identifying health outcomes,10 its role in ECD research remains underexplored. This scoping review aims to map the existing literature, identify research gaps and highlight opportunities for leveraging ML to support young children’s full developmental potential. The primary research question guiding this review is: what is the current state of the use of ML in ECD research? As a secondary focus, we were also keen to explore the types of ML techniques applied, where they have been used, key variables examined (outcomes and features), data types, data splitting and validation strategies used, key themes, model performance, model explainability, clinical relevance and reported limitations. This review seeks to identify research gaps and inform future research directions that can maximise ML’s potential for enhancing early childhood outcomes globally.
Methods
This scoping review followed the methodological framework outlined by Arksey and O’Malley,11 which was later enhanced by Levac et al.12 To begin the scoping review, a suitable team was assembled consisting of experts in ML, ECD research and research synthesis.12 The team agreed on the broad research question to be addressed and the overall study protocol. The review includes the following five key steps: (1) identifying the research question, (2) identifying relevant studies, (3) study selection, (4) charting the data and (5) collating, summarising and reporting the results. The last optional step of the framework, the ‘consultation exercise’, was not conducted, as stakeholder engagement was not required for this scoping review. A detailed review protocol was registered with the Open Science Framework on 2 July 2024. The protocol can be accessed using this link: https://osf.io/tdh86.
Research question
This review was guided by the question, ‘What is the current state of the use of ML techniques in ECD research?’ The purpose of this review was to map the existing literature on the use of ML in ECD research, including its geographical distribution, to identify research gaps and inform future directions.13
Data sources and search strategy
The initial search was conducted on 16 June 2024, in four electronic databases: PubMed, Web of Science, IEEE Xplore and PsycINFO, online supplemental text 1. Other than age limit filters placed on PubMed (birth to 18 years) and PsycINFO (birth to 12 years), no limits were applied to the database search. The search query was tailored to the specific requirements of each database and consisted of terms deemed by the authors to describe the scope of the review. These terms included “child”, “machine learning”, “physical development”, “language development”, “cognitive development”, “socioemotional development” and variants of these main terms. Boolean operators “OR” and “AND” were used to connect and refine the search results. The databases were selected due to their comprehensive coverage of peer-reviewed literature and their ability to allow for complex search string construction, Boolean operators and filtering options to refine search results effectively. Additionally, we conducted searches on grey literature databases, such as OpenGrey, on 7 September 2024, as well as reference tracking and hand-searching to identify relevant articles not captured in the initial search.
On search completion, all the identified citations were imported into Rayyan,14 a software used to organise and manage literature reviews. Duplicates were identified and removed using the software. The remaining citations were managed for subsequent title and abstract relevance screening, as well as full-text screening by the software.
Eligibility criteria
To be included in the review, studies needed to focus on children in their early years of life (0–8 years).4 Eligible articles had to be written in English, use ML techniques or their variants such as deep learning (DL) and natural language processing techniques, and examine child development in specific domains such as cognitive, language, social-emotional and physical or motor development. There were no restrictions on the geographical location of the studies to allow for a comprehensive understanding of where ML has been applied across diverse populations and settings. Articles were excluded if the full text could not be obtained, if they examined the area of robotics; or if they focused on neurodevelopmental disorders such as autism spectrum disorder, attention-deficit/hyperactivity disorder, communication disorders or any disease condition such as malaria. Review articles were also excluded from consideration. Titles and abstracts were screened based on the inclusion criteria. The full text of the selected citations was then carefully evaluated by two independent reviewers (FNB and DC) to ensure they met the inclusion criteria, with reasons for excluding any sources that did not qualify being documented and reported. Any disagreements between the reviewers were resolved through discussions involving additional members of the author team.
Data extraction and charting
Data from the articles included in the review were extracted and organised by three independent reviewers (FNB, DC and PNM) using a data extraction template created in MS Excel. Each reviewer completed the template independently. The extracted data covered various aspects, such as the country of origin, study aim, study theme, sample size, participants’ age, study outcomes, outcome measures, features, algorithms used, data types, data splitting and validation strategies, model performance, model explainability, clinical relevance and study limitations. The MS Excel template was refined as needed during the data extraction process. Studies were also excluded at this phase if they were found to not meet the eligibility criteria. The data extracted was further summarised to enhance the synthesis process.
A narrative approach, along with descriptive statistics and visualisation techniques, was employed to synthesise the extracted data from the included studies. This analysis encompassed several key areas, including publication trends over time, distribution of study themes and clinical relevance, types of algorithms used, countries represented, data types/sources used, outcomes and key features examined, data splitting and validation strategies, model performance and explainability and reported limitations. The goal was to understand the scope and focus of the studies across different geographical and methodological contexts. Visual representations such as flow charts and bar graphs were used to enhance the clarity and interpretation of these findings. The analysis was conducted using Python software V.3.11.4.
Patent and public involvement
Patients or the public were not involved in the design, or conduct, or reporting, or dissemination plans of this research.
Results
Selection of sources of evidence
The search process identified 759 unique records. After screening and evaluating the articles for eligibility, 27 were selected for the final analysis. The search results are illustrated in a flow diagram in figure 1, while the data extracted from the included studies is presented in online supplemental tables 1,2.
Figure 1. Preferred Reporting Items for Systematic Review and Meta-Analysis flow diagram showing the process of the scoping review.
Geographical representation and publication trend
The 27 studies included in this scoping review originated from 12 different countries, with the USA leading in publications (n=11, 40.7%),15,25 followed by the Republic of Korea (n=4, 14.8%).26,29 Sri Lanka and China each contributed two publications (n=2, 7.4%),30 31 while Finland, Bangladesh, Sweden, the Netherlands, Japan, the UK, Ireland and India each contributed one publication (n=1, 3.7%).32,39 Over the 7 years covered in this review, publication numbers increased from one publication in 2018 to seven publications in 2023. As of the time of study retrieval, 16 June 2024, only one study had been published. The geographical distribution of studies is shown in figure 2.
Figure 2. A map showing the distribution of included studies worldwide.
Algorithms
To analyse the types of algorithms in online supplemental table 1, six categories were established: (1) supervised ML classifiers (SMLC), (2) DL, (3) supervised machine learning regressors (SMLR), (4) DL combined with SMLC, (5) DL combined with transfer learning (TL) and (6) TL alone. Some studies employed a mix of approaches, using either a single type of algorithm or multiple algorithms. The most frequently used ML algorithms were SMLC (n=11, 40.7%),16 20 22 24 26 30 32 34 36 37 40 followed by DL (n=6, 22.2%)25 27 29 31 39 41 and SMLR (n=4, 14.8%).15 19 21 38 Additionally, some studies combined DL with other algorithms (SMLC, SMLR and TL) (n=5, 18.5%),17 18 23 28 33 and TL alone (n=1, 3.7%).35
Outcomes and important predictors/features
To analyse the outcomes in online supplemental table 1, seven categories were established: (1) cognitive, language and motor; (2) cognitive, language and emotional; (3) cognitive and motor development; (4) cognitive; (5) language; (6) motor; and (7) overall development. The most frequently investigated outcome in the studies was cognitive development (n=9, 33.3%),1618 23 25 30 35,38 followed by motor development (n=8, 29.6%)15 19 26 29 32 33 39 41 and language development (n=6, 22.2%).20 22 27 28 34 40 Some studies examined combinations of outcomes, including cognitive, language and motor development (n=1, 3.7%)17; cognitive, language and emotional development (n=1, 3.7%)31; and cognitive and motor development (n=1, 3.7%).21 One study focused on overall development (n=1, 3.7%),24 while another did not specify any outcomes (n=1, 3.7%),28 (figure 3).
Figure 3. Outcomes examined in the included studies.

A range of features was identified across the included studies as predictive of motor, cognitive, language or overall developmental outcomes in early childhood. The most frequently reported predictors fell into six broad categories.
First, brain features, both functional and structural, were predictive of motor, cognitive and language development delays. These included the frontal, limbic, occipital and parietal lobes; postcentral gyrus; superior occipital gyrus; subcortical grey features; thalamic features; curvature of the temporal lobe and insula; as well as morphometric and white matter features.
Second, anthropometric and clinical/biological indicators, such as gestational age, birth weight, head circumference, length, weight and Apgar score, were frequently reported in studies predicting motor and cognitive outcomes.
Third, motor indicators were commonly used to assess both motor and overall development. These included movement in prone, supine, sitting and standing positions, as well as key functional skills like sitting up, walking, running, jumping and climbing stairs. Additionally, some studies employed skeletal movement analysis and sensor-augmented toys to capture motor performance more objectively.
Fourth, linguistic and expressive features, such as grammatical and lexico-semantic complexity, utterance patterns, words and part-of-speech usage, were central to predicting language outcomes.
Fifth, socio-demographic and environmental factors, including maternal education, maternal age, socioeconomic status, sex, exposure to non-family language and maternal behaviours such as alcohol use and smoking, were linked to cognitive and language development.
Lastly, medical history and nutritional variables, such as a history of pulmonary hypertension, blood transfusions, antenatal corticosteroid exposure, patent ductus arteriosus, breastfeeding status, oxygen support, parenteral alimentation, maternal body mass index and twin status, were reported to be associated with cognitive and language development, online supplemental table 2.
Outcome measures
The most frequently used measure for determining outcomes was Bayley Scales of Infant and Toddler Development-third edition (n=9, 33.3%).15 17 18 22 23 30 34 37 38 Five studies did not specify the measures used (n=5, 18.5%).24 27 28 31 35 Two studies employed the Korean Developmental Screening Test for Infants and Children (n=2, 7.4%),26 29 while another two used the Mullen Scales of Early Learning (n=2, 7.4%).21 25 Other measures reported include the Kaufman Brief Intelligence Test, second edition (n=1, 3.7%),36 Test of Gross Motor Development, third edition (n=1, 3.7%),39 Movement Assessment Battery for Children, second edition (n=1, 3.7%),19 MacArthur-Bates Communicative Development Inventory (MBCDI) (n=1, 3.7%),40 MBCDI/Clinical Evaluation of Language Fundamentals-Preschool 2 (n=1, 3.7%),20 Alberta Infant Motor Scale (n=1, 3.7%),19 Test of Gross Motor Development, second edition (n=1, 3.7%),41 Reading and Math Assessments (n=1, 3.7%)16 and Developmental Reference Age Prediction (n=1, 3.7%),33 (figure 4).
Figure 4. Outcome measures used in the included studies. AIMS, Alberta Infant Motor Scale; Bayley-III, Bayley Scales of Infant and Toddler Development-third edition; CELF-P2, Clinical Evaluation of Language Fundamentals-Preschool 2; DAP, Developmental Reference Age Prediction; KBIT-2, Kaufman Brief Intelligence Test, second edition; K-DST, Korean Developmental Screening Test for Infants and Children; MABC-2, Movement Assessment Battery for Children, second edition; MBCDI, MacArthur-Bates Communicative Development Inventory; MSEL, Mullen Scales of Early Learning; RMA, Reading and Math Assessments; TGMD-2, Test of Gross Motor Development, second edition; TGMD-3, Test of Gross Motor Development, third edition.

Data type/source
To analyse the data types in online supplemental table 1, eight categories were established: video, image, image-based, sensor-based, survey-based, text, clinical and game-based data. The most frequently used type of data in the included studies was image data (n=6, 22.2%),1516 21,23 35 followed by both video data19 26 29 30 39 and image-based data,17 25 31 34 41 both representing 18.5% of the studies (n=5). Image-based data encompassed a variety of other types, including socio-demographic, clinical, audio and game data. Sensor-based data, which included socio-demographic and physical data, was used in four articles (14.8%).24 32 33 40 Survey-based data, incorporating environmental data, totalled to three articles (11.1%).16 20 36 Additionally, text data accounted for two articles (7.4%),27 28 while clinical37 and game-based data38 were each used in one article (3.7%), (figure 5).
Figure 5. Type of data used in the included studies.

Theme and clinical use/relevance
The majority of the 27 studies included in this review focused solely on prediction tasks (n=20, 74.1%).1517,27 30 32 Other studies addressed both prediction and additional tasks, such as data curation (n=2, 7.4%),28 29 data curation combined with object recognition (n=2, 7.4%),31 35 causal inference (n=1, 3.7%),16 object recognition (n=1, 3.7%)41 and anomaly detection (n=1, 3.7%).39 Out of the 27 studies reviewed, only 1 study (3.7%)33 had been practically implemented in the context of clinical trials to quantify infants’ motor performance without healthcare worker supervision, using infant wearables comprising a multisensor garment paired with an automated DL-based algorithm. The remaining 26 studies (96.3%) were still in the research or development phase. While several studies focused on early screening, identification or intervention, none had progressed to clinical application, online supplemental table 2.
Age at prediction
To analyse the ages in online supplemental table 2, five categories corresponding to the most common categories seen in the literature were established: older than 2 years, 2 years, 1–23 months, 20–71 months and 18–35 months. Most of the included studies focused on predictions made for children older than 2 years (n=10, 37.0%),16 20 27 28 31 32 35 36 38 41 followed by predictions for those at age 2 years (n=8, 29.6%)15 17 18 21 23 25 34 37 and between 1 and 23 months (n=6, 22.2%).19 22 24 30 33 40 Other studies made predictions at ages ranging from 20 to 71 months (n=1, 3.7%)29 and 18 to 35 months (n=1, 3.7%).26 One study did not specify the age (n=1, 3.7%).39
Data splitting and validation strategies
Nine studies1517 18 26 29,31 37 38 adopted a common approach of splitting the data into training and test sets using a 70:30 or 80:20 ratio. In these cases, 70% or 80% of the data was used for training and cross-validation for hyperparameter tuning, typically employing k-fold, repeated or stratified cross-validation, while the remaining 30% or 20% served as an internal test set. In contrast, other studies1621,24 33 34 36 40 relied solely on cross-validation techniques, such as k-fold or leave-one-out cross-validation, without including a separate internal test set. Some studies (n=2, 7.4%)20 25 did not specify the type of cross-validation used but reported using an independent test set for external validation. A few studies applied uncommon data split ratios, such as 98:2,32 93:728 and 85:15,39 often due to small sample sizes or other dataset-specific constraints. Some studies19 27 35 41 did not report whether or how they split their data for model training and evaluation, online supplemental table 2.
Model performance and explainability
The model performance analysis focused on the types of algorithms used, the evaluation metrics applied and the corresponding results. For SMLC, the most reported metric was the area under the curve (AUC), which appeared in six studies with values ranging from 0.75 to 0.92.23 26 30 36 37 40 One study30 also reported recall (0.832) and precision (0.823). Additionally, two studies23 36 provided specificity, with values ranging from 96% to 98%, sensitivity (86% to 89%) and accuracy values between 91% and 95%.
Accuracy was the primary performance metric in four studies,20 24 34 36 with values ranging from 76% to 94.4%. One study combined accuracy with sensitivity (0.89) and specificity (0.86). Another study22 exclusively reported sensitivity (89%–100%) and specificity (86%–100%). However, one study16 did not report conclusive results.
In studies using DL techniques, accuracy was the primary performance metric in six studies,2527,29 31 39 with reported values ranging from 78% to 99.3%. One study41 did not report performance details. Another study33 provided an R² value of 0.99. A study combining DL and TL reported AUC values of 0.85 and 0.87,17 while a study applying TL alone35 did not specify performance metrics.
For SMLR, two studies19 33 reported R² values of 0.95 and 0.99, while another study29 reported a root mean square error (RMSE) of 0.18. Other studies15 38 did not provide any performance measures.
Out of the 27 studies reviewed, approximately 41%15,1820 24 applied explanation methods to interpret model outputs. These included Random forest feature importance, permutation importance, SHapley Additive exPlanations, random stump analysis, backtracked model weights, Gradient-weighted Class Activation Mapping, causal inference techniques, Least Absolute Shrinkage and Selection Operator regression and Circos plots. The remaining 59% of studies did not use any explanation method, online supplemental table 2.
Limitations of the included studies
To analyse the limitations in online supplemental table 2, 12 categories were identified, including small sample size, lack of external validation, imbalanced datasets, a combination of small sample size and imbalanced dataset, small sample size with validation issues, unobserved confounding factors (eg, child’s age, birth weight and family income), internet connectivity, processing limitations, lack of evaluation of study features, sampling bias (eg, over-representation of low-risk, more advantaged participants, such as older, well-educated mothers with higher birth weight infants), lack of expert in labelling outcomes and missing data. 10 out of the 27 of the included studies were limited by small sample sizes (n=10, 37.0%),1519 24 25 27,30 34 38 followed by issues with external validation (n=5, 18.5%)18 23 26 40 41 and imbalanced datasets (n=2, 7.41%).20 37 Some studies faced multiple constraints, such as both a small sample size and an imbalanced dataset (n=1, 3.7%),17 or a small sample size alongside validation problems (n=1, 3.7%).33 Additional challenges included assumptions about unobserved confounding factors (n=1, 3.7%),16 internet connectivity limitations (n=1, 3.7%),31 processing issues,35 lack of evaluation of study features (n=1, 3.7%),22 sampling bias (n=1, 3.7%),36 expert labelling (n=1, 3.7%)32 and missing data (n=1, 3.7%).21 One study did not report any limitations (n=1, 3.7%).39
Discussion
This scoping review examined the application of ML techniques in ECD research, identifying existing evidence and highlighting gaps for further exploration. The results indicate a growing interest in using ML methodologies to enhance the understanding and prediction of early childhood developmental outcomes. However, notable limitations in the existing literature require attention.
One significant limitation highlighted by the review is the geographical disparity in study distribution, with most of the research originating from high-income countries, such as the USA, a number from European countries and the Republic of Korea. To the best of our knowledge, this is the first review to explicitly examine geographical disparities in the application of ML to ECD research. It reveals a striking under-representation of studies from LMICs, particularly in SSA. This persistent geographical imbalance limits the contextual relevance and scalability of current ML models for global child development. Addressing this gap will require building research capacity, improving data infrastructure and adapting ML models to low-resource contexts.42 43
The review also found that studies frequently relied on image, video and sensor-based data, while traditional survey and text-based data were less commonly used. Consistent with a prior review,44 we observed a heavy reliance on high-tech modalities. This preference for image and video data may limit the incorporation of important risk factors, such as maternal mental health and parental economic variables, which are crucial for understanding child development.3 4 Incorporating additional data sources, including parental inputs and environmental indicators, could provide a more comprehensive view of child development and enhance the robustness of ML models.
Furthermore, the findings indicate that most studies employed SMLC, DL and a combination of these approaches, a pattern that aligns with existing literature.44 45 Despite the variety of algorithms used, most studies focused primarily on prediction tasks, with limited exploration of causal inference. ML with causal inference could yield deeper insights into ECD research.46 Additionally, the emphasis on specific domains, such as cognitive, language and motor development, suggests that ML techniques have not yet been widely applied to comprehensively assess overall child development. For a holistic evaluation, it is essential to consider other domains of development, such as socio-emotional skills, alongside physical, language and cognitive growth.5 47
Diverse features, such as brain data, clinical indicators, medical history, nutrition and socio-demographic factors, were found to be predictive of the same developmental outcomes across studies. For example, cognitive development was predicted by brain-based measures, clinical/biological indicators, medical history, nutrition and socio-demographic variables. These findings mirror those in a prior neurodevelopmental review44 and suggest that future research should prioritise comprehensive, multimodal approaches that integrate diverse types of predictors to improve accuracy and generalisability. For a long time, the expenses related to brain imaging have been a barrier to its inclusion in many child development data collection processes. However, with the emergence of faster and relatively affordable imaging technologies, there is hope that more studies in LMICs will begin to incorporate imaging data.
Across all the reviewed studies, only one demonstrated practical implementation and clinical use, while the rest, despite exploring early screening, identification or intervention, remain in the research or development phase. Several studies aimed to support early screening and intervention, but none delivered integrated or actionable frameworks to guide care. Only one study indirectly supported early cognitive development, without a direct pathway to clinical implementation. Consistent with a prior review,44 this reveals a major gap between tool development and real-world application. Future directions should prioritise validating these models in diverse, real-world populations, embedding them into healthcare and education systems and linking screening outcomes to concrete early interventions. There is also a need for implementation studies that address feasibility, cost and scalability, especially in low-resource settings where early detection tools are most urgently needed.
The age distribution of predictive modelling studies in ECD reveals a concentration on children older than 2 years, with relatively fewer studies focusing on infants and younger toddlers. This presents a limitation in the context of early identification and intervention, as the optimal window for neuroplasticity occurs during infancy.5 To maximise the impact of early detection and support timely care, future research should focus on developing models tailored to younger age groups, using accessible and scalable data sources.
Data splitting and validation approaches varied widely across studies. Notably, few studies used standard train-validation-test splits, while others relied only on cross-validation or did not report their methods. Consistent with a prior review,44 external validation was rarely performed and a few studies used unconventional split ratios due to small sample sizes. Future studies should adopt standardised and transparent validation methods. Clear reporting practices and the inclusion of external validation are essential to enhance the generalisability and comparability of ML models across studies.
In evaluating model performance, the included studies used a range of metrics including accuracy, sensitivity, specificity, AUC, recall, precision, RMSE and R² values. Overall, the performance of the SMLC was reasonable and aligned with previous studies,44 45 with the most reported metric being AUC, which ranged from 0.75 to 0.92, demonstrated good discriminative ability. Accuracy, sensitivity and specificity were reported in the selected studies and demonstrated good prediction performance of the models. Similarly, studies using SMLR and DL techniques showed comparable performance, with accuracy ranging from 78% to 99.3%, and a few studies reported R² values close to 1, indicating high predictive power. Although some studies did not provide complete performance details, the reported results suggest that these algorithms have the potential to support accurate prediction. Echoing previously observed trends in,44 more than half of the studies did not apply any model explanation methods, limiting transparency and interpretability of their findings. This reduces the ability to understand which features drive predictions, especially in clinical or developmental contexts. Future studies should prioritise the use of explanation techniques to enhance model interpretability. This will support better understanding, trust and potential clinical adoption of predictive models, particularly when used by non-technical stakeholders like clinicians or caregivers.
Similar to previous findings,44 many studies were limited by small sample sizes, which could reduce the reliability and scalability of the findings.48 49 Small sample sizes are often associated with overfitting, leading to models that perform well on training data but poorly on new, unseen data.49 To mitigate this issue, future research should aim to increase sample sizes to improve the generalisability and robustness of ML models, reducing the risk of overfitting and ensuring that findings are applicable to broader populations. Importantly, even large datasets may yield poor generalisability if not accompanied by proper validation, including careful tuning and the use of held-out test sets. Another common limitation identified was imbalanced datasets, which can affect model performance and hinder accurate prediction of developmental delays (eg, cognitive, motor delays), typically the minority class in most studies. In such contexts, relying on accuracy alone can be misleading. Evaluation metrics such as precision, recall and F1-score are recommended to better capture performance in the presence of class imbalance.50 Additionally, researchers may employ sampling techniques, such as oversampling, undersampling or synthetic data generation (eg, Synthetic Minority Over-sampling Technique) to improve model performance and fairness, especially when the goal is early detection of rare but critical outcomes.51
Additional challenges noted in the review include unobserved confounding factors such as child’s age, birth weight and family income; missing data; and sampling biases due to the overrepresentation of low-risk, more advantaged participants, such as older, well-educated mothers with higher birth weight infants, which may limit the generalisability and equity of predictive models. Overcoming these issues requires rigorous data pre-processing and the integration of domain knowledge. Furthermore, computational and internet connectivity limitations were also identified as practical barriers to implementing ML models in ECD research.
Lastly, it should be acknowledged that the scoping review process had two main limitations. First, only English-language articles were included, which may have led to the exclusion of relevant studies published in other languages. This introduces a potential language bias, particularly in global research areas such as ECD and ML, where valuable insights may be published in non-English journals. Second, the review excluded studies focusing on neurodevelopmental disorders such as autism spectrum disorder, attention-deficit/hyperactivity disorder and communication disorders. While this exclusion helped maintain a clear focus on general child development, it may limit the applicability of the findings to children with developmental challenges who might also benefit from early prediction and intervention tools. Future reviews could broaden the scope to include such populations for a more comprehensive understanding.
Conclusion
This scoping review underscores the growing interest in applying ML to ECD, with a noticeable increase in publications over the past 7 years. However, the geographical representation remains heavily skewed toward high-income countries, with minimal contributions from LMICs, particularly in SSA. The most used algorithms were supervised ML classifiers and DL methods, with cognitive, motor and language development being the primary predicted outcomes. Yet, predictions were often made after the age of two, limiting opportunities for earlier intervention.
The predictors used spanned brain imaging, clinical indicators, behavioural milestones, language features, socio-demographic factors and nutrition. However, there was a reliance on high-tech data types like imaging and video, which may not be feasible in low-resource settings. Most studies focused on prediction rather than causal inference or real-world implementation, and nearly half did not apply any explainability methods, hindering model transparency. Moreover, small sample sizes, lack of external validation and imbalanced datasets were frequent limitations that threaten the robustness and generalisability of current models.
Future research should focus on collecting comprehensive, longitudinal data that captures developmental trajectories and contextually relevant factors and on expanding outcome domains to include socioemotional development and overall child development. Priority should be given to explainability, standardised validation methods and the development of scalable tools suitable for frontline use in diverse global contexts. Additionally, expanding research efforts in LMICs is critical to ensuring that ML tools for ECD are equitable, culturally relevant and practically useful in addressing developmental challenges worldwide.
Supplementary material
Footnotes
Funding: This research was supported by the Office of The Director, National Institutes of Health (OD), the National Institute of Biomedical Imaging and Bioengineering (NIBIB), the National Institute of Mental Health (NIMH) and the Fogarty International Center (FIC) of the National Institutes of Health under award number U54TW012089 (AA and AKW). The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
Prepub: Prepublication history and additional supplemental material for this paper are available online. To view these files, please visit the journal online (https://doi.org/10.1136/bmjopen-2025-100358).
Provenance and peer review: Not commissioned; externally peer reviewed.
Patient consent for publication: Not applicable.
Ethics approval: Not applicable.
Map disclaimer: The depiction of boundaries on this map does not imply the expression of any opinion whatsoever on the part of BMJ (or any member of its group) concerning the legal status of any country, territory, jurisdiction or area or of its authorities. This map is provided without any warranty of any kind, either express or implied.
Patient and public involvement: Patients and/or the public were not involved in the design, or conduct, or reporting, or dissemination plans of this research.
Data availability statement
No data are available.
References
- 1.Richter L, Black M, Britto P, et al. Early childhood development: an imperative for action and measurement at scale. BMJ Glob Health. 2019;4:e001302. doi: 10.1136/bmjgh-2018-001302. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Arnold DH, Doctoroff GL. The early education of socioeconomically disadvantaged children. Annu Rev Psychol. 2003;54:517–45. doi: 10.1146/annurev.psych.54.111301.145442. [DOI] [PubMed] [Google Scholar]
- 3.Maggi S, Irwin LJ, Siddiqi A, et al. The social determinants of early child development: an overview. J Paediatr Child Health. 2010;46:627–35. doi: 10.1111/j.1440-1754.2010.01817.x. [DOI] [PubMed] [Google Scholar]
- 4.Likhar A, Baghel P, Patil M. Early Childhood Development and Social Determinants. Cureus. 2022;14:e29500. doi: 10.7759/cureus.29500. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Black MM, Walker SP, Fernald LCH, et al. Early childhood development coming of age: science through the life course. The Lancet. 2017;389:77–90. doi: 10.1016/S0140-6736(16)31389-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Grantham-McGregor S, Cheung YB, Cueto S, et al. Developmental potential in the first 5 years for children in developing countries. The Lancet. 2007;369:60–70. doi: 10.1016/S0140-6736(07)60032-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Rahman A, Harrington R, Bunn J. Can maternal depression increase infant risk of illness and growth impairment in developing countries? Child Care Health Dev . 2002;28:51–6. doi: 10.1046/j.1365-2214.2002.00239.x. [DOI] [PubMed] [Google Scholar]
- 8.Walker SP, Wachs TD, Meeks Gardner J, et al. Child development: risk factors for adverse outcomes in developing countries. The Lancet. 2007;369:145–57. doi: 10.1016/S0140-6736(07)60076-2. [DOI] [PubMed] [Google Scholar]
- 9.Krishnan G, Singh S, Pathania M, et al. Artificial intelligence in clinical medicine: catalyzing a sustainable global healthcare paradigm. Front Artif Intell . 2023;6:1227091. doi: 10.3389/frai.2023.1227091. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Ahsan MM, Luna SA, Siddique Z. Machine-Learning-Based Disease Diagnosis: A Comprehensive Review. Healthcare (Basel) 2022;10:541. doi: 10.3390/healthcare10030541. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Arksey H, O’Malley L. Scoping studies: towards a methodological framework. Int J Soc Res Methodol. 2005;8:19–32. doi: 10.1080/1364557032000119616. [DOI] [Google Scholar]
- 12.Levac D, Colquhoun H, O’Brien KK, et al. Scoping studies: advancing the methodology. Implement Sci. 2010;5:69. doi: 10.1186/1748-5908-5-69. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Daudt HML, van Mossel C, Scott SJ. Enhancing the scoping study methodology: a large, inter-professional team’s experience with Arksey and O’Malley’s framework. BMC Med Res Methodol. 2013;13:1–9. doi: 10.1186/1471-2288-13-48. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Ouzzani M, Hammady H, Fedorowicz Z, et al. Rayyan-a web and mobile app for systematic reviews. Syst Rev. 2016;5:210.:210. doi: 10.1186/s13643-016-0384-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Kline JE, Sita Priyanka Illapani V, He L, et al. Automated brain morphometric biomarkers from MRI at term predict motor development in very preterm infants. Neuroimage Clin. 2020;28:102475. doi: 10.1016/j.nicl.2020.102475. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Wodtke GT, Ard K, Bullock C, et al. Concentrated poverty, ambient air pollution, and child cognitive development. Sci Adv. 2022;8:eadd0285. doi: 10.1126/sciadv.add0285. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.He L, Li H, Chen M, et al. Deep Multimodal Learning From MRI and Clinical Data for Early Prediction of Neurodevelopmental Deficits in Very Preterm Infants. Front Neurosci. 2021;15:753033. doi: 10.3389/fnins.2021.753033. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Chen M, Li H, Wang J, et al. Early Prediction of Cognitive Deficit in Very Preterm Infants Using Brain Structural Connectome With Transfer Learning Enhanced Deep Convolutional Neural Networks. Front Neurosci. 2020;14:858. doi: 10.3389/fnins.2020.00858. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Modayur B, Fair-Field T, Komori S. Enhancing motor screening efficiency: Toward an empirically derived abridged version of the Alberta Infant Motor Scale. Early Hum Dev. 2023;177–178:105723. doi: 10.1016/j.earlhumdev.2023.105723. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Borovsky A, Thal D, Leonard LB. Moving towards accurate and early prediction of language delay with network science and machine learning approaches. Sci Rep. 2021;11:8136. doi: 10.1038/s41598-021-85982-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Adeli E, Meng Y, Li G, et al. Multi-task prediction of infant cognitive scores from longitudinal incomplete neuroimaging data. Neuroimage. 2019;185:783–92. doi: 10.1016/j.neuroimage.2018.04.052. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Vassar R, Schadl K, Cahill-Rowley K, et al. Neonatal Brain Microstructure and Machine-Learning-Based Prediction of Early Language Development in Children Born Very Preterm. Pediatr Neurol. 2020;108:86–92. doi: 10.1016/j.pediatrneurol.2020.02.007. [DOI] [PubMed] [Google Scholar]
- 23.Ali R, Li H, Dillman JR, et al. A self-training deep neural network for early prediction of cognitive deficits in very preterm infants using brain functional connectome data. Pediatr Radiol. 2022;52:2227–40. doi: 10.1007/s00247-022-05510-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Fitter NT, Funke R, Pulido JC, et al. Toward Predicting Infant Developmental Outcomes From Day-Long Inertial Motion Recordings. IEEE Trans Neural Syst Rehabil Eng. 2020;28:2305–14. doi: 10.1109/TNSRE.2020.3016916. [DOI] [PubMed] [Google Scholar]
- 25.Girault JB, Munsell BC, Puechmaille D, et al. White matter connectomes at birth accurately predict cognitive abilities at age 2. Neuroimage. 2019;192:145–55. doi: 10.1016/j.neuroimage.2019.02.060. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Chun S, Jang S, Kim JY, et al. Comprehensive Assessment and Early Prediction of Gross Motor Performance in Toddlers With Graph Convolutional Networks-Based Deep Learning: Development and Validation Study. JMIR Form Res. 2024;8:e51996. doi: 10.2196/51996. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Oh B-D, Lee Y-K, Kim J-D, et al. Deep Learning-Based End-to-End Language Development Screening for Children Using Linguistic Knowledge. Appl Sci (Basel) 2022;12:4651. doi: 10.3390/app12094651. [DOI] [Google Scholar]
- 28.Choi J-M, Lee Y-K, Kim J-D, et al. Machine Learning-Based Automatic Utterance Collection Model for Language Development Screening of Children. Appl Sci (Basel) 2022;12:4747. doi: 10.3390/app12094747. [DOI] [Google Scholar]
- 29.Kim HH, Kim JY, Jang BK, et al. Multiview child motor development dataset for AI-driven assessment of child development. Gigascience. 2023;12:giad039. doi: 10.1093/gigascience/giad039. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Ji S, Ma D, Pan L, et al. Automated prediction of infant cognitive development risk by video: A pilot study. IEEE J Biomed Health Inform. 2023;28:690–701. doi: 10.1109/JBHI.2023.3266350. [DOI] [PubMed] [Google Scholar]
- 31.Oshadhini P, Siriwardana K, Ranmini HP, et al. In: Kelegama T, editor. KidSwatch: Children Development Observation System. 2023 5th International Conference on Advancements in Computing (ICAC); 2023. In. ed. [Google Scholar]
- 32.Brons A, de Schipper A, Mironcika S, et al. Assessing Children’s Fine Motor Skills With Sensor-Augmented Toys: Machine Learning Approach. J Med Internet Res. 2021;23:e24237. doi: 10.2196/24237. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Airaksinen M, Taylor E, Gallen A, et al. Charting infants’ motor development at home using a wearable system: validation and comparison to physical growth charts. EBioMedicine. 2023;92:104591. doi: 10.1016/j.ebiom.2023.104591. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Valavani E, Blesa M, Galdi P, et al. Language function following preterm birth: prediction using machine learning. Pediatr Res. 2022;92:480–9. doi: 10.1038/s41390-021-01779-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Mitra A, Mostafiz T, Rashid RU, editors. Photoplay: an android application to stimulate children’s cognitive development. 2017 IEEE Region 10 Humanitarian Technology Conference (R10-HTC); 2017: IEEE. In. [Google Scholar]
- 36.Bowe AK, Lightbody G, Staines A, et al. Predicting Low Cognitive Ability at Age 5-Feature Selection Using Machine Learning Methods and Birth Cohort Data. Int J Public Health. 2022;67:1605047. doi: 10.3389/ijph.2022.1605047. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Bowe AK, Lightbody G, Staines A, et al. Prediction of 2-Year Cognitive Outcomes in Very Preterm Infants Using Machine Learning Methods. JAMA Netw Open. 2023;6:e2349111. doi: 10.1001/jamanetworkopen.2023.49111. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Mukherjee D, Bhavnani S, Swaminathan A, et al. Proof of Concept of a Gamified DEvelopmental Assessment on an E-Platform (DEEP) Tool to Measure Cognitive Development in Rural Indian Preschool Children. Front Psychol. 2020;11:1202. doi: 10.3389/fpsyg.2020.01202. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Suzuki S, Amemiya Y, Sato M, editors. Skeleton-based visualization of poor body movements in a child’s gross-motor assessment using convolutional auto-encoder. 2021 IEEE International Conference on Mechatronics (ICM); 2021: IEEE. In. [Google Scholar]
- 40.Wong PCM, Lai CM, Chan PHY, et al. Neural Speech Encoding in Infancy Predicts Future Language and Communication Difficulties. Am J Speech Lang Pathol. 2021;30:2241–50. doi: 10.1044/2021_AJSLP-21-00077. [DOI] [PubMed] [Google Scholar]
- 41.Andarage ISN, Fernando D, Lokuarachchi BA, et al. In: Wijewickrama P, editor. Early Childhood Action Monitoring and Analytics System (ECAMS). 2023 IEEE 28th Pacific Rim International Symposium on Dependable Computing (PRDC); 2023: IEEE. In. ed. [Google Scholar]
- 42.Caruana R, Niculescu-Mizil A, editors. Data mining in metric space: an empirical analysis of supervised learning performance criteria. Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining; 2004. In. [Google Scholar]
- 43.Okolo CT, Aruleba K, Obaido G, et al. Responsible AI in Africa—Challenges and opportunities. Responsible AI in Africa: Challenges and opportunities. Nat Mach Intell. 2022;4:35–64. [Google Scholar]
- 44.Bowe AK, Lightbody G, Staines A, et al. Big data, machine learning, and population health: predicting cognitive outcomes in childhood. Pediatr Res. 2023;93:300–7. doi: 10.1038/s41390-022-02137-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Nyamapfene A. Computational investigation of early child language acquisition using multimodal neural networks: a review of three models. Artif Intell Rev. 2009;31:35–44. doi: 10.1007/s10462-009-9125-6. [DOI] [Google Scholar]
- 46.Cui P, Athey S. Stable learning establishes some common ground between causal inference and machine learning. Nat Mach Intell. 2022;4:110–5. doi: 10.1038/s42256-022-00445-z. [DOI] [Google Scholar]
- 47.Spittle AJ, Orton J, Doyle LW, et al. Early developmental intervention programs post hospital discharge to prevent motor and cognitive impairments in preterm infants. Cochrane Database Syst Rev. 2007;2007:CD005495. doi: 10.1002/14651858.CD005495.pub2. [DOI] [PubMed] [Google Scholar]
- 48.Ho SY, Phua K, Wong L, et al. Extensions of the External Validation for Checking Learned Model Interpretability and Generalizability. Patterns (N Y) 2020;1:100129. doi: 10.1016/j.patter.2020.100129. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Vabalas A, Gowen E, Poliakoff E, et al. Machine learning algorithm validation with a limited sample size. PLoS One. 2019;14:e0224365. doi: 10.1371/journal.pone.0224365. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One. 2015;10:e0118432. doi: 10.1371/journal.pone.0118432. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Chawla NV, Bowyer KW, Hall LO, et al. SMOTE: Synthetic Minority Over-sampling Technique. jair. 2002;16:321–57. doi: 10.1613/jair.953. [DOI] [Google Scholar]


