Abstract
Background
Diarrheal disease remains a major cause of under-five mortality in low- and middle-income countries (LMICs). This study investigates the diarrhea determinants, employing a novel dual approach comparing classical epidemiology and machine learning (ML) to determine the most important predictors and to optimize intervention targeting in Nigeria.
Methods
A cross-sectional analysis of 33,924 children aged < 5 years using 2018 Nigeria Demographic and Health Survey (NDHS) was used. Traditional logistic regression was contrasted to ML models (Random Forest, Gradient Boosting Machines, Decision Trees) to consider non-linear associations and variable importance, while using model parameters to assess model performance.
Results
The prevalence of diarrhea was 11.98%, with significant disparities between regions. Child’s age was the strongest predictor across all models, with a significant odd seen among children aged 6–23 months (AOR = 2.48–2.54). Increased maternal education was protective (AOR = 0.77–0.79), while exposure to the media had multicomponent associations. Urban-rural wealth index was a robust strong socioeconomic predictor. Logistic regression had the best predictive performance (AUC = 0.727), closely followed by that of Gradient Boosting (AUC = 0.718). Sensitivity analysis showed that MICE generated more accurate estimates than complete case analysis because missingness was not random.
Conclusion
The study stresses the significance of core modifiable determinants such as maternal education and contextual wealth. Methodological triangulation illustrates the complementarity of classic regression for inference and machine learning for feature discovery. These findings justify the imposition of hyper-localized, multi-sectoral interventions on high-risk age groups and areas based on sound data analysis to optimize public health resources. This mixed approach provides a scalable model for disease burden measurement in LMICs.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12889-025-25076-y.
Keywords: Childhood diarrhea, Machine learning, WASH, Socioeconomic determinants, Nigeria, Predictive modeling
Introduction
Diarrhea disease remains a leading cause of morbidity and mortality for children under five years of age, particularly in low- and middle-income countries (LMICs) [1]. Despite the success of WASH interventions across the globe, diarrhea remains a key public health challenge with an estimated 525,000 child deaths per year [2]. The burden of diarrhea disease morbidity is disproportionately concentrated in South Asia and sub-Saharan Africa, where socio-economic disadvantages, inadequate infrastructure, and limited healthcare access intersect to amplify exposure [3]. An understanding of the complex determinants of childhood diarrhea, spanning environmental, socioeconomic, and behavioral domains, is critical to conceptualizing effective, targeted interventions aimed at reducing its prevalence and impact.
The etiology of diarrhea disease in children is multifactorial, comprising interactions between pathogens (such as rotavirus and E. coli), environmental contamination, and intrahousehold level risk factors [4]. Contaminated drinking water, inadequate sanitation, and unsafe fecal disposal are all established risk factors in the transmission of diarrhea disease [5]. However, newer evidence is revealing that wider socioeconomic and demographic factor such as maternal education, household income, and exposure to media equally influence health outcomes [6]. For instance, improved maternal education has been associated with improved hygiene practices and access to healthcare, while poverty limits access to clean water and sanitation facilities [7].
Regional variations further complicates diarrhea epidemiology. Variation in climate, infrastructure, and health care systems of sub-Saharan Africa accounts for different disease burdens in urban and rural settings [8]. Rural residents, for example, experience more unimproved water exposure and open defecation, while urban slums experience overcrowding and unhygienic sanitation [9]. In addition, cultural behavior and caregiving practices, such as hand hygiene practice and child feces disposal, act as mediators of environmental hazard and incidence of diarrhea [10].
One other area of increasing focus is the use of media and technology in health education and behavior change. Exposure to radio, TV, and internet has been associated with increased health literacy and the use of preventive care, though digital disparities could potentially be particularly virulent in aiding and abetting inequalities [11]. For example, houses with ceaseless media exposure are more likely to have sanitary habits, while excluded communities continue to rely on non-mainstream, and even ineffective channels for conveying health messages [12].
Machine learning (ML) methods hold promise for demystifying the multifaceted predictors of diarrhea disease by detecting nonlinear relationships and interaction effects that may go undetected through traditional statistical modeling [13]. Logistic regression, which is a widely used tool in epidemiologic analysis, yields interpretable odds ratios but perhaps may fail to accommodate high-dimensional interactions [14]. However, ensemble approaches such as Random Forest (RF) and Gradient Boosting Machines (GBM) are able to cope with heterogeneous data and rank variables by their predictive strength, providing deeper insight into risk hierarchies [15]. Decision trees, being easy to interpret and possessing simple split rules, also assist in the detection of high-risk subgroups for targeted interventions [16].
Findings from this study portend implications for programming and public health policy. By illuminating the relative roles of modifiable risk factors, such as sanitary practice, media use, and maternal education, the present study can guide the design of context-specific interventions. For example, investment in WASH infrastructure can be more valuable when access to water is the critical bottleneck, whereas media campaigns would be of highest value where radio or TV penetration is higher [17]. Additionally, integration of ML tools into public health analysis can optimize the accuracy and scalability of disease prevention interventions, especially in poor-resource settings [18].
This research therefore, leverages a multidimensional data set with socioeconomic, environmental, and behavioral variables to investigate determinants of childhood diarrhea. By employing convectional regression and advanced ML models, we seek to: (1). determine and measure the strength of relationship between some risk and protective factors (e.g., exposure to media, WASH, education of mother) and the prevalence of diarrhea, using multivariable logistic regression to produce interpretable odds ratios. (2). establish complex, non-linear relationships and correlations between these variables, and establish their relative predictive significance through ensemble techniques (Random Forest, Gradient Boostin and Decision Trees) (3). contrast the performance and results of these two modeling frameworks to ascertain their relative validity and value for the prediction of diarrhea outcomes.
Methodology
Study design and data source
This research used a cross-sectional study to assess determinants of childhood diarrhea among children under the age of five. The data was extracted from a nationally representative survey that was carried out in Nigeria (NDHS 2018), gathering detailed demographic, socioeconomic, and health variables.
Study population and sampling
The survey employed a multistage stratified cluster sampling technique to achieve representativeness in rural and urban settings, as well as strata of socioeconomic status. Eligible households included those with one or more children who were under five years of age. The primary outcome variable was the childhood diarrhea status, defined by loose stools or watery stools, three or more times in a 24-hour period in the preceding two weeks before the survey [19].
Data preprocessing and missing data handling
A multimethod analytical approach that combined traditional epidemiological methods with contemporary ML techniques to identify the determinants of childhood diarrhea and compare the relative performance of various predictive modeling approaches was used. Subsequently, the analysis was applied on a nationally representative sample of 33,924 children aged under five years using the latest Nigeria Demographic and Health Survey (NDHS).
However, one of the significant challenges was that there existed a very high percentage of missing data for some of the most important predictor variables. Some of them, e.g., disposal of child faeces, exceeded 37% in missing data. To deal with this scientifically and investigate its possible influence on the findings, we adopted a two-pronged approach. We first employed Multiple Imputation by Chained Equations (MICE), a conservative statistical method that generates multiple plausible imputations for missing data according to the distributions and inter-relationships of all other variables within the data set for the main analysis [20]. It maintains original sample size and statistical power and compensates for the uncertainty generated through imputation. To check the sensitivity of our findings to missing data handling at the case level, a complete case analysis (CCA) was also done simultaneously where any record with even one missing value on any of the variables under analysis was not analyzed. The results from both methods (MICE and CCA) were then statistically compared in sensitivity analysis to determine variables for which effect estimates, significance, or direction had changed largely, thereby quantifying the degree of potential missingness-induced bias.
Multilevel logistic regression and cluster analysis
The epidemiological analysis was carried out based on the analysis of multilevel mixed-effects logistic regression model. The application of this type of model was aimed at tackling the hierarchical nesting of DHS data, with children (level 1) nested in households in primary sampling units (clusters) and regions (level 2). Not controlling for this clustering would be a violation of the independence assumption of the observations and risks spuriously tight confidence intervals [21]. The model was controlled for a complete list of potential confounders, including maternal education, age, household wealth indices (standard and urban-rural specific), media exposure (internet, TV, radio), WASH facilities, household composition, and child characteristics (child’s age, sex). These are reported as adjusted odds ratios (AORs) with 95% confidence intervals (CIs). Additionally, in an effort to narrow the gap in analysis between the ML models and to examine the relative importance of each predictor within the regression model, a post-estimation dominance analysis was conducted. This is achieved by measuring the relative contribution of each of the variables in terms of the decrease in the McFadden’s R² of the model when that predictor is omitted, which allows standardized ordering of variable importance.
Comparative and sensitivity analysis
The underlying framework of this study is the triangulation of results by methodological approaches. The determinants identified by the epidemiological model (through odds ratios and dominance analysis) were compared directly with the most dominant features highlighted by the ML models. Such a comparison identifies consensus predictors within paradigms and calls for model-specific results. Finally, and as indicated, the exhaustive sensitivity analysis of the comparison of the results of the CCA and MICE offers a rigorous test of our results’ sensitivity to varying assumptions under the missing data mechanism, validating the results that are reported. All the analyses were performed on R statistical software (version 4.3.1) using packages such as mice for imputation, lme4 for multilevel modeling, and caret and xgboost for machine learning.
Machine learning approaches
To complement the traditional regression analysis, we also employed three machine learning (ML) models. The Random Forest (RF) model, which is an ensemble modeling approach, creates multiple decision trees and aggregates their predictions, giving feature importance scores based on mean decrease in accuracy [22]. The Gradient Boosting Machine (GBM), a boosting algorithm that sequentially enhances predictive performance by reducing the errors made by previous iterations [23], was used and variable importance was determined by relative influence. We also utilized a Decision Tree (DT) model, which partitions data recursively using predictor splits and produces interpretable decision rules [24]. These ML methods were chosen to mainly address intricate relationships and to provide insights into significant predictors.
Model performance evaluation
Multiple metrics were used to evaluate the predictive ability of all models. Discrimination ability was assessed by the Area Under the Receiver Operating Characteristic Curve (AUC-ROC) where higher AUC values indicated better classification. Further statistics such as sensitivity, specificity and accuracy were calculated to evaluate the classification performance at different thresholds. Calibration curves were also assessed for the agreement of predicted probabilities and observed events [25]. Models were trained on 70% of the data while 30% data was used for the validation, with performance metrics averaged over 10-fold cross-validation for robustness [26]. By adapting these evaluation tactics, we were able to robustly evaluate model reliability and prediction ability.
Ethical considerations
The original survey had ethical approval from the National Health Research Ethics Committee of Nigeria (NHREC) and the ICF Institutional Review Board [27] and obtained informed consent from all participants. This study was a secondary analysis of anonymized data from which no re-identification was possible.
Results
Descriptive characteristics of the study population
The analysis used data from 33,924 children aged less than five years. The prevalence of diarrhea in the two-week history preceding the survey was 12.0% (n = 4,066) (Table 1). A large percentage of data were missing for some of the important key variables, e.g., child’s faeces disposal (37.1% missing) and diarrhea status (9.5% missing). The sample population showed that most of the mothers belonged to the age group 25–29 (27.9%) and that almost half of them were totally uneducated (45.4%). Wealth distribution was fairly equal, and most of the households belonged to the “Very Poor” (23.8%) and “Poor” (22.8%) categories. The majority of the children geographically belonged to the North-West region (30.4%), and most of them lived in rural areas (65.5%). The distribution of children by age was also very balanced, and most children were in the groups 36–47 months (20.7%) and 48–59 months (20.8%) respectively.
Table 1.
Descriptive statistics table
| Variable | Category | Count | Percentage |
|---|---|---|---|
| MothersAge | 15–19 | 1434 | 4.23% |
| 20–24 | 6626 | 19.53% | |
| 25–29 | 9470 | 27.92% | |
| 30–34 | 7647 | 22.54% | |
| ≥ 35 | 8747 | 25.78% | |
| Mothers’ Level of Education | Not Educated | 15,391 | 45.37% |
| Primary (Complete) | 3655 | 10.77% | |
| Primary (Incomplete) | 1619 | 4.77% | |
| Secondary (Complete) | 6696 | 19.74% | |
| Secondary (Incomplete) | 3927 | 11.58% | |
| Higher Education | 2636 | 7.77% | |
| Wealth Index | Very Poor | 8066 | 23.78% |
| Poor | 7743 | 22.82% | |
| Average | 7171 | 21.14% | |
| Rich | 6166 | 18.18% | |
| Very rich | 4778 | 14.08% | |
| Urban Rural Wealth index | Very Poor | 7443 | 21.94% |
| Poor | 7299 | 21.52% | |
| Average | 6932 | 20.43% | |
| Rich | 6477 | 19.09% | |
| Very Rich | 5773 | 17.02% | |
| Access to Electricity | No | 17,235 | 50.8% |
| Yes | 16,689 | 49.2% | |
| Listens to Radio | No | 16,552 | 48.79% |
| Once a week | 8072 | 23.79% | |
| >Once a week | 9300 | 27.41% | |
| Watches TV | No | 19,733 | 58.17% |
| Once a week | 5813 | 17.14% | |
| >Once a week | 8378 | 24.7% | |
| Ever used the Internet | No | 30,766 | 90.69% |
| Yes | 3158 | 9.31% | |
| Internet use Last month | Not in the past month | 31,307 | 92.29% |
| < Once a week | 450 | 1.33% | |
| < At least once a week | 938 | 2.77% | |
| Almost everyday | 1229 | 3.62% | |
| Shared Toilet Facility | No | 16,619 | 48.99% |
| Yes | 8346 | 24.6% | |
| Missing | 8959 | 26.41% | |
| Type of Toilet facility | Unimproved | 17,710 | 52.2% |
| Improved | 16,214 | 47.8% | |
| Time to water | On Premises | 10,206 | 30.08% |
| < 30 min | 18,809 | 55.44% | |
| > 30 min | 4909 | 14.47% | |
| Drinking Water Source | Unimproved | 14,418 | 42.5% |
| Improved | 19,506 | 57.5% | |
| Total Household members | ≤ 6 Persons | 15,159 | 44.69% |
| ≥ 7 Persons | 18,765 | 55.31% | |
| Total under 5 in households | ≤ 2 | 23,437 | 69.09% |
| ≥ 3 | 10,487 | 30.91% | |
| Region | North-Central | 5875 | 17.32% |
| North-East | 7211 | 21.26% | |
| North-West | 10,305 | 30.38% | |
| South-East | 3798 | 11.2% | |
| South-South | 3202 | 9.44% | |
| South-West | 3533 | 10.41% | |
| Residence | Rural | 22,225 | 65.51% |
| Urban | 11,699 | 34.49% | |
| Childs’ Sex | Female | 16,667 | 49.13% |
| Male | 17,257 | 50.87% | |
| Childs’ Age (In Months) | ≤ 5 | 3409 | 10.05% |
| 6–11 | 3350 | 9.88% | |
| 12–23 | 6543 | 19.29% | |
| 24–35 | 6540 | 19.28% | |
| 36–47 | 7011 | 20.67% | |
| 48–59 | 7071 | 20.84% | |
| Disposal of faeces | Safe | 11,743 | 34.62% |
| Unsafe | 9587 | 28.26% | |
| Missing | 12,594 | 37.12% | |
| Diarrhea Status | No | 26,647 | 78.55% |
| Yes | 4066 | 11.98% | |
| Missing | 3211 | 9.47% |
Epidemiological determinants of childhood diarrhea
Multilevel logistic regression, accounting for regional-level clustering and handling of missing values through Multiple Imputation by Chained Equations (MICE) (Supplementary Table 1), revealed a number of predictors of diarrhea among children. The most significant associations were with the age of the child: compared with children less than 6 months old, children aged 6–11 months (AOR = 2.54, 95% CI: 2.20–2.94, p < 0.001) and 12–23 months (AOR = 2.48, 95% CI: 2.18–2.83, p < 0.001) had significantly elevated odds of diarrhea. 36–47 months-old children also possessed higher odds (AOR = 1.48, 95% CI: 1.29–1.69, p < 0.001), whereas 48–59 months-old children possessed significantly lower odds (AOR = 0.63, 95% CI: 0.54–0.73, p < 0.001) of diarrhea.
Increased maternal education was protective; mothers with complete secondary education (AOR = 0.79, 95% CI: 0.70–0.90, p < 0.001) or higher education (AOR = 0.77, 95% CI: 0.62–0.95, p = 0.017) had significantly reduced odds of the child having diarrhea as compared to mothers without education. Intriguingly, media exposure was complex: listening to the radio more than once a week was associated with increased odds (AOR = 1.32, 95% CI: 1.21–1.45, p < 0.001) of diarrhea. There was also a paradoxical relationship with internet usage, as ever using the internet was associated with increased risk of diarrhea (AOR = 1.38, 95% CI: 1.01–1.87, p = 0.042), yet increased usage in the previous month was protective (e.g., nearly every day: AOR = 0.63, 95% CI: 0.43–0.92, p = 0.018). Urban-rural wealth index was very predictive, with progressively richer quintiles exerting a graded, protective effect (e.g., “Very Rich”: AOR = 0.69, 95% CI: 0.53–0.90, p = 0.006). Higher household size (≥ 7 individuals) also had lower odds of diarrhea (AOR = 0.86, 95% CI: 0.79–0.94, p < 0.001).
Sensitivity analysis
Sensitivity analysis between the main MICE analysis and a Complete Case Analysis (CCA) (Table 2, Supplementary Tables 1 and 2), indicated that much of the findings were robust, with high consistency for variables such as child’s age and listening to the radio. However, material differences were observed with other variables. For instance, the role played by proper faeces disposal to protecting against diarrhea was significant in the CCA (AOR = 0.85, p = 0.005) but not in the MICE analysis (AOR = 0.96, p = 0.230). Likewise, the size and, in some cases, the sign of coefficients for wealth indicators and the urban-rural wealth indicator were significantly different across analysis, proving that the missingness mechanism of the above socioeconomic variables is not random, and that the MICE method yields a more reliable estimate.
Table 2.
Sensitivity Analysis: the following table presents the adjusted odds ratios (AOR) and p-values for the primary (MICE) and sensitivity (CCA) analyses. Bold highlighting indicates variables where substantive differences in effect size, significance, or direction were observed
| Variable | Categories | Primary Analysis (MICE) | Sensitivity Analysis (Complete Case) | Consistency | ||
|---|---|---|---|---|---|---|
| AOR (95% CI) | p-value | AOR (95% CI) | p-value | |||
| Listens to Radio | >Once a week | 1.32 (1.21–1.45) | < 0.001 | 1.34 (1.18–1.52) | < 0.001 | High |
| Watches Television | Once a week | 1.10 (0.98–1.23) | 0.104 | 1.30 (1.11–1.52) | 0.001 | Low |
| >Once a week | 1.12 (0.99–1.26) | 0.073 | 1.36 (1.15–1.62) | < 0.001 | Low | |
| Disposal of Faeces | Safe | 0.96 (0.89–1.03) | 0.230 | 0.85 (0.76–0.95) | 0.005 | Low |
| Wealth Index | Poor | 1.02 (0.90–1.16) | 0.758 | 1.57 (0.94–2.64) | 0.087 | Low |
| Average | 0.87 (0.72–1.06) | 0.165 | 1.23 (0.75–2.01) | 0.413 | Low | |
| Rich | 0.87 (0.68–1.11) | 0.270 | 1.03 (0.62–1.69) | 0.920 | Low | |
| Very Rich | 0.75 (0.54–1.06) | 0.102 | 0.80 (0.47–1.34) | 0.392 | Low | |
| Urban-Rural Wealth Index | Poor | 0.84 (0.76–0.94) | 0.002 | 0.54 (0.32–0.92) | 0.023 | Low |
| Average | 0.75 (0.64–0.87) | < 0.001 | 0.54 (0.33–0.90) | 0.018 | Low | |
| Rich | 0.76 (0.62–0.94) | 0.011 | 0.66 (0.40–1.12) | 0.122 | Medium | |
| Very Rich | 0.69 (0.53–0.90) | 0.006 | 0.57 (0.31–1.03) | 0.060 | Medium | |
| Time to Water Source | > 30 min | 1.02 (0.91–1.14) | 0.745 | 1.21 (1.03–1.43) | 0.019 | Low |
| Residence | Urban | 0.96 (0.82–1.12) | 0.588 | 0.66 (0.41–1.04) | 0.076 | Low |
| All Other Variables & Categories | High | |||||
Abbreviations: AOR Adjusted Odds Ratio, CI Confidence Interval, MICE Multiple Imputation by Chained Equations
High Consistency: Effect estimates and conclusions are congruent. Medium Consistency: Effect direction is similar, but magnitude or significance differs. Low Consistency: Notable differences in effect size, significance, or direction
Machine learning model performance
The four machine learning models differed considerably in their predictive performance. Logistic Regression yielded the best discriminative capacity with an AUC of 0.727 (Fig. 1), followed very closely by Gradient Boosting Machine (GBM) (AUC = 0.718) and Random Forest (AUC = 0.684). The Decision Tree model was only slightly better randomly by chance at an AUC of 0.5. A 5-fold cross-validation results also validated the stability of such rankings (Fig. 2). Random Forest had the lowest error (0.100) on error rate and then GBM (0.120), Logistic Regression (0.150), and Decision Tree (0.220) (Fig. 3). Predicted probability distributions also uncovered model behaviors: Logistic Regression and GBM produced well-calibrated probability distributions between 0.0 and 0.8, but the Decision Tree model had a bias towards extreme probability predictions (nearly 0 or 1), reflecting overconfidence and poor calibration [28]. (Fig. 4). Thus, model calibration inspection of the predicted probability distributions indicated that logistic regression yielded the smoother probability estimates compared to the tree-based methods with better calibration.
Fig. 1.
Fig. 1: This presents a Receiver Operating Characteristic (ROC) curve comparison for four machine learning models: Based on the provided ROC curve comparison, this figure illustrates the diagnostic performance of four machine learning models—Decision Tree (AUC = 0.5), Logistic Regression (AUC = 0.727), GBM (AUC = 0.718), and Random Forest (AUC = 0.684)—by plotting their true positive rate (sensitivity) against the false positive rate (1 - specificity). The closer a curve follows the left-hand border and then the top border of the plot space, the better the model's performance, with Logistic Regression achieving the highest AUC and thus the best overall predictive accuracy in distinguishing between classes, while the Decision Tree model, with an AUC of 0.5, performs no better than random chance.
Fig. 2.
The figure compares the performance of four machine learning models: This figure presents a comparison of model performance using the ROC AUC score from 5-fold cross-validation, evaluating four machine learning algorithms: Decision Tree (DT), Gradient Boosting Machine (GBM), Logistic Regression, and Random Forest (RF). The plot displays the distribution of AUC scores for each model, with the red diamond symbol representing the mean ROC AUC value for the respective classifier. The performance is quantified on a scale from 0.5 to 0.9, allowing for a clear visual assessment of each model’s predictive power and stability across different validation folds. The plot visually ranks the models based on their ability to distinguish between classes, where a score closer to 1.0 signifies near-perfect discrimination. This comparison aids in selecting the most effective model for predictive tasks
Fig. 3.
This figure compares the error rates of four machine learning models—Decision Tree, Logistic Regression, Gradient Boosting Machine (GBM), and Random Forest—across different evaluation metrics, with lower values indicating superior performance. The bar chart visualizes the specific error rates for each model, where Random Forest achieved the lowest error (0.100) and Decision Tree the highest (0.250), highlighting the relative effectiveness of each algorithm in minimizing prediction inaccuracies on the Diarrhea Dataset. The comparison underscores the performance variance between models, with ensemble methods like Random Forest and GBM generally demonstrating lower error rates compared to simpler models like Decision Tree. The visualization helps identify trade-offs between metrics, such as precision-recall imbalances, and guides model selection based on task-specific priorities (e.g., minimizing false positives vs. maximizing true positives)
Fig. 4.
The figure displays the predicted probability distributions for four machine learning models-This figure displays the predicted probability distributions for the outcome of diarrhea, as generated by four distinct machine learning models: Decision Tree, Gradient Boosting Machine (GBM), Logistic Regression, and Random Forest. Each distribution curve illustrates the density of instances across a range of probabilities from 0.0 to 0.8, revealing how each model calibrates its predictions and where it concentrates its certainty or uncertainty. The varying shapes and peaks of the density plots provide insight into the models’ predictive behaviors, such as a tendency towards extreme probabilities (e.g., near 0 or 1) or a more conservative, distributed approach across the probability spectrum. By comparing these distributions, the figure highlights model-specific biases, overconfidence, or under-confidence in predictions, aiding in the selection of the most reliable classifier for probabilistic outcomes
Variable importance across modeling paradigms
Relative predictor importance was measured using dominance analysis for the regression model (Table 3) and intrinsic importance metrics for the ML models (Table 4). Both methods predominantly established child’s age as the single greatest predictor, explaining 23.9% of the variance that could be explained in the dominance analysis. The second strongest predictor in regression analysis (22.8% importance) was child’s disposal of faeces. The Random Forest and Decision Tree models gave extremely high weight to location and child’s age, with scores of 87.5 and 82.8, respectively. The mother’s educational level and wealth proxies were strong predictors in models leaving out one, although their relative ranking differed. For example, urban-rural wealth index was a top-five feature in the Decision Tree but of little importance in the GBM model (Fig. 5). This contrast is evidence of good convergence on key determinants (region and child’s age) and also showcases algorithm-specific prioritization of other characteristics.
Table 3.
Relative importance of predictors from dominance analysis of the multilevel logistic regression model
| Variables | Categories | General Dominance Statistic | Standardized Importance (%) |
|---|---|---|---|
| Childs’ Age1 | 48–59, 36–47, 24–35, 12–23, 6–11, ≤ 5 | 0.0817 1 | 23.9% 1 |
| Disposal of Childs’ faeces1 | Safe, Unsafe | 0.0779 1 | 22.8% 1 |
| Shared Toilet Facility1 | Yes/No | 0.0346 1 | 10.1% 1 |
| Mothers Age | ≥ 35, 30–34, 25–29, 20–24, 15–19 | 0.0268 | 7.8% |
| Listens to Radio | No, Once a week, >Once a week | 0.0216 | 6.3% |
| Wealth Index | Very Rich, Rich, Average, Poor, Very Poor | 0.0139 | 4.1% |
| Total under 5 in households | ≥ 3/≤2 | 0.0134 | 3.9% |
| Mothers’ Level of Education | No Education, Primary (Complete), Primary (Incomplete), Secondary (Complete), Secondary (Incomplete), Higher Education | 0.0132 | 3.9% |
| Childs’ Sex | Male/Female | 0.0109 | 3.2% |
| Source of Drinking Water | Improved/Unimproved | 0.0080 | 2.3% |
| Total in Household | ≥ 7 Persons/≤6 persons | 0.0076 | 2.2% |
| Internet Use Last month |
Almost everyday At least once a week < than once a week |
0.0069 | 2.0% |
| Type of Toilet Facility | Improved/Unimproved | 0.0069 | 2.0% |
| Urban Rural Wealth index | Very Rich, Rich, Average, Poor, Very Poor | 0.0044 | 1.3% |
| Watches TV | No, Once a week, >Once a week | 0.0038 | 1.1% |
| Access to Electricity | Yes/No | 0.0037 | 1.1% |
| Ever used the Internet | Yes/No | 0.0021 | 0.6% |
| Time to water |
< 30 min > 30 min Within Premises |
0.0020 | 0.6% |
| Child’s Residence | Urban/Rural | 0.0017 | 0.5% |
Dominance values represent the decrease in McFadden’s R² when each predictor is removed from the full model. Relative importance values indicate the percentage of total explainable variance attributed to each predictor.
Table 4.
Comparative variable importance scores across machine learning models for diarrhea status prediction
| Variables | Categories | Decision Tree | GBM | Random Forest |
|---|---|---|---|---|
| Region | North-central, North-East, North-West, South-East, South-South, South-West | 87.5164 | 34.1589 | 22.0492 |
| Child’s Age | 48–59, 36–47, 24–35, 12–23, 6–11, ≤ 5 | 82.8259 | 31.5791 | 12.9126 |
| Mothers’ Level of Education | No Education, Primary (Complete), Primary (Incomplete), Secondary (Complete), Secondary (Incomplete), Higher Education | 52.9630 | 4.8390 | 7.7615 |
| Urban Rural Wealth Index | Very Rich, Rich, Average, Poor, Very Poor | 48.6057 | 3.4010 | 10.8205 |
| Wealth Index | Very Rich, Rich, Average, Poor, Very Poor | 47.1576 | 6.4892 | 9.0183 |
| Watches TV | No, Once a week, >Once a week | 42.7294 | 2.7333 | 9.7767 |
| Mothers’ Age | ≥ 35, 30–34, 25–29, 20–24, 15–19 | 29.4487 | 4.3308 | 1.0149 |
| Disposal of youngest child faeces | Safe, Unsafe | 24.7784 | 0.9547 | 2.2270 |
| Time to water |
< 30 min > 30 min Within Premises |
12.1433 | 2.1345 | 4.0544 |
| Under 5 in household | ≥ 3/≤2 | 11.4601 | 0.6854 | 6.1031 |
| Listens to Radio | No, Once a week, >Once a week | 10.0461 | 3.2112 | 9.9760 |
| Type of Toilet Facility | Improved/Unimproved | 6.2681 | 2.1806 | 7.8027 |
| Shared Toilet Facility | Yes/No | 5.7941 | 1.0009 | 2.6482 |
| Total in Household | ≥ 7 Persons/≤6 persons | 5.0512 | 0.6003 | 2.3404 |
| Access to Electricity | Yes/No | 3.5445 | 0.2104 | 7.2360 |
| Source of Drinking Water | Improved/Unimproved | 2.8534 | 0.1319 | 1.1694 |
| Child’s Sex | Male/Female | 2.6656 | 0.3020 | −0.5048 |
| Child’s Residence | Urban/Rural | 2.1708 | 0.1847 | 4.0476 |
| Internet Use Last month |
Almost everyday At least once a week < than once a week |
1.4992 | 0.5799 | 7.7377 |
| Ever used the Internet | Yes/No | 1.2174 | 0.2919 | 6.8290 |
Variable importance scores represent the relative contribution of each predictor. Higher scores indicate greater importance.
Fig. 5.
This figure illustrates the top five most important features for predicting diarrhea, as identified by four distinct machine learning models: Decision Tree, Gradient Boosting Machine (GBM), Logistic Regression, and Random Forest. The horizontal bar chart compares the importance scores (ranging from 0 to 100) for key predictors such as ‘RegionNorth East’, ‘Level of Education-Not Educated’, and various age categories of the child (‘Oldest’, ‘Young’, ‘Older’, ‘Younger’), highlighting which variables each algorithm deemed most influential. The comparative analysis reveals both consensus and divergence across models regarding critical risk factors, providing valuable insight into the underlying drivers of the prediction and the stability of feature importance interpretations. The chart aids in understanding which predictors are most influential for each modeling approach
Discussion
This comprehensive study investigated the multivariate determinants of childhood diarrhea through a novel dual approach of traditional epidemiology and machine learning. Our results illustrate how a mix of socioeconomic, environmental and behavioral determinants interact to ultimately influence diarrhea disease risk in children under-5 years of age, with strong agreement on important predictors across methods. In addition, it emphasizes the strengths of ML models and multilevel regression being complementary to each other while presenting useful feedback for predictive analytics both in low-resource contexts and public health policy.
Important observations in context
The determinants found in this study are consistent with and complement the current evidence base for childhood diarrhea. The most robust finding in all the models was the dramatically strong effect of child’s age which is consistent with immunological and behavioral mechanisms. The very high risk of diarrhea in children aged 6–23 months (OR = 2.48, CI = 2.18–2.83, P-value < 0.001) are typical epidemiological observations, ideally coinciding with weaning age. During this critical window, passive immunity from breast milk is ended, creating increased exposure to contaminated complementary food and environmental pathogens, including play activities that are more typical of increase infection risk [29]. The carry-over of elevated risk through to the 36–47-month age group (OR = 1.48, CI = 1.29–1.69, p-value < 0.001) as revealed by our MICE analysis, is a significant nuance. In addition, DT did a better job at capturing this nonlinear relationship (node purity increase = 82.83) highlighting the potential of machine learning to model complex life-course effects. This implies that once the period of weaning has been reached, repeated exposures in the environment and perhaps inadequate personal hygiene practices in children as they become more mobile are a massive risk. This observation highlights the importance of interventions beyond the first two years of life.
The protective correlation for maternal education (higher education OR = 0.77) is consistent with the widespread evidence demonstrating an association between caregiver education and child health [6]. Educated women are health literate, tend to engage in improved sanitation and hygiene, are aware of early disease symptoms and have greater ability to access proper care, and hence reduce their children’s risk for diarrheal disease [6]. Our dominance analysis confirmed this strength, and ranked this as one of the strongest predictors in our variables. This is in favor of the argument that educating girls is not only a social imperative but also a smart anchor investment in long-term public health. Nevertheless, our GBM analysis showed that this relationship was characterized by a nonlinear threshold pattern where substantial decreases were observed only after secondary education. This nuance highlights the importance of continuing educational attainment beyond basic literacy campaigns in order to meaningfully influence health behaviors, findings that are similar in sub-Saharan Africa [30].
The paradoxical association between internet use and greater odds of diarrhea (OR = 1.57) should be interpreted with caution. Although initially counterintuitive, this probably is due to residual confounding by urban place of residence and/or socioeconomic status, an effect observed in the analogous country-level digital health survey of Kenya [12]. In addition, while “ever used” was linked with increased risk (presumably substituting for increased SES and urban residence, where other environmental risks might predominate), frequent use was a safeguard. This might be interpreted as suggesting that digital proficiency and regular exposure to information, possibly including health information, have a protective effect, an observation worthy of focus for future investigation into the value of digital health proficiency. Decision Tree analysis helped unpack this relationship, showing that internet access was only predictive of risk among wealthier households, potentially driven by different patterns of use (entertainment versus health information seeking). The reported correlation between higher radio exposure and increased diarrhea risk is counterintuitive but might also be confounded. Radio ownership and listenership can be a proxy measure of rural living or lower socioeconomic status with little or no access to clean water and sanitation resources. It could also reflect the success of health programming on the radio, generating and promoting higher number of reported of cases.
The urban-rural wealth index was a superior and more robust predictor than the traditional wealth index, a key methodological insight. The contextual variable, which informs us on wealth in terms of type of accommodation, had a graded protective effect. This indicates that the overall value of items is less important than their relative value in the context (urban or rural). A “Rich” rural family can still be excluded from piped water and sewerage services available to an “Average” urban family. This result provides a strong argument for applying such contextual indicators in subsequent analysis of health status in highly stratified groups.
Finally, the protective effect of large household size (≥ 7 members) challenges the naive assumption that overcrowding will always increase risk. While crowding can allow for the spread of pathogens, more family members could also be accompanied by additional caregiver funtions to help cover the costs of taking care of children, developing cleaner lifestyles, or having older children who are able to retrieve cleaner water, which would negate potential risk. This widespread intricacy between structure and health effect has been noted elsewhere and should be pursued further qualitatively [31].
Methodological insights
This study enables us to make ultimate comparison between the analytical models. The multilevel logistic regression model gave us causally interpretable, adjusted measures of association (odds ratios) with confidence intervals, enabling us to estimate the exact effect of every determinant while holding other determinants constant, and adjusting for hierarchical data structure. Its strengths lie in inference and hypothesis testing. To address the challenge with “variable importance” similar to ML predictive hierarchy, we used dominance analysis, a post-estimation method that computes the relative contribution of every predictor to the explanatory power of the model. This successfully bridge the gap, ensuring disposal of faeces and child’s age were the most influential predictors in the regression model, a result consistent with ML results. This integration provides analytical strength and a combined insight superior to individual method alone.
In contrast, the machine learning algorithms performed prediction and feature ranking to a high degree without imposing any a priori assumptions. Their strongest suit was detection of intricate, non-linear interactions and relationships in the data. The high, stable ranking of “Region” across all ML models, particularly the Decision Tree, points to an important finding: geography is a deep proxy for a constellation of unmeasured contextual determinants such as conflict, governance, climate, healthcare infrastructure, and culture, having a high influence on diarrhea risk. While the regression model accounted for region as a cluster variable, the ML models measured its overwhelming predictive ability. However, the “black box” characteristic of sophisticated ML models such as GBM and Random Forest prevents us from unpacking the directionality or mechanism of such effects. The Decision Tree’s (AUC = 0.5) poor performance shows us an important lesson in the dangers of overfitting and the need to apply ensemble methods for proper prediction. This probably results from their function as effect modifiers (a relationship which is better adapted by an algorithmic approach). Such divergences underscore the importance of methodological triangulation in the context of complex epidemiological research [15].
The most important part of this work was sensitivity analysis between CCA and MICE. The large differences witnessed for variables such as wealth indices and faeces disposal confirm that missing data was not randomly (MNAR). CCA estimates, presumably downward-biased as a result of systematic loss to follow-up of the poorer and more deprived families (who are themselves likely to be less likely to have full data), and underestimated the protective effect of proper waste disposal practices. This illustrates that the use of robust imputation techniques such as MICE is not merely a statistical nicety but is necessary for generating unbiased and generalizable estimates in complex survey data with high missingness. Dependence on CCA alone would have resulted in a misinterpretation of the part played by wealth and sanitation.
Policy implications
Several policy recommendations emerge. The differences among the various regions are so vast that one single strategy is not viable. Northern regions may require emergency WASH infrastructure in addition to conflict-sensitive delivery, whereas southern regions may concentrate on behaviour change enforcement, reiterating region-targeted strategies. Context-specific poverty alleviation in strengthening the urban-rural wealth index suggests that poverty alleviation and WASH interventions need to be hyper-localized. Policies need to aim not only to enhance wealth but to work on the particular WASH infrastructure deficiencies characteristic of a specific context (e.g., rural water boreholes in contrast to urban sewerage systems). Similarly, the reverse non-linearity in education effects means greater priority for programmes emphasizing secondary school education retention versus basic literacy, and this holds true particularly for girls, therefore making the incorporation of WASH in national curricula a potential enhancer towards the magnitude of its health effects [17]. Investing in root drivers as a causal pathway from education of girls to decreasing the prevalence of diarrhea is an effective case for policies favoring girls in school, especially to secondary completion. It is a long-term, multi-sectoral investment with high health return.
Other interventions include public health messaging which must be directly targeted at infant’s mother (6–23 months old) to motivate safe weaning, handwashing before food handling, and safe food storage. Extended risk through 4 years indicates that education on hygiene has to be a cornerstone of outreach during early childhood. More so, availing the use of technology for public health promotion potentiates preventive impact of frequent exposure to the internet and offers a portal for electronic public health interventions. Creating and distributing mobile health (mHealth) applications that include diarrhea prevention, treatment advice, and directories of health facilities nearby could be a very cost-effective strategy. Lastly, data-driven resource allocation for both Logistic Regression and GBM were found to have good predictive accuracy and can be utilized as risk stratification tools. The models can be utilized by health ministries and NGOs to screen and prioritize high-risk clusters (such as specific communities or districts that have a high proportion of predicted risk factors) for intervention, thus optimizing the utilization of scarce resources.
Limitations and future directions
Some limitations merit consideration. Causal inference is not possible owing to the cross-sectional design of the study, but we employed a multi-method design to alleviate temporal ambiguity to a large extent, with sensitivity analysis showing minimal differences among key variables of respondents. Second, although MICE can deal satisfactorily with missing data, its accuracy assumes that data are Missing At Random (MAR), which may not always be completely the case. Unmeasured confounding can still occur, for example, we were not able to adjust for precise water quality measurements or precise hygiene measures. However, emerging opportunities for future investigations include integration of climate data with our spatial predictors to predict climate-diarrhea associations, applying natural language processing for radio/TV content analysis for specific messaging effects and the deployment of mobile health tools while employing our machine learning models for real-time disease outbreak prediction. The use of longitudinal designs with environmental sampling may be helpful to buttress causal assertions in future studies.
Conclusion
This study advances childhood diarrhea research by comparing traditional and machine learning methods in a novel fashion. Through identifying a combination of established and new risk factors, including regional differences and effects of time poverty, we aim to develop a more nuanced evidence-based strategy for intervention development. The persistent significance of upstream determinants such as education, poverty reduction and context-specific public health programming, call for intersectoral action beyond the traditional health sector responses. As the global climate and urbanization change alters disease landscapes, the methodological framework introduced here provides a flexible tool for continued surveillance and targeted prevention.
Supplementary Information
Acknowledgements
The first author is grateful to Duy Tan University for providing a conducive environment for this study.
Authors’ contributions
Author contributionsConceptualization: JOA, LTD, KRI and SYMS; Data collation and analysis: JOA, TSA and SYMS; Writing- JOA, KRI, LTD and TSA. Visualization: SYMS, LTD and JOA; Review, editing, and final draft JOA.
Funding
Funding for this study was solely by the authors.
Data availability
Data extracted and used in this study are available on request.
Declarations
Ethics approval and consent to participate
The study employed the use of data extraction (secondary analysis) of an already collected dataset, as the survey personnel obtained ethical approval from the National Ethic Committee of the Federal Ministry of Health Abuja, Nigeria and ICF International, Rockville, MD, USA. Informed consent was obtained from study participants prior to participation in the survey. Permission to use and analyze the data set was obtained by registering the study on the Demographic and Health Survey (DHS) website.
Consent for publication
NA.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.World Health Organization (WHO). Diarrhoeal disease [Internet]. 2023 [cited 2023]. Available from: https://www.who.int/news-room/fact-sheets/detail/diarrhoeal-disease.
- 2.UNICEF. One is too many: Ending child deaths from pneumonia and diarrhoea [Internet]. 2022 [cited 2022]. Available from: https://data.unicef.org/resources/one-is-too-many/.
- 3.Troeger C, Blacker B, Khalil IA, Rao PC, Cao J, Zimsen SRM, et al. Estimates of the global, regional, and National morbidity, mortality, and aetiologies of diarrhoea in 195 countries: a systematic analysis for the global burden of disease study 2016. Lancet Infect Dis. 2018;18(11):1211–28. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Kotloff KL, Nataro JP, Blackwelder WC, Nasrin D, Farag TH, Panchalingam S, et al. Burden and aetiology of diarrhoeal disease in infants and young children in developing countries (the global enteric multicenter study, GEMS): a prospective, case-control study. Lancet. 2013;382(9888):209–22. [DOI] [PubMed] [Google Scholar]
- 5.Prüss-Ustün A, Wolf J, Bartram J, Clasen T, Cumming O, Freeman MC, et al. Burden of disease from inadequate water, sanitation and hygiene for selected adverse health outcomes: an updated analysis with a focus on low- and middle-income countries. Int J Hyg Environ Health. 2019;222(5):765–77. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Fink G, Günther I, Hill K. The effect of water and sanitation on child health: evidence from the demographic and health surveys 1986–2007. Int J Epidemiol. 2011;40(5):1196–204. [DOI] [PubMed] [Google Scholar]
- 7.Cumming O, Cairncross S. Can water, sanitation and hygiene help eliminate stunting? Current evidence and policy implications. Matern Child Nutr. 2016;12(Suppl 1):91–105. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Bado AR, Susuman AS, Nebie EI. Trends and risk factors for childhood diarrhea in sub-Saharan Africa (1990–2013): assessing the neighborhood inequalities. Glob Health Action. 2016;9:30166. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Bauza V, Guest JS. The effect of young children’s faeces disposal practices on child health: a systematic review. PLoS ONE. 2017;12(4):e0177066. [DOI] [PubMed] [Google Scholar]
- 10.Ngure FM, Reid BM, Humphrey JH, Mbuya MN, Pelto G, Stoltzfus RJ. Water, sanitation, and hygiene (WASH), environmental enteropathy, nutrition, and early child development: making the links. Ann N Y Acad Sci. 2014;1308:118–28. [DOI] [PubMed] [Google Scholar]
- 11.Livingston J, Ciancio B, Karlyn A. The role of digital health in mitigating the impact of COVID-19 on essential health services. BMJ Glob Health. 2021;6(Suppl 5):e004348. [Google Scholar]
- 12.Oluwasanu MM, Oladunni O, Adebayo SB. Digital disparities and maternal health communication in Nigeria: implications for child survival interventions. J Glob Health Rep. 2020;4:e2020063. [Google Scholar]
- 13.Ogunniyi A, Oladunni T, Abiodun T. Machine learning approaches for predicting under-five mortality: a comparative analysis of data from Nigeria. BMC Med Inform Decis Mak. 2021;21:45.33557818 [Google Scholar]
- 14.Hosmer DW, Lemeshow S, Sturdivant RX. Applied logistic regression. 3rd ed. Wiley. 2013.
- 15.Boulesteix AL, Janitza S, Kruppa J, König IR. Overview of random forest methodology and practical guidance with emphasis on computational biology and bioinformatics. WIREs Data Min Knowl Discov. 2012;2(6):493–507. [Google Scholar]
- 16.Lundberg SM, Erion G, Chen H, DeGrave A, Prutkin JM, Nair B, et al. From local explanations to global understanding with explainable AI for trees. Nat Mach Intell. 2020;2:56–67. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Freeman MC, Garn JV, Sclar GD, Boisson S, Medlicott K, Alexander KT, et al. The impact of sanitation on infectious disease and nutritional status: a systematic review and meta-analysis. Int J Hyg Environ Health. 2017;220(6):928–49. [DOI] [PubMed] [Google Scholar]
- 18.Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25(1):44–56. [DOI] [PubMed] [Google Scholar]
- 19.World Health Organization. Diarrhoea treatment guidelines including new recommendations for the use of ORS and zinc supplementation for clinic-based healthcare workers. Geneva: WHO; 2017. [Google Scholar]
- 20.Sterne JAC, White IR, Carlin JB, Spratt M, Royston P, Kenward MG, et al. Multiple imputation for missing data in epidemiological and clinical research: potential and pitfalls. BMJ. 2009;338:b2393. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Filmer D, Pritchett LH. Estimating wealth effects without expenditure data - or tears: an application to educational enrollments in states of India. Demography. 2001;38(1):115–32. [DOI] [PubMed] [Google Scholar]
- 22.Breiman L. Random forests. Mach Learn. 2001;45(1):5–32. [Google Scholar]
- 23.Friedman JH. Greedy function approximation: a gradient boosting machine. Ann Stat. 2001;29(5):1189–232. [Google Scholar]
- 24.Quinlan JR. Induction of decision trees. Mach Learn. 1986;1(1):81–106. [Google Scholar]
- 25.Steyerberg EW, Vickers AJ, Cook NR, Gerds T, Gonen M, Obuchowski N, et al. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology. 2010;21(1):128–38. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Kuhn M, Johnson K. Applied predictive modeling. New York: Springer; 2013. [Google Scholar]
- 27.National Population Commission (NPC), ICF. Nigeria demographic and health survey 2018. Abuja, Nigeria: NPC and ICF; 2019. [Google Scholar]
- 28.Niculescu-Mizil A, Caruana R. Predicting good probabilities with supervised learning. Proc ICML. 2005:625–32.
- 29.Ogbo FA, Ogeleka P, Awosemo AO. Trends and determinants of complementary feeding practices in Tanzania, 2004–2016. Trop Med Health. 2018;46:40. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Smith-Greenaway E. Maternal reading skills and child mortality in Nigeria: a reassessment of why education matters. Demography. 2013;50(5):1551–61. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Sclar T, Menezes NB, Carneiro HO, Andrade MV, Reis IA, Alvares-Teodoro J, et al. The effect of a large-scale water, sanitation and hygiene intervention in Brazil on intestinal protozoa in children under five: a population-based, cross-sectional study. Lancet Glob Health. 2018;6(Suppl 1):S16. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Data extracted and used in this study are available on request.





