Abstract
Anemia remains a major global public health problem, affecting an estimated 1.93 billion people worldwide. Pregnant women in sub-Saharan Africa (SSA) are disproportionately affected, placing them at greater risk of adverse maternal and neonatal outcomes. This study used machine learning to predict anemia and identify key predictors among pregnant women in SSA. We analyzed the most recent Demographic and Health Survey (DHS) data from 24 SSA countries, including a weighted sample of 14,569 pregnant women. Data were prepared in SPSS version 27 and analyzed in Python version 3.12. An Extreme Gradient Boosting (XGBoost) classifier was developed to predict anemia, while Shapley Additive Explanations (SHAP) identified influential predictors and improved model interpretability. Model performance was assessed using accuracy, recall, F1 score, and area under the receiver operating characteristic curve (AUC). The XGBoost model achieved 87% accuracy, 85% recall, a 73% F1 score, and an AUC of 95%, indicating excellent predictive performance. SHAP analysis identified unimproved water sources, lack of antenatal care visits, absence of mobile phone ownership, older maternal age, financial barriers to healthcare, cigarette smoking, khat chewing, limited media exposure, urban residence, home delivery, and use of unimproved cooking fuel as the most influential predictors of anemia. These findings demonstrate that machine learning can accurately identify pregnant women at high risk of anemia using routinely collected sociodemographic, behavioral, and maternal health data. Integrating ML-based prediction models into antenatal care and digital maternal health platforms could support timely risk identification, personalized interventions, and targeted strategies to reduce maternal anemia across SSA.
Keywords: Machine learning, SHAP, Anemia, Pregnant women, DHS
Introduction
In public health, Machine Learning (ML) is essential for predictive analytics, health surveillance, and disease outbreak prediction. By analyzing population-level data, ML models can forecast the spread of infectious diseases, optimize resource allocation, and identify at-risk populations for targeted interventions 1 Machine Learning can handle large-scale, heterogeneous datasets and reveal underlying patterns that may not be seen using traditional approaches, it has the potential to revolutionize healthcare and public health. 2 Healthcare and public health systems can provide more effective, individualized, and fair care by utilizing ML, which will ultimately improve health outcomes. 3
Anemia is a condition characterized by a decrease in the number of red blood cells or a reduction in hemoglobin levels below the body’s physiological requirements. This leads to an impaired ability to transport oxygen effectively throughout the body. 4 Hemoglobin concentration is the main hematological method for assessing anemia, but the cut-off point can fluctuate due to multiple factors. 5 According to the World Health Organization (WHO), anemia in the case of pregnant women is diagnosed when the hemoglobin concentration falls below 11 g/dl. 6 Iron deficiency is often accountable for the primary cause of anemia in pregnant women, followed by folic acid and vitamin B-12 deficiencies and parasitic diseases. 7
Maternal anemia has afflicted nearly 32 million pregnant women worldwide. 8 In low- and middle-income countries, as high as fifty percent of pregnant women are diagnosed with this condition. 9 Almost half of anemia cases are observed in Africa. 5 This high rate of anemia in developing countries can be attributed to higher rates of dietary micronutrient deficiencies, such as iron and folic acid, as well as infections like malaria, HIV, and hookworm infestation, compared to developed countries. 10 Other factors contributing to anemia in developing countries include low socioeconomic status, rural residence, short birth intervals, delayed initiation of antenatal care, and grand multi-parity. 11
The prevalence of maternal anemia among pregnant women varies across the globe. In developed nations, anemia rates are relatively lower; for example, the rate is 18% in the United States and 20% in Australia. 12 However, in developing countries, anemia rates have increased significantly, with the highest prevalence in sub-Saharan Africa, ranging from 38.9% to 48.7%. 9 Specifically, the prevalence of anemia among pregnant women was 50.1% in Ethiopia, 53% in Sudan, and 71% in Guinea.13–15 Furthermore, cohort studies have demonstrated a direct relationship between maternal anemia and mortality, with each 0.1 g/dl increase in maternal hemoglobin associated with a 29% decrease in maternal mortality. 15 Maternal anemia can also slow down psychomotor development, impair cognitive performance, and reduce intelligence test scores in newborns, potentially affecting their long-term development. 16 Moreover, Women with anemia will be at higher risk of maternal morbidities, including miscarriages, antepartum hemorrhage, postpartum hemorrhage, preeclampsia, and delayed labor. 17
Several studies have utilized classical statistical methods to identify the factors influencing anemia status among pregnant women.18–31 Classical data analysis and associated factor identification have been accomplished through statistical models. 32 Statistical modeling cautious about uncertainty, it requires a lot of attention to confidence intervals and hypothesis tests. In contrast, ML modeling embraces uncertainty with little to no assumptions being made. 33 Classical statistics relying solely on assumption can limit the ability to discover new insights and information. ML algorithms are typically designed to make accurate predictions by learning from data rather making prior assumptions. 34 This approach enables algorithms to reveal hidden knowledge and patterns that may not be obvious based on prior assumptions. 35 Machine learning models also offer the capability to interact seamlessly with digital systems, enabling organizations to leverage insights from research to address practical problems and real-world challenges. By bridging the gap between theoretical research and practical applications, ML models drive innovations and advancements in the healthcare industry. 36 Therefore, this study aimed to address the evidence gap on anemia by developing a predictive model, and by identifying the most predictors among pregnant women in sub-Saharan African countries using machine learning algorithm.
Method
Study period and setting
This study used data from the Demographic and Health Surveys (DHS), which are nationally representative, cross-sectional household surveys done in more than 85 countries since 1984. The DHS gathers a wide range of objective and self-reported data, with a particular emphasis on indicators of fertility, reproductive health, mother and child health, mortality, nutrition, and adult health behaviors. The DHS’ key strengths include high response rates, wide national coverage, thorough interviewer training, uniform data collection techniques across nations, and consistent content across time. The study used data from recent surveys conducted in 24 SSA countries between 2016 and 2024 G.C. These countries were Benin (2017-18), Burkina Faso (2021), Burundi (2016-17), Cameroon (2018), Côte d'Ivoire (2021), Ethiopia (2016), Gabon (2019-21), Ghana (2022), Gambia (2018), Guinea (2018), Lesotho (2023-24), Liberia (2021-19), Madagascar (2021), Malawi (2015-16), Mali (2018), Mauritania (2019-21), Mozambique (2023), Nigeria (2021), Rwanda (2016), Sierra Leone (2019), South Africa (2016), Tanzania (2022), Uganda (2016), and Zambia (2018). For this analysis, the study focused on pregnant women by appending individual records from each country and identifying respondents who were pregnant at the time of data collection. The data used in the study is publicly available and can be accessed at https://dhsprogram.com/data/available-datasets.cfm. In DHS, the study participants were selected using a two-stage stratified sampling procedure. First the enumeration area was selected based on each country’s frame from the previous census performed. Second, households in each enumeration area were selected. After extracting the dataset, we included 14,569 weighted pregnant women form 24 SSA countries.
Study variables
The dependent variable in this study was anemia status. Anemia is classified as a categorical variable in each country’s DHS datasets: non-anemic, mild, moderate, and severe. To better suit the XGBoost classifier, we reclassified anemia as a binary variable. The mild, moderate, and severe anemia categories were merged and labeled as “anemic”, 1 whereas the non-anemic group was marked as “non-anemic” (0). This recategorization was necessary because the number of cases in the severe and moderate anemia categories was too small to provide reliable estimates and meaningful insights when modeled separately. By combining these categories, we guaranteed that the dependent variable had a more balanced distribution and improved the classifier’s capacity to effectively identify patterns associated with anemia. This technique is consistent with best practices for machine learning models, where adequate representation across classes is required for reliable prediction and performance evaluation.37–39 The independent variables for this study were selected based on existing literature on issues influencing anemia. The independent variables included source of drinking water, chewing, smoking, delivery by Cesarean Section (CS), having a bank account, Antenatal care (ANC) visit, husband’s occupation, type of fuel used, type of toilet facility, place of delivery, ability to get money for medical care, mobile phone ownership, place of residence, respondent’s occupation, media exposure, distance from a health facility, health facility visit, age, wealth index, internet use, educational status, and marital status.18–31
Data management and analysis
Individual-level datasets from the 24 countries DHS surveys were appended and weighted using SPSS version 27 and Microsoft Excel 2019. All further analysis, including data preprocessing, model development, and interpretation, was conducted using Python version 3.12, with the following principal libraries: Pandas (v2.2.0) and NumPy (v1.26.4) for data manipulation; Matplotlib (v3.8.2) and Seaborn (v0.13.2) for visualisation; Scikit-learn (v1.4.0) for preprocessing, model evaluation, and Recursive Feature Elimination; XGBoost (v2.0.3) for classifier training; imbalanced-learn (v0.12.0) for SMOTE oversampling; and SHAP (v0.44.0) for model interpretability. To account for the complex, multistage stratified cluster design of the DHS and ensure representative estimates, three survey design variables were incorporated: V005 (sampling weight), V021 (primary sampling unit/cluster), and V022 (stratification variable). In accordance with DHS guidelines, V005 was normalised by dividing by 1,000,000 and applied as a sample weight during model training, thereby adjusting for differential selection probabilities across strata. 40 The figure below summarizes the analytical workflow used in this study, from DHS data pooling through model interpretation using SHAP (Figure 1).
Figure 1.

Methodological flowchart for XGBoost-based prediction of anemia among pregnant women using pooled DHS data (24 SSA countries).
Data division
The study employed a hold-out validation strategy in which 80% of the dataset was allocated for model training and the remaining 20% was reserved for independent testing. The test set was held out entirely from all preprocessing and model development steps. In addition, stratified 10-fold cross-validation was applied within the training set to obtain a robust and stable estimate of model generalization performance. In this procedure, the training data were partitioned into 10 equal folds; the model was trained on nine folds and evaluated on the remaining fold iteratively until each fold had served as the validation set once. 41 Data splitting was performed prior to any preprocessing step that could to prevent information leakage, including imputation, SMOTE oversampling, encoding, and scaling. A fixed random seed (seed = 42) was used throughout to ensure full reproducibility.
Data refinement and feature processing
Missing values were handled using K-nearest neighbor (KNN) imputation with k = 5, applied within the training data to preserve data integrity and prevent information leakage. To address class imbalance in the outcome variable, where the majority of women were non-anemic, the Synthetic Minority Oversampling Technique (SMOTE) was applied exclusively to the training dataset to avoid biasing the evaluation process. Categorical variables were transformed into numerical format using one-hot encoding. 42
Dimensionality reduction
Recursive Feature Elimination (RFE) was employed to identify and rank the most predictive features. RFE was performed exclusively on the training dataset after the train/test split. RFE iteratively removes the least important features and refits the model, thereby identifying the subset of predictors that maximizes model performance. A Random Forest classifier was selected as the base estimator for RFE owing to its robustness in capturing non-linear relationships, handling feature interactions, and generating reliable importance rankings.
RFE identified “source of drinking water,” “khat chewing,” and “cigarette smoking” as among the most influential predictors, while “marital status,” “educational status,” and “internet use” were found to have low predictive importance and were excluded, reducing the feature set from 22 to 19 predictors. This approach improved computational efficiency while preserving model interpretability and predictive performance.
XGBoost classifier
Extreme Gradient Boosting (XGBoost) is an ensemble learning technique that belongs to the gradient boosting framework. It constructs prediction models by sequentially combining multiple decision trees (base learners), with each successive tree correcting the errors of its predecessors. 43 XGBoost is recognized for its computational efficiency, scalability, and strong predictive performance on structured tabular data. It incorporates built-in L1 (Lasso) and L2 (Ridge) regularization mechanisms to reduce overfitting and handles missing values natively during training. 44 In addition, XGBoost can effectively capture complex, non-linear relationships and high-order interactions among predictors, which are common in population health and demographic datasets. This capability makes it particularly suitable for modeling multifactorial public health conditions such as anemia, where biological, behavioral, socioeconomic, and environmental predictors may interact in complex ways. 45 The selection of XGBoost in this study was further justified by the nature of the Demographic and Health Survey data, which consist of high-dimensional, heterogeneous variables with potentially complex dependencies. XGBoost has demonstrated robust performance in epidemiological and healthcare prediction studies involving tabular and imbalanced datasets, making it an appropriate choice for anemia prediction in a large multi-country population. 46 Furthermore, the model integrates well with interpretable artificial intelligence approaches such as Shapley Additive Explanations (SHAP), enabling transparent identification and ranking of influential predictors. This interpretability was essential for the present study, as the objective extended beyond prediction to understanding the relative contribution of predictors of anemia among pregnant women. Rather than benchmarking multiple algorithms, this study prioritized the application of a theoretically justified, computationally efficient, and interpretable model suitable for large-scale maternal health data. 47 The decision trees in XGBoost are trained using the following objective function:
| (1) |
where l ( , ) is the training loss function measuring the difference between predicted and actual values, and Ω(fk) is the regularization term that penalizes model complexity to prevent overfitting.
Hyperparameter optimization
Although XGBoost default parameters provide a reasonable starting point, fine-tuning is necessary to optimize performance for specific datasets and reduce the risk of overfitting or underfitting. Hyperparameter optimization was conducted using Bayesian optimization, implemented via the scikit-optimize library. Bayesian optimization constructs a probabilistic surrogate model of the objective function and uses this to guide the search toward promising parameter configurations, balancing exploration of the parameter space and exploitation of known good regions. This approach is substantially more efficient than grid search (exhaustive) or random search (non-adaptive) for the large, continuous parameter space of XGBoost. Optimization was run for 50 iterations on the training dataset using 10-fold cross-validation as the performance estimator, with AUC as the primary optimization criterion. Key parameters optimized included: learning_rate (0.01-0.30), n_estimators (50-300), max_depth (3-10), min_child_weight (1-10), subsample (0.5-1.0), colsample_bytree (0.5-1.0), gamma (0-1), reg_alpha (0-1), and reg_lambda (0-1). The decision threshold for binary classification was set at 0.5, consistent with balanced class probabilities after SMOTE oversampling. Sensitivity analyses conducted at thresholds of 0.40 and 0.60 produced comparable classification patterns, supporting the appropriateness of the 0.5 threshold in this context.
Model evaluation measures
The approach’s output was evaluated using a confusion matrix, which records the number of correctly and incorrectly determined occurrences. The classifiers’ efficiency evaluated based on accuracy, sensitivity, precision, f1 score and AUC.
Accuracy
Accuracy is the percentage of true events among the total number of cases tested. In this study, it was used to determine model efficacy. 48
| (2) |
Sensitivity/recall
Sensitivity is a test that counts the number of correctly predicted positive events among all positive events. This provides us with the number of predicted positives relative to the total number of positive classes. This is referred to as recall, and it can be computed using the provided formula. 49
| (3) |
Precision
Precision is determined by dividing the total number of positive events predicted by the classifier by the number of correct events. Another name for this is positive predictive value. This study utilized the following formula, which was derived from the confusion matrix, to verify the model output. 50
| (4) |
F1-score
The F1-score is a harmonic mean of precision and sensitivity (recall), providing a balanced measure of model performance, particularly when dealing with imbalanced datasets. In this study, the F1-score was calculated using the following formula derived from the confusion matrix. 51
| (5) |
Model interpretation
Although machine learning models’ growing complexity has greatly increased prediction accuracy, it has also made them more challenging to understand. We are unable to completely explain how models use characteristics and samples to generate judgments because of this lack of interpretability. 52 This limitation has therefore hindered the use of machine learning techniques in some fields. Many applied researchers continue to favor simpler, easier-to-understand models, such logistic and linear regression, to solve this. A number of interpretation algorithms have been put forth in an effort to address the interpretability issue. For example, SHAP 53 and Local Interpretable Model-Agnostic Explanations (LIME) 54 are frequently used to understand individual predictions. Similar to this, Accumulated Local Effects Plots 55 and Partial Dependence Plots 56 aid in describing patterns between variables and results. In order to mimic complex models and offer insights into their decision-making processes, some researches have also used more straightforward, interpretable models, including decision trees. Shapley values, which come from coalitional game theory, are the foundation of the SHA strategy. The contribution of each feature to a model’s prediction for a particular instance is measured by Shapley values in machine earning. This contribution is determined by subtracting a feature’s impact from its average impact over all instances. Shapley values, which provide a reliable and consistent method of measuring feature relevance in complex models, are derived by weighting and adding up the marginal contributions of each possible combination of feature values. 57 In this study the link between the predictors and the outcome was assessed using the SHAP feature selection approach. SHAP was chosen because it provides clear and interpretable insights into how each feature contributes to model decisions, which is crucial in healthcare applications where interpretability is important.
Results
Socio-demographic characteristics
A total of 14,569 respondents were included in the analysis. The prevalence of anemia among pregnant women was 42%. Of these, 9,962 (68.4%) resided in rural areas, while 4,607 (31.6%) lived in urban settings. Regarding educational status, 5,335 participants (36.6%) had no formal education, and 4,663 (32.0%) had only primary education. Marital status data revealed that 12,931 respondents (88.8%) were married, and 1,184 (8.1%) were single. In terms of mobile phone ownership, 7,500 participants (51.5%) lacked access to a mobile phone. Additionally, only 2,253 respondents (16.8%) reported using the internet. Regarding healthcare, 13,422 participants (92.1%) had attended at least one antenatal care (ANC) visit, while 11,638 deliveries (79.9%) took place in a health facility. A significant financial barrier to healthcare was reported by 7,273 respondents (49.9%). Notably, 13,276 participants (91.1%) did not have a bank account (Table 1).
Table 1.
Socio-demographic characteristics of Study participants in 24 selected Sub-Saharan Africa Countries by using recent DHS dataset 2016-2024.
| Variable | Category | Frequency | Percentage (%) |
|---|---|---|---|
| Age | 15-19 | 2141 | 14.7 |
| 20-24 | 3836 | 26.3 | |
| 25-29 | 3560 | 24.4 | |
| 30-34 | 2712 | 18.6 | |
| 35-39 | 1678 | 11.5 | |
| 40-44 | 538 | 3.7 | |
| 45-49 | 104 | 0.7 | |
| Residence | Urban | 4607 | 31.6 |
| Rural | 9962 | 68.4 | |
| Educational Status | No Education | 5335 | 36.6 |
| Primary | 4663 | 32.0 | |
| Secondary | 4015 | 27.6 | |
| Higher | 556 | 3.8 | |
| Marital Status | Single | 1184 | 8.1 |
| Married | 12931 | 88.8 | |
| Widowed | 47 | 0.3 | |
| Divorced | 407 | 2.8 | |
| Mobile Phone Ownership | Yes | 7069 | 48.5 |
| No | 7500 | 51.5 | |
| Internet Use | Yes | 2253 | 16.8 |
| No | 12116 | 83.2 | |
| Cigarette Smoking | No | 14421 | 99.0 |
| Yes | 148 | 1.0 | |
| ANC | No | 1147 | 7.9 |
| Yes | 13422 | 92.1 | |
| Health Facility Visit in Last 12 Months | No | 4472 | 30.6 |
| Yes | 10097 | 69.3 | |
| Place of Delivery | Home | 2931 | 20.1 |
| Facility | 11638 | 79.9 | |
| Has Bank Account | No | 13276 | 91.1 |
| Yes | 1293 | 8.9 | |
| Getting Money Needed for Treatment | Big problem | 7273 | 49.9 |
| Not a big problem | 7296 | 50.1 |
Table 1 Socio-demographic and Related Characteristics of Study Participants.
Machine learning analysis
In this study, the XGBoost model’s effectiveness in predicting anemia among pregnant women was examined using two approaches, 10-fold cross-validation and the hold-out technique (which used 80% of the dataset for training and 20% for testing). Following hyperparameter optimization, the XGBoost model achieved strong predictive performance with an accuracy of 87%, precision of 65%, recall of 85%, F1 score of 73%, and AUC of 95%. The high recall demonstrates the model’s effectiveness in identifying true anemia cases, minimizing missed diagnoses, a critical feature for clinical screening applications. The moderate precision reflects some false positives, while the excellent AUC confirms the model’s reliable discriminative ability between anemic and non-anemic pregnant women. Model performance was further validated through 10-fold cross-validation, confirming stability and generalizability across data subsets (Figure 2). The 10-fold cross-validation results show that the model consistently performed well across all folds, with accuracy ranging from 0.85 to 0.90, suggesting a high overall correctness in predictions. The Recall score averaged about 0.70, demonstrating the model’s ability to detect the vast majority of genuine positive instances, which is critical for reducing missed anemia diagnoses. However, Precision was quite low (about 0.55), indicating that the model tended to overestimate positive cases. The F1 Score, which balances Precision and Recall, stayed steady between 0.60 and 0.65 over all folds, illustrating the model’s balanced performance in managing false positives and false negatives. It was particularly successful in detecting anemia in pregnant women, with a high Recall rate, which is critical for public health to avoid missed diagnoses (Figure 3).
Figure 2.

XGBoost model performance for predicting anemia status and identifying its predictors among pregnant women in 24 sub-Saharan Africa: Evidence form dataset DHS 2016-2024 G.C.
Figure 3.

Crossvalidation metrics for classifier to predict anemia status and identifying its predictors among pregnant women in 24 sub-Saharan Africa: Evidence from dataset DHS 2016-2024 G.C.
Model optimization
The model improves significantly in most major performance indicators after hyperparameter adjustment with Bayesian Optimization. Initially, the model had 84% accuracy, 51% precision, 80% recall, and a 63% F1 score. Following adjustment, the accuracy rises to 87%, indicating improved categorization and a more accurate model overall. Precision improves significantly, going from 51% to 65%, owing to better management of false positives via adjusted parameters such as min_child_weight, learning_rate, and gamma. Recall also rises from 80% to 85%, demonstrating that the model is becoming more adept at detecting positive occurrences. The F1 score improves from 63% to 73%, demonstrating a better balance between precision and recall, which is critical for tasks requiring a trade-off between minimizing false positives and false negatives. , The ROC AUC incease from 88% to 95%. indicate that the model’s discriminative ability has been slightly enhanced in favor of better precision and recall, which can be a common trade-off during hyperparameter tuning. Overall, hyperparameter tuning has led to a more robust model, enhancing precision, recall, and F1 score, making it better at correctly identifying and classifying instances (Figure 4, Table 2).
Figure 4.

Model performance before and after tuning XGB classifier for predicting anemia status and identifying its associated factors among pregnant women in 24 sub-Saharan Africa: Evidence from dataset DHS 2016-2024 G.C.
Table 2.
Default and tunned parameter value for XGBoost model to predict anemia status among pregnant women in Sub-Saharan Africa.
| Parameter | Default value | Tuned value | Rationale |
|---|---|---|---|
| learning_rate | 0.3 | 0.09 | Small learning rates are typically more robust, but slower |
| max_depth | 6 | 5 | Deeper trees capture more complex interactions but risk overfitting |
| min_child_weight | 1 | 9 | Larger values prevent overfitting by requiring more samples per leaf |
| Subsample | 1 | 0.8 | Controls overfitting by randomly sampling training data |
| colsample_bytree | 1 | 0.5 | Controls feature selection per tree, helps reduce overfitting |
| Gamma | 0 | 0.5 | Controls complexity; higher values prevent overly complex models |
| n_estimators | 100 | 200 | Controls the number of boosting rounds |
| reg_alpha | 0 | 0.8 | L1 regularization to prevent overfitting |
| reg_lambda | 1 | 0.5 | L2 regularization to prevent overfitting |
SHAP analysis
The SHAP bar plot rank features by average influence (mean |SHAP value|) on model predictions. “Source of drinking water” was the most influential predictor. “ANC visit” ranked next, reflecting prenatal care’s role in early detection and prevention through supplementation and education. “Mobile phone ownership” was also among the top predictors, likely linked to better access to health information and appointment reminders. “Age” ranked highly too, consistent with age-related physiological susceptibility to anemia in pregnancy. Economic and lifestyle features “getting money for medical care,” “smoking,” and “chewing khat” showed moderate relevance, pointing to links between financial constraints, health behaviors, and anemia risk. Lower-ranked features, including “place of delivery,” “health facility visit,” and “type of toilet facility,” contributed comparatively little, indicating a smaller role relative to top predictors in this population and context (Figure 5).
Figure 5.

SHAP global importance plot of XGBoost classifier model for identifying its associated factors of anemia during pregnancy in selected 24 Sub-Saharan African using recent DHS dataset from 2016-2024G.C.
The SHAP summary plot ranks predictors of anemia in pregnant women by importance, with SHAP values showing direction and magnitude of association: positive values relate to higher anemia risk, negative values to lower risk. Unsafe drinking water source was the strongest correlate, linked with higher anemia rates. Fewer or no ANC visits were associated with higher risk, while more visits correlated with lower risk. Mobile phone ownership was associated with lower risk, and media exposure showed a similar protective association. Smoking and khat-chewing were linked with higher risk. Rural residence and greater distance from health facilities were associated with higher anemia rates, reflecting geographic disparities. Solid fuel use and non-institutional delivery were associated with higher anemia frequency. Bank account ownership and health facility visits were associated with lower risk, pointing to links between financial inclusion, healthcare access, and anemia outcomes (Figure 6).
Figure 6.

Beeswarm plot, ranked by mean absolute SHAP value generated by XGBoost model for identifying predictors of anemia during pregnancy in selected 24 Sub-Saharan African using recent DHS dataset from 2016-2024 G.C.
Discussion
The aim of this study was to develop ML model, to uncover major factors of anemia among pregnant women in 24 selected Sub-Saharan African countries. XGBoost classifier was trained on balanced training data using train test split and 10-fold cross-validation. Classification accuracy, AUC, precision, recall, and F1 score were used to evaluate the performance the classifier. XGBoost classifier achieved an accuracy of 87%, AUC of 95%, precision of 65%, recall of 85% and F1 score of 63%. To date, no studies have been conducted utilizing ML algorithms to anemia status among pregnant women in SSA level. Compared to related studies, a study conducted in Ethiopia reported that the Random Forest classifier achieved superior performance, with an accuracy of 97%, precision of 93%, recall of 93%, and an F1-score of 93%. The difference in performance could be attributed to variations in study objectives, methodological approaches, and the homogeneity of the datasets used. The Ethiopian study might have benefited from focusing on a more uniform population with fewer confounding variables, thereby enhancing the classifier’s ability to identify patterns effectively. 58 Similarly, a study in Bangladesh using machine learning techniques reported an accuracy of 81.4% with the Random Forest classifier for predicting anemia among women. The current study outperformed this, potentially due to the integration of advanced preprocessing techniques, such as SMOTE for addressing class imbalance, and the use of state-of-the-art classifiers like XGBoost, which excel in capturing complex, non-linear relationships within the dataset 59
In this study, SHAP was used to interpret the model by identifying the relative contribution and direction of each predictor to the predicted risk of anemia among pregnant women. The SHAP analysis based on the XGBoost model, revealed that source of drinking water, ANC visit, mobile phone ownership, Age, got money for medical help, smoking cigarette, Khat chewing, media exposure, types of fuel used, distance form health facility and, place of delivery were the most relevant predictors of anemia among pregnant women in Sub-Saharan Africa.
Anemia is a leading cause of morbidity worldwide and is recognized as a major public health issue in Sub-Saharan Africa regions. 60 Particularly, reproductive-age women are more susceptible to anemia owing to the high iron during pregnancy, lactation, menstruation, and undernutrition throughout their reproductive period. 61 In this study, the prevalence of anemia among women of reproductive age was 42%. This prevalence is consistent with studies conducted in Nepal, 31 Myanmar, 62 and the Democratic Republic of Congo. 63 However, this prevalence of anemia is higher than in studies done in Brazil, 64 Iran, 65 and Turkey. 66 These disparities indicate that the prevalence of anemia can be attributed to geographical, cultural, and dietary factors in these countries. Specifically, developing countries face social and biological vulnerabilities within households and society. 67 Furthermore, developing countries suffer nutritional deficiencies, including nutrition-related anemia, primarily due to poverty and social position within the household. 68 Especially, in Eastern Africa, low socioeconomic status, high prevalence of communicable diseases, and limited healthcare access lead to inadequate access to iron-rich foods, resulting in anemia.
This study discovered that pregnant women using unsafe drinking water in their households had a higher risk of developing anemia. 69 This finding aligns with research from Uganda, 18 Rwanda, 19 Mali, 20 Tanzania, 20 and East Africa. 21 This could be due to women in this region having significant vulnerability to contracting various infectious diseases, as they lack access to safe water and adequate sanitation. Hence, they face a high risk of exposure to human feces through drinking water. 70 Human feces are highly contaminated with pathogens, significantly increasing the risk of waterborne and foodborne diseases. 71 These open-defecation practices in low-income regions can easily contaminate drinking water, often leading to bacterial diarrhea and parasitic diseases like hookworm and schistosomiasis, which are responsible for nutrient-deficient anemia. 72 This situation is worsened due to shifting the immune system during pregnancy to tolerate the developing fetus. 69 Recently, to mitigate this problem, iron interventions such as iron fortification and supplementation, and chemotherapy prevention have been widely employed to manage anemia. Nevertheless, these measures have proven inadequate in sub-Saharan Africa, where the burden of infectious diseases persists. 73 Fortunately, studies have highlighted that improving sanitation facilities and ensuring safe drinking water in developing countries can significantly reduce the prevalence of anemia in pregnant women. 74 Therefore, nutrient absorption, including iron, can be improved by avoiding these infections, especially crucial to preventing anemia.
The current study found that pregnant women with no repeated antenatal care (ANC) visits are highly prone to developing anemia. This finding was concordant with studies conducted in Ethiopia, 75 Ghana, 76 Tanzania, 77 and Benin 78, which have confirmed that anemia is common among women with less frequent ANC visits. This resemblance might be because most women in developing countries often attend fewer than four antenatal care visits throughout their pregnancy, facing increased risks of maternal and neonatal complications. 79 Consequently, they are more prone to anemia due to an inability to remember to take iron-folic acid pills and poorer nutritional status from a lack of counseling about iron-rich foods. However, as per WHO recommendations, pregnant women with four or more ANC visits were less likely to be anemic due to preventive measures such as malaria prophylaxis, iron and folate supplements, insecticide-treated nets, and nutritional and vaccination programs provided during these visits. 80 Furthermore, increasing awareness about the importance of ANC services can boost the number of ANC visits among pregnant women.
This study also found that pregnant women living in rural areas tend to have higher rates of anemia than those in urban settings. This finding is congruent with studies conducted in Ethiopia, 81 Malawi, 82 and East Africa. 21 This similarity could be attributed to socio-cultural differences. In rural areas, the prevalence of anemia was increased due to restricted access to micronutrient-rich diets due to food taboos, limited mass media access, insufficient nutrition information, and the aggravating effect of multiple nutrient deficiencies. 83 Besides, farming in rural areas may heighten the risk of anemia due to infectious and parasitic infestations such as malaria, schistosomiasis, and hookworm. 84 Our study also identified that women who considered the distance to a health facility to be a significant issue were more likely to experience higher levels of anemia compared to those who did not view it as a major problem, consistent with other findings from Gambia, 85 East Africa 21 and sub-Saharan Arica. 24 The observed outcome can be explained by the lack of access to treatments for anemia-causing disorders and maternal health services, including prenatal and postnatal care, due to the distance from health facilities contributing to higher anemia rates. Women living far from health facilities find it difficult to obtain preventive and therapeutic care, such as timely iron and folate supplementation during pregnancy, modern contraceptives, and other essential services, increasing their susceptibility to anemia. 86 Therefore, to address anemia among women living in rural and far from health facilities, strategies such as mobile health clinics, community health workers, dietary counseling, improved sanitation facilities, awareness campaigns, and telemedicine can be employed.
This study also discovered that women in households exposed to solid cooking fuel were more likely to experience anemia than those using clean cooking fuel. This conclusion was in agreement with studies conducted in Ethiopia, 22 Sri Lanka, 23 Sub-Saharan Africa, 24 and China. 25 This similarity can be attributed to pregnant women’s exposure to pollutants from solid cooking fuel, which has been associated with systemic inflammation, reducing red blood cell formation and iron balance. This may be linked with low serum iron levels and anemia. Moreover, indoor air pollution from solid fuel combustion increases carbon monoxide levels, which inhibits iron absorption in the gastrointestinal tract, contributing to iron-deficiency anemia.22,87 Therefore, adopting clean cooking fuels like biogas or electricity and improving kitchen ventilation are essential strategies to prevent anemia in pregnant women.
This study also demonstrated that pregnant women with inadequate media exposure and mobile phone ownership had a higher incidence of anemia. This finding is consistent with studies done in Indonesia, 55 Nepal, 26 and low and middle-income countries, 27 which showed that pregnant women exposed to either newspapers, radio, mass media, or social media (e.g. Facebook, WhatsApp, and Twitter) were less likely to develop anemia. An experimental study conducted in Jordan found that 24% of pregnant women in the control group and 54% in the experimental group were non-anemic after receiving media exposure. 88 Similarly, a quasi-experimental study in Indonesia indicated that health education and counseling through pictorial handbooks effectively reduced the number of women who became anemic. 89 These studies underscore that educating women through health training programs is an effective strategy for teaching proper dietary practices. Such programs can help them become familiar with food classifications and improve their ability to select foods that increase hemoglobin levels, like those rich in iron, protein, and vitamin C. 88 In addition, mass media channels, including posters, press and article advertisements, television, radio communications, and social media platforms, are being promoted to raise awareness about public health issues. 90 Therefore, raising awareness and providing information on health-related matters through various communication channels can be a crucial strategy for reducing anemia prevalence among women of reproductive age.
In this study, pregnant women who had a history of smoking and chewing khat were more likely to have anemia compared to women who had not used them. This result was similar to previous studies in India, 28 Ghana, 29 Ethiopia, 30 and Nepal, 31 which showed that smoking and khat use increased the risk of women having anemia. This similarity may be because smoking adversely affects overall health and nutrient absorption. When women use tobacco, it can cause deficiencies in essential nutrients like iron, vitamin C, and folate. These nutrients are vital for healthy red blood cell production. 91 Additionally, smoking causes chronic inflammation and oxidative stress, impairing red blood cell function and production, thereby increasing the risk of anemia. 92 On the other hand, smokers may show elevated hemoglobin levels due to the erythropoietin-stimulating effects of increased carbon monoxide from smoking, potentially leading to a lower observed prevalence of anemia among them. 93 Therefore, it’s essential to enhance the participation in smoking cessation programs and support groups. These initiatives are crucial in providing the guidance, encouragement, and strategies necessary to quit smoking successfully.
Strengths and limitations
Strengths
This study leveraged a large, nationally representative pooled dataset from 24 Sub-Saharan African countries, improving the generalizability of the findings. The use of advanced machine learning techniques, including XGBoost and Recursive Feature Elimination, allowed for robust prediction and identification of key predictors of anemia. Careful data preprocessing, balancing, and hyperparameter tuning further enhanced model performance and reliability.
Limitations
The cross-sectional design prevents establishing causality between predictors and anemia. Some variables were self-reported, which may introduce recall or reporting bias. Differences in survey years and country-specific factors were not fully accounted for, which may influence the results. The pooled prevalence may mask heterogeneity across countries and time periods. The binary classification of anemia may have obscured distinctions between mild, moderate, and severe cases. Although SHAP was used to improve model interpretability, the stability of SHAP feature importance rankings across different data splits or repeated model training was not formally evaluated. Finally, while machine learning models provide high predictive accuracy, their complexity may limit full interpretability, and unmeasured confounders could affect the findings.
Conclusion
This study demonstrated the efficacy of the XGBoost classifier in predicting anemia and identifying its key predictors among pregnant women in Sub-Saharan Africa. The model’s high accuracy and interpretability, achieved through advanced preprocessing and SHAP analysis, revealed critical factors such as unsafe drinking water, inadequate antenatal care visits, and reliance on solid cooking fuels. Behavioral factors like smoking and chewing khat, along with socioeconomic and geographic barriers, also significantly contributed to anemia risk. These findings underscore the need for targeted interventions, including improving water quality, expanding antenatal care coverage, promoting clean cooking fuels, and leveraging digital platforms for health education. By addressing these multifaceted predictors, policymakers and healthcare practitioners can implement effective strategies to reduce anemia prevalence and improve maternal health outcomes in resource-constrained regions. The study highlights the transformative potential of machine learning in public health, advocating for its broader application to inform evidence-based decision-making and resource allocation.
Recommendation
This study recommends integrating ML and AI-based anemia prediction tools, such as XGBoost with SHAP explanations, into maternal and child health information systems to enable early identification of high-risk pregnant women and guide targeted interventions. Generative AI could be leveraged to simulate personalized intervention scenarios and optimize preventive strategies. Digital health platforms and mobile applications should implement these models to provide real-time risk alerts during ANC visits, enhancing timely decision-making. Priority actions include improving access to safe drinking water, promoting clean cooking fuels, increasing ANC attendance, and using AI-driven mobile health education to reduce risky behaviors such as smoking and khat chewing. Policymakers are encouraged to adopt data-driven, multi-sectoral strategies, strengthen health informatics infrastructure, and build ML/AI capacity across Sub-Saharan Africa to ensure sustainable, scalable implementation of predictive and preventive maternal health programs.
Acknowledgments
The authors would like to thank the Measure DHS Program for granting access to the Demographic and Health Survey (DHS) datasets used in this study. We also acknowledge the University of Gondar and the Center for Digital Health Implementation Science for providing an academic and research-supportive environment that facilitated the completion of this work. We further extend our sincere appreciation to the Sankofa Book Readers Association for providing a supportive community and a “home away from home” that encouraged learning, reflection, and intellectual growth throughout this work.
Footnotes
Author contributions: EA, and AT developed the concept for the study. EA and AT reviewed the literature. EA conducted the data analysis. GY, AT, and EA discussed the findings. All authors proofread the manuscript for spelling and grammar, and all approved the final version for submission.
Funding: No financial support was received by the authors for the conduct of the research, authorship, or publication of this article.
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
ORCID iDs
Eliyas Addisu Taye https://orcid.org/0009-0005-1561-072X
Abel Temeche Kassaw https://orcid.org/0009-0008-3431-8187
Ethical considerations
This study used secondary data analysis, hence no direct participation from individuals was required. A consent letter for data access was obtained from a major health and demographic survey via a web-based request submitted to https://www.dhsprogram.com.
Consent to participate
This study used exclusively de-identified information, ensuring full compliance with ethical standards for participant privacy and confidentiality.
Data Availability Statement
The datasets analyzed in the current study are available in the public domain through the Measure DHS website https://dhsprogram.com/data/available-datasets.cfm.
References
- 1.dos Santos BS, Steiner MTA, Fenerich AT, et al. Data mining and machine learning techniques applied to public health problems: A bibliometric analysis from 2009 to 2018. Computers & Industrial Engineering 2019; 138: 106120. 10.1016/j.cie.2019.106120 [DOI] [Google Scholar]
- 2.Biswas A, Tucker J, Bauhoff S. Performance of predictive algorithms in estimating the risk of being a zero-dose child in India, Mali and Nigeria. BMJ Global Health 2023; 8(10): e012836. 10.1136/bmjgh-2023-012836 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Ftye M, Letta A, Achamyeleh B. DESIGNING A PREDICTIVE MODEL FOR THE LIKELIHOOD OF CONTRACEPTIVE METHOD USAGE IN ETHIOPIA. Cosmos Journal of Engineering & Technology 2022; 12(1): 1-6. [Google Scholar]
- 4.Balarajan Y, Ramakrishnan U, Fau - Ozaltin E, et al. Anaemia in low-income and middle-income countries. The Lancet, 2011, pp. 1474. (Electronic). [DOI] [PubMed] [Google Scholar]
- 5.World Health O . Haemoglobin concentrations for the diagnosis of anaemia and assessment of severity. World Health Organization, 2011. [Google Scholar]
- 6.Rahman MM, Abe SK, Rahman MS, et al. Maternal anemia and risk of adverse birth and health outcomes in low- and middle-income countries: systematic review and meta-analysis. The American Journal of Clinical Nutrition, 2016, pp. 1938–3207. (Electronic). [DOI] [PubMed] [Google Scholar]
- 7.Ibrahim ZM, El-Hamid SA, Mikhail H, et al. Assessment of adherence to iron and folic acid supplementation and prevalence of anemia in pregnant women. Med J Cairo Univ 2011; 79(2): 115–121. [Google Scholar]
- 8.Daru J, Zamora J, Fernández-Félix BM, et al. Risk of maternal mortality in women with severe anaemia during pregnancy and post partum: a multilevel analysis. The Lancet Golbal Health, 2018, pp. 2214. (Electronic). [DOI] [PubMed] [Google Scholar]
- 9.Stevens GA, Finucane MM, De-Regil LM, et al. Global, regional, and national trends in haemoglobin concentration and prevalence of total and severe anaemia in children and pregnant and non-pregnant women for 1995-2011: a systematic analysis of population-representative data. The Lancet Golbal Health, 2013, pp. 2214. (Electronic). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.McLean E, Cogswell M, Fau - Egli I, et al. Worldwide prevalence of anaemia, WHO Vitamin and Mineral Nutrition Information System, 1993-2005. Cambridge University Press, 2009, pp. 1368–9800. (Print). [DOI] [PubMed] [Google Scholar]
- 11.Kassa GM, Muche AA, Berhe AK, et al. Prevalence and determinants of anemia among pregnant women in Ethiopia; a systematic review and meta-analysis. BMC Hematology, 2017, pp. 1839–2052. (Print). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Fite MB, Assefa N, Mengiste B. Prevalence and determinants of Anemia among pregnant women in sub-Saharan Africa: a systematic review and Meta-analysis. Archives of Public Health 2021; 79(1): 219. 10.1186/s13690-021-00711-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Haider BA, Olofin I, Fau - Wang M, et al. Anaemia, prenatal iron use, and risk of adverse pregnancy outcomes: systematic review and meta-analysis. The BMJ, 2013, pp. 1756–1833. (Electronic). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Adam I, Ibrahim Y, Elhardello O. Prevalence, types and determinants of anemia among pregnant women in Sudan: a systematic review and meta-analysis. BMC hematology 2018; 18: 1–8. 10.1186/s12878-018-0124-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Fowkes FJI, Moore KA, Opi DH, et al. Iron deficiency during pregnancy is associated with a reduced risk of adverse birth outcomes in a malaria-endemic area in a longitudinal cohort study. BMC medicine 2018; 16: 1–10. 10.1186/s12916-018-1146-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Obeagu E. Maternal Anemia and Its Impact on Fetal Growth and Development: A Review. Annals of Hematology & Oncology 2024; 11: 1468. [Google Scholar]
- 17.Suryanarayana R, Chandrappa M, Santhuram AN, et al. Prospective study on prevalence of anemia of pregnant women and its outcome: A community based study. Medknow Publications, 2017, pp. 2249–4863. (Print)). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Nankinga OA-O, Aguta D. Determinants of Anemia among women in Uganda: further analysis of the Uganda demographic and health surveys. BMC Public Health, 2019, pp. 1471–2458. (Electronic)). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Habyarimana F, Zewotir T, Ramroop S. Spatial Distribution and Analysis of Risk Factors Associated with Anemia Among Women of Reproductive Age: Case of 2014 Rwanda Demographic and Health Survey Data. The Open Public Health Journal 2018; 11: 425–437. 10.2174/1874944501811010425 [DOI] [Google Scholar]
- 20.Kothari M, Coile A, Huestis A, et al. Exploring associations between water, sanitation, and anemia through 47 nationally representative demographic and health surveys. Annals of the New York Academy of Sciences 2019; 1450: 249–267. 10.1111/nyas.14109 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Teshale A, Tesema G, Worku M, et al. Anemia and its associated factors among women of reproductive age in eastern Africa: A multilevel mixed-effects generalized linear model. PLOS ONE 2020; 15: e0238957. 10.1371/journal.pone.0238957 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Kanno GG, Geremew T, Diro T, et al. The link between indoor air pollution from cooking fuels and anemia status among non-pregnant women of reproductive age in Ethiopia. Sage Open Medicine, 2022; 10: 20503121221107466. 10.1177/20503121221107466 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Pathirathna MA-O, Samarasekara BA-OX, Mendis C, et al. Is biomass fuel smoke exposure associated with anemia in non-pregnant reproductive-aged women? PLOS One, 2022, pp.1932–6203. (Electronic). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Tirore LL, Areba AS, Habte A, et al. Prevalence and associated factors of severity levels of anemia among women of reproductive age in sub-Saharan Africa: a multilevel ordinal logistic regression analysis. Frontiers in Public Health, 2024, pp. 2296–2565. (Electronic). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.He Y, Liu X, Zheng Y, et al. Lower socioeconomic status strengthens the effect of cooking fuel use on anemia risk and anemia-related parameters: Findings from the Henan Rural Cohort. Science of The Total Environment, 2022, pp. 1879. (Electronic). [DOI] [PubMed] [Google Scholar]
- 26.Acharya D, Adhikari R, Simkhada P. Prevalence and Determinants of Anaemia among Women aged 15-49 in Nepal: A Trend Analysis from Nepal Demographic and Health Surveys from 2006 to 2016. Asian Journal of Population Sciences 2022; 1: 32–48. 10.3126/ajps.v1i1.43593 [DOI] [Google Scholar]
- 27.Alem AZ, Efendi F, McKenna L, et al. Prevalence and factors associated with anemia in women of reproductive age across low- and middle-income countries based on national data. Scientific Reports, 2023, pp. 2045–2322. (Electronic). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Kuppusamy PA-O, Prusty RA-O, Khan SA-O. Assessing the prevalence and predictors of anemia among pregnant women in India: findings from the India National Family Health Survey 2019-2021. Taylor & Francis online, 2024, pp. 1473–4877. (Electronic)). [DOI] [PubMed] [Google Scholar]
- 29.Appiah-Dwomoh C, Tettey P, Akyeampong E, et al. Smoke exposure, hemoglobin levels and the prevalence of anemia: a cross-sectional study in urban informal settlement in Southern Ghana. BMC Public Health 2024; 24(1): 854. 10.1186/s12889-024-18304-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Misgana T, Tesfaye D, Alemu D, et al. Khat use and associated factors during pregnancy in eastern Ethiopia: A community-based cross-sectional study. Frontiers in Global Women's Health, 2022, pp. 2673–5059 (Electronic). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Gautam S, Min H, Kim H, et al. Determining factors for the prevalence of anemia in women of reproductive age in Nepal: Evidence from recent national survey data. PLOS ONE, 2019, pp. 1932–6203. (Electronic)). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Bizzego A, Gabrieli G, Bornstein MH, et al. Predictors of contemporary under-5 child mortality in low-and middle-income countries: A machine learning approach. International journal of environmental research and public health 2021; 18(3): 1315. 10.3390/ijerph18031315 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Ij H. Statistics versus machine learning. Nat Methods 2018; 15(4): 233. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Sarker IH. Machine learning: Algorithms, real-world applications and research directions. SN computer science 2021; 2(3): 160. 10.1007/s42979-021-00592-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Dhar V. Data science and prediction. Communications of the ACM 2013; 56(12): 64–73. 10.1145/2500499 [DOI] [Google Scholar]
- 36.Zhang A, Xing L, Zou J, et al. Shifting machine learning for healthcare from development to deployment and from models to data. Nature Biomedical Engineering 2022; 6(12): 1330–1345. 10.1038/s41551-022-00898-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Liyew AM, Tesema GA, Alamneh TS, et al. Prevalence and determinants of anemia among pregnant women in East Africa; A multi-level analysis of recent Demographic and Health Surveys. PLOS ONE 2021; 16(4): e0250560. 10.1371/journal.pone.0250560 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Woldegebriel AG, Gebregziabiher Gebrehiwot G, Aregay Desta A, et al. Determinants of anemia in pregnancy: findings from the Ethiopian health and demographic survey. Anemia 2020; 2020(1): 2902498. 10.1155/2020/2902498 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Gebremedhin S, Enquselassie F. Correlates of anemia among women of reproductive age in Ethiopia: evidence from Ethiopian DHS 2005. Ethiopian Journal of Health Development 2011; 25(1): 22–30. 10.4314/ejhd.v25i1.69842 [DOI] [Google Scholar]
- 40.Wu X, Zheng B. Whom Does Algorithmic Risk Stratification Miss? A Fairness Audit of Machine Learning Targeting for Concurrent Maternal–Child Double Burden of Malnutrition Across 30 Low-and Middle-Income Countries. medRxiv 2026; 26352000. [Google Scholar]
- 41.Brownlee J. Statistical methods for machine learning: Discover how to transform data into knowledge with Python. Machine Learning Mastery, 2018. [Google Scholar]
- 42.Kuhn M, Kuhn M, Johnson K, et al. An introduction to feature selection. Applied predictive modeling, 2013, pp. 487–519. [Google Scholar]
- 43.Chen T, Guestrin C. (eds). Xgboost: A scalable tree boosting system. Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining. Association for Computing Machinery, 2016. [Google Scholar]
- 44.Chen T. Xgboost: extreme gradient boosting. R package version 04-2 2015; 1(4): 785-794. [Google Scholar]
- 45.Lai SBS, Shahri N, Mohamad MB, et al. Comparing the performance of AdaBoost, XGBoost, and logistic regression for imbalanced data. Mathematics and Statistics 2021; 9(3): 379–385. 10.13189/ms.2021.090320 [DOI] [Google Scholar]
- 46.Shao Z, Ahmad MN, Javed A. Comparison of Random Forest and XGBoost Classifiers Using Integrated Optical and SAR Features for Mapping Urban Impervious Surface. Remote Sensing 2024; 16(4): 665. 10.3390/rs16040665 [DOI] [Google Scholar]
- 47.Machine Learning for Engineers Available from. https://apmonitor.com/pds/index.php/Main/XGBoostClassifier [Google Scholar]
- 48.Šimundić AM. Measures of Diagnostic Accuracy. Basic Definitions. Ejifcc 2009; 19(4): 203–211. [PMC free article] [PubMed] [Google Scholar]
- 49.Santini A, Man A, Voidăzan S. Accuracy of diagnostic tests. The Journal of Critical Care Medicine 2021; 7(3): 241–248. 10.2478/jccm-2021-0022 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Flach P, Kull M. Precision-recall-gain curves: PR analysis done right. Advances in neural information processing systems 2015; 28: 838-846. [Google Scholar]
- 51.Powers DM. Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation. arXiv preprint arXiv:201016061 2020; 2: 37-63. [Google Scholar]
- 52.Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature machine intelligence 2019; 1(5): 206–215. 10.1038/s42256-019-0048-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Nohara Y, Matsumoto K, Soejima H, et al. Explanation of machine learning models using shapley additive explanation and application for real data in hospital. Computer Methods and Programs in Biomedicine 2022; 214: 106584. 10.1016/j.cmpb.2021.106584 [DOI] [PubMed] [Google Scholar]
- 54.Palatnik de Sousa I, Maria Bernardes Rebuzzi Vellasco M, Costa da Silva E. Local interpretable model-agnostic explanations for classification of lymph node metastases. Sensors 2019; 19(13): 2969. 10.3390/s19132969 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Danesh T, Ouaret R, Floquet P, et al. Interpretability of neural networks predictions using Accumulated Local Effects as a model-agnostic method. Computer aided chemical engineering 2022; 51: 1501–1506, Elsevier. [Google Scholar]
- 56.Wright R. Interpreting black-box machine learning models using partial dependence and individual conditional expectation plots. Exploring SAS® Enterprise Miner Special Collection, 2018, vol 1950. [Google Scholar]
- 57.Bifarin OO. Interpretable machine learning with tree-based shapley additive explanations: Application to metabolomics datasets for binary classification. Plos one 2023; 18(5): e0284315. 10.1371/journal.pone.0284315 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Kitaw B, Asefa C, Legese F. Leveraging machine learning models for anemia severity detection among pregnant women following ANC: Ethiopian context. BMC Public Health 2024; 24(1): 3500. 10.1186/s12889-024-21039-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Islam MM, Rahman MJ, Roy DC, et al. Risk factors identification and prediction of anemia among women in Bangladesh using machine learning techniques. Current Women's Health Reviews 2022; 18(1): 118–133. 10.2174/1573404817666210215161108 [DOI] [Google Scholar]
- 60.Mwangi MN, Mzembe G, Moya E, et al. Iron deficiency anaemia in sub-Saharan Africa: a review of current evidence and primary care recommendations for high-risk groups. The Lancet Haematology 2021; 8(10): e732–e743. 10.1016/s2352-3026(21)00193-9 [DOI] [PubMed] [Google Scholar]
- 61.Mawani M, Aziz Ali S, Ali G, et al. Iron Deficiency Anemia among Women of Reproductive Age, an Important Public Health Problem: Situation Analysis. Reproductive System & Sexual Disorders 2020; 5: 1-6. [Google Scholar]
- 62.Win H, Ko M. Geographical disparities and determinants of anaemia among women of reproductive age in Myanmar: analysis of the 2015–2016 Myanmar Demographic and Health Survey. WHO South-East Asia Journal of Public Health 2018; 7: 107–113. 10.4103/2224-3151.239422 [DOI] [PubMed] [Google Scholar]
- 63.Kandala NA-O, Pallikadavath S, Amos Channon A, et al. A multilevel approach to correlates of anaemia in women in the Democratic Republic of Congo: findings from a nationally representative survey. European Journal of Clinical Nutrition, 2023, pp. 1476–5640. (Electronic). [DOI] [PubMed] [Google Scholar]
- 64.Bezerra AGN, Leal VS, Lira PIC, et al. Anemia and associated factors in women at reproductive age in a Brazilian Northeastern municipality. Associação Brasileira de Saúde Coletiva, 2018, pp. 1980–5497. (Electronic). [Google Scholar]
- 65.Sadeghian M, Fatourechi A, Lesanpezeshki M, et al. Prevalence of anemia and correlated factors in the reproductive age women in rural areas of tabas. Tehran University of Medical Sciences, 2013, pp. 1735–8949. (Print). [PMC free article] [PubMed] [Google Scholar]
- 66.Saydam BK, Genc RE, Sarac F, et al. Prevalence of anemia and related factors among women in Turkey. Professional Medical Publications, 2017. (Print). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Bentley M, Griffiths P. The burden of anemia among women in India. European journal of clinical nutrition 2003; 57: 52–60. 10.1038/sj.ejcn.1601504 [DOI] [PubMed] [Google Scholar]
- 68.Chalise BA-O, Aryal KK, Mehta RK, et al. Prevalence and correlates of anemia among adolescents in Nepal: Findings from a nationally representative cross-sectional survey. PLOS ONE, 2018, pp. 1932–6203. (Electronic). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Kmush B, Walia B, Neupane A, et al. Community-level impacts of sanitation coverage on maternal and neonatal health: A retrospective cohort of survey data. BMJ global health 2021; 6: e005674. 10.1136/bmjgh-2021-005674 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Mara D, Lane J, Scott B, et al. Sanitation and Health. PLoS medicine 2010; 7: e1000363. 10.1371/journal.pmed.1000363 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Inah A, Goodness C, Kalu R, et al. Assessment Of Access To Safe Drinking Water And Water Quality Of Rural Communities Of Akpabuyo Local Government Area Of Cross River State. International Journal of Environment and Pollution Research 2020;13:50–58. [Google Scholar]
- 72.Abate M, Kinfe B, Tadesse D, et al. Intestinal parasitosis and anaemia among patients in a Health Center, North Ethiopia. BMC Research Notes 2017; 10: 632. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Pasricha SR, Armitage AE, Prentice AM, et al. Reducing anaemia in low income countries: control of infection is essential. The BMJ, 2018, pp. 1756–1833. (Electronic)). [DOI] [PubMed] [Google Scholar]
- 74.Chanimbe B, Issah A-N, Mahama AB, et al. Access to basic sanitation facilities reduces the prevalence of anaemia among women of reproductive age in sub-saharan Africa. BMC Public Health 2023; 23(1): 1999. 10.1186/s12889-023-16890-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Alemu T, Umeta M. Reproductive and Obstetric Factors Are Key Predictors of Maternal Anemia during Pregnancy in Ethiopia: Evidence from Demographic and Health Survey. Wiley Online Library, 2011, pp. 2090. (Print). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Akowuah JA, Owusu-Addo E, Antwiwaa AO. Predictors of Anaemia prevalence among pregnant women in urban Ghana: a cross-sectional study. INQUIRY: The Journal of Health Care Organization, Provision, and Financing, 2019. [Google Scholar]
- 77.Sunguya BF, Ge Y, Mlunde L, et al. High burden of anemia among pregnant women in Tanzania: a call to address its determinants. Nutrition Journal, 2021, pp. 1475. (Electronic). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Bodeau-Livinec F, Briand V, Fau - Berger J, et al. Maternal anemia in Benin: prevalence, risk factors, and association with low birth weight. American Society of Tropical Medicine and Hygiene, 2011, pp. 1476–1645. (Electronic). [Google Scholar]
- 79.Kuhnt J, Vollmer S. Antenatal care services and its implications for vital and health outcomes of children: evidence from 193 surveys in 69 low-income and middle-income countries. BMJ Open 2017; 7(11): e017122. 10.1136/bmjopen-2017-017122 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80.Verma R, Verma L. A Literature Review on Emerging Factors Affecting Antenatal Care Utilization of Pregnant Women. International Journal of Food and Nutritional Sciences 2022; 11: 58–66. [Google Scholar]
- 81.Kenea A, Negash E, Bacha L, et al. Magnitude of Anemia and Associated Factors among Pregnant Women Attending Antenatal Care in Public Hospitals of Ilu Abba Bora Zone, South West Ethiopia: A Cross-Sectional Study. Wiley Online Library, 2018, pp. 1267–2090. (Print). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82.Adamu AL, Crampin A, Kayuni N, et al. Prevalence and risk factors for anemia severity and type in Malawian men and women: urban and rural differences. Population Health Metrics, 2017, pp. 1478–7954. (Electronic). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Abate TW, Getahun B, Birhan MM, et al. The urban–rural differential in the association between household wealth index and anemia among women in reproductive age in Ethiopia, 2016. BMC Women's Health 2021; 21(1): 311. 10.1186/s12905-021-01461-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Berhe K, Fseha B, Gebremariam G, et al. Risk factors of anemia among pregnant women attending antenatal care in health facilities of Eastern Zone of Tigray. Ethiopia, case-control study; 2018; 34: 1937–8688. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85.Shitu K, Terefe B. Anaemia and its determinants among reproductive age women (15–49 years) in the Gambia: a multi-level analysis of 2019–20 Gambian Demographic and Health Survey Data. Archives of Public Health 2022; 80(1): 228. 10.1186/s13690-022-00985-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86.Dotse-Gborgbortsi W, Nilsen K, Ofosu A, et al. Distance is “a big problem”: a geographic analysis of reported and modelled proximity to maternal health services in Ghana. BMC Pregnancy and Childbirth 2022; 22(1): 672. 10.1186/s12884-022-04998-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87.Zou Y-Y, Yuan Y, Kan EM, et al. Combustion smoke-induced inflammation in the olfactory bulb of adult rats. Journal of Neuroinflammation 2014; 11: 1–16. 10.1186/s12974-014-0176-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Abujilban S, Hatamleh R, Al-Shuqerat S. The impact of a planned health educational program on the compliance and knowledge of Jordanian pregnant women with anemia. Women Health , 2019, pp. 1541. (Electronic). [DOI] [PubMed] [Google Scholar]
- 89.Nahrisah PA-O, Somrongthong R, Viriyautsahakul NA-O, et al. Effect of Integrated Pictorial Handbook Education and Counseling on Improving Anemia Status, Knowledge, Food Intake, and Iron Tablet Compliance Among Anemic Pregnant Women in Indonesia: A Quasi-Experimental Study. Dove Medical Press, 2020, pp. 1178–2390. (Print). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.Al-Dmour HA-O, Masa'deh RA-O, Salman AA-O, et al. Influence of Social Media Platforms on Public Health Protection Against the COVID-19 Pandemic via the Mediating Effects of Public Health Awareness and Behavioral Changes: Integrated Model. 2020, pp. 1438–8871. (Electronic). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91.Sma W, Alvi A. Correlation between anemia and smoking: Study of patients visiting different outpatient departments of Integral Institute of Medical Science and Research, Lucknow. National Journal of Physiology, Pharmacy and Pharmacology 2019. 9: 1. [Google Scholar]
- 92.Malenica M, Prnjavorac B, Bego T, et al. Effect of Cigarette Smoking on Haematological Parameters in Healthy Population. Avicena Media, 2017. (Print). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 93.Nordenberg D, Yip R, Fau - Binkin NJ, et al. The effect of cigarette smoking on hemoglobin levels and anemia screening. Journal of the American Medical Association, 1990. (Print). [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The datasets analyzed in the current study are available in the public domain through the Measure DHS website https://dhsprogram.com/data/available-datasets.cfm.
