Skip to main content
One Health logoLink to One Health
. 2024 Aug 13;19:100874. doi: 10.1016/j.onehlt.2024.100874

Contribution of artificial intelligence for understanding animal rabies epidemiology in Morocco: What are the perspectives of an innovative and predictive approaches?

Ilham Ahamjik a,, Ayman Agbani b, Mounia Abik c, Mounir Khayli a, Naima Galzim a, Jaouad Berrada b, Mohammed Bouslikhane b
PMCID: PMC11378930  PMID: 39247759

Abstract

Rabies is a major zoonotic disease legally notifiable in Morocco and elsewhere. Given the burden of rabies and its impact on public health, several national control programs have been implemented since 1986, without achieving their expected objectives.

The aim of this study was to design a predictive analysis of rabies in Morocco. The expected outcome was the construction of probabilistic diagrams that can guide actions for the integrated control of this disease, involving all stakeholders, in the country. Such modeling is an essential step in operational epidemiology to optimize expenditure of public funds allocated to the integrated strategy for fighting this disease.

The methodology employed combined the use of geospatial analysis tools (kriging) and artificial intelligence models (Machine Learning). In order to investigate the link between the risk of rabies within a territorial municipality (commune) and its socio-economic situation, the following data were analyzed: (1) health data: reported animal cases of rabies between 2004 and 2021 and data obtained through the ArcGIS kriging tool (Geospatial data); (2) demographic and socio-economic data. We compared several Machine Learning models. Of these, the “Imbalanced-Xgboost” model associated with kriging yielded the best results. After optimizing this model, we mapped our results for better visualization.

The obtained results complement and consolidate previous study in this field with greater accuracy, showing a strong correlation between a commune's socio-economic status, its geographical location and its risk level of rabies. From this, 399 out of the 1546 communes have been identified as high-risk areas, accounting for 25.8% of the total number of communes. Under this risk-based approach, the results of these analyses make it practical to take targeted decisions for rabies prevention and control, as well as canine population control, in a territorial commune according to its risk level. Such an approach allows obvious optimized distribution of financial resources and adaptation of the control actions to be taken.

The study highlights also the importance of using innovative technologies to refine epidemiological approaches and fill gaps in field data. Through this study, we hope to contribute to eradication of rabies in Morocco by providing reliable data and practical recommendations for control actions against rabies.

Keywords: Rabies, Epidemiology, Public health, Morocco, Kriging, Artificial intelligence, Machine learning

1. Introduction

Rabies is a viral zoonosis that affects the central nervous system. This disease affects all warm-blooded mammals. The virus is in particular present in the saliva and brain of infected animals, most often dogs. It is generally transmitted by the bite of a sick animal. Bats are also a major reservoir of the virus in certain regions [1]. Its strong presence in animal populations creates multiple opportunities for transmission from one species to another, which can mainly affect domestic animals and humans [1].

Rabies is one of the most fatal zoonoses. The disease has been documented as far back as 2300 BCE [2]. For over 4000 years, rabies has affected almost every region of the world, and important efforts have been made to eradicate it. Rabies has disappeared from Western Europe, North America, Japan, South Korea and parts of Latin America, but persists in many parts of Africa and Asia [1]. Unlike many other diseases, the tools to eliminate dog-borne rabies are available. 100% prevention of the disease is possible [1], and canine rabies vaccines effectively eliminate the disease.

In Morocco, rabies is a contagious disease, which legally must be reported and is subject to the veterinary health measures set out in current legislation and regulations (https://www.onssa.gov.ma/reglementation/reglementation-sectorielle/sante-animale/maladies-animales). This zoonosis is endemic in the country and constitutes a public health problem. The main reservoir and source of contamination for humans and animals is the dog. The disease is reported in most of the provinces, but to varying degrees depending on dog population density and the socio-cultural and economic conditions. Several plans to control the disease have been implemented in Morocco since the 1980s, but the disease has not yet been eradicated.

Around 20,000 people a year receive post-exposure rabies treatment. The average cost of rabies treatment is 700 Dirhams (around $71) per person, and is paid for entirely by local authorities. Human rabies prevention costs the state around 14.5 million Dirhams (around $1.49 million) each year. The costs include biological analysis, rabies centers and costs of rehabilitation of exposed people [3]. Also, the cost of an annual rabies eradication campaign is around 11,446,475 dirhams (around $1.17 million) [4]. As rabies remains a One Health problem involving several departments, its solution can only be multi-sectoral [4]. Controlling rabies in dogs, and consequently in humans, will certainly have very positive repercussions for the population, the livestock sector and country-based tourism activities. In addition, the success of any rabies control strategy depends on taking into account the epidemiological situation of the disease, a good knowledge of the risk factors associated with the occurrence of the disease, and the rigorous application of the proposed control measures. Moreover, integrated approaches for rabies control based on a One Health approach offer clear advantages for disease elimination [5]. Integrated control of canine rabies and human exposure, through closer cooperation between public health and animal health, increases the sensitivity of control and avoids overuse or unnecessary use of post-exposure prophylaxis (PEP) treatments. One Health approaches provide further evidence in favor of eliminating canine rabies [6].

In order to solve the problem of rabies in Morocco, we need to adapt rabies control measures at local level, starting by targeting the communes most at risk. As part of this, we used artificial intelligence tools to identify the most important factors in the spread and persistence of rabies in Morocco.

Artificial Intelligence (AI) is the implementation of a number of techniques designed to enable machines to imitate a real form of intelligence. AI is being implemented in a growing number of fields of application. Aimed at simulating human intelligence, artificial intelligence has been emerging since the early 2010s driven by Deep Learning, Big Data and the increasing of computing power [7]. AI is used in many fields, such as medicine, finance, manufacturing, gaming, security, logistics, cars, education and many others. AI is seen as an ever-evolving technology with the potential to revolutionize many aspects of everyday life and transform businesses and industries [8].

2. Methods

The methodology adopted in the present study is spatial modeling using geostatistical and Machine Learning tools to design a predictive model of disease clusters to help implement an appropriate integrated control strategy at the local level. As part of this study, we performed a number of analyses. The methods below are organized in terms of coverage and in-depth analysis.

First, we analyzed the spatial distribution of animal rabies in Morocco over the period 2004–2021, taking into account that some cases of animal rabies in Morocco, in particular among free roaming dogs, are not systematically detected by the surveillance system. Then, we applied a predictive model using interpolation methods to improve the accuracy of our data. For this, we used Kriging, a geostatistical analysis tool [9]. This method is based on the assumption that the spatial field of the variable under study is a completion of a random function. The basic idea of kriging is to predict the value of the regionalized variable studied at an unsampled site by a linear combination of data from adjacent points. In the rabies case, each data point is referenced by its spatial location [9].

Then we used Machine Learning tools to design a model for classifying municipalities ‘communes’ at risk of rabies.

In order to understand the problem of animal rabies in Morocco and to propose an appropriate control strategy, a set of data was collected to cover a wide range of indicators relating to the environment and human activities, as well as to reflect the reality in the field. These data are as follows:

  • A collection of data from the database of the Office National de Sécurité Sanitaire des Produits Alimentaires (ONSSA) over the period 2004 to 2021. A total of 4587 cases of animal rabies were recorded, with an annual average of 255 cases.

  • A dataset of 42,024 demographic, sociological and economic records for 1546 Moroccan communes at national level from the Haut-Commissariat au Plan (2014).

  • Data on the risk factors of animal rabies highlighted in the study on the determining factors of rabies in Morocco [4], taking into account several factors such as geographical location, socio-economic and demographic characteristics of the environment in Morocco.

  • A shape file representing Morocco's 1546 communes.

2.1. Crisp-DM method for machine learning

Cross-Industry Standard Process for Data Mining (CRISP-DM) is a process model for data mining projects and is considered the standard technic for data mining projects [10]. CRISP-DM was published in 2000 and includes the following stages: understanding the problem and data, data preparation, modeling, evaluation and deployment (Fig. 1, [11]).

Fig. 1.

Fig. 1

Diagram of CRISP-DM method [11]. CRISP_DM is a field-tested method for guiding data mining work. This diagram is a description of the data mining lifecycle.

https://www.ibm.com/docs/fr/spss-modeler/saas?topic=dm-crisp-help-overview

In the following, we will detail the stages of this methodology:

Stage 1: Understanding the problem, the data and data processing.

This stage involves identifying the problem to be solved, defining the project objectives and establishing the success criteria. It is also necessary to identify and gather data sources, then carry out an initial exploratory analysis to assess their quality and relevance to the analysis. Data are cleaned, transformed and pre-processed to prepare them for modeling.

  • a.

    Available data: Health and socio-economic data

The field data are frequently incomplete and often contain errors. For this reason, it is important to pre-process data in order to resolve the above problems and prepare it for further processing [12]. For our study, we used data from different source, namely the ONSSA database, related to registered cases of animal rabies., and demographic and socio-economic data, from the Haut-Commissariat au Plan (https://www.hcp.ma).

Data description (Fig. 2, Fig. 3, Fig. 4): We used data from 1546 communes in Morocco. These data were divided into two distinct categories: firstly, there were 32 variables related to population demography and socio-economics:

  • Volumetric index (in%)

  • Poverty rate (%)

  • Poverty severity index (%)

  • Vulnerability rate (%)

  • Average number of cattle in the commune

  • Average number of small ruminants in the commune

  • Average accessibility rate

  • Inequality

  • Human Development Index

  • Communal Development Index

  • Population and households Rural Population

  • Population and households Rural Households

  • Population and households Rural Average size

  • Average accessibility rate

  • Total slaughterhouses

  • Secondary slaughterhouses

  • Number of markets

  • Land use

  • Secondary land use

  • Tertiary land use

  • Poverty

  • Vulnerability

  • Severity of poverty

  • Occupancy rate Rural

  • Population and households Urban Population

  • Population and households Urban Households

  • Population and households Urban Average size

  • Occupancy rate Urban

  • Male Population

  • Male Illiteracy rate

  • Female population

  • Female illiteracy rate

Fig. 2.

Fig. 2

Histogram of the Occurrence of variable ‘yearly average of rabies cases over 18-years’. This histogram displays the frequency distribution of the yearly average number of rabies cases in animals over an 18-year period. The majority of the communes had <1 case in the 18 years studied.

Fig. 3.

Fig. 3

Density plot for the variable ‘yearly average of rabies cases over 18 years’. This density plot illustrates the density of the yearly average number of rabies cases over an 18-year period.

Fig. 4.

Fig. 4

Box plot for the variable ‘yearly average of rabies cases over 18 years’. This box plot summarizes the distribution of the yearly average number of rabies cases over an 18-year period.

Secondly, we exploited health data on annual cases of rabies in dogs and other species, for the period from 2004 to 2021 and we used geo-spatial data obtained by interpolation of rabies cases by location using the ArcGIS kriging tool.

Rabies cases were taken for the period from 2004 to 2021, and we used the annual average of rabies cases. This new variable has a minimum value of 0 and a maximum value of 4.3. The mean of this variable is 0.19. The median of the variable is 0.055. Finally, the third quartile of this variable is 0.22, indicating that 75% of values are below this threshold.

When there is a large number of variables, the final model can become complex and the risk of overlearning can increase. It is therefore crucial to reduce the dimensionality of our study while retaining essential information. To this end, we have carried out a selection of variables to retain only those that are necessary. Based on the correlation between the different variables, we only kept those that were independent of each other. Examples of correlated variables: Poverty rate, poverty index, vulnerability rate, poverty, vulnerability, poverty severity, and inequality.

Population and household data were merged into a single variable called “Total_Population”, while male and female illiteracy rates were combined into a single variable called “Illiteracy_Rate”. We merged data on slaughterhouses and markets “souks” into a single variable called ‘Feed’ in our database, as they represent an important part of feed accessibility for dogs.

Having made this preliminary selection and taking into account the qualitative variable ‘Type’, we were able to reduce our demographic and socioeconomic data to a total of 11 variables (Fig. 5). We transformed the total cases variable into an average over an 18-year period. Then, in order to use classification models, we determined significant thresholds to divide these averages into several classes.

Fig. 5.

Fig. 5

Correlation matrix of quantitative variables retained after preliminary selection. This correlation matrix shows the relationships between various quantitative variables used in the study. It shows a strong correlation between the illiteracy rate and both the human development index and the CDL index. This analysis helped in the choice of which variable to use in our model.

  • b.

    Geo-spatial data: interpolation using the Kriging tool

For the integration of geo-spatial data in our analysis, we relied on a study carried out [9], where the interest of interpolation via Kriging applied to rabies data was shown. The linear estimation method known as “Kriging” ensures minimum variance. It performs spatial interpolation of a regionalized variable, based on the calculation of the mathematical expectation of a random variable, and using modeling and interpretation of the experimental variogram. This method is considered the best linear unbiased estimator, as it takes into account both the distance between the data and the point of estimation, as well as the distances between the data themselves [9].

We used the ArcGIS tool to perform spatial interpolation. We compiled a database listing the communes that had reported >6 cases of rabies during the period under review. This database was used as input for Kriging. This method enabled us to identify 160 communes that had never reported a case of rabies during the study period but which, according to kriging, were at risk with a threshold set at >0.69 on average rabies cases (at least 13 cases over 18 years) (Fig. 6).

Fig. 6.

Fig. 6

Map of communes predicted to be at risk by kriging tool without clinical cases reported in the field. This map visualizes the communes predicted to be at risk of rabies based on kriging tool (interpolative method), despite no clinical cases being reported in these areas.

We withdrew the 160 predicted communes from our database to be used for model evaluation, as well as the 84 communes located in the south of the country, which were automatically eliminated by kriging due to their distance from the communes where rabies is declared. This exclusion was made to avoid noise in our model.

Finally, we compared this prediction with the results obtained by interpolation (kriging). The new variable obtained by interpolation represents the 18-year averages of rabies cases. This variable was then divided into two classes: low risk (<0.69) and high risk (greater than or equal to 0.69). Among the communes studied, we identified 947 communes belonging to the low-risk class, and 296 communes classified as high-risk (Fig. 7).

Fig. 7.

Fig. 7

Distribution of communes according to their level of rabies risk. The data here shows its unbalanced nature, while 947 communes are in the class 0 (low risk) and 296 are in class 1 (high risk).

Stage 2: Modeling.

In this stage, various modeling methods are applied to the prepared data to create a predictive or descriptive model. In our case, we need to create a supervised learning model and, more specifically, a classification model for an imbalanced data set. There are several supervised classification models for imbalanced data, and we chose to compare some of the most used:

  • -

    Logistic regression: this is a model that predicts the probability of an observation belonging to a particular class, based on explanatory variables [13].

  • -

    Support vector machine (SVM): this is a model that separates observations according to their class. It can be useful for imbalanced data, as it allows greater emphasis to be placed on the minority class [14].

  • -

    Neural networks: These are mathematical models inspired by human brain. They are capable of processing complex data by using layers of artificial neurons to perform calculations and transformations on the input data [15].

  • -
    Random forest: this is a model that combines several decision trees to improve classification accuracy. It is particularly useful for imbalanced data, as it is less likely to over fit than other models [16]. In our case, we tried two other types of extension to this model:
    • Balanced random forests: These are used when the output classes are imbalanced. In an imbalanced data set, there is a prevalence of one class over the others. The model corrects it by using random subsampling of the majority class, or by oversampling the minority class to balance the classes.
    • Weighted random forests: These are used to solve problems where certain characteristics are more important than others in predicting the output variable. Weighted random forests give more weight to important characteristics and less weight to less important ones.
  • -
    XGBoost model (eXtreme Gradient Boosting): This is a set-based machine learning algorithm used for classification, regression and other predictive analysis tasks. It is an implementation of the gradient boosting algorithm that combines several weak decision tree models to form a stronger prediction model [17]. We also tested another extension of this model:
    • The Imbalanced-Xgboost extension: A new extension to the Xgboost model introduced in 2020 for imbalanced data with binary classification [18]. This method consists of giving more weight to the minority class and less weight to the majority class when training the model. Thus, the performance of the prediction is improved.

The use of variable selection methods during modeling can improve model performance. In our study, we used the RFE (Recursive feature elimination) method. This is a feature selection method that iteratively eliminates the least important features until an optimal number of features is reached.

In our study, we have used several methods to calculate the importance of variables. Some methods are provided by the models themselves, while other methods can be used in addition. We focus on two methods: The Permutation Feature Importance (PFI) method, which evaluates the importance of each feature by measuring the effect of its random permutation on model performance, and the SHapley Additive exPlanations (SHAP) method, which explains the predictions of Machine Learning models by assigning importance to each feature in the overall model prediction.

Stage 3: Evaluation.

In this stage, the model is evaluated according to the objectives and success criteria defined in the first phase. The results are interpreted and additional measures are defined. In our case, it was important use binary classification models for imbalanced data. When a class distribution is imbalanced, this can adversely affect current evaluation measures and lead to biased classification. The model may favor the majority class and perform poorly on the minority class. Therefore, it is important to use alternative evaluation measures to evaluate models with imbalanced data [19].

There are several alternative measures that can be used to evaluate models with imbalanced data [19]. The measures we have used are:

  • -

    G-means: a measure that takes into account the balance of classification performance between the majority and minority classes. It is calculated by taking the square root of the product of sensitivity and specificity.

  • -

    F-Balanced accuracy measure: a measure that combines precision and recall to give an overall assessment of model performance.

  • -

    ROC curve (Receiver Operating Characteristic): a graphical representation of the performance of a binary classifier. The ROC curve plots the rate of true positives (sensitivity) against the rate of false positives (1-specificity) for different threshold values.

  • -

    Precision-Recall (PR) curves: another graphical representation of binary classifier performance, plotting precision (positive predictive value) versus recall (sensitivity) for different threshold values. PR curves and ROC curves should be used together to obtain a more complete picture of model performance.

  • -

    Area under the ROC curve (AUC): a commonly used measure of overall classification performance. A higher AUC value indicates better classification performance, with a value of 1 indicating perfect classification performance and a value of 0.5 indicating random choice.

In order to have a reliable model, the data is divided into three sets (training, cross-validation and the test set used for model evaluation) at a percentage of 60%, 20%, 20% respectively. To be even more conservative, we withdrew a random sample (n = 60) with a representation of both classes. This sample was not used during the training of our model, but was used to test its performance on new data.

Stage 4: Deployment.

In this final stage, the model was implemented in a production environment, where it was used to solve real-life problems. This involves deployment planning, performance monitoring and ongoing model maintenance.

The CRISP-DM methodology is commonly used in business, but also in scientific research. However, it should be noted that most research studies using it do not include a deployment phase [10].

2.2. Employed tools

  • -

    Rstudio: is an integrated development environment (IDE) designed specifically for the R programming language. It offers a user-friendly interface for writing, debugging and testing R code, making it a popular tool among data scientists, statisticians and researchers working with R. So we used the R programming language, with the help of this tool, to prepare our databases.

  • -

    Jupyter Notebook: is an open source web application for creating and sharing documents containing interactive code, visualizations and text explanations. It is widely used in the scientific community for data analysis, mathematical modeling, machine learning and scientific research. We have used the Python programming language, thanks to this tool, for our data analysis and modeling.

  • -

    ArcGis (Kriging and mapping): is a GIS (Geographic Information System) software platform developed by Esri. It lets you create, manage, analyze and visualize geographic data through interactive maps and web applications. ArcGIS is a powerful tool for managing and analyzing geographic data, and is widely used in many industries for decision support and planning. This tool has been used to perform kriging interpolation.

3. Results and interpretation

3.1. Modeling results

To model our data, we implemented two distinct options:

  • -

    Option 1: we tested different models using only: reported animal cases of rabies between 2004 and 2021 and demographic and socio-economic data;

  • -

    Option 2: we integrated geo-spatial data (obtained through the ArcGIS kriging tool) into our analysis and assessed the impact of this new added data on model performance.

Option 1: We used the demographic and socio-economic data as model features, and rabies cases data as the target. We implemented two distinct scenarios:

  • Scenario 1: Binary classification (Presence/Absence).

The classification was based on the presence or absence of rabies. Table 1 compares the different models used.

Table 1.

Accuracy and F1-score of different models in binary classification.

Database Evaluation measures Models
Logistic regression Neural networks Random forest Xgboost
Scenario 1 Accuracy 70.1% 58.7% 73.3% 79.2%
F1-score 68% 36.7% 71% 78%

The XGBoost model performs best for this classification, followed by the “Random Forest” and “Logistic Regression” models. Although these results are acceptable in general, they do not reach the specific needs of our field problem. Indeed, a binary classification by presence or absence of rabies cases does not allow us to accurately characterize the risk level of a commune with only one case in 18 years.

This is why our next step was to test a three-class model (low, medium, high) to better define the level of risk associated with each commune.

  • Scenario 2: Classification into three classes.

We categorized our target variable according to the number of cases over an 18-year period, as follows: No risk (0 case), Medium risk (1 to 13 cases) and High risk (>13 cases). The Table 2 shows the best results obtained by various models.

Table 2.

Accuracy and F1-score of the different models for three-class classification.

Database Evaluation measures Models
Logistic regression Neural networks Random forest Xgboost
Scenario 2 Accuracy 64.5% 52.19% 69.3% 71.7%
F1-score 51.3% 22.8% 54.6% 59%

Once again, the Xgboost model showed the best results, but its F1 evaluation score remains low. Despite our efforts to optimize different models and methods, we may have reached the limit of what the available data allowed us to achieve. Indeed, some cases of rabies in free roaming dogs may be underestimated.

Option 2: The aim was to develop a predictive model that takes into account both geo-spatial data (based on the role of geographical proximity in the spread of the disease) and demographic, socioeconomic data.

We therefore considered integrating geo-spatial data obtained through kriging, as the proximity of communes plays a crucial role in the spread of the disease, given that dogs, the main vectors of rabies, move in search of feed to new territories.

For this matter, more reliable variables, we used only those communes where we were 100% sure of the presence of the disease (setting a threshold of 6 declared cases over an 18-year period) to perform our interpolation using Kriging.

Therefore, we evaluated the performance of several models, using the demographic and socioeconomic data as features and geo-spatial data associated with rabies cases data as the target variable. The Table 3 compares the performance of the different models used, as well as their ability to adapt to the introduction of new data.

Table 3.

Different Models performance (accuracy and F1-score) and their adaptability to the introduction of new data.

Databases Evaluation Measures Models
Logistic regression SVM Neural Networks Random Forest Balanced Random Forest Weighted Random Forest Xgboost Imbalanced
Xgboost
Option 2 Accuracy 76.68% 77.1% 76.68% 82.5% 78.4% 82.5% 84.3% 90.2%
F1-score 59.5% 67% 43.4% 71.5% 75% 71% 76% 90.5%
Withdrawn Data n = 60 Accuracy 70% 71.6% 70% 75% 70% 75% 81.6%
F1-score 62.5% 66.6% 64% 78.2% 65.3% 71.69 80.7%

We saw a clear improvement in our results with this new analysis. Most models showed good performances, with a relatively small drop-off when tested on new data. These encouraging results underline the correlation between the all types of data used, but also highlight the crucial importance of geo-spatial data.

The imbalanced Xgboost model proved to be the best and demonstrated its ability to adapt to new data. We will therefore look to achieve better results through variable selection and model optimization.

  • -

    Variable selection using the RFE (Recursive feature elimination) method [20]. The variables retained are as follows:

  • Volumetric index (in%);

  • Vulnerability rate (in%);

  • Bovin_average: average number of cattle in the commune;

  • Petit Rum: average number of small ruminants in the commune;

  • Access_Moym: average accessibility rate;

  • CDL Index: Communal Development Index;

  • Total_Population;

  • Feed: the sum of slaughterhouses and markets (souks) in the commune;

  • Type: Urban or rural Commune.

  • -

    The XGBoost model provides a variable (or feature) importance measure that indicates the contribution of each variable to the construction of the algorithm's decision trees. This measure is based on the reduction in error obtained by using each variable to divide the data into subgroups during tree construction (Fig. 8).

Fig. 8.

Fig. 8

Classification of selected variables by importance. This measure is based on the reduction in error obtained by using each variable to divide the data into subgroups during tree construction. It measures the performance that each variable adds or deducts to enable us to select the most important variable.

We used the PFI and SHAP methods to check the importance of all the remaining variables previously used (Fig. 9). We also tried to remove the variables with the lowest scores, and found that the model's performance dropped. This proved that all these variables have an impact on prediction and that it is important to keep them all.

Fig. 9.

Fig. 9

Importance of variables calculated by Shap and PFI methods. The use of two different methods Shapley additive exPlanations and permutation feature importance. These methods list the variables that the model considers important. This shows what variables played the biggest roles in the development of the model and have an impact on prediction.

Each of the selected variables has a valid justification for its importance. For instance, the variables “Bovin_average and ‘Small_Ruminant’ indicate the presence of species particularly susceptible to rabies, and doing so their density plays a crucial role in the risk of the disease spreading in a commune. In addition, the variables ‘Total_Population’, ‘CDL Index’, ‘Volumetric Index’ and ‘Vulnerability Rate’ highlight the fact that if the population is large and have a low awareness of health risks, combined with difficult living conditions, this increases the amount of waste and other feed sources available to free roaming dogs. For other aspects, the ‘Access_Moym’ variable reflects the commune's accessibility and its connectivity to road networks since we observe that dog packs follow the main roads to move around, due to the availability of feed sources. Ultimately, some facilities are an important part of the dogs” access to feed and therefore the variable “Feed” represented for the model as the sum of “Slaughterhouses” and “Market” variables.

To optimize our model, we used the “GridSearch” method. This involved defining a series of values for each parameter of the algorithm, then testing all possible combinations of these values to determine which produces the best performance results. We used this method to find optimal values for the key hyper-parameters as the Objective, the Evaluation metric, the Tree construction method, the Gamma, the Learning rate, the Maximum depth, the Number of estimators the Reg_alpha and Reg_lambda, and finaly the positive scaling weight. We used this method to find optimal values for the following hyper-parameters as below:

  • Objective: We defined this as “binary:logistic”, meaning that the model performs binary classification and uses a logistic cost function for learning.

  • Evaluation metric: The metric used to evaluate the model's performance. Here, it is defined as “AUC”, which measures the area under the ROC curve.

  • Tree construction method: The method used to construct the decision trees in the model.

  • Gamma: The regularization parameter of the decision tree. It controls the minimum reduction in loss required to make a further division of the tree. A higher value of gamma leads to stronger regularization.

  • Learning rate: Controls the size of weight updates at each learning stage. A higher learning rate value can speed up learning, but can also lead to over-fitting.

  • Maximum depth: This controls the complexity of the model and can help avoid overfitting.

  • Number of estimators: The number of decision trees to be built in the model.

  • Reg_alpha and Reg_lambda: The L1 and L2 regularization parameters, respectively. They control the penalty added to the variable weights to avoid overfitting.

  • Positive scaling weight: This can be used to help compensate for the effect of class imbalance in the data, by giving more weight to the minority sample. For our model it is calculated as a function of the ratio between classes.

The Table 4 shows the performance of the final model.

Table 4.

Final model performance and adaptability to new data before and after optimization.

Database Evaluation Measures Model: Imbalanced XGboost
Before Optimization After Optimization
Option 2 Accuracy 90.2% 92%
F1-score 90.5% 92%
Withdrawn Data n = 60 Accuracy 81.6% 86.6%
F1-score 80.7% 87.4%

By adjusting the hyperparameters to optimize and regularize the model, we were able to improve its performance, particularly when testing new data, which also enhanced its reliability.

3.2. Evaluation

The Table 5, Table 6 and Fig. 10, Fig. 11 show the performance of our final model according to the different evaluation measures chosen.

Table 5.

Accuracy and F1-score of the optimized Imbalanced-Xgboost model.

Accuracy 92%
F1-score 92%

Table 6.

Different evaluation measures for each of the two classes.

Class Metric
Precision Recall F1-score G-means
0 (low risk) 93% 91% 92% 0.919
1 (High risk) 91% 93% 92% 0.919

Fig. 10.

Fig. 10

ROC curve for the different model sets (Training, Validation, Test) and measurement of the area under the curve AUC. These AUC values are 1 for the training set, 0.97 for the validation set and 0.97 for the test set, given that a value of 1 represents perfect performance and that a random model would have a value of 0.5.

Fig. 11.

Fig. 11

Precision-Recall curve for the optimized Imbalanced-Xgboost model. As we can observe, during the evaluation phase of the model, the area under the precision-recall curve shows a good performance of the model, the larger this area is, the model's performance is increased.

Taking into account the unbalanced nature of the data we have used, our model cannot be evaluated strictly by accuracy and F1 score. We therefore used other more reliable evaluation measures (Table 6).

The measures were calculated for each class to determine whether the model works equally well for the minority class (class 1). The G-means measure is determined by multiplying precision and recall, and taking the square root of the result. A G-means value of 1 indicates perfect model performance.

The ROC curve was also used to evaluate the model (Fig. 10). A perfect classifier would have an ROC curve that passes through the upper left corner of the graph, indicating high sensitivity and specificity for all threshold values [19]. In our case, the curves for the different sets indicate good model performance. The area under the curve (AUC) values are 1 for the training set, 0.97 for the validation set and 0.97 for the test set, given that a value of 1 represents perfect performance and that a random model would have a value of 0.5.

The fact that there were no major differences between the AUC values of the 3 sets indicates that the risk of overlearning the model is low. When evaluating our model, an important performance measure is the area under the precision-recall curve. The larger this area, the better the performance of our model (Fig. 11).

Finally, we tested the performance of our model on the previously withdrawn data. We note a performance drop of <5.7%, which is a very good indicator (Table 7) and prove that there is no overfitting.

Table 7.

Accuracy and F1-score of our model on the new data.

Accuracy 86.6%
F1-score 87.4%

After evaluating the model and noting its high and consistent performance on different evaluation measures, we are able to state that the results obtained are reliable and can be used to predict at-risk areas. These results indicate that the precision of the prediction of at-risk communes is 82%, and that we could have identified 93% of communes actually at risk (recall) (Table 8).

Table 8.

Different evaluation measures of our model for the two classes on new data.

Class Metric
Precision Recall F1-score G-means
Low risk 92% 80% 86% 0.853
High risk 82% 93% 87.5% 0.87

3.3. Prediction

After confirming the performance of our model, we carried out the prediction on the 160 communes (initially withdrawn) which were predicted by interpolation using the kriging tool. Only 145 communes were retained due to missing data. Our model, evaluated on new data (60 communes), gave us a precision of 82% and a recall of 93%:

  • Precision = TP / (TP + FP) = 82%

  • Recall = TP / (TP + FN) = 93%

  • TP: True Positive; FP: False Positive; FN: False Negative

This means that of the 145 communes assessed, 69 were predicted to be at risk. With a Precision of 82%, this means that 57 of these 69 communes are actually at risk (true positives). A recall of 93% means that of the 61 communes actually at risk (true positives and false negatives), we were able to detect 57, the other communes that were predicted as not at risk are probably close to contaminated areas and have probably had cases of the disease, but our model considers them not to be at high risk.

Our model predicted that 69 of the 145 communes to be predicted were high-risk and 76 communes were low-risk (Table 9). The geographical distribution of these communes doesn't provide much information (Fig. 12), so mapped all 1546 communes of Morocco, including the 160 we withdrawn at the outset. The 15 communes that were not predicted due to missing values were considered high-risk.

Table 9.

Prediction results for communes predicted initially by the kriging tool.

Class Prediction
Low risk 76 communes
High risk 69 communes

Fig. 12.

Fig. 12

Mapping of the 145 communes predicted by our two-class model. The model predicted that 69 of the 145 communes to be predicted were high-risk and 76 communes were low-risk.

When we displayed our forecasts on the map, we noticed that high-risk areas were grouped in clusters (Fig. 13), mainly on the coasts. The southern communes were not concerned because no rabies cases were reported in these communes during the study period. In addition, these communes have a low population density and climatic conditions that are unsuitable for the spread of rabies.

Fig. 13.

Fig. 13

Map of risk communes, divided into two classes. The map shows high-risk areas grouped into clusters. No rabies cases were reported in the southern communes of the country during the study period (low population density and climatic conditions are unsuitable for the spread of rabies).

4. Discussion

This study focusses on data analysis and modeling related to animal rabies in Morocco. The data were organized into two categories: (1) health data: reported animal cases of rabies between 2004 and 2021 and data obtained through the ArcGIS kriging tool (geospatial data); (2) demographic and socio-economic data which were used to test different classification models (Fig. 14).

Fig. 14.

Fig. 14

Summary diagram of the workflow used in this study.

Results showed that the XGBoost model was the most efficient for binary classification (presence or absence of rabies), but did not fully reach the specific objectives of the study, as it did not precisely characterize the rabies risk level associated to each commune. Consequently, a three-class model was tested (high, medium and low risk) and the results showed that the XGBoost model still performed best, but with a relatively low F1 score (59%).

To overcome the underestimation of rabies cases, particularly among free roaming dogs, with regard to the disease epidemiology, geo-spatial data obtained by kriging tool were incorporated into the analysis in order to compensate the lack of some health data [9]. The obtained results led to a significant improvement in model performance, highlighting the critical importance of geo-spatial data. The “Imbalanced XGBoost” model showed better performance, with a high capacity to adapt to new data.

To improve the performance and reliability of our rabies propagation model, we have performed variables selection, as first step, followed by the optimization and regularization steps. This selection allowed us to filter the 32 initial variables down to the following nine quantitative variables: volumetric index, vulnerability rate, average number of cattle in the commune (Bovin_Moym), number of small ruminants (PetitRum), average accessibility rate (Access_Moym), commune development index (CDL Index), total population, sum of slaughterhouses, and markets present in the commune (Feed) and the qualitative variable type of area (commune rural or urban).

Each selected variable is important for understanding the risk of disease spread in a given commune. For instence, the variables “Bovin-Moym” and “PetitRum” are indicators of the presence of susceptible species to rabies, while “Access_Moym” reflects the commune's connectivity to road networks as a factor that influences the movements of free roaming dogs [21]. The variables “Total_Population”, “CDL Index”, “Volumetric Index” and “Vulnerability Rate” provide information on living conditions in the commune, such as population density and level of development, which can have significant impact on disease incidence. Finally, the “Feed” variable is the sum of the “Slaughterhouses” and “markets (Souk)” variables, which represent an important part of free roaming dogs' access to food [22].

During the evaluation, the model showed high and consistent performance on various evaluation measures. The results obtained testify to the model's effectiveness in classifying high-risk communes in Morocco. These promising results can serve as a basis for further studies to further improve the modeling of disease spread and help develop more targeted prevention policies.

After validating the performance of our model, we proceeded to predict the communes that had been withdrawn (145 communes). Indeed, taking into account the precision and recall values obtained. From a total of 145 communes, our model succeeded in predicting 57 communes as being truly at risk from 69 predicted as at risk communes.

After displaying our projections on the map, we found that high-risk areas were concentrated in clusters. The most represented communes are on the Atlantic coast, followed by Marrakech and the surrounding area. The third cluster is located in the Middle Atlas region, while the last cluster is in the Eastern region, particularly around the city of Oujda.

According to the field observations and risk mapping, the reasons for the high risk of rabies vary from a cluster to another. This also reinforces our choice of variables; as all the factors mentioned are involved in the spread of rabies, while their importance varies according to each cluster.

In fact, we noticed that on the Atlantic coast we have areas with the most infrastructure and major construction sites. This validates the results obtained by [9], who found a strong association between the presence of building sites and infrastructure projects (railroads, freeways, etc.) and dog rabies [21].

In the case of Marrakech, a tourist city that has a large number of visitors every year, human density and dog's accessibility to food are the factors that play the most important role in the spread of rabies.

As for the Middle Atlas, the area has a fairly low development index such as lack of infrastructure to limit access to food for free roaming dogs.

Finally, the Eastern region is characterized by a high presence of ruminants, particularly sheep and goats flocks, conducted in pastoral system, so that a single dog carrying the virus can cause considerable damage.

Given that the model's ability to discriminate and predict the spread of rabies in Morocco depends on the volume and quality of the data, the results showed an improvement in model performance when more complete data were incorporated into the analysis. Indeed, the study identified disease niches, particularly in the same areas previously identified by interpolative analysis tool (Kriging) [9], providing a relevant validation of our XGBoost model for assessing rabies risk levels.

The risk maps obtained by this model also revealed an extension of the spatial spectrum of the disease in comparison with the cartographic approximations provided by interpolation tools. The mapping obtained using the XGBoost tool revealed more calibrated and precise spatial distributions of the disease than those elaborated in a previous study using the Kriging tool [9].

The results of the study also suggest that understanding the distribution of animal rabies in Morocco, as deduced from the XGBoost data analysis, initially relates to the accurate identification and improvement of geographical models of disease risk. Given that, each of the geographical areas of rabies distribution is conditioned by a number of risk factors. The geographical clusters of rabies identified by artificial intelligence methods therefore encourage the deployment of control measures adapted to each geographical risk typology.

However, it has proved necessary to continue strengthening state-of-the-art methods for analyzing health data, so that the resulting distribution mapping can be significantly improved and the spatial analysis more precise. New technologies are constantly evolving, and it is essential to use this evolution to optimize disease control and its subsequent costs.

We also demonstrated the importance of using new technologies in the control disease strategy. The results of the study showed that conventional analysis methods were not sufficiently effective for accurate analysis of the disease's risk and distribution profiles. Consequently, it has become crucial to use innovative tools such as geo-spatial mapping and artificial intelligence to improve the modeling of rabies spread.

A better understanding and assessment of the complexity of the risk factors associated with the occurrence and spread of the disease will lead to better management, and therefore to a reduction in the risk of disease occurrence. The integration of artificial intelligence as a tool for epidemiological analysis of rabies determinants fully demonstrated its interest and relevance in the present study, insofar as it provided clarifications and improvements in the spatial pattern of canine and animal rabies in Morocco. In addition, the use of artificial intelligence will enable a better understanding of the disease's propagation patterns and obviously the early detection of high-risk areas.

In general terms, it is widely believed that AI can significantly enhance animal disease management through innovative approaches such as predictive modeling [7]. However, AI-based methods can face several limitations, primarily related to data availability. In our case, incomplete field data and under-reporting of animal rabies cases, especially among free-roaming dogs, pose significant challenges. In addition, missing demographic and socio-economic data related to some communes further complicate the analysis. Furthermore, the “feed” variable was limited to data from slaughterhouses and markets, neglecting other sources such as public landfills and household garbage where free-roaming dogs may also feed. These factors highlight important limitations in our study.

In an endemic health context, the identification of risk profiles that predispose dog population or livestock to rabies contamination is necessary to reinforce control and prevention measures and thus limit the incidence of the disease. The ultimate aim of our research is to provide animal health managers and decision-makers with decision-support tools incorporating these risk assessment tools at different levels (data collection, analysis, modeling and decision-support) to better deploy and implement the concepts of operational epidemiology, and thus better target at-risk areas and guide disease surveillance and control actions.

5. Conclusion

This study implements various innovative methods and technologies to analyze socio-economic and health data, in order to detect areas at high risk of disease transmission, as well as the most relevant and closely related variables. This analysis provides a better understanding of the correlation between the distribution of rabies in Morocco and human activities at community level. The success of any rabies control strategy depends on understanding of the disease associated risk factors, the rigorous implementation of the selected measures options, and consideration of the epidemiological situation. Consequently, it will be possible to design and implement effective intervention measures if these relationships are clearly identified.

Significant outcomes were obtained: Firstly, we have proved the strong correlation between health data, socio-economic data, human activity and geographical location. Additionally, we were able to classify our communes according to their respective risk levels and we identified as well 399 of Morocco's 1546 communes as of high risk, representing 25.8% of the areas initially targeted. This information shows us clearly where the intervention might be most needed to control the disease.

Moreover, we've created spatially informed risk matrices accompanied by detailed explanations of the underlying integrated assessments. This visualization enhances the comprehensibility of the data for decision-makers. In fact, we have been able to determine which communes should be targeted first, enabling us to adapt control actions, better targeting the most at risk areas and optimizing resources.

Lastly, the artificial intelligence technology and geo-spatial tool we have explored provide scope for developing epidemiological analysis tools, which better serve the principle of evidence-based decision-making.

Funding statement

This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.

CRediT authorship contribution statement

Ilham Ahamjik: Writing – review & editing, Writing – original draft, Visualization, Validation, Software, Project administration, Methodology, Investigation, Formal analysis, Data curation, Conceptualization. Ayman Agbani: Writing – review & editing, Writing – original draft, Visualization, Validation, Software, Methodology, Formal analysis. Mounia Abik: Writing – review & editing, Writing – original draft, Visualization, Validation, Supervision, Project administration, Methodology, Formal analysis. Mounir Khayli: Software, Formal analysis. Naima Galzim: Software, Formal analysis. Jaouad Berrada: Writing – review & editing, Writing – original draft, Visualization, Validation, Supervision, Project administration, Methodology, Formal analysis. Mohammed Bouslikhane: Writing – review & editing, Writing – original draft, Visualization, Validation, Supervision, Project administration, Methodology, Formal analysis.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Data availability

The authors do not have permission to share data.

References

  • 1.World Organisation for Animal Health Rabies. 2024. https://www.woah.org/en/disease/rabies/ (Accessed: march 18th 2022)
  • 2.Bourhy H., Reynes J.M., Dunham E.J., Dacheux L., Larrous F., Huong V.T.Q., Xu G., Yan J., Miranda M.E.G., Holmes E.C. The origin and phylogeography of dog rabies virus. J. Gen. Virol. 2008;89:2673–2681. doi: 10.1099/vir.0.2008/003913-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.El Harrak M. Compendium of the OIE Global Conference on Rabies Control; 2008. Epidemiological situation of rabies in Morocco and control strategy. Rabies control – towards sustainable prevention at the source.https://www.woah.org/en/produit/rabies-control-towards-sustainable-prevention-at-the-source/ [Google Scholar]
  • 4.Khayli M., Lhor Y., Derkaoui S., Lezaar Y., Elharrak M., Sikly L., Bouslikhane M. Determinants of canine rabies in Morocco: how to make pertinent deductions for control? Epidemiol Open J. 2019;4:1–11. doi: 10.17140/EPOJ-4-113. [DOI] [Google Scholar]
  • 5.Zinsstag J., Schelling E., Waltner-Toews D., Whittaker M.A., Tanner M. 2020. One health, Une Seule Santé. Théorie et pratique des approches intégrées de la santé Éditions Quæ. [DOI] [Google Scholar]
  • 6.Léchenne M., Miranda M.E., Zinsstag J. One health. Éditions Quæ. 2020. Chapitre 16 - Lutte intégrée contre la rage; pp. 245–262.https://books.openedition.org/quae/36160 [Google Scholar]
  • 7.Yongjun Xu, Liu Xin, Cao Xin, Huang Changping, Liu Enke, Qian Sen, Liu Xingchen, Yanjun Wu, Dong Fengliang, Qiu Cheng-Wei, Qiu Junjun, Hua Keqin, Wentao Su, Jian Wu, Huiyu Xu, Han Yong, Chenguang Fu, Yin Zhigang, Liu Miao, Roepman Ronald, Dietmann Sabine, Virta Marko, Kengara Fredrick, Zhang Ze, Zhang Lifu, Zhao Taolan, Dai Ji, Yang Jialiang, Lan Liang, Luo Ming, Liu Zhaofeng, An Tao, Zhang Bin, He Xiao, Cong Shan, Liu Xiaohong, Zhang Wei, Lewis James P., Tiedje James M., Wang Qi, An Zhulin, Wang Fei, Zhang Libo, Huang Tao, Chuan Lu, Cai Zhipeng, Wang Fang, Zhang Jiabao. Artificial intelligence: a powerful paradigm for scientific research. The Innovation. 2021;2 doi: 10.1016/j.xinn.2021.100179. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Rong G., Mendez A., Bou Assi E., Zhao B., Sawan M. Artificial intelligence in healthcare: review and prediction case studies. Engineering. 2020;6:291–301. doi: 10.1016/j.eng.2019.08.015. [DOI] [Google Scholar]
  • 9.Khayli M., Lhor Y., Bengoumi M., Zro K., El Harrak M., Bakkouri A., Akrim M., Yaagoubi R., El Berbri I., Kichou F., Berrada J., Bouslikhane M. Using geostatistics to better understand the epidemiology of animal rabies in Morocco: what is the contribution of the predictive value? Heliyon. 2021;7 doi: 10.1016/j.heliyon.2021.e06019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Schröer C., Kruse F., Gómez J.M. A systematic literature review on applying CRISP-DM process model. Procedia Comput. Sci. 2021;181:526–534. doi: 10.1016/j.procs.2021.01.199. [DOI] [Google Scholar]
  • 11.Sigwadi Loshani, Muka Mr. J, Breytenbach J. University of the Western Cape; 2021. The traversal data analysis and the predictive analyis of customer behaviour in a convenience store chain. [DOI] [Google Scholar]
  • 12.Kang M., Tian J. John Wiley & Sons, Ltd; 2018. Machine Learning: Data Pre-Processing in Prognostics and Health Management of Electronics; pp. 111–130. [DOI] [Google Scholar]
  • 13.Hand D.J. Vol. 79. 2011. Logistic regression models by Joseph M. Hilbe. International Statistical Review; pp. 287–288. [DOI] [Google Scholar]
  • 14.Suthaharan S. vol. 36. Springer; Boston, MA: 2016. Support Vector Machine, in Machine Learning Models and Algorithms for Big Data Classification: Thinking with Examples for Effective Learning; pp. 207–235. Integrated Series in Information Systems. [DOI] [Google Scholar]
  • 15.Domany E., Van Hemmen J.L., Schulten K. Springer Science & Business Media; 2012. Models of Neural Networks I. [Google Scholar]
  • 16.Steven J., Rigatti Random Forest. J. Insur. Med. 2017;47:31–39. doi: 10.17849/insm-47-01-31-39.1. [DOI] [PubMed] [Google Scholar]
  • 17.Quinto B. 2020. Next-generation machine learning with spark: covers XGBoost. LightGBM, spark NLP, distributed deep learning with Keras, and more. Apress. [DOI] [Google Scholar]
  • 18.Wang Chen, Deng Chengyuan, Wang Suzhen. Imbalance-XGBoost: leveraging weighted and focal losses for binary label-imbalanced classification with XGBoost. Pattern Recogn. Lett. 2020;136:190–197. doi: 10.1016/j.patrec.2020.05.035. [DOI] [Google Scholar]
  • 19.Bekkar M., Djema H.K., Alitouche T.A. Evaluation measures for models assessment over imbalanced data sets. J. Inform. Engineer. and Applications. 2013;3:27–38. https://www.iiste.org/Journals/index.php/JIEA/article/view/7633/8051 [Google Scholar]
  • 20.Garg P., Singh S.N. 11th International Conference on Cloud Computing, Data Science & Engineering (Confluence), Noida, India. 2021. Analysis of ensemble learning models for identifying spam over social networks using recursive feature elimination; pp. 713–718. [DOI] [Google Scholar]
  • 21.Talbi C., Holmes EC., de Benedictis P., Faye O., Nakouné E., Gamatié D., Diarra A., Elmamy BO., Sow A., Adjogoua EV., Sangare O., Dundon WG., Capua I., Sall AA., Bourhy H. Evolutionary history and dynamics of dog rabies virus in western and Central Africa. J. Gen. Virol. 90 (2009) pp. 783–791 doi: 10.1099/vir.0.007765-0. [DOI] [PubMed]
  • 22.Office National de Sécurité Sanitaire des Produits Alimentaires La rage animale. 2024. https://www.onssa.gov.ma/sante-animale-dsa/programme-de-prophylaxie/rage-animale/ (Accessed: January 25th 2022)

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The authors do not have permission to share data.


Articles from One Health are provided here courtesy of Elsevier

RESOURCES