Skip to main content
BMC Public Health logoLink to BMC Public Health
. 2025 Jan 4;25:34. doi: 10.1186/s12889-024-21187-0

Forecasting cardiovascular disease mortality using artificial neural networks in Sindh, Pakistan

Moiz Qureshi 1,4, Khushboo Ishaq 2, Muhammad Daniyal 3, Hasnain Iftikhar 4,6,, Mohd Ziaur Rehman 5, S A Atif Salar 6
PMCID: PMC11699765  PMID: 39754102

Abstract

Cardiovascular disease (CVD) is a leading cause of death and disability worldwide, and its incidence and prevalence are increasing in many countries. Modeling of CVD plays a crucial role in understanding the trend of CVD death cases, evaluating the effectiveness of interventions, and predicting future disease trends. This study aims to investigate the modeling and forecasting of CVD mortality, specifically in the Sindh province of Pakistan. The civil hospital in the Nawabshah area of Sindh province, Pakistan, provided the data set used in this study. It is a time series dataset with actual cardiovascular disease (CVD) mortality cases from 1999 to 2021 included. This study analyzes and forecasts the CVD deaths in the Sindh province of Pakistan using classical time series models, including Naïve, Holt-Winters, and Simple Exponential Smoothing (SES), which have been adopted and compared with a machine learning approach called the Artificial Neural Network Auto-Regressive (ANNAR) model. The performance of both the classical time series models and the ANNAR model has been evaluated using key performance indicators such as Root Mean Square Deviation Error, Mean Absolute Error (MAE), and Mean Absolute Percentage Error (MAPE). After comparing the results, it was found that the ANNAR model outperformed all the selected models, demonstrating its effectiveness in predicting CVD mortality and quantifying future disease burden in the Sindh province of Pakistan. The study concludes that the ANNAR model is the best-selected model among the competing models for predicting CVD mortality in the Sindh province. This model provides valuable insights into the impact of interventions aimed at reducing CVD and can assist in formulating health policies and allocating economic resources. By accurately forecasting CVD mortality, policymakers can make informed decisions to address this public health issue effectively.

Keywords: Cardiovascular disease, Analyzing and forecasting, Mortality, Time series models, Artificial neural network approach

Introduction

Cardiovascular disease (CVD) is a broad term encompassing various heart and blood vessel conditions. It is a leading cause of death worldwide, accounting for an estimated 17.9 million deaths each year, according to the World Health Organization (WHO). The most common types of CVD include coronary artery disease, stroke, heart failure, and peripheral artery disease. These conditions can develop over time due to a combination of factors, including high blood pressure, high cholesterol, smoking, diabetes, obesity, physical inactivity, and a family history of heart disease (https://www.who.int). In Sindh, CVDs are a serious public health issue. Age, gender, obesity, hypertension, hyperglycemia, and hyperlipidemia are the main risk factors for cardiovascular disease. Nonetheless, insufficient is known about the frequency and risk factors related to Hyderabad's population in both urban and rural areas [1]. The burden of CVD is characterized by its immense prevalence, as it remains the leading cause of death globally. This encompasses the human toll in terms of lives lost and affected and the economic burden of managing and treating these conditions. In addition, early diagnosis and treatment are crucial for managing CVD and reducing the risk of serious complications. Besides unhealthy lifestyle choices, other risk factors for heart disease include smoking, alcohol, high cholesterol levels, obesity, high blood pressure, and diabetes [2]. According to the WHO, CVDs are a group of disorders of the heart and blood vessels. Collectively, CVD accounted for an estimated 17.9 million deaths worldwide in 2019. Of these, representing 32% of all global deaths. According to the most recent WHO data on the heart attack ratio in Pakistan, 240,720 people died from coronary heart disease in Pakistan in 2020, accounting for 16.49 percent of all fatalities. The intensity and ratio of deaths are increasing, which is dangerous for public health in Pakistan and Southeast Asia (https://www.who.int/data/gho/data/countries) [3, 4].

Modeling CVD cases is a multidisciplinary endeavor that leverages various approaches, from epidemiology and statistics to cutting-edge machine learning and artificial intelligence techniques. These models are instrumental in unraveling the intricate web of factors contributing to CVD, including genetics, lifestyle choices, environmental factors, and healthcare interventions. In this era of data-driven healthcare, CVD modeling approaches are at the forefront of efforts to predict, prevent, and manage this pervasive and life-altering disease. They offer valuable insights into disease trends, risk factors, and the effectiveness of interventions [5, 6]. Several models have been applied in the literature to model and predict the death rates due to CVD. Three distinct statistical models were employed to forecast heart disease, specifically utilizing a Support Vector Machine (SVM), Decision Tree (DT), and Logistic Regression model (LR) [7]. The investigation revealed that through the application of the 'C-Rule' and employing various combinations, it is possible to improve predictive accuracy [8]. A time series model was introduced, proposing a novel approach known as the combined reinforcement multitask progressive time series model for CHD prediction [9]. The findings indicated that deep reinforcement learning (DRL) pre-training and multitasking exhibited superior performance in CHD prediction. In another study, five machine learning models were harnessed to predict the daily admissions for cardiovascular disease (CVD) [10]. After subjecting them to key performance indicators for comparison, it was evident that the Random Forest (RF) model surpassed its peers in forecasting daily CVD admissions. To further enhance the forecasting of heart disease, a hybrid time series modeling approach was applied, combining the Support Vector Machine model (SVM) and Random Forest (RF) [11]. This hybrid model significantly boosted forecasting efficiency by up to 88.7%.

Moreover, the ARIMA model was deployed to predict the mortality rate of CVD patients, and its efficiency in this context was duly noted [12]. Researchers employed the Lee-Carter and Bayesian Age Period Cohort (BAPC) models for broader mortality trend projections, extending their forecasts to 2030 in England [13]. In the context of disease classification and forecasting, various machine learning models were applied, particularly for Romania's International Classification of Disease (ICT) [14], with the findings emphasizing the models' significance in this predictive task. The authors [15] made a comparative analysis based on linear and non-linear time series models to predict the stay at the ICU using a sample size of 6064. The data was thoughtfully partitioned into training and testing sets, ultimately revealing that the Gaussian Naive Bayes and Logistic Regression hybrid model (GB + LR) exhibited superior performance in predicting the overall survival of cardiac patients. A comparative evaluation to predict CVD based on machine learning and conventional logistics regression is examined [16]. Results showed that the machine learning models predict more accurately than the classical method. The authors [17] also used time series regression in epidemiology studies to investigate the short-term association. Hypertension is considered one of the significant risk factors in developing CVD, as studied by [18]. The study [19] aimed to explore the relation between the risk factors of myocardial infarction (MI), and to achieve this end, binary logistic regression is applied. The study found that gender, family history, and other related variables are statistically significant for MI. A survey of risk factors [20] found that multiple risk factors are statistically significant in developing CVD in Pakistan.

Moreover, the authors [21] conducted a cross-sectional study at NICVD on CVD knowledge. Logistic regression [22] is used and compared with machine learning models in predicting chronic disease. For further details on these studies, interested readers are encouraged to refe`r to the respective citations [23, 24].

The ARIMA and Seasonal-ARIMA are time series forecasting models used in statistics and econometrics to analyze and predict patterns in time-dependent data. ARIMA is significant because it provides a flexible framework for modeling and forecasting time series data, making it a valuable tool. At the same time, SARIMA allows for more accurate modeling and forecasting of time series data that exhibit both short-term fluctuations and longer-term seasonal patterns. Time series models, exemplified by ARIMA and SARIMA, find widespread utility in analyzing CVD incidence data and making short-term predictions. ARIMA models are the best-known model for time series forecasting and have been used by many researchers to predict infectious diseases with characteristic seasonal outbreaks [25, 26].

Almosova et al., [27] indicate the superiority of the machine learning model over than existing classical model in forecasting inflation variables. However, the studies in [28, 29] indicates the superiority of the ANNAR model over the classical approaches in modeling and forecasting the death of COVID-19 patients. Similarly, the article [30] has proposed ANNAR-based ensemble models for influenza incidents. The study found that the NNAR-based model results in the lowest accuracy error.

A study in Shandong, China, harnessed SARIMA modeling to aptly capture the seasonal and trend patterns in stroke incidence data, showcasing the model's ability to characterize such temporal dynamics effectively [31]. Given the shortcomings of ARIMA models, there is increasing interest in using ANN models for epidemiological time series forecasting [32] because these models account for nonlinearities in the data. Machine learning models, including artificial neural networks (ANNs) and support vector machines (SVMs), have emerged as valuable tools for CVD prediction. An A N-based approach, implemented in Shanghai, China, demonstrated superior predictive accuracy for stroke incidence compared to conventional time-series models [33].

This study aims to model and forecast annual mortality rates for CVD using various stochastic time series models by comparing the conventional linear and non-linear machine learning models. The paper is structured as follows: First, we provide a review of related literature review and previous work on CVD prediction. Furthermore, we provide details related to data description and methodology, followed by an overview of the results and their interpretation. Finally, we conclude with policy recommendations for future research direction.

Materials and methods

The dataset in this study was collected from the Civil Hospital in the Nawabshah district of Sindh province, Pakistan. It includes actual cases of CVD related deaths that occurred between 1999 and 2021. A solid time series dataset that covers trends and patterns in CVD mortality over more than two decades is provided by this extensive collection, which contains yearly data. A thorough examination of the evolution of CVD-related mortality in response to numerous factors, including alterations in healthcare infrastructure, public health campaigns, lifestyle adjustments, and socioeconomic advancements in the area, is made possible by the dataset's wide temporal scope. Through the utilization of this abundant dataset, scholars can acquire a significant understanding of the epidemiology of cardiovascular disorders, pinpoint plausible risk factors, and formulate focused treatments aimed at alleviating the prevalence of CVD in Sindh province. This data set is free from any missing value and used with actual observations. Also, the ethical approval of this data set is received from the hospital administration.

Statistical analysis

The data has been collected from Nawabshah, Sindh, Pakistan Civil Hospital, spanning the years 1999 to 2021, focusing on the number of people affected by CVD. To understand this data, which changes over time, a specific approach known as "time series analysis" is being employed. The analysis process involves several stages. Initially, the data will be subjected to a descriptive analysis to identify patterns and essential characteristics. This lays the foundation for subsequent phases. Following this, the data will be visualized by creating time series plots, enabling the observation of trends and changes in CVD deaths over the years. These visual representations help identify patterns such as seasonality and anomalies. Various time series models will be applied to make meaningful predictions and forecasts, each with its unique approach. These models include the straightforward "Naive" model, the more sophisticated "Simple Exponential Smoothing (SES)" method, "Holt's Linear Exponential Smoothing," and a highly advanced ANNAR Model." These models allow the data to be explored, patterns to be captured, and educated predictions about future trends in heart-related health issues to be made. The flow chart is given below for the data processing.graphic file with name 12889_2024_21187_Figa_HTML.jpg

Data description

A time series plot was constructed using a graphical representation of data points collected over a period, then used to analyze and visualize changes in data over time. Figure 1 shows the visual display of the yearly time series of deaths from cardiovascular disease. Inspection of the time series plot in Fig. 1 suggests an increasing and decreasing trend.

Fig. 1.

Fig. 1

Time series of Yearly death cases of CVD

The main focus of applying these models is to capture the data-generating process of the series. Both classical time series and machine learning models focused on short-term forecasting and then provided a comparison by indicator testing. The data was divided into 80% training (1999–2016) and 20% (2017–2021) testing for validation. The summary statistics are shown in Table 1 that the maximum number of death cases of CVD in Sindh was observed to be 107 in 1999, which continued to rise due to several factors until 2018 when the deaths rose to 408 cases. The average number of deaths was 236, with a median of 231. Figure 3 illustrates the residual diagnostic plots, which play a vital role in time series analysis as they help evaluate the adequacy of the time series model. Residuals in a time series model represent the discrepancies between observed values and the predicted values generated by the model. These plots are employed to visually assess whether the residuals adhere to certain assumptions: normal distribution, homoscedasticity (constant variance), and independence (lack of autocorrelation). The Histogram of residuals is used to verify the normal distribution of residuals. Additionally, the R function checkresiduals() is utilized to perform these diagnostic checks, generating a time plot, an ACF plot, a histogram of the residuals, and a normal curve.

Table 1.

Summary statistics of CVD death cases

Minimum 107.0
Maximum 408.0
1st Quartile 181.0
3rd Quartile 268.0
Median 231.0
Mean 236.6
Skewness 0.42
Kurtosis −0.79

Fig. 3.

Fig. 3

Residual Diagnostics of CVD death cases for ANNAR, SES, Holt, and Naïve

In the fields of time series analysis and artificial intelligence, greater caution must be taken when working with small sample sets to ensure robustness and prevent overfitting. Below the methods are mentioned that highlight their use and robustness.

Naïve method

In time series analysis, a naive method is the simplest forecasting approach where the next value in the time series is predicted to be equal to the current value. It can be expressed using the following equation:

y^(t+1)=y(t) 1

where y^(t+1) represents the predicted value for the next time period, and y(t) represents the actual value for the current time period. The naive method assumes that the time series is stationary and there are no trends or seasonal patterns in the data. It is a useful benchmark for evaluating the performance of more advanced forecasting models. The Naïve method makes minimal assumptions about the data and Less likely to overfit due to its simplicity [34, 35].

Holt-winter exponential smoothing method

Holt-Winters forecasting, also known as triple exponential smoothing, is a popular time series forecasting method that uses exponential smoothing to capture trends and seasonality in the data. This method can handle small sample sizes effectively, but careful tuning of parameters is required to avoid overfitting [36]. The method involves using three smoothing equations, one for the level, one for the trend, and one for the seasonality, to produce a forecast.

The equations are:

Levelequation:Tt=αYt+(1-α)(Lt-1+Tt-1) 2
Trendequation:Tt=β(Lt-Lt-1)+(1-β)Tt-1 3
Seasonalequation:St=γ(Yt-Lt)+(1-γ)St-m 4
Forecastequation:Ft+k=Lt+kTt+St-m+1+k 5

where: Yt = the actual value at time t, Lt = the level at time t, Tt = the trend at time t,St = the seasonal component at time t, m = the number of seasons in a year α, β, and γ are smoothing parameters between 0 and 1, which control the amount of smoothing applied to each component. In forecast equation, Ft represents the forecast for period t and k is the lag parameter or time shift i.e. how many periods ahead the forecast is made[37].

By estimating the smoothing parameters and applying these equations, a Holt-Winters forecast can be generated for future time periods.

Artificial Neural Network Autoregressive (ANNAR) model

Neural network models are a type of machine learning model inspired by the structure and function of the human brain. They are designed to recognize patterns in data and make predictions or decisions based on that data. Neural networks consist of layers of interconnected nodes, called artificial neurons, inspired by the structure of neurons in the human brain. Each neuron receives input from other neurons, processes that input, and passes the result on to other neurons in the next layer.

Conventional time series models, such as exponential smoothing or ARIMA, assume that the data have linear connections. Nonetheless, a lot of time series from the real world show intricate, non-linear patterns. ANNAR models are good at capturing these complex interactions because of their non-linear activation functions [38]. ANNAR models are very adaptable and have a broad variety of function approximations. With sufficient data and processing power, they can model any underlying process thanks to their universal approximation capacity. ANNAR models can learn and identify patterns in the data without explicit specification of the model form [39].

Neural network models have been used in a variety of applications, such as image and speech recognition, natural language processing, and game playing, amongst other uses. ANNAR models have been employed to predict disease outbreaks, patient admissions, and other health metrics. Their effectiveness in handling complex, multi-factorial data makes them suitable for these applications [40, 41]. They are a powerful tool for solving complex problems and have achieved state-of-the-art performance on many tasks [42, 43]. This network methodology permits to model of any linear and non-linear phenomena. Neural network autoregressive models are a type of model that is based on a simple neuronal structure that is organized in layers. These neural networks are classified further into two categories. The first category consists of the simplest neural network and the second category is a complex neural network. No hidden layer is involved in a simple neural network, while in a complex network, more than one hidden layer is used. In these neural networks, different methods are used to fit the data. The most commonly used procedure to fit the data is the feed-forward method. The graph of the feed-forward method is given in Fig. 2. The structure of a feed-forward network is composed of three parts, namely, the input layer that is used to process the observation, the hidden layer that is used to weigh these observations, and the output layer which is the gateway of the result. The hidden layer is responsible for processing being linked to a mathematical function according to suitable weight-age. The mathematical equation for the Neural Network autoregression can be written as.

fw=i=1KWi,jYi 6

Fig. 2.

Fig. 2

ANNAR model with four inputs one hidden layer with three hidden neurons

In Eq. 1 the variable Yi stands for the hidden layer algorithm which uses the sigmoid function. The sigmoid function is given in Eq. 7. Here Wi,j represents the weight or coefficient associated with the ith element in the vector W of the jth element in the Y vector

gy=11+e-y 7

The graph of the feed-forward method is given in Fig. 2.

Simple exponential smoothing method

Simple exponential smoothing (SES) is a technique used for the forecasting of time series data. This method is usually applicable for forecasting of any time series. This method is assumed to fit best when the time series data has no seasonality or trend. This method assumes a weighting procedure for the successive time series observations, as it assigns weights in exponentially decreasing form over time. The mathematical formula of simple exponential smoothing can be written as [44].

fst=αxt+1-αst-1 8

After simplification Eq. 8 results in

fst=st-1+α(xt-st-1) 9

Here the

st= Smoothed statistic or the weighted average of current observation xt

st-1=One-time lagged smoothed statistic
α=Smoothingparameterrangesfrom0<α<1

xt = Current time period.

Testing indicators

The most important task in time series analysis is the evaluation and selection of a suitable model. This is because the researcher assumes that the chosen model works more efficiently than others. Criteria exist for the application of a suitable model for forecasting, some of which are given below [45].

MSE=1ni=1net2 10
RMSE=1ni=1net2 11
MAE=1ni=1net 12
MAPE=1ni=1n|et||Yt|100 13

where et stands for the error terms of yearly death cases and Yt stands for the observed time series at a point in time t. Based on these criteria, we select the model which results in the lowest number.

Results

The study commenced by considering the number of modeling samples. It was determined that a sample size larger than 50 would be optimal for effectively capturing the statistical properties of the time series data. When the sample size is small, the parameters of the ARIMA model may become more inaccurate, leading to unreliable forecasts. In such instances, alternative methods such as simple exponential smoothing, naïve, and Holt-Winter models may be more appropriate [46, 47]. Simple exponential smoothing is particularly useful for forecasting time series data, especially when the sample size is small. This approach involves calculating weighted averages of past observations, with the weights gradually decreasing exponentially as the observations become older. To validate the model, the data was divided into an 80% training set and a 20% testing set.

It can be noted from Table 2, that in training 80% of the data the ANNAR model showed the lowest values of all KPIs. Since the observations are few and the rest 20% will be very few in numbers, we will be applying the rest of the techniques to the complete dataset. Neural networks can be applied when the sample size is small. In machine learning, neural networks can be applied to datasets of different sizes, ranging from a few data points to millions of data points. The size of the dataset does not determine the applicability of neural networks, but rather the complexity of the problem you are trying to solve and the architecture of the neural network [48, 49].

Table 2.

Candidate models for the yearly CVD death cases Split data technique 80%

Candidate Models MSE RMSE MAE MAPE
Naïve 1790.14 42.31 33.23 14.11
SES 1690.85 41.12 31.39 13.32
Holt 1645.11 40.56 30.68 13.18
ANNAR 1572.12 39.65 29.31 12.70

Results from Table 3 show that the neural network autoregressive model (ANNAR) outperformed all the candidate models in testing on the complete dataset. The root means square error for the death case is 38.86 and the mean absolute error is 13.08 which indicates the dominancy of the ANNAR upon all the other selected models. The Naïve model showed the maximum values of KPIs as MSE = 2304, RMSE = 48, MAE = 33.86, and MAPE = 13.86 followed by the SES method showing MSE 2124.29, RMSE = 46.09, MAE = 33.85, and MAPE = 13.22. ANNAR showed the lowest values of KPIs among all candidate models applied and proved to be a better-performing methodology for modeling and forecasting the CVD death cases in Sindh. Further, we converted Table 2 into a visual form to enable a better understanding using residual analysis based on exploratory statistics. The residual plots of the death series and further showed the fitted versus the observed values of the CVD death cases.

Table 3.

Candidate models for the yearly CVD death cases for 20% of data

Candidate Models MSE RMSE MAE MAPE
Naïve 2304.00 48.00 33.86 13.86
SES 2124.29 46.09 33.85 13.77
Holt 2082.09 45.63 32.213 13.22
ANNAR 1510.10 38.86 30.04 13.08

The Ljung-Box test is also conducted, wherein the null hypothesis assumes no autocorrelation among the residual terms, indicating a lack of model fit. The diagnostic examinations of the residuals concluded that the chosen ANNAR model exhibits a satisfactory fit without any autocorrelation among the residuals [47]. Figure 3 presents the diagnostic results, including the ACF plot and the plot of residuals overlaid with the normal curve, demonstrating the normality of the residuals. In a normally distributed scenario, the histogram should display a bell-shaped distribution, resembling the density plot. Notably, the histogram fitted using the ANNAR method aligns well with the residual data compared to other candidate models [50]. Furthermore, the lag values of the residuals generated by the ANNAR model fall within the acceptable probability limits. To further illustrate the closeness between observed and fitted observations, we present a plot depicting the observed versus fitted values of the series. Figure 4 shows the deviation from the observed is less through the neural network autoregressive approach which is an indication that this model is the best fit for the series. This is then used to make a next five-year forecast with the forecasted values given in Table 4 (95% confidence interval). Moreover the QQ-norm plots are given in Appendix.

Fig. 4.

Fig. 4

Observed versus fitted graph of CVD using ANNAR and SES

Table 4.

Long-term (5-year) forecasted values using neural network autoregressive

Year Forecasted value 95% Lower C.I 95% Upper C.I
2022 236.32 158.13 314.29
2023 247.67 137.23 358.90
2024 263.06 122.28 384.12
2025 282.72 117.20 393.24
2026 304.09 120.54 392.08

Discussion

In recent years, a model called the Artificial Neural Network Auto-Regressive (ANNAR) has become quite powerful [51, 52]. It's good at understanding complicated, non-straightforward connections in data. ANNAR has a track record of success in various situations, like predicting disease outbreaks and assessing the effectiveness of drugs, which has made it popular among researchers [5355]. In this study, this work used traditional methods for analyzing time series data and more modern machine-learning techniques to the small sample size. Regarding the prediction of long-term cardiovascular disease cases, the ANNAR model outperformed the other techniques [5658]. Further, this technique can be extended to trend analysis and other machine learning models based on a non-linear approach. The ANNAR model's improved performance is probably caused by its strong training and optimization strategies as well as its capacity to accurately capture and predict temporal dependencies and complicated, non-linear interactions seen in the data [5961]. For a variety of time-series prediction tasks, its versatility, adaptability, and sophisticated feature processing make it a good fit [6264].

In light of the results, it is recommended that preventive measures be taken to reduce the burden of CVDs in Sindh, including promoting healthy diets, increasing physical activity, and reducing tobacco use. Screening and early detection programs can also help diagnose and manage CVDs, reducing the risk of complications and improving outcomes. Individuals in Sindh need to take an active role in managing their heart health by making healthy lifestyle choices and seeking medical care when needed.

Implication of the Study

The present study focused on modeling CVD death cases in the Sindh Province of Pakistan by conventional and non-linear time series models ANNAR. The outcomes demonstrated that the ANNAR model proposed in this research performed exceptionally well compared to the conventional methods. This study is unique in its nature as no such study has been performed for modeling the CVD death cases in the Sindh province of Pakistan. This study can help identify high-risk areas and populations that require greater attention and resources. This information can be used to prioritize resource allocation toward prevention and treatment strategies for those populations. Modelling CVD death cases can help estimate the economic impact of the disease in the region, and the cost-effectiveness of different interventions.

In the light of above suggestions, the following policy measures can be taken at the government level as well provincial level.

  1. Encouraging healthy lifestyle choices: One of the best ways to prevent CVD is by adopting a healthy lifestyle, which includes regular exercise, healthy eating habits, quitting smoking, and reducing alcohol consumption. The government can launch public awareness campaigns to promote healthy living.

  2. Providing access to preventive care: Early detection and treatment of CVD can significantly reduce mortality rates. Therefore, it is crucial to provide access to preventive care, such as regular check-ups, blood pressure and cholesterol screenings, and other diagnostic tests.

  3. Improving the quality of healthcare services: Healthcare facilities in Pakistan need to be improved, and the quality of care should be enhanced. The government can invest in healthcare infrastructure, equip hospitals with modern technology, and train healthcare workers to provide better care.

  4. Increasing taxes on unhealthy products: Taxes on unhealthy products, such as tobacco and sugary drinks, can reduce their consumption and promote healthier choices.

  5. Providing access to affordable healthy foods: The government can encourage the production and consumption of healthy foods, such as fruits, vegetables, and whole grains, and make them more affordable for the general population.

These policies can significantly reduce the number of CVD deaths in Pakistan. However, implementing these policies requires a sustained effort and collaboration between the government, healthcare providers, and the public.

Conclusion

Cardiovascular disease is one of the leading causes of death for humankind. This work aims to predict the yearly CVD patients in one Pakistani city Nawabshah, in the Sindh province. According to the findings, Sindh urgently needs focused public health programs to increase public knowledge of CVD risk factors such as tobacco use, physical inactivity, and unhealthy diets. For maximum effect, educational activities can be customized to target specific groups. In modeling diseases, there is no single method that is considered definitively better than all others. The choice of method depends on the specific requirements of the problem being addressed and the availability of data. Exponential models, such as simple exponential growth models, have been commonly used in the past to model the spread of diseases. These models are relatively straightforward to implement, but they have limitations when it comes to capturing complex disease dynamics. Neural network models, such as ANNAR, have gained popularity in recent years due to their ability to capture complex nonlinear relationships in data. ANNAR has been applied to a range of disease modeling problems, including predicting disease outbreaks and drug efficacy. This work applied different time series methodologies, one based on the classical time series method and the second based on the machine learning technique. The results found that the ANNAR outperformed for the long-term forecasting period.

Future research could focus on comparing a wider variety of modeling approaches outside of ANNAR and traditional time series methods. Identifying the best strategy for forecasting CVD death rates in Nawabshah, could involve utilizing additional machine learning models such as Random Forests, Support Vector Machines (SVM), and ensemble approaches.

Acknowledgements

The authors extend their sincere appreciation to the Researchers Supporting Project number (RSPD2025R1038), King Saud University, Riyadh, Saudi Arabia.

Abbreviations

CVD

Cardiovascular disease

ANNAR

Artificial neural network auto-regressive

SES

Exponential smoothing

ARIMA

Auto-regressive integrated moving average

SVM

Support Vector Machines

RMSE

Root Mean Square Deviation Error

MAE

Mean Absolute Error

MAPE

Mean Absolute Percentage Error

Appendix

Fig. 5.

Fig. 5

QQ norm Plot of Naïve method

5,

Fig. 6.

Fig. 6

QQ norm Plot of Holts method

6,

Fig. 7.

Fig. 7

QQ norm Plot of SES method

7 and

Fig. 8.

Fig. 8

QQ norm Plot of ANNAR method

8

Authors’ contributions

All the authors (Moiz Qureshi, Khushboo Ishaq, Muhammad Daniyal, Hasnain Iftikhar, Mohd Ziaur Rehman, and S. A. Atif Salar) read and approved the final manuscript. They have also equally contributed to this research, approved its claims, and agreed to be authors.

Funding

Researchers Supporting Project number (RSPD2025R1038), King Saud University, Riyadh, Saudi Arabia.

Data availability

The data that support the findings of this study are available from the corresponding author upon reasonable request.

Declarations

Ethics approval and consent to participate

This study has been approved by the ethical review board of the district office, Sindh with approval no 32E/4/2021.

Consent for publication

The authors declare no conflict of interest for this article.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1. Balouch, F. G., Laghari, D. Z. A., Baig, N. M., & Samo, A. A. (2022). Prevalence of cardiovascular disease risk factors in urban and rural areas of Hyderabad, Sindh, Pakistan.
  • 2.Malav A, Kadam K, Kamat P. Prediction of heart disease using k-means and artificial neural network as hybrid approach to improve accuracy. International Journal of Engineering and Technology. 2017;9(4):3081–5. [Google Scholar]
  • 3.Cuba WM, Huaman Alfaro JC, Iftikhar H, López-Gonzales JL. Modeling and analysis of monkeypox outbreak using a new time series ensemble technique. Axioms. 2024;13(8):554. [Google Scholar]
  • 4.Iftikhar H, Khan M, Khan Z, Khan F, Alshanbari HM, Ahmad Z. A comparative analysis of machine learning models: a case study in predicting chronic kidney disease. Sustainability. 2023;15(3):2754. [Google Scholar]
  • 5.Iftikhar H, Khan M, Khan MS, Khan M. Short-term forecasting of monkeypox cases using a novel filtering and combining technique. Diagnostics. 2023;13(11):1923. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Zhao Y, Xiong W, Li C, Zhao R, Lu H, Song, S.,... Ge, J. Hypoxia-induced signaling in the cardiovascular system: pathogenesis and therapeutic targets. Signal Transduct Target Ther. 2023;8(1):431. 10.1038/s41392-02. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Mythili T, Mukherji D, Padalia N, Naidu A. A heart disease prediction model using SVM-decision trees-logistic regression (SDL). International Journal of Computer Applications. 2013;68(16).
  • 8.Sajid MR, Muhammad N, Zakaria R, Shahbaz A, Nauman A. Associated factors of cardiovascular diseases in Pakistan: Assessment of path analyses using warp partial least squares estimation. Pakistan Journal of Statistics and Operation Research. 2020:265–77.
  • 9.Akhtar S, Asghar N. Risk factors of cardiovascular disease in district Swat. J Pak Med Assoc. 2015;65(9):1001–4. [PubMed] [Google Scholar]
  • 10.Hu Z, Qiu H, Su Z, Shen M, Chen Z. A stacking ensemble model to predict daily number of hospital admissions for cardiovascular diseases. IEEE Access. 2020;8:138719–29. [Google Scholar]
  • 11.Mohan S, Thirumalai C, Srivastava G. Effective heart disease prediction using hybrid machine learning techniques. IEEE Access. 2019;7:81542–54. [Google Scholar]
  • 12.McNown R, Rogers A. Forecasting cause-specific mortality using time series methods. Int J Forecast. 1992;8(3):413–32. [Google Scholar]
  • 13.Guzman Castillo M, Gillespie DO, Allen K, Bandosz P, Schmid V, Capewell S, et al. Future declines in coronary heart disease mortality in England and Wales could counter the burden of population aging. PLoS ONE. 2014;9(6):e99482. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Olsavszky V, Dosius M, Vladescu C, Benecke J. Time series analysis and forecasting with automated machine learning on a national ICD-10 database. Int J Environ Res Public Health. 2020;17(14):4979. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Konar S, Auluck N, Ganesan R, Goyal AK, Kaur T, Sahi M, et al. A non-linear time series based artificial intelligence model to predict outcome in cardiac surgery. Heal Technol. 2022;12(6):1169–81. [Google Scholar]
  • 16.Suzuki S, Yamashita T, Sakama T, Arita T, Yagi N, Otsuka T, et al. Comparison of risk models for mortality and cardiovascular events between machine learning and conventional logistic regression analysis. PLoS ONE. 2019;14(9):e0221911. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Bhaskaran K, Gasparrini A, Hajat S, Smeeth L, Armstrong B. Time series regression studies in environmental epidemiology. Int J Epidemiol. 2013;42(4):1187–95. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Iqbal R, Ahmad Z, Malik F, Mahmood S, Shahzadi N, Mehwish S, et al. A statistical analysis of hypertension as a cardiovascular risk factor. Middle East J Sci Res. 2012;12(1):19–22. [Google Scholar]
  • 19.Khan MZ, Pervaiz MK, Javed I. Biostatistical study of clinical risk factors of myocardial infarction: a case-control study from Pakistan. Pakistan Armed Forces Medical Journal. 2016;66(3):354–60. [Google Scholar]
  • 20.Zulfiqar, N., Razzaq, S., & Satti, S.Risk Factors of Cardiac Diseases In Pakistan. (2019) >17th, 373.
  • 21.Khan MS, et al. Knowledge of modifiable risk factors of heart disease among patients with acute myocardial infarction in Karachi, Pakistan: A cross-sectional study. BMC Cardiovasc Disord. 2006;6:1–9. 10.1186/1471-2261-6-18. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Nusinovici S, Tham YC, Yan MYC, Ting DSW, Li J, Sabanayagam C, et al. Logistic regression was as good as machine learning for predicting major chronic diseases. J Clin Epidemiol. 2020;122:56–69. [DOI] [PubMed] [Google Scholar]
  • 23.Ahmed R, Rizwan-ur-Rashid MP, Ahmed SW. Prevalence of cigarette smoking among young adults in Pakistan. J Pak Med Assoc. 2008;58(11):597–601. [PubMed] [Google Scholar]
  • 24.Kanwal T, Manzoor S, Firdos M, Hassan I, Aslam S. Regression Methods for Analyzing Risk Factors Causing Cardiovascular Disease in Muzaffarabad AJ&K. Pakistan Pakistan Journal of Medical Research. 2019;58(4):180–6. [Google Scholar]
  • 25.Jahangeer SMA, Ikram A, Anmol A, Lashari MN, Kataria K, Turk E, et al. The relationship of lifestyle and dietary habits of southeast Asian (Pakistani) population with cardiovascular diseases: a case-control study. Pakistan Heart Journal. 2022;55(4):396–403. [Google Scholar]
  • 26.Huang, Y., Wang, C., Zhou, T., Xie, F., Liu, Z., Xu, H.,... Xu, K. (2024). Lumican promotes calcific aortic valve disease through H3 histone lactylation. European Heart Journal, ehae407. 10.1093/eurheartj/ehae407. [DOI] [PubMed]
  • 27.Almosova A, Andresen N. Nonlinear inflation forecasting with recurrent neural networks. J Forecast. 2023;42(2):240–59. [Google Scholar]
  • 28.Perone, G. (2021). Comparison of ARIMA, ETS, NNAR, TBATS and hybrid models to forecast the second wave of COVID-19 hospitalizations in Italy. The European Journal of Health Economics, 1–24. [DOI] [PMC free article] [PubMed]
  • 29.Alshanbari HM, Iftikhar H, Khan F, Rind M, Ahmad Z, El-Bagoury AAAH. On the implementation of the artificial neural network approach for forecasting different healthcare events. Diagnostics. 2023;13(7):1310. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Sultana, N., Sharma, N., & Sharma, K. P. (2019, April). Ensemble model based on NNAR and SVR for predicting influenza incidences. In Proceedings of the International Conference on Advances in Electronics, Electrical & Computational Intelligence (ICAEEC).
  • 31.Li X, Fan J, Wang Y. Time series analysis of stroke incidence in Shandong, China. J Epidemiol Community Health. 2015;69(5):450–6. [Google Scholar]
  • 32.Rapsomaniki E, Timmis A, George J, Pujades-Rodriguez M, Shah AD, Denaxas S, et al. Blood pressure and incidence of twelve cardiovascular diseases: lifetime risks, healthy life-years lost, and age-specific associations in 1· 25 million people. The Lancet. 2014;383(9932):1899–911. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Kaji DA, Zech JR, Kim JS, Cho SK, Dangayach NS, Costa AB, et al. An attention based deep learning model of clinical events in the intensive care unit. PLoS ONE. 2019;14(2):e0211057. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Brodie RJ, De Kluyver CA. A comparison of the short term forecasting accuracy of econometric and naive extrapolation models of market share. Int J Forecast. 1987;3(3–4):423–37. [Google Scholar]
  • 35.Octiva, C. S., Nuryanto, U. W., Eldo, H., & Tahir, A. (2024). Application of Holt-Winter Exponential Smoothing Method to Design a Drug Inventory Prediction Application in Private Health Units. Jurnal Informasi Dan Teknologi, 1–6.
  • 36.Syafei, A. D., Ramadhan, N., Hermana, J., Slamet, A., Boedisantoso, R., & Assomadi, A. F. (2018). Application of Exponential Smoothing Holt Winter and ARIMA Models for Predicting Air Pollutant Concentrations. EnvironmentAsia, 11(3).
  • 37.VP, V. (2024). Non-linear time series Models and their applications (Doctoral dissertation, Department of Statistics, Farook College.).
  • 38.Fortuna L, Nunnari G, Nunnari S. Nonlinear modeling of solar radiation and wind speed time series, vol. 10. Berlin, Germany: Springer; 2016. [Google Scholar]
  • 39.Demir İ, Kirisci M. Forecasting COVID-19 disease cases using the SARIMA-NNAR hybrid model. Universal Journal of Mathematics and Applications. 2022;5(1):15–23. [Google Scholar]
  • 40.Carbo-Bustinza N, Iftikhar H, Belmonte M, Cabello-Torres RJ, De La Cruz ARH, López-Gonzales JL. Short-term forecasting of Ozone concentration in metropolitan Lima using hybrid combinations of time series models. Appl Sci. 2023;13(18):10514. [Google Scholar]
  • 41.Che Z, Purushotham S, Cho K, Sontag D, Liu Y. Recurrent neural networks for multivariate time series with missing values. Sci Rep. 2018;8(1):6085. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Tang Z, Fishwick PA. Feed-forward neural nets as models for time series forecasting. ORSA Journal of Computing. 1993;5:374–85. [Google Scholar]
  • 43.Leslie N. Smith. A disciplined approach to neural network hyper-parameters: Part 1 – learning rate, batch size, momentum, and weight decay, 2018.
  • 44.Gardner ES Jr. Exponential smoothing: The state of the art. J Forecast. 1985;4(1):1–28. [Google Scholar]
  • 45.Zhou L, Zhao P, Wu D, Cheng C, Huang H. Time series model for forecasting the number of new admission inpatients. BMC Med Inform Decis Mak. 2018;18:1–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Hyndman RJ, Kostenko AV. Minimum sample size requirements for seasonal forecasting models. foresight. 2007;6(Spring):12–5.
  • 47.Box GE, Jenkins GM, Reinsel GC, Ljung GM. Time series analysis: forecasting and control: John Wiley & Sons; 2015.
  • 48.Bengio Y. Practical recommendations for gradient-based training of deep architectures. Neural Networks: Tricks of the Trade: Second Edition: Springer; 2012. p. 437–78.
  • 49.Wang X, Chen X, Tang Y, Wu J, Qin D, Yu, L.,... Wu, A. The Therapeutic Potential of Plant Polysaccharides in Metabolic Diseases. Pharmaceuticals. 2022;15(11):1329. 10.3390/ph15111329. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Jung RC, Kukuk M, Liesenfeld R. Time series of count data: modeling, estimation and diagnostics. Comput Stat Data Anal. 2006;51(4):2350–64. [Google Scholar]
  • 51.Khan A, Qureshi M, Daniyal M, Tawiah K. A novel study on machine learning algorithm-based cardiovascular disease prediction. Health Soc Care Community. 2023;2023(1):1406060. [Google Scholar]
  • 52.Jiang C, Xie N, Sun T, Ma W, Zhang, B.,... Li, W. Xanthohumol Inhibits TGF-β1-Induced Cardiac Fibroblasts Activation via Mediating PTEN/Akt/mTOR Signaling Pathway. Drug Des Dev Ther. 2020;14:5431–9. 10.2147/DDDT.S282206. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Tawiah K, Daniyal M, Qureshi M. Pakistan CO2 emission modelling and forecasting: a linear and nonlinear time series approach. J Environ Public Health. 2023;2023(1):5903362. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Qureshi M, Khan S, Bantan RA, Daniyal M, Elgarhy M, Marzo RR, Lin Y. Modeling and forecasting monkeypox cases using stochastic models. J Clin Med. 2022;11(21):6555. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Li H, Wang Y, Fan R, Lv H, Sun H, Xie, H.,... Xia, Z. The effects of ferulic acid on the pharmacokinetics of warfarin in rats after biliary drainage. Drug Des Dev Ther. 2016;10:2173–80. 10.2147/DDDT.S107917. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Zhao Y, Hu J, Sun X, Yang K, Yang L, Kong, L.,... Ge, J. Loss of m6A demethylase ALKBH5 promotes post-ischemic angiogenesis via post-transcriptional stabilization of WNT5A. Clin Transl Med. 2021;11(5):e402. 10.1002/ctm2.402. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Deng J, Liu Q, Ye L, Wang S, Song Z, Zhu, M.,... Chen, T. The Janus face of mitophagy in myocardial ischemia/reperfusion injury and recovery. Biomed Pharmacother. 2024;173: 116337. 10.1016/j.biopha.2024.116337. [DOI] [PubMed] [Google Scholar]
  • 58.Iftikhar H, Daniyal M, Qureshi M, Tawiah K, Ansah RK, Afriyie JK. A hybrid forecasting technique for infection and death from the mpox virus. Digital Health. 2023;9:20552076231204748. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Wynants, L., Van Calster, B., Collins, G. S., Riley, R. D., Heinze, G., Schuit, E., ... & van Smeden, M. (2020). Prediction models for diagnosis and prognosis of covid-19: systematic review and critical appraisal. bmj, 369. [DOI] [PMC free article] [PubMed]
  • 60.Gan W, Koehoorn M, Davies H, Demers P, Tamburic L, Brauer M. Long-term exposure to traffic-related air pollution and the risk of coronary heart disease hospitalization and mortality. Epidemiology. 2011;22(1):S30. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Wen J, Li S, Lin Z, Hu Y, Huang C. Systematic literature review of machine learning based software development effort estimation models. Inf Softw Technol. 2012;54(1):41–59. [Google Scholar]
  • 62.Goldenberg, A., Zheng, A. X., Fienberg, S. E., & Airoldi, E. M. (2010). A survey of statistical network models. Foundations and Trends® in Machine Learning, 2(2):129–233.
  • 63.Iftikhar, H., Qureshi, M., Zywiołek, J., López-Gonzales, J. L., & Albalawi, O. (2024). Short-term PM 2.5 forecasting using a unique ensemble technique for proactive environmental management initiatives. Frontiers in Environmental Science. 12:1442644.
  • 64.Almarashi AM, Daniyal M, Jamal F. A novel comparative study of NNAR approach with linear stochastic time series models in predicting tennis player’s performance. BMC Sports Sci Med Rehabil. 2024;16(1):28. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The data that support the findings of this study are available from the corresponding author upon reasonable request.


Articles from BMC Public Health are provided here courtesy of BMC

RESOURCES