Abstract
Background and Aims
Like early diagnosis, predicting the survival of patients with Coronavirus Disease 2019 (COVID‐19) is of great importance. Survival prediction models help doctors be more cautious to treat the patients who are at high risk of dying because of medical conditions. This study aims to predict the survival of hospitalized patients with COVID‐19 by comparing the accuracy of machine learning (ML) models.
Methods
It is a cross‐sectional study which was performed in 2022 in Fasa city in Iran country. The research data set was extracted from the period February 18, 2020 to February 10, 2021, and contains 2442 hospitalized patients' records with 84 features. A comparison was made between the efficiency of five ML algorithms to predict survival, includes Naive Bayes (NB), K‐nearest neighbors (KNN), random forest (RF), decision tree (DT), and multilayer perceptron (MLP). Modeling steps were done with Python language in the Anaconda Navigator 3 environment.
Results
Our findings show that NB algorithm had better performance than others with accuracy, precision, recall, F‐score, and area under receiver operating characteristic curve of 97%, 96%, 96%, 96%, and 97%, respectively. Based on the analysis of factors affecting survival, heart disease, pulmonary diseases and blood related disease were the most important disease related to death.
Conclusion
The development of software systems based on NB will be effective to predict the survival of COVID‐19 patients
Keywords: COVID‐19, decision tree, K‐nearest neighbors, machine learning, Naive Bayes, random forest
Key points
We evaluated the role of clinical data in the survival of hospitalized COVID‐19 patients.
NB classifier can classify all the cases of death correctly and its False Negative Rate is zero.
History of high blood pressure and heart disease was important underlying diseases factors in the survival of COVID‐19 patients.
1. INTRODUCTION
One of the health threats to countries is the spread of infectious diseases, which, besides causing international crises, 1 creates many health, economic and social challenges and problems. 2 The most important part of health care support services is continuous monitoring of outbreaks of infectious and fatal respiratory diseases. 3 , 4 Coronaviruses (CoV) are a large family of infectious viruses which cause numerous problems with a wide range of illnesses from the common cold to more severe illnesses such as Middle East Respiratory Syndrome (MERS‐CoV) and Acute Respiratory Syndrome (SARS‐CoV). 5 The new coronavirus (nCoV) first emerged in Wuhan, China on December 19, 2019, and quickly spread to other countries. 6 COVID‐19 has emerged as a major threat to health care systems worldwide 7 and has affected various aspects of medical science, from health to education. 8
From the initial reports of the outbreak of this disease, until December 19, 2022, the number of confirmed global cases was 649,038,437 and the number of deaths because of Covid‐19 was 6,645,812. 9 This statistic for Iran was 7,560,444 confirmed cases and 144,664 deaths. 10 Because of the rapid spread of this disease, in many countries and regions, the number of patients infected with this virus has exceeded the capacity of hospitals and has imposed a heavy burden on health systems and resources in many countries. 11 Timely diagnosis of infectious diseases leads to further optimization of treatment and thus containment of outbreaks. 12 To reduce part of this burden, the use of clinical prediction models can be helpful and provide better decisions 13 , 14 , 15 , 16 , 17 and play an effective role in controlling disease outbreaks. 18 In recent years, the applications of artificial intelligence (AI) in various fields of health have received more attention. 19 Artificial intelligence is a subbranch of computer science that can make smart decisions and help predict by learning from data. 20 AI can help professionals calculate risk factors, 21 process medical images (including chest images), 22 , 23 classify drugs, even analyze them and ultimately respond to a crisis. 21 Data mining is a very popular method that uses statistical methods, visualization, machine learning, and other techniques of data manipulation and knowledge extraction to analyze, diagnose, predict and control various diseases, including infectious diseases to gain insight into data and hidden patterns. 24 , 25 , 26 Machine learning as a learning method by automating knowledge acquisition has important applications in AI research. 27 Using various data mining techniques to the data collected about infectious diseases, besides a better understanding of the effective factors in the occurrence of infection, it is possible to recover from these diseases. 18
By the spread of COVID‐19, using ML algorithms for various applications, 28 including disease diagnosis, 29 prediction of readmission, 30 predicting of survival based on data, specific clinical symptoms, and parameters 21 , 31 , 32 , 33 and providing of decision‐making models 20 were developed. Early diagnosis of COVID‐19 patients is essential for the identification and survival of vulnerable patients. 31 With the improvement of diagnostic accuracy of covid‐19 patients, it is very necessary to identify and predict survival and cases leading to death. 33 Therefore, identifying factors affecting the level of risk and determining the probability of survival is very important. Moreover, evaluating the symptoms of COVID‐19 because of the identification of the early signs of danger can also help the doctor make historic decisions about the methods of hospitalization, treatment, and discharge of the patient. 34
Therefore, this study aims to provide a prediction model for the survival of COVID‐19 patients in Fasa city. Based on our current knowledge, no model was done on clinical data of this region. The descriptions of the used algorithms and data set are presented in Section 2. In Section 3, the evaluation results are shown. The comparison of the results with similar research was discussed in Section 4. Finally, the conclusion and future works is presented in Section 5.
2. METHODS
The protocol of this study was approved by the Ethics Committee of the Fasa University of Medical Sciences (FUMS) with code (IR.FUMS.REC.1399.059). Protected personal health information was removed. The block diagram is shown in Figure 1.
Figure 1.

Block diagram of this research.
2.1. Data set description
The data set of this research was with 2442 hospitalized patients related to Valiasr Hospital, which were classified into two classes: death (187 cases) and survival (2255 cases).
All models are implemented using the Python programming language and in the Anaconda Navigator 3 environment. Python is used as a dynamic and general‐purpose programming language in various areas of ML. Machine Learning algorithms in Python are implemented using special libraries. 35
2.2. Data preparation
The preprocessing process includes checking for outliers and noise, and missing values. Because of the large volume of electronic health record data that is produced in different places, data loss, duplication, and even wrong data are inevitable. Therefore, the data are preprocessed, and then subjected to the building model stage. After examining the initial data set, it was found that numerous records contain missing values, so that many futures lacked values. These records were discarded. In the next step, some features contained a few missing values, which were replaced by the Mode value. Moreover, to check the duplicate data, based on the patients' ID, this case was checked and the duplicate data were excluded from the data set. After cleaning the data set, the number of samples was reduced to 948 cases. The number of samples of death and survival class was equal to 786 and 162, respectively. It is an imbalanced data set. The synthetic Minority Oversampling Technique (SMOTE) was used to create the balanced data set. This method synthetically generates samples from the minority class to create balance. A balanced data set helps the classifier be informed about both classes for prediction and avoids distributional bias. 36
2.3. Feature selection
In the used data set, all features except age are of qualitative type. Therefore, χ 2 analysis was used to check the role of the variables in predicting survival. In this algorithm, the influential variables to predict the outcome feature (survival) are selected based on the p value. χ 2 test is one of the most reliable statistical tests that can find out whether there is a systematic relationship between the two variables. 37 This test is usually used for relationships where both variables are nonparametric. 38 If there is no systematic relationship between two variables in a sample, it can be concluded that the two variables are independent of each other, which is called statistical independence. 39 There is a direct relationship between the value of χ 2 statistic and significance, that is, the larger the value of χ 2, the greater the probability of the relationship between the two variables. More precisely, if p < 0.05, the significance of the relationship between two variables is accepted with 95% confidence. 38
From the initial 84 features (83 independent and 1 dependent), contact and address‐related information were removed, and the number of features reached 59, among which 17 features were removed due to missing a lot (>70% missing value), and 42 features remained, of which 36 independent features are listed in the Table 1.
Table 1.
The features along with the corresponding p Value.
| Feature | p Value | Feature | p Value |
|---|---|---|---|
| Age | 0.226 | Nausea | 0.209 |
| Sex | 0.061 | Decreased consciousness | 0.273 |
| Contact history | 0.948 | Respiratory distress | <0.001 |
| Intubation | <0.001 | Muscular pain | <0.005 |
| Pregnancy | 0.135 | Diabetes | 0.600 |
| Vomit | 0.736 | Immunodeficiency | 0.639 |
| Cough | 0.649 | Blood disease | <0.005 |
| Vertigo | 0.079 | Neurological disease | 0.319 |
| Headache | 0.376 | Pulmonary disease | <0.001 |
| Fever | 0.428 | Heart disease | <0.001 |
| Diarrhea | 0.064 | Asthma | 0.740 |
| Anorexia | 0.461 | Kidney disease | 0.338 |
| Decreased smell | 0.416 | Blood pressure | <0.005 |
| Chest pain | 0.634 | Number breaths | 0.671 |
| Decreased taste | 0.639 | Po2 level | <0.001 |
| Stomachache | 0.660 | CT scan | <0.001 |
| limb_ paralysis | 0.848 | HIV/AIDS | 0.650 |
| Skin lesion | 0.322 | limb plague | 0.420 |
“Other Chronic Diseases” because of the incomprehensibility of the feature was deleted. “Neurological chronic diseases,” “blood chronic diseases,” and “convulsions” were deleted because the value of these features for all the instants was no. We used the “po2_number” feature only to check the correctness of determining the level of normal and non‐normal.
2.4. Modeling
After preprocessing the data and selecting the last features, the data sets are divided into two categories: training data sets (75%) and testing data sets (25%). Prediction models are created and evaluated with training and test data sets, respectively. Five predictive techniques to predict the survival of COVID‐19 patients were investigated. These algorithms include NB, K‐Nearest Neighbor (KNN), random forest (RF), decision tree (DT), and multilayer perceptron (MLP). The optimal performance of an algorithm can be achieved if the best configuration has been used in the model's build. So, we performed hyper parameter optimization to specify the values at which the ML algorithm will perform best. For each algorithm, we mentioned the hyperparameters that improved the performance of the model by adjusting the values. Other hyperparameters with the default values maintained the performance of the model.
We considered these five algorithms according to the following characteristics:
2.4.1. DT
It is used for classification tasks which was widely appreciated for their ability to manage classified data, simplicity, and comprehensibility. 40 , 41 This algorithm comprises three types of nodes (root, internal node and external node called leaf) in its tree structure. 1 Based on the tree structure, the internal node is a decision point and each leaf is a class. It uses a set of decision rules to assign data to classes. 42 One of the DT classifiers is Classification and regression tree (CART). It is a well‐known decision tree technique. We tuned the CART with max_depth = 7, class_weight = “balanced” and random_state = 42.
2.4.2. RF
This algorithm produces a tree like the DT algorithm, with the difference that in the process of tree generation, several trees are produced from the values of random samples in the data set, and the result will be based on the results of most developed trees. 43 It is a powerful statistical classifier. 44 This classifier has a high ability to manage high‐dimensional data. 45 Compared to other classifiers, RF has much higher classification accuracy, is a powerful method to determine variable significance, models complex interactions among predictive features, and can perform several types of statistical data analysis, It is more flexible to handle missing values. 44 There is a hyperparameter “n_estimators,” which is actually the number of trees the algorithm builds before receiving maximum votes or averaging predictions. More trees increase efficiency and make predictions more stable. “min_sample_leaf” as another hyperparameter specifies the minimum number of leaves required to split an external node. The best performance of this algorithm was evaluated with the values of 1000 and 2 for hyperparameters n_estimators and min_sample_leaf, respectively.
2.4.3. MLP
MLP is a class of feedforward artificial neural networks (ANN) that have at least three layers of nodes (an input layer, a hidden layer, and an output layer). Except for the input nodes, each node is a neuron that uses a nonlinear activation function. For training, MLP uses a supervised ML approach called backpropagation. 46 , 47 Neurons are the computing units in a neural network. In the MLP, the outputs of the first layer are used as the inputs of the next hidden layer; this continues until, after a certain number of layers, the outputs of the last hidden layer are used as inputs to the output layer. 46 The hyperparameters set in this algorithm are:
Number of hidden layers = 2, epochs = 200, activation function = “Rectified Linear Unit (ReLU)” and “softmax” for hidden and output layer, respectively, neurons in each hidden layer = 6, batch size = 64, layer type = dense, optimizer = “Adam.”
2.4.4. KNN
It is a simple but effective algorithm for classification. 48 KNN regression is used to transform data with higher dimensions into a one‐dimensional space and is used as an efficient computational tool. 49 KNN uses numerous sets of the nearest neighbors for the modeling. KNN is easy to implement and intuitive to understand. 48 We tuned the KNN with n neighbors = 4 and metric = “Euclidean.”
2.4.5. NB
It is one of the popular classification methods used in various fields, including medical diagnosis. 50 NB Methods are based on applying Bayes' theorem with the “naive” assumption of conditional independence among every pair of features, given the value of the class variable. 50 , 51 We used Gaussian NB algorithm for classification. The hyperparameters of Gaussian NB model would have default.
2.5. Model assessment
An important part of building an effective ML prediction model is evaluating the model's performance. One tool for measuring the performance of models is the confusion matrix (Table 1). We used some evaluation criteria, including accuracy, sensitivity, readability (Table 2), and area under the receiver operating characteristic curve (ROC) or under ROC curve (AUC) measurements. ROC curve is a useful tool for measuring the performance of models, used to predict the probability of a binary outcome. For several candidate threshold values between 0.0 and 1.0, this plot shows the false positive rate versus true positive rate and how accurate the model is in predicting the positive class when the correct result is positive.
Table 2.
Confusion matrix for measuring models' performance.
| Predicted class | |||
|---|---|---|---|
| Death | Survival | ||
| Actual class | Death | True positive (TP) | False negative (FN) |
| Survival | False positive (FP) | True negative (TN) | |
In the confusion matrix, each column shows the number of test samples in which the model predicted their class labels while each row shows the number of test samples based on actual labels (Table 2).
Terminology from a confusion matrix is as follows:
TP: the number of test samples, class labels of which are predicted by the classification model as “Death” correctly.
TN: the number of test samples, class labels of which are predicted by the classification model as “Survival” correctly.
FP: the number of test samples, class labels of which are predicted by the classification model as “Death” incorrectly; however, their actual class label is “Survival.”
FN: the number of test samples, class labels of which are predicted by the classification model as “Survival” incorrectly; however, their actual class label is “Death.”
Based on these definitions, fundamental evaluation metrics can be defined (Table 3).
Table 3.
Calculating the evaluation criteria.
| Criteria | Equation |
|---|---|
| Accuracy | (TP+TN)/(TP +TN+FP+FN) |
| Recall | TP/(TP+FN) |
| Precision | TP/(TP+ FP) |
| F‐ measure | 2* Recall * Precision/(Recall+ Precision) |
AUC metric is better than accuracy for balanced and imbalanced data sets because it uses probabilities of predictions. 52
3. RESULTS
After preprocessing, the number of records was reduced to 1572. The number of samples in both classes was equal (786:786). After measuring the characteristics of data set using χ 2 test, 9 features for survival of COVID‐19 had a significant correlation with the output class at p < 0.05. Table 4 shows the last features and the p value for each of them.
Table 4.
Features, percentage level of the COVID‐19 symptoms in the updated data set.
| Feature name | Total (n = 1572) | Survived (n = 786) | Death (n = 786) | p Value |
|---|---|---|---|---|
| Intubation | <0.001 | |||
| Yes | 772 | 32 (4.14%) | 740 (95.85%) | |
| No | 800 | 754 (94.25%) | 46 (5.75%) | |
| CT scan | <0.001 | |||
| Yes | 993 | 374 (37.66%) | 619 (165.50%) | |
| No | 579 | 412 (71.15%) | 167 (28.84%) | |
| Blood_ pressure | <0.001 | |||
| Yes | 597 | 227 (38.02%) | 370 (61.97%) | |
| No | 975 | 559 (57.33%) | 416 (42.66%) | |
| Heart disease | <0.001 | |||
| Yes | 271 | 92 (33.94%) | 179 (66.05%) | |
| No | 1301 | 694 (53.34%) | 607 (46.65%) | |
| Respiratory disease | <0.001 | |||
| Yes | 1064 | 579 (54.41%) | 485 (45.58%) | |
| No | 508 | 301 (59.25%) | 207 (40.74%) | |
| Pulmonary disease | <0.001 | |||
| Yes | 47 | 33 (70.21%) | 14 (29.78%) | |
| No | 1525 | 772 (50.62%) | 753 (49.37%) | |
| PO2 | <0.001 | |||
| Yes | 310 | 240 (77.41%) | 70 (22.58%) | |
| No | 1262 | 546 (43.26%) | 716 (56.73%) | |
| Blood disease | <0.005 | |||
| Yes | 20 | 5 (25%) | 15 (75%) | |
| No | 1552 | 781 (50.32%) | 771 (49.67%) | |
| Muscular pain | <0.005 | |||
| Yes | 301 | 170 (56.47%) | 131 (43.52%) | |
| No | 1271 | 616 (48.46%) | 655 (51.53%) | |
The data type of all data set features is nominal. In Table 3, separately for survival and death classes, the number of samples for having (“yes”) or not having (“no”) that feature is specified. For example, 66.05% of the people who died had a history of heart disease, and 33.94% of the samples that survived did not have a history of heart disease.
Experimental results show that NB classifier had better performance than others with accuracy, precision, recall, F‐score, and ROC of 97%, 96%, 96%, 96%, and 97.0%, respectively (Figures 2 and 3).
Figure 2.

Performance comparison of survival prediction models.
Figure 3.

Comparison of AUC values in different classifiers.
Most prediction models have a ROC curve value above 0.9 which indicated the high classification power of the algorithms. KNN algorithm has a ROC curve value under 0.9.
NB classifier can classify all the cases of death correctly and its False Negative Rate (FNR) rate is zero (Figure 4).
Figure 4.

Confusion matrix for Naive Bayes classifier.
4. DISCUSSION
NB, KNN, RF, DT, and MLP algorithms were used to predict the survival of hospitalized COVID‐19 patients. Several supervised machine learning methods were used to build a disease prediction model, and an algorithm can perform differently depending on the data set. 53 Hyperparameter optimization was performed by handling multiple training sessions with the same data set to determine the best configuration for each algorithm. Using underlying diseases and symptoms as input independently, the NB model achieved 97% predictive accuracy and correctly classified all patients whose actual class was death. Based on χ 2 analysis, the most important factors affecting the death of patients have been reported. History of high blood pressure and heart disease was important underlying diseases factors in the survival of COVID‐19 patients. Patients who were intubation and whose CT‐scan results were abnormal are more likely to die.
Various studies 54 , 55 , 56 , 57 , 58 , 59 , 60 , 61 were conducted in predicting the survival and mortality of COVID‐19 patients using ML techniques. Krajah's et al. 59 study showed intubation was the most important features for the survival prediction. These findings are consistent with the results of our study.
Similar to our study, Mohammad et al. 62 in their study classified the mortality rate of COVID‐19 patients who suffer from underlying diseases or comorbidities. The results showed that the diseases, such as diabetes, asthma, cancer, high blood pressure, heart disease, and chronic obstructive pulmonary disease all affect the performance of the classifiers to detect mortality. As our study showed that high blood pressure, heart disease, and COPD were important predictors of death.
As in our study, one of the important underlying disease factors in the survival prediction was the history of hypertension and heart disease, Ikemura et al. 57 obtained the same results in their study. Likewise, the results of the study by Castagna et al. 63 showed that the history of heart disease is important in the death of patients and in the severity of deterioration of patients with COVID‐19.
The results of Khozeimeh et al.'s 58 study also showed that one of the most important features related to the mortality of patients with COVID‐19 was heart disease. The results of studies of Ahmed et al. 54 and Cisterna‐García et al. 56 similar to the results of our study, identified the features of lung disease, blood pressure, and acute respiratory distress syndrome as effective factors in the development of a predictive survival system.
Verity et al. 64 concluded that age, especially over 60‐year‐old, and underlying diseases, such as high blood pressure, cardiovascular diseases, diabetes, chronic respiratory disease, and cancer are among the most important factors associated with deaths caused by Corona. Also, the results of Moulai et al.'s 65 study showed that shortness of breath and underlying diseases are the most effective factors in predicting the mortality of patients with COVID‐19. However, the name of the underlying disease was not clearly stated. The importance of underlying diseases is such that Huyut 66 states that the lack of attention to the patient's comorbidities was one limitation of his study.
As the results of our study showed that NB has the highest performance. Similarly, Shanbehzadeh et al. 67 in their study, examined 54 factors to predict mortality because of COVID‐19; including demographics, clinical manifestations, risk factors, laboratory tests, and treatment plans. Ten ML algorithms, such as Kstar, LWL, MLP, Naive Bayes, SVM, Bayesian network, J48, RF, OneR, and PART were developed. The results showed that the best performance was related to the Bayesian network algorithm with accuracy and sensitivity of 89.31% and 64.2% respectively, in predicting mortality. From 54 predictors, the most important predictors of death and mortality from COVID‐19 were determined absolute neutrophil count, lymphocytes, loss of smell, and taste. The number of platelets, magnesium, and headache respectively obtained the least importance for predicting mortality from COVID‐19.
In the study of Krajah et al. 59 ML algorithms, such as CART, LR (logistic regression), and linear discriminant analysis (LDA), SVM, KNN and simple Gaussian NB were compared. The results showed that the most accurate model in the presented COVID‐19 data set for predicting death or survival belonged to logistic regression, in which precision, recall, and f1 score were 83%, 96%, and 90%, respectively. This model was also the simplest and fastest algorithm and provided satisfactory results, while the SVM had the highest accuracy but the slowest execution.
The results of the study of Muhammad et al. 60 showed that among models, such as SVM, simple Bayes, LR, RF, and KNN, the developed model with decision tree algorithm was more efficient in predicting the recovery of COVID‐19 patients with an accuracy of 99.85%. The results of Mohammad et al. 62 showed that among the seven ML bagging algorithm, J48, LR, RF and SVM, NB and threshold selection, the best performance related to the bagging algorithm with the highest accuracy, TPR and FPR were 83.55%, 0.835, and 0.160, respectively.
Atiyah et al. 68 showed that the Stochastic Gradient Descent (SGD) algorithm performed best among the classification algorithms (LR, GNB, RF, SGD, KNN, SVM, XGBoost, and DT) used based on ICU labeling, and KNN showed the worst performance.
Moulaei et al. 65 compared ML techniques (DT, MLP, J48, RF, KNN, and SVM) to predict mortality in COVID‐19 patients. The best result was related to the random forest model with ROC (1.000), precision (99.74%), accuracy (99.23%), specificity (99.84%), and sensitivity (98.25%). Moreover, the second most effective factor to predict the mortality was underlying diseases.
Atlam et al. 55 compared deep learning and Cox regression models to predict the survival of people infected with coronavirus and showed that Deep_Cox_COVID_19 was better than Cox_COVID_19 in terms of adaptation, precision, and accuracy. A threshold was used to evaluate the accuracy of the model. Furthermore, the most important factors affecting mortality were age, pneumonia, muscle pain, and sore throat. So that the patients whose survival probability was higher than the threshold were more likely to survive than the patients whose survival probability was lower than the threshold.
Osi et al. 31 predicted the survival outcomes of hospitalized COVID‐19 patients by comparing three machine learning algorithms (LDA, RF, and SVM) using patient demographic characteristics and clinical data. The best algorithm with 100% prediction accuracy belonged to the RF classifier, and LDA and SVM algorithms showed 95.2% and 90.9% accuracy, respectively.
In this study, comorbidities and old age had the lowest chance of survival for patients from COVID‐19 compared to other variables. The study showed that hypertension (57.5%) was the most common underlying disease and recorded a higher mortality rate, similar to the results.
Quiroz‐Juárez et al. 61 developed a ML model based on a multilayer feed‐forward network to determine the survival or death of COVID‐19 patients using 21 features including information about comorbidities, demographic information, and information about COVID‐19.
The results showed that compared to three other ML algorithms (SVM, LR, and KNN), this algorithm showed better performance and accuracy with accuracy = 93.5%, specificity = 90.9%, and sensitivity = 96.1%. Similar to the results of our study, the probability of death was higher in those whose treatment process led to intubation and ICU admission.
Examining previous studies and comparing their results shows that ML techniques have a wonderful ability to predict survival. By comparing the known effective factors in the COVID‐19 patients' survival with the factors identified in this study, it can be seen that the importance of risk factors related to the survival or mortality of the disease may have different importance from country to country and even from a geographical region to another region.
Figure 5 shows the application of the proposed model as a decision‐making module during medical decision‐making. The proposed model can assist health care teams to evaluate the severity status of COVID‐19 patients. The survival prediction models can only help health care providers in sorting ICU care for patients, not invalidate their judgment. Shortage of resources needed for health care in times of pandemics can be a serious challenge for hospitals. Prioritizing the treatment of patients is of great importance in the COVID‐19 pandemic. Deployment of the proposed model, experts can study the relationships among the patients' medical conditions (e.g., heart disease, and blood pressure, etc.), and the likelihood of death from COVID‐19. This allows them to pay more attention to the treatment of patients with critical conditions. Predicted probability by survival prediction models at baseline times, such as diagnosis or treatment can be used by the health care team to make important decisions about patient care, such as increasing the frequency of monitoring or implementing specific treatments which leads to the reduction of patient mortality. 69
Figure 5.

Practical application of the proposed model.
A hospital information system (HIS) is a comprehensive software program that integrates patient information and enables the communication between different departments of a hospital and other health care facilities. 70 Integrating the proposed machine learning model as a decision support module with HIS provides real‐time decision‐making based on the history of patients' recorded data.
4.1. Limitations
Most studies have used chest CT image and laboratory data sets to predict survival. Therefore, the comparison between the valuable risk factors in the survival of patients in different studies and our study was limited. Another limitation of this research is using a data set from one center. It is meaning that our proposed model's effectiveness in other data sets might differ. 71 Therefore, to develop a comprehensive survival model that can be used at the national level, it is necessary to use the data set of patients from different centers with more excellent geographical distribution in the model training process.
Another limitation of this study is that our database lacked patient treatment data. Features related to treatment, medicine, and procedures undertaken in the hospital could lead to more useful results. These features would be relevant to analyze since they can convey important information that can help health professionals to reduce death by COVID‐19. Unfortunately, these data were not available in the database. For this reason, in our proposed model, the influence of these factors on the survival of patients has been ignored.
5. CONCLUSION
To predict the chance of survival of patients with COVID‐19, the possibility of training ML algorithms on clinical data was investigated. Moreover, we evaluated the role of clinical data in the survival of hospitalized COVID‐19 patients. We believe that developing a decision support system based on the proposed model can help the health care team make important decisions about patient care, such as increasing the frequency of monitoring or implementing specific treatments, by identifying patients at risk of death. In future work, software development based on the NB model will be investigated. In the next step, research works should be conducted to obtain evidence based on which using death warning systems can lead to improved patient outcomes.
Further research are encouraged to try and apply more data sets instead of fixed data sets to generalize the survival models for COVID‐19 survival prediction.
AUTHOR CONTRIBUTIONS
Azita Yazdani: Conceptualization; methodology; supervision; writing—original draft; writing—review & editing. Somayeh Kianian Bigdeli: Methodology; writing—review & editing. Maryam Zahmatkeshan: Writing—original draft; writing—review & editing.
CONFLICT OF INTEREST STATEMENT
The authors declare no conflict of interest.
ETHICS STATEMENT
This study was extracted from a research supported financially by the Fasa University of Medical Sciences with the ethics code of IR.FUMS.REC.1399.059. https://ethics.research.ac.ir/ProposalCertificateEn.php?id=136619&Print=true&NoPrintHeader=true&NoPrintFooter=true&NoPrintPageBorder=true&LetterPrint=true
TRANSPARENCY STATEMENT
The lead author Maryam Zahmatkeshan affirms that this manuscript is an honest, accurate, and transparent account of the study being reported; that no important aspects of the study have been omitted; and that any discrepancies from the study as planned (and, if relevant, registered) have been explained.
ACKNOWLEDGMENTS
In this section, Fasa University of Medical Sciences was thanked for providing the data.
Yazdani A, Bigdeli SK, Zahmatkeshan M. Investigating the performance of machine learning algorithms in predicting the survival of COVID‐19 patients: a cross section study of Iran. Health Sci Rep. 2023;6:e1212. 10.1002/hsr2.1212
DATA AVAILABILITY STATEMENT
All the data of this research are available to Zahmatkeshan M and she is responsible for the accuracy of the data and the accuracy of the data analysis. Ethical permission has not been obtained to provide the data set used to other researchers. Individuals' demographic information was concealed to protect privacy and confidentiality in data collection.
REFERENCES
- 1. Afrash MR, Kazemi‐Arpanahi H, Shanbehzadeh M, Nopour R, Mirbagheri E. Predicting hospital readmission risk in patients with COVID‐19: a machine learning approach. Inform Med Unlocked. 2022;30:100908. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Chowdhury SF, Sium SMA, Anwar S. Research and management of rare diseases in the COVID‐19 pandemic era: challenges and countermeasures. Front Public Health. 2021;9:640282. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Sandhu R, Gill HK, Sood SK. Smart monitoring and controlling of pandemic influenza A (H1N1) using social network analysis and cloud computing. J Comput Sci. 2016;12:11‐22. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Singh S, Bansal A, Sandhu R, Sidhu J. Fog computing and IoT based healthcare support service for dengue fever. Int J Pervasive Comput Commun. 2018;14:197‐207. [Google Scholar]
- 5. WHO . Coronavirus disease (COVID‐19) pandemic. 2022. https://www.euro.who.int/en/health-topics/health-emergencies/coronavirus-covid-19
- 6. Lu H, Stratton CW, Tang YW. Outbreak of pneumonia of unknown etiology in Wuhan, China: the mystery and the miracle. J Med Virol. 2020;92:401‐414. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Shanbehzadeh M, Nopour R, Kazemi‐Arpanahi H. Developing an artificial neural network for detecting COVID‐19 disease. J Educ Health Promot. 2022;11:2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Erfannia L, Sharifian R, Yazdani A, Sarsarshahi A, Rahati R, Jahangiri S. Students' satisfaction and e‐learning courses in Covid‐19 pandemic era: a case study. Stud Health Technol Inform. 2022;289:180‐183. [DOI] [PubMed] [Google Scholar]
- 9. WHO . WHO Coronavirus (COVID‐19) Dashboard. WHO; 2022. https://covid19.who.int/ [Google Scholar]
- 10. WHO . Iran (Islamic Republic of), Statistics. WHO; 2022. https://www.who.int/countries/irn/ [Google Scholar]
- 11. Zhai Y, Wang Y, Zhang M, et al. From isolation to coordination: how can telemedicine help combat the COVID‐19 outbreak? MedRxiv; 2020. [Google Scholar]
- 12. Jain S, Nehra M, Kumar R, et al. Internet of medical things (IoMT)‐integrated biosensors for point‐of‐care testing of infectious diseases. Biosens Bioelectron. 2021;179:113074. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Churpek MM, Snyder A, Sokol S, Pettit NN, Edelson DP. Investigating the impact of different suspicion of infection criteria on the accuracy of quick sepsis‐related organ failure assessment, systemic inflammatory response syndrome, and early warning scores. Crit Care Med. 2017;45(11):1805‐1812. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Green M, Lander H, Snyder A, Hudson P, Churpek M, Edelson D. Comparison of the between the flags calling criteria to the MEWS, NEWS and the electronic Cardiac Arrest Risk Triage (eCART) score for the identification of deteriorating ward patients. Resuscitation. 2018;123:86‐91. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Henry KE, Hager DN, Pronovost PJ, Saria S. A targeted real‐time early warning score (TREWScore) for septic shock. Sci Transl Med. 2015;7(299):299ra122. [DOI] [PubMed] [Google Scholar]
- 16. Shafiei F, Fekri‐Ershad S. Detection of lung cancer tumor in CT scan images using novel combination of super pixel and active contour algorithms. Traitement du Signal. 2020;37(6):1029‐1035. [Google Scholar]
- 17. Wang J, Yu H, Hua Q, et al. A descriptive study of random forest algorithm for predicting COVID‐19 patients outcome. PeerJ. 2020;8:e9945. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Al‐Turaiki I, Alshahrani M, Almutairi T. Building predictive models for MERS‐CoV infections using data mining techniques. J Infect Public Health. 2016;9(6):744‐748. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Moulaei K, Shanbehzadeh M, Mohammadi‐Taghiabad Z, Kazemi‐Arpanahi H. Comparing machine learning algorithms for predicting COVID‐19 mortality. BMC Med Inform Decis Mak. 2022;22(1):2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Khan M, Mehran MT, Haq ZU, et al. Applications of artificial intelligence in COVID‐19 pandemic: a comprehensive review. Expert Syst Appl. 2021;185:115695. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. AlMoammar A, AlHenaki L, Kurdi H, eds. Selecting accurate classifier models for a MERS‐CoV dataset. Proceedings of SAI Intelligent Systems Conference. Springer; 2018. [Google Scholar]
- 22. Almalki YE, Qayyum A, Irfan M, Haider N, Glowacz A, Alshehri FM, eds. A novel method for COVID‐19 diagnosis using artificial intelligence in chest X‐ray images. Healthcare. Multidisciplinary Digital Publishing Institute; 2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Irfan M, Iftikhar MA, Yasin S, et al. Role of hybrid deep neural networks (HDNNs), computed tomography, and chest X‐rays for the detection of COVID‐19. Int J Environ Res Public Health. 2021;18(6):3056. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Belostecinic G, Mogoș RI, Popescu ML, et al. Teleworking—an economic and social impact during COVID‐19 pandemic: a data mining analysis. Int J Environ Res Public Health. 2022;19(1):298. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Lavrač N, ed. Machine learning for data mining in medicine. Joint European Conference on Artificial Intelligence in Medicine and Medical Decision Making. Springer; 1999. [Google Scholar]
- 26. Pushparaj MS. Machine learning for data mining in medicine: SHIVAJI UNIVERSITY, KOLHAPUR. 2008.
- 27. Teng X, Gong Y, eds. Research on application of machine learning in data mining. IOP Conference Series: Materials Science and Engineering. IOP Publishing; 2018. [Google Scholar]
- 28. Nemati M, Ansary J, Nemati N. Machine‐learning approaches in COVID‐19 survival analysis and discharge‐time likelihood prediction using clinical data. Patterns. 2020;1(5):100074. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Civit‐Masot J, Luna‐Perejón F, Domínguez Morales M, Civit A. Deep learning system for COVID‐19 diagnosis aid using X‐ray pulmonary images. Appl Sci. 2020;10(13):4640. [Google Scholar]
- 30. Shanbehzadeh M, Haghiri H, Afrash MR, Amraei M, Erfannia L, Kazemi‐Arpanahi H. Comparison of machine learning tools for the prediction of ICU admission in COVID‐19 hospitalized patients. Shiraz E‐Med J. 2022;23(5):e117849. [Google Scholar]
- 31. Osi AA, Abdu M, Muhammad U, Ibrahim A, Isma'il LA, Suleiman AA. A classification approach for predicting COVID‐19 patient's survival outcome with machine learning techniques. medRxiv; 2020. [Google Scholar]
- 32. Reddy CK, Li Y, Aggarwal C. A review of clinical prediction models. Healthcare Data Analy. 2015;36:343‐378. [Google Scholar]
- 33. Yan L, Zhang H‐T, Xiao Y, et al. Prediction of survival for severe Covid‐19 patients with three clinical features: development of a machine learning‐based prognostic model with clinical data in Wuhan. medRxiv; 2020. [Google Scholar]
- 34. Kim G, Yoo CD, Yang SJ. Survival analysis of COVID‐19 patients with symptoms information by machine learning algorithms. IEEE Access. 2022;10:62282‐62291. [Google Scholar]
- 35. Sarkar T. How to use python data science packages more productively. Productive and Efficient Data Science with Python. Springer; 2022:47‐84. [Google Scholar]
- 36. D'Angela A. Why weight? The importance of training on balanced datasets. Accessed January 28, 2021. https://towardsdatascience.com/why-weight-the-importance-of-training-on-balanced-datasets-f1e54688e7df
- 37. Poonia RC, Gupta MK, Abunadi I, Albraikan AA, Al‐Wesabi FN, Hamza MA, eds. Intelligent diagnostic prediction and classification models for detection of kidney disease. Healthcare. MDPI; 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Shen C, Panda S, Vogelstein JT. The chi‐square test of distance correlation. J Comput Graph Stat. 2022;31(1):254‐262. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Nugroho WH, Handoyo S, Akri YJ, Sulistyono AD. Building multiclass classification model of logistic regression and decision tree using the chi‐square test for variable selection method. J Hunan Univ Nat Sci. 2022;49(4):172‐181. [Google Scholar]
- 40. Muhammad L, Haruna AA, Mohammed IA, Abubakar M, Badamasi BG, Amshi JM, eds. Performance evaluation of classification data mining algorithms on coronary artery disease dataset. 9th International Conference on Computer and Knowledge Engineering (ICCKE). IEEE; 2019. [Google Scholar]
- 41. Muhammad LJ, Salisu S, Yakubu A, et al. Using decision tree data mining algorithm to predict causes of road traffic accidents, its prone locations and time along Kano–Wudil highway. Int J Database Theory Appl. 2017;10(11):197‐206. [Google Scholar]
- 42. Bhattacharya S, Agarwala A, Roy S. Mood detection and prediction using conventional machine learning techniques on COVID19 data. Soc Netw Anal Min. 2022;12(1):139. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Breiman L. Random forests. Mach Learn. 2001;45(1):5‐32. [Google Scholar]
- 44. Cutler DR, Edwards TC Jr., Beard KH, et al. Random forests for classification in ecology. Ecology. 2007;88(11):2783‐2792. [DOI] [PubMed] [Google Scholar]
- 45. Hemalatha M. A hybrid random forest deep learning classifier empowered edge cloud architecture for COVID‐19 and pneumonia detection. Expert Syst Appl. 2022;210:118227. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46. Darvishan A, Bakhshi H, Madadkhani M, Mir M, Bemani A. Application of MLP‐ANN as a novel predictive method for prediction of the higher heating value of biomass in terms of ultimate analysis. Energ Source, Part A: Recovery, Util Environ Effects. 2018;40(24):2960‐2966. [Google Scholar]
- 47. Ghafouri‐Fard S, Mohammad‐Rahimi H, Motie P, Minabi MAS, Taheri M, Nateghinia S. Application of machine learning in the prediction of COVID‐19 daily new cases: a scoping review. Heliyon. 2021;7(10):e08143. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48. Zhang W, Chen X, Liu Y, Xi Q. A distributed storage and computation k‐nearest neighbor algorithm based cloud‐edge computing for cyber‐physical‐social systems. IEEE Access. 2020;8:50118‐50130. [Google Scholar]
- 49. Devi T, Gopalan K. A statistical model of COVID‐19 infection incidence in the Southern Indian State of Tamil Nadu. Int J Environ Res Public Health. 2022;19(17):11137. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50. Mansour NA, Saleh AI, Badawy M, Ali HA. Accurate detection of Covid‐19 patients based on feature correlated naïve Bayes (FCNB) classification strategy. J Ambient Intell Humaniz Comput. 2022;13(1):41‐73. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51. Oğuz Ç, Yağanoğlu M. Detection of COVID‐19 using deep learning techniques and classification methods. Inf Process Manage. 2022;59:103025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52. Halimu C, Kasem A, Newaz SS, eds. Empirical comparison of area under ROC curve (AUC) and Mathew correlation coefficient (MCC) for evaluating machine learning algorithms on imbalanced datasets for binary classification. Proceedings of the 3rd international conference on machine learning and soft computing; 2019. [Google Scholar]
- 53. Villavicencio CN, Macrohon JJE, Inbaraj XA, Jeng J‐H, Hsieh J‐G. COVID‐19 prediction applying supervised machine learning algorithms with comparative analysis using WEKA. Algorithms. 2021;14(7):201. [Google Scholar]
- 54. Ahmed TU, Jamil MN, Hossain MS, Islam RU, Andersson K. An integrated deep learning and belief rule base intelligent system to predict survival of COVID‐19 patient under uncertainty. Cognit Comput. 2022;14(2):660‐676. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55. Atlam M, Torkey H, El‐Fishawy N, Salem H. Coronavirus disease 2019 (COVID‐19): survival analysis using deep learning and Cox regression model. Pattern Anal Appl. 2021;24(3):993‐1005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56. Cisterna‐Garcia A, Guillen‐Teruel A, Caracena M, et al. A predictive model for hospitalization and survival to COVID‐19 in a retrospective population‐based study. medRxiv; 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57. Ikemura K, Bellin E, Yagi Y, et al. Using automated machine learning to predict the mortality of patients with COVID‐19: prediction model development study. J Med Internet Res. 2021;23(2):e23458. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58. Khozeimeh F, Sharifrazi D, Izadi NH, et al. Combining a convolutional neural network with autoencoders to predict the survival chance of COVID‐19 patients. Sci Rep. 2021;11(1):15343. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59. Krajah A, Almadani YF, Saadeh H, Sleit A, eds. analyzing Covid‐19 data using various algorithms. IEEE Jordan International Joint Conference on Electrical Engineering and Information Technology (JEEIT). IEEE; 2021. [Google Scholar]
- 60. Muhammad LJ, Islam MM, Usman SS, Ayon SI. Predictive data mining models for novel coronavirus (COVID‐19) infected patients' recovery. SN Comput Sci. 2020;1(4):206. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61. Quiroz‐Juárez MA, Torres‐Gómez A, Hoyo‐Ulloa I, León‐Montiel RJ, U'Ren AB. Identification of high‐risk COVID‐19 patients using machine learning. PLoS One. 2021;16(9):e0257234. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62. Mohammad MA, Aljabri M, Aboulnour M, Mirza S, Alshobaiki A. Classifying the mortality of people with underlying health conditions affected by COVID‐19 using machine learning techniques. Appl Comput Intell Soft Comput. 2022;2022. [Google Scholar]
- 63. Castagna F, Kataria R, Madan S, et al. A history of heart failure is an independent risk factor for death in patients admitted with coronavirus 19 disease. J Cardiovasc Dev Dis. 2021;8(7):77. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64. Verity R, Okell LC, Dorigatti I, et al. Estimates of the severity of coronavirus disease 2019: a model‐based analysis. Lancet Infect Dis. 2020;20(6):669‐677. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65. Moulaei K, Ghasemian F, Bahaadinbeigy K, Ershad Sarbi R, Mohamadi Taghiabad Z. Predicting mortality of COVID‐19 patients based on data mining techniques. J Biomed Phys Eng. 2021;11(5):653‐662. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66. Huyut M. Automatic detection of severely and mildly infected COVID‐19 patients with supervised machine learning models. IRBM. 2022;44(1):100725. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67. Shanbehzadeh M, Orooji A, Kazemi‐Arpanahi H. Comparing of data mining techniques for predicting in‐hospital mortality among patients with covid‐19. J Biostat Epidemiol. 2021;7(2):154‐173. [Google Scholar]
- 68. Atiyah OS, Thalij SH. A comparison of Coivd‐19 cases classification based on machine learning approaches. IJEEE. 2022;18(1):139‐143. [Google Scholar]
- 69. Suresh K, Severn C, Ghosh D. Survival prediction models: an introduction to discrete‐time modeling. BMC Med Res Methodol. 2022;22(1):207. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70. Sadoughi F, Sarsarshahi A, Eerfannia I, Firouzabad SA. Ranking evaluation factors in hospital information systems. Hum Vet Med. 2016;8(2):92‐97. [Google Scholar]
- 71. Monjur O, Preo RB, Shams AB, Raihan MMS, Fairoz F. COVID‐19 prognosis and mortality risk predictions from symptoms: a cloud‐based smartphone application. BioMed. 2021;1(2):114‐125. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
All the data of this research are available to Zahmatkeshan M and she is responsible for the accuracy of the data and the accuracy of the data analysis. Ethical permission has not been obtained to provide the data set used to other researchers. Individuals' demographic information was concealed to protect privacy and confidentiality in data collection.
