Skip to main content
AMIA Annual Symposium Proceedings logoLink to AMIA Annual Symposium Proceedings
. 2018 Dec 5;2018:998–1007.

A Hybrid Residual Network and Long Short-Term Memory Method for Peptic Ulcer Bleeding Mortality Prediction

Qingxiong Tan 1, Andy Jinhua Ma 1,2, Huiqi Deng 1,2, Vincent Wai-Sun Wong 3, Yee-Kit Tse 3, Terry Cheuk-Fung Yip 3, Grace Lai-Hung Wong 3, Jessica Yuet-Ling Ching 3, Francis Ka-Leung Chan 3, Pong-Chi Yuen 1
PMCID: PMC6371275  PMID: 30815143

Abstract

The prediction of patient mortality, which can detect high-risk patients, is a significant yet challenging problem in medical informatics. Thanks to the wide adoption of electronic health records (EHRs), many data-driven methods have been proposed to forecast mortality. However, most existing methods do not consider correlations between static and dynamic data, which contain significant information about mutual influences between these data. In this paper, we utilize a deep Residual Network (ResNet) consisting of many convolution units, which can jointly analyze different variables, to capture correlation information in and between static and dynamic variables. Furthermore, the Long Short-Term Memory (LSTM) method is used to extract temporal dependencies information from dynamic data. Finally, a deep fusion method is used to integrate these different types of information to improve mortality prediction. Experiment results on Peptic Ulcer Bleeding (PUB) mortality prediction show that the proposed method outperforms existing methods and achieves an AUC (area under the receiver operating characteristic curve) score of 0.9353.

Introduction

The prediction of patient mortality, enabling detection of patients with high-risk of unfavorable clinical outcomes, is a significant yet challenging research problem. The wide application of electronic health records1 (EHRs) in hospitals has produced a large volume of detailed digital health data (diagnosis codes, laboratory parameters, medications, etc.) of patients. It provides an unprecedented opportunity for developing data-driven methods, especially deep learning techniques2,3, to promote clinical predictive research4,5, such as patient mortality prediction6,7, early detection of disease8,9, and clinical decision making10. EHRs often contain both static data (e.g., birthdate, gender, total doses of concomitant drugs) and dynamic sequence data (e.g., lab test results).

Medical dynamic time series are observations of several important indexes changing over time and contain significant information about changes in patient health status. Extracting temporal dependencies information from medical time series to build clinical prediction models has become a hot topic in recent years3,11,12,13. There are also some machine learning approaches which make use of static information to build medical prediction approaches14,15. However, further improvement of these single dynamic-information-based or single static-information-based methods with respect to their prediction accuracy is highly challenging, because of the inherently limited value of the information contained in the input variables.

Some new methods also combine static and dynamic data to build medical forecasting models16,17. However, these methods treat static and dynamic data as two isolated objects and simply concatenate features extracted from static data and features extracted from dynamic time series data. Then, this concatenation is mapped to the outputs without considering the correlations between the static and dynamic variables. However, the correlations between static and dynamic variables are important because of the relationship of mutual influences between them. Static variables affect the future trajectories of dynamic sequence variables. For example, differences in static variables, such as gender and birthdate, can correspond to differences among patients with respect to their personal physiques, leading in turn to differences in their dynamic time series. Conversely, dynamic time series can influence the value of some static variables. For example, a doctor may use the dynamic changes of a patient’s measurements to determine the patient’s total doses of drugs during treatment.

In this paper, we propose a hybrid Residual Network (ResNet) and Long Short-Term Memory (LSTM) method, which integrates correlation information between static and dynamic data, and temporal dependencies information contained in dynamic data, to improve Peptic Ulcer Bleeding (PUB) mortality prediction accuracy. Specifically, the deep ResNet contains many convolution units to analyze different variables jointly and has a residual learning framework to enable the ResNet to become deeper than traditional networks without the vanishing gradient problem. So it has stronger modeling capacity than traditional networks. We utilize the deep ResNet method to capture information about inner-channel correlations among different original static variables (static variables collected from EHRs directly), different new static variables (static variables refined from dynamic data), and moreover cross-channel correlations between original and new static variables. Furthermore, to extract temporal dependencies information from irregularly sampled PUB dynamic data, we modify the standard Dynamic Time Warping (DTW) model to align a large quantity of irregularly sampled time series. Then, the LSTM model, which avoids the vanishing gradient problem by controlling different gates in the memory cell, is used to model the aligned sequence data. Finally, we use a deep Multi-Residual Multi-Scale Network to fuse the correlation information extracted from diversity layers of ResNet and temporal dependencies information extracted from various layers of the LSTM to improve patient mortality prediction. Experiment results demonstrate that our hybrid ResNet and LSTM approach can well integrate different types of information to achieve higher prediction accuracy than other state-of-the-art methods.

Related Algorithms

Dynamic Time Warping Classical Dynamic Time Warping18 (DTW) is an algorithm designed to align a single pair of series. Suppose there are two time series T and O with lengths of n and m respectively:

T=t1,t2,⋯,ti,⋯,tnO=o1,o2,⋯,oj,⋯,om

To align these two time series, we build an n-by-m matrix, at which the (ith, jth) element contains the distance between ti and oj. The Euclidean distances are often used to measure this distance, namely d(ti, oj) = (ti - oj)2. The warping path W consisting of a contiguous set of matrix elements is used to define the aligning relationship between series T and O, whose kth element is defined as wk=(i,j)k. W can be represented as:

W=w1,w2,⋯,wk,⋯,wKmax(n,m)≤K<n+m−1

W is subjected to three conditions:

Boundary: w1 = (1,1) and wK = (n, m).

Monotonicity: Given wk = (i, j)k then wk−1=(i′,j′)k−1 where i−i′≥0 and j−j′≥0.

Continuity: Given wk=(i,j)k then wk−1=(i′,j′)k−1 where i−i′≤1 and j−j′≤1.

The optimal path can be efficiently found by using dynamic programming to minimize the dissimilarity distance of aligned series under the three conditions given above.

Deep Residual Network Deep Residual Networks19 are neural networks containing many basic blocks that consist of a residual module F, namely a stack of several layers, and a shortcut connection. By connecting several layers with shortcuts, the network can avoid the vanishing gradient problem, which is a problem that very deep plain neural networks often suffer from. As a result, the Residual Network can train very deep networks and gain stronger modeling capacity than other plain neural networks. Using Xi to represent the input of the ith basic block, the output is:

Yi=Fi(Xi)+Xi

where Fi(·) is the connection of some batch normalization units20, convolution units and rectified linear units.

Long Short-Term Memory Long Short-Term Memory3 (LSTM) is a modified version of recurrent neural networks (RNNs), specifically introduced to solve the vanishing gradient problem. LSTM has a memory unit to encode what knowledge has been learned, and learn when to forget and update hidden states when new information is given. The operation of the memory unit is controlled by three gates: input gate i, forget gate f, and output gate o. The update function and output are defined as follows:

it=σ(Wixxt+Wimmt−1+bi)ft=σ(Wfxxt+Wfmmt−1+bf)ot=σ(Woxxt+Wommt−1+b0)gt=φ(Wcxxt+Wcmmt−1+bc)ct=ft⊙ct−1+it⊙gtht=ot⊙ct

where σ(x)=(1+e−x)−1 is a sigmoid nonlinearity that maps inputs into the range of [0, 1] and φ(x) is a hyperbolic tangent nonlinearity. W are matrices that represent the parameters of the gates that are trained in the learning process and ⊙ means the product with the value of a gate. By controlling multiple gates, the LSTM method can prevent the vanishing gradient problem and capture temporal dependencies.

Methods

Figure 1 gives the overview of our proposed method for PUB patient mortality prediction. We first extract several types of important static information, which can represent the general situation of the time series data, from the irregularly sampled dynamic data and use a modified DTW method to align this huge amount of irregularly sampled PUB sequences data. Then, we utilize the deep Residual Network to capture information about correlations in and between the original static data and the new static data refined from the dynamic data. Furthermore, we use the LSTM method to extract temporal dependencies information from the aligned time series data. Finally, we utilize a deep fusion method to integrate these different types of information to improve mortality prediction.

Figure 1.

Figure 1.

The overview of the proposed method.

Data processing

The irregularly sampled PUB dynamic dataset causes difficulty in building a prediction model because time series with different sampling intervals are not comparable. We transform the original data into regularly spaced time series by calculating the average value of the time series in time intervals with equal length (in the experiments, we set the length of the time interval to half a year or one year, and compare their respective results). This averaging process helps us to capture the essential trends of original series, thus reducing sensitivity to measurement error and improving robustness. In the next step, we will extract several types of static information and align dynamic data.

1) Extracting Static Information

To mine correlations between static and dynamic variables, we extract several new static variables that can represent the general state of the dynamic data. This will allow us to capture information about correlations between static data and dynamic data by analyzing the relationships between the original and newly extracted static variables. The first variable we extract is the average value of each dynamic time series. This information is crucial because it provides a basic overall idea of the total time series. The second variable we extract is whether the dynamic series have missing data at different time intervals. Most existing methods solve the missing data problem by simply using zero or mean values to impute missing values16. In this study, we believe that missing data can also be valuable, because they reflect doctors’ judgments about the health of their patients. For example, if a doctor believes that a patient is in a good state, the doctor will most likely arrange for the patient to take fewer measurements to spare the patient pain and expense. Based on this consideration, we utilize moving window periods to generate labels indicating whether time series have missing data at different time intervals so as to describe the changing characteristics of the dynamic time series.

2) Time Series Alignment

The PUB dynamic series data have different lengths. However, the modeling of machine learning algorithms for time series data is often built on the assumption that the data are aligned. Temporal discrepancies in series data will lead to ill-generalizable models, which in turn affect forecasting performance. Thus, it is necessary for dynamic series data to be aligned. There are many alignment algorithms21, among which DTW is one of the most popular algorithms. DTW aims at aligning two sequences by warping the time axis iteratively until an optimal match is found. However, standard DTW is designed to align a single pair of time series22 and is not practical for multiple time series (MTS) alignment because it cannot ensure that different pairs of aligned time series have the same length. Although some enhanced methods have been proposed for MTS alignment23,24, they can only handle a small number of time series. However, for EHR dataset, the number of time series data is often thousands or even millions.

To solve this problem, we propose a simple but practical aligning framework based on two basic ideas: 1) For each dynamic feature, there exists a template time series that can represent the essential changing tendency of this feature for this specific group of people; 2) The standard DTW method is modified to align time series to the length of the template time series. To identify a suitable template time series, we first filter out dynamic data that are affected by the missing data problem, and then calculate mean values of remaining time series at different moving window periods and use these mean values as the template time series, which has strong robustness and can represent the essential changing trends of the chosen feature within this group of people. Then, we modify the standard DTW method by adjusting its continuity constraint. For instance, we set the step size as {[1, 0], [1, 1]}. Taking [1, 0] as an example, the value 1 means that the warping path moves one step along the direction of the template time series T, while the value 0 at the second element means that the warping path does not move along the direction of the time series O. It can be seen that all these step size settings have a value of 1 at the first element, which means that no matter whether the warping path is [1, 0] or [1, 1], it always has a sub-movement along the direction of T. Together with the boundary condition that the start and end cells of the warping path are the start and end points of these aligning time series, the length of aligned time series is equal to the length of T. This enables us to align a large quantity of PUB time series.

Extracting Correlation and Temporal Information

Through the above data processing, the data are now regularly spaced and have the same length. Next, we extract correlation and temporal information from the PUB data. First, we extract correlation information among different types of data. As mentioned above, the PUB data include both static and dynamic data. There are correlations both in and between different static and dynamic variables. Thus, we utilize a deep Residual Network, which has many convolution units to jointly analyze several variables, to capture information on the correlations among the original static variables and the new static variables refined from the dynamic data. Furthermore, Residual Network (ResNet) has a residual learning framework, which facilitates the training process and allows it to become deeper than traditional networks, and hence has stronger modeling ability than traditional methods. The architecture of the ResNet that we use to extract correlation features is given in Figure 2. The schemas of four kinds of basic blocks with different numbers of filters are shown in Figure 2(a)-(d) and the overall structure is given in Figure 2(e). Each basic block consists of a stack of 3 convolution layers, namely 1×1, 3×3, and 1×1 convolution units, and a shortcut connection. These convolution units can analyze several variables at the same time and thus capture the correlation information between static and dynamic variables. Meanwhile, the shortcuts ensure that these basic blocks can be stacked to increase the depth of the network, increasing the modeling capacity of the ResNet while avoiding the vanishing gradient problem. Furthermore, to ensure that the ResNet can extract enough correlation information features at the last few layers, several full connection layers with different numbers of units are connected to the averaging pool. We implement the ResNet in Keras platform and train the network for 100 epochs by using Adam25 to optimize the model parameters. The learning rate starts from 0.001 and then decays by 50% every 10 epochs.

Figure 2.

Figure 2.

The architecture of the Residual Network. (a)-(d) Schemas of four kinds of single block with different numbers of filters. (e) The overall Residual Network structure.

Second, we extract temporal information from the dynamic time series data. The temporal dependencies information in medical time series data can reflect important information about changes in patient conditions and hence help detect high-risk patients. However, traditional methods tend to summarize the sequence ensemble into aggregate features, which ignores the temporal relationships among different elements. To improve prediction accuracy, we utilize an LSTM method to extract temporal information from the time series data. The LSTM model has memory cells that capture information about what has been calculated so far and decide how to use this information for further calculations. That is, the model uses sequential information for modeling instead of assuming that all inputs are independent of each other as traditional neural networks do. Furthermore, by controlling the opening and closing of input gate, forget gate and output gate, the LSTM method avoids the vanishing gradient problem, thus capturing temporal dependencies from sequence data more accurately. This motivates our choice of utilizing LSTM to capture temporal information from dynamic medical series. The LSTM used in our method contains five LSTM layers, followed by seven full connection layers whose numbers of units are 512, 256, 128, 64, 32, 16, and 2 respectively. Similarly to the ResNet, the LSTM is also implemented in Keras platform and trained by using Adam. We train the LSTM for 500 epochs with learning rate starting from 0.001 and then decaying by 50% every 50 epochs.

Feature Fusion

With more layers, the deep neural networks can go deeper and extract features that have more informative relationships with the outputs. However, it is inevitable that some important information may be lost during information processing in the early layers of the deep neural networks. Considering this problem, we not only extract features from the last few layers of the deep neural network, but also from several shallow layers of the network, i.e., the correlation information and temporal dependencies information are extracted from multiple layers of the Residual Network method and the LSTM method respectively.

Figure 3 shows the schema for the Multi-Residual Multi-Scale Network that we use to fuse correlation and temporal features information. As shown in Figure 3(a), the basic unit of this fusion method is a Multi-Scale residual block. Unlike ordinary residual blocks, this block has different sizes of convolution units and thus can analyze different variables at multiple scales and from different aspects at the same time, thus fusing different features more comprehensively. Furthermore, as shown in Figure 3(b), this deep feature fusion method has a Multi-Residual network architecture. By increasing the number of residual functions, this deep Multi-Residual Multi-Scale feature fusion method becomes wider and can utilize more residual information in model learning, thus improving final prediction accuracy. Again, we implement the Multi-Residual Multi-Scale Network in Keras platform and use Adam as the optimizer for training the network. We train the network for 50 epochs. The learning rate starts at 0.0003 and is divided by 5 every 10 epochs.

Figure 3.

Figure 3.

The schema for the Multi-Residual Multi-Scale feature fusion method

Results

Dataset and Experimental Design

The experiments are performed on a Peptic Ulcer Bleeding (PUB) EHR dataset collected at the Endoscopy Center, Prince of Wales Hospital, Hong Kong, from January 1, 2007 to December 31, 2016. This dataset is a retrospective cohort of 6,367 patients who were diagnosed with PUB. The dataset contains 35 types of static variables (e.g., birthdate, gender, and total doses of some concomitant drugs that patients took during the treatment) and 7 types of dynamic time series of lab test results. We use these 35 types of static variables as original static data and the 7 types of lab test results as dynamic data. The goal of our research is to predict whether the patient dies within 10 years after being diagnosed with PUB. 48.77% of all the patients are mortality positive (i.e., the patient died). Accurate mortality prediction can help detect high-risk patients and optimize the patient management plan. We conduct this research using a two-fold cross-validation methodology. The experiments are evaluated using two indexes: receiver operator characteristic (ROC) curves and the area under the ROC curves (AUC).

As mentioned before, our experiments explore the influences on prediction accuracy of adding mean values and missing data labels to the original static data. Therefore, we use three kinds of combinations of static data: Static-I includes the original static data only; Static-II includes Static-I and mean values; Static-III includes Static-II and missing data labels. Our experiments also explore the influence of time interval length when resampling the irregularly sampled data into regularly spaced time series. We experimentally compare two resampling frequencies: Frequency-I (twice a year) and Frequency-II (once a year). As baseline methods we use Logistic Regression (LR) and Random Forests (RF), using the dynamic data with the resampling frequency of Frequency-II as the input. We implement these baseline methods using the scikit-learn26 package.

Time Series Alignment Results

Figure 4 shows an example of using the modified DTW method to align two time series. As shown in Figure 4(a), the original time series and template time series vary in speeds of changes. We can observe that both time series contain a trend of first decreasing and then increasing. For the template time series, this trend appears between 2015 and 2017. However, for the original time series, the trend appears between 2014 and 2016. This difference causes an unequal length problem, making it difficult to mine temporal dependencies information from the original time series directly. By utilizing the modified DTW method to align these series, the unequal length problem can be solved. As shown in Figure 4(b), the aligned time series have the same length and trend. At the same time, the aligned time series still preserve some important information about specific changes in the health status of the patient. This demonstrates that the modified DTW method can successfully align a large quantity of PUB dynamic series, such that they all have the same length as the template series and can well align irregular time series data into regular time series data. These capabilities are very important for the subsequent modeling using machine learning methods.

Figure 4.

Figure 4.

Example of using the modified DTW method to align time series. (a) Template time series and original time series before applying the modified DTW method. (b) Original time series is stretched to the length of the template time series after applying the modified DTW method.

Experiment Results

Table 1 demonstrates the AUC scores of different methods for the PUB patient mortality prediction. It can be seen that the proposed hybrid Residual Network and LSTM method achieves the best prediction performance among all of these methods with the highest AUC score, i.e., 0.9353. Furthermore, the statistical analysis demonstrates that the 95% confidence interval (CI) of the AUC scores of our proposed method is distributed between 0.9261 and 0.9440, which is higher than 95% CI of the AUC scores of other methods. These results statistically prove that our hybrid method outperforms other comparison methods. Similarly, by comparing the ROC curves of different methods, as given in Figure 5, we can see that the proposed hybrid method has better ROC curve than all the other comparison methods, which further demonstrates that the proposed method has the best mortality forecasting performance. All of these results support that by capturing the correlation information between different variables and integrating this information with temporal information, our proposed method can improve PUB mortality forecasting accuracy. Besides this, several other conclusions can also be obtained.

Table 1.

AUC scores of different methods

graphic file with name 2975069t1.jpg

Figure 5.

Figure 5.

ROC curves of different methods.

Firstly, prediction results of Residual Network (ResNet) for different static data are different. It can be observed from Table 1 that when the mean values of the dynamic time series are added as new static data, the AUC score increases from 0.8683 to 0.8937, which demonstrates that prediction accuracy increases after adding information about the mean values of the dynamic data. Similarly, as shown in Figure 5(a), the ROC curve of the ResNet for static variables becomes higher when the mean values of the dynamic time series are added as new static data. It can be inferred that the ResNet can jointly analyze static and dynamic variables to improve the prediction accuracy of PUB patient mortality. Furthermore, by comparing the prediction results of the ResNet for Static-II and Static-III, as shown in Table 1 and Figure 5(a), we can see that the ResNet has higher accuracy for the latter, which demonstrates that the use of missing-data labels can improve forecasting performance. This proves that missing data labels contain valuable information about the judgments of doctors regarding the health conditions of patients.

Secondly, the prediction results for data with different resampling resolutions are different. By comparing the ROC curves of the ResNet for Static-III with Frequency I and II (Figure 5(a)), of LSTM for dynamic variables with Frequency I and II (Figure 5(b)), and of our proposed hybrid method with Frequency II (Figure 5(d)), we can find that for each method, the ROC curve of a higher resampling resolution is better than for the corresponding method with a lower resampling resolution. Likewise, when the resampling resolution increases, the AUC scores of the same method also increase. This demonstrates that with a higher resampling resolution, more valuable information is preserved in the medical time series data regarding the dynamics of the disease, resulting in better prediction performance. For dynamic data with the same resampling frequency, it can be seen from Figure 5(c) that LSTM outperforms Logistic Regression (LR) and Random Forests (RF), which is consistent with the corresponding AUC scores obtained by these methods given in Table 1. These results demonstrate that deep learning models have stronger modeling capacity for the PUB clinical data than the traditional baseline models.

Thirdly, experiment results demonstrate that the integration of correlation information among different variables and temporal dependencies information in the dynamic time series data can increase the prediction performance for PUB patient mortality. From Figure 5(e) we can see that our proposed hybrid method has a higher ROC curve than the methods built on single correlation information or single temporal information. A similar trend can be observed from Figure 5(d) by comparing the ROC curves of different methods for data resampled at Frequency II. It can be concluded that for both resampling frequencies, the integration of correlation information and temporal dependencies information improves the mortality forecasting performance, and that the higher the resampling resolution, the better the forecasting.

Finally, experiment results demonstrate that our method outperforms two state-of-the-art approaches described in references [16] and [17], which also build prediction models utilizing both static and dynamic data. As shown in Figure 5(f), our method has a higher ROC curve than these two methods. The AUC score of our method is also larger. Both ROC curve and AUC score prove that our method outperforms these two state-of-the-art methods. It can be explained as follows. The methods in references [16] and [17] simply concatenate the static and dynamic features, and use a prediction layer to map these features to the output. In contrast, our method utilizes a deep ResNet to jointly analyze the original static variables and the new static variables extracted from the dynamic data to capture correlation information between these data. Therefore, although both our hybrid method and these two methods use the same data to build models, and both utilize similar methods, which all are special kinds of recurrent neural network (RNN), to extract temporal dependencies information, our proposed method utilizes the ResNet to capture correlation information and as a result achieves better mortality forecasting performance.

Discussion

In this study, we demonstrate that our proposed hybrid ResNet and LSTM method can successfully capture the correlation information between static data and dynamic data, and integrate this correlation information with temporal information extracted from dynamic data to improve mortality prediction of PUB patients. There are some critical discoveries in this study. Firstly, cross-channel correlations between static and dynamic data can improve forecasting performance. It is intuitively plausible that there should be correlations between static and dynamic variables, because they can influence each other. In recent years, although some methods that combine static and dynamic data to predict patient mortality have been proposed, these methods tend to simply concatenate static and dynamic features without considering their correlations. In our research, we utilize a ResNet with many convolution units to analyze different variables jointly, enabling us to capture correlation information between different data. Experiment results demonstrate that our method outperforms two other state-of-the-art methods, which combine static and dynamic data to build medical forecasting models without considering correlations between them. This proves that the mining of correlation information helps to improve mortality prediction performance, which can better detect patients with high risks of negative outcomes. Secondly, our research exploits the missing data phenomenon to improve mortality forecasting accuracy. This strongly implies that missing data contain information about the judgments of doctors regarding the conditions of patients. Experiment results prove the effectiveness of utilizing missing data labels to improve forecasting performance. Therefore, instead of simply using zero or mean value to implement missing values, we can find out the mechanism that causes the missing data phenomenon, and in turn make use of this phenomenon to improve our medical forecasting models. Thirdly, our research confirms an intuitive assumption about the influence of sampling resolution on the prediction performance, i.e., increasing sampling resolution can to some extent help reserve more valuable information about the conditions of patients and thus improve forecasting accuracy. Our experiments compare two resampling frequencies, i.e., twice a year and once a year, and find that models using twice-yearly resampled data achieve higher prediction accuracy. This suggests that if patients undergo health tests at the hospital more frequently, the recorded data will contain more information about their health conditions. Finally, our research proves that the integration of correlation information and temporal information can effectively improve the forecasting accuracy of patient mortality. It can be inferred that the EHR data contain multiple kinds of information, including temporal and correlation information. Mining and integrating different useful information will enable us to model changes of disease more accurately and obtain better clinical prediction results.

Limitations and Future Work

Despite promising performances for mortality prediction, there are still aspects in which the proposed method can be improved, which we will address in future work. Firstly, in this study, we only use numeric data records to build the models. Besides numeric datasets, there are also some text datasets, such as diagnoses records. The utilization of these records will tap new sources of information on patients’ conditions and improve prediction performance. In our future work, we will utilize text mining and natural language processing methods to extract useful information from text datasets of medical diagnoses and use this text information to improve patient health status prediction models. Secondly, we compare prediction performances for dynamic time series data resampled at two resampling frequencies, namely twice a year and once a year. However, we have not conducted experiments at higher resampling resolution. In our future work, we will increase the sampling resolution and compare corresponding prediction performances, to identify an optimal sampling frequency that can preserve all of the necessary information about the dynamics of the disease, while minimizing times that patients visit hospitals to record their health data, which can be expensive and painful. Thirdly, we believe that by performing sub-analyses on the population, we will be able to group patients into different types according to their disease characteristics and build a personalized method according to the characteristics of each sub-group of patients to improve prediction performance. However, due to the limited size of the PUB dataset, such sub-analyses would divide the dataset into relatively small groups, which may impair the training of the model. In our future work, we will perform sub-analyses on bigger clinical datasets to build more accurate personalized prediction methods.

Conclusion

The forecasting of patient mortality, which helps detect high-risk patients, is a significant yet challenging research problem. In this study, we propose a hybrid method that makes use of a deep Residual Network and the LSTM method to extract correlation information between static and dynamic data, and temporal dependencies information contained in dynamic time series data, and utilizes a deep Multi-Residual Multi-Scale Network to fuse these different types of information to improve mortality prediction of PUB patients. Experiment results demonstrate that our proposed hybrid approach outperforms several existing approaches.

References

  • 1.Pivovarov R, Elhadad N. Automated methods for the summarization of electronic health records. Journal of the American Medical Informatics Association. 2015;22(5):938–947. doi: 10.1093/jamia/ocv032. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Jagannatha AN, Yu H. Bidirectional RNN for medical event detection in electronic health records. In: Proceedings of the conference. Association for Computational Linguistics. North American Chapter. Meeting. NIH Public Access; 2016. pp. 473–482. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Lipton ZC, Kale DC, Elkan C, et al. 2015. Learning to diagnose with LSTM recurrent neural networks. arXiv preprint arXiv:1511.03677. [Google Scholar]
  • 4.Yadav P, Steinbach M, Kumar V, et al. Mining electronic health records (EHRs): a survey. ACM Computing Surveys (CSUR) 2018;50(6):85. [Google Scholar]
  • 5.Shickel B, Tighe P, Bihorac A, et al. 2017. Deep EHR: a survey of recent advances in deep learning techniques for electronic health record (EHR) analysis. arXiv preprint arXiv:1706.03446. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Walker AS, Mason A, Quan TP, et al. Mortality risks associated with emergency admissions during weekends and public holidays: an analysis of electronic health records. The Lancet. 2017;390(10089):62–72. doi: 10.1016/S0140-6736(17)30782-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Gong JJ, Naumann T, Szolovits P. Predicting clinical outcomes across changing electronic health record systems. In: Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM; 2017. pp. 1497–1505. [Google Scholar]
  • 8.Choi E, Schuetz A, Stewart WF, et al. Using recurrent neural network models for early detection of heart failure onset. Journal of the American Medical Informatics Association. 2016;24(2):361–370. doi: 10.1093/jamia/ocw112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Soguero-Ruiz C, Hindberg K, Rojo-Álvarez JL, et al. Support vector feature selection for early detection of anastomosis leakage from bag-of-words in electronic health records. IEEE Journal of Biomedical and Health Informatics. 2016;20(5):1404–1415. doi: 10.1109/JBHI.2014.2361688. [DOI] [PubMed] [Google Scholar]
  • 10.Visvanathan K, Levit LA, Raghavan D, et al. Untapped potential of observational research to inform clinical decision making: American Society of Clinical Oncology Research Statement. Journal of Clinical Oncology. 2017;35(16):1845–1854. doi: 10.1200/JCO.2017.72.6414. [DOI] [PubMed] [Google Scholar]
  • 11.Razavian N, Sontag D. 2015. Temporal convolutional neural networks for diagnosis from lab tests. arXiv preprint arXiv:1511.07938. [Google Scholar]
  • 12.Choi E, Bahadori MT, Schuetz A, et al. Doctor AI: Predicting clinical events via recurrent neural networks. In: Machine Learning for Healthcare Conference; 2016. pp. 301–318. [PMC free article] [PubMed] [Google Scholar]
  • 13.Yang S, Santillana M, Brownstein JS, et al. Using electronic health records and Internet search information for accurate influenza forecasting. BMC Infectious Diseases. 2017;17(1):332. doi: 10.1186/s12879-017-2424-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Wang J, Yang X, Cai H, et al. Discrimination of breast cancer with microcalcifications on mammography by deep learning. Scientific Reports. 2016;6:27327. doi: 10.1038/srep27327. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Gorunescu F, Belciug S. Evolutionary strategy to develop learning-based decision systems. Application to breast cancer and liver fibrosis stadialization. Journal of Biomedical Informatics. 2014;49:112–118. doi: 10.1016/j.jbi.2014.02.001. [DOI] [PubMed] [Google Scholar]
  • 16.Che Z, Purushotham S, Khemani R, et al. 2016. Interpretable deep models for icu outcome prediction. In: AMIA Annual Symposium Proceedings. American Medical Informatics Association; pp. 371–380. [PMC free article] [PubMed] [Google Scholar]
  • 17.Esteban C, Staeck O, Baier S, et al. Predicting clinical events by combining static and dynamic information using recurrent neural networks. In: 2016 IEEE International Conference on Healthcare Informatics (ICHI). IEEE; 2016. pp. 93–101. [Google Scholar]
  • 18.Keogh E, Ratanamahatana C A. Exact indexing of dynamic time warping. Knowledge and Information Systems. 2005;7(3):358–386. [Google Scholar]
  • 19.He K, Zhang X, Ren S, et al. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 2016. pp. 770–778. [Google Scholar]
  • 20.Ioffe S, Szegedy C. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In: International Conference on Machine Learning; 2015. pp. 448–456. [Google Scholar]
  • 21.Vu TN, Laukens K. Getting your peaks in line: a review of alignment methods for NMR spectral data. Metabolites. 2013;3(2):259–276. doi: 10.3390/metabo3020259. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Senin P. Vol. 855. USA: 2008. Dynamic time warping algorithm review. Information and Computer Science Department University of Hawaii at Manoa Honolulu; pp. 1–23. [Google Scholar]
  • 23.Xi X, Keogh E, Shelton C, et al. Fast time series classification using numerosity reduction. In: Proceedings of the 23rd International Conference on Machine Learning. ACM; 2006. pp. 1033–1040. [Google Scholar]
  • 24.Zhou F, De la Torre F. Generalized canonical time warping. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2016;38(2):279–294. doi: 10.1109/TPAMI.2015.2414429. [DOI] [PubMed] [Google Scholar]
  • 25.Kingma D P, Ba J. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980. [Google Scholar]
  • 26.Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research. 2011;12:2825–2830. [Google Scholar]

Articles from AMIA Annual Symposium Proceedings are provided here courtesy of American Medical Informatics Association

RESOURCES