Skip to main content
PLOS One logoLink to PLOS One
. 2021 Jun 24;16(6):e0253125. doi: 10.1371/journal.pone.0253125

Machine learning-based glucose prediction with use of continuous glucose and physical activity monitoring data: The Maastricht Study

William P T M van Doorn 1,2,#, Yuri D Foreman 1,3,#, Nicolaas C Schaper 1,4,5, Hans H C M Savelberg 6, Annemarie Koster 5,7, Carla J H van der Kallen 1,3, Anke Wesselius 8, Miranda T Schram 1,3,9, Ronald M A Henry 1,3,9, Pieter C Dagnelie 1,3, Bastiaan E de Galan 1,4,10, Otto Bekers 1,2, Coen D A Stehouwer 1,3, Steven J R Meex 1,2,, Martijn C G J Brouwers 1,4,‡,*
Editor: Chi-Hua Chen11
PMCID: PMC8224858  PMID: 34166426

Abstract

Background

Closed-loop insulin delivery systems, which integrate continuous glucose monitoring (CGM) and algorithms that continuously guide insulin dosing, have been shown to improve glycaemic control. The ability to predict future glucose values can further optimize such devices. In this study, we used machine learning to train models in predicting future glucose levels based on prior CGM and accelerometry data.

Methods

We used data from The Maastricht Study, an observational population‐based cohort that comprises individuals with normal glucose metabolism, prediabetes, or type 2 diabetes. We included individuals who underwent >48h of CGM (n = 851), most of whom (n = 540) simultaneously wore an accelerometer to assess physical activity. A random subset of individuals was used to train models in predicting glucose levels at 15- and 60-minute intervals based on either CGM data or both CGM and accelerometer data. In the remaining individuals, model performance was evaluated with root-mean-square error (RMSE), Spearman’s correlation coefficient (rho) and surveillance error grid. For a proof-of-concept translation, CGM-based prediction models were optimized and validated with the use of data from individuals with type 1 diabetes (OhioT1DM Dataset, n = 6).

Results

Models trained with CGM data were able to accurately predict glucose values at 15 (RMSE: 0.19mmol/L; rho: 0.96) and 60 minutes (RMSE: 0.59mmol/L, rho: 0.72). Model performance was comparable in individuals with type 2 diabetes. Incorporation of accelerometer data only slightly improved prediction. The error grid results indicated that model predictions were clinically safe (15 min: >99%, 60 min >98%). Our prediction models translated well to individuals with type 1 diabetes, which is reflected by high accuracy (RMSEs for 15 and 60 minutes of 0.43 and 1.73 mmol/L, respectively) and clinical safety (15 min: >99%, 60 min: >91%).

Conclusions

Machine learning-based models are able to accurately and safely predict glucose values at 15- and 60-minute intervals based on CGM data only. Future research should further optimize the models for implementation in closed-loop insulin delivery systems.

Introduction

The increasing prevalence of diabetes entails an increase in debilitating complications, such as retinopathy, neuropathy, and cardiovascular disease [13]. Maintaining plasma glucose levels within the reference range is essential for the prevention of diabetes-related complications, which are generally attributable to chronic hyperglycaemia, although hypoglycaemia has been suggested to contribute to cardiovascular disease risk as well [35]. One of the most promising developments to minimize hyperglycaemia and hypoglycaemia–and, hence, to increase time in range–in individuals with diabetes who require insulin treatment is a closed-loop insulin delivery system (also known as the artificial pancreas). Such a system integrates continuous glucose monitoring (CGM), insulin (with or without glucagon) infusion, and a control algorithm to continuously regulate blood glucose levels [6, 7]. Multiple studies have shown the merit of incorporating the artificial pancreas into clinical care of individuals with type 1 or type 2 diabetes [8, 9].

Despite prior efforts, there are still numerous points that need to be addressed in order to improve the individual components of closed-loop systems [6, 10]. With regard to CGM, this includes overcoming sensor delay (i.e., the inherent ~10-minute discrepancy between interstitially measured and actual plasma glucose values), and sensor malfunctions (i.e., periods during which no glucose values are recorded) [6, 10, 11]. Continuous glucose prediction is a potentially viable strategy to both handle sensor delay and bridge periods of sensor malfunction. The use of machine learning has yielded encouraging glucose prediction accuracy results in relatively small study populations (mostly individuals with type 1 diabetes) or in silico studies, as extensively reviewed elsewhere [12]. Large, human-based study populations are now needed to reliably assess to what extent and within what time interval (i.e., prediction horizon) glucose values can be accurately predicted by use of machine learning. Additionally, incorporation of physical activity, which is considered an important factor for glucose control in daily life, could further improve glucose prediction [6].

In this study, we investigated to what extent glucose values can be accurately predicted at intervals of 15 and 60 minutes by a machine learning model that has been trained with a sliding time window of glucose values preceding the predicted values at a fixed interval. Additionally, we studied whether glucose prediction can be further improved by incorporation of accelerometer-measured physical activity, and to what extent the results differ in a subgroup analysis of individuals with type 2 diabetes only. For this, we used a large population of individuals with either normal glucose metabolism (NGM), prediabetes, or type 2 diabetes who simultaneously underwent CGM and continuous accelerometry during a one-week period. Last, we used the publicly available OhioT1DM Dataset to explore whether CGM-based prediction models would translate to individuals with type 1 diabetes, the primary target population for closed-loop insulin delivery.

Methods

Study population and design

We used data from The Maastricht Study, an observational, prospective, population-based cohort study. The rationale and methodology have been described previously [13]. In brief, The Maastricht Study focuses on the aetiology, pathophysiology, complications and comorbidities of type 2 diabetes, and is characterized by an extensive phenotyping approach. All individuals aged between 40 and 75 years and living in the southern part of the Netherlands were eligible for participation. Participants were recruited through mass media campaigns and from the municipal registries and the regional Diabetes Patient Registry via mailings. For reasons of efficiency, recruitment was stratified according to known type 2 diabetes status, with an oversampling of individuals with type 2 diabetes. In general, the examinations of each participant were performed within a time window of three months. From 19 September 2016 until 13 September 2018, participants were invited to also undergo CGM [14]. During this period, a selected group of recently included participants were invited to return for CGM. In these participants only, there was a median time interval of 2.1 years between CGM and all other measurements. The present report includes cross-sectional data of the 851 participants who had at least 48h of CGM data available and were classified with NGM, prediabetes, or type 2 diabetes. The Maastricht Study has been approved by the institutional medical ethical committee (Medisch-ethische toetsingscommissie aZM/UM [METC]; NL31329.068.10) and the Minister of Health, Welfare and Sports of the Netherlands (Permit 131088-105234-PG). All participants gave written informed consent.

Continuous glucose monitoring

The rationale and methodology of CGM (iPro2 and Enlite Glucose Sensor; Medtronic, Tolochenaz, Switzerland) have been described previously [14]. In brief, the CGM device was worn abdominally and recorded subcutaneous interstitial glucose values (range: 2.2–22.2 mmol/L) every five minutes for a seven-day period. For calibration purposes, participants were asked to perform self-measurements of blood glucose four times daily (Contour Next; Ascensia Diabetes Care, Mijdrecht, the Netherlands). Participants were blinded to the CGM recording, but not to self-measured values. Diabetes medication use was allowed and no dietary instructions were given. We only included individuals with at least 48h of CGM, but excluded the first 24h of CGM from analysis because of insufficient calibration. For the glucose prediction analyses, all remaining glucose data points were used. We additionally calculated mean sensor glucose, standard deviation (SD), and coefficient of variation (CV) with the use of Glycemic Variability Research Tool (GlyVaRT; Medtronic) software.

Accelerometry

As described previously, daily physical activity was measured with use of the triaxial activPAL3 accelerometer (PAL technologies; Glasgow, United Kingdom) [13, 15]. The accelerometer was, just as the CGM device, attached during the first research visit; participants wore the accelerometer on the front of the right thigh for eight consecutive days. No physical activity instructions were given. PAL Software Suite version 8 (PAL technologies) was used to convert the event-based accelerometry data files into 15-second interval data files. We used the composite of X, Y, and Z accelerations for each 15-second interval as the measure of physical activity.

Assessment of participant characteristics

As described previously [13], we classified glucose metabolism status (GMS) as either NGM, prediabetes, or type 2 diabetes based on both a standardized 2-hour 75 gram oral glucose tolerance test and use of glucose-lowering medication [16]. We assessed medication use as part of a medication interview. Additionally, we determined smoking status and history of diabetes based on questionnaires, measured weight and height–to calculate body mass index (BMI)–and office blood pressure during a physical examination, and measured HbA1c as well as lipid profile in fasting venous blood.

Dataset construction

An overview of data preprocessing, model development, and model evaluation is given in Fig 1. In order to train our models in predicting future glucose values, we constructed two separate datasets (Fig 1, panel a). The first dataset consisted of only the participants’ six-day, five-minute interval CGM data (n = 851). The second dataset consisted of both CGM and accelerometry data (n = 540). To synchronize CGM (determined at 5-minute intervals) and accelerometry data (determined at 15-second intervals) in the second dataset, we linearly interpolated glucose values between two glucose data points with a frequency of 15 seconds. Consistent and aligned frequency intervals across these parameters are a statistical precondition for this type of model development [17]. The study populations were randomly split into a training (70%), tuning (10%), and evaluation (20%) dataset such that data from a given individual were present only in one set. The training set was used to train the proposed models. The tuning set was used to iteratively improve the models by selecting the best model architectures and hyperparameters. Finally, the best models were evaluated on the independent evaluation set that was retained during model development.

Fig 1. Overview of data preprocessing, model development and evaluation.

Fig 1

Data was used from The Maastricht Study, an observational population-based cohort that comprises individuals with normal glucose metabolism (NGM), prediabetes, or type 2 diabetes (panel A). We included 851 individuals who underwent continuous glucose monitoring (CGM), most of whom simultaneously wore an accelerometer to assess physical activity (X, Y, and Z accelerations). Models developed with the long-short term memory (LSTM) architecture were trained in predicting glucose levels at 15- and 60-minute intervals with either CGM data only (1) or both CGM and accelerometer data (2) (panel B). Finally, model performance was evaluated by glucose profile analysis, performance metrics (root-mean-square error [RMSE]; Spearman’s correlation coefficient [rho]; proportions), and clinical error grids (panel C).

Model development and design

Our proposed predictive model operates sequentially over CGM and accelerometry data (Fig 1, panel b). At each individual time point, 30 minutes of prior time series data were provided to the statistical model (e.g., six CGM-based glucose values), based on which it predicted glucose values at specified time intervals. For this study, we set these time intervals at 15 and 60 minutes. The nature of this prediction task can be solved by a variety of statistical and machine learning models. In the current study, we assessed autoregressive integrated moving average, support vector regression, gradient-boosting systems, shallow and deep multi-layer perceptron neural networks, and several recurrent neural network (RNN) architectures, including classical RNN [18, 19], gated recurrent units [20], long-short term memory (LSTM) networks [21], and all of its bi-directional variants [22, 23] (S1 File).

Model selection and training

The classical RNN architecture had superior performance at the 15-minute prediction interval (Table 1, RMSE: 0.485 [0.481–0.490]), whilst the LSTM network outperformed all other architectures at the 60-minute prediction interval (Table 1, RMSE: 0.941 [0.937–0.945]). Considering the performance of the LSTM network at a 15-minute prediction interval was nearly as good as the classical RNN, we selected the multi-task LSTM network among several alternatives as architecture of choice to continue our investigations(S1 File and Table 1). This architecture runs sequentially over time series data and is able to implicitly model the historical context of an individual by modifying an internal state through time. Specifically, we designed this architecture to predict both time intervals simultaneously, often referred to as “multi-task learning”, which aims to share knowledge amongst prediction tasks.

Table 1. Baseline statistical and machine learning model comparison for predicting glucose values.

Prediction window and baseline model CGM-based glucose prediction Combined glucose prediction
Rho RMSE, mmol/L Rho RMSE, mmol/L
15 minutes ARIMA 0.842 [0.837–0.848] 0.504 [0.490–0.518] 0.834 [0.829–0.840] 0.498 [0.492–0.505]
SVR 0.791 [0.781–0.802] 0.558 [0.549–0.567] 0.703 [0.694–0.712] 0.612 [0.601–0.622]
LightGBM 0.783 [0.767–0.795] 0.589 [0.577–0.601] 0.783 [0.771–0.794] 0.497 [0.582–0.613]
Shallow MLP 0.810 [0.804–0.816] 0.517 [0.506–0.529] 0.763 [0.754–0.772] 0.592 [0.581–0.603]
Deep MLP 0.807 [0.797–0.818] 0.511 [0.504–0.518] 0.828 [0.819–0.837] 0.510 [0.503–0.517]
RNN 0.894 [0.8870.902] 0.485 [0.4810.490] 0.890 [0.8820.898] 0.477 [0.4720.482]
LSTM 0.872 [0.865–0.879] 0.482 [0.477–0.487] 0.884 [0.878–0.890] 0.501 [0.496–0.506]
60 minutes ARIMA 0.307 [0.284–0.329] 1.543 [1.489–1.623] 0.303 [0.283–0.322] 1.502 [1.455–1.568]
SVR 0.388 [0.376–0.398] 1.386 [1.322–1.452] 0.394 [0.382–0.405] 1.412 [1.350–1.475]
LightGBM 0.500 [0.491–0.508] 1.118 [1.098–1.136] 0.498 [0.485–0.511] 1.128 [1.107–1.148]
Shallow MLP 0.503 [0.495–0.511] 1.081 [1.074–1.088] 0.483 [0.470–0.495] 1.081 [1.070–1.092]
Deep MLP 0.496 [0.484–0.509] 1.108 [1.100–1.115] 0.515 [0.502–0.528] 1.108 [1.099–1.017]
RNN 0.591 [0.581–0.600] 0.989 [0.983–0.995] 0.596 [0.589–0.603] 0.992 [0.984–0.998]
LSTM 0.605 [0.5930.616] 0.941 [0.9370.945] 0.602 [0.5950.609] 0.922 [0.9190.926]

Performance was assessed by Spearman’s rank correlation coefficient (rho) and root-mean-square error (RMSE). Data are reported as median [95% confidence intervals], calculated using 1,000 bootstraps.

Next, we evaluated a broad spectrum of hyperparameter combinations for this network (S1 Table). This resulted in a multi-task LSTM architecture, consisting of three layers, including a dropout layer with a total of 56–104 neurons (S2 Table). During training, we used exponential learning-rate decay via the Adam optimization scheme [24]. The best validation results were achieved by use of an initial learning rate with a decay of 0.001 every 1,000 training steps, with a batch size of 1024, and a back-propagation through a time window of 30 minutes. This defines the amount of historic data the model uses, which in our case translates to six (first dataset) or 120 (second dataset) glucose data points, for the model to provide a prediction. The loss function during training was the mean average of the mean-squared error function of all predictions. The maximum amount of epochs was 50.000 with an early stopping criterion (based on 20% hold-out data) set to 250 epochs. We performed data preprocessing, model development, selection, and training using Python programming language (version 3.7.1) with the use of packages Numpy (version 1.17), Pandas (version 0.24), Keras (version 2.2.2), Scikit-learn (version 0.22.0) and Tensorflow (version 2.0.1, beta).

Translation of the prediction models to the OhioT1DM Dataset

We used data from the OhioT1DM Dataset to explore whether our CGM-based prediction models would translate to individuals with type 1 diabetes. The OhioT1DM Dataset is freely available for scientific purposes and contains data of 6 individuals with type 1 diabetes who were all using insulin pump therapy and CGM [25]. The participants provided interstitial glucose values every five minutes for an eight-week period. First, in order to also include 30-minute prediction, we retrained our main CGM-based models on the main study population with identical hyperparameters and settings (S2 Table). Then, we evaluated the main CGM-based model on the test portion of the OhioT1DM Dataset (20%). Next, we aimed to optimize our main CGM-based model by training it on the train portion of the OhioT1DM Dataset. Specifically, we trained the model using an Adam optimizer with a learning rate of 10−4, a batch size of 1024, a maximum of 10.000 epochs and an early stopping criterion (based on 20% of the training data) set to 100 epochs. Last, we evaluated this optimized model on the test portion using performance metrics and safety error grids, as described previously.

Model evaluation and statistical analysis

Model evaluation was performed in the independent evaluation sets of individuals that were not used during model development (Fig 1, panel c). We employed several metrics to assess the performance of our models: root-mean-square error (RMSE), proportion of predicted values within 5% or 10% of actual glucose values, and Spearman’s rank correlation coefficient (rho) (S2 File). Bootstrapping was performed to obtain 95% confidence intervals for each of these metrics [26]. In addition, we used error grids that are classically used for assessment of blood glucose monitor safety (i.e., surveillance error grid, Parkes error grid) to evaluate the safety of our glucose prediction models [27, 28]. Last, we performed several sensitivity analysis in our main study population by stratifying model performance for: (1) GMS (i.e., separate results for NGM and prediabetes); (2) day (06.00 to 24.00h) and night (24.00 to 06.00h); and (3) low or high glucose variability, defined as the 97.5th percentile of CGM-assessed SD in individuals with NGM (SD > 1.37 mmol/L) [14].

Normally distributed data are presented as mean ± SD, non-normally distributed data as median and interquartile range, and categorical data as n (%). Statistical analyses were performed using the Statistical Package for Social Sciences (version 25.0; IBM, Chicago, Illinois, USA) and the Python programming language (version 3.7.1).

Results

Main study population characteristics

In total, 896 individuals underwent CGM as part of The Maastricht Study’s extensive phenotyping approach. We included participants with at least 48h of CGM data and either NGM, prediabetes, or type 2 diabetes. This resulted in the final study population of 851 individuals. Of this population, 540 participants (63.5%) simultaneously underwent CGM and accelerometry.

Table 2 shows the overall and type 2 diabetes-stratified characteristics of the two study populations (CGM-based as well as CGM- and accelerometry-based glucose prediction). The overall participant characteristics of both populations were generally comparable with regard to age, sex, BMI, glycaemic indices, blood pressure, and lipid profile, although the latter contained fewer participants with prediabetes or type 2 diabetes. Additionally, the participants with type 2 diabetes in the CGM- and accelerometry-based glucose prediction population were more often newly diagnosed with type 2 diabetes. Accordingly, these participants less often used glucose-lowering medication. Participant characteristics of the NGM and prediabetes subgroups are described in S3 Table.

Table 2. Participant characteristics of the CGM-based and CGM- and accelerometry-based glucose prediction study populations.

CGM-based glucose prediction CGM- and accelerometry-based glucose prediction
Characteristic Total (n = 851) T2D (n = 197) Total (n = 540) T2D (n = 68)
Age, years 59.9 ± 8.7 62.4 ± 7.8 59.1 ± 8.7 62.0 ± 6.9
Women, n (%) 418 (49.1) 69 (35.0) 276 (51.1) 22 (32.4)
BMI, kg/m2 27.2 ± 4.4 29.7 ± 4.7 26.5 ± 4.0 28.6 ± 4.1
Newly diagnosed T2D, n (%) 70 (8.2) 70 (35.5) 35 (6.5) 35 (51.5)
Glucose metabolism status
    NGM/PreD/T2D, n 470/184/197 - 372/99/68 -
    NGM/PreD/T2D, % 55.2/21.6/23.1 - 69.1/18.3/12.6 -
Fasting plasma glucose, mmol/L 5.4 [5.0–6.2] 7.3 [6.5–8.4] 5.3 [4.9–5.8] 7.2 [6.3–8.4]
2-h post-load glucose, mmol/L 6.7 13.6 6.2 12.5
[5.2–9.1] [11.7–16.2] [5.0–7.7] [11.3–16.6]
HbA1c, % 5.7 ± 0.8 6.7 ± 1.0 5.6 ± 0.6 6.4 ± 0.9
HbA1c, mmol/mol 39.1 ± 8.3 49.2 ± 10.8 37.3 ± 6.2 46.9 ± 10.2
Sensor glucose
    Mean, mmol/L 6.1 [5.7–6.7] 7.5 [6.8–8.7] 5.9 [5.6–6.4] 7.3 [6.5–8.2]
    SD, mmol/L 0.84 1.51 0.79 1.46
[0.68–1.18] [1.14–1.95] [0.66–1.01] [0.94–1.99]
    SD > 1.37 mmol/L, n (%) 142 (16.7) 115 (58.4) 50 (9.3) 36 (52.9)
    CV, % 14.0 19.3 13.3 19.2
[11.6–17.6] [15.9–24.0] [11.2–16.8] [14.5–24.1]
Diabetes medication use, n (%) 109 (12.8) 109 (55.6) 27 (4.8) 27 (39.7)
    Insulin 19 (2.2) 19 (9.6) 4 (0.7) 4 (5.9)
    Metformin 104 (12.2) 104 (53.1) 27 (5.0) 27 (39.7)
    Sulfonylureas 21 (2.5) 21 (10.7) 6 (1.1) 6 (8.8)
    Thiazolidinediones 0 (0) 0 (0) 0 (0) 0 (0)
    GLP-1 analogues 3 (0.4) 3 (1.5) 1 (0.2) 1 (1.5)
    DDP-4 inhibitors 1 (0.1) 1 (0.5) 0 (0) 0 (0)
    SGLT-2 inhibitors 1 (0.1) 1 (0.5) 0 (0) 0 (0)
Office SBP, mmHg 133.3 ± 18.0 139.4 ± 15.6 132.2 ± 17.9 137.7 ± 15.3
Office DBP, mmHg 75.2 ± 10.2 77.7 ± 10.5 74.7 ± 10.1 77.7 ± 9.6
Antihypertensive medication use, n (%) 305 (35.9) 126 (64.3) 162 (30.0) 41 (60.3)
Total-to-HDL cholesterol ratio 3.5 [2.8–4.3] 3.6 [2.9–4.3] 3.4 [2.8–4.3] 3.7 [2.8–4.6]
Triglycerides, mmol/L 1.3 [0.9–1.8] 1.5 [1.0–2.1] 1.2 [0.9–1.7] 1.6 [1.0–2.3]
Lipid-modifying medication use, n (%) 212 (24.9) 115 (58.4) 100 (18.5) 39 (57.4)
Smoking status
    Never/former/current, n 327/415/106 67/104/26 214/253/70 19/36/13
    Never/former/current, % 38.6/48.9/12.5 34.0/52.8/13.2 39.9/47.1/13.0 27.9/52.9/19.1

Data are reported as mean ± SD, median [interquartile range], or number (percentage [%]) as appropriate. CGM, continuous glucose monitoring; BMI, body mass index; T2D, type 2 diabetes; NGM, normal glucose metabolism; PreD, prediabetes; HbA1c, glycated haemoglobin A1c; SD, standard deviation; CV, coefficient of variation; GLP-1, glucagon-like peptide-1; DPP-4, dipeptidase-4; SGLT-2, sodium-glucose cotransporter 2; SBP, systolic blood pressure; DBP, diastolic blood pressure; HDL, high-density lipoprotein.

Overall performance of machine learning-based glucose prediction

We trained two machine learning models (i.e., CGM-based; CGM- and accelerometry-based) in predicting glucose levels at 15- and 60-minute intervals. Visually, both models appeared capable of accurately predicting the real glucose profiles, as illustrated by the representative examples in S1 and S2 Figs. Next, we assessed the performance of our models in our evaluation datasets with a variety of metrics, including an average error term (RMSE), the proportion of predictions within 5% or 10% deviation of the actual value, and correlation (rho). The evaluation datasets comprise 20% of the original or stratified study populations and thus vary in sample size (n = 13–170).

Overall, our models demonstrated high prediction accuracy, supported by low RMSE values and high proportions of predicted glucose values within 5% and 10% deviation (Table 3). Model performance in the type 2 diabetes subgroup was generally lower compared to the overall group, except for correlation coefficients, which were often higher in individuals with type 2 diabetes. This phenomenon can be largely attributed to the lower correlation coefficients of individuals with NGM and prediabetes (S4 Table), which is caused by range restriction (i.e., smaller glucose ranges attenuate the correlation coefficients) [29]. Consequently, the correlation coefficients are valid for the comparison of CGM-based glucose prediction to CGM- and accelerometry-based glucose prediction, but not for comparison of the overall study population to the type 2 diabetes subgroup. In addition, we observed short-to-moderate time lags for the 15- and 60-minute predictions (S5 Table).

Table 3. Overall performance in the main study population of CGM-based and CGM- and accelerometry-based machine learning models trained in predicting glucose values at time intervals of 15 and 60 minutes.

CGM-based glucose prediction CGM- and accelerometry-based glucose prediction
Total (n = 170) T2D (n = 43) Total (n = 109) T2D (n = 13)
15 minutes RMSE, mmol/L 0.188 [0.186–0.191] 0.288 [0.281–0.306] 0.184 [0.177–0.189] 0.271 [0.260–0.282]
< 5%, % 92.98 [92.87–93.05] 92.02 [91.83–92.25] 93.06 [93.03–93.09] 92.04 [91.99–92.11]
< 10%, % 99.17 [99.13–99.23] 98.88 [98.82–98.94] 99.25 [99.21–99.28] 98.90 [98.83–98.97]
Rho 0.961 [0.959–0.962] 0.987 [0.985–0.989] 0.968 [0.964–0.970] 0.990 [0.988–0.993]
60 minutes RMSE, mmol/L 0.589 [0.582–0.592] 0.701 [0.692–0.711] 0.582 [0.579–0.586] 0.700 [0.693–0.708]
< 5%, % 70.22 [70.09–70.41] 66.23 [66.13–66.33] 70.11 [70.05–70.17] 66.17 [66.09–66.22]
< 10%, % 87.39 [87.24–87.53] 85.82 [85.70–85.93] 87.44 [87.38–87.50] 86.11 [86.01–86.20]
Rho 0.721 [0.719–0.722] 0.781 [0.779–0.782] 0.725 [0.721–0.729] 0.790 [0.782–0.799]

Data are reported as mean [95% confidence interval]. CGM, continuous glucose monitoring; T2D, type 2 diabetes; RMSE, root-mean-square error; < 5%, percentage of predicted values within 5% of actual glucose values; < 10%, percentage of predicted values within 10% of actual glucose values; rho, Spearman’s rank correlation coefficient.

In general, incorporation of accelerometry data in the models only slightly improved performance metrics at both prediction intervals (Table 3). S4 Table shows the model performance in NGM and prediabetes subgroups. Glucose prediction was most precise in individuals with NGM. Of note, the ML-based models substantially outperformed a naive approach that used t0 as predicted glucose value (S6 Table, S3 and S4 Figs).

Safety evaluation with clinical error grids

We assessed the safety of our machine learning-based glucose prediction using two clinical error grids (i.e., surveillance and Parkes error grids). Fig 2 depicts the safety results for individuals with type 2 diabetes according to the surveillance error grid. At the 15-minute interval, almost all predictions (>99.9%) were clinically safe (i.e., a risk score between 0 and 1.0) (Fig 2, panels A and B). At the extended prediction window of 60 minutes, clinical safety was slightly lower (98.4–99.2%) (Fig 2, panels C and D). Parkes error grid assessment yielded similar results (S5 Fig). Of note, less accurate predictions were more often in the vertical B-D zones than in the horizontal B-E zones (e.g., S4 Fig, panel C: 11.80% versus 4.24%), which suggests a model tendency to underestimate rather than overestimate actual glucose values, the latter of which being more dangerous.

Fig 2. Surveillance error grid evaluation of glucose prediction safety at time intervals of 15 and 60 minutes in the main study population.

Fig 2

Assessment of CGM-based glucose prediction safety in individuals with type 2 diabetes (n = 43) at 15 minutes (panel A) and 60 minutes (panel C). Assessment of CGM- and accelerometry-based glucose prediction safety in individuals with type 2 diabetes (n = 13) at 15 minutes (panel B) and 60 minutes (panel D). The risk score values translate to the following degrees of risk: 0–0.5, none; 0.5–1.0, slight (lower); 1.0–1.5, slight (higher); 1.5–2.0, moderate (lower); 2.0–2.5, moderate (higher); 2.5–3.0, great (lower); 3.0–3.5, great (higher); > 3.5 extreme [27].

Additional analyses

To further obtain insights into our model predictions, we assessed performance metrics stratified by day and night (S7 Table). Fifteen-minute predictions did not materially differ between day and night. By contrast, accuracy of 60-minute predictions was lower during the day than at night. In addition, we stratified the results by high or low glucose variability (i.e., SD cut-off of 1.37 mmol/L) (S8 Table). Model performance was slightly lower at higher glucose variability, at both time intervals of 15 and 60 minutes.

Translation of the prediction models to the OhioT1DM Dataset

The prediction accuracy of the CGM-based model that was developed with our main study population was moderate in individuals with type 1 diabetes (RMSEs at 15, 30, and 60 min: 0.689 [0.685–0.693], 1.189 [1.183–1.195], and 1.918 [1.910–1.926] mmol/L), but substantially improved after being trained on data from each individual with type 1 diabetes (RMSEs at 15, 30, and 60 min: 0.426 [0.422–0.430], 1.046 [1.039–1.052], and 1.733 [1.725–1.741] mmol/L; S9 Table). Accordingly, clinical safety was substantial as shown by the high percentages of clinically safe predictions (15-minute: >99%, 30-minute: >97%, and 60-minute: >91%; Fig 3).

Fig 3. Surveillance error grid evaluation of glucose prediction safety at time intervals of 15, 30, and 60 minutes in individuals with type 1 diabetes.

Fig 3

Assessment of CGM-based glucose prediction safety in individuals with type 1 diabetes (n = 6) at 15 (panel A), 30 (panel B), and 60 minutes (panel C). The risk score values translate to the following degrees of risk: 0–0.5, none; 0.5–1.0, slight (lower); 1.0–1.5, slight (higher); 1.5–2.0, moderate (lower); 2.0–2.5, moderate (higher); 2.5–3.0, great (lower); 3.0–3.5, great (higher); > 3.5 extreme [27].

Discussion

In this study with 851 individuals and almost 1.4 million glucose measurements, we investigated whether glucose values can be accurately predicted by using machine learning-based models that utilise recently measured CGM and physical activity data with the prospect of improving closed-loop insulin delivery systems. Our study has several important findings and unique characteristics. First, the machine learning-based models are capable of accurately predicting the actual glucose profiles at 15 minutes, as reflected by several objective performance metrics (e.g., RMSE, rho; Table 2) and visual illustrations (S1 and S2 Figs). Despite prediction accuracy being moderately lower at 60 minutes, more than 98% of the predicted values remained sufficiently accurate to be deemed clinically safe based on surveillance error grids (Fig 2). Second, glucose prediction only improved slightly when accelerometer-assessed physical activity data was incorporated in the models. Third, translation of our CGM-based glucose prediction models to individuals with type 1 diabetes yielded encouraging results (i.e., ample prediction accuracy and clinical safety).

Although most research has thus far focused on type 1 diabetes [12], several efforts have been made to use machine learning for glucose prediction in individuals with type 2 diabetes [3034]. Most of these studies assessed technical aspects of glucose prediction in relatively small (n = 1 to 50) or even virtual, in silico populations. Such studies provide valuable comparisons of models, but show suboptimal and highly variable performance in predicting glucose values. To our knowledge, this is the first study to report this level of performance in a large, population-based sample of individuals with NGM, prediabetes, or type 2 diabetes. Our CGM-based models were able to accurately predict glucose values at 15 (RMSEs, overall/type 2 diabetes: 0.19/0.29 mmol/L) and 60 minutes (RMSEs, overall/type 2 diabetes: 0.59/0.70 mmol/L). These results surpass previously reported RMSE values for a sample of 50 individuals with type 2 diabetes, which were 0.65 and 1.50 mmol/L for 15- and 60-minute CGM-based glucose prediction, respectively [34]. We expect this difference to, in part, stem from our much larger sample size. To our knowledge, our exploratory translation to individuals with type 1 diabetes (S9 Table) showed that our models perform equally well as recent publications in the field [12, 3538]. For example, the best performing model of the Blood Glucose Level Prediction Challenge 2018, which was also based on a LSTM architecture as well as was trained on and evaluated in the OhioT1DM Dataset, reported 30-minute and 60-minute RMSEs of 1.05 and 1.74 mmol/L [35]. Additionally, Kriventsov et al. recently described large-scale application of glucose prediction in a smartphone app (Diabits) and reported a comparable RMSE at 30 minutes (1.04 mmol/L) [36]. We anticipate that further technical development of our prediction models, while using a larger sample of individuals with type 1 diabetes, will advance performance even more.

We integrated physical activity, which we assessed via accelerometry, into our glucose prediction model, because of its short- and long-term effects on daily glucose patterns. Whereas an acute bout of physical activity can either decrease or increase serum glucose levels, prolonged exercise improves insulin sensitivity, and thus insulin-stimulated glucose uptake [39]. While it should be noted that CGM- and accelerometry-based glucose prediction yielded larger improvements relative to CGM-based glucose prediction for the 60-minute interval, most notably during the day (S7 Table) and in individuals with higher glucose variability (S9 Table), incorporation of physical activity generally only marginally improved glucose prediction. This can be explained by the observation that the models based on CGM data only already performed very well, which limits the ability to achieve additional improvements [40]. Also, the effect of physical activity on serum glucose levels is relatively small in people with beta-cell function that is either normal or only mildly deficient. Given the absence of pancreatic glucoregulation in individuals with type 1 diabetes, it is conceivable that incorporation of accelerometry data leads to more substantially improved model performance in this patient group [40], which, at present, we were not able to further explore. In addition, a time interval of 15 or 60 minutes could be too short to incorporate long-term physical activity effects into the prediction model.

The closed-loop insulin delivery system has been shown to improve glycaemic control in individuals with type 1 or type 2 diabetes [8, 9, 41]. Nevertheless, several aspects of the artificial pancreas require further enhancement [6, 10]. Our results demonstrate that machine learning-based glucose prediction has the promise of being a valid and safe strategy to both overcome ~10-minute sensor delay and bridge prolonged periods of sensor malfunction. Not only are more than 99% of the predicted glucose values in clinically safe zones (i.e., Parkes error grid zone A and B), the model also tended to slightly underestimate rather than overestimate the actual glucose values. In case the prediction model were to be implemented, this would further reduce the risk of iatrogenic hypoglycaemia. Nevertheless, future research is needed to assess whether incorporation of these prediction models in a closed-loop insulin delivery system safely improves glycaemic control.

This proof-of-principle study has several strengths and limitations. Strengths are 1) the largest well-characterized, population-based study sample thus far, which ensured sufficient statistical power; 2) the unique large-scale combination of CGM and continuous accelerometry, which enabled us to study to what extent incorporation of data on physical activity would improve prediction in this population; 3) the gold-standard assessment of GMS, which allowed for the comparison of performance in NGM, prediabetes and type 2 diabetes; 4) the broad and solid evaluation of various statistical and machine learning architectures for this prediction task; and 5) result robustness, as reflected by the consistency of several statistical and clinical performance metrics.

Our research had certain limitations. First, the main study population comprised individuals with NGM, prediabetes, or type 2 diabetes, who are generally not the target population for closed-loop insulin delivery systems. We, therefore, exploratively investigated whether our prediction models would translate to individuals with type 1 diabetes using the OhioT1DM Dataset, which yielded encouraging results. Nevertheless, we underscore the importance of extensive evaluation of the models in a larger sample of individuals with type 1 diabetes, insulin-treated type 2 diabetes, or both. Second, we were unable to factor in other important elements pertaining to glycaemic control (e.g., diet or medication use) [6]. In automated, self-regulatory closed-loop systems, utilization of these kinds of data requires manual input, which is less convenient and reliable than CGM. In addition, since glucose prediction was only slightly improved by incorporating physical activity, we expect relatively little gain from including such factors into our models, at least in individuals with type 2 diabetes. However, given the results of several small studies that have incorporated diet and medication use [12], we acknowledge that this may not hold true for individuals with type 1 diabetes. In this regard, large-scale studies are required to reach more definitive conclusions. If diet, medication use, or other factors were to be incorporated, it is necessary to evaluate whether LSTM remains the best-performing machine learning architecture.

Conclusion

In this study, we show that our machine learning-based models are able to accurately and safely predict glucose values for up to 60 minutes in individuals with, NGM, prediabetes, or type 2 diabetes. In addition, translation of our prediction models to individuals with type 1 diabetes showed encouraging results. We observed particularly high precision at a 15-minute prediction window, which is a clinically relevant timespan to align interstitially measured glucose values by continuous glucose measurement systems with actual plasma glucose values. As such, the prediction model can be used to improve closed-loop insulin delivery systems by overcoming sensor delay. In addition, longer prediction intervals may be used to safely bridge periods of sensor malfunction. Last, our current findings question the use of accelerometry to substantially improve prediction. Future research should validate our findings by replicating the results in a larger sample of individuals with type 1 diabetes and studying the effects of implementing the prediction model in a closed-loop insulin delivery system.

Supporting information

S1 Fig. Illustrative examples of continuous glucose monitoring-based machine learning model predictions compared to actual values.

(DOCX)

S2 Fig. Illustrative examples of continuous glucose monitoring- and accelerometry-based machine learning model predictions compared to actual values.

(DOCX)

S3 Fig. Surveillance error grid evaluation of glucose prediction safety at time intervals of 15 and 60 minutes using glucose value t0 as predictor.

(DOCX)

S4 Fig. Performance characteristics of a prediction model using t0 as predictor across time horizons between 0 and 120 minutes.

(DOCX)

S5 Fig. Parkes error grid evaluation of glucose prediction safety at time intervals of 15 and 60 minutes.

(DOCX)

S1 Table. Hyperparameter combinations evaluated in current experiments.

(DOCX)

S2 Table. Final set of hyperparameters for each of the machine learning models.

(DOCX)

S3 Table. Extended baseline characteristics.

(DOCX)

S4 Table. Extended analysis of model performance in normal glucose metabolism and prediabetes subgroups.

(DOCX)

S5 Table. Extended analysis on time lag between predicted and actual glucose values.

(DOCX)

S6 Table. Extended analysis of model performance with t0 glucose value as predictor.

(DOCX)

S7 Table. Model performance stratified by day and night.

(DOCX)

S8 Table. Model performance stratified by low versus high glucose variability.

(DOCX)

S9 Table. Extended analysis of model performance in the Ohio T1DM Dataset.

(DOCX)

S1 File. Background information on machine learning models reviewed in current study.

(DOCX)

S2 File. Background information on metrics used in the current study.

(DOCX)

Acknowledgments

Prior presentation

An abstract of this study was submitted to the Annual Meeting of the European Association for the Study of Diabetes (Vienna, Austria, 21–25 September 2020). The conference abstract has been published by EMJ Diabetes.

Data Availability

Data are unsuitable for public deposition due to ethical restriction and privacy of participant data. The study has been approved by the medical ethical committee of the Maastricht University Medical Center (NL31329.068.10/ MEC 10-2-023) and the Netherlands Health Council under the Dutch “Law for Population Studies” (Permit 131088-105234-PG). Data are available from The Maastricht Study for any interested researchers who meet the criteria for access to confidential data. The Maastricht Study Management Team (research.dms@mumc.nl) and the corresponding author (Martijn C.G.J. Brouwers) may be contacted to request data.

Funding Statement

The Maastricht Study was supported by the European Regional Development Fund via OP-Zuid, the Province of Limburg, the Dutch Ministry of Economic Affairs (grant 31O.041), Stichting De Weijerhorst (Maastricht, the Netherlands), the Pearl String Initiative Diabetes (Amsterdam, the Netherlands), School for Cardiovascular Diseases (CARIM, Maastricht, the Netherlands), School for Public Health and Primary Care (CAPHRI, Maastricht, the Netherlands), School for Nutrition and Translational Research in Metabolism (NUTRIM, Maastricht, the Netherlands), Stichting Annadal (Maastricht, the Netherlands), Health Foundation Limburg (Maastricht, the Netherlands), and by unrestricted grants from Janssen-Cilag B.V. (Tilburg, the Netherlands), Novo Nordisk Farma B.V. (Alphen aan den Rijn, the Netherlands), Sanofi-Aventis Netherlands B.V. (Gouda, the Netherlands), and Medtronic (Tolochenaz, Switzerland). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

References

  • 1.Collaboration NCDRF. Worldwide trends in diabetes since 1980: a pooled analysis of 751 population-based studies with 4.4 million participants. Lancet. 2016;387(10027):1513–30. doi: 10.1016/S0140-6736(16)00618-8 ; PubMed Central PMCID: PMC5081106. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Emerging Risk Factors C, Sarwar N, Gao P, Seshasai SR, Gobin R, Kaptoge S, et al. Diabetes mellitus, fasting blood glucose concentration, and risk of vascular disease: a collaborative meta-analysis of 102 prospective studies. Lancet. 2010;375(9733):2215–22. doi: 10.1016/S0140-6736(10)60484-9 ; PubMed Central PMCID: PMC2904878. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Forbes JM, Cooper ME. Mechanisms of diabetic complications. Physiol Rev. 2013;93(1):137–88. doi: 10.1152/physrev.00045.2011 . [DOI] [PubMed] [Google Scholar]
  • 4.American Diabetes A. 6. Glycemic Targets: Standards of Medical Care in Diabetes-2019. Diabetes Care. 2019;42(Suppl 1):S61–S70. Epub 2018/12/19. doi: 10.2337/dc19-S006 . [DOI] [PubMed] [Google Scholar]
  • 5.International Hypoglycaemia Study G. Hypoglycaemia, cardiovascular disease, and mortality in diabetes: epidemiology, pathogenesis, and management. Lancet Diabetes Endocrinol. 2019;7(5):385–96. Epub 2019/03/31. doi: 10.1016/S2213-8587(18)30315-2 . [DOI] [PubMed] [Google Scholar]
  • 6.Cobelli C, Renard E, Kovatchev B. Artificial pancreas: past, present, future. Diabetes. 2011;60(11):2672–82. Epub 2011/10/26. doi: 10.2337/db11-0654 ; PubMed Central PMCID: PMC3198099. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Bruttomesso D. Toward Automated Insulin Delivery. N Engl J Med. 2019;381(18):1774–5. Epub 2019/10/17. doi: 10.1056/NEJMe1912822 . [DOI] [PubMed] [Google Scholar]
  • 8.Weisman A, Bai JW, Cardinez M, Kramer CK, Perkins BA. Effect of artificial pancreas systems on glycaemic control in patients with type 1 diabetes: a systematic review and meta-analysis of outpatient randomised controlled trials. Lancet Diabetes Endocrinol. 2017;5(7):501–12. Epub 2017/05/24. doi: 10.1016/S2213-8587(17)30167-5 . [DOI] [PubMed] [Google Scholar]
  • 9.Kumareswaran K, Thabit H, Leelarathna L, Caldwell K, Elleri D, Allen JM, et al. Feasibility of closed-loop insulin delivery in type 2 diabetes: a randomized controlled study. Diabetes Care. 2014;37(5):1198–203. Epub 2013/09/13. doi: 10.2337/dc13-1030 . [DOI] [PubMed] [Google Scholar]
  • 10.Blauw H, Keith-Hynes P, Koops R, DeVries JH. A Review of Safety and Design Requirements of the Artificial Pancreas. Ann Biomed Eng. 2016;44(11):3158–72. Epub 2016/11/04. doi: 10.1007/s10439-016-1679-2 ; PubMed Central PMCID: PMC5093196. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Rodbard D. Continuous Glucose Monitoring: A Review of Successes, Challenges, and Opportunities. Diabetes Technol Ther. 2016;18 Suppl 2:S3–S13. Epub 2016/01/20. doi: 10.1089/dia.2015.0417 ; PubMed Central PMCID: PMC4717493. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Woldaregay AZ, Arsand E, Walderhaug S, Albers D, Mamykina L, Botsis T, et al. Data-driven modeling and prediction of blood glucose dynamics: Machine learning applications in type 1 diabetes. Artif Intell Med. 2019;98:109–34. Epub 2019/08/07. doi: 10.1016/j.artmed.2019.07.007 . [DOI] [PubMed] [Google Scholar]
  • 13.Schram MT, Sep SJ, van der Kallen CJ, Dagnelie PC, Koster A, Schaper N, et al. The Maastricht Study: an extensive phenotyping study on determinants of type 2 diabetes, its complications and its comorbidities. Eur J Epidemiol. 2014;29(6):439–51. doi: 10.1007/s10654-014-9889-0 . [DOI] [PubMed] [Google Scholar]
  • 14.Foreman YD, Brouwers M, van der Kallen CJH, Pagen DME, van Greevenbroek MMJ, Henry RMA, et al. Glucose variability assessed with continuous glucose monitoring: reliability, reference values and correlations with established glycaemic indices—The Maastricht Study. Diabetes Technol Ther. 2019. Epub 2019/12/31. doi: 10.1089/dia.2019.0385 . [DOI] [PubMed] [Google Scholar]
  • 15.van der Berg JD, Stehouwer CD, Bosma H, van der Velde JH, Willems PJ, Savelberg HH, et al. Associations of total amount and patterns of sedentary behaviour with type 2 diabetes and the metabolic syndrome: The Maastricht Study. Diabetologia. 2016;59(4):709–18. Epub 2016/02/03. doi: 10.1007/s00125-015-3861-8 ; PubMed Central PMCID: PMC4779127. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.WHO. Definition and diagnosis of diabetes mellitus and intermediate hyperglycaemia: report of a WHO/IDF consultation. WHO. 2006. [Google Scholar]
  • 17.Staudemeyer RC, Rothstein Morris E. Understanding LSTM—a tutorial into Long Short-Term Memory Recurrent Neural Networks. arXiv e-prints [Internet]. 2019. September 01, 2019. Available from: https://ui.adsabs.harvard.edu/abs/2019arXiv190909586S. [Google Scholar]
  • 18.Sherstinsky A. Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) Network. arXiv e-prints. 2018:arXiv:1808.03314.
  • 19.Rumelhart DE, Hinton GE, Williams RJ. Learning representations by back-propagating errors. Nature. 1986;323(6088):533–6. doi: 10.1038/323533a0 [DOI] [Google Scholar]
  • 20.Chung J, Gulcehre C, Cho K, Bengio Y. Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. arXiv e-prints. 2014:arXiv:1412.3555.
  • 21.Hochreiter S, Schmidhuber J. Long Short-Term Memory. Neural Comput. 1997;9(8):1735–80. doi: 10.1162/neco.1997.9.8.1735 [DOI] [PubMed] [Google Scholar]
  • 22.Graves A, Fernández S, Schmidhuber J. Bidirectional LSTM Networks for Improved Phoneme Classification and Recognition 2005. 799–804 p. [Google Scholar]
  • 23.Schuster M, Paliwal K. Bidirectional recurrent neural networks. Signal Processing, IEEE Transactions on. 1997;45:2673–81. doi: 10.1109/78.650093 [DOI] [Google Scholar]
  • 24.Kingma DP, Ba J. Adam: A Method for Stochastic Optimization. arXiv e-prints [Internet]. 2014 December 01, 2014:[arXiv:1412.6980 p.]. Available from: https://ui.adsabs.harvard.edu/abs/2014arXiv1412.6980K.
  • 25.Marling C, Bunescu RC, editors. The OhioT1DM Dataset For Blood Glucose Level Prediction. KHD@IJCAI; 2018. [PMC free article] [PubMed]
  • 26.Efron B, Tibshirani RJ. An introduction to the bootstrap. New York, N.Y.; London: Chapman & Hall; 1993. [Google Scholar]
  • 27.Klonoff DC, Lias C, Vigersky R, Clarke W, Parkes JL, Sacks DB, et al. The surveillance error grid. J Diabetes Sci Technol. 2014;8(4):658–72. Epub 2015/01/07. doi: 10.1177/1932296814539589 ; PubMed Central PMCID: PMC4764212. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Pfutzner A, Klonoff DC, Pardo S, Parkes JL. Technical aspects of the Parkes error grid. J Diabetes Sci Technol. 2013;7(5):1275–81. Epub 2013/10/16. doi: 10.1177/193229681300700517 ; PubMed Central PMCID: PMC3876371. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Bland JM, Altman DG. Correlation in restricted ranges of data. BMJ. 2011;342:d556. doi: 10.1136/bmj.d556 . [DOI] [PubMed] [Google Scholar]
  • 30.Sudharsan B, Peeples M, Shomali M. Hypoglycemia prediction using machine learning models for patients with type 2 diabetes. J Diabetes Sci Technol. 2015;9(1):86–90. Epub 2014/10/16. doi: 10.1177/1932296814554260 ; PubMed Central PMCID: PMC4495530. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Georga E, Protopappas V, Fotiadis D. Glucose Prediction in Type 1 and Type 2 Diabetic Patients Using Data Driven Techniques. 2011. [Google Scholar]
  • 32.Faruqui SHA, Du Y, Meka R, Alaeddini A, Li C, Shirinkam S, et al. Development of a Deep Learning Model for Dynamic Forecasting of Blood Glucose Level for Type 2 Diabetes Mellitus: Secondary Analysis of a Randomized Controlled Trial. JMIR Mhealth Uhealth. 2019;7(11):e14452. Epub 2019/11/05. doi: 10.2196/14452 ; PubMed Central PMCID: PMC6858613. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Albers DJ, Levine M, Gluckman B, Ginsberg H, Hripcsak G, Mamykina L. Personalized glucose forecasting for type 2 diabetes using data assimilation. PLoS Comput Biol. 2017;13(4):e1005232. Epub 2017/04/28. doi: 10.1371/journal.pcbi.1005232 ; PubMed Central PMCID: PMC5409456. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Mohebbi A, Johansen AR, Hansen N, Christensen PE, Tarp JM, Jensen ML, et al. Short Term Blood Glucose Prediction based on Continuous Glucose Monitoring Data. arXiv e-prints [Internet]. 2020 February 01, 2020:[arXiv:2002.02805 p.]. Available from: https://ui.adsabs.harvard.edu/abs/2020arXiv200202805M. [DOI] [PubMed]
  • 35.Martinsson J, Schliep A, Eliasson B, Mogren O. Blood Glucose Prediction with Variance Estimation Using Recurrent Neural Networks. Journal of Healthcare Informatics Research. 2020;4(1):1–18. doi: 10.1007/s41666-019-00059-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Kriventsov S, Lindsey A, Hayeri A. The Diabits App for Smartphone-Assisted Predictive Monitoring of Glycemia in Patients With Diabetes: Retrospective Observational Study. JMIR Diabetes. 2020;5(3):e18660. Epub 2020/09/23. doi: 10.2196/18660 ; PubMed Central PMCID: PMC7539161. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Li K, Liu C, Zhu T, Herrero P, Georgiou P. GluNet: A Deep Learning Framework for Accurate Glucose Forecasting. IEEE J Biomed Health Inform. 2020;24(2):414–23. Epub 2019/08/02. doi: 10.1109/JBHI.2019.2931842 . [DOI] [PubMed] [Google Scholar]
  • 38.Chen J, Li K, Herrero P, Zhu T, Georgiou P, editors. Dilated Recurrent Neural Network for Short-time Prediction of Glucose Concentration. KHD@IJCAI; 2018.
  • 39.Stanford KI, Goodyear LJ. Exercise and type 2 diabetes: molecular mechanisms regulating glucose uptake in skeletal muscle. Adv Physiol Educ. 2014;38(4):308–14. Epub 2014/12/01. doi: 10.1152/advan.00080.2014 ; PubMed Central PMCID: PMC4315445. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Pencina MJ, D’Agostino RB, Pencina KM, Janssens AC, Greenland P. Interpreting incremental value of markers added to risk prediction models. Am J Epidemiol. 2012;176(6):473–81. Epub 2012/08/10. doi: 10.1093/aje/kws207 ; PubMed Central PMCID: PMC3530349. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Blauw H, Onvlee AJ, Klaassen M, van Bon AC, DeVries JH. Fully Closed Loop Glucose Control With a Bihormonal Artificial Pancreas in Adults With Type 1 Diabetes: An Outpatient, Randomized, Crossover Trial. Diabetes Care. 2021. Epub 2021/01/06. doi: 10.2337/dc20-2106 . [DOI] [PubMed] [Google Scholar]

Decision Letter 0

Chi-Hua Chen

17 Dec 2020

PONE-D-20-30681

Machine learning-based glucose prediction with use of continuous glucose and physical activity monitoring data: The Maastricht Study

PLOS ONE

Dear Dr. Brouwers,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please submit your revised manuscript by Jan 31 2021 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: http://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols

We look forward to receiving your revised manuscript.

Kind regards,

Chi-Hua Chen, Ph.D.

Academic Editor

PLOS ONE

Journal Requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2. Thank you for stating the following in the Financial Disclosure * (delete as necessary) section:

"The Maastricht Study was supported by the European Regional Development Fund via OP-Zuid, the Province of Limburg, the Dutch Ministry of Economic Affairs (grant 31O.041), Stichting De Weijerhorst (Maastricht, the Netherlands), the Pearl String Initiative Diabetes (Amsterdam, the Netherlands), School for Cardiovascular Diseases (CARIM, Maastricht, the Netherlands), School for Public Health and Primary Care (CAPHRI, Maastricht, the Netherlands), School for Nutrition and Translational Research in Metabolism (NUTRIM, Maastricht, the Netherlands), Stichting Annadal (Maastricht, the Netherlands), Health Foundation Limburg (Maastricht, the Netherlands), and by unrestricted grants from Janssen-Cilag B.V. (Tilburg, the Netherlands), Novo Nordisk Farma B.V. (Alphen aan den Rijn, the Netherlands), Sanofi-Aventis Netherlands B.V. (Gouda, the Netherlands), and Medtronic (Tolochenaz, Switzerland). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript."

We note that you received funding from a commercial source: Janssen-Cilag B.V., Novo Nordisk Farma B.V., Sanofi-Aventis Netherlands B.V., Medtronic.

Please provide an amended Competing Interests Statement that explicitly states this commercial funder, along with any other relevant declarations relating to employment, consultancy, patents, products in development, marketed products, etc.

Within this Competing Interests Statement, please confirm that this does not alter your adherence to all PLOS ONE policies on sharing data and materials by including the following statement: "This does not alter our adherence to PLOS ONE policies on sharing data and materials.” (as detailed online in our guide for authors http://journals.plos.org/plosone/s/competing-interests).  If there are restrictions on sharing of data and/or materials, please state these. Please note that we cannot proceed with consideration of your article until this information has been declared.

Please include your amended Competing Interests Statement within your cover letter. We will change the online submission form on your behalf.

Please know it is PLOS ONE policy for corresponding authors to declare, on behalf of all authors, all potential competing interests for the purposes of transparency. PLOS defines a competing interest as anything that interferes with, or could reasonably be perceived as interfering with, the full and objective presentation, peer review, editorial decision-making, or publication of research or non-research articles submitted to one of the journals. Competing interests can be financial or non-financial, professional, or personal. Competing interests can arise in relationship to an organization or another person. Please follow this link to our website for more details on competing interests: http://journals.plos.org/plosone/s/competing-interests

3. We note that you have indicated that data from this study are available upon request. PLOS only allows data to be available upon request if there are legal or ethical restrictions on sharing data publicly. For information on unacceptable data access restrictions, please see http://journals.plos.org/plosone/s/data-availability#loc-unacceptable-data-access-restrictions.

In your revised cover letter, please address the following prompts:

a) If there are ethical or legal restrictions on sharing a de-identified data set, please explain them in detail (e.g., data contain potentially identifying or sensitive patient information) and who has imposed them (e.g., an ethics committee). Please also provide contact information for a data access committee, ethics committee, or other institutional body to which data requests may be sent.

b) If there are no restrictions, please upload the minimal anonymized data set necessary to replicate your study findings as either Supporting Information files or to a stable, public repository and provide us with the relevant URLs, DOIs, or accession numbers. Please see http://www.bmj.com/content/340/bmj.c181.long for guidelines on how to de-identify and prepare clinical data for publication. For a list of acceptable repositories, please see http://journals.plos.org/plosone/s/data-availability#loc-recommended-repositories.

We will update your Data Availability statement on your behalf to reflect the information you provide.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Partly

Reviewer #3: No

**********

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #2: No

Reviewer #3: No

**********

3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: No

Reviewer #3: No

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: One issue of people with T2D using CGM is that, usually they do not need a CGM daily to monitor their glucose level all the time because they do not need to inject insulin like people with T1D. Please address this point to clarify the motivation and contribution of this work.

The references cited in this paper is not state-of-the-art. Many important related references using CGM, wearables in diabetes management using machine learning, are missing, such as

Data-driven modeling and prediction of blood glucose dynamics: Machine learning applications in type 1 diabetes

Convolutional recurrent neural networks for glucose prediction

Prediction of hypoglycemia during aerobic exercise in adults with type 1 diabetes

Normally people investigate the prediction of next 15, 30 and 60 mins. Why only 15 and 60 mins results were discussed in this paper

How to you deal with meal and insulin data in the prediction model? If they are not included, it seems the prediction can merely follow the trend of real glucose value to achieve an acceptable accuracy.

Not clear about the dataset. It says ‘From September 19, 2016 until September 13, 2018, participants were invited to undergo CGM.’ So how many dates of CGM data does the dataset have?

People cannot tell details in Figure S1, S2. Could you please zoom in so readers can see the difference? In addition, it is better to compare the results of different algorithms in figures.

How the extra accelerometer data contribute the accuracy of glucose prediction, co

mparing to the accuracy of sole CGM-based glucose prediction? Happy to see a concrete discussion to address this.

Besides RMSE, can you please calculate the time lag between the real and predicted glucose curve, in terms of different algorithms used in the paper? Because it is an important feature to measure the performance.

The results of RMSE (at 15 (RMSE:0.19mmol/L; rho:0.96) and 60 minutes (RMSE:0.59mmol/L, rho:0.72).), are too good to be true, from my point of view. For example, even give meal and insulin, exercise info, the RMSE of 60 mins prediction for T1D is larger than 30 mg/dL. For T2D the results will be better, but 0.59mmol/L is still very small. Can you compare your results to other existing algorithms, and convince readers that this good results is in feasible.

Reviewer #2: The study proposes a straight forward strategy of predicting blood glucose levels using ML models. The models are trained with a large dataset of 851 patients. The dataset contains data from T2Ds, prediabetics and normal individuals. The forecasting is done for a PH of 15 and 60 minutes. The results show almost perfect prediction, this is due to methodological errors.

The authors claim to have split training, cross-validation and test data randomly. This could prove to be a wrong strategy in time-series forecasting as there is a chance of the model getting trained on the future data.

The authors trained multiple models for prediction purpose. It is seen in the performance comparison table that classical RNN performs best for 15 min PH and LSTM performs best for 60 min PH. The manuscript, however, only contains details about the LSTM model.

Since, the proposed study does exactly the same what various other works have been doing for BG prediction during the past decade, no attempt at performance comparison with prior work has been made.

Performance improvement depicted in the CGM+PA dataset is not significant and hence provides no motive for designers to prefer one over the other.

Since it is understood that the glucose variability in NGM is low, and the number of individuals with NGM in both datasets are the largest, the underlying trends being identified by the ML model are overwhelmed by such data. It explains why the ML model are predicting almost perfectly.

Reviewer #3: This article presents the work on the application of different machine learning techniques for glucose prediction from CGM and physical activity bracelets.

This is a study with a very large number of patients but in my opinion the article is not interesting for the journal for several reasons.

Firstly, the data are not available to the research community, which makes it difficult to check whether the techniques presented can be overcome by the countless number of papers in the area.

Secondly, no new techniques are proposed, there are many studies in the area and the techniques of machine learning have been studied in depth, the authors can see for example all the articles in results were reported in different previous publications recently and some years ago. You can see, for instance the works presented at the last two workshops on Blood Glucose Level Prediction (BGLP) Challenge

http://ceur-ws.org/Vol-2675/

You can also find several journal papers

Hidalgo, J. I., Colmenar, J. M., Kronberger, G., Winkler, S. M., Garnica, O., & Lanchares, J. (2017). Data based prediction of blood glucose concentrations using evolutionary methods. Journal of medical systems, 41(9), 142.

Woldaregay, A. Z., Årsand, E., Walderhaug, S., Albers, D., Mamykina, L., Botsis, T., & Hartvigsen, G. (2019). Data-driven modeling and prediction of blood glucose dynamics: Machine learning applications in type 1 diabetes. Artificial intelligence in medicine, 98, 109-134.

Velasco, J. M., Garnica, O., Lanchares, J., Botella, M., & Hidalgo, J. I. (2018). Combining data augmentation, EDAs and grammatical evolution for blood glucose forecasting. Memetic Computing, 10(3), 267-277.

De Falco, I., Della Cioppa, A., Giugliano, A., Marcelli, A., Koutny, T., Krcma, M., ... & Tarantino, E. (2019). A genetic programming-based regression for extrapolating a blood glucose-dynamics model from interstitial glucose measurements and their first derivatives. Applied Soft Computing, 77, 316-328.

Contreras, I., & Vehi, J. (2018). Artificial intelligence for diabetes management and decision support: literature review. Journal of medical Internet research, 20(5),

And even more from the last two years on NNs and DL approaches

Moreover a prediction horizon of 15 minutes is so short that any Naive approach could reach a 95% of safe predictions, I recommend the authors the exercise of predicting the glucose value for 15 minutes as the value a t=0.

Experimental results are not useful as they are presented in the paper. All the techniques are summarized in just one table and no discrimination among them is done. The main conclusion of the paper is so general that is obvious. Is something like the affirmation " Medicine is good" or something similar.

Last but no least, the study affirm that, although it was made with T2 diabetes patients, it could be extrapolated to other T1 patients. I am sure that conclusions for T2 can not directly extrapolated to other type of patients. It has been shown in the past that in-silico results are not extensible to T1 real patients nor to T2 and vice versa. Glucose Variability of one T1 or T2 patients are different, T1 can produce little amounts of insulin or not, T2 insulin resistance could be heavier for one patient than for other....

So in this conditions the study is of little interest for the journal, In my humble opinion, I think that the data set has a great potential and that the research team is capable of prepare and in depth analysis of machine learning technique, I would recommend to separate and configure ML techniques for the different types of patients, and of course when presenting the results separate by ML techniques and data sets.

**********

6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

Reviewer #2: No

Reviewer #3: No

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

PLoS One. 2021 Jun 24;16(6):e0253125. doi: 10.1371/journal.pone.0253125.r002

Author response to Decision Letter 0


8 Feb 2021

We would like to thank the reviewers for their positive feedback on our study and for the time spent on our manuscript. We believe that their comments have given us the opportunity to substantially improve our manuscript. Notably, we have now included a first step of model validation in individuals with type 1 diabetes (OhioT1DM Dataset). As the reviewers can appreciate, our prediction models translate quite well to individuals with type 1 diabetes and are competitive with current studies in type 1 diabetes (S10 Table and Figure 3). Please find below our point-by-point rebuttal.

Reviewer #1

1. One issue of people with T2D using CGM is that, usually they do not need a CGM daily to monitor their glucose level all the time because they do not need to inject insulin like people with T1D. Please address this point to clarify the motivation and contribution of this work.

We certainly agree with the reviewer on this point. At present, most of the individuals with type 2 diabetes do not have an indication to wear CGM for a long time period. Furthermore, the number of individuals with type 2 diabetes who have an indication for a closed-loop insulin delivery system is even lower, although the use of such systems in type 2 diabetes has been investigated(1) and may become more frequent in the future.

Hence, we acknowledge that individuals with type 1 diabetes are, at present, the main target population for closed-loop insulin delivery systems, and had already detailed this in the limitations part of our Discussion (Page 18, Lines 395-397). We have now added the point that individuals with type 1 diabetes are the main population eligible for closed-loop insulin delivery to our Introduction (Page 6, Lines 103-105).

2. The references cited in this paper is not state-of-the-art. Many important related references using CGM, wearables in diabetes management using machine learning, are missing, such as Data-driven modeling and prediction of blood glucose dynamics: Machine learning applications in type 1 diabetes; Convolutional recurrent neural networks for glucose prediction; Prediction of hypoglycemia during aerobic exercise in adults with type 1 diabetes

Since we have now expanded our work to include translation of our prediction models to individuals with type 1 diabetes, and have updated our references accordingly, we would like to thank the reviewer for the literature suggestions.

3. Normally people investigate the prediction of next 15, 30 and 60 mins. Why only 15 and 60 mins results were discussed in this paper

Indeed, several previous studies have included 30 minutes as a prediction interval in addition to the 15 and 60 minutes that we used for our main results. Our reasoning for this was based on clinical applicability: 15 minutes to reflect overcoming sensor delay (i.e., the inherent ~10-minute discrepancy between interstitially measured and actual plasma glucose values) and 60 minutes to reflect bridging a relatively long period of sensor malfunction. We have outlined this in our Introduction (Pages 5-6, Lines 81-93). Moreover, if prediction accuracy and clinical safety are very high at 60 minutes, there is little reason to also investigate 30 or 45 minutes.

However, as advised by the reviewer, we have included the 30-minute interval for the prediction model validation in individuals with type 1 diabetes (S10 Table) in order to allow comparison with the current literature. This turned out to be more justified for this population, since clinical prediction safety was not as high at 60 minutes as compared to type 2 diabetes.

4. How to you deal with meal and insulin data in the prediction model? If they are not included, it seems the prediction can merely follow the trend of real glucose value to achieve an acceptable accuracy.

Unfortunately, we were unable to incorporate meal and therapy data in our prediction models, since they were not available for the seven-day recording period. We have discussed this in the Limitations section (Pages 18-19, Lines 421-431). We agree that incorporation of these data would seem logical from a physiological viewpoint. However, in automated, self-regulatory closed-loop systems, it would require manual input. Also, prediction in individuals with type 2 diabetes was accurate and safe to such an extent that addition of meal and therapy data is expected to lead to only a slight improvement of the prediction models. Indeed, the 15- and 60-minute prediction models are based on only previous glucose values (in case of the CGM-based approach), but we do not regard this as problematic, since the accuracy and clinical safety are nevertheless high. Still, we acknowledge that this may be different for individuals with type 1 diabetes (Page 18-19, Lines 417-421).

5. Not clear about the dataset. It says ‘From September 19, 2016 until September 13, 2018, participants were invited to undergo CGM.’ So how many dates of CGM data does the dataset have?

We apologize for the misunderstanding. The sentence referred to the total inclusion period of participants. All participants underwent seven-day CGM. We have rewritten this part of the study inclusion process and moved it to the Study population and design (Page 7, Lines 117-124).

6. People cannot tell details in Figure S1, S2. Could you please zoom in so readers can see the difference? In addition, it is better to compare the results of different algorithms in figures.

We adjusted S1 and S2 Figure to include a certain region that is zoomed in, so the actual and predicted glucose profiles can be examined. To ensure a fair comparison, we also reduced the line width of both profiles by a small margin.

7. How the extra accelerometer data contribute the accuracy of glucose prediction, comparing to the accuracy of sole CGM-based glucose prediction? Happy to see a concrete discussion to address this.

As already discussed on Pages 18-19 (Lines 378-393), we propose the following explanations for the very modest improvement in the prediction model after incorporating the accelerometer data: 1) the CGM-only models perform very well, and as such, substantial further improvement is very difficult to achieve; 2) the contribution of accelerometer data may physiologically be greater in type 1 diabetes (which, unfortunately, we were not able to investigate at present); and 3) the time intervals used may be too short to for the model to incorporate sustained physical activity effects into the prediction.

8. Besides RMSE, can you please calculate the time lag between the real and predicted glucose curve, in terms of different algorithms used in the paper? Because it is an important feature to measure the performance.

As suggested by the reviewer, we calculated the time lag between the real and predicted glucose value. We calculated the prediction time lag by measuring the time-shift that results in the highest cross correlation coefficient between them, according to the formula(2, 3):

τ_delay= 〖arg max┬k〗⁡〖 ((y_k ) ˇ(k│k-PH)*y(k))〗

These results have now been added to the supplemental materials (S6 Table) and we have updated our S2 Supporting information accordingly.

9. The results of RMSE (at 15 (RMSE:0.19mmol/L; rho:0.96) and 60 minutes (RMSE:0.59mmol/L, rho:0.72).), are too good to be true, from my point of view. For example, even give meal and insulin, exercise info, the RMSE of 60 mins prediction for T1D is larger than 30 mg/dL. For T2D the results will be better, but 0.59mmol/L is still very small. Can you compare your results to other existing algorithms, and convince readers that this good results is in feasible.

It should be acknowledged that our study was based on individuals with normal glucose metabolism (NGM), prediabetes, and type 2 diabetes, not type 1 diabetes. Therefore, we primarily compared our results to studies with a comparable study population (i.e., individuals with type 2 diabetes). The current literature in type 2 diabetes is limited, which may explain why our results seem so good. Nevertheless, when comparing our results to the best algorithm in type 2 diabetes published to date (with at least a sample size > 10 participants)(4), our prediction results are indeed better. We expect this to be mainly due to a large sample size difference (i.e., a larger sample yields better and more reliable prediction), which we have now added to the Discussion section (Page 18, Lines 365-366). Moreover, as our approach ensured that we evaluated our models in individuals completely retained from model development, we prevented our models from recognizing data patterns on which they were trained.

As we have now added an exploratory validation in individuals with type 1 diabetes, we have added a comparison of these result to the current literature on individuals with type 1 diabetes as well (Page 18, Lines 367-376). This shows that our findings in individuals with type 1 diabetes are comparable to the best studies in the field.

Reviewer #2

1. The study proposes a straight forward strategy of predicting blood glucose levels using ML models. The models are trained with a large dataset of 851 patients. The dataset contains data from T2Ds, prediabetics and normal individuals. The forecasting is done for a PH of 15 and 60 minutes. The results show almost perfect prediction, this is due to methodological errors. The authors claim to have split training, cross-validation and test data randomly. This could prove to be a wrong strategy in time-series forecasting as there is a chance of the model getting trained on the future data.

We agree that certain training strategies can cause the model to –during evaluation– recognize data on which it had been trained, which indeed would lead to near perfect prediction. However, as explained in the Methods section under dataset construction (Page 9, Lines 170-172), the datasets were split in such a way that any given participant was present in only the training, tuning, or evaluation set. As such, the models have been trained in other individuals than those who are present in the evaluation set. Therefore, it is not possible for the model to have been trained on ‘future data’.

2. The authors trained multiple models for prediction purpose. It is seen in the performance comparison table that classical RNN performs best for 15 min PH and LSTM performs best for 60 min PH. The manuscript, however, only contains details about the LSTM model.

We agree with the reviewer’s comment that the classical RNN just outperformed all other models at a prediction horizon of 15 minutes. Considering 1) the differences between the RNN and LSTM (RMSE: 0.485 [0.481-0.490] vs. 0.482 [0.477-0.487]) at 15 minutes were negligible and not statistically significant; and 2) RNN was substantially worse than LSTM (RMSE: 0.989 [0.983-0.995] vs. 0.941 [0.937-0.945]) at 60 minutes, we decided to use the LSTM architecture for predictions at both time horizons. The details about the RNN and LSTM used in the baseline comparison are described in S1 Supporting Information (Page 2).

3. Since, the proposed study does exactly the same what various other works have been doing for BG prediction during the past decade, no attempt at performance comparison with prior work has been made.

We believe that our work is unique in several respects. First, our study features one of the largest study populations to date. Second, the large-scale combination of CGM and accelerometry is unique. Third, we train glucose prediction models in individuals with NGM, prediabetes, or type 2 diabetes, research on which is notably scarce.

As our initial findings were obtained from a population ranging from normal glucose metabolism to type 2 diabetes, we compared our performance metrics in the type 2 diabetes subgroup with the largest study in type 2 diabetes to date (n=50)(4). This can be found in the Discussion (Page 17-18, Lines 355-367). Initially, we did not set our findings against studies in individuals with type 1 diabetes because comparison would not be valid (individuals with type 1 diabetes experience much greater daily glucose variability). As we now have extended our results to individuals with type 1 diabetes, we have also included a comparison to the most recent and best-performing studies in type 1 diabetes (Page 18, Lines 367-376).

4. Performance improvement depicted in the CGM+PA dataset is not significant and hence provides no motive for designers to prefer one over the other.

We agree with the reviewer that model performance improves only slightly when accelerometer data is incorporated, as we have delineated this in our Discussion accordingly (e.g., Page 17, Lines 350-351; Page 18, Lines 378-393). Nevertheless, our study is the first to use such a large population to make this important comparison.

5. Since it is understood that the glucose variability in NGM is low, and the number of individuals with NGM in both datasets are the largest, the underlying trends being identified by the ML model are overwhelmed by such data. It explains why the ML model are predicting almost perfectly.

We assent with the reviewer that prediction is (expected to be) most accurate in individuals with NGM. Hence, we stratified the performance metrics based on the participants’ glucose metabolism status in order to assess how accuracy would fare for participants with prediabetes or type 2 diabetes (Table 2, S5 Table). Furthermore, the clinical safety results are shown only for individuals with type 2 diabetes (Figure 2). Based on these data, we believe that the high accuracy and clinical safety in individuals with type 2 diabetes are not caused by the model being overwhelmed by data of participants with NGM. This is now further supported by the finding that our models translate quite well to individuals with type 1 diabetes.

Reviewer #3

1. This article presents the work on the application of different machine learning techniques for glucose prediction from CGM and physical activity bracelets. This is a study with a very large number of patients but in my opinion the article is not interesting for the journal for several reasons. Firstly, the data are not available to the research community, which makes it difficult to check whether the techniques presented can be overcome by the countless number of papers in the area.

Data of The Maastricht Study are certainly available to researchers who meet the criteria for access to confidential data. For the safety and privacy of the participants, as requested by law and ethical regulations, strict procedures to obtain data are in place. The implication of this is that the data have been deemed unsuitable for public deposition by The Board of The Maastricht Study, as described in detail under Data availability (Page 22, Lines 465-472). This should not be a ground to preclude publication of our work in PLOS ONE. Accordingly, multiple manuscripts that used data from The Maastricht Study and, thus, were under the same restrictions regarding data availability have been published in PLOS ONE(5-9).

2. Secondly, no new techniques are proposed, there are many studies in the area and the techniques of machine learning have been studied in depth, the authors can see for example all the articles in results were reported in different previous publications recently and some years ago. You can see, for instance the works presented at the last two workshops on Blood Glucose Level Prediction (BGLP) Challenge http://ceur-ws.org/Vol-2675/ You can also find several journal papers

Hidalgo, J. I., Colmenar, J. M., Kronberger, G., Winkler, S. M., Garnica, O., & Lanchares, J. (2017). Data based prediction of blood glucose concentrations using evolutionary methods. Journal of medical systems, 41(9), 142.

Woldaregay, A. Z., Årsand, E., Walderhaug, S., Albers, D., Mamykina, L., Botsis, T., & Hartvigsen, G. (2019). Data-driven modeling and prediction of blood glucose dynamics: Machine learning applications in type 1 diabetes. Artificial intelligence in medicine, 98, 109-134.

Velasco, J. M., Garnica, O., Lanchares, J., Botella, M., & Hidalgo, J. I. (2018). Combining data augmentation, EDAs and grammatical evolution for blood glucose forecasting. Memetic Computing, 10(3), 267-277.

De Falco, I., Della Cioppa, A., Giugliano, A., Marcelli, A., Koutny, T., Krcma, M., ... & Tarantino, E. (2019). A genetic programming-based regression for extrapolating a blood glucose-dynamics model from interstitial glucose measurements and their first derivatives. Applied Soft Computing, 77, 316-328.

Contreras, I., & Vehi, J. (2018). Artificial intelligence for diabetes management and decision support: literature review. Journal of medical Internet research, 20(5),

And even more from the last two years on NNs and DL approaches

We thank the reviewer for these literature suggestions. We would like to point out that such references were not included because of our focus on type 2 diabetes. Since we have now expanded our work to include translation of our prediction models to individuals with type 1 diabetes, we have updated our references accordingly. Our findings are comparable to most recent findings, as summarized in the table below (available in the reviewer document).

In addition, we would like to point out that the purpose of our study was not to develop a completely new glucose prediction technique. On the one hand, we aimed to investigate how well a glucose prediction model would fare in a large study population with actual CGM data (in contrast to in silico data). On the other hand, we aimed to assess whether incorporation of simultaneously assessed accelerometry data would improve glucose prediction. We certainly have the ambition to develop our models further (e.g., by using more complex or novel algorithms), but view such endeavors to be out of scope for the current manuscript.

3. Moreover a prediction horizon of 15 minutes is so short that any Naive approach could reach a 95% of safe predictions, I recommend the authors the exercise of predicting the glucose value for 15 minutes as the value a t=0.

As the reviewer suggested, we calculated the performance characteristics of a model predicting the glucose value at t0. The results are shown in S7 Table and S3 Figure. As the reviewer can appreciate, naive prediction performance is substantially worse than our ML-based prediction models, especially in the type 2 diabetes.

We further analyzed the effect of using t0 as prediction value between prediction horizons 0 and 120 minutes. S4 Figure illustrates the effect for each of the performance measures.

4. Experimental results are not useful as they are presented in the paper. All the techniques are summarized in just one table and no discrimination among them is done.

First, we want to highlight that we have compared a large number of different machine learning models (i.e., ARIMA, support vector regression, gradient-boosting trees, feed-forward neural networks, and recurrent neural networks [RNN]), which is outlined in the Methods (Page 9, Lines 183-187). The results of the comparisons are indeed shown in one supplementary table (S1 Table). However, we do not concur that this is a drawback of our study. The aims of our study were to assess to what extent glucose values can be accurately predicted at 15- and 60-minute intervals, and whether it could be further improved by incorporation of accelerometer-measured physical activity data. We intended to find the best machine learning-based prediction model as a mains to an end, not as a goal in itself. After we concluded that LSTM performed best, especially at a prediction interval of 60 minutes, there was no need to further address or compare all other possible techniques. Still, we did further optimize the architecture of the LSTM-based models; the hyperparameters that were studied and chosen are shown in S2 and S3 Table.

5. The main conclusion of the paper is so general that is obvious. Is something like the affirmation " Medicine is good" or something similar.

We have now rewritten the main conclusion in order to incorporate our validation in type 1 diabetes. We believe that this overcomes the point made by the reviewer.

6. Last but no least, the study affirm that, although it was made with T2 diabetes patients, it could be extrapolated to other T1 patients. I am sure that conclusions for T2 can not directly extrapolated to other type of patients. It has been shown in the past that in-silico results are not extensible to T1 real patients nor to T2 and vice versa. Glucose Variability of one T1 or T2 patients are different, T1 can produce little amounts of insulin or not, T2 insulin resistance could be heavier for one patient than for other....

We agree with the reviewer that our findings cannot automatically be translated to a T1D population. Based on the reviewer’s remark, we have therefore extended our results to a small study population of individuals with type 1 diabetes (OhioT1DM Dataset). As can be appreciated from Figure 3, the clinical safety of our prediction models is indeed quite high, even at the 60-minute interval (>91% predictions are highly clinically safe). Furthermore, the accuracy of our models is comparable to current studies in type 1 diabetes. We have rewritten our discussion to include our reflection on the findings for type 1 diabetes (Page 17, Lines 367-376).

7. So in this conditions the study is of little interest for the journal, In my humble opinion, I think that the data set has a great potential and that the research team is capable of prepare and in depth analysis of machine learning technique, I would recommend to separate and configure ML techniques for the different types of patients, and of course when presenting the results separate by ML techniques and data sets.

We would like to thank the reviewer for acknowledging the great potential of our dataset. Still, we want to reiterate what the main purposes of our study were. First, to investigate to what extent glucose values can be accurately predicted at 15- and 60-minute intervals, while using the best performing ML-based model in a large sample of participants with either NGM, prediabetes, or type 2 diabetes. Second, whether prediction could be further improved by incorporation of accelerometer-measured physical activity data. Finally, we believe that the (clinical) interest has been augmented by inclusion of a T1D dataset in the revised version of the manuscript, for which the reviewer is greatly acknowledged. 

Reviewer #4

1. In this paper, the authors used machine learning to train models in predicting future glucose levels based on prior CGM and accelerometry data. According to experiments the authors conducted, the results showed that machine learning-based models are able to accurately and safely predict glucose values both at 15- and 60-minute intervals with only CGM data; And incorporation of accelerometer data slightly improved prediction. It is interesting and of great value to utilize machine learning-based models to predict future glucose levels. At the same time, I have several major concerns about this study. First, the authors did not conduct a literature review of researches on future glucose level predictions. Based on this manuscript, we do not know which models are used to predict future glucose levels, and how the current progress is, especially the use of machine learning-based models that this paper employed. Further, it is also difficult to determine whether this paper has enough innovations and contributions, as the authors did not list in detail.

In the revised Discussion of our manuscript, we now extensively compare our results with literature results (Page 17-18, Lines 355-376). We also refer to a recent review paper on this topic in the Introduction section (Page 5, Lines 87-89)(13). Of note, the true purposes of our study were to assess to what extent glucose values can be accurately predicted at 15- and 60-minute intervals in individuals with NGM, prediabetes, or type 2 diabetes, and whether the prediction could be further improved by incorporation of accelerometer-measured physical activity data. Unique in this regard are our large study population of individuals with NGM, prediabetes, or type 2 diabetes (n=851) and the large-scale combination of simultaneously performed CGM and accelerometry (n=540).

2. Second, in the “Model development and design” part of the article, the authors mentioned that the prediction task of future glucose levels can be solved by a variety of statistical and machine learning models, and the authors assessed various models, such as autoregressive integrated moving average, support vector regression, etc. Finally, a LSTM architecture was chosen as this had the best performance in the tuning dataset. There are two questions of the model selection:

1) Why the authors selected these statistical and machine learning models? What are the applicability and advantages of these models?

The statistical and machine learning models used were selected on the basis of previously published literature in relation to time-series forecasting and glucose prediction. Details and applicability for each of these models were briefly discussed in S1 Supporting Information (Page 2, Lines 24-56).

2) The authors utilized the multi-task LSTM network to predict future glucose levels finally as it had the best performance in the tuning dataset. If the dataset is modified or other features like diet, medication use are added in the models, whether the multi-task LSTM network can also achieve the best performance is unknown. Therefore, except for the reason of best performance, the authors should detail other reasons for choosing the multi-task LSTM network.

Our major rationale to select the LSTM architecture is the baseline performance described in S1 Supporting Information (Page 2, Lines 24-56) and S1 Table. Additional advantages include the option to incorporate explainability into our LSTM models (14, 15), and the relatively low computing cost compared to complex, deep neural networks which could potentially hinder application of these networks in relatively simple devices such as closed-loop insulin delivery systems. Nonetheless, we agree with the author that in case additional features, such as diet or medication use, would be added to the dataset, we would have to reevaluate our current architecture. We have included this in our Limitations section (Page 20, Lines 430-431).

3. Third, the authors verified the performance of the multi-task LSTM network with 851 individuals. As is known to all, deep learning models are usually validated on a large number of datasets. The amount of data in this paper may not be convincing enough for model validation.

We agree with the reviewer that deep learning models should be validated on a large number of samples in order to confidently assess its generalizability. In our specific study, it is important to realize that, besides the number of individuals, the number of glucose measurements are also critical in the training of these models. We had almost 1.4 million glucose measurements available in the our study. Currently, this is the largest study that described the application of ML-based glucose prediction in individuals with type 2 diabetes. Nevertheless, we acknowledge that future studies should assess to what extent our models generalize to other populations.

References

1. Kumareswaran K, Thabit H, Leelarathna L, Caldwell K, Elleri D, Allen JM, et al. Feasibility of closed-loop insulin delivery in type 2 diabetes: a randomized controlled study. Diabetes Care. 2014;37(5):1198-203.

2. Li K, Liu C, Zhu T, Herrero P, Georgiou P. GluNet: A Deep Learning Framework for Accurate Glucose Forecasting. IEEE J Biomed Health Inform. 2020;24(2):414-23.

3. Perez-Gandia C, Facchinetti A, Sparacino G, Cobelli C, Gomez EJ, Rigla M, et al. Artificial neural network algorithm for online glucose prediction from continuous glucose monitoring. Diabetes Technol Ther. 2010;12(1):81-8.

4. Mohebbi A, Johansen AR, Hansen N, Christensen PE, Tarp JM, Jensen ML, et al. Short Term Blood Glucose Prediction based on Continuous Glucose Monitoring Data. arXiv e-prints [Internet]. 2020 February 01, 2020:[arXiv:2002.02805 p.]. Available from: https://ui.adsabs.harvard.edu/abs/2020arXiv200202805M.

5. de Rooij BH, van der Berg JD, van der Kallen CJ, Schram MT, Savelberg HH, Schaper NC, et al. Physical Activity and Sedentary Behavior in Metabolically Healthy versus Unhealthy Obese and Non-Obese Individuals - The Maastricht Study. PLoS One. 2016;11(5):e0154358.

6. Sorensen BM, Houben A, Berendschot T, Schouten J, Kroon AA, van der Kallen CJH, et al. Cardiovascular risk factors as determinants of retinal and skin microvascular function: The Maastricht Study. PLoS One. 2017;12(10):e0187324.

7. Elissen AMJ, Hertroijs DFL, Schaper NC, Bosma H, Dagnelie PC, Henry RM, et al. Differences in biopsychosocial profiles of diabetes patients by level of glycaemic control and health-related quality of life: The Maastricht Study. PLoS One. 2017;12(7):e0182053.

8. Martens RJH, van der Berg JD, Stehouwer CDA, Henry RMA, Bosma H, Dagnelie PC, et al. Amount and pattern of physical activity and sedentary behavior are associated with kidney function and kidney damage: The Maastricht Study. PLoS One. 2018;13(4):e0195306.

9. Consolazio D, Koster A, Sarti S, Schram MT, Stehouwer CDA, Timmermans EJ, et al. Neighbourhood property value and type 2 diabetes mellitus in the Maastricht study: A multilevel study. PLoS One. 2020;15(6):e0234324.

10. Kriventsov S, Lindsey A, Hayeri A. The Diabits App for Smartphone-Assisted Predictive Monitoring of Glycemia in Patients With Diabetes: Retrospective Observational Study. JMIR Diabetes. 2020;5(3):e18660.

11. Martinsson J, Schliep A, Eliasson B, Mogren O. Blood Glucose Prediction with Variance Estimation Using Recurrent Neural Networks. Journal of Healthcare Informatics Research. 2020;4(1):1-18.

12. Chen J, Li K, Herrero P, Zhu T, Georgiou P, editors. Dilated Recurrent Neural Network for Short-time Prediction of Glucose Concentration. KHD@IJCAI; 2018.

13. Woldaregay AZ, Arsand E, Walderhaug S, Albers D, Mamykina L, Botsis T, et al. Data-driven modeling and prediction of blood glucose dynamics: Machine learning applications in type 1 diabetes. Artif Intell Med. 2019;98:109-34.

14. Thorsen-Meyer H-C, Nielsen AB, Nielsen AP, Kaas-Hansen BS, Toft P, Schierbeck J, et al. Dynamic and explainable machine learning prediction of mortality in patients in the intensive care unit: a retrospective study of high-frequency data in electronic patient records. The Lancet Digital Health. 2020.

15. Lauritsen SM, Kristensen M, Olsen MV, Larsen MS, Lauritsen KM, Jorgensen MJ, et al. Explainable artificial intelligence model to predict acute critical illness from electronic health records. Nat Commun. 2020;11(1):3852.

Attachment

Submitted filename: Reponse to Reviewers.docx

Decision Letter 1

Chi-Hua Chen

19 Mar 2021

PONE-D-20-30681R1

Machine learning-based glucose prediction with use of continuous glucose and physical activity monitoring data: The Maastricht Study

PLOS ONE

Dear Dr. Brouwers,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please submit your revised manuscript by May 03 2021 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: http://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols

We look forward to receiving your revised manuscript.

Kind regards,

Chi-Hua Chen, Ph.D.

Academic Editor

PLOS ONE

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #1: All comments have been addressed

Reviewer #3: (No Response)

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #3: No

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #3: No

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #3: No

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #3: Yes

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: The paper has been improved significantly. All my comments have been addressed with clear explanation and all updates have been shown explicitly in the paper.

Reviewer #3: I think that my concerns have not been addressed.

The paper has little interest for the reader of the journal. In the case of people not working in the field, the contribution is so poor that no extrapolation to other works can be done. On the other hand for people working on this problem, conclusion are known, statistical validation is not made and conclusions are not fully supported by experiments.

The inclusion of T1D patients is forced in my opinion and does not make much sense with the other results.

My questions are again the same

2. Secondly, no new techniques are proposed,....

Analyses are mere description of the results, see for instance lines 323 to 329:

323 Additional analyses

324 To further obtain insights into our model predictions, we assessed performance metrics

325 stratified by day and night (S8 Table). Fifteen-minute predictions did not materially differ

326 between day and night. By contrast, accuracy of 60-minute predictions was lower during the

327 day than at night. In addition, we stratified the results by high or low glucose variability (i.e.,

328 SD cut-off of 1.37 mmol/L) (S9 Table). Model performance was slightly lower at higher

329 glucose variability, at both time intervals of 15 and 60 minutes.

4. Experimental results are not useful as they are presented in the paper. All the techniques are summarized in just one table and no discrimination among them is done.

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

Reviewer #3: No

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

PLoS One. 2021 Jun 24;16(6):e0253125. doi: 10.1371/journal.pone.0253125.r004

Author response to Decision Letter 1


9 Apr 2021

We would like to thank the reviewers for their feedback on our study and for the time spent on our manuscript. Please find below our point-by-point rebuttal.

Reviewer #1

The paper has been improved significantly. All my comments have been addressed with clear explanation and all updates have been shown explicitly in the paper.

We once more would like to thank the reviewer for his/her positive feedback on our study and for the time spent on our manuscript.

Reviewer #3

I think that my concerns have not been addressed. The paper has little interest for the reader of the journal. In the case of people not working in the field, the contribution is so poor that no extrapolation to other works can be done. On the other hand for people working on this problem, conclusion are known, statistical validation is not made and conclusions are not fully supported by experiments.

We respectfully disagree with the comments made by the reviewer. We would like to reiterate the main aims of our study. First, we investigated to what extent glucose values can be accurately predicted at 15- and 60-minute intervals, while using the best performing ML-based model in a large sample of participants with either normal glucose metabolism (NGM), prediabetes, or type 2 diabetes. Such large scale data has not yet been used in the context of a glucose prediction study. Second, we assessed whether prediction could be further improved by incorporation of accelerometer-measured physical activity data. Such a large-scale combination of simultaneously collected continuous glucose monitoring and activity tracker has not yet been used. Finally, we augmented the (clinical) interest of our study by including model validation in a type 1 diabetes dataset in the revised version of the manuscript, for which the reviewer is greatly acknowledged. In conclusion, the combination of our large cohort of individuals with NGM, prediabetes or type 2 diabetes as well as the examination of glucose measurements combined with physical activity data has never been described to date and can, therefore, be of great interest to the readers of PLOS ONE.

The inclusion of T1D patients is forced in my opinion and does not make much sense with the other results.

We chose to include individuals with type 1 diabetes in order to examine the performance of our models in this subgroup as part of a proof-of-concept analysis. As described in our Discussion, we agree that further, comprehensive evaluation is necessary in order to examine the real clinical benefit and performance of our models in individuals with type 1 diabetes. Notably, previous studies have also employed this dataset in order to validate their models (1-7).

My questions are again the same 2. Secondly, no new techniques are proposed,....

Analyses are mere description of the results, see for instance lines 323 to 329:

Additional analyses

To further obtain insights into our model predictions, we assessed performance metrics stratified by day and night (S8 Table). Fifteen-minute predictions did not materially differ between day and night. By contrast, accuracy of 60-minute predictions was lower during the day than at night. In addition, we stratified the results by high or low glucose variability (i.e., SD cut-off of 1.37 mmol/L) (S9 Table). Model performance was slightly lower at higher glucose variability, at both time intervals of 15 and 60 minutes.

We again would like to stress that the main objective of the current work is not the technological advancement of algorithms per se. We certainly have the ambition to develop our models further (e.g., by using more complex or innovative algorithms), but view such endeavors to be out of scope for the current manuscript.

4. Experimental results are not useful as they are presented in the paper. All the techniques are summarized in just one table and no discrimination among them is done.

We assume the reviewer is referencing to the baseline comparison of various statistical and machine learning models as depicted in S1 Table. In line with our previous comment, we would like to point out that we did not aim to achieve technological advancements. Hence, we intended to find the best machine learning-based prediction model as a means to an end, not as a goal in itself. After we concluded that LSTM performed best, especially at a prediction interval of 60 minutes, there was no need to further address or compare all other possible techniques. As such, the comprehensive baseline evaluation of different statistical and machine learning models can even be considered a strength of our study.

References

1. Marling C, Bunescu RC, editors. The OhioT1DM Dataset For Blood Glucose Level Prediction. KHD@IJCAI; 2018.

2. Marling C, Bunescu RC, editors. The OhioT1DM Dataset For Blood Glucose Level Prediction: Update 2020. KHD@IJCAI; 2020.

3. Pappada SM, Owais MH, Cameron BD, Jaume JC, Mavarez-Martinez A, Tripathi RS, et al. An Artificial Neural Network-based Predictive Model to Support Optimization of Inpatient Glycemic Control. Diabetes Technol Ther. 2020;22(5):383-94.

4. Li K, Liu C, Zhu T, Herrero P, Georgiou P. GluNet: A Deep Learning Framework for Accurate Glucose Forecasting. IEEE J Biomed Health Inform. 2020;24(2):414-23.

5. Chen J, Li K, Herrero P, Zhu T, Georgiou P, editors. Dilated Recurrent Neural Network for Short-time Prediction of Glucose Concentration. KHD@IJCAI; 2018.

6. Martinsson J, Schliep A, Eliasson B, Meijner C, Persson S, Mogren O. Automatic blood glucose prediction with confidence using recurrent neural networks. 3rd International Workshop on Knowledge Discovery in Healthcare Data, KDH@IJCAI-ECAI 2018, 13 July 2018; 20182018. p. 64-8.

7. Kriventsov S, Lindsey A, Hayeri A. The Diabits App for Smartphone-Assisted Predictive Monitoring of Glycemia in Patients With Diabetes: Retrospective Observational Study. JMIR Diabetes. 2020;5(3):e18660.

Attachment

Submitted filename: Response to Reviewers.docx

Decision Letter 2

Chi-Hua Chen

23 Apr 2021

PONE-D-20-30681R2

Machine learning-based glucose prediction with use of continuous glucose and physical activity monitoring data: The Maastricht Study

PLOS ONE

Dear Dr. Brouwers,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please submit your revised manuscript by Jun 07 2021 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: http://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Chi-Hua Chen, Ph.D.

Academic Editor

PLOS ONE

Journal Requirements:

Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #1: All comments have been addressed

Reviewer #3: (No Response)

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #3: Yes

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #3: Yes

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: No

Reviewer #3: No

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #3: Yes

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: My all comments have been addressed properly. The only issue is that, Plus One requires all data underlying the findings in the manuscript fully available. It is better to address this issue before publish.

Reviewer #3: I really appreciate the efforts of the authors for improving the paper and answering my questions. I would like to explain better which is my opinion about the great potential of the work.

What I would expect of such amount of data is to obtain guidelines for selecting and designing better ML (or not ML) algorithms, based on the precision needed, the time for response, the data availability and of course the features of the patient.

For me, what it would be useful for the journal readers is a combination of tables 1 and 2 with information provided as supporting information.

As the supporting information is going to be publish, a summary of this information should be included in the paper, in order to highlight the insights of this study. I would like to see a Table with a summary of the supporting information in the main paper

The paper is is a very interesting work

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

Reviewer #3: No

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

PLoS One. 2021 Jun 24;16(6):e0253125. doi: 10.1371/journal.pone.0253125.r006

Author response to Decision Letter 2


20 May 2021

We would like to thank the reviewers for their feedback on our study and for the time spent on our manuscript. Please find below our point-by-point rebuttal. The page and line numbers refer to the manuscript version with track changes.

Reviewer #1

My all comments have been addressed properly. The only issue is that, Plus One requires all data underlying the findings in the manuscript fully available. It is better to address this issue before publish.

Data of The Maastricht Study are certainly available to researchers who meet the criteria for access to confidential data. For the safety and privacy of the participants, as requested by law and ethical regulations, strict procedures to obtain data are in place. The implication of this is that the data have been deemed unsuitable for public deposition by The Board of The Maastricht Study, as described in detail under Data availability (Page 24, Lines 493-498). This should not be a ground to preclude publication of our work in PLOS ONE. Accordingly, multiple manuscripts that used data from The Maastricht Study and, thus, were under the same restrictions regarding data availability have been published in PLOS ONE(1-5).

Reviewer #3

I really appreciate the efforts of the authors for improving the paper and answering my questions. I would like to explain better which is my opinion about the great potential of the work. What I would expect of such amount of data is to obtain guidelines for selecting and designing better ML (or not ML) algorithms, based on the precision needed, the time for response, the data availability and of course the features of the patient. For me, what it would be useful for the journal readers is a combination of tables 1 and 2 with information provided as supporting information. As the supporting information is going to be publish, a summary of this information should be included in the paper, in order to highlight the insights of this study. I would like to see a Table with a summary of the supporting information in the main paper. The paper is is a very interesting work

As suggested by the reviewer, we have now included the baseline comparison of statistical and machine learning models for glucose prediction in our main manuscript as Table 1 (Page 12, Lines 223-225). Accordingly, we made various changes in the methods section of our main manuscript (Page 10, Lines 195-205; Page 22, Lines 439-441) and supplementary information (Page 4, Lines 92-98).

References

1. de Rooij BH, van der Berg JD, van der Kallen CJ, Schram MT, Savelberg HH, Schaper NC, et al. Physical Activity and Sedentary Behavior in Metabolically Healthy versus Unhealthy Obese and Non-Obese Individuals - The Maastricht Study. PLoS One. 2016;11(5):e0154358.

2. Sorensen BM, Houben A, Berendschot T, Schouten J, Kroon AA, van der Kallen CJH, et al. Cardiovascular risk factors as determinants of retinal and skin microvascular function: The Maastricht Study. PLoS One. 2017;12(10):e0187324.

3. Elissen AMJ, Hertroijs DFL, Schaper NC, Bosma H, Dagnelie PC, Henry RM, et al. Differences in biopsychosocial profiles of diabetes patients by level of glycaemic control and health-related quality of life: The Maastricht Study. PLoS One. 2017;12(7):e0182053.

4. Martens RJH, van der Berg JD, Stehouwer CDA, Henry RMA, Bosma H, Dagnelie PC, et al. Amount and pattern of physical activity and sedentary behavior are associated with kidney function and kidney damage: The Maastricht Study. PLoS One. 2018;13(4):e0195306.

5. Consolazio D, Koster A, Sarti S, Schram MT, Stehouwer CDA, Timmermans EJ, et al. Neighbourhood property value and type 2 diabetes mellitus in the Maastricht study: A multilevel study. PLoS One. 2020;15(6):e0234324.

Attachment

Submitted filename: Response to Reviewers.docx

Decision Letter 3

Chi-Hua Chen

1 Jun 2021

Machine learning-based glucose prediction with use of continuous glucose and physical activity monitoring data: The Maastricht Study

PONE-D-20-30681R3

Dear Dr. Brouwers,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice for payment will follow shortly after the formal acceptance. To ensure an efficient process, please log into Editorial Manager at http://www.editorialmanager.com/pone/, click the 'Update My Information' link at the top of the page, and double check that your user information is up-to-date. If you have any billing related questions, please contact our Author Billing department directly at authorbilling@plos.org.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Chi-Hua Chen, Ph.D.

Academic Editor

PLOS ONE

Additional Editor Comments (optional):

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #1: All comments have been addressed

Reviewer #3: All comments have been addressed

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #3: Yes

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #3: Yes

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #3: No

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #3: Yes

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: All comments have been nicely addressed. Data of The Maastricht Study are certainly available to researchers who meet the criteria for access to confidential data

Reviewer #3: All my comments have been addressed. I really appreciate the efforts made to include the tables I suggested.

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

Reviewer #3: No

Acceptance letter

Chi-Hua Chen

15 Jun 2021

PONE-D-20-30681R3

Machine learning-based glucose prediction with use of continuous glucose and physical activity monitoring data: The Maastricht Study

Dear Dr. Brouwers:

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now with our production department.

If your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information please contact onepress@plos.org.

If we can help with anything else, please email us at plosone@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Professor Chi-Hua Chen

Academic Editor

PLOS ONE

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 Fig. Illustrative examples of continuous glucose monitoring-based machine learning model predictions compared to actual values.

    (DOCX)

    S2 Fig. Illustrative examples of continuous glucose monitoring- and accelerometry-based machine learning model predictions compared to actual values.

    (DOCX)

    S3 Fig. Surveillance error grid evaluation of glucose prediction safety at time intervals of 15 and 60 minutes using glucose value t0 as predictor.

    (DOCX)

    S4 Fig. Performance characteristics of a prediction model using t0 as predictor across time horizons between 0 and 120 minutes.

    (DOCX)

    S5 Fig. Parkes error grid evaluation of glucose prediction safety at time intervals of 15 and 60 minutes.

    (DOCX)

    S1 Table. Hyperparameter combinations evaluated in current experiments.

    (DOCX)

    S2 Table. Final set of hyperparameters for each of the machine learning models.

    (DOCX)

    S3 Table. Extended baseline characteristics.

    (DOCX)

    S4 Table. Extended analysis of model performance in normal glucose metabolism and prediabetes subgroups.

    (DOCX)

    S5 Table. Extended analysis on time lag between predicted and actual glucose values.

    (DOCX)

    S6 Table. Extended analysis of model performance with t0 glucose value as predictor.

    (DOCX)

    S7 Table. Model performance stratified by day and night.

    (DOCX)

    S8 Table. Model performance stratified by low versus high glucose variability.

    (DOCX)

    S9 Table. Extended analysis of model performance in the Ohio T1DM Dataset.

    (DOCX)

    S1 File. Background information on machine learning models reviewed in current study.

    (DOCX)

    S2 File. Background information on metrics used in the current study.

    (DOCX)

    Attachment

    Submitted filename: Reponse to Reviewers.docx

    Attachment

    Submitted filename: Response to Reviewers.docx

    Attachment

    Submitted filename: Response to Reviewers.docx

    Data Availability Statement

    Data are unsuitable for public deposition due to ethical restriction and privacy of participant data. The study has been approved by the medical ethical committee of the Maastricht University Medical Center (NL31329.068.10/ MEC 10-2-023) and the Netherlands Health Council under the Dutch “Law for Population Studies” (Permit 131088-105234-PG). Data are available from The Maastricht Study for any interested researchers who meet the criteria for access to confidential data. The Maastricht Study Management Team (research.dms@mumc.nl) and the corresponding author (Martijn C.G.J. Brouwers) may be contacted to request data.


    Articles from PLoS ONE are provided here courtesy of PLOS

    RESOURCES