Skip to main content
Advances in Wound Care logoLink to Advances in Wound Care
. 2024 Jun 13;13(6):281–290. doi: 10.1089/wound.2023.0194

An Improved Clinical and Genetics-Based Prediction Model for Diabetic Foot Ulcer Healing

Gary Hettinger 1, Nandita Mitra 1, Stephen R Thom 2, David J Margolis 1,3,*
PMCID: PMC11339549  PMID: 38258807

Abstract

Objective:

The goal of this investigation was to use comprehensive prediction modeling tools and available genetic information to try to improve upon the performance of simple clinical models in predicting whether a diabetic foot ulcer (DFU) will heal.

Approach:

We utilized a cohort study (n = 206) that included clinical factors, measurements of circulating endothelial precursor cells (CEPCs), and fine sequencing of the NOS1AP gene. We derived and selected relevant predictive features from this patient-level information using statistical and machine learning techniques. We then developed prognostic models using machine learning approaches and assessed predictive performance. The presentation is consistent with TRIPOD requirements.

Results:

Models using baseline clinical and CEPC data had an area under the receiver operating characteristic curve (AUC) of 0.73 (0.66–0.80). Models using only single nucleotide polymorphisms (SNPs) of the NOS1AP gene had an AUC of 0.67 (95% confidence interval, CI: [0.59–0.75]). However, models incorporating baseline and SNP information resulted in improved AUC (0.80, 95% CI [0.73–0.87]).

Innovation:

We provide a rigorous analysis demonstrating the predictive potential of genetic information in DFU healing. In this process, we present a framework for using advanced statistical and bioinformatics techniques for creating superior prognostic models and identify potentially predictive SNPs for future research.

Conclusion:

We have developed a new benchmark for which future predictive models can be compared against. Such models will enable wound care experts to more accurately predict whether a patient will heal and aid clinical trialists in designing studies to evaluate therapies for subjects likely or unlikely to heal.

Keywords: prognostic models, diabetic foot ulcers, machine learning, epidemiology


graphic file with name wound.2023.0194_figure2.jpg

David J. Margolis, MD, PhD

INTRODUCTION

A recent study of chronic wound prevalence among Medicare beneficiaries noted a 13% increase in chronic wounds in the United States between the years of 2014 and 2019.1 The investigators estimated that between 8.2 million and 10.5 million U.S. Medicare beneficiaries have a chronic wound (between 400 and 500,000 with a diabetic foot ulcer [DFU]), and the prevalence of chronic wounds had increased from 14.5% to 16.4% in that time frame.1 The total cost to Medicare for the treatment of individuals whose primary diagnosis was a chronic wound ranged from 24.7 to 33.6 billion dollars per year from 2014 to 2019.1

Diabetes mellitus (DM) is a worldwide problem. Much of the morbidity and mortality associated with DM is due to complications of DM like DFUs and lower extremity amputations (LEAs).2 LEAs mostly occur in individuals with a DFU.3,4 The annual incidence of DFU in Medicare beneficiaries with diabetes is about 6 per hundred and for LEA about 5 per thousand, but rates of DFU and LEA can vary up to two- to threefold by race/ethnicity, geographic location, and gender.4–6 LEA is an important burdensome complication of DM that carries an annual risk of mortality of about 15–20%, and medical costs for those with an LEA range between $30,000 and $60,000 per annum.4,6–9

In a study that analyzed 100% of the Medicare population, ∼11% of patients with a DFU and 22% of those with an LEA died within the same calendar year as their first diagnosis/procedure.4 At any time, those with a DFU have 2.48 times the odds of dying compared to those without DM after adjusting for common causes of death among those with diabetes like cardiovascular disease, renal disease, peripheral vascular disease, and stroke. A similar study showed that those with diabetes and an LEA were thrice more likely to die than those with diabetes and no history of LEA.8,9 The 5-year mortality associated with LEA is greater than the risk of death from breast cancer.10 LEAs are often classified as minor (e.g., toes, foot, and so on ) or major (i.e., transtibial or higher). The 5-year mortality after a major LEA is greater than the average risk from all cancers and only about 25% less than from lung cancer, one of the deadliest cancers.10

CLINICAL PROBLEM ADDRESSED

The goal of every clinician when treating a patient with a DFU is to get the wound to heal. About 90% of individuals with LEA have histories of DFUs that were not healing.11,12 Of those with a DFU, more than half will have a second or recurrent DFU in the next 12 months.13

Several tools have been created that can be helpful to a clinician when trying to predict which wounds will heal. These tools mostly incorporate clinical factors associated with the DFU that are simple to measure at the bedside.14–21 The use of these tools in a clinical setting has many purposes, including providing better information to the patient about their prognosis and helping clinicians to improve treatment management decisions through risk stratification. However, most of these tools are old, and newer tools tend to rely solely on clinical observations utilized in the older tools.16,17,20–22 At best, previously described prognostic tools successfully predict which wounds will heal about 60–70% of the time.16,17,20–22

Several previous investigations have shown that genetic variation of the NOS1AP gene and factors associated with the nitric oxide synthase-one accessory protein (NOS1AP) might be associated with DFUs that are more likely to heal.23–26 Why these associations exist is not fully known but could be related to interactions between capon, a protein coded for by the NOS1AP gene, and endothelial precursor cells which have been associated with wound repair.23–26

Overall, the goal of DFU care is to both heal the wound and prevent the medical sequelae associated with an unhealed DFU and/or LEA. To that end, the goal of this investigation was to demonstrate the use of patient-level genetic information and advanced prognostic modeling methods in predicting if a DFU will heal. For this task, we improved upon simple models using well-known predictive clinical information with complex single nucleotide polymorphism (SNP) data from the NOS1AP gene, measurements of circulating endothelial precursor cells (CEPCs), and state-of-the-art machine learning models.14–17,26

MATERIALS AND METHODS

Cohort

We investigated a cohort of patients from a multicenter study, called the Diabetic Foot Ulcer Consortium (DFUC), composed of wound care centers at academic institutions: University of Miami, Icahn School of Medicine at Mount Sinai, and University of Pennsylvania, as well as the community based-MVS Wound Care in Maryland. The DFUC was designed to evaluate circulating cellular markers, CEPCs, and microparticles (MP) as prognostic factors associated with the healing of DFU.14 All subjects were examined by a DFUC collaborator. All subjects were at least 40 years of age at the time of the original DFU diagnosis. Enrollment into DFUC and DFUC subject assessments occurred between October 2018 and January 2021.

Enrollment was based on an investigator's determination that the study subject had a history of adult-onset DM, had adequate arterial flow for healing (e.g., ankle brachial index of >0.80 and <1.20, palpable foot pulses, and so on), had a sensor neuropathy (e.g., sensory defect noted by Semmes Weinstein monofilaments), and had a wound on the plantar aspect of the foot that was eligible for DFU standard care.26 Standard care routinely included a baseline history and physical examination, including evaluation of lower extremity arterial flow, ascertainment of wound duration, wound area and depth, and an assessment of sensory neuropathy, sharp debridement, off-loading, treatment of infection (if present), a primary bandage, and then recurring periodic evaluation.26 All clinical information was well as materials for biochemical and genetic testing occurred at the baseline visit. Only outcome status (e.g., whether a wound healed) was determined at subsequent visits. As part of standard care, based on progress over the first few weeks, changes in the treatment plan did occur based on local care decisions.27 A previous study using this cohort confirmed that in this cohort the simple wound assessments of wound duration and wound size were highly predictive of a healed DFU.14

Analysis plan

We evaluated data from 206 patients in DFUC. In addition to NOS1AP variation, we assessed previously evaluated measures such as simple wound measures and CEPCs known to be prognostic factors associated with the healing of DFUs.14,26 We explored the predictive potential of the NOS1AP in two main steps. First, after next generation sequencing of the NOS1AP gene, we identified that NOS1AP SNPs highly correlated with healing and evaluated the performance of predictive models incorporating these SNPs to quantify the predictive potential of NOS1AP in our cohort.

We then compared the most correlated SNPs in the entire cohort to the most correlated within each fold of our model-training process to both understand potential generalizability of results from our first step and identify the most relevant SNPs in our data. Finally, we explored the sensitivity of our main results to missing data using multiply imputed datasets and the use of unsupervised feature engineering techniques to extract relevant information from the SNPs.

SNP variables

We first categorized the 98 measured SNPs from blood obtained at the baseline visit by the number of alleles at the SNP locus (0, 1, or 2). Only SNP categories with prevalence between 10% and 90% in our cohort were included. This resulted in 94 primary SNP variables. Finally, we added interaction terms between each SNP and six subpopulation indicators (Black, White, Hispanic, not Hispanic, Male, Female) since genetic effects have been known to vary by race, ethnicity, and sex. In a secondary analysis, we repeated this process for SNP-SNP interactions.

To explore whether automated feature engineering tools would identify relevant SNP information, we additionally developed unsupervised features from the 94 primary SNP variables and the 6 subpopulation indicators using 3 approaches: principal component analysis (PCA), sparse PCA (SPCA), and autoencoders.28–31 In addition to PCA, we considered features derived from SPCA to allow for principal components derived from fewer original covariates and autoencoders to allow for nonlinear relationships among variables. We included enough principal components to explain at least 60% of the variance and included any additional principal components that explained more than 4% of the variance. The dimension of the autoencoder hidden layer was chosen to match the number of selected principal components from SPCA.

Feature selection techniques

To create reduced feature sets (FSs) for our models, we used the following approach. First, for each candidate variable, we developed logistic regression models for the probability of healing using the variable as a predictor and adjusting for an indicator of the relevant subpopulation, if the variable included an interaction with a subpopulation. For each genetic component (i.e., SNP or SNP-SNP category), we then selected among the maximum seven variables (one main term and six subpopulation interactions) the term with the lowest p-value in the respective model.

We then created two FSs from this initial set. First, we selected variables with a p-value below 0.05 (Feature Set A [FS-A]). We then created general linear model (GLM) for these variables as before but now additionally adjusting for three baseline features previously found to be predictive of healing—the natural log of wound duration, the natural log of wound area, and the natural log of CEPC-CD34+45dim.26 Our second set then included features with p-values in the second model below 0.1 (FS-B). For our secondary analysis, we repeated this process for the larger variable set, including both SNP and SNP-SNP interactions using p-value thresholds of 0.01 and 0.05 for the two steps (FS-C and FS-D). For sets with more than 10 features, we also created sets ranking the lowest 10 p-values (FS-A-10, FS-C-10, FS-D-10).

Predictive models

We developed prediction models using SuperLearner (SL).32,33 SL works by including a set of candidate models in its library, fitting and generating prediction models for each of these candidates, and then estimating the optimal combination of these models by selecting a set of weights (ensemble SL) or the best individual model (discrete SL). To achieve a diverse set of candidate learners in our model, we included a GLM, two elastic net regression models (one with full weight to the L1 penalty, i.e., least absolute shrinkage and selection operator [LASSO], and one with L1 and L2 penalties weighted evenly), a generalized additive model (GAM) for potential nonlinear relationships between continuous covariates and the outcome, and a Bayesian Additive Regression Tree (BART).34–36

For training and evaluation, the dataset was first split into 10 randomized folds by stratifying on the binary healing outcome. Each fold was then split into 20 randomized folds again stratifying on the outcome. For each of the 10 outer folds, SL parameters were estimated by optimizing for the non-negative binomial likelihood using cross-validation on the 20 inner folds. Elastic net and BART learners were given the full FS for each model. GLM and GAM were given the features selected by LASSO within a given fold when the full FS contained more than 10 variables.

Even though model training was conducted within each fold with testing outcomes hidden from the training data, relevant SNP and SNP-SNP variables were selected using the full data, which may result in overconfidence that our observed SNPs are generalizable to the larger population. Therefore, in secondary analyses, we applied our feature selection techniques within each fold, instead of on the entire dataset, to assess the stability of our selection approach. In this study, only the top 10 variables were selected in each fold by p-value. Primary and secondary models were trained and tested on the dataset with complete clinical and SNP information. As a sensitivity analysis, we also developed 10 imputed datasets for patients with incomplete clinical and demographic information using chained equations and replicated our primary analyses across these imputations. Statistical analyses were performed in R version 4.0.2 (R Group for Statistical Computing).

Evaluation

Prediction performance was evaluated according to area under the receiver operating characteristic curve (AUC), Brier Score, and calibration intercept and slope.37–39 Ninety-five percent confidence intervals (CIs) are calculated using an influence curve-based approach for AUC and normal approximations assuming independent folds for Brier Scores.40 We assessed the stability of SNP feature selection by calculating the percentage of cross-validation folds where a given feature was one of the 10 selected.

Electronic laboratory notebook was not used for this investigation.

RESULTS

Patient population

Our patient population included 206 patients with healing information, 134 of which healed. Of these, 204 contained NOS1AP SNP information (used for SNP variable selection and multiple imputation analyses) and 179 contained complete information regarding baseline predictors and subpopulation indicators (used for primary model development and evaluation). Summaries of these baseline predictors, subpopulation indicators, and additional baseline information are shown in Table 1.

Table 1.

Characteristics at the baseline study visit

Covariate N Unhealed (n = 132) Healed (n = 72)
Age at enrollment, mean (SD) 202 57.51 (9.56) 58.32 (9.92)
Age of diabetes onset, mean (SD) 180 39.55 (13.16) 39.28 (12.57)
Black (%) 204 76 (57.6) 43 (59.7)
Male (%) 204 98 (74.2) 51 (70.8)
Ethnicity (not Hispanic) (%) 203 112 (84.8) 68 (94.4)
Chronic kidney diseasea (%) 166 34 (30.1) 15 (28.3)
Peripheral arterial diseasea (%) 198 31 (24.0) 13 (18.8)
Previous amputation (%) 202 62 (47.7) 35 (48.6)
Neuropathya (%) 196 109 (86.5) 65 (92.9)
No. of wounds, mean (SD) 202 1.39 (0.72) 1.26 (0.67)
HbA1c, mean (SD) 157 7.92 (2.72) 8.05 (2.16)
ln CEPC-CD34+45dim, mean (SD) 190 3.11 (1.04) 3.29 (0.94)
ln Wound duration months, mean (SD) 193 3.08 (1.09) 2.59 (1.23)
ln Wound area, mean (SD) 200 1.39 (1.85) 0.05 (1.89)
Wound depth: dermis only (%) 204 40 (30.3) 16 (22.2)
a

History of specific illness as reported by the subject.

CEPC, circulating endothelial precursor cell; SD, standard deviation.

Models with preselected SNPs

Using our selection approaches, 13 NOS1AP SNP variables were selected for FS-A, 5 for FS-B, 31 for FS-C, and 15 for FS-D. Those from FS-A-10, FS-B, FS-C-10, and FS-D-10 are summarized in Table 2. Model performance is summarized in Table 3. The previous standard baseline model with only the three baseline variables achieved an AUC of 0.73 (95% CI: [0.66–0.80]). Models using only SNP features achieved only slightly worse performance with AUCs of 0.67 (95% CI: [0.59–0.75]) and 0.72 (95% CI: [0.64–0.80]) for FS-A and FS-C, respectively. Trimming FS-C to the top 10 substantially reduced performance (AUC 0.66, 95% CI: [0.58–0.74]) but trimming FS-A did not (AUC 0.68, 95% CI: [0.60–0.76]). Most of these SNP-only models exhibited diminished Brier scores and calibration metrics from the baseline model.

Table 2.

Summary of top-correlated preselected single nucleotide polymorphisms

SNP (No. of Alleles) SNP-Interaction Population N FS Population Adjusted Odds Ratio (95% CI) Population Adjusted p-Value Population+Base Adjusted Odds Ratio (95% CI) Population+Base Adjusted p-Value
rs749042615 (0) Full 176 A-10, B 0.29 (0.13–0.66) 0.004 0.24 (0.08–0.65) 0.006
rs3795646 (2) Not Hispanic 31 A-10, B 0.27 (0.09–0.68) 0.010 0.21 (0.06–0.60) 0.007
rs11588227 (2) Male 35 A-10 2.60 (1.20–5.70) 0.016 1.67 (0.66–4.18) 0.270
rs7521206 (0) Not Hispanic 111 A-10 0.47 (0.25–0.87) 0.016 0.44 (0.21–0.92) 0.030
rs56294486 (1) Male 28 A-10, B 0.26 (0.07–0.73) 0.019 0.31 (0.08–0.93) 0.053
rs10800426 (2) Not Hispanic 42 A-10 2.22 (1.10–4.52) 0.025 1.49 (0.63–3.49) 0.354
rs56294486 (2) Male 110 A-10 2.49 (1.09–6.27) 0.039 2.17 (0.84–6.19) 0.122
rs880296 (1) Not Hispanic 79 A-10 0.52 (0.28–0.96) 0.040 0.48 (0.23–1.01) 0.056
rs3795646 (1) White 29 A-10 2.72 (1.04–7.33) 0.043 1.03 (0.31–3.38) 0.961
rs3795646 (0) Black 39 B 2.22 (1.01–4.92) 0.048 3.85 (1.48–10.44) 0.007
rs880296 (0) Not Hispanic 25 B 2.40 (1.03–5.77) 0.045 2.6 (0.82–8.65) 0.107
rs749042615 (0) rs7521206 (0) Not Hispanic 93 C-10, D-10 0.31 (0.17–0.58) <0.001 0.22 (0.10–0.47) <0.001
rs56294486 (2) rs6427662 (1) Male 46 C-10, D-10 3.53 (1.71–7.41) <0.001 2.79 (1.18–6.70) 0.020
rs56294486 (2) rs7521206 (1) Male 39 C-10, D-10 3.61 (1.70–7.85) <0.001 4.06 (1.67–10.24) 0.002
rs11588227 (0) rs749042615 (0) Full 131 C-10, D-10 0.39 (0.21–0.71) 0.002 0.40 (0.19–0.81) 0.012
rs11588227 (2) rs347278 (2) Full 34 C-10 3.24 (1.53–7.03) 0.002 2.81 (1.11–7.26) 0.030
rs10753755 (2) rs749042615 (0) Not Hispanic 135 C-10 0.35 (0.17–0.69) 0.003 0.25 (0.10–0.59) 0.002
rs10737537 (2) rs749042615 (1) Full 21 C-10, D-10 4.31 (1.7–11.91) 0.003 4.78 (1.56–16.16) 0.008
rs749042615 (0) rs880296 (1) Not Hispanic 67 C-10 0.37 (0.19–0.72) 0.004 0.35 (0.15–0.76) 0.009
rs347278 (2) rs56294486 (2) Male 69 C-10 2.78 (1.39–5.67) 0.004 2.06 (0.91–4.71) 0.083
rs7513166 (2) rs7521206 (0) Not Hispanic 72 C-10 0.39 (0.20–0.75) 0.005 0.36 (0.16–0.78) 0.012
rs2812136 (2) rs3795646 (0) Black 30 D-10 3.17 (1.36–7.58) 0.008 5.36 (1.94–15.79) 0.002
rs35222272 (0) rs3795646 (2) Not Hispanic 30 D-10 0.21 (0.06–0.57) 0.005 0.18 (0.05–0.55) 0.006
rs347278 (2) rs56294486 (1) Male 26 D-10 0.13 (0.02–0.45) 0.006 0.14 (0.02–0.54) 0.012
rs35222272 (0) rs749042615 (0) Female 31 D-10 0.21 (0.06–0.65) 0.009 0.08 (0.02–0.34) <0.001
rs1848843 (0) rs56294486 (2) Male 67 D-10 2.67 (1.34–5.42) 0.006 2.72 (1.20–6.34) 0.018

Population Adjusted Odds Ratios adjust for population-specific rates. Population+Base Adjusted Odds Ratios adjust for population-specific rates and three baseline variables.

CI, confidence interval; FS, feature set; SNP, single nucleotide polymorphism.

Table 3.

Performance of predictive models developed using preselected features

FS AUC (95% CI) Brier Score (95% CI) Calibration Intercept Calibration Slope
FS-A 0.67 (0.59–0.75) 0.21 (0.17–0.24) 0.00 0.78
FS-A-10 0.68 (0.6–0.76) 0.21 (0.18–0.24) 0.02 0.76
FS-C 0.72 (0.64–0.80) 0.19 (0.15–0.23) 0.03 0.83
FS-C-10 0.66 (0.58–0.74) 0.22 (0.19–0.24) 0.02 0.49
FS-A-10+FS-C-10 0.70 (0.62–0.78) 0.21 (0.18–0.24) −0.01 0.64
Base 0.73 (0.66–0.80) 0.19 (0.18–0.21) 0.01 0.98
Base+FS-B 0.80 (0.73–0.87) 0.17 (0.15–0.20) 0.05 1.02
Base+FS-D 0.79 (0.73–0.86) 0.17 (0.15–0.20) 0.05 0.89
Base+FS-D-10 0.81 (0.74–0.87) 0.17 (0.14–0.20) 0.05 0.89
Base+FS-B+FS-D-10 0.81 (0.74–0.87) 0.16 (0.14–0.19) 0.07 0.93

AUC, area under the receiver operating characteristic curve.

Including both SNP and baseline features in our models substantially improved prediction performance (Table 3). Adding SNP variables to the baseline model (Base+FS-B) resulted in an AUC of 0.80 (95% CI: [0.73–0.87]). The inclusion of all selected SNP-SNP interactions (Base+FS-D) did not improve AUC (AUC 0.79, 95% CI: [0.73–0.86]). However, including the top 10 SNP-SNP features from FS-D in addition to the Base+FS-B model resulted in the best AUC (0.81, 95% CI [0.74–0.87]) and Brier Score (0.16, 95% CI: [0.14–0.19]) (Table 3). A calibration curve comparing the predicted healing probabilities from the highest performing model to observed proportions of the healing outcome is presented in Supplementary Fig. S1.

SNP selection stability

A visualization of SNP selection is presented in Fig. 1. Eight of the 10 SNPs selected on the full cohort for FS-A-10 were selected by LASSO in at least 70% of the outer cross-validation folds. When replicating the FS-A-10 selection process within each fold, seven of these eight SNPs were selected in at least 60% of the folds. Among the five SNPs selected on the full cohort for FS-B, all were selected by LASSO in 100% of the training folds and four SNPs were selected in at least 70% of the within-fold replicated selection processes. Selection among SNPs and SNP-SNP interactions was less stable.

Figure 1.

Figure 1.

A visualization of feature selection stability. We consider sets of SNP features selected using data from 179 patients with complete data and then create 10 cross-validation training-testing splits. Within each training fold, we assess whether each feature was also selected by a LASSO model trained on these preselected SNP and baseline features (LASSO) or identified when replicating our feature selection approach (Within-Fold Selection). LASSO, least absolute shrinkage and selection operator; SNP, single nucleotide polymorphism.

While all 10 variables from FS-C-10 were selected by LASSO in at least half of the training folds, only 5 satisfied this condition when moving variable selection within each fold. Similarly, 9 of the 10 variables from FS-D-10 were selected by LASSO in at least half of the training folds, but only 3 satisfied the condition when using within-fold variable selection. The SNPs and SNP-SNP interactions consistently selected in these respective variable sets may have a higher likelihood of generalizability to other cohorts and are listed in Table 4.

Table 4.

Single nucleotide polymorphism and single nucleotide polymorphism-single nucleotide polymorphism interaction variables selected in at least 50% of training folds by both a least absolute shrinkage and selection operator model trained on preselected single nucleotide polymorphism and baseline features and when replicating our entire feature selection approach within a given training fold

FS SNP (No. of Alleles) SNP (No. of Alleles) Population Within-Fold Selection (%)
FS-A-10 rs10800426 (2) Not Hispanic 80
FS-A-10 rs11588227 (2) Male 60
FS-A-10, FS-B rs3795646 (2) Not Hispanic 90, 90
FS-A-10, FS-B rs56294486 (1) Male 100, 90
FS-A-10, FS-B rs749042615 (0) Full 100, 100
FS-A-10 rs7521206 (0) Not Hispanic 90
FS-A-10 rs880296 (1) Not Hispanic 60
FS-B rs3795646 (0) Black 70
FS-C-10 rs11588227 (2) rs347278 (2) Full 60
FS-C-10 rs10753755 (2) rs749042615 (0) Not Hispanic 50
FS-C-10, FS-D-10 rs749042615 (0) rs7521206 (0) Not Hispanic 90, 90
FS-C-10 rs56294486 (2) rs6427662 (1) Male 80
FS-C-10, FS-D-10 rs56294486 (2) rs7521206 (1) Male 80, 70
FS-D-10 rs2812136 (2) rs3795646 (0) Black 50

Exploratory models

Exploratory models using unsupervised feature engineering techniques were largely unsuccessful as the addition of PCA, SPCA, or auto-encoded features to the baseline model did not improve performance (Table 5). This may be due to the limited sample size resulting in added noise compared to relevant signal in feature derivation. Models on the 204 patients considered in the imputed datasets demonstrated similar performance to those on the 179 patients with complete baseline information (Table 5). There was a slight decrease in baseline model performance, as our imputed baseline predictors are quite variable. However, when we included the fully-captured SNP features, the models using imputed baseline and SNP features had comparable performance to models using the subset of patients with complete information.

Table 5.

Model performance of exploratory analyses using baseline features and single nucleotide polymorphism features derived from unsupervised training techniques (N = 179) and multiple imputation analyses (N = 204)

N FS AUC (95% CI) Brier Score (95% CI) Calibration Intercept Calibration Slope
179 Base+PCA 0.72 (0.64–0.79) 0.20 (0.18–0.22) −0.01 0.98
179 Base+SPCA 0.68 (0.6–0.76) 0.21 (0.19–0.22) 0.01 0.97
179 Base+AE 0.68 (0.59–0.76) 0.20 (0.19–0.22) 0.00 0.87
204 FS-A-10+FS-C-10 0.72 (0.65–0.79) 0.20 (0.17–0.23) −0.01 0.88
204 Base-Imputed 0.71 (0.63–0.79) 0.20 (0.18–0.23) 0.00 0.93
204 Base-Imputed+FS-B+FS-D-10 0.81 (0.74–0.87) 0.17 (0.14–0.21) 0.00 0.84

PCA, principal component analysis; SPCA, sparse principal component analysis.

DISCUSSION

DFUs are one of the most serious comorbidities of DM. Over an individual's lifetime with DM, about 30% will have a DFU and DFUs are the strongest predictor for LEA.3 Very few therapies for DFU have received FDA approval to heal foot ulcers. Over the past 20 years, several investigators have created prognostic models to help wound care providers care for individuals with DFU. Treatment management decisions are often based on the likelihood that an individual will heal with more advanced therapies like tissue substitutes being used for harder to heal wounds.3,41,42

Prognostic models are also used to help providers inform patients about the likelihood that they will heal with standard care. The majority of the currently used prognostic tools, including previous studies by the authors of this study, rely solely on simple-to-measure clinical evaluations like wound area and wound duration and standard regression techniques.16,20–22 In our current investigation, we were able to improve upon traditional models by identifying potentially-relevant NOS1AP SNPs and using advanced statistical and machine learning techniques. Our resulting prognostic models had AUC between 0.79 and 0.81 and Brier scores between 0.16 and 0.17, a 10–16% absolute improvement over previous models that rely only on wound size, wound duration, and wound depth.15–17,20,43

We know of no previous chronic wound care study that has used similar statistical and machine learning techniques to rigorously evaluate the predictive potential of NOS1AP information or variation in any other gene.15–17,20,43 The precision and accuracy of the models presented in this study are superior to simpler models developed previously, thus demonstrating the potential importance of genetic information and the described computer-based technologically advanced prognostic modeling methods in predicting DFU healing.44

External valıdation of new prognostic models on independent cohorts is essential when deploying models for clinical practice, and therefore, this is an important step for future work. For example, a recent study by Jull et al. demonstrated the generalizability and reproducibility of using simple clinical evaluations to predict the likelihood that a venous leg ulcer (another chronic wound) might heal.22 Still, both CEPCs and NOS1AP variants have been associated with DFU healing in previous studies using independent cohorts,14,23–26,45 and many prior publications have evaluated simple wound measures like wound size, wound duration, and depth.16,21,22,46,47 Furthermore, like most studies of prognostic models, our models were created by subdividing our dataset into a training dataset and validation dataset.20–22,46,47 While this form of validation cannot replicate the rigor achieved using an entirely independent external dataset, it does enhance the internal and external validity of prognostic modeling methods.44

While genetic assays of NOS1AP and serum evaluations of CEPC might be beyond the scope of routine clinical practice, the methods described in this study can be used to evaluate other potential risk factors as well.24,26 As there has been a paucity of recent therapeutic advances in DFU treatment, creative new approaches like those described in our study which identify genetic components of healing may unlock important insights for developing new therapies. The models that we present in this study demonstrate the importance of interactions between different factors associated with wound healing.

There are numerous challenges when building prediction models incorporating high-dimensional genetic information. First, one must balance exploration of flexible interactions between SNPs, clinical measures, and healing with sample size constraints. In our approach, we exploit prior knowledge that genetic components are known to be important for similar subpopulations, including differing ethnicity, race, and sex, by specifying interaction terms between SNPs and subpopulation indicators as predictors but allowing data-adaptive learners like BART to explore little-known complex relationships between genetic and clinical predictors according to their informativeness within our cohort.

In secondary and exploratory analyses, we assessed different approaches for representing genetic structure. We found similar performance to primary analyses when directly including SNP-SNP interactions as candidate predictors, but reduced performance with unsupervised feature engineering techniques. Development of larger cohorts or automated approaches that exploit known genetic relationships may improve upon these results and help identify important mechanisms for healing. Learners designed to counter overfitting like BART and Elastic Net performed better than alternatives without these mechanisms like Gradient Boosting Machines (GBM) and Logistic Regression, which should be noted for future DFU prediction models in high-dimensional settings.36

As noted, previously described prognostic tools tended to focus on using data elements easy to collect at the time of patient evaluation and then successfully predict which wounds will heal about 60–70% of the time.16,17,20–22 Compared to previous publications, the models presented in this report used more complex statistical and informatics tools, as well as more robust forms of data collection (e.g., genetic and biochemical biomarkers). Using these methods, the likelihood of successful prediction was >80%. Predictive models with AUC >80% are considered to be almost perfect.44,48

The measurement of CEPCs and genetic variation of NOS1AP may not be practically feasible in clinical practice at this time. However, these improved prediction modeling techniques can still be used to develop simplified models for practical use. Furthermore, by presenting these new models, we are creating a new benchmark for which future models can be compared against. Such predictive models will enable wound care experts to more accurately prognosticate whether a patient will heal and will aid clinical trialists in designing better studies to evaluate therapies for subjects likely or unlikely to heal. For example, Jull et al. demonstrated the importance of stratifying the likelihood that a venous leg ulcer might heal with respect to the investigative agent.22 Stratification and risk modeling may enhance both patient care and the efficiency of randomized clinical trial designs for investigational agents. Finally a recent publication by Bender et al. used a case based reasoning approach to create an algorithm-operated learning system.46 This artificial intelligence technique is an interesting approach that could be helpful in discovering new prognostic variables that could be included in advanced prognostic models like that described in our study.46

INNOVATION

Determining whether a patient with a DFU will heal is not an easy task. Researchers are yet to substantially improve upon prognostic models generally built using a few, simple-to-measure clinical factors. This study introduces a novel analysis of biochemical and genetic information as potential predictors and provides both a framework and benchmark for future research in this area using sophisticated statistical and bioinformatics tools. In the cohort used for this study, our methodology to incorporate genetic information led to an absolute 11% improvement in AUC over baseline methods. The goal of this article was to introduce computer-based technologically advanced prognostic modeling techniques that allow for use of large messy datasets, thereby improving our ability to prognostic the outcome for an individual with a DFU.

KEY FINDINGS

  • Prediction models have an important role with respect to clinical care, design of studies, and risk stratification of patients with chronic wounds such as DFUs.

  • Previous models have primarily used simple to measure clinical wound factors with moderate success.

  • We demonstrate the use of blood based biomarkers and statistical and machine learning techniques can markedly improve the prognostic accuracy of DFU prognostic models.

DATA AVAILABILITY

The NOS1AP genomic dataset is available on Sequence Read Archive (SRA) maintained by the National Library of Medicine. It is submission SUB11343510 and was released on May 30, 2023.

ACKNOWLEDGMENTS AND FUNDING SOURCES

This work was supported, in part, by a grant from the National Institutes for Health (NIDDK) R01-DK116199 (Multiple principle investigators D.J.M./S.R.T.). The sponsors had no role in the design and conduct of the study; collection, management, analysis, and interpretation of the data; preparation, review, or approval of the article; and the decision to submit the article for publication.

AUTHOR DISCLOSURE AND GHOSTWRITING

The authors have no competing financial interest with respect to the content of the article. The article was written only by the authors listed.

ABOUT THE AUTHORS

Gary Hettinger, PhD is a Doctoral Student in Biostatistics, Nandita Mitra, PhD is a Professor of Biostatistics, Stephen R. Thom, MD, PhD is a Professor of Emergency Medicine, David J. Margolis, MD, PhD is a Professor of Dermatology and Epidemiology.

SUPPLEMENTARY MATERIAL

Supplementary Figure S1

Abbreviations and Acronyms

AUC

area under the receiver operating characteristic curve

BART

Bayesian Additive Regression Tree

CEPC

circulating endothelial precursor cell

CI

confidence interval

DFU

diabetic foot ulcer

DFUC

Diabetic Foot Ulcer Consortium

DM

diabetes mellitus

FS

feature set

GAM

generalized additive model

GLM

generalized linear model

LASSO

least absolute shrinkage and selection operator

LEA

lower extremity amputation

NOS1AP

nitric oxide synthase-one accessory protein

PCA

principal component analysis

SD

standard deviation

SL

SuperLearner

SNP

single nucleotide polymorphism

SPCA

sparse principal component analysis

REFERENCES

  • 1. Carter MJ, DaVanzo J, Haught R, et al. Chronic wound prevalence and the associated cost of treatment in Medicare beneficiaries: Changes between 2014 and 2019. J Med Econ 2023;26(1):894–901; doi: 10.1080/13696998.2023.2232256 [DOI] [PubMed] [Google Scholar]
  • 2. Boulton AJ, Vileikyte L, Ragnarson-Tennvall G, et al. The global burden of diabetic foot disease. Lancet 2005;366(9498):1719–1724. [DOI] [PubMed] [Google Scholar]
  • 3. Armstrong DG, Boulton AJM, Bus SA. Diabetic foot ulcers and their recurrence. N Engl J Med 2017;376(24):2367–2375; doi: 10.1056/NEJMra1615439 [DOI] [PubMed] [Google Scholar]
  • 4. Margolis D, Malay DS, Hoffstad OJ, et al. Incidence of Diabetic Foot Ulcer and Lower Extremity Amputation Among Medicare Beneficiaries, 2006 to 2008. Agency for Healthcare Research and Quality: Rockville, MD; 2010. [PubMed] [Google Scholar]
  • 5. Margolis D, Malay DS, Hoffstad OJ, et al. Prevalence of Diabetes, Diabetic Foot Ulcer, and Lower Extremity Amputation Among Medicare Beneficiaries, 2006 to 2008. Agency for Healthcare Research and Quality: Rockville, MD; 2010. [PubMed] [Google Scholar]
  • 6. Margolis D, Malay DS, Hoffstad OJ, et al. Economic Burden of Diabetic Foot Ulcers and Amputations Among Medicare Beneficiaries, 2006 to 2008. Agency for Healthcare Research and Quality: Rockville, MD; 2010. [PubMed] [Google Scholar]
  • 7. Rice JB, Desai U, Ristovska L, et al. Economic outcomes among Medicare patients receiving bioengineered cellular technologies for treatment of diabetic foot ulcers. J Med Econ 2015;18(8):586–595; doi: 10.3111/13696998.2015.1031793 [DOI] [PubMed] [Google Scholar]
  • 8. Walsh JW, Hoffstad OJ, Sullivan MO, et al. Association of diabetic foot ulcer and death in a population-based cohort from the United Kingdom. Diabet Med 2016;33(11):1493–1498; doi: 10.1111/dme.13054 [DOI] [PubMed] [Google Scholar]
  • 9. Hoffstad O, Mitra N, Walsh J, et al. Diabetes, lower-extremity amputation, and death. Diabetes Care 2015;38(10):1852–1857; doi: 10.2337/dc15-0536 [DOI] [PubMed] [Google Scholar]
  • 10. Armstrong DG, Swerdlow MA, Armstrong AA, et al. Five year mortality and direct costs of care for people with diabetic foot complications are comparable to cancer. J Foot Ankle Res 2020;13(1):16; doi: 10.1186/s13047-020-00383-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Margolis DJ, Jeffcoate W. Epidmeiology of foot ulcerationand amputation: Can global variation be explained? Med Clin North Am 2013;97:791–805. [DOI] [PubMed] [Google Scholar]
  • 12. Boulton AJM. The pathway to foot ulceration in diabetes. Med Clin North Am 2013;97:775–790. [DOI] [PubMed] [Google Scholar]
  • 13. Malay D, Margolis DJ, Hofstad OJ, et al. The incidence and risks of failure to heal following lower extremity amputation for the treatment of diabetic neuropathic foot ulcer. J Foot Ankle Surg 2006;45:366–375. [DOI] [PubMed] [Google Scholar]
  • 14. Margolis DJ, Mitra N, Malay DS, et al. Further evidence that wound size and duration are strong prognostic markers of diabetic foot ulcer healing. Wound Repair Regen 2022;30(4):487–490; doi: 10.1111/wrr.13019 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Margolis DJ, Kantor J, Santanna J, et al. Risk factors for delayed healing of neuropathic diabetic foot ulcers: A pooled analysis. Arch Dermatol 2000;136(12):1531–1535. [DOI] [PubMed] [Google Scholar]
  • 16. Margolis DJ, Taylor LA, Hofstad O, et al. Diabetic neuropathic foot ulcer: The association of wound size, wound duration, and wound grade. Diabetes Care 2002;25:1835–1839. [DOI] [PubMed] [Google Scholar]
  • 17. Margolis DJ, Taylor LA, Hofstad O, et al. Diabetic neuropathic foot ulcers: Predicting which ones will heal. Am J Med 2003;115:627–631. [DOI] [PubMed] [Google Scholar]
  • 18. Kurd SK, Hoffstad OJ, Bilker WB, et al. Evaluation of the use of prognostic information for the care of individuals with venous leg ulcers or diabetic neuropathic foot ulcers. Wound Repair Regen 2009;17(3):318–325; doi: 10.1111/j.1524-475X.2009.00487.x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Fife CE, Eckert KA, Carter MJ. Publicly reported wound healing rates: The fantasy and the reality. Adv Wound Care (New Rochelle) 2018;7(3):77–94; doi: 10.1089/wound.2017.0743 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Fife CE, Horn SD. The wound healing index for predicting venous leg ulcer outcome. Adv Wound Care (New Rochelle) 2020;9(2):68–77; doi: 10.1089/wound.2019.1038 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Fife CE, Horn SD, Smout RJ, et al. A predictive model for diabetic foot ulcer outcome: The Wound Healing Index. Adv Wound Care (New Rochelle) 2016;5(7):279–287; doi: 10.1089/wound.2015.0668 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Jull A, Lu H, Jiang Y. A simple index to predict healing in venous leg ulcers: A secondary analysis from four randomised controlled trials. J Wound Care 2023;32(10):657–664; doi: 10.12968/jowc.2023.32.10.657 [DOI] [PubMed] [Google Scholar]
  • 23. Margolis DJ, Hoffstad O, Thom SR. NOS1AP is associated with impaired healing of diabetic foot ulcer and diminished response to healing of circulating stem/progenitor cells. Wound Repair Regen 2017;25(4):733–736. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Thom SR, Bhopale VM, Margolis DJ. Neutrophil microparticle production and inflammasome activation by hyperglycemia due to cytoskeletal instability. J Biol Chem 2017;292(44):18312–18324. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Margolis DJ, Hampton M, Hoffstad O, et al. NOS1AP genetic variation is associated with impaired healing of diabetic foot ulcers and diminished response to healing of circulating stem/progenitor cells. Wound Repair Regen 2017;25(4):733–736; doi: 10.1111/wrr.12564 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Margolis DJ, Mitra N, Hoffstad O, et al. Circulating endothelial precursor cells are associated with a healed diabetic foot ulcer evaluated in a prospective cohort study. Wound Repair Regen 2023;31(1):128–134; doi: 10.1111/wrr.13055 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Thum T, Fraccarollo D, Schultheiss M, et al. Endothelial nitric oxide synthase uncoupling impairs endothelial progenitor cell mobilization and function in diabetes. Diabetes 2007;56(3):666–674; doi: 10.2337/db06-0699 [DOI] [PubMed] [Google Scholar]
  • 28. Husson F, Le S, Pagès J.. Exploratory Multivariate Analysis by Example Using R. Chapman and Hall: London; 2010. [Google Scholar]
  • 29. Lammers B. ANN2: Artificial Neural Networks for Anomaly Detection. R package version 2.3.4. 2020. Available from: https://CRANR-projectorg/package=ANN2 [Last accessed: January 6, 2024].
  • 30. Erichson NB, Zheng P, Manohoar K, et al. Sparse principal component analysis via variable projection. SIAM J Appl Maths 2020;80(2); doi: 10.1137/18M1211350 [DOI] [Google Scholar]
  • 31. Erichson NB, Zheng P, Aravkin S. sparsepca: Sparse Principal Component Analysis (SPCA). R package version 0.1.2. 2018. Available from: https://CRANR-projectorg/package=sparsepca [Last accessed: January 6, 2024].
  • 32. Polley E, LeDell E, Kennedy C, et al. SuperLearner: Super Learner Prediction. R package version 2.0-28. 2021. Available from: https://CRANR-projectorg/package=SuperLearner [Last accessed: January 6, 2024].
  • 33. van der Laan MJ, Polley EC, Hubbard AE. Super learner. Stat Appl Genet Mol Biol 2007;6:Article 25; doi: 10.2202/1544-6115.1309 [DOI] [PubMed] [Google Scholar]
  • 34. Sparapani R, Spanbauer C, McCulloch R. Nonparametric machine learning and efficient computation with Bayesian additive regression trees: The BART R package. J Stat Softw 2021;97:1–66. [Google Scholar]
  • 35. Friedman J, Hastie T, Tibshirani R. Regularization paths for generalized linear models via coordinate descent. J Stat Softw 2010;33(1):1–22. [PMC free article] [PubMed] [Google Scholar]
  • 36. Hastie T. gam: Generalized Additive Models. R package version 1.20. 2020. Available from: https://CRANR-projectorg/package=gam [Last accessed: January 6, 2024].
  • 37. Assel M, Sjoberg DD, Vickers AJ. The Brier score does not evaluate the clinical utility of diagnostic tests or prediction models. Diagn Progn Res 2017;1:19; doi: 10.1186/s41512-017-0020-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Van Calster B, McLernon DJ, van Smeden M, et al. Calibration: The Achilles heel of predictive analytics. BMC Med 2019;17(1):230; doi: 10.1186/s12916-019-1466-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Moons KG, Altman DG, Reitsma JB, et al. Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD): Explanation and elaboration. Ann Intern Med 2015;162(1):W1–W73; doi: 10.7326/m14-0698 [DOI] [PubMed] [Google Scholar]
  • 40. LeDell E, Petersen M, van der Laan M. Computationally efficient confidence intervals for cross-validated area under the ROC curve estimates. Electron J Stat 2015;9(1):1583–1607; doi: 10.1214/15-ejs1035 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Steed DL, Attinger C, Colaizzi T, et al. Guidelines for the treatment of diabetic ulcers. Wound Repair Regen 2006;14(6):680–692. [DOI] [PubMed] [Google Scholar]
  • 42. Everett E, Mathioudakis N. Update on management of diabetic foot ulcers. Ann N Y Acad Sci 2018;1411(1):153–165; doi: 10.1111/nyas.13569 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Kantor J, Margolis DJ. The efficacy and prognostic value of simple wound measurements. Arch Dermatol 1998;134:(12):1571–1574. [DOI] [PubMed] [Google Scholar]
  • 44. Harrell FE Jr. Regression Modeling Strategies. Springer: New York; 2001. [Google Scholar]
  • 45. Thom SR, Bhopale VM, Arya AK, et al. Blood-borne microparticles are an inflammatory stimulus in type 2 diabetes mellitus. Immunohorizons 2023;7(1):71–80; doi: 10.4049/immunohorizons.2200099 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Bender C, Cichosz SL, Malovini A, et al. Using case-based reasoning in a learning system: A prototype of a pedagogical nurse tool for evidence-based diabetic foot ulcer care. J Diabetes Sci Technol 2021;16(2):454–459; doi: 10.1177/1932296821991127 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. Bender C, Cichosz SL, Pape-Haugaard L, et al. Assessment of simple bedside wound characteristics for a prediction model for diabetic foot ulcer outcomes. J Diabetes Sci Technol 2020;15(5):1161–1167; doi: 10.1177/1932296820942307 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics 1977;33(1):159–174; doi: 10.2307/2529310 [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Figure S1

Data Availability Statement

The NOS1AP genomic dataset is available on Sequence Read Archive (SRA) maintained by the National Library of Medicine. It is submission SUB11343510 and was released on May 30, 2023.


Articles from Advances in Wound Care are provided here courtesy of SAGE Publications

RESOURCES