Abstract
INTRODUCTION
Most people with Alzheimer's disease and related dementia (ADRD) also suffer from two or more chronic conditions, known as multiple chronic conditions (MCC). While many studies have investigated the MCC patterns, few studies have considered the synergistic interactions with other factors (called the syndemic factors) specifically for people with ADRD.
METHODS
We included 40,290 visits and identified 18 MCC from the National Alzheimer's Coordinating Center. Then, we utilized a multi‐label XGBoost model to predict developing MCC based on existing MCC patterns and individualized syndemic factors.
RESULTS
Our model achieved an overall arithmetic mean of 0.710 AUROC (SD = 0.100) in predicting 18 developing MCC. While existing MCC patterns have enough predictive power, syndemic factors related to dementia, social behaviors, mental and physical health can improve model performance further.
DISCUSSION
Our study demonstrated that the MCC patterns among people with ADRD can be learned using a machine‐learning approach with syndemic framework adjustments.
Highlights
Machine learning models can learn the MCC patterns for people with ADRD.
The learned MCC patterns should be adjusted and individualized by syndemic factors.
The model can predict which disease is developing based on existing MCC patterns.
As a result, this model enables early specific MCC identification and prevention.
Keywords: ADRD, dementia, disease patterns, early prevention, machine learning, syndemic framework
1. BACKGROUND
Alzheimer's disease and related dementia (ADRD) is a set of progressive neurodegenerative diseases that impair a person's cognitive behavior, memory, problem‐solving, and other thinking skills, primarily in older adults. With the growing aging population in the United States (U.S.), ADRD is currently impacting 6.7 million people, costing $345 billion from total Medicare and Medicaid payments in 2023. 1 It is a devastating and costly disease because 87% of people with ADRD (roughly 2.7 times more than people without ADRD) also suffer from two or more chronic conditions, 1 , 2 commonly known as multimorbidity or multiple chronic conditions (MCC). Studies have shown that MCC among people with ADRD not only heightened mortality risks, 3 but also increased costs substantially for healthcare utilization. 3 , 4 , 5 , 6 In particular, they found that people with ADRD and more than five comorbidities had 1.31 times more primary care consultations, 3 had 1.56 times higher risks of death, 3 and incurred $27–$54K more in healthcare costs than people without ADRD per year. 4 Thus, understanding the MCC patterns is crucial to managing healthcare for people with ADRD.
The MCC patterns can be referred to as which diseases are directly responsible for other diseases (causal multimorbidity) 7 or which groups of diseases simultaneously occur more than often by chance (cluster multimorbidity). 8 , 9 Many have identified MCC clusters using factor analysis, network and cluster analysis, and association rules. 9 , 10 , 11 , 12 However, this type of analysis does not identify which disease exactly within the MCC cluster would happen next. For instance, knowing that cluster A‐B‐C is the commonly reported MCC cluster, it is still challenging to identify whether disease B or C will happen next if a patient has disease A. Others have explored the MCC patterns to predict which disease would happen next using artificial intelligence/machine learning (AI/ML) models. This includes link predictions with graphical models 13 and multi‐label classifiers such as logistic regression, random forest, and support vector machines. 14 , 15 These studies have provided an early prevention tool for healthcare professionals to personalize effective care for their patients. Yet, studies that investigated the MCC patterns for people with ADRD are limited.
Current literature has suggested that the interactions between MCC and ADRD are far more complicated than other chronic conditions. Specifically, cognitive impairment may complicate clinical care and make it harder to detect other conditions; patients have more difficulty in expressing their discomfort while clinicians might be predominantly focusing on behavioral or psychological symptoms. 16 This relationship is also further complicated by polypharmacy and psychosocial factors. 16 To conceptualize this complex relationship, Dunn et al. proposed a theoretical syndemic framework that emphasized eight major syndemic factors specifically for people with ADRD: lifestyles, social, environmental, carer/family, mental health, physical health, polypharmacy, and severity and type of dementia. 17 In their proposed syndemic framework, 18 they highlighted that chronic conditions do not just exist in parallel, but involve complex interactions that frequently amplify each other's effects, consequently leading to more severe conditions. With that in mind, to understand the MCC patterns for people with ADRD, one would need to learn the complex relationships between MCC and ADRD, and their intrinsic interactions with eight other major syndemic factors. However, current models have yet to apply the theoretical syndemic framework for people with ADRD using AI/ML models to learn the complex relationships of MCC patterns and syndemic factors.
Hence, in this study, we proposed to learn the complex MCC patterns adjusted by the theoretical syndemic framework using a multi‐label AI/ML approach with imbalanced learning methods. Specifically, we proposed to predict the developing (future) MCC for people with ADRD based on interactions between existing MCC patterns and individualized syndemic factors. As a proof‐of‐concept, we hypothesized that a multi‐label classifier with imbalanced learning methods could capture the complex MCC patterns for multiple diseases at once. Furthermore, our secondary objective is to understand how the syndemic factors affect the AI/ML model performance.
2. METHODS
2.1. Data source
This study used survey data called Uniform Data Sets (UDS) from the National Alzheimer's Coordinating Center (NACC). 19 NACC collaborates with more than 42 Alzheimer's Disease Research Centers throughout the United States to collect data from volunteers or referral‐based participants with ranging cognitive statuses, from normal cognition to dementia. The UDS consists of a total of 2166 survey data elements across 20 types of survey forms (Table S1 ) that were completed by the clinicians at each patient's annual visit, including demographics and clinical data since 2005.
2.2. Cohort selection
We included patients whose data were collected using the latest version of the UDS—Version 3. We chose the latest version because it contains more information about the patient's health conditions. For instance, Form D2: Clinician‐assessed Medical Conditions (a major source for our outcomes) is only available in Version 3. We also further excluded patients who died.
As people with ADRD often develop MCC over a number of years, the MCC patterns can be learned for each individual throughout their medical history and across individuals. In other words, observing newly developed MCC at multiple time points for an individual would assist machine learning models in discovering the complex MCC patterns for that individual, as well as generalized MCC patterns for the entire ADRD cohort. Moreover, people with MCC can also have higher risks of developing ADRD. Therefore, we will include as many annual visits as possible for each patient regardless of their cognitive status. Using the previous annual visit (hereby referred to as the baseline visits) information from the patient, the AI/ML model will learn to predict the MCC that are presented in the patient's next visit (hereby referred to as the follow‐up visits). As a result, we excluded patients who only had one visit record.
2.3. Label processing
To predict future MCC (i.e., new MCC) in the follow‐up visits, we first identified a total of 18 MCC as labels to be predicted. These 18 MCC were chosen from two of the survey forms that contain patient diagnoses assessed by clinicians: (1) Form D1 ‐ Clinician Diagnosis that contains dementia‐related diseases, and (2) Form D2 ‐ Clinician‐assessed Medical Conditions that contains other chronic conditions. From Form D1, we chose three dementia‐related disease indications (any dementia, mild cognitive impairment (MCI), Alzheimer's disease) because these diseases are more prevalent and related to our study. On the other hand, we identified 15 other chronic conditions from Form D2 (cancer, diabetes, myocardial infarction, congestive heart failure, atrial fibrillation, angina, hypercholesterolemia, vitamin B12 deficiency, thyroid disease, arthritis, urinary incontinence, bowel incontinence, sleep apnea, rapid eye movement (REM) sleep behavior disorder, hyposomnia/insomnia) after excluding procedure‐related indications from this form. These disease indications were first processed into binary flags of 0 or 1. Then, the disease indications identified at the follow‐up visits were compared to previous baseline visits. If any of the diseases existed previously in the patient history records, the disease would be considered as a pre‐existing disease and, hence, would be labeled as 0 in the outcome label since it was not a new disease that happened in the follow‐up visits. Therefore, an outcome label of 1 indicates that this patient has not had the disease before but has developed the disease during the follow‐up visit. Although we only included visit records that used the latest form (V3), we searched through the entire patient history to find these disease indications. We also further included disease‐related medications. For example, any use of U.S. Food and Drug Administration (FDA)‐approved medication for Alzheimer's disease symptoms, including cholinesterase inhibitors and memantine would be considered a positive indication of Alzheimer's disease (i.e., 1 for Alzheimer's disease flag). Hypertension was excluded because every participant had hypertension at baseline; hence, it can be assumed they still have hypertension in the follow‐ups.
RESEARCH IN CONTEXT
Systematic review: The authors reviewed the literature using traditional sources (e.g., PubMed). While many have learned the multiple chronic conditions (MCC) patterns, few have incorporated the interactions with syndemic factors for people with Alzheimer's disease and related dementia (ADRD).
Interpretation: Our proof‐of‐concept study showed that advanced machine learning models can achieve moderately good performance in predicting specific developing MCCs for people with ADRD using existing MCC patterns and syndemic factors. This demonstrated that the complex MCC patterns can be learned with adjustments of the syndemic factors. By learning the MCC patterns, our model may assist clinicians in early disease identification and ultimately enable early prevention for people with ADRD.
Future directions: While our model performed well in predicting dementia‐related MCCs, there are also areas for improvement in predicting other chronic conditions. Due to the limitations of the dataset, future directions include utilizing electronic health records as they contain more comprehensive healthcare data variables.
2.4. Feature selection
In order to learn future MCC from existing MCC patterns and syndemic factors, we included the 18 outcome variables that we identified earlier as features using patients’ baseline visits. In other words, instead of collecting the 18 MCC in the follow‐up visits, we collect these indicators for each patient in their baseline visits as features. If patients had any history of the MCC on or before the baseline visits, they would have a value of 1, else 0. We also extracted patient characteristics and information about their caretakers in their baseline visits from across the 20 available survey forms. These variables were further manually grouped into the eight major categories described by Dunn et al. 17 We identified a total of 1,281 variables: 18 MCC, 13 lifestyles, 88 social, 1 environmental, 212 carer/family, 139 mental health, 294 physical health, 29 polypharmacy, and 487 severity and type of dementia. As expected, we identified the most features related to ADRD since NACC is collecting data for ADRD purposes. We were only able to identify 1 environmental‐related variable as environmental factors usually involve pollution levels that are difficult to obtain in medical data.
There are four different data types in the UDS dataset: (1) binary variables, (2) categorical variables, (3) continuous variables, and (4) free‐text variables. We excluded free‐text variables as they often contain very specific information that is not generalizable to many patients. To prepare the data for our AI/ML model, we further processed categorical data using one‐hot encoding and replaced outliers in continuous variables using the provided reference range to missing values. Any missing values in the binary data were treated as “no” or “not found”, that is, 0.
2.5. Model development
To predict multiple diseases (i.e., MCC) at once, we adopted a multi‐label classification problem setting. Unlike multi‐class problems, multi‐label allows for more than one label (i.e., MCC in our study) per sample. However, if a patient is well‐managed, rarely does a patient develop new diseases at every single visit (i.e., every year). This poses a new challenge on top of the multi‐label problem, that is, imbalanced data. Imbalanced data are a type of data distribution where the cases (patients with the disease) are significantly less than the control (patients who do not have the disease). It is difficult to handle as AI/ML models as it can achieve high accuracy by simply classifying all input as control. However, this kind of model will not be helpful as they cannot identify which future MCC will occur.
To identify the best AI/ML model, we carried out a sensitivity analysis by comparing the model performance across seven commonly used AI/ML models: Logistic Regression, Naïve Bayes, Random Forest, AdaBoost, XGBoost, Multi‐layer Perceptron Classifier, and Deep Neural Networks (five hidden layers ). As we have previously discussed, imbalanced data pose a special challenge where accuracy is not the best metric to evaluate model performance. Hence, to select the best model, we used the area under the receiver operating characteristic curve (AUROC) to measure the degree of separability among cases and control. The higher the AUROC (closer to 1), the better the model can discern cases and control. On the other hand, if the AUROC is 0.5, the model performance is as good as a random guess. Through this experiment, we found Logistic Regression had the worst performance based on AUROC, while boosted trees (AdaBoost and XGBoost) and Deep Neural Networks had better AUROC (Table S2). We decided to implement XGBoost because it has higher performance than Deep Neural Networks but has more flexibility in hyperparameter fine‐tuning than AdaBoost. Moreover, XGBoost has also been the state‐of‐the‐art model for structural data due to its ability to convert weak learners to strong learners with sequential ensemble learning techniques.
The hyperparameters of the XGBoost model were set fixed across MCC. Through empirical testing, we found that the following threshold for their respective hyperparameters can achieve reasonable performance without overfitting: 1000 number of estimators, 25 minimum child weight, and a maximum depth of 5; 70% of features and 70% of samples were randomly chosen in each iteration; focal loss function was used as the objective function to address the class imbalanced issue; early stopping round was set at 25 to avoid overfitting (i.e., stop training after model did not improve after 25 iterations), and the best iteration was retained as the final model. No further hyperparameter finetuning for each MCC as further finetuning did not result in better model performance. To handle the potentially skewed distribution of MCC probabilities, we also chose a disease‐specific optimal threshold for each MCC using a geometric mean of sensitivity and specificity (). Sensitivity is a metric that ranged from 0 to 1, it tests the ability of the model to detect cases while specificity (also ranging from 0 to 1) tests the model ability to detect control. If a model has high sensitivity (closer to 1) and high specificity (closer to 1), the model can detect cases and control accurately. Therefore, a higher G‐Mean means the model has a good balanced ability in detecting both control and cases. The optimal threshold was set at the threshold level that has the highest G‐Mean across all training data.
2.6. Model evaluation
Recall that we need to use patients’ multiple visit time points to understand which disease will happen at the next visit. This poses a problem in our NACC example as we may violate the assumption of data independence as one patient might contribute as more than one sample. As a result, it can cause data leakage during a random train‐test split as the testing dataset might be predicting the patient's MCC profile at visit 2 (labels) using the visit 1(features) record, while the patient's complete record (at visit 3 or later) might be included in the training dataset (has already included visits 1 and 2). To overcome this, we chose cluster sampling grouped by patients on top of the conventional random train‐test split. Cluster sampling identifies unique patient clusters and randomly chooses clusters in the train‐test split. This will not only avoid data leakage but also allow us to study the full MCC pattern of individuals.
Given that the goal of our study is to identify developing (future) MCC, it would be redundant to predict the likelihood of developing MCC that a patient already has. For example, if a patient already has dementia in baseline, it would be redundant for the model to predict if this patient will develop dementia in the next follow‐up since this is a chronic disease and we can assume that the patient will still suffer from dementia. Hence, we also removed patients who already have the disease from the particular MCC prediction model (both training and testing).
To ensure the generalizability of the model, we performed five‐fold cross‐validation for each MCC. On top of accuracy, we reported the average and standard deviation of the AUROC score, sensitivity, and specificity for each MCC across all five folds to understand the model performance. We also summarized the overall model performance using an arithmetic average across all MCC.
2.7. Secondary analysis
To achieve our secondary objective of understanding the effect of each syndemic factor on MCC prediction model performance, we evaluated the model performance using forward selection of syndemic factors to add to the baseline model incrementally (Table 1). We defined the baseline model (V1) as the model that only fitted with the 18 existing MCC to predict future MCC. The baseline model was evaluated using AUROC. With each addition of syndemic factors, in a particular order of lifestyles, social, environmental, carer/family, mental health, physical health, polypharmacy, and dementia‐related features, the model was evaluated with AUROC and compared to the previous model performance. The increment or decrease in AUROC due to the addition of the syndemic factor was visualized in a stacked bar chart across all 18 MCC.
TABLE 1.
Model versions and definitions
| Model version | Model definition |
|---|---|
| V1 | Included 18 existing multiple chronic conditions (MCC) only |
| V2 | V1 + 13 lifestyles features |
| V3 | V2 + 88 social features |
| V4 | V3 + 1 environmental features |
| V5 | V4 + 212 carer/family features |
| V6 | V5 + 139 mental health features |
| V7 | V6 + 311 physical health features |
| V8 | V7 + 29 polypharmacy features |
| V9 | V8 + 492 severity and types of dementia features |
Note: The model versions are numbered from V1 to V9. At each of the subsequent models, an additional group of features (of the identified eight syndemic factors) is added to the previous model. The type of syndemic factors and the total number of features are described in the Model Definition.
3. RESULTS
3.1. Study cohort baseline characteristics
This study included a total of 40,290 visit records among 11,538 unique patients. On average, each patient had three baseline visits, and the average length of the next visits (days between baseline and follow‐up visits) is a little over a year (430.3 days). Table 2 shows their baseline characteristics at their baseline visits. The average age of our study cohort was 74.6 years old. The proportion of females (58.7%) was slightly more than males (41.3%). The majority of the patients were White (83.4%), followed by Black (11.9%), Asian (2.9%), American Indian or Alaska Native (0.6%), Native Hawaiian or Other Pacific Islander (0.1%), and others/unknown (1.0%) patients. At baseline, 28.5% of our study cohort had mild cognitive impairment, 38.8% had Alzheimer's disease, and 20.7% had dementia. Among the other non‐dementia‐related chronic conditions, the most common diseases were hypercholesterolemia (67.7%), followed closely by arthritis (64.7%), and REM sleep behavior disorder (52.3%). On the other hand, congestive heart failure (3.2%) and myocardial infarction (5.4%) were the least prevalent diseases among our study cohort at their baseline visits.
TABLE 2.
Patient characteristics at baseline visits
| Parameter | Overall |
|---|---|
| No. of samples (n) | 40,290 |
| No. of patients (n) | 11,538 |
| Time to next follow‐up, mean in days (SD) | 430.3 (147.5) |
| Age, mean in years (SD) | 74.6 (9.7) |
| Female, n (%) | 23,650 (58.7) |
| Race | |
| White, n (%) | 33,593 (83.4) |
| Black, n (%) | 4,812 (11.9) |
| American Indian or Alaska Native, n (%) | 253 (0.6) |
| Native Hawaiian or Other Pacific Islander, n (%) | 33 (0.1) |
| Asian, n (%) | 1,177 (2.9) |
| Other/unknown, n (%) | 422 (1.0) |
| Dementia, n (%) | 8,354 (20.7) |
| Mild cognitive impairment, n (%) | 11,474 (28.5) |
| Alzheimer's disease, n (%) | 15,651 (38.8) |
| Cancer, n (%) | 8,310 (20.6) |
| Diabetes, n (%) | 6,525 (16.2) |
| Myocardial infarction, n (%) | 2,174 (5.4) |
| Congestive heart failure, n (%) | 1,304 (3.2) |
| Atrial fibrillation, n (%) | 4,348 (10.8) |
| Angina, n (%) | 4,620 (11.5) |
| Hypercholesterolemia, n (%) | 27,291 (67.7) |
| Vitamin B12 deficiency, n (%) | 3,344 (8.3) |
| Thyroid disease, n (%) | 9,718 (24.1) |
| Arthritis, n (%) | 26,048 (64.7) |
| Urine incontinence, n (%) | 11,131 (27.6) |
| Bowel incontinence, n (%) | 3,698 (9.2) |
| Sleep apnea, n (%) | 8,084 (20.1) |
| REM sleep behavior disorder, n (%) | 21,067 (52.3) |
| Hyposomnia/Insomnia, n (%) | 8,330 (20.7) |
Note: Patient characteristics at the baseline visits are summarized in either as number of samples (n) and/or percentage of the total samples (%), or mean and standard deviation (SD).
Abbreviation: REM, rapid eye movement.
3.2. Model performance
Table 3 shows the performance of our multi‐label MCC prediction model using XGBoost for each of the 18 identified MCC across all five‐fold cross‐validations. As we expected, the disease prevalence for each MCC was very low in follow‐up visits, resulting in a very imbalanced class distribution. The most prevalent MCC at follow‐ups was arthritis (15.7%), followed by hypercholesterolemia (8.7%), and urine incontinence (5.8%). On the other hand, the least prevalent MCC at follow‐ups was myocardial infarction (0.5%), followed by congestive heart failure (0.8%), and diabetes (1.2%).
TABLE 3.
Multi‐label prediction model performance across 18 multiple chronic conditions
| Multiple chronic conditions (MCC) | Prevalence | AUROC | Sensitivity | Specificity |
|---|---|---|---|---|
| Dementia | 3.7% (0.3%) | 0.949 (0.006) | 0.904 (0.022) | 0.870 (0.016) |
| Mild cognitive impairment | 3.8% (0.2%) | 0.818 (0.012) | 0.763 (0.018) | 0.727 (0.032) |
| Alzheimer's disease | 4.6% (0.4%) | 0.808 (0.013) | 0.710 (0.052) | 0.766 (0.040) |
| Cancer | 3.9% (0.3%) | 0.604 (0.004) | 0.591 (0.028) | 0.576 (0.012) |
| Diabetes | 1.2% (0.1%) | 0.681 (0.031) | 0.651 (0.072) | 0.624 (0.087) |
| Myocardial infarction | 0.5% (0.1%) | 0.737 (0.027) | 0.686 (0.059) | 0.683 (0.070) |
| Congestive heart failure | 0.8% (0.1%) | 0.806 (0.027) | 0.784 (0.069) | 0.667 (0.060) |
| Atrial fibrillation | 1.5% (0.1%) | 0.691 (0.013) | 0.625 (0.041) | 0.670 (0.045) |
| Angina | 1.6% (0.1%) | 0.680 (0.021) | 0.634 (0.033) | 0.639 (0.057) |
| Hypercholesterolemia | 8.7% (0.8%) | 0.597 (0.015) | 0.538 (0.059) | 0.619 (0.044) |
| Vitamin B12 deficiency | 2.3% (0.1%) | 0.589 (0.008) | 0.597 (0.037) | 0.550 (0.034) |
| Thyroid disease | 1.4% (0.2%) | 0.572 (0.020) | 0.463 (0.068) | 0.657 (0.098) |
| Arthritis | 15.7% (0.7%) | 0.655 (0.013) | 0.627 (0.077) | 0.605 (0.055) |
| Urine incontinence | 5.8% (0.3%) | 0.732 (0.010) | 0.646 (0.056) | 0.694 (0.061) |
| Bowel incontinence | 2.7% (0.2%) | 0.841 (0.019) | 0.725 (0.033) | 0.796 (0.021) |
| Sleep apnea | 2.8% (0.3%) | 0.678 (0.019) | 0.615 (0.060) | 0.654 (0.064) |
| REM sleep behavior disorder | 1.6% (0.2%) | 0.705 (0.017) | 0.694 (0.071) | 0.647 (0.069) |
| Hyposomnia/Insomnia | 5.4% (0.4%) | 0.641 (0.023) | 0.597 (0.031) | 0.621 (0.046) |
Note: The prevalence of the disease and the multi‐label XGBoost prediction model performance (in AUROC, sensitivity, and specificity) at predicting each of the 18 multiple chronic conditions (MCC) are summarized in the format of mean (standard deviation).
Abbreviations: AUROC, area under the receiver operating characteristic curve; REM, rapid eye movement.
By adjusting existing MCC patterns by syndemic factors at baseline, our prediction model achieved an overall arithmetic average AUROC score of 0.836 (SD = 0.069) during training and 0.710 (SD = 0.100) during testing across the five‐fold cross‐validation. Overall average specificity (0.670) during testing was slightly higher than sensitivity (0.658). Our model achieved higher performance in predicting dementia‐related chronic conditions than other chronic conditions. The model AUROC score in predicting dementia was the highest at 0.949, followed by bowel incontinence at 0.841, mild cognitive impairment at 0.818, and Alzheimer's disease at 0.808. Despite its lower prevalence, our prediction model was able to predict severe heart diseases moderately well: AUROC score of 0.806 for congestive heart failure and 0.737 for myocardial infarction. Most of the other chronic conditions had an average AUROC score between the ranges of 0.60 and 0.79, while the model performed worst at predicting hypercholesterolemia (0.597), vitamin B12 deficiency (0.589), and thyroid disease (0.572).
3.3. The effect of syndemic factors
Figure 1 shows the performance of our prediction model for each 18 MCC as we adjust existing MCC patterns (V1) by adding more syndemic factors to the model. Our model was able to achieve a relatively high AUROC score in predicting dementia (0.865) using the existing MCC patterns only (V1), as well as bowel incontinence (0.814) and congestive heart failure (0.754). The overall model performance was only moderate at predicting other MCC (average AUROC = 0.619) using existing MCC patterns alone (V1). When adding the subsequent syndemic factors to the MCC patterns‐only model (V1), the overall model performance showed little to moderate improvements, but significantly more when the model was adjusted with social‐related (+0.016 AUROC), physical health (+0.013 AUROC), dementia‐related (+0.010 AUROC), and mental health (+0.009 AUROC) features. These improvements in model performance also vary across MCC. The features related to the severity and type of dementia improved the AUROC for dementia‐related chronic conditions the most: dementia (+0.033), mild cognitive impairment (+0.089), and Alzheimer's disease (+0.065). Social‐related features improved model performance the most for metabolic diseases such as diabetes (+0.034) and thyroid disease (+0.021), as well as cancer (+0.033), sleep apnea (+0.054), and others. Vitamin B12 deficiency (+0.010), bowel incontinence (+0.010), and hyposomnia/insomnia (+0.024) experienced the most improvement when mental health features were added. Physical health features improved model performance the most for cardiovascular diseases such as myocardial infarction (+0.016) and congestive heart failure (+0.027), as well as hypercholesterolemia (+0.026) and REM sleep behavior disorder (+0.024).
FIGURE 1.

Model performance of 18 multiple chronic conditions (MCC) for each addition of syndemic factors. The x‐axis represents 18 MCC (labels of the model) and the y‐axis represents the model performance in AUROC (area under the receiver operating characteristic curve). The stacked bar charts represent the improvement in AUROC for each MCC when compared to the previous model version where the model versions (V1‐V9) are identified in different colors.
The adjustments of most of the syndemic factors improved the model performance on average, some syndemic factors harmed the model performance. For instance, dementia‐related features negatively impacted the model performance for cardio‐metabolic diseases such as congestive heart failure (−0.006), atrial fibrillation (−0.010), angina (−0.006), diabetes (−0.006), and thyroid disease (−0.010), and harmed REM sleep behavior disorder (−0.017) the most. Similarly, mental health features also negatively impacted model performance on cardio‐metabolic diseases such as myocardial infarction (−0.017), atrial fibrillation (−0.003), angina (−0.001), and thyroid disease (−0.007), as well as sleep apnea (−0.008). Nonetheless, these negative impacts were minimal. Altogether, syndemic factors generally improved the overall model performance.
4. DISCUSSION
The current literature has found that the interactions between MCC and ADRD are more complex than other chronic conditions due to the unique problems experienced by patients with ADRD, such as memory loss and impairment in cognitive behavior. Moreover, this interaction is further complicated by the synergistic interactions with other syndemic factors such as patient's lifestyles, social activities, polypharmacy, and others. This idea was first proposed by Dunn et al. 17 , however, current literature has yet to incorporate the unique interactions between MCC and ADRD with these syndemic factors. The goal of our study was to learn the MCC patterns adjusted by the theoretical syndemic framework proposed by Dunn et el. 17 for people with ADRD to early identify developing MCC. By understanding the MCC patterns, one can have a more concrete idea of which of the many possible diseases (918 MCC 20 ) to focus on prevention. Specifically, as a proof‐of‐concept, we tested the concept of learning MCC patterns and adjustment by syndemic factors (individualize the prediction model) by using a multi‐label prediction model with a gradient‐boosted tree algorithm, XGBoost to predict developing MCC. In our NACC example, we focused on 18 MCC (due to the data limitation in NACC) and built our prediction model based on their interactions with each other (including ADRD) and syndemic factors. We initially hypothesized that an AI/ML approach could capture the complex MCC patterns, and we showed that our prediction model performance can achieve moderately good performance (AUROC = 0.710; specificity = 0.670; sensitivity = 0.658) by adjusting existing MCC patterns with syndemic factors. In particular, our model performed best in predicting dementia‐related MCC, such as dementia (AUROC = 0.949), mild cognitive impairment (AUROC = 0.818), and Alzheimer's disease (AUROC = 0.808), as well as bowel incontinence (AUROC = 0.841) and congestive heart failure (AUROC = 0.806). On the other hand, our model performance was only modest for other chronic conditions. Nonetheless, our findings showed that, with multi‐label settings and imbalanced learning methods, AI/ML models can predict multiple diseases at once unlike risk factor approaches which can only target a handful of diseases; and identify specifically which diseases are more likely to happen next even with less prevalent diseases, unlike association analysis which usually rely on more prevalent diseases.
Our secondary objective was to explore the effect of each syndemic factor on the model performance. We found that the performance of the MCC prediction model with existing MCC patterns alone was considerably high for dementia (0.865) and bowel incontinence (0.814) using the 18 existing MCC patterns alone. In further investigation, we found that mild cognitive impairment diagnosis at baseline was a strong predictor for dementia while dementia diagnosis at baseline was a strong predictor for bowel incontinence. The strong predictive power for these diseases concurs with the current literature that there exist strong associations between diseases among the MCC patterns: 10%–20% of people with mild cognitive impairment will develop dementia in a year, 21 while bowel diseases are found to be highly associated with dementia. 22 Although we saw higher performance in predicting serious heart diseases, such as congestive heart failure (0.754) and myocardial infarction (0.729), the AUROC scores for other MCC using existing MCC patterns only were average. This lower performance in other MCC was also consistent when adjusting the existing MCC patterns with syndemic factors. One possible reason might be that the NACC centers focus on ADRD‐related data; hence, its collected data did not have sufficient signals for these diseases. In fact, no strong predictors were found in models predicting cancer, hypercholesterolemia, vitamin B12 deficiency, and thyroid disease.
In addition, we also found that the inclusion of four major categories of syndemic factors improved the overall model performance the most: social‐related, the severity and type of dementia, mental health, and physical health features. Predictably, dementia‐related features improved the model performance for dementia‐related MCC the most. These features partially contributed to the better model performance in predicting dementia and dementia‐related chronic conditions because NACC collects a very detailed description of the disease progression such as the patient's ability in speech, memory, orientation, and daily activities, as well as the level of tremors or rigidity in faces, hands, and feet. In fact, impairment in cognition, orientation, and memory, and difficulty in traveling at baseline were some of the strong predictors for dementia‐related MCC. Our model in predicting the incidence of dementia achieved a similar performance (AUROC = 0.949) with a previous study that also utilized NACC data 23 (AUROC = 0.92) and gradient‐boosted trees algorithms. However, the progression of dementia did not help but slightly harmed model performance for some of the cardio‐metabolic MCC. This is unexpected as previous literature has shown that cardio‐metabolic has a heightened associated risk of dementia incidence. 24 The same decline in model performance in predicting cardio‐metabolic diseases can also be seen when adding mental health features. This might suggest that the relationship between cardio‐metabolic and dementia might be unidirectional. Also, there might exist some noise in these detailed progressions of dementia and mental health status among patients who will develop cardio‐metabolic diseases.
To the authors’ best knowledge, no studies have utilized AI/ML approaches to unravel the complex MCC patterns adjusted by syndemic theory among people with ADRD. Nonetheless, this study had a few limitations: (1) This study used annual survey data collected from volunteering participants of ADRD research centers. As a result, only data collected at the annual visit were available to the authors. Any information that was not collected at the centers, such as the pollution index was not included in this study. (2) Most of the ADRD research centers are associated with academic medical centers that are in metropolitan areas. Hence, our findings might not generalize to patients living in rural communities. (3) The majority of the study cohort was White individuals. With the lack of racial diversity, the findings of our study might not generalize to other races. (4) This study used patients’ previous visit information to predict their new incidence of MCC in the next visits. However, the time to disease was not considered because the actual diagnosis date will not be very accurate. (5) The inclusion of syndemic factors in the secondary objective was chosen manually and cumulatively. In other words, any change in the order in which syndemic factors were first added might change the result of this study. However, we saw the greatest improvements from factors added last, so the change in result is not highly probable. (6) With the common understanding that increasing features results in increasing performance, the improvements we found with each addition of syndemic factors might be the result of overfitting the noise in the data. We have taken tremendous precautions in preventing overfitting in our model development, such as early stopping and restricting the complexity of the model (i.e., maximum depth, and sampling features). (7) Similar to other data‐driven AI/ML models, our model can only learn patterns that exist within the training dataset. As a result, while our model performed well for dementia‐related diseases (since we used one of the most powerful ADRD datasets), our model did not perform well for chronic conditions that lack strong predictors. (8) Interactions between prediction models were not included as the initial experiment showed signs of overfitting.
5. CONCLUSION
Despite the increased prevalence of ADRD and the severity of MCC in people with ADRD, current clinical practice guidelines for MCC have yet to address priority changes for people with ADRD. 25 , 26 , 27 Our study demonstrated the potential of utilizing AI/ML approaches to unravel the complex MCC patterns among people with ADRD adjusted by the theoretical syndemic framework proposed by Dunn et al. 17 Our work provided a useful clinical tool to help early identify MCC that has not happened (or developing) based on MCC and syndemic factors an ADRD patient has had, by providing a list of MCC that might happen in the next visit. Traditionally, through family history and patient complaints, healthcare professionals will conduct further diagnostic testing such as laboratory testing to confirm a diagnosis that they are suspecting. However, the unique problems experienced by people with ADRD make the diagnostic process difficult. As a result, healthcare professionals may miss potential diagnoses or need to perform additional diagnostic testing to confirm their diagnosis. By providing a data‐driven suggestion for developing MCC, this tool can help avoid missed diagnoses and also reduce the financial burdens from extensive diagnostic testing. Furthermore, we also elucidated important syndemic factors for the 18 identified MCC in this study. Although our MCC prediction model performed well in predicting dementia‐related conditions due to the power of NACC data, the model has room for improvement in predicting other chronic conditions. Our future directions include improving our current model by utilizing more complete health data (such as electronic health records); considering disease pathways, disease interactions, and variations in MCC patterns influenced by the progression of dementia; and providing transparency to the complex MCC patterns by visualizing the MCC patterns. This can be achieved using graph neural network (GNN) models as GNNs can handle intricate relationships while providing a clear view of the modeled relationships.
CONFLICT OF INTEREST STATEMENT
The authors declare that they have no conflict of interest. Author disclosures are available in the supporting information.
CONSENT STATEMENT
Written informed consent is obtained from all NACC participants and co‐participants.
Supporting information
Supporting information
Supporting information
ACKNOWLEDGMENTS
Many thanks to NACC for providing data. The NACC database is funded by NIA/NIH Grant U24 AG072122. NACC data are contributed by the NIA‐funded ADRCs: P30 AG062429 (PI James Brewer, MD, PhD), P30 AG066468 (PI Oscar Lopez, MD), P30 AG062421 (PI Bradley Hyman, MD, PhD), P30 AG066509 (PI Thomas Grabowski, MD), P30 AG066514 (PI Mary Sano, PhD), P30 AG066530 (PI Helena Chui, MD), P30 AG066507 (PI Marilyn Albert, PhD), P30 AG066444 (PI John Morris, MD), P30 AG066518 (PI Jeffrey Kaye, MD), P30 AG066512 (PI Thomas Wisniewski, MD), P30 AG066462 (PI Scott Small, MD), P30 AG072979 (PI David Wolk, MD), P30 AG072972 (PI Charles DeCarli, MD), P30 AG072976 (PI Andrew Saykin, PsyD), P30 AG072975 (PI David Bennett, MD), P30 AG072978 (PI Neil Kowall, MD), P30 AG072977 (PI Robert Vassar, PhD), P30 AG066519 (PI Frank LaFerla, PhD), P30 AG062677 (PI Ronald Petersen, MD, PhD), P30 AG079280 (PI Eric Reiman, MD), P30 AG062422 (PI Gil Rabinovici, MD), P30 AG066511 (PI Allan Levey, MD, PhD), P30 AG072946 (PI Linda Van Eldik, PhD), P30 AG062715 (PI Sanjay Asthana, MD, FRCP), P30 AG072973 (PI Russell Swerdlow, MD), P30 AG066506 (PI Todd Golde, MD, PhD), P30 AG066508 (PI Stephen Strittmatter, MD, PhD), P30 AG066515 (PI Victor Henderson, MD, MS), P30 AG072947 (PI Suzanne Craft, PhD), P30 AG072931 (PI Henry Paulson, MD, PhD), P30 AG066546 (PI Sudha Seshadri, MD), P20 AG068024 (PI Erik Roberson, MD, PhD), P20 AG068053 (PI Justin Miller, PhD), P20 AG068077 (PI Gary Rosenberg, MD), P20 AG068082 (PI Angela Jefferson, PhD), P30 AG072958 (PI Heather Whitson, MD), P30 AG072959 (PI James Leverenz, MD). The authors received no specific grants or any other financial support for the research and publication of this article.
Yew PY, Devera R, Liang Y, et al. Unraveling the multiple chronic conditions patterns among people with Alzheimer's disease and related dementia: A machine learning approach to incorporate synergistic interactions. Alzheimer's Dement. 2024;20:4818–4827. 10.1002/alz.13923
REFERENCES
- 1. Alzheimer's Association . 2023 Alzheimer's disease facts and figures. Alzheimers Dement 2023;19:1598‐6195. doi: 10.1002/alz.13016 [DOI] [PubMed] [Google Scholar]
- 2. Beerten SG, Helsen A, De Lepeleire J, Waldorff FB, Vaes B. Trends in prevalence and incidence of registered dementia and trends in multimorbidity among patients with dementia in general practice in Flanders, Belgium, 2000‐2021: a registry‐based, retrospective, longitudinal cohort study. BMJ Open. 2022;12:e063891. doi: 10.1136/bmjopen-2022-063891 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Browne J, Edwards DA, Rhodes KM, Brimicombe DJ, Payne RA. Association of comorbidity and health service usage among patients with dementia in the UK: a population‐based study. BMJ Open. 2017;7:e012546. doi: 10.1136/bmjopen-2016-012546 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Tonelli M, Wiebe N, Joanette Y, et al. Age, multimorbidity and dementia with health care costs in older people in Alberta: a population‐based retrospective cohort study. Can Med Assoc Open Access J. 2022;10:E577‐E588. doi: 10.9778/cmajo.20210035 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. MacNeil‐Vroomen JL, Thompson M, Leo‐Summers L, Marottoli RA, Tai‐Seale M, Allore HG. Health‐care use and cost for multimorbid persons with dementia in the National Health and Aging Trends Study. Alzheimers Dement J. 2020;16:1224‐1233. doi: 10.1002/alz.12094 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Salber PR, Selecky CE, Soenksen D, Wilson T. Impact of dementia on costs of modifiable comorbid conditions. Am J Manag Care. 2018;24(11):e344‐e351. [PubMed] [Google Scholar]
- 7. van den Akker M, Buntinx F, Roos S, Knottnerus JA. Problems in determining occurrence rates of multimorbidity. J Clin Epidemiol. 2001;54:675‐679. doi: 10.1016/S0895-4356(00)00358-9 [DOI] [PubMed] [Google Scholar]
- 8. Marengoni A, Roso‐Llorach A, Vetrano DL, et al. Patterns of multimorbidity in a population‐based cohort of older people: sociodemographic, lifestyle, clinical, and functional differences. J Gerontol A Biol Sci Med Sci. 2020;75:798‐805. doi: 10.1093/gerona/glz137 [DOI] [PubMed] [Google Scholar]
- 9. Guisado‐Clavero M, Roso‐Llorach A, López‐Jimenez T, et al. Multimorbidity patterns in the elderly: a prospective cohort study with cluster analysis. BMC Geriatr. 2018;18:16. doi: 10.1186/s12877-018-0705-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Held FP, Blyth F, Gnjidic D, et al. Association rules analysis of comorbidity and multimorbidity: the concord health and aging in men project. J Gerontol A Biol Sci Med Sci. 2016;71:625‐631. doi: 10.1093/gerona/glv181 [DOI] [PubMed] [Google Scholar]
- 11. Prados‐Torres A, Poblador‐Plou B, Calderón‐Larrañaga A, et al. Multimorbidity patterns in primary care: interactions among chronic diseases using factor analysis. PloS One. 2012;7:e32190. doi: 10.1371/journal.pone.0032190 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Lee Y, Kim H, Jeong H, Noh Y. Patterns of multimorbidity in adults: an association rules analysis using the Korea health panel. Int J Environ Res Public Health. 2020;17:2618. doi: 10.3390/ijerph17082618 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Dong G, Zhang Z‐C, Feng J, Zhao X‐M. MorbidGCN: prediction of multimorbidity with a graph convolutional network based on integration of population phenotypes and disease network. Brief Bioinform. 2022;23:bbac255. doi: 10.1093/bib/bbac255 [DOI] [PubMed] [Google Scholar]
- 14. PolessaPaula D, Barbosa Aguiar O, Pruner Marques L, et al. Comparing machine learning algorithms for multimorbidity prediction: an example from the Elsa‐Brasil study. PloS One. 2022;17:e0275619. doi: 10.1371/journal.pone.0275619 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Alsaleh MM, Allery F, Choi JW, et al. Prediction of disease comorbidity using explainable artificial intelligence and machine learning techniques: a systematic review. Int J Med Inf. 2023;175:105088. doi: 10.1016/j.ijmedinf.2023.105088 [DOI] [PubMed] [Google Scholar]
- 16. Calderón‐Larrañaga A, Vetrano DL, Ferrucci L, et al. Multimorbidity and functional impairment–bidirectional interplay, synergistic effects and common pathways. J Intern Med. 2019;285:255‐271. doi: 10.1111/joim.12843 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Dunn R, Clayton E, Wolverson E, Hilton A. Conceptualising comorbidity and multimorbidity in dementia: a scoping review and syndemic framework. J Multimorb Comorbidity. 2022;12:26335565221128432. doi: 10.1177/26335565221128432 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Manderson L. Introduction to syndemics: a critical systems approach to public and community health by Merrill singer. Med Anthropol Q. 2012;26:643‐645. doi: 10.2307/41811623 [DOI] [Google Scholar]
- 19. Beekly DL, Ramos EM, Lee WW, et al. The National Alzheimer's Coordinating Center (NACC) database: the Uniform Data Set. Alzheimer Dis Assoc Disord. 2007;21:249‐258. doi: 10.1097/WAD.0b013e318142774e [DOI] [PubMed] [Google Scholar]
- 20. Calderón‐Larrañaga A, Vetrano DL, Onder G, et al. Assessing and measuring chronic multimorbidity in the older population: a proposal for its operationalization. J Gerontol Ser A. 2017;72:1417‐1423. doi: 10.1093/gerona/glw233 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Etgen T, Sander D, Bickel H, Förstl H. Mild cognitive impairment and dementia. Dtsch Ärztebl Int. 2011;108:743‐750. doi: 10.3238/arztebl.2011.0743 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Santiago JA, Potashkin JA. The impact of disease comorbidities in Alzheimer's disease. Front Aging Neurosci. 2021;13. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. James C, Ranson JM, Everson R, Llewellyn DJ. Performance of machine learning algorithms for predicting progression to dementia in memory clinic patients. JAMA Netw Open. 2021;4:e2136553. doi: 10.1001/jamanetworkopen.2021.36553 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Calvin CM, Conroy MC, Moore SF, Kuźma E, Littlejohns TJ. Association of multimorbidity, disease clusters, and modification by genetic factors with risk of dementia. JAMA Netw Open. 2022;5:e2232124. doi: 10.1001/jamanetworkopen.2022.32124 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Tsunawaki S, Abe M, DeJonckheere M, et al. Primary care physicians’ perspectives and challenges on managing multimorbidity for patients with dementia: a Japan–Michigan qualitative comparative study. BMC Prim Care. 2023;24:132. doi: 10.1186/s12875-023-02088-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. National Institute for Health and Care Excellence . Multimorbidity: clinical assessment and management. 2016.
- 27. Guiding principles for the care of older adults with multimorbidity: an approach for clinicians . Guiding principles for the care of older adults with multimorbidity: an approach for clinicians: american geriatrics society expert panel on the care of older adults with multimorbidity. J Am Geriatr Soc 2012;60:E1‐25. doi: 10.1111/j.1532-5415.2012.04188.x [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supporting information
Supporting information
