Abstract
The development of chronic disease is a long-term process involving multiple endpoints. Although multi-state Cox models can estimate state-specific survival risks over time, they are not well suited for comparing the effectiveness of treatment regimes. A discrete-time split-state framework has been proposed, which divides disease states into substates by conditioning on past history. As this framework is both “memoryless” and “memorable,” the transition rates can be synthesized into summary measures, multimorbidity-adjusted life year (MALY) and substate-specific life year (SSLY). Building on this framework, we propose to investigate the causal effects of static and dynamic treatment regimes over the disease course under the assumptions of constant baseline confounders and instantaneous effects of interventions on transition rates. We identify the optimal treatment regime as the one that maximizes MALY and use SSLY to elucidate the mechanisms of how treatments influence disease progression. In the application, we identified the optimal weight targets in the ARIC study by modeling the disease course in healthy, at-metabolic-risk, coronary heart disease, heart failure, and mortality states. The estimated MALY was 1.80 years higher (95% CI: 0.62, 2.78) under regime “Normal weight initially and change to overweight if age >65 y” compared to regime “Normal weight across all states.” SSLY decomposition indicates that this gain arises from increased life year in all substates except the healthy state. In summary, our method provides a framework to evaluate the health benefits of treatment regimes over the disease course and has the potential to improve the precision prevention of chronic diseases.
Keywords: Multi-state modeling, static and dynamic treatment regime, chronic disease, precision prevention, life expectancy
1. Introduction
Comparative effectiveness of interventions has been widely studied in medical research. The interventions can be categorized into static regimes, where all treatments during follow-up are predetermined at baseline, and dynamic regimes, where treatment at each time point depends on an individual’s response to prior treatments. The health effects of treatment regimes are commonly measured by estimating the risk of disease development using the parametric g-formula or the marginal structure model (MSM).1,2 These models first estimate conditional survival risks within each stratum of confounders, and then derive the marginal survival risk for the overall population through standardization (g-formula) or inverse probability weighting (MSM).
Rather than being defined by a single disease endpoint, the development of chronic disease is a long-term process involving multiple endpoints. Taking heart disease as an example, it begins with the development of hypertension, hypercholesterolemia, or hyperglycemia (i.e. an at-metabolic-risk state), followed by the onset of coronary heart disease (CHD), which may progress to heart failure and eventually end in mortality.3,4 The disease states in this multi-state process are interconnected, and treatment at each state can influence the entire course of disease progression. However, existing methods that assess the effects of treatment regimes based on a single endpoint fail to capture the complexity of the disease course and may therefore misidentify the optimal treatment regime.
Multi-state modeling has been used to describe the course of chronic diseases. For example, multi-state Cox models use multiple Cox models to characterize transitions between disease states. 5 However, the estimated transition parameters are state- and time-specific, making it difficult to use them for evaluating the effects of treatment regimes. To address this, a discrete-time split-state framework for multi-state modeling has been proposed, 6 which divides disease states into substates by conditioning on past history. As the substates contain information on past history, the time-specific transition parameters can be synthesized into summary measures, multimorbidity-adjusted life year (MALY), and substate-specific life year (SSLY). 7 In this article, we extend this into a multi-state causal framework to evaluate the health benefits of treatment regimes over the disease course. Our method identifies the optimal treatment regime that generates the most benefits, as indicated by MALY, and elucidates the mechanisms through which treatment regimes affect disease progression, as indicated by SSLY. Our method is developed under the assumptions of constant baseline confounders and instantaneous effects of interventions on transition rates.
2. Methods
2.1. An introduction to the discrete-time multi-state framework
We applied a discrete-time multi-state framework to estimate the health benefits of treatment regimes over the disease course.6,7 While this framework can be applied to chronic disease with any number of states, we illustrate it here using five states S0-S4 (Figure 1). In brief, the framework divides each disease state into substates by conditioning on past history and estimates transition rates between substates using cause-specific Cox models. For example, S2 can be divided into substates S and S , where S represents transitions from S0 and S represents transitions from S0 and S1. The framework has two unique features. First, it is “memoryless.” By conditioning on past states, the newly created substates are independent of past history and exhibit the Markov property, regardless of whether the original process is Markovian. Thus, regardless of whether the original process is Markov, the Aalen–Johansen estimator can be used to estimate state occupation probabilities (or survival risks) based on transition rates. Second, it is “memorable,” as the substates contain information of past states and indicate multimorbidity status. Thus, the state- and time-specific survival risks can be synthesized into summary measures of MALY and SSLY. 7 Specifically, MALY is used to identify the optimal treatment regime, and SSLY is used to elucidate the mechanisms.
Figure 1.

Construction of the process of disease substates by conditioning on past states. Note: This figure is cited from our previous publication. 6
Suppose we have a population of N participants. Let denote disease states, and consider a disease process that includes states , and . By conditioning on past history, the five states , , , , and can be divided into nine disease substates, , , , , , , , and . Let denote the exposures in states , and at time (note that there is no exposure in the final state ). After dividing states into substates, the corresponding exposures are denoted as , , , , , , and . Let represent baseline confounders. Our method is developed under two assumptions: (a) are constant over the follow-up period; (b) The effects of exposure on the transition rates to subsequent states are instantaneous.
2.2. Estimating conditional transition rates
As described in detail in a previous publication, cause-specific Cox models are applied to examine the associations between exposure and transition rates. In brief, the data are divided into subsets by disease state: subset 1 starts from and ends with the incidence of S1, S2, S3, S4, or the end of follow-up, whichever comes first; subset 2 starts from and ends with the incidence of S2, S3, S4 or the end of follow up; subset 3 starts from and ends with the incidence of S3, S4 or the end of follow up; and subset 4 starts from and ends with the incidence of S4 or the end of follow up. Each subset is restructured from wide to long format by age. Cause-specific Cox models (Models 1–4) are fitted to each subset to estimate conditional transition rates. Model 1 in subset 1 starting from S0 state:
where m ranges from 1 to 4, indicating the development of S1, S2, S3, or S4 states from S0. is a function of age. is the baseline hazard of developing state m. is time to event in each time interval and ranges from [0, 1] (we use age as the time scale, and thus the maximum length of each time interval is 1). Model 2 in subset 2 starting from S1 state:
where m can be 2–4, indicating the development of S2, S3, or S4 states. Model 3 in subset 3 starting from S2 state:
where m can be 3 and 4, indicating the development of S3 or S4 states from S2. is a joint indicator of the past states S0 and S1, and has two categories, “0” and “0_1.” When , corresponds to , and when , corresponds to . Model 4 in subset 4 starting from S3 state:
where m indicates the development of S4 state from S3. J is a joint indicator of the past states S0, S1 and S2, and has 4 categories, “0,” “0_1,” “0_2,” “0_1_2.” When , “0_1,” “0_2,” “0_1_2,” corresponds to , , , and , respectively.
The discrete-time hazard can be approximated by the age-specific risk at the end of each time interval, 8 which can be estimated using the cumulative incidence function (CIF) derived from cause-specific Cox models.9,10 Specifically, Model 1 estimates , , , and , the conditional transition rates from S0 to states S1–S4, respectively. Model 2 estimates , , and , the conditional transition rates from S1 to states S2–S4, respectively. Model 3 estimates and ), the conditional transition rates from S to S3 and S4, respectively, and and , the conditional transition rates from S to S3 and S4, respectively. Model 4 estimates , , , and , the conditional transition rates to S4 from S , S , S , and S , respectively. These conditional transition rates are used to construct the conditional transition rate matrix as shown below, where . Each entry in the matrix represents the average transition rate from the substate in the row to the substate in the column, at time , in a population.
2.3. Static treatment regime and conditional SSLY and MALY
Let denote the treatment regime, where .
For a static regime, the hypothetical intervention is constant over time. Let denote the conditional state occupation probabilities at time . For a five-state transition, , where each component represents the proportion of participants in at time in substates , , , , , , , and , respectively. By using the Aalen–Johansen estimator, 11 the conditional state occupation probability can be estimated as , where is an identity matrix, and is the start of follow-up.
The estimation of SSLY and MALY based on has been described in detail in a previous publication. 7 In brief, SSLY can be estimated as the sum of the probabilities of being in each substate (except S4) from the starting age until death. Let denote a matrix representing life years spent in substates, where , corresponding to the substates , , , , , , and , respectively. The SSLY in substate is estimated as , where is the start of follow-up and is the end.
As these substates indicate multimorbidity status, a disability weight (DW) is assigned to each substate to account for the severity of multimorbidity.12,13 DW quantify the severity of a disease, ranging from 0 (full health) to 1 (death), 12 and the DW for each disease can be obtained from the Global Burden of Disease Study. 14 For disease substates with multimorbidity, DW values can be calculated using a multiplicative approach.12,13 For example, for the substate S , the DW is calculated as . The weighted SSLY in substate is then computed as .
As all substates are mutually exclusive, the conditional MALY can be calculated as the weighted average of SSLY across all substates. MALY takes the multimorbidity of each substate into consideration and estimates the adjusted life years in full health.
2.4. Static treatment regime and counterfactual SSLY and MALY
To compare health benefits of static treatment regimes, we hypothetically assign each regime to all individuals in the target population. The regime yielding the highest mean MALY is the optimal treatment. Theoretically, the MALY used for comparison is the counterfactual MALY, defined as the average MALY in a population had everyone received that treatment regime. This is essentially to standardize the conditional MALY according to the distribution of confounders L in a population. The counterfactual MALY is estimated as
| (1) |
However, in practice, it is difficult to estimate when contains continuous variables. In fact, (1) is equivalent to an individual-level estimation, with the illustration and proof shown below.
For person , we only observe the person being in one substate with exposure at time , and the estimated conditional transition rate is , which is estimated from Models 1–4. To construct the counterfactual transition rate matrix under the hypothetical intervention, four assumptions of consistency and exchangeability are needed.
Assumption 1 (Consistency assumption for treatment) —
The observed transition rate for every treated participant equals the hypothetical transition rate had the person received treatment. For example, for participant in observed with , if the person received intervention , the counterfactual transition rate would be = .
Assumption 2 (Exchangeability assumption for treatment) —
The untreated would have the same transition rate as the treated had the person received treatment. For example, for participant in observed with and participant in observed with , and , the counterfactual transition rate would be = , had person received intervention .
Assumption 3 (Consistency assumption for disease substate) —
The observed transition rate for every participant experienced a disease substate equals the hypothetical transition rate had the person experienced the disease substate. For example, for participant in , if the person received intervention , the counterfactual transition rate would be = .
Assumption 4 (Exchangeability assumption for disease substate) —
Participants would have experienced the same transition rate as those who have an alternative disease history, had they experienced the alternative disease history. For example, if person experiences and person does not, and , the counterfactual transition rate would be = , had the person experienced .
Using these assumptions, the conditional transition rate estimated from Models 1–4 is used to construct the counterfactual transition rate matrix .
For person , the counterfactual state occupation probability is estimated as
For person , the counterfactual SSLY in substate is estimated as
For person , the counterfactual MALY is estimated as
The counterfactual SSLY in substate in the whole population is estimated as
The counterfactual MALY in the whole population is estimated as
| (2) |
In fact, the counterfactual MALY and SSLY estimated by standardizing conditional measures are equivalent to those obtained by averaging the corresponding measures over the total population. Take MALY as an example, (1) and (2) are equivalent, with proof shown below.
Proof.
We divide L into stratum , with participants in each stratum.
2.5. Dynamic treatment regime and individual-level estimation of counterfactual MALY
The dynamic treatment regime can vary over time by person, depending on the decision criteria, such as the occurrence of side effects of treatment. Let denote the side effects at time , which may be affected by and baseline confounders . We assume that the intervention has only instantaneous effect on (i.e. the effect is only at time and does not persist beyond that time point).
Step 1. For person starting from , assign treatment based on their initial state occupation probability ).
Step 2. Use the assigned treatment to fit Models 1–4 and construct the counterfactual transition rate matrix .
Step 3. Estimate the state occupation probabilities at time as
Step 4. Estimate the adverse effect at , , as by fitting models regressing on and .
Step 5. Based on the updated state occupation probability, , and side effect, , apply the decision criteria to define the next treatment , . Then refit Models 1–4 to construct and estimate . Then estimate and determine treatment — . Repeat this process iteratively at each time t until the end of follow-up .
For participant , the counterfactual SSLY in substate under the dynamic treatment regime is estimated as
For participant , the counterfactual MALY is estimated as
The counterfactual SSLY in the whole population is estimated as
The counterfactual MALY in the whole population is estimated as
2.6. Compare counterfactual MALY across treatment regimes
The difference in MALY between two treatment regimes can be estimated as
The 95% confidence interval (CI) can be estimated using bootstrap. We randomly draw samples with replacement to create a new sample that are the same size as the original population. The counterfactual MALY for each treatment regime and the difference in MALY between two regimes can be estimated from each bootstrap sample. Statistically, the resulting parameter estimates across all bootstrap samples represent the empirical distribution and can be summarized with median values and a 95% CI defined by 2.5th to 97.5th percentiles of the empirical distribution.
3. Application
We applied our method to identify optimal body mass index (BMI) range across the course of heart disease in the Atherosclerosis Risk in Communities (ARIC) study. The AHA has recommended maintaining a kg/m across life course for prevention of CVD. 15 However, the associations of BMI with CVD-related endpoints vary by disease state. Studies have shown that overweight (BMI 25–30 kg/m ) and obesity ( kg/m ) are associated with higher risks of cardiometabolic disease compared to normal weight.16,17 However, among older people or participants with CVD, elevated BMI is associated with lower risk of mortality compared to those of normal weight.18–21 Thus, it is imperative to identify disease-state specific BMI range that generates the most health benefits across the course of heart disease.
3.1. Study population
We used data from the ARIC study.22,23 The enrolled participants underwent a phone interview and clinic visit at baseline in 1987 and were followed up by telephone calls and re-examinations until 2019. Participants were contacted periodically by phone and interviewed about interim hospital admissions, cardiovascular outpatient diagnoses, and deaths. We obtained ARIC data through Biologic Specimen and Data Repository Information Coordinating Center (BioLINCC), an open repository managed by the National Heart, Lung, and Blood Institute (NHLBI). The study was deemed exempt by the Institutional Review Board (IRB) of the University of North Carolina (UNC) at Chapel Hill.
3.2. Assessment of exposure, outcome, and covariates
Exposure. BMI was measured on weight scales by medical professionals. Normal weight is defined as kg/m , overweight is defined as BMI between 25–30 kg/m , and obesity is defined as kg/m .
Variables used to define states of heart disease. Low density lipoprotein cholesterol (LDL-c), high density lipoprotein cholesterol (HDL-c), triglycerides, and blood glucose were assessed in laboratories at each visit. Blood pressure was measured using electronic monitor by medical professionals at each visit. Participants who reported CVD-related events were asked to provide medical records that were reviewed by physicians. 23 Deaths were identified from systematic searches of vital records in states and the National Death Index, supplemented by reports from family members and postal authorities. 24
Covariates. Age, sex, race, education level, smoking status, and family history of heart disease were measured using questionnaires at baseline.
3.3. Definition of the course of heart disease
We modeled the course of heart disease in five states: Healthy, at metabolic risk, coronary heart disease (CHD), heart failure, and mortality. The at-metabolic-risk state was defined as the development of hypertension, hyperlipidemia, or diabetes. Hypertension was defined as blood pressure 140/90 mmHg or a history of hypertension or use of blood pressure medications. 25 Hyperlipidemia included primary hypertriglyceridemia ( 175 mg/dL) or primary hypercholesterolemia (LDL-c 160–189 mg/dL, and/or non-HDL-c 190–219 mg/dL). 26 Diabetes was defined as fasting glucose 7.7 mmol/L. 27
3.4. Regimes of BMI range
We compared four BMI regimes as below. Regimes 1–3 were static regimes which predefined BMI range before intervention. Regime 4 was a dynamic regime which defined BMI range by a person’s characteristics that can change over time.
Regime 1: Normal weight ( kg/m ) in all disease states.
Regime 2: Normal weight ( kg/m ) in healthy, at-metabolic-risk, and CHD states, and overweight (BMI 25–30 kg/m ) if the person develops heart failure.
Regime 3: Normal weight ( kg/m ) in healthy and at-metabolic-risk states, and overweight (BMI 25–30 kg/m ) if the person develops CHD or heart failure.
Regime 4: Normal weight ( kg/m ) at the start of intervention, and change to overweight (BMI 25–30 kg/m ) if age >65y.
By coding normal weight, overweight, and obesity as 0, 1, and 2, respectively, the four regimes can be rewritten as follows.
Regime 1:
Regime 2:
Regime 3:
Regime 4: For , the intervention was 0 at each time interval in each of the substate if age 65 years, and changed to 1 if age >65 years.
3.5. Statistical analysis
We applied our method to examine the effects of BMI regimes on the course of heart disease. In brief, the transition rates between substates were estimated using cause-specific Cox models, adjusting for age (continuous), sex (women, men), race (White, Black), education (less than high school, high school/vocational school, more than high school [college or professional school]), family history of heart disease (yes, no), and baseline smoking status (never, past, current). For each treatment regime, the age-specific transition rates and state occupation probabilities were estimated, which were used to synthesize MALY and SSLY. Specifically, for the estimation of MALY, we used DWs derived from the Global Burden of Disease Study 2019. 14 The DWs were 0.049, 0.072, and 0.074 for at-metabolic-risk, CHD, and heart failure, respectively. For states with multimorbidity, a multiplicative approach was used to assign disability weights. For example, the DW for CHD transitioning from healthy and at metabolic risk states was calculated as: . The most beneficial BMI regime was identified as the one associated with the highest counterfactual MALY. More details of the statistical analysis of the application study can be found in the supplemental materials.
3.6. Results
Figure 2 shows the associations of BMI with the rates of transitioning to subsequent states. Starting from the healthy state, overweight and obesity were associated with higher rate ratios of transitioning to at-metabolic-risk states (1.21 [95% CI: 1.10, 1.32] for overweight; 1.48 [95% CI: 1.32, 1.65] for obesity). The positive associations of overweight with transitions to the next states were attenuated to null starting from the at-metabolic-risk or CHD state. The association became inverse starting from the heart failure state, where overweight was associated with a lower rate ratio of transitioning to death (0.82 [95% CI: 0.73, 0.92]).
Figure 2.

Hazard ratios (95% CIs) for state transitions by obesity status across disease substates in the Atherosclerosis Risk in Communities (ARIC) study, 1987–2019 ( ). Normal weight: body mass index (BMI) < 25 kg/m ; Overweight: 25 kg/m BMI < 30 kg/m ; Obesity: BMI 30 kg/m . Cause-specific Cox models adjusted for age (continuous), race (White, Black), sex (women, men), education (less than high school, high school/vocational school, more than high school [college or professional school]), family history of cardiovascular disease (yes, no), and baseline smoking status (never, past, current).
The cardiovascular benefits of the four BMI regimes are presented in Table 1. For example, the counterfactual MALY was 35.08 years (95% CI: 22.63, 47.11) under the hypothetical Regime 1, meaning that the mean healthy life expectancy accounting for multimorbidity would be 35.08 years had everyone in the population maintained a normal weight across all disease states. Compared to Regime 1, the MALY was 0.38 years (95% CI: 0.07, 0.90) higher under Regime 3, and 1.80 years higher (95% CI: 0.62, 2.78) under Regime 4. These results suggest that starting with a normal weight and changing to overweight after age 65y (Regime 4) yields the most cardiovascular benefits.
Table 1.
Difference in multimorbidity-adjusted life years (MALY) across body mass index (BMI) regimes in the Atherosclerosis Risk in Communities (ARIC) study, 1987–2019 ( ).
| MALY,life year (95% CI) | Difference in MALY,life year (95% CI) | |
|---|---|---|
| Regime 1 | 35.08 (22.63, 47.11) | Reference |
| Regime 2 | 35.52 (23.02, 47.39) | 0.37 (0.10, 0.86) |
| Regime 3 | 35.54 (23.04, 47.40) | 0.38 (0.07, 0.90) |
| Regime 4 | 36.95 (24.28, 48.56) | 1.80 (0.62, 2.78) |
To elucidate the mechanisms by which the BMI regimes affect disease progression, we estimated SSLY under each treatment regime. Compared to Regime 1, the SSLY of the optimal regime (Regime 4) was higher in all substates except the healthy state, with the largest gain observed in the at-metabolic-risk state (Table 2). This finding suggests that maintaining a BMI between 25 and 30 kg/m may elongates life expectancy among older adults, even for those who start from the at-metabolic-risk state.
Table 2.
Estimated substate-specific life years (SSLY) by body mass index (BMI) regimes in the Atherosclerosis Risk in Communities (ARIC) study, 1987–2019 ( ).
| Disease substates | Life year (95% CI) | |||
|---|---|---|---|---|
| Regime 1 | Regime 2 | Regime 3 | Regime 4 | |
| Healthy (S0) | 27.73 (17.60, 35.48) | 27.73 (17.60, 35.48) | 27.73 (17.60, 35.48) | 27.67 (17.69, 35.38) |
| At metabolic risk (S1) | 28.22 (3.20, 44.27) | 28.22 (3.20, 44.27) | 28.22 (3.20, 44.27) | 29.16 (3.69, 45.44) |
| CHD without metabolic risk factors (S ) | 0.33 (0.05, 1.83) | 0.33 (0.05, 1.84) | 0.33 (0.05, 1.86) | 0.36 (0.05, 1.97) |
| CHD with metabolic risk factors (S ) | 0.98 (0.13, 2.61) | 0.98 (0.13, 2.63) | 0.99 (0.13, 2.67) | 1.15 (0.17, 3.04) |
| Heart failure with no metabolic risk factors or CHD (S ) | 1.40 (0.47, 3.33) | 1.65 (0.54, 3.84) | 1.65 (0.54, 3.84) | 1.82 (0.64, 4.22) |
| Heart failure only with metabolic risk factors (S ) | 1.94 (0.29, 3.52) | 2.29 (0.35, 4.16) | 2.30 (0.35, 4.16) | 2.65 (0.45, 4.67) |
| Heart failure only with CHD (S ) | 0.08 (0.01, 0.48) | 0.09 (0.01, 0.56) | 0.09 (0.01, 0.56) | 0.10 (0.01, 0.64) |
| Heart failure with metabolic risk factors and CHD (S ) | 0.20 (0.02, 0.65) | 0.23 (0.02, 0.79) | 0.23 (0.02, 0.79) | 0.27 (0.03, 0.93) |
3.7. Discussion
The proposed method is developed based on a discrete-time split-state framework for multi-state modeling. 6 In brief, this framework splits disease states into substates by conditioning on past history. The newly created substates have two unique features. First, they are “memoryless,” meaning that they exhibit the Markov property, regardless of whether the original process is Markovian. Second, they are “memorable, as being in a substate indicates multimorbidity status. The features of being “memoryless” and “memorable” enable the synthesis of the substate- and time-specific transition rates into summary measures (i.e. SSLY and MALY) to characterize the entire disease course. In this article, building on this framework, we propose a multi-state causal approach to identify the optimal treatment regime that generates the most benefits, as indicated by the maximum MALY. Our method also provides insights into the underlying mechanisms of treatment regime on disease progression, as indicated by SSLY.
From a methodology perspective, our method is a causal inference approach that evaluates the counterfactual effect of a treatment regime on the disease course, assuming all participants receive the treatment. Existing methods that examine the effects of treatment regimes on disease risk include the parametric g-formula and the MSM.1,2 Both methods control for time-varying confounding by conditioning on past confounders under treatment and estimate conditional survival probabilities using the Kaplan–Meier estimator. They derive counterfactual effects in the entire population through standardization (g-formula) or inverse probability weighting (MSM). However, these methods focus on the risk of a single endpoint and ignore the entire disease course. This limitation may result in the misidentification of optimal treatment regimes, especially when treatments exert long-term effects across multiple disease states. Our method incorporates the entire disease course and provides a more comprehensive evaluation of treatment effectiveness. It can be considered as an extension of the g-formula to the multi-state setting. In addition to causal inference methods, Sequential Multiple Assignment Randomized Trials (SMARTs) have been used to evaluate the causal effect of treatment regimes on an outcome. 28 In SMART designs, participants are randomized at multiple decision points to construct and compare dynamic treatment regimes. However, SMART trials are often impractical for chronic diseases, given the extreme long follow-up needed for multiple disease states to develop. By estimating the counterfactual effect under hypothetical interventions, our method emulates a sequential trial that would otherwise be infeasible to implement in real-world settings.
From a public health perspective, our method has the potential to advance the precision prevention of chronic diseases by identifying the optimal treatment for each person at each (sub)state that maximizes the precise benefits over the disease course. First, our method enables the identification of (sub)state-specific optimal treatments. Traditional chronic disease prevention is generally divided into primary prevention—aimed at reducing disease incidence—and secondary prevention—focused on preventing complications once a disease has developed. 29 However, due to the complex and heterogeneous progression of chronic diseases, prevention strategies may need to be further tailored to specific disease states. Our method identifies the optimal substate-specific strategies for precision prevention. Second, our method allows for dynamic, individualized interventions. Conventional public health prevention strategies are typically applied to the general population, with limited consideration for individual variability. In this article, our dynamic treatment regime accounts for individual heterogeneity, enabling treatment decisions to be made based on a person’s current disease state and response over time. This facilitates the development of personalized prevention strategies that are most beneficial across the entire disease course. 2 Third, because MALY incorporates multimorbidity and quality of life over the disease course, our method provides a more precise estimation of prevention benefits. This is particularly valuable for cost-effectiveness analyses and public health decision-making. Moreover, the estimated SSLY offer insights into the mechanisms through which treatment impacts disease progression. As suggested in the application section, starting with a normal weight and changing to overweight after age 65y (Regime 4) achieved significantly higher MALY compared to other regimes. This was due to prolonging time spent in all disease substates, particularly the at-metabolic-risk substate.
We acknowledge that our method needs the assumption that the effects of treatments on transition rates are instantaneous. For treatments that had cumulative effects, to estimate the cumulative treatment effect at time t, we would need to trace the transitions between substates from time 0 to t to calculate the time spent in each substate from time 0 to t, as the treatments may vary by disease substate. The disease substates created by our framework record multimorbidity status at each time point and are used to estimate MALY, however, they do not trace the transitions between substates from time 0 to t. Although accommodating cumulative effects is beyond the scope of this article, our proposed method can still be applied under two specific scenarios involving cumulative treatments: a static treatment regime with same treatment applied to all disease states, and a two-state transition that is equivalent to a conventional single time-to-event outcome.
Another assumption we assume in this article is that confounders are constant over follow-up. Time-varying confounders can be influenced by treatment and disease state. Under the assumption of instantaneous effect of treatment on confounders, time-varying confounders can be modeled by tracing the transition between time and . It is of high interest to extend our causal framework to allow for time-varying confounding for future research.
In summary, we developed a multi-state causal framework to estimate the health benefits of both static and dynamic treatment regimes over the disease course. This method enables the identification of the optimal treatment regime that maximizes health benefits, as measured by MALY, and reveals the mechanisms through which treatment regimes influence disease progression, as captured by SSLY. Our method has the potential to advance the precision prevention of chronic diseases.
Supplemental Material
Supplemental material, sj-docx-1-smm-10.1177_09622802261445822 for Estimating the effects of treatment regimes over the course of chronic disease: A multi-state causal framework with baseline confounding by Ming Ding in Statistical Methods in Medical Research
Footnotes
ORCID iD: Ming Ding https://orcid.org/0000-0001-9947-3446
Funding: The author received no financial support for the research, authorship, and/or publication of this article.
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data availability: The R scripts for the application study is available in the GitHub repository (https://github.com/mingding-hsph/multistate_causal). The ARIC data were obtained from NHLBI BioLINCC under a data usage agreement that restricts open access. Researchers interested in accessing the ARIC data can request it from NHLBI BioLINCC.
Supplemental material: Supplemental material for this article is available online.
References
- 1.Robins JM, Hernan MA, Brumback B. Marginal structural models and causal inference in epidemiology. Epidemiology 2000; 11: 550–560. [DOI] [PubMed] [Google Scholar]
- 2.Young JG, Cain LE, Robins JM, et al. Comparative effectiveness of dynamic treatment regimes: an application of the parametric g-formula. Stat Biosci 2011; 3: 119–143. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Olafiranye O, Zizi F, Brimah P, et al. Management of hypertension among patients with coronary heart disease. Int J Hypertens 2011; 2011: 653903. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Viau DM, Sala-Mercado JA, Spranger MD, et al. The pathophysiology of hypertensive acute heart failure. Heart 2015; 101: 1861–1867. [DOI] [PubMed] [Google Scholar]
- 5.Meira-Machado L, de Una-Alvarez J, Cadarso-Suarez C, et al. Multi-state models for the analysis of time-to-event data. Stat Methods Med Res 2009; 18: 195–222. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Ding M, Chen H, Lin FC. A discrete-time split-state framework for multi-state modeling with application to describing the course of heart disease. BMC Med Res Methodol 2025; 25: 54. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Ding M, Lin FC, Meyer ML. Summary estimates derived from a multi-state non-Markov framework to characterize the course of heart disease. BMC Med Res Methodol (in revision) medRxiv (2024). [DOI] [PMC free article] [PubMed]
- 8.Hernán MA JMR. Causal inference: what If. Boca Raton, FL: Chapman & Hall/CRC, 2020. [Google Scholar]
- 9.Gray RG. A class of K-sample tests for comparing the cumulative incidence of a competing risk. Ann Stat 1988; 16: 1141–1154. [Google Scholar]
- 10.Gerds TA, Ohlendorff JS, Blanche P, et al. Risk regression models and prediction scores for survival analysis with competing risks. https://githubcom/tagteam/riskRegression (2023).
- 11.Aalen OO, Johansen S. An empirical transition matrix for non-homogeneous Markov chains based on censored observations. Scand J Stat 1978; 5: 141–150. [Google Scholar]
- 12.Murray CJ. Quantifying the burden of disease: the technical basis for disability-adjusted life years. Bull World Health Organ 1994; 72: 429–445. [PMC free article] [PubMed] [Google Scholar]
- 13.Mathers CD, Sadana R, Salomon JA, et al. Healthy life expectancy in 191 countries, 1999. Lancet 2001; 357: 1685–1691. [DOI] [PubMed] [Google Scholar]
- 14.Diseases GBD, Injuries C. Global burden of 369 diseases and injuries in 204 countries and territories, 1990–2019: a systematic analysis for the Global Burden of Disease Study 2019. Lancet 2020; 396: 1204–1222. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Lloyd-Jones DM, Allen NB, Anderson CAM, et al. Life’s essential 8: updating and enhancing the American Heart Association’s construct of cardiovascular health: a presidential advisory from the American Heart Association. Circulation 2022; 146: e18–e43. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Global BMIMC, Di Angelantonio E, Bhupathiraju ShN, et al. Body-mass index and all-cause mortality: individual-participant-data meta-analysis of 239 prospective studies in four continents. Lancet 2016; 388: 776–786. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Khan SS, Ning H, Wilkins JT, et al. Association of Body Mass Index with lifetime risk of cardiovascular disease and compression of morbidity. JAMA Cardiol 2018; 3: 280–287. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Curtis JP, Selter JG, Wang Y, et al. The obesity paradox: body mass index and outcomes in patients with heart failure. Arch Intern Med 2005; 165: 55–61. [DOI] [PubMed] [Google Scholar]
- 19.Flegal KM, Kit BK, Orpana H, et al. Association of all-cause mortality with overweight and obesity using standard body mass index categories: a systematic review and meta-analysis. JAMA 2013; 309: 71–82. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Kalantar-Zadeh K, Rhee CM, Chou J, et al. The obesity paradox in kidney disease: how to reconcile it with obesity management. Kidney Int Rep 2017; 2: 271–281. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Carnethon MR, De Chavez PJ, Biggs ML, et al. Association of weight status with mortality in adults with incident diabetes. JAMA 2012; 308: 581–590. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.The ARIC investigators. The atherosclerosis risk in communities (ARIC) study: design and objectives. Am J Epidemiol 1989; 129: 687–702. [PubMed] [Google Scholar]
- 23.Rose GA. Cardiovascular survey methods. Geneva Albany, NY: World Health Organization; WHO Publications Centre distributor, 1982. [Google Scholar]
- 24.Fung TT, Chiuve SE, McCullough ML, et al. Adherence to a DASH-style diet and risk of coronary heart disease and stroke in women. Arch Intern Med 2008; 168: 713–720. [DOI] [PubMed] [Google Scholar]
- 25.Unger T, Borghi C, Charchar F, et al. 2020 International Society of Hypertension Global Hypertension Practice guidelines. Hypertension 2020; 75: 1334–1357. [DOI] [PubMed] [Google Scholar]
- 26.Grundy SM, Stone NJ, Bailey AL, et al. 2018 AHA/ACC/AACVPR/AAPA/ABC/ACPM/ADA/AGS/APhA/ASPC/NLA/PCNA guideline on the management of blood cholesterol: executive summary: a report of the American College of Cardiology/American Heart Association Task Force on clinical practice guidelines. Circulation 2019; 139: e1046–e1081. [DOI] [PubMed] [Google Scholar]
- 27.American Diabetes A. Diagnosis and classification of diabetes mellitus. Diabetes Care 2010; 33(Suppl. 1): S62–S69. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Murphy SA. An experimental design for the development of adaptive treatment strategies. Stat Med 2005; 24: 1455–1481. [DOI] [PubMed] [Google Scholar]
- 29.Arnett DK, Blumenthal RS, Albert MA, et al. 2019 ACC/AHA guideline on the primary prevention of cardiovascular disease: a report of the American College of Cardiology/American Heart Association task force on clinical practice guidelines. Circulation 2019; 140: e596–e646. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplemental material, sj-docx-1-smm-10.1177_09622802261445822 for Estimating the effects of treatment regimes over the course of chronic disease: A multi-state causal framework with baseline confounding by Ming Ding in Statistical Methods in Medical Research
