Skip to main content
Springer logoLink to Springer
. 2026 Jan 19;44(3):329–341. doi: 10.1007/s40273-025-01580-2

Development of Cardiovascular Risk Equations in People with Overweight or Obesity and Established Cardiovascular Disease Without Diabetes Based on the SELECT Trial

Martin Bøg 1,, Anders Bo Bojesen 1, Scott Emerson 2, Milana Ivkovic 1, Nia C Jenkins 3, Christopher Lübker 1, Muhammad Mamdani 4, Thomas Padgett 3, Luc Van Gaal 5, Peter E Weeke 1, A Michael Lincoff 6
PMCID: PMC12916528  PMID: 41553702

Abstract

Background

Overweight and obesity is a prevalent and growing global health concern associated with a significant healthcare burden. With recent advancements in weight management interventions demonstrating cardioprotective benefits, there is a need for risk equations that can accurately predict cardiovascular event risk in health economic models to support healthcare decision making and resource allocation.

Objective

We aimed to derive risk equations using SELECT trial data that estimate the risk of acute coronary syndrome and stroke in people with established cardiovascular disease and overweight or obesity but without diabetes mellitus for use in health economic modelling.

Methods

Risk equations estimating first in-trial observed acute coronary syndrome and stroke were derived from patient-level data from the SELECT trial, a double-blind randomised placebo-controlled trial comparing semaglutide 2.4 mg with placebo. Risk factors were identified from the literature and by clinical experts. Risk equations were developed using cause-specific Cox proportional hazard models with all-cause mortality as a competing risk. Least absolute shrinkage and selection operator (LASSO) penalisation was applied 100 times in bootstrap samples to refine the risk equations. Final risk equations were validated with clinical experts and using SELECT trial data.

Results

Sex, coronary heart disease, a history of acute coronary syndrome and use of angina medications had the strongest associations with acute coronary syndrome (hazard ratio 1.67–1.82). For stroke, a history of atrial fibrillation, cerebrovascular disorder, stroke and transient ischaemic attack had the strongest associations (hazard ratio 1.40–1.70). In terms of discrimination, the risk equations had an area under the receiver operating characteristic curve (AUC) of 0.66–0.68 for acute coronary syndrome and 0.68–0.71 for stroke, and C-indices of 0.67–0.69 for acute coronary syndrome and 0.66–0.69 for stroke. The risk equations showed good cohort-level predictive capabilities over a 4-year time horizon.

Conclusions

These novel, treatment-specific, trial-derived risk equations are the first to predict the risk of secondary cardiovascular events in people with overweight or obesity and established cardiovascular disease but without diabetes. Incorporating these risk equations into health economic models may improve the accuracy of economic evaluations in this population.

Supplementary Information

The online version contains supplementary material available at 10.1007/s40273-025-01580-2.

Key Points for Decision Makers

Health economic models for overweight and obesity rely upon effective prediction models to estimate the risk of events. However, risk equations predicting secondary cardiovascular events in populations with overweight or obesity and established cardiovascular disease but without diabetes mellitus are lacking.
The recent availability of randomised controlled trial data conducted in such populations provides the opportunity to develop risk equations to support decision making in this specific population.
The risk equations from this study are appropriate for predicting the risk of secondary cardiovascular outcomes, suitable for economic evaluations in populations represented by the SELECT trial, i.e. those with overweight or obesity and established cardiovascular disease but without diabetes, and are the first to include semaglutide treatment as a covariate capturing observed treatment-specific effects.

Introduction

The World Obesity Federation estimated that by 2035 more than half of the global population—over 4 billion people—will be living with either overweight or obesity [1]. Despite both clinical and epidemiological evidence linking obesity to the development of health complications, including cardiovascular disease (CVD) and mortality, effective pharmacological interventions that result in durable weight loss and reduced cardiovascular (CV) outcomes have not been forthcoming until recently [24].

As the global obesity epidemic continues to grow and place increasing demand on healthcare systems, robust economic evaluations are needed to assess the value of emerging technologies and support decision makers in allocating healthcare resources. Health economic models are often used for this purpose and estimate the overall number of clinical events and associated costs and utilities to determine the cost effectiveness of interventions. As such, accurate risk prediction of clinical events is essential for the accuracy of these models.

Health economic models used in obesity apply risk equations to estimate the onset of clinical events, including type 2 diabetes (T2D), acute coronary syndrome (ACS), stroke, sleep apnoea, knee replacement and obesity-linked cancers, based on modifiable cardio-metabolic risk factors [57]. The Framingham Recurring Coronary Heart Disease (CHD) Study risk equation is commonly used to estimate the risk of recurrent CVD events [57]. However, the risk equation was derived from data collected from the 1970s, in a population not limited to people with overweight or obesity, and does not adjust for weight or body mass index (BMI). When available, using individual-level data from the appropriate trial to derive risk equations may produce more accurate predictions of event rates than using pre-existing risk equations from broader populations.

Prior CV outcomes trials (CVOTs) conducted in people with cardio-renal metabolic diseases have largely been focused on populations with T2D, limiting the opportunities to develop CV risk equations in populations with obesity and CVD, but without pre-existing diabetes. The SELECT trial (NCT03574597 [8]) was the first CVOT in people with overweight or obesity and established CVD but without diabetes. Therefore, risk equations derived from the SELECT trial data may yield greater risk prediction accuracy of secondary CV outcomes in similar populations than the Framingham Recurring CHD Study risk equation. This study aimed to derive risk equations for the first occurrence of ACS and stroke in people with overweight or obesity and established CVD, but without diabetes, with high population-level prediction accuracy, suitable for use in health economic modelling.

Methods

Data Source

Risk equations were derived from patient-level data from the SELECT trial, a multicentre, double-blind, randomised, placebo-controlled, event-driven superiority trial, in people with overweight or obesity and established CVD (prior myocardial infarction [MI], stroke or symptomatic peripheral arterial disease), but without diabetes. Eligible participants were required to be aged ≥ 45 years with BMI ≥ 27 kg/m2. Those with haemoglobin A1c (HbA1c) ≥ 48 mmol/mol (6.5%) or a history of type 1 diabetes or T2D, who had MI, stroke, hospitalisation for unstable angina pectoris or a transient ischaemic attack (TIA) within 60 days of screening, or who had New York Heart Association class IV heart failure were excluded. The SELECT trial demonstrated subcutaneous semaglutide 2.4 mg weekly was superior to standard of care (hazard ratio [HR], 0.80; 95% confidence interval, 0.72–0.90; p < 0.001) in reducing the incidence of death from CV causes, non-fatal MI or non-fatal stroke at a mean follow-up of 39.8 months, with a maximum follow-up of 56 months. Full details on the design of the SELECT clinical trial have been published previously [8, 9].

Outcomes

Outcomes of interest were defined as first in-trial observed ACS (including fatal and non-fatal MI and hospitalisation for unstable angina, as in the Core Obesity Model [5]) or first in-trial observed stroke event (fatal and non-fatal; defined as any event of an acute ischaemic or haemorrhagic stroke, not including TIA). Definitions for MI, hospitalisation for unstable angina and stroke events from the SELECT trial have been published previously [8]. Analyses were based on the intention-to-treat principle, assessing effectiveness in all subjects who underwent randomisation irrespective of adherence to treatment or changes to background medication.

Clinical Risk Factors

An initial list of candidate risk factors was compiled based on demonstrated links to the outcomes, or to comorbid kidney disease, CVD or diminished glycaemic control in the literature [1013], or risk factors used in established CV risk prediction equations or risk calculators [1420]. Clinical experts were asked in a questionnaire to rank the candidate risk factors by importance for each outcome, and to specify additional risk factors that were omitted, to inform a discussion. Risk factors were identified before the read-out of SELECT trial data to reduce the impact of trial outcomes influencing selection; this list was then refined based on the availability of data within the SELECT trial dataset and the availability of measurements at, or prior to, baseline. Further details of expert elicitation are outlined in the Electronic Supplementary Material (ESM), and the refined list of candidate risk factors is available in Table S1 of ESM 1.

Statistical Analysis and Derivation of Risk Equations

Risk equations for first in-trial ACS or stroke events were developed using cause-specific Cox proportional hazard models for survival data. The Cox proportional hazards model was chosen as it does not make any assumptions about the underlying distribution of the baseline hazard. All-cause mortality was included as a competing risk in the models to account for censoring due to death before the event of interest occurred. The ACS and stroke elements of the models were adjusted for risk factors; the all-cause mortality competing risk element was unadjusted. Administrative censoring was applied at the end of the study or upon patient withdrawal from the study. Patients discontinuing from their assigned treatment were not censored.

The risk equations were derived by repeated application of a four-step process (Sect. 2.1 of the ESM) that tested combinations of the candidate risk factors to determine the most influential combinations, in order to obtain parsimonious final risk equations. The regression algorithm least absolute shrinkage and selection operator (LASSO) penalisation was applied to all candidate risk factors and interaction pairs, to remove the least important risk factors for predicting the outcome. This was repeated 100 times in bootstrap samples, with risk factors remaining in at least 50 iterations retained to form the preliminary risk equations. Treatment allocation was retained in all 100 iterations for both risk equations by design (Figs. S2 and S3 of ESM 1). These preliminary risk equations were discussed with clinical experts with extensive experience in obesity, cardiology and statistics, to validate the risk factors identified by LASSO penalisation and to identify any additional risk factors to be included. Further details of the expert elicitation are outlined in the ESM, and suggestions for additions and removals of risk factors are outlined in Table S2 of ESM 1. Based on these discussions, the refined list of risk factors was updated and the derivation process re-run, with the input from clinical experts retained to form the final risk equations. This approach to risk factor selection limits overfitting to maximise the external validity of the risk equations. The proportional hazards assumption of the cause-specific Cox proportional hazards model was tested using the global Schoenfeld residual test; a statistically significant p value indicates non-proportional hazards and requires alternative modelling techniques to be explored. All analyses were carried out in R v4.2.

Missing Data

Missing data for baseline characteristics were populated with mean imputed values from multiple imputation sets derived from predictive mean matching in chained equations using all available baseline data. No imputation was used on missing outcome variables where missingness was due to loss to follow-up, as this was handled via censoring in the survival models. Missing data for time since CV events at baseline were not imputed.

Internal Validation

The derived risk equations for ACS and stroke were internally validated against the SELECT trial data to assess sensitivity, specificity and calibration. The predictive performance of the risk equations was evaluated by testing their discrimination, i.e. the ability to discriminate between those who experienced the event and those who did not. Receiver operating characteristic curves and the related area under the receiver operating characteristic curve (AUC) metrics were prepared for 1-, 2-, 3-, and 4-year time horizons by plotting true-positive rate (sensitivity) against the false-positive rate (1-specificity) for event prediction. A diagonal line through the origin represents a model that is no better than random prediction, with a corresponding AUC of 0.5. Harrell’s concordance (C)-index measures the proportion of patient pairs whose predicted event times had the same relative order as their true event times (the higher the value, the better the model’s predictive capability). Over ten iterations, AUC and C-indices were derived by fitting the model to 90% of the SELECT trial data, then the predictions validated against the remaining 10% (leave-one-group-out cross-validation). The leave-one-group-out cross-validation allows a more realistic measure of discrimination than when the validation sample is the same as the derivation sample.

Risk equation performance was also assessed based on their calibration, i.e. how accurately absolute risk was predicted. Calibration was assessed in two ways: plotting the observed cumulative incidence curve over time, split by the patient-predicted risk decile and overlaying the predicted incidence curves; and plotting predicted incidence against observed incidence per decile at 1-, 2-, 3- and 4-year prediction horizons for the whole population as well as split by treatment arm.

Comparison to Published Risk Equations

The risk equations for ACS and stroke derived from the SELECT trial data were compared with the Framingham Recurring CHD Study risk equation, which predicts a composite outcome in which only MI and angina pectoris are considered components of ACS. The combined proportions of the MI and angina pectoris events observed in the Framingham Recurring CHD study (0.639 for male individuals, 0.692 for female individuals) were used to reweight predicted probabilities of non-fatal ACS from the Framingham Recurring CHD risk equation. As there is no stroke component in the Framingham Recurring CHD risk equation, only the ACS risk equation was compared with the ACS risk equation derived from the SELECT trial. Furthermore, only the predicted risk in a population receiving standard of care (placebo) was compared, as the Framingham Recurring CHD risk equation  does not capture the treatment effects of semaglutide; predictions under the standard-of-care arm provides more comparable conditions and provides a fairer comparison to demonstrate differences in performance. Risk equation performance was compared by testing the discrimination in terms of the C-index, and calibration via plotting predicted incidence against observed incidence per decile at 1-, 2-, 3- and 4-year prediction horizons.

Results

Final Risk Equation Development

Baseline characteristics of the risk factors included in the final risk equations are provided in Table 1 (note: full baseline characteristics of the SELECT trial population have been previously reported [9]). The mean age of patients was 61.6 years, and 72.3% were male. Most patients were either current or former smokers (16.8% and 48.5%, respectively). At baseline, 76.4% had a history of ACS, and 23.3% had a history of stroke. During the SELECT trial, 789 (4.5%) patients experienced an ACS event, and 338 (1.9%) experienced a stroke event. Levels of missing data, requiring imputation, were low (0.1–2.6%).

Table 1.

Baseline characteristics for risk factors from the SELECT trial used in the final risk equations

Variables Total population
Patients: N 17,604
ACS events, n (%) 789 (4.5)
 MIa 577 (3.3)
 Hospitalisation for unstable anginaa 233 (1.3)
Stroke events, n (%) 338 (1.9)
Demographics/cohort summary/anthropometrics
Treatment allocation, n (%)
 Semaglutide 8803 (50.0)
 Placebo 8801 (50.0)
Age (years)
 Mean (SD) 61.6 (8.9)
 Median (IQR) 61.0 (55.0, 68.0)
Sex, n (%)
 Female 4872 (27.2)
 Male 12,732 (72.3)
Tobacco use, n (%)
 Current smoker 2950 (16.8)
 Previous smoker 8532 (48.5)
 Never smoked 6120 (34.8)
 Missing 2 (0.01)
Comorbidity, n (%)
 Sleep apnoea 2571 (14.6)
 COPD 1471 (8.4)
 CHD 14,452 (82.1)
 History of ACS 13,452 (76.4)
 Revascularisation 11,849 (67.3)
 AF 1613 (9.2)
 Cerebrovascular disorder 1176 (6.7)
 History of stroke 4110 (23.3)
 TIA 764 (4.3)
Lipids/vital signs
Systolic blood pressure (mmHg)
 Mean (SD) 130.9 (15.0)
 Median (IQR) 130.0 (120.0, 140.0)
 Missing, n (%) 10 (0.1%)
Total/HDL cholesterol ratio
 Mean (SD) 3.7 (1.1)
 Median (IQR) 3.5 (2.9, 4.3)
 Missing, n (%) 486 (2.8)
Biochemistry/glucose metabolism/haematology/urinalysis
log(serum hsCRP) (mg/L)
 Mean (SD) 0.66 (1.13)
 Median (IQR) 0.60 (− 0.13, 1.41)
 Missing, n (%) 119 (0.7)
Urine albumin/creatinine ratio (mg/mmol)
 Mean (SD) 3.12 (9.0)
 Median (IQR) 0.85 (0.51, 1.77)
 Missing: n (%) 463 (2.6)
Medications
 Anti-angina medications: n (%) 7030 (39.9)
 Platelet aggregation inhibitors: n (%) 15,181 (86.2)

ACS acute coronary syndrome, AF atrial fibrillation, CHD coronary heart disease, COPD chronic obstructive pulmonary disease, HDL high-density lipoprotein, hsCRP high-sensitivity C-reactive protein, IQR interquartile range, MI myocardial infarction, SD standard deviation, TIA transient ischaemic attack

aA patient may experience both MI and hospitalisation for unstable angina

The estimated cause-specific HRs for the final risk equations are listed in Table 2. Equations used for the calculation of absolute risk are provided in ESM 1; associated baseline hazards are provided in ESM 2. For both ACS and stroke, treatment with semaglutide was associated with a lower cause-specific hazard of the event compared with placebo, demonstrating a 23% reduction in the cause-specific hazard of ACS and a 10% reduction in the cause-specific hazard of stroke over the trial horizon (up to 56 months).

Table 2.

Estimated coefficients for final risk equations

Variable ACS Stroke
HR (95% CI) HR (95% CI)
Demographics/cohort summary/anthropometrics
 Treatment arm (vs placebo) 0.77 (0.67, 0.88) 0.90 (0.72, 1.11)
 Age 0.99 (0.99, 1.00) 1.03 (1.02, 1.04)
 Currently smoking 1.45 (1.23, 1.72) NA
 Male (vs female) 1.67 (1.26, 2.22) NA
Comorbidities
 Apnoea 1.22 (1.01, 1.46) NA
 COPD 1.48 (1.20, 1.84) NA
 CHD 1.77 (1.11, 2.84) NA
 History of ACS 1.82 (0.79, 4.17) NA
 Revascularisation 1.49 (1.18, 1.87) NA
 AF NA 1.51 (1.09, 2.08)
 Cerebrovascular disorder NA 1.41 (1.02, 1.95)
 History of stroke NA 1.40 (0.88, 2.22)
 TIA NA 1.70 (1.19, 2.41)
Lipids/vital signs
 SBP (mmHg) NA 1.01 (1.00, 1.02)
 Total/HDL cholesterol ratio 1.25 (1.06, 1.49) NA
Biochemistry/glucose metabolism/haematology/urinalysis
 log (hsCRP (mg/L)) 1.15 (1.08, 1.23) NA
 uACR (mg/mmol) 1.01 (1.00, 1.01) 1.00 (0.99, 1.01)
Medications
 Angina medications 1.71 (1.21, 2.41) NA
 Platelet aggregation inhibitors NA 1.13 (0.78, 1.64)
Interactions
 Angina medications × male (interaction) 0.79 (0.54, 1.15) NA
 Total/HDL cholesterol ratio × history of ACS (interaction) 0.94 (0.78, 1.12) NA
 Platelet aggregation inhibitors × no stroke (interaction) NA 0.55 (0.33, 0.93)

ACS acute coronary syndrome, AF atrial fibrillation, CHD coronary heart disease, CI confidence intervals, COPD chronic obstructive pulmonary disease, HDL high-density lipoprotein, HR hazard ratio, hsCRP high-sensitive C-reactive protein in serum, NA not applicable, SBP systolic blood pressure, TIA transient ischaemic attack, uACR urine albumin creatinine ratio

Sex, CHD, a history of ACS and angina medications had the strongest associations with ACS (HR > 1.5), with a history of ACS having the strongest association (HR 1.82). The increased risk associated with angina medication (HR 1.71) was reduced for male patients (angina medications × male; HR 0.79). A history of sleep apnoea, chronic obstructive pulmonary disease, CHD and revascularisation were associated with an increase in the cause-specific hazard of ACS. Of the continuous variables, total/high-density lipoprotein (HDL) cholesterol ratio and log (high-sensitivity C-reactive protein) had the strongest associations (a one standard deviation increase increased the cause-specific hazard by 28% and 17%, respectively). The increased risk associated with a unit increase in total/HDL cholesterol (HR 1.25) was reduced for those with a history of ACS (total/HDL cholesterol ratio × history of ACS; HR 0.94). Note, the associations observed between variables and outcomes should be interpreted with caution, these variables may act as markers of other factors rather than as direct causal effects.

For stroke, a history of atrial fibrillation and a history of TIA showed a HR > 1.5, with TIA having the strongest association (HR 1.70). All included comorbidity histories (atrial fibrillation, cerebrovascular disease, stroke and TIA), platelet aggregation inhibitors, and increasing age and systolic blood pressure (SBP) were associated with an increase in cause-specific hazard of stroke. The increased risk associated with platelet aggregation inhibitors (HR 1.13) was reduced for patients without a history of stroke (platelet aggregation inhibitors × no stroke; HR 0.55). A 1 standard deviation increase in age was shown to increase the cause-specific hazard of stroke by 28%, while a 1 standard deviation increase in SBP increased the cause-specific hazard of stroke by 18%.

Internal Validation

Receiver operating characteristic curves and AUCs showed little variation between 1-, 2-, 3-, and 4-year time horizons (0.66–0.68 for ACS and 0.68–0.71 for stroke), showing consistent discrimination ability (Fig. 1). The 4-year C-indices for the total population, semaglutide arm and placebo arm were 0.68, 0.69 and 0.67, respectively, for ACS, and 0.68, 0.69 and 0.66, respectively, for stroke.

Fig. 1.

Fig. 1

Receiver operating characteristic curves for (A) acute coronary syndrome (ACS) and (B) stroke, for varying time horizons, with associated area under the receiver operating characteristic curve (AUC) values, in the total population. The 45° grey line represents the line of random probabilities

The ACS risk equation broadly predicts absolute risk accurately across the whole trial time period for deciles 1, 3, 4, 7 and 10, with underprediction for deciles 2 and 5, and overprediction for deciles 6, 8 and 9 (Fig. 2a), where decile 1 represents the 10% of the population with the lowest risk and 10 represents the 10% of the population with the highest risk. For stroke, predicted and observed curves generally align across the trial time period for most deciles (Fig. 2b).

Fig. 2.

Fig. 2

Cumulative incidence curves for (A) acute coronary syndrome (ACS) and (B) stroke, for observed and predicted data, split by decile of predicted risk

Calibration plots of predicted incidence versus observed incidence per decile for 1-, 2-, 3- and 4-year prediction time horizons are presented in Fig. 3. For both ACS and stroke, across all time horizons, the observed risk increased with increasing risk decile and there was an approximately even distribution of points along the diagonal line, indicating that predictive performance was not dependent on risk decile. When splitting by treatment arm (Fig. 4), both arms had points lying close to and on either side of the diagonal line for all time horizons, indicating that error was not correlated to incidence and was consistent between treatment arms.

Fig. 3.

Fig. 3

Calibration plot for (A) acute coronary syndrome (ACS) and (B) stroke of predicted incidence versus observed incidence per risk decile. The dashed line represents the line of perfect calibration. Each dot represents one decile. y years

Fig. 4.

Fig. 4

Calibration plots for (A) acute coronary syndrome (ACS) and (B) stroke of predicted incidence versus observed incidence per risk quintile per treatment arm. The dashed line represents the line of perfect calibration. Each dot represents one quintile. y years

Comparison to Published Risk Equations

The predictive ability of the ACS risk equation derived from the SELECT trial was compared with the Framingham Recurring CHD Study risk equation on SELECT trial data. The SELECT trial derived risk equation had a higher C-index than the Framingham equation (0.67 vs 0.57) and demonstrated better calibration. Calibration plots in Fig. 5 show that the SELECT trial derived risk equation has an approximately even distribution of points along the diagonal lines, whereas the Framingham risk equation has more points below the diagonal line with predicted incidence deviating further below with an increasing time horizon, demonstrating that the Framingham equation consistently overpredicted risk. However, it is important to note that models are expected to perform better in the population where they were derived.

Fig. 5.

Fig. 5

Calibration plots for acute coronary syndrome (ACS) of predicted incidence versus observed incidence per risk decile for the Framingham Recurring CHD Study and SELECT trial. Each dot represents one decile. y years

Discussion

Health economic models used to evaluate the cost effectiveness of healthcare interventions in chronic diseases such as obesity often incorporate risk equations to predict event risk within specific populations, to enable the translation of clinical outcomes into costs and health consequences in models. In this study, we report the development of separate risk equations for the prediction of ACS and stroke in people with overweight or obesity and established CVD but without diabetes, derived in a multi-step process from SELECT trial data. Utilising patient-level data from a cohort that closely aligns with the target population to derive risk equations can improve the predictive ability, accuracy and robustness of health economic models compared with pre-existing risk equations. The Framingham Recurring CHD Study risk equations have been used in health economic models of obesity, such as the Core Obesity Model [5, 6]; however, they are not reflective of people with overweight or obesity but without diabetes. Compared with the Framingham Recurring CHD Study risk equations, our novel, treatment-specific, trial-derived risk equations more accurately predict the risk of secondary CV events in patient groups with overweight or obesity and established CVD without pre-existing diabetes, as represented in the SELECT population. It is important to note that this comparison only reflects the predictive performance in the SELECT population; it is anticipated that our derived risk equations will have superior performance in the SELECT population than published risk equations, such as Framingham, as predictive performance on the derivation population is usually higher than on external data sets.

Discrepancies in the definition of outcomes between the Framingham Recurring CHD Study risk equations and the SELECT trial should be noted. Specifically, SELECT considered both fatal and non-fatal ACS events; however, in the Framingham risk equation, components of ACS considered are non-fatal MI and angina, only. Furthermore, the Framingham composite outcome of CHD requires splitting into its component parts, with the assumption that the distribution of risk across the component parts corresponds to the number of events observed for each component part.

Treatment with semaglutide was associated with a reduction in cause-specific hazard for both ACS (23% reduction) and stroke (10% reduction) compared with placebo; therefore, treatment with semaglutide is strongly prognostic, requiring treatment-specific risk equations. In both risk equations, for ACS and stroke, the presence of specific comorbidities (e.g. a history of CHD, ACS, atrial fibrillation, TIA and angina medication as a risk marker for angina) as well as biomarkers and other metrics (e.g. total/HDL cholesterol, high-sensitivity C-reactive protein, age and SBP) were shown to be influential. Among published risk equations, there is variation around the inclusion of comorbidities when predicting risk of CV events [1517, 21]. This could be due to the differences in population (e.g. diabetes vs non-diabetes), the data source from which risk equations were derived (e.g. robust patient biomarker data are likely to be more limited in real-world datasets compared with randomised controlled trials), the relationship between the risk factors and treatment effect, or the methodological approach taken to derive the risk equations (e.g. this study used a data-driven process [LASSO penalisation] combined with clinical elicitation to identify influential risk factors, whereas an alternative approach may place different emphasis on parsimony or derive models based primarily on clinical opinion). It is important to note when interpreting these results that risk factors included within the risk equations, such as angina medication or revascularisation, may reflect the risk of the underlying condition, rather than an increased risk associated with the interventions.

A key strength of this study lies in the methodology used to derive the risk equations, combining statistical robustness with clinical validity. As far as the authors are aware, this is an approach not previously used in the derivation of similar CV risk equations. Known risk factors associated with CV events were identified, and model selected risk factors were assessed for face validity with clinical experts, reinforcing the balance between clinical and statistical validity. Incorporating robust statistical methods, such as LASSO penalisation, mitigated against the risk of overfitting and ensured strong statistical relationships present in the data were captured in the risk equations.

The risk equations included factors influenced by adiposity, including comorbidities (e.g. sleep apnoea) and biomarkers (e.g. total/HDL cholesterol, high-sensitivity C-reactive protein and SBP). Despite strong evidence describing the relationship between adiposity and CV events [4], weight-based measures such as BMI, while considered as candidate risk factors, did not have strong predictive value in the SELECT population. Clinical experts were able to validate this as likely due to the smaller differential impact of BMI increasing risk in the SELECT population, which is already at an elevated risk because of overweight or obesity, compared with the general population. Further, this may indicate that CV risk related to adiposity operates primarily through its effects on metabolic and inflammatory pathways; once these mediators are accounted for, BMI provides limited predictive value. As such, if these measures or their trajectories are not available, then BMI could function as a noisy proxy. In economic models where such CV biomarkers are tracked longitudinally, CV benefits are captured through improvements in these factors and risk is recalculated over time; however, residual adiposity-related risk may exist, potentially limiting the equations’ ability to fully capture full CV benefits.

The recent focus on demonstrating the cardioprotective effects of new drugs in the cardio-renal-metabolic disease space has highlighted that older approaches to modelling CV event risk, which pre-date dedicated CVOTs, may fail to capture treatment-related CV benefits. The cardioprotective impact demonstrated in CVOTs in T2D populations extends beyond the improvements in known biomarker risk factors, suggesting health economic modelling would benefit from risk equations that include treatment effect, unless this effect can be reflected in measurable risk factors [2224]. In the SELECT trial, the magnitude of the reduction in the three-point MACE endpoint was shown to be driven by factors beyond weight loss, suggesting that alternative mechanisms beyond reduced adiposity influence CV risk reduction [25]. Classic diabetes models often incorporate CV risk via changes in surrogate risk factors, which may not be adequate to fully capture changes in CV risk, and post-hoc efforts to improve predictive capability via calibration have shown only partial success [22, 24, 26]. The risk equations we present are derived directly from SELECT trial data, and are therefore expected to accurately capture the cardioprotective benefits of semaglutide and provide a better estimate of the value of therapy. However, further work is needed to validate findings and determine their real-world applicability, acceptability and generalisability in modelling outcomes in populations with overweight or obesity with established CVD without T2D.

The primary limitation of this study is uncertainty as to the translation of trial-derived risk equations to real-world populations. The SELECT trial was highly monitored and the extent to which the survival estimates generalise to a broader population is unknown. Analytical steps have been taken in the design of the study to increase the generalisability of the final models, such as incorporating additional risk factors that are known to be clinically associated with the CV event in question. Further research is ongoing to externally validate these risk equations in order to test the generalisability to real-world populations. Overall, the derived risk equations in the current work were calibrated well to the SELECT population (i.e. predict events well at the cohort level) for all time horizons, and the ability to discriminate (i.e. predict events well at the individual level) was in line with published secondary CV risk equations [16, 20, 27, 28]. Tools used in clinical practice to predict risk outcomes in individuals require parsimonious risk equations with high discrimination [2932]; however, the risk equations derived in this study are intended for modelling health outcomes, which tend to consider population-level outcomes for which good calibration is most critical. Other limitations include the relatively small number of observed SELECT in-trial events of ACS and stroke, and the lack of follow-up beyond 56 months to inform predictions exceeding this time frame, which will require further work to investigate extrapolations using parametric models. Additionally, this analysis only considered risk factors at baseline; as part of future work, alternatives such as landmarking to incorporate post-baseline risk factors may be explored. Alternative modelling methods to Cox proportional hazard methods using machine learning are also available and may provide improved performance. Methods used for economic evaluations to support health technology assessments need to be independently assessed for suitability; machine learning models are often perceived to lack transparency by assessors, and non-standard methodologies may limit the ability of review groups to decide if they are appropriate. During the development of these risk equations, the practical implications of the methods were also considered to balance overall performance with other elements such as transparency and reproducibility.

Conclusions

This study outlines an innovative approach to developing risk equations that incorporates statistical approaches for risk factor selection (LASSO penalisation and 100 repeated runs) with expert clinical knowledge to support the identification and validation of risk factors, resulting in parsimonious, representative and clinically valid risk equations that demonstrated acceptable accuracy and good predictive capability over a 4-year time horizon. These risk equations are the first to include semaglutide as a covariate, capturing observed treatment-specific effects that would not be reflected otherwise.

Supplementary Information

Below is the link to the electronic supplementary material.

Acknowledgements

The authors thank James Dennis and Victoria O’Morain for providing advice and support in manuscript preparation, which was funded by Novo Nordisk in accordance with Good Publication Practice (GPP3) guidelines (https://www.ismpp.org/gpp3).

Funding

This work was supported by Novo Nordisk who provided support for analysis and medical writing for this study.

Declarations

Conflict of Interest

Martin Bøg, Anders Bo Bojesen, Milana Ivkovic, Christopher Lübker and Peter E. Weeke are employees of Novo Nordisk. Nia C. Jenkins and Thomas Padgett are employees of Health Economics and Outcomes Research Ltd. Health Economics and Outcomes Research Ltd. received fees from Novo Nordisk in relation to this study. Scott Emerson has received consulting honoraria from Amylyx, AstraZeneca, Avillion, Ayala, Bayer, BeiGene, Boehringer Ingelheim, 89 Bio, BioAge, BioAtla, Bristol Meyer Squibb, BridgeBio, Daiichi Sankyo, Denovo, Fore Therapeutics, GlaxoSmithKline, Inovio, Insmed, Ipsen, Karuna, Lilly, Lundbeck, Mirati, Moderna, Novartis, Novavax, Novo Nordisk, NSABP, Pfizer, Principia, Reata, Rebiotx, Roche, Sanofi, SOLVD, Sutro Biopharma and TG Therapeutics. Muhammad Mamdani has received grant funding from Roche, honoraria for speaking engagements and one-time advisory boards from Takeda and Baxter, and is an advisor to Signal 1 and Mutual Health. A. Michael Lincoff has received research grants paid to his institution from AbbVie Inc., AstraZeneca, CSL Behring, Eli Lilly and Company, Esperion Therapeutics, Inc. and Novartis and has served as a consultant for Akebia Therapeutics Inc., Alnylam Pharmaceuticals Inc., Ardelyx, Canary Cure, Eli Lilly and Company, FibroGen, GlaxoSmithKline, Intarcia, Medtronic Vascular, Inc., Novartis Pharmaceuticals Corporation, Novo Nordisk, Provention Bio, Entity and ReCor Medica. Luc Van Gaal is a member of the speakers bureau and/or advisory board for Bayer Pharma, Boehringer Ingelheim, Currax Pharma, Daiichi-Sankyo, Eli Lilly & Co., Nestlé Health Science, Novo Nordisk and Regeneron Pharma.

Ethical Approval

The CPRD Independent Scientific Advisory Committee for Medicines & Healthcare products Regulatory Agency database research granted ethics approval (protocol number 22_002360).

Consent to Participate

Not applicable.

Consent for Publication

Not applicable.

Availability of Data and Material

The datasets generated and/or analysed during the current study are available from the corresponding author on reasonable request.

Code Availability

Not applicable.

Author Contributions

MB, ABB and MI conceptualised and designed the study. NJ and TP were responsible for the data analysis. All authors contributed to interpretation of the results, review of the manuscript and approval of the final manuscript for publication.

References

  • 1.World Obesity Federation. World obesity atlas 2023. 2023. https://data.worldobesity.org/publications/?cat=19. Accessed 27 Dec 2025.
  • 2.Wilding J, Jacob S. Cardiovascular outcome trials in obesity: a review. Obes Rev. 2020;22:e13112. [DOI] [PubMed] [Google Scholar]
  • 3.Ryan DH, Lingvay I, Colhoun HM, et al. Semaglutide effects on cardiovascular outcomes in people with overweight or obesity (SELECT) rationale and design. Am Heart J. 2020;229:61–9. [DOI] [PubMed] [Google Scholar]
  • 4.Lopez-Jimenez F, Almahmeed W, Bays H, et al. Obesity and cardiovascular disease: mechanistic insights and management strategies: a joint position paper by the World Heart Federation and World Obesity Federation. Eur J Prev Cardiol. 2022;29(17):2218–37. [DOI] [PubMed] [Google Scholar]
  • 5.Lopes S, Meincke HH, Lamotte M, et al. A novel decision model to predict the impact of weight management interventions: the Core Obesity Model. Obes Sci Pract. 2021;7(3):269–80. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Hoog MM, Kan H, Deger KA, et al. Modeling potential cost-effectiveness of tirzepatide versus lifestyle modification for patients with overweight and obesity. Obesity. 2025;33(7):1297–308. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Lopes S, Johansen P, Lamotte M, et al. External validation of the core obesity model to assess the cost-effectiveness of weight management interventions. Pharmacoeconomics. 2020;38(10):1123–33. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Lincoff AM, Brown-Frandsen K, Colhoun HM, et al. Semaglutide and cardiovascular outcomes in obesity without diabetes. N Engl J Med. 2023;389(24):2221–32. [DOI] [PubMed] [Google Scholar]
  • 9.Lingvay I, Brown-Frandsen K, Colhoun HM, et al. Semaglutide for cardiovascular event reduction in people with overweight or obesity: SELECT study baseline characteristics. Obesity. 2023;31(1):111–22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Regnault V, Thomas F, Safar ME, et al. Sex difference in cardiovascular risk: role of pulse pressure amplification. J Am Coll Cardiol. 2012;59(20):1771–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Franklin SS, Khan SA, Wong ND, et al. Is pulse pressure useful in predicting risk for coronary heart disease? The Framingham Heart Study. Circulation. 1999;100(4):354–60. [DOI] [PubMed] [Google Scholar]
  • 12.Arnett DK, Blumenthal RS, Albert MA, et al. 2019 ACC/AHA guideline on the primary prevention of cardiovascular disease: executive summary: a report of the American College of Cardiology/American Heart Association Task Force on clinical practice guidelines. Circulation. 2019;140(11):e563–95. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Visseren FLJ, Mach F, Smulders YM, et al. 2021 ESC guidelines on cardiovascular disease prevention in clinical practice: developed by the Task Force for cardiovascular Disease Prevention in Clinical Practice with representatives of the European Society of Cardiology and 12 medical societies with the special contribution of the European Association of Preventive Cardiology (EAPC). Eur Heart J. 2021;42(34):3227–37. [DOI] [PubMed] [Google Scholar]
  • 14.Hippisley-Cox J, Coupland C, Brindle P. Development and validation of QRISK3 risk prediction algorithms to estimate future risk of cardiovascular disease: prospective cohort study. BMJ. 2017;357:j2099. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.D’Agostino RB, Russell MW, Huse DM, et al. Primary and subsequent coronary risk appraisal: new results from the Framingham study. Am Heart J. 2000;139(2 Pt 1):272–81. [DOI] [PubMed] [Google Scholar]
  • 16.Klooster CCV, Bhatt DL, Steg PG, et al. Predicting 10-year risk of recurrent cardiovascular events andcardiovascular interventions in patients with established cardiovascular disease: results from UCC-SMART and REACH. Int J Cardiol. 2021;325:140–8. [DOI] [PubMed] [Google Scholar]
  • 17.Cederholm J, Eeg-Olofsson K, Eliasson B, et al. Risk prediction of cardiovascular disease in type 2 diabetes: a risk equation from the Swedish National Diabetes Register. Diabetes Care. 2008;31(10):2038–43. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Hayes AJ, Leal J, Gray AM, et al. UKPDS outcomes model 2: a new version of a model to simulate lifetime health outcomes of patients with type 2 diabetes mellitus using data from the 30 year United Kingdom Prospective Diabetes Study: UKPDS 82. Diabetologia. 2013;56(9):1925–33. [DOI] [PubMed] [Google Scholar]
  • 19.Steen DL, Khan I, Andrade K, et al. Event rates and risk factors for recurrent cardiovascular events and mortality in a contemporary post acute coronary syndrome population representing 239 234 patients during 2005 to 2018 in the United States. J Am Heart Assoc. 2022;11(9):e022198. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Dorresteijn JA, Visseren FL, Wassink AM, et al. Development and validation of a prediction rule for recurrent vascular events based on a cohort study of patients with arterial disease: the SMART risk score. Heart. 2013;99(12):866–72. [DOI] [PubMed] [Google Scholar]
  • 21.Talha I, Elkhoudri N, Hilali A. Major limitations of cardiovascular risk scores. Cardiovasc Ther. 2024;2024:4133365. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Si L, Willis MS, Asseburg C, et al. Evaluating the ability of economic models of diabetes to simulate new cardiovascular outcomes trials: a report on the Ninth Mount Hood Diabetes Challenge. Value Health. 2020;23(9):1163–70. [DOI] [PubMed] [Google Scholar]
  • 23.Cefalu WT, Kaul S, Gerstein HC, et al. Cardiovascular outcomes trials in type 2 diabetes: where do we go from here? Reflections from a Diabetes Care Editors’ Expert Forum. Diabetes Care. 2017;41(1):14–31. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Willis M, Asseburg C, Nilsson A, et al. Challenges and opportunities associated with incorporating new evidence of drug-mediated cardioprotection in the economic modeling of type 2 diabetes: a literature review. Diabetes Ther. 2019;10(5):1753–69. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Deanfield J, Lincoff AM, Kahn SE, et al. Semaglutide and cardiovascular outcomes by baseline and changes in adiposity measurements: a prespecified analysis of the SELECT trial. Lancet. 2025;406(10516):2257–68. [DOI] [PubMed] [Google Scholar]
  • 26.McEwan P, Chubb B, Bennett Wilton H. Modelling cardiovascular outcomes in type 2 diabetes in the era of cardiovascular outcomes trials. Value Health. 2017;20:A747. [Google Scholar]
  • 27.Poppe KK, Wells S, Jackson R, et al. Predicting cardiovascular disease risk across the atherosclerotic disease continuum. Eur J Prev Cardiol. 2022;28(18):2010–7. [DOI] [PubMed] [Google Scholar]
  • 28.Holt A, Batinica B, Liang J, et al. Development and validation of cardiovascular risk prediction equations in 76 000 people with known cardiovascular disease. Eur J Prev Cardiol. 2024;31(2):218–27. [DOI] [PubMed] [Google Scholar]
  • 29.Betts MB, Milev S, Hoog M, et al. Comparison of recommendations and use of cardiovascular risk equations by health technology assessment agencies and clinical guidelines. Value Health. 2019;22(2):210–9. [DOI] [PubMed] [Google Scholar]
  • 30.Li X, Li F, Wang J, et al. Prediction of complications in health economic models of type 2 diabetes: a review of methods used. Acta Diabetol. 2023;60(7):861–79. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Alba AC, Agoritsas T, Walsh M, et al. Discrimination and calibration of clinical prediction models: users’ guides to the medical literature. JAMA. 2017;318(14):1377–84. [DOI] [PubMed] [Google Scholar]
  • 32.Emamipour S, Pagano E, Di Cuonzo D, et al. The transferability and validity of a population-level simulation model for the economic evaluation of interventions in diabetes: the MICADO model. Acta Diabetol. 2022;59(7):949–57. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials


Articles from Pharmacoeconomics are provided here courtesy of Springer

RESOURCES