Abstract
Background and objectives
The History, ECG, Age, Risk factors and Troponin (HEART) score is widely used for emergency department (ED) chest pain risk stratification, but its fixed-point structure may not optimally represent relationships between its components and major adverse cardiac events (MACE). We evaluated whether increasingly flexible modelling approaches improve prediction when restricted to the same five HEART components.
Methods
We conducted a retrospective multicentre cohort study of 47 853 adult ED patients with chest pain and documented HEART scores across six EDs from 2021 to 2023. Data from five EDs (n=39 587) were used for model development and one geographically held-out ED (n=8266) was used for validation. The conventional HEART score was compared with categorical cross-validated logistic regression (CVLR) and extreme gradient boosting (XGBoost) for predicting 30-day MACE. Performance was assessed using the area under the receiver operating characteristic (AUROC), the area under the precision–recall curve (AUPRC), the Brier score, paired bootstrap comparisons and clinical performance at comparable-sensitivity operating points.
Results
In the held-out validation cohort, 267 patients (3.23%) experienced 30-day MACE. AUROCs were similar for HEART (0.855), CVLR (0.857) and XGBoost (0.856; pairwise p>0.05). In contrast, both data-driven models demonstrated significantly higher AUPRC than HEART: CVLR (0.237 vs 0.187; Δ=0.050, p<0.001) and XGBoost (0.244 vs 0.187; Δ=0.057, p<0.001) with no significant difference between CVLR and XGBoost (p=0.305). At comparable-sensitivity operating points, HEART, CVLR and XGBoost classified 63.74%, 59.39% and 61.58% of patients as low risk, respectively, with corresponding low-risk MACE rates of 0.53%, 0.51% and 0.57%.
Conclusions
Data-driven reweighting of HEART components improved precision–recall performance and risk ranking compared with fixed HEART scoring. However, XGBoost provided no incremental benefit over categorical CVLR, and neither approach expanded low-risk classification at comparable sensitivity.
Keywords: Chest Pain, Translational Medical Research, Acute Coronary Syndrome
WHAT IS ALREADY KNOWN ON THIS TOPIC
Machine learning models have been reported to outperform the History, ECG, Age, Risk factors and Troponin (HEART) score for chest pain risk stratification, but most studies used additional predictors, making it unclear whether improvements were due to greater model complexity or more clinical information.
WHAT THIS STUDY ADDS
Using identical HEART components across all models, we found that increasing model complexity did not improve discrimination or clinically meaningful low-risk classification beyond simpler data-driven modelling.
HOW THIS STUDY MIGHT AFFECT RESEARCH, PRACTICE OR POLICY
These findings suggest that future efforts should prioritise identifying more informative predictors rather than relying solely on increasingly complex machine learning algorithms.
Introduction
Chest pain remains one of the most common reasons for emergency department (ED) evaluation and a persistent diagnostic challenge. Although most patients do not have acute coronary syndrome (ACS), missed major adverse cardiac events (MACE) carry substantial morbidity and mortality.1 2 Accurate risk stratification is therefore central to ED chest pain pathways, guiding disposition decisions while balancing patient safety, diagnostic testing, hospital admission and healthcare resource utilisation.
The HEART score, based on History, ECG, Age, Risk factors and Troponin, is widely used to estimate short-term risk among ED patients with suspected ACS. Its clinical appeal lies in its simplicity, transparency and bedside applicability.3 4 Validation studies across diverse populations have demonstrated strong discrimination for 30-day MACE, establishing the HEART as an important tool for ED chest pain risk stratification.5 Patients with a HEART score <4 are typically considered low risk, whereas scores ≥7 often prompt additional testing, observation, admission or referral to cardiology.5 6 Despite its clinical utility, the HEART score is intentionally simplified. Each component is assigned a fixed weight (0–2 points), implicitly assuming uniform and linear relationships between predictors and outcomes across heterogeneous patient populations.7 8
Machine learning (ML) and statistical prediction models may address some limitations of fixed-point scoring by allowing data-driven weighting of predictors, modelling non-linear relationships and accounting for interactions among variables.9–12 Prior studies have reported improved discrimination using ML models for chest pain and ACS risk stratification.13–16 However, many such models incorporate a large number of additional variables, making it difficult to determine whether improved performance reflects a better modelling strategy or simply greater access to clinical information.13–16 Moreover, added model complexity may reduce transparency and may not translate into clinically meaningful improvement at decision thresholds relevant to ED disposition.
Accordingly, the incremental value of model complexity remains uncertain when the available clinical information is held constant. A fair comparison requires evaluating traditional HEART scoring, simpler data-driven reweighting and non-linear ML using the same five HEART components. Such a design directly tests whether replacing fixed scoring with increasingly complex prediction models improves discrimination, calibration and clinically actionable low-risk classification.
In this multicentre study, we developed models using data from five EDs and evaluated performance in a geographically held-out sixth ED. We compared the conventional HEART score with cross-validated logistic regression (CVLR) and extreme gradient boosting (XGBoost) models trained exclusively on the same five HEART components for predicting 30-day MACE. We hypothesised that while data-driven models may modestly improve statistical performance, increasing model complexity would provide limited incremental clinical value over the traditional HEART score when no additional predictors are incorporated.
Methods
Study design and setting
We conducted a retrospective, multicentre observational cohort study using routinely collected electronic health record (EHR) data across six EDs between 1 January 2021 and 31 December 2023. Two sites are tertiary referral, Level I trauma centres (>100 000 annual ED visits) and four are community hospitals (60 000–90 000 annual visits). All sites share an integrated EHR system, enabling standardised data capture and outcome ascertainment across centres. The study was approved by the Institutional Review Boards with a waiver of informed consent due to its retrospective design.
Study population
We included adult patients (≥18 years) presenting to the ED with a chief concern of chest pain who had a documented HEART score during the index ED visit. During the study period, the HEART score was calculated as part of routine ACS evaluation at all participating EDs. Patients were excluded if they (1) left before completion of evaluation (including left without being seen, eloped or left against medical advice), (2) were transferred directly to another facility, taken emergently to cardiac catheterisation or died in the ED, or (3) had missing HEART score documentation. To avoid correlated observations, only the first (index) ED visit was included for each patient.
Outcome definition
The primary outcome was 30-day MACE following the index ED visit, defined as a composite of acute myocardial infarction (AMI), coronary revascularisation (percutaneous coronary intervention or coronary artery bypass grafting) or all-cause mortality. AMI was identified using International Classification of Diseases, Tenth Revision, Clinical Modification (ICD-10-CM) codes and revascularisation procedures were identified using ICD-10-Procedure Coding System and Current Procedural Terminology codes. Mortality was ascertained from EHR records and supplemented by available administrative data sources. Outcome ascertainment was based on structured coding and applied uniformly across sites. Because outcomes were derived from administrative data, no subjective adjudication or assessor blinding was performed.
Traditional HEART score
The HEART score comprises five components: History, ECG, Age, Risk factors and Troponin, each scored from 0 to 2, yielding a total score ranging from 0 to 10. All participating sites used high-sensitivity cardiac troponin assays, with scoring standardised using assay-specific 99th percentile upper reference limits. Consistent with established clinical practice, a HEART score of 0–3 was considered low risk, a HEART score of 4–6 was considered intermediate risk and a score of 7–10 was considered high risk.
Data preparation and predictor specification
To permit a controlled comparison of modelling strategies, all prediction approaches were restricted to the same five HEART components: History, ECG, Age, Risk factors and Troponin. No additional demographic, clinical, laboratory, imaging or comorbidity variables were included. Each HEART component was represented using its recorded 0-point, 1-point or 2-point value. Because all five components were required for calculating the documented HEART score, complete component data were available for the analytical cohort and no predictor imputation was performed.
Development and held-out ED validation design
To evaluate model transportability across clinical sites and reduce potential optimism associated with random patient-level partitioning, data from five EDs were used for model development and data from the remaining ED was reserved as a geographically held-out validation cohort. This site-level holdout was selected to provide a more stringent assessment of model transportability than a random patient-level split within the pooled health-system cohort, while preserving differences in patient case mix and outcome prevalence that may occur across ED settings. The held-out ED was not used for model fitting, hyperparameter selection, probability calibration, threshold selection or any other aspect of model development. The development cohort included 39 587 patients from five EDs, and the held-out validation cohort included 8266 patients from the sixth ED. All final model performance and clinical operating characteristics were evaluated in the held-out validation cohort.
Prediction models
Three approaches of increasing modelling complexity were evaluated using identical predictor information: (1) the conventional HEART score, representing fixed-point additive scoring; (2) CVLR, representing data-driven statistical reweighting of the individual HEART components; and (3) XGBoost, representing a more flexible non-linear modelling approach capable of capturing complex predictor relationships and interactions.
Cross-validated logistic regression
A penalised logistic regression model was developed using the five individual HEART components as predictors and 30-day MACE as the binary outcome. Model regularisation was selected using fivefold cross-validation within the development cohort. Predictor scaling and all model-fitting procedures were performed within the development data. Predicted probabilities were calibrated using cross-validated Platt scaling within the development cohort before application to the held-out validation cohort.
XGBoost
An XGBoost classifier was developed using the same five HEART components and 30-day MACE outcome. Model hyperparameters were specified and optimised within the development cohort, and probability calibration was performed using cross-validated Platt scaling. The final calibrated model was then applied without further modification to the held-out ED validation cohort.
HEART-based probability estimation
To permit probability-based evaluation of calibration and Brier scores, HEART-based predicted probabilities were derived using logistic regression fitted in the development cohort, with 30-day MACE as the outcome and total HEART score as the sole predictor. The fitted model was then applied without refitting to the held-out validation cohort. This probability transformation was used only for probability-based performance assessment and did not alter the conventional HEART score or its clinical low-risk threshold of 0–3.
Model performance evaluation
All model performance evaluations were conducted exclusively in the held-out ED validation cohort. Discrimination was assessed using the area under the receiver operating characteristic curve (AUROC). Because 30-day MACE was relatively uncommon, the area under the precision–recall curve (AUPRC) was also evaluated. Overall probabilistic accuracy was assessed using the Brier score. Non-parametric bootstrap resampling of the validation cohort was used to estimate 95% CIs for AUROC, AUPRC and Brier scores. Pairwise differences in AUROC and AUPRC were evaluated using paired bootstrap resampling, preserving the paired predictions of all models within each resampled patient set. Comparisons were performed for CVLR versus HEART, XGBoost versus HEART and CVLR versus XGBoost. Two-sided bootstrap p values were calculated for each pairwise difference.
Clinical operating-point comparison
Because comparisons at arbitrary probability thresholds may favour models with different calibration characteristics or implicit safety trade-offs, clinical performance was compared at operating points with sensitivity approximately comparable to that of the conventional low-risk HEART threshold. The conventional HEART score defined patients with scores of 0–3 as low risk, 4–6 as intermediate risk and 7–10 as high risk. In the held-out validation cohort, model-specific probability thresholds for CVLR and XGBoost were selected to achieve sensitivity as close as possible to that of the conventional HEART 0–3 threshold. Because sensitivity changes discretely as individual events cross a classification threshold, these operating points were considered matched at comparable sensitivity. At each operating point, we calculated the number and proportion of patients classified as low risk, the number and observed rate of 30-day MACE among low-risk patients, sensitivity, specificity, positive predictive value and negative predictive value (NPV).
To evaluate whether the data-driven models provided additional risk stratification beyond binary low-risk classification, we also performed a post hoc analysis among patients with conventional intermediate-risk (HEART 4–6) and high-risk (HEART 7–10) scores in the held-out validation cohort. The original CVLR and XGBoost models, along with their predicted probabilities, were retained without retraining. Within each conventional HEART category, model discrimination was assessed using AUROC and AUPRC with bootstrap 95% CIs, and CVLR and XGBoost were compared using paired bootstrap differences. Among patients with HEART scores of 4–6, we additionally examined observed 30-day MACE rates across clinically interpretable model-predicted risk strata of <2%, 2–5% and ≥5%.
Statistical analysis
Continuous variables were summarised using means with SD or medians with IQRs, as appropriate. Categorical variables were summarised using counts and percentages. Comparisons between the development and held-out validation cohorts were performed using the Wilcoxon rank-sum test for HEART score and χ² tests for categorical variables.
All statistical tests were two-sided, with p<0.05 considered statistically significant. Model development and performance evaluation were conducted using Python (V.3.8), and descriptive statistical analyses were conducted using Stata (V.14; StataCorp, College Station, Texas, USA).
Reporting guideline
This study was conducted with reference to the Standards for Reporting Diagnostic Accuracy Studies (STARD) guideline and Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis (TRIPOD+AI) guideline.17 18
Results
Study population and cohort characteristics
A total of 47 853 patients met the study inclusion criteria. Of these, 39 587 patients (82.7%) from five EDs comprised the development cohort, and 8266 patients (17.3%) from the geographically held-out sixth ED comprised the validation cohort. The 30-day MACE rate was higher in the validation cohort than in the development cohort (3.23% vs 1.98%; p<0.001).
The validation cohort also had a higher median HEART score than the development cohort (3 (IQR, 2–4) vs 2 (IQR, 1–4); p<0.001) and a lower proportion of patients classified as low risk by the conventional HEART score (63.74% vs 71.08%; p<0.001). Despite these differences, the 30-day MACE rate among patients with low-risk HEART scores remained low in both cohorts (0.53% vs 0.36%; p=0.058). The distributions of individual HEART components also differed between the development and validation cohorts (table 1).
Table 1. Variables used for 30-day MACE outcome predictions.
| Training data | Held-out data | P value | |
|---|---|---|---|
| Number (%) | 39 587 (82.73) | 8266 (17.27) | |
| HEART score—median (IQR) | 2 (1–4) | 3 (2–4) | <0.001 |
| Low-risk HEART (0–3)—n (%) | 28 137 (71.08) | 5269 (63.74) | <0.001 |
| 30-day MACE—n (%) | 784 (1.98) | 267 (3.23) | <0.001 |
| 30-day MACE in low-risk HEART—n (%) | 100 (0.36) | 28 (0.53) | 0.058 |
| H-score (History) | |||
| 0 | 24 602 (87.44) | 4731 (89.79) | <0.001 |
| 1 | 3452 (12.27) | 512 (9.72) | |
| 2 | 83 (0.29) | 26 (0.49) | |
| E-score (ECG) | |||
| 0 | 22 844 (81.19) | 3569 (67.74) | <0.001 |
| 1 | 5261 (18.70) | 1688 (32.04) | |
| 2 | 32 (0.11) | 12 (0.23) | |
| A-score (Age) | |||
| 0 | 14 955 (53.15) | 2616 (49.65) | <0.001 |
| 1 | 10 910 (38.77) | 2438 (46.27) | |
| 2 | 2272 (8.07) | 215 (4.08) | |
| R-score (Risk factor) | |||
| 0 | 10 463 (37.19) | 1354 (25.70) | <0.001 |
| 1 | 15 747 (55.97) | 3441 (65.31) | |
| 2 | 1927 (6.85) | 474 (9.00) | |
| T-score (Troponin) | |||
| 0 | 27 980 (99.44) | 5219 (99.05) | 0.003 |
| 1 | 145 (0.52) | 45 (0.85) | |
| 2 | 12 (0.04) | 5 (0.09) |
HEART, History, ECG, Age, Risk factors and Troponin; MACE, major adverse cardiac events.
Overall predictive performance
In the held-out ED validation cohort, the three approaches demonstrated similar discrimination for 30-day MACE (table 2). AUROC was 0.855 (95% CI 0.836 to 0.875) for the conventional HEART score, 0.857 (95% CI 0.836 to 0.878) for CVLR and 0.856 (95% CI 0.835 to 0.877) for XGBoost. Precision–recall performance was numerically higher for the two data-driven models. AUPRC was 0.187 (95% CI 0.151 to 0.231) for HEART, 0.237 (95% CI 0.191 to 0.293) for CVLR and 0.244 (95% CI 0.196 to 0.300) for XGBoost. Overall probabilistic accuracy was similar across approaches, with Brier scores of 0.028 for all three models.
Table 2. Predictive performance and calibration of the HEART score, CVLR and XGBoost models in the held-out ED validation cohort.
| Model | AUROC (95% CI) | AUPRC (95% CI) | Brier score (95% CI) |
|---|---|---|---|
| HEART Score | 0.855 (0.836 to 0.875) | 0.187 (0.151 to 0.231) | 0.028 (0.025 to 0.031) |
| CVLR | 0.857 (0.836 to 0.878) | 0.237 (0.191 to 0.293) | 0.028 (0.025 to 0.031) |
| XGBoost | 0.856 (0.835 to 0.877) | 0.244 (0.196 to 0.300) | 0.028 (0.025 to 0.032) |
AUPRC, area under the precision–recall curve; AUROC, area under the receiver operating characteristic; CVLR, cross-validated logistic regression; ED, emergency department; HEART, History, ECG, Age, Risk factors and Troponin; XGBoost, extreme gradient boosting.
Clinical performance at comparable-sensitivity operating points
Clinical performance was compared at model-specific operating points selected to provide sensitivity approximately comparable to that of the conventional HEART low-risk threshold of 0–3 (table 3). Using HEART 0–3 as the low-risk definition, the conventional HEART score classified 5269 patients (63.74%) as low risk, among whom 28 experienced 30-day MACE, corresponding to an observed event rate of 0.53%. Sensitivity was 89.51% (95% CI 85.26% to 92.64%), specificity was 65.52% (95% CI 64.47% to 66.55%) and NPV was 99.47% (95% CI 99.23% to 99.63%). Comparable sensitivity and similarly high NPV were observed for CVLR and XGBoost (table 3). Thus, at comparable-sensitivity operating points, the three approaches identified broadly similar proportions of patients as low risk and achieved similar low-risk MACE rates and NPVs. Neither data-driven approach substantially expanded the low-risk population compared with the conventional HEART score.
Table 3. Clinical performance of the three approaches at comparable-sensitivity operating points in the held-out ED cohort.
| Model | Matched threshold | Low-risk—n (%) | MACE in low-risk—n (%) | Sen (%, 95% CI) | Spe (%, 95% CI) | PPV (%, 95% CI) | NPV (%, 95% CI) |
|---|---|---|---|---|---|---|---|
| HEART | 0–3 | 5269 (63.74) | 28 (0.53) | 89.51 (85.26 to 92.64) | 65.52 (64.47 to 66.55) | 7.97 (7.06 to 9.00) | 99.47 (99.23 to 99.63) |
| CVLR | 0.0108 | 4909 (59.39) | 25 (0.51) | 90.64 (86.54 to 93.58) | 61.06 (59.98 to 62.12) | 7.21 (6.38 to 8.13) | 99.49 (99.25 to 99.65) |
| XGBoost | 0.0128 | 5090 (61.58) | 29 (0.57) | 89.14 (84.84 to 92.33) | 63.27 (62.21 to 64.32) | 7.49 (6.63 to 8.46) | 99.43 (99.18 to 99.60) |
CVLR, cross-validated logistic regression; ED, emergency department; HEART, History, ECG, Age, Risk factors and Troponin; MACE, major adverse cardiac events; NPV, negative predictive value; PPV, positive predictive value; Sen, sensitivity; Spe, specificity; XGBoost, extreme gradient boosting.
Pairwise comparison of model performance
Paired bootstrap comparisons demonstrated no significant differences in AUROC among the three approaches (table 4). Compared with HEART, the difference in AUROC was 0.002 for CVLR (95% CI −0.005 to 0.008; p=0.604) and 0.001 for XGBoost (95% CI −0.006 to 0.008; p=0.816). AUROC also did not differ between CVLR and XGBoost (difference, −0.001; 95% CI −0.005 to 0.004; p=0.711). Both data-driven models demonstrated significantly higher AUPRC than the conventional HEART score. Compared with HEART, AUPRC was higher by 0.050 for CVLR (95% CI 0.033 to 0.071; p<0.001) and by 0.057 for XGBoost (95% CI 0.036 to 0.082; p<0.001). However, the difference in AUPRC between CVLR and XGBoost was not statistically significant (difference, 0.008; 95% CI −0.006 to 0.021; p=0.305). These findings indicate that data-driven reweighting of the five HEART components improved precision–recall performance relative to the conventional fixed-point score, but increasing model complexity from logistic regression to XGBoost did not provide measurable incremental improvement in discrimination, precision–recall performance or clinical low-risk classification.
Table 4. Paired bootstrap comparisons of discrimination and precision–recall performance among the three prediction models.
| Comparison | Δ AUROC (95% CI) | P value | Δ AUPRC (95% CI) | P value |
|---|---|---|---|---|
| CVLR vs HEART | 0.002 (−0.005 to 0.008) | 0.604 | 0.050 (0.033 to 0.071) | <0.001 |
| XGBoost vs HEART | 0.001 (−0.006 to 0.008) | 0.816 | 0.057 (0.036 to 0.082) | <0.001 |
| CVLR vs XGBoost | −0.001 (−0.005 to 0.004) | 0.711 | 0.008 (−0.006 to 0.021) | 0.305 |
AUPRC, area under the precision–recall curve; AUROC, area under the receiver operating characteristic; CVLR, cross-validated logistic regression; HEART, History, ECG, Age, Risk factors and Troponin; XGBoost, extreme gradient boosting.
Risk stratification within non-low-risk HEART categories
Among 2705 patients with conventional HEART scores of 4–6, 165 (6.10%) experienced 30-day MACE. Both data-driven models further differentiated risk within this intermediate-risk group. For CVLR, observed MACE rates were 3.06%, 3.97% and 12.17% among patients with model-predicted risks of<2%, 2–5% and ≥5%, respectively. Corresponding rates for XGBoost were 3.80%, 8.14% and 20.53%. Within HEART 4–6, AUROC was 0.674 (95% CI 0.629 to 0.722) for CVLR and 0.667 (95% CI 0.620 to 0.717) for XGBoost, with AUPRCs of 0.155 (95% CI 0.116 to 0.210) and 0.156 (95% CI 0.117 to 0.215), respectively. Neither AUROC nor AUPRC differed significantly between the two models (online supplemental table S1).
Among 292 patients with HEART scores of 7–10, 74 (25.34%) experienced MACE. Within this high-risk group, AUROC was 0.695 (95% CI 0.625 to 0.764) for CVLR and 0.676 (95% CI 0.604 to 0.746) for XGBoost; corresponding AUPRCs were 0.445 (95% CI 0.346 to 0.571) and 0.462 (95% CI 0.354 to 0.577). Pairwise differences between the models were again not statistically significant. These findings indicate that data-driven modelling provided additional graded risk information within conventional HEART categories, while greater non-linear model complexity did not improve overall within-category performance.
Discussions
In this multicentre study, we evaluated whether increasing model complexity improved 30-day MACE prediction when all models were restricted to the same clinical information as the HEART score. Using five EDs for model development and a geographically held-out sixth ED for validation, we found that the conventional HEART score, CVLR and XGBoost demonstrated nearly identical AUROC and Brier scores. Both data-driven models achieved significantly higher AUPRC than the conventional HEART score; however, XGBoost did not significantly outperform CVLR in either AUROC or AUPRC. At operating points with comparable sensitivity, neither CVLR nor XGBoost increased the proportion of patients classified as low risk. These findings suggest that data-driven reweighting of HEART components may improve risk ranking, but increasing model complexity alone provides limited incremental clinical value when no additional predictor information is introduced.
Our findings highlight the distinction between clinical information content and modelling complexity. The conventional HEART score combines five clinical domains using fixed integer points, whereas CVLR allows separate data-driven effects for each component level and XGBoost additionally accommodates non-linear relationships and interactions. Despite this increasing flexibility, overall discrimination remained similar. This finding is consistent with the broader observation that apparent improvements from ML may reflect either more extensive predictor information or more flexible modelling.19 20 By holding predictor information constant, our study suggests that, for HEART-based chest pain risk stratification, flexible reweighting may capture much of the available predictive information without requiring additional non-linear model complexity.
The higher AUPRCs of CVLR and XGBoost suggest that data-driven weighting may improve risk ranking across a broader spectrum of predicted risk, even when binary low-risk classification is not expanded. Because 30-day MACE was uncommon, precision–recall performance provides complementary information to AUROC by emphasising performance for the event class.21 This improved risk ranking did not translate into a larger safely classified low-risk population at comparable sensitivity. Instead, maintaining comparable sensitivity was accompanied by a modest reduction in the proportion classified as low risk with CVLR (59.4%) and XGBoost (61.6%) compared with HEART (63.7%), representing a potential trade-off between improved overall risk ranking and low-risk classification efficiency at the examined operating points. Additionally, such improvement may be clinically relevant beyond binary rule-out classification, because chest pain risk stratification also informs graded management decisions across the non-low-risk spectrum, including observation, additional diagnostic testing, expedited outpatient evaluation and cardiology assessment. However, among patients with HEART scores 4–6, both data-driven models identified subgroups with substantially different observed 30-day MACE rates while retaining moderate within-category discrimination. Both models also retained discrimination among patients with HEART scores 7–10, although XGBoost did not significantly outperform CVLR in either subgroup. These findings suggest that flexible reweighting may provide additional graded risk information beyond the conventional HEART categories, while greater non-linear model complexity offers little measurable incremental benefit.
The comparable performance of CVLR and XGBoost also has implications for model selection. CVLR relaxes the fixed weighting assumptions of the conventional HEART score while retaining a relatively transparent additive structure.22 XGBoost can represent more complex interactions and non-linear relationships but at the cost of greater complexity and reduced transparency.23 In the absence of demonstrable improvement in discrimination, precision–recall performance or low-risk classification, additional model complexity may be difficult to justify solely on statistical grounds.
The site-level holdout provided a more stringent assessment of within-system geographical transportability than a random patient-level split, because the validation ED contributed no data to model fitting and differed from the development cohort in case mix, HEART score distributions and MACE prevalence. Nevertheless, this should not be considered fully external validation because all six EDs belonged to the same integrated healthcare system and shared an EHR infrastructure and a general clinical environment. Validation across independent health systems and more diverse populations remains necessary to establish broader generalisability.
We deliberately restricted the data-driven models to the same five HEART components to isolate the effect of modelling complexity from that of additional predictor information. This design does not imply that additional clinical information lacks predictive value. Contemporary approaches such as Collaboration for the Diagnosis and Evaluation of Acute Coronary Syndrome (CoDE-ACS) have already demonstrated the potential value of incorporating more detailed and continuous clinical information for individualised chest pain risk assessment.24 Instead, our findings suggest that future improvements may depend more on incorporating informative predictors, such as serial troponin measurements, detailed ECG features, medications and longitudinal EHR data, than on applying increasingly complex algorithms to the same five components. Future work should determine whether such additional information provides transportable improvements in prediction and clinical decision-making.
Several limitations should be considered. First, the retrospective observational design is subject to selection bias and misclassification, and inclusion was limited to patients with documented HEART scores. The History component is subjective and may vary across clinicians and sites, and all models remain dependent on the quality of the underlying clinical assessments. Future studies should evaluate the inter-rater reliability of the History component and determine how variability in its assessment affects the performance and transportability of HEART-based risk models. Second, patients transferred directly to another facility, those undergoing emergent cardiac catheterisation or those dying during the index ED encounter were excluded because HEART-based disposition risk stratification is less relevant for these immediately recognised high-risk presentations. However, this narrows the population to which our findings apply. Third, the geographically held-out validation ED remained within the same health system, limiting the generalisability of the conclusions. Fourth, EHR-based and administrative-based outcome ascertainment may miss events that occur outside the available data networks. Fifth, 30-day MACE was evaluated as a composite outcome; its individual components differ in clinical importance and mechanism, and their relatively small event counts precluded reliable separate modelling. Finally, restricting the models to the five HEART components constrained the potential predictive performance. Accordingly, our results should not be interpreted as evidence that ML has no role in chest pain risk stratification; rather, they suggest that algorithmic complexity without additional informative predictors may offer limited incremental value.
Conclusion
In this multicentre development and geographically held-out ED validation study, data-driven reweighting of the five HEART components improved precision–recall performance compared with conventional fixed-point HEART scoring, suggesting more effective risk ranking across the spectrum of predicted risk. However, increasing model complexity from categorical CVLR to XGBoost provided no incremental benefit, and neither data-driven approach expanded low-risk classification at comparable sensitivity. These findings suggest that flexible weighting of established predictors may improve risk stratification, whereas additional non-linear model complexity offers limited incremental value when predictor information is held constant. Future work should focus on identifying additional clinically informative predictors that provide transportable incremental value and on determining whether improvements in statistical performance translate into better risk-informed clinical management.
Supplementary material
Footnotes
Funding: The authors have not declared a specific grant for this research from any funding agency in the public, commercial or not-for-profit sectors.
Provenance and peer review: Not commissioned; externally peer reviewed.
Patient consent for publication: Not applicable.
Data availability free text: The dataset and code used and analysed in the current study are available from the corresponding author upon reasonable request.
Ethics approval: This study has been carried out in accordance with The Code of Ethics of the World Medical Association (Declaration of Helsinki) for studies involving human subjects. This study was approved by the University of North Texas Health Science Center regional Institutional Review Board, with a full waiver of participation and informed consent from human subjects (IRB No.1541042-5).
Data availability statement
Data are available upon reasonable request.
References
- 1.Gulati M, Levy PD, Mukherjee D, et al. 2021 AHA/ACC/ASE/CHEST/SAEM/SCCT/SCMR guideline for the evaluation and diagnosis of chest pain: a report of the American College of Cardiology/American Heart Association Joint Committee on Clinical Practice Guidelines. Circulation. 2021;144:e368–454. doi: 10.1161/CIR.0000000000001029. [DOI] [PubMed] [Google Scholar]
- 2.Kawatkar AA, Sharp AL, Baecker AS, et al. Early noninvasive cardiac testing after emergency department evaluation for suspected acute coronary syndrome. JAMA Intern Med. 2020;180:1621–9. doi: 10.1001/jamainternmed.2020.4325. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Six AJ, Backus BE, Kelder JC. Chest pain in the emergency room: value of the HEART score. Neth Heart J. 2008;16:191–6. doi: 10.1007/BF03086144. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Meyering SH, Schrader CD, Kumar D, et al. Role of HEART score in evaluating clinical outcomes among emergency department patients with different ethnicities. J Int Med Res. 2021;49 doi: 10.1177/03000605211010638. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Laureano-Phillips J, Robinson RD, Aryal S, et al. HEART score risk stratification of low-risk chest pain patients in the emergency department: a systematic review and meta-analysis. Ann Emerg Med. 2019;74:187–203. doi: 10.1016/j.annemergmed.2018.12.010. [DOI] [PubMed] [Google Scholar]
- 6.Sharp AL, Wu Y-L, Shen E, et al. The HEART score for suspected acute coronary syndrome in U.S. emergency departments. J Am Coll Cardiol. 2018;72:1875–7. doi: 10.1016/j.jacc.2018.07.059. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Green SM, Schriger DL. A methodological appraisal of the HEART score and its variants. Ann Emerg Med. 2021;78:253–66. doi: 10.1016/j.annemergmed.2021.02.007. [DOI] [PubMed] [Google Scholar]
- 8.Tomaszewski CA, Nestler D, American College of Emergency Physicians Clinical Policies Subcommittee (Writing Committee) on Suspected Non–ST-Elevation Acute Coronary Syndromes Clinical policy: critical issues in the evaluation and management of emergency department patients with suspected non-ST-elevation acute coronary syndromes. Ann Emerg Med. 2018;72:e65–106. doi: 10.1016/j.annemergmed.2018.07.045. [DOI] [PubMed] [Google Scholar]
- 9.Leopold JA, Loscalzo J. Emerging role of precision medicine in cardiovascular disease. Circ Res. 2018;122:1302–15. doi: 10.1161/CIRCRESAHA.117.310782. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Wang Y, Aivalioti E, Stamatelopoulos K, et al. Machine learning in cardiovascular risk assessment: towards a precision medicine approach. Eur J Clin Invest. 2025;55 Suppl 1:e70017. doi: 10.1111/eci.70017. [DOI] [PubMed] [Google Scholar]
- 11.Armoundas AA, Narayan SM, Arnett DK, et al. Use of artificial intelligence in improving outcomes in heart disease: a scientific statement from the American Heart Association. Circulation. 2024;149:e1028–50. doi: 10.1161/CIR.0000000000001201. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Khera R, Oikonomou EK, Nadkarni GN, et al. Transforming cardiovascular care with artificial intelligence: from discovery to practice. J Am Coll Cardiol. 2024;84:97–114. doi: 10.1016/j.jacc.2024.05.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Bouzid Z, Sejdic E, Martin-Gill C, et al. Electrocardiogram-based machine learning for risk stratification of patients with suspected acute coronary syndrome. Eur Heart J. 2025;46:943–54. doi: 10.1093/eurheartj/ehae880. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Chou EH, Lu T-C, Chiu Y-T, et al. Using machine learning to risk stratify emergency department patients with chest pain but no acute myocardial infarction: a multicenter retrospective analysis. J Am Heart Assoc. 2025;14:e041915. doi: 10.1161/JAHA.125.041915. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Alkhamis MA, Al Jarallah M, Attur S, et al. Interpretable machine learning models for predicting in-hospital and 30 days adverse events in acute coronary syndrome patients in Kuwait. Sci Rep. 2024;14:1243. doi: 10.1038/s41598-024-51604-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Yu M-Y, Yoo HY, Han GI, et al. Comparing the performance of machine learning models and conventional risk scores for predicting major adverse cardiovascular cerebrovascular events after percutaneous coronary intervention in patients with acute myocardial infarction: systematic review and meta-analysis. J Med Internet Res. 2025;27:e76215. doi: 10.2196/76215. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi: 10.1136/bmj-2023-078378. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Bossuyt PM, Reitsma JB, Bruns DE, et al. STARD 2015: an updated list of essential items for reporting diagnostic accuracy studies. BMJ. 2015;351:h5527. doi: 10.1136/bmj.h5527. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Steele AJ, Denaxas SC, Shah AD, et al. Machine learning models in electronic health records can outperform conventional survival models for predicting patient mortality in coronary artery disease. PLoS One. 2018;13:e0202344. doi: 10.1371/journal.pone.0202344. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Christodoulou E, Ma J, Collins GS, et al. A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models. J Clin Epidemiol. 2019;110:12–22. doi: 10.1016/j.jclinepi.2019.02.004. [DOI] [PubMed] [Google Scholar]
- 21.Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One. 2015;10:e0118432. doi: 10.1371/journal.pone.0118432. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Shipe ME, Deppen SA, Farjah F, et al. Developing prediction models for clinical use using logistic regression: an overview. J Thorac Dis. 2019;11:S574–84. doi: 10.21037/jtd.2019.01.25. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Al’Aref SJ, Maliakal G, Singh G, et al. Machine learning of clinical variables and coronary artery calcium scoring for the prediction of obstructive coronary artery disease on coronary computed tomography angiography: analysis from the CONFIRM registry. Eur Heart J. 2020;41:359–67. doi: 10.1093/eurheartj/ehz565. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Boeddinghaus J, Doudesis D, Lopez-Ayala P, et al. Machine learning for myocardial infarction compared with guideline-recommended diagnostic pathways. Circulation. 2024;149:1090–101. doi: 10.1161/CIRCULATIONAHA.123.066917. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Data are available upon reasonable request.
