Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Aug 13.
Published before final editing as: J Card Fail. 2026 Jul 15:S1071-9164(26)00400-8. doi: 10.1016/j.cardfail.2026.07.003

Deep learning interpretation of echocardiographic images predicts incident heart failure and subtypes

Emily S Lau 1,2,3,4,5,*, Tal Shnitzer 5,*, Athar Roshandelpoor 3,7, Samuel Friedman 5, Carl T Andrews 1,3, Mostafa Al-Alusi 1,2,3,5,6, Jonathan W Cunningham 1,3,5, Arash Nargesi 1,3,5, Steven A Lubitz 1,3,4,5, Michael H Picard 1,3, Shaan Khurshid 1,2,3,4,5,6, Mahnaz Maddah 5, Patrick T Ellinor 1,2,3,4,5,6, Jennifer E Ho 3,5,7
PMCID: PMC13463178  NIHMSID: NIHMS2196559  PMID: 42457081

Abstract

Background:

Accurate prediction of incident heart failure (HF) may help prioritize HF preventive therapies. Deep learning interpretation of echocardiograms may improve HF risk prediction beyond clinical risk models. We trained and validated a deep learning model to predict incident HF from transthoracic echocardiographic images (Echocardiogram-to-Heart Failure, “Echo2HF”)

Methods:

Echo2HF was developed using 4,057,664 echocardiogram videos from 70,763 patients receiving longitudinal ambulatory care at Massachusetts General Hospital (MGH). Performance for 10-year incident HF was evaluated in an internal MGH test set and an external test set of 34,802 individuals without prevalent HF from Brigham and Women’s Hospital (BWH). Model performance was evaluated using area under the receiver operating characteristic curves (AUROC) and compared with the Pooled Cohorts Equations to Prevent Heart Failure (PCP-HF) and the Predicting Risk of cardiovascular disease EVENTs (PREVENT) clinical risk scores.

Results:

Echo2HF was trained in 64,167 individuals and evaluated in a hold-out sample of 6,394 individuals from MGH (279 HF events, age 62 ± 17 years, 48% women) and 34,802 individuals from BWH (1280 events, age 62 ± 15 years, 56% women). Echo2HF discriminated incident HF, with 10-year AUROC of 0.83 (95% CI0.81–0.85] and 0.82 (95% CI 0.81–0.83) at BWH, with numerically higher discrimination versus both PCP-HF and PREVENT.

Conclusion:

Deep learning analysis of echocardiograms accurately discriminated future HF risk, with favorable performance over current clinical HF scores. Future work should assess whether broader use of artificial intelligence-enabled echocardiographic risk stratification may improve HF prevention and clinical outcomes, including among individuals who do not have a clinical indication for echocardiography.

Introduction

Heart failure (HF) affects over 64 million individuals worldwide, with prevalence and costs continuing to rise.1,2,3 The emergence of HF preventive therapies, including sodium-glucose cotransporter-2 (SGLT2) inhibitors and incretin-based therapies including glucagon-like peptide-1 receptor agonists (GLP-1RA) and glucose dependent insulinotropic polypeptide (GIP)/GLP-1R dual agonists, has heightened the need to accurately identify high-risk individuals for targeted prevention.4 Reflecting this unmet need, HF prevention was included as a key provision for the first time in the 2022 American Heart Association (AHA)/American College of Cardiology (ACC)/Heart Failure Society of America (HFSA) Guidelines for the Management of Heart Failure.4

Several clinical models have been developed and validated for the prediction of incident HF, but adoption in practice has been limited by the need for multiple clinical variables and heterogeneous performance across external validation samples.57 Current HF guidelines also identify at-risk individuals with ‘pre-HF’ or stage B HF based on echocardiographic evidence of abnormal cardiac structure or function, but reproducible ascertainment echocardiographic measures at scale remain challenging. Deep learning echocardiographic analysis can reliably assess cardiac structure and function, detect myocardial diseases, and predict systemic cardiovascular risk factors.811 Recent studies suggest that deep learning approaches applied to transthoracic echocardiogram (TTE) images may capture latent information relevant to HF risk prediction beyond traditional clinical variables.1214

We previously demonstrated that deep learning-estimated left heart measures were associated with future clinical outcomes including HF.15 Here, we developed Echocardiogram-to-Heart Failure (Echo2HF), a deep learning model applied to clinically obtained echocardiographic images to directly estimate future HF risk using two electronic health record (EHR) ambulatory care samples comprising 105,565 individuals with longitudinal follow-up.

Methods

Study Cohorts.

Training sample.

The Community Care Cohort Project (C3PO, n=523,445) and Enterprise Warehouse of Cardiology (EWOC, n=99,253) cohorts are retrospective ambulatory primary care and cardiology EHR cohorts within the Mass General Brigham (MGB) network, respectively.16,17 Both cohorts have been previously described. Briefly, adults aged 18–90 years with ≥2 office visits between 1–3 years apart in a qualifying clinic between 2000–2019 were identified from a common EHR. Echo2HF was trained and validated on the subset of individuals in the C3PO and EWOC cohorts who had ≥1 transthoracic echocardiogram (TTE) performed at Massachusetts General Hospital (MGH). Individuals with prevalent HF were included in the training set to enrich learning of HF-related echocardiographic phenotypes.

Test samples.

Echo2HF was then evaluated among individuals without prevalent HF in two independent test samples: a randomly selected MGH internal test set not used in model training (MGH Internal Test) and in an external sample of C3PO participants who underwent ≥1 TTE at Brigham and Women’s Hospital (BWH) (BWH External Test), a separate institution within the same integrated MGB healthcare system. Individuals who underwent TTE at both MGH and BWH were included in MGH only. All partitions were performed at the patient level, with no individual-level overlap between the training and test sets. The MGB Institutional Review Board approved both C3PO and EWOC study protocols.

Incident heart failure

The primary outcome was incident HF. HF was defined using a previously validated approach that combines (1) ≥1 International Classification of Diseases, 9th or 10th revision (ICD 9/10) code for HF and (2) HF classification using a transformer-based NLP model trained to identify HF hospitalization from free text discharge summaries.18 This definition has been validated against expert adjudication and clinical trial clinical endpoint committees.1820

For HF subtype analyses, incident HF events were categorized as HF with preserved ejection fraction (HFpEF, LV ejection fraction [LVEF] ≥50%), HF with reduced ejection fraction (HFrEF, LVEF<50%), or unclassified when no LV function assessment available. LVEF was ascertained from echocardiograms within 1 year of the HF date, prioritizing studies within 6 months after diagnosis, then 6 months before diagnosis, 6–12 months after diagnosis, and 6–12 months before diagnosis.

Echocardiogram data processing and Echo2HF architecture

Details of TTE acquisition, data processing, de-identification, and view classification model have been previously reported and are summarized in the Supplemental Methods.15 Echo2HF is a three-dimensional (3D) convolutional neural network (CNN) based on the Dimensional Reconstruction of Imaging Data (DROID) architecture and designed to predict 10-year incident HF from TTE videos. The model uses a loss function that accounts for survival time and censoring, with time to HF is defined from TTE date. In the MGH training set only, prevalent HF was encoded as a time-to-event of zero (Supplemental Methods).15,21

Parasternal long-axis, apical 4-chamber, and apical 2-chamber views were used. Each video was represented as 224×224×16 matrix, retaining every fourth frame and truncating after 16 retained frames. The model generated a prediction for each individual video, and the final study-level prediction was defined as the median of the video-level predictions across all videos within that TTE study. To optimize model performance, Echo2HF was trained using incident HF as the primary target and age and sex as secondary prediction targets, an approach that modestly outperformed a model without secondary targets (Table S1).22 The final Echo2HF model architecture comprises incident HF, age, and sex predictions (Figure S1). In secondary analyses, we developed Echo2HF-5 year to predict 5-year incident HF from the TTE. Saliency maps were generated to identify image regions with greatest influence on model predictions (Supplemental Methods).

All MGH TTEs were randomly split at the patient level into independent training (3,861,653 echocardiogram videos/n individuals=54,305), validation (117,949 echocardiogram videos/n=9,862), and testing (internal test: 78,062 echocardiogram videos/n=6,596) sets. The TTEs performed at BWH were used as external testing (314,164 echocardiogram videos/n=34,802) (Figure 1). Multiple TTEs per patient were used during training; all validation and test analyses were restricted to individuals without prevalent HF at the index TTE.

Figure 1. Development and test samples.

Figure 1.

Depicted is an overview of the training, validation, and test sets. Echo2HF was derived in the Massachusetts General Hospital (MGH) development set and tested in an internal sample within MGH and an external hospital sample (Brigham and Women’s Hospital, BWH). *includes individuals with no HF, incident HF, and prevalent HF. ** excludes prevalent HF.

Comparison with established clinical risk scores

Echo2HF was compared with two validated 10-year HF risk scores: the Pooled Cohorts Equations to Prevent Heart Failure (PCP-HF) and the Predicting Risk of cardiovascular disease EVENTs (PREVENT-HF).5,23 Because electrocardiogram data were frequently missing, we used a modified version of the PCP-HF equations without QRS duration for comparison.24 Clinical factors required for PCP-HF and PREVENT scores were ascertained from EHR data within 3 years prior to and up to 1 year following the TTE study date (definitions listed in Table S2). When fasting glucose was unavailable, glycated hemoglobin was converted to estimated average blood glucose levels using published American Diabetes Association equations and calibrated to a fasting level.25 Comparisons with PCP-HF and PREVENT used a complete case approach restricted to individuals aged 30–79 years.

Statistical analysis

Follow-up began at the date of index TTE study date and continued until incident HF, death, last clinical encounter, or age 90. Death was ascertained from the Social Security Death Index or the MGB internal documentation of death. To reduce misclassification of prevalent or clinically evolving HF as potential incident HF, HF events within 30 days after the TTE date were considered prevalent and excluded from incident analyses. Sensitivity analyses extended this blanking period to 90 days.

Model discrimination was evaluated using time-dependent AUROC and average precision (AP) and integrated calibration indices (ICI), with 10-year AUROC as the primary metric. Because incident HF was uncommon, time-dependent AP was examined as a complementary measure of precision and recall. Inverse probability of censoring weights were used for both AUROC and AP.26 Calibration was assessed using observed-versus-predicted risk plots with adaptive hazard regression and integrated calibration index (ICI) (primary calibration metric). Curves were fit using default parameters, with spline complexity limited when event counts were low. For AP, AUROC, and ICI estimates, confidence intervals were estimated using bootstrapping. We compared cumulative HF incidence among individuals above versus below the median Echo2HF risk score. Because clinically actionable HF risk thresholds have not been established, the median HF risk score was used to illustrate risk separation.

In secondary analyses, we evaluated Echo2HF-5-year model discrimination and calibration. In exploratory analyses, we performed a subgroup analysis examining model performance by HF subtype (HFrEF and HFpEF) and evaluated performance of models combining Echo2HF with PCP-HF and PREVENT probabilities (Supplemental Methods). We also calculated continuous and category-based net reclassification improvement (NRI) using 5-year HF risk thresholds of <5%, 5–10%, and >10%.” Finally, we performed sensitivity analyses, including assessment of Echo2HF performance among individuals with vs without common cardiac comorbidities and echocardiograms with normal (LVEF ≥ 50%) vs impaired (LVEF <50%) LVEF.

All tests were 2-sided and p values <0.05 were considered significant. Analyses were performed using Python v3.6.9 with Tensorflow, Pandas, and Scikit-learn. The study adheres to the PRIME (Proposed Requirements for Cardiovascular Imaging-Related Machine Learning Evaluation) guidelines (Supplemental Methods).

Data availability

MGH and BWH data contain protected health information and are not publicly available because individuals did not consent to data sharing. Source code and model weights for Echo2HF are available at: https://github.com/broadinstitute/ml4h/tree/master/model_zoo/Echo2HF

Results

Echo2HF was trained, validated, and internally tested on 70,763 patients with echocardiograms from MGH, including an MGH Internal Test set of 6,596 individuals with mean age 62 years, 48% women. The BWH External Test set comprised 34,802 individuals with mean age 62 years, 56% women. Baseline characteristics were similar across test sets (Table 1).

Table 1.

Baseline Characteristics of Training, Validation, and Test Sets

MGH Training (n=54,305) MGH Validation (n=9,862) MGH Internal Test (n=6,596) BWH External Test (n=34,802)
No. of echocardiogram studies 132,577 19,880 13,449 71,397
No. of echocardiogram videos 3,861,653 117,949 78,062 314,164
Clinical characteristics
Age, years 62.8±16.9 62.8±16.9 61.6±16.9 61.9±15.2
Female, n (%) 25,902 (48%) 4,804 (49%) 3,191 (48%) 19,609 (56%)
Race and ethnicity, n (%)
 Asian or Pacific Islander 1,696 (3%) 308 (3%) 202 (3%) 881 (2%)
 Black 2,442 (5%) 453 (5%) 278 (4%) 3,362 (10%)
 Hispanic or Latino 1,217 (2%) 216 (2%) 139 (2%) 1,704 (5%)
 Other or Unknownb 5,567 (10%) 1,066 (11%) 599 (9%) 2,151 (6%)
 White 43,386 (80%) 7,819 (79%) 5,378 (82%) 26,704 (77%)
Body mass index, kg/m2 30.2±6.9 30.4±7.3 30.2±6.8 30.0±7.4
Systolic blood pressure, mmHg 129.5±18.6 129.6±18.2 130.4±18.7 132.5±18.9
Diabetes, n (%) 7,172 (13%) 822 (8%) 553 (8%) 3,506 (10%)
Hyperlipidemia, n (%) 15,546 (29%) 2,420 (25%) 1,626 (25%) 8,296 (24%)
Myocardial infarction, n (%) 4,915 (9%) 806 (8%) 553 (8%) 4,024 (11%)
Atrial fibrillation, n (%) 10,154 (19%) 1,711 (17%) 1,243 (19%) 5,536 (16%)
Current smoker, n (%) 1,636 (3%) 224 (2%) 145 (2%) 800 (2%)
Hypertension medication, n (%) 14,009 (26%) 1,740 (18%) 1,153 (17%) 7,769 (22%)
Glucose-lowering medication, n (%) 4,555 (8%) 520 (5%) 358 (5%) 2,282 (7%)
Echocardiographic measures
LV ejection fraction, % 60 ± 15 63 ± 12 63 ± 12 58 ± 11
LV end diastolic diameter, mm 46 ± 7 45 ± 7 45 ± 7 45 ± 7
IVS thickness, mm 11 ± 3 10 ± 2 11 ± 3 11 ± 3
PWT thickness, mm 10 ± 2 10 ± 2 10 ± 2 10 ± 2
LA dimension, mm* 39 ± 7 37 ± 6 37 ± 6 36 ± 7

Results are displayed as means ± standard deviation or n (%) unless otherwise noted.

Abbreviations: IVS = interventricular septal, LA = left atrial, LV = left ventricular, PWT = posterior wall thickness

a-Refinement and validation cohorts exclude people with prevalent heart failure diagnosis.

b-Other is defined as either mixed, unknown, or specified race not included in other categories.

*

LA dimension was available for n=26,457 in the training set, n=3,687 in the refinement set, n=2,557 in the MGH internal test set, and n=9,706 in the BWH external test set.

Echo2HF Model Discrimination of 10-year Incident Heart Failure

Over 10 years of follow-up, there were 279 incident HF events in the MGH Test set (10-year cumulative risk 9.1%, 95% CI 7.9–10.4%) and 1,280 events in the BWH Test set (8.5%, 95% CI 7.9–9.1%).

Echo2HF consistently discriminated 10-year incident HF risk in both test sets with AUROC [95% CI] of 0.84 [0.81–0.87] in MGH Test and 0.84 [0.82–0.85] in BWH Test (p=0.04 versus age and sex for both) and AP of 0.28 [0.24–0.35] in MGH Test and 0.26 [0.24–0.29] in BWH Test (p<0.001 versus age and sex) (Table 2 and Figures 2 & 3). Echo2HF also stratified absolute HF risk. In the BWH test set, individuals with predicted risk above the median had 10-year incidence of 16.4%, compared with 1.2% among those below the median (Figure 2). Higher Echo2HF risk was associated with substantially greater future HF risk in the BWH Test (HR 15.8, 95% CI 12.7–19.7, p-value <0.001). Echo2HF performed well across key subgroups, including age, sex, cardiovascular risk factor burden, and baseline LVEF (Figures S2-5). Results were similar using a 90-day blanking period after the index TTE (Table S3).

Table 2.

Model discrimination and calibration performance

MGH Internal Test (n=6,596) BWH External Test (n=34,802)
AUROC [95% CI] AP [95% CI] ICI [95% CI] AUROC [95% CI] AP [95% CI] ICI [95% CI]
Overall
Echo2HF 10y [95% CI] 0.84 [0.81–0.87] 0.28 [0.24–0.35] 0.05 [0.031–0.060] 0.84 [0.82–0.85] 0.26 [0.24–0.29] 0.04 [0.03–0.04]
Echo2HF 5y [95% CI] 0.86 [0.84–0.88] 0.19 [0.16–0.23] 0.05 [0.04–0.06] 0.86 [0.85–0.87] 0.20 [0.18–0.22] 0.03 [0.03–0.04]
Age and sex 10y 0.80 [0.76, 0.83] 0.13 [0.11, 0.16] 0.017 [0.007, 0.029] 0.73 [0.71, 0.75] 0.11 [0.10–0.12] 0.01 [0.01–0.02]
PCP-HF and PREVENT Comparison *
Echo2HF 10y [95% CI] 0.86 [0.75–0.94] 0.28 [0.23–0.34] 0.03 [0.01–0.06] 0.81 [0.77–0.85] 0.34 [0.28–0.42] 0.02 [0.01–0.04]
PCP-HF [95% CI] 0.77 [0.66–0.87] 0.38 [0.23–0.68] 0.0 4 [0.01–0.11] 0.75 [0.69–0.80] 0.19 [0.15–0.25] 0.03 [0.02–0.07]
PREVENT [95% CI] 0.79 [0.68–0.88] 0.31 [0.21–0.55] 0.05 [0.01–0.11] 0.77 [0.72–0.82] 0.18 [0.15–0.22] 0.021 [0.01–0.08]

Abbreviations: AUROC = area under the receiver operating curve, AP = average precision, ICI = integrated calibration index. The ICI is a measure of the average prediction error weighted by the empirical risk distribution and where an ICI of zero indicates perfect estimates.20

*

Echo2HF performance was compared with PCP-HF and PREVENT-HF in n=1,116 individuals from MGH and n=7,914 individuals from BWH with complete data for PCP-HF and PREVENT score calculations.

Figure 2. Echocardiogram to Heart Failure (Echo2HF).

Figure 2.

Echo2HF is a 3D convolutional neural network designed to estimate future risk of heart failure (HF) directly from the transthoracic echocardiogram. HF events were defined using a validated natural language processing (NLP) algorithm shown to adjudicate HF events with superior performance compared to diagnosis codes alone. Echo2HF showed consistent HF risk stratification performance in two test sets separate from model training: Massachusetts General Hospital (MGH) and Brigham and Women’s Hospital (BWH).

Figure 3. Echo2HF discrimination of incident heart failure.

Figure 3.

Figure 3 depicts model discrimination for Echo2HF (green) and the age and sex model (gray), measured using the time-dependent area under the receiver operating characteristic curve (AUROC, top 2 panels), and average precision (bottom 2 panels) and in the test sets. A model with no discriminative power (“No Skill”), or the performance observed when all individuals are classified as cases, is depicted by black triangles in the bottom panels. The x-axis depicts an increasing prediction window.

In HF subtypes analyses, Echo2HF discriminated 10-year incident HF risk for both future HFrEF (MGH Test: AUROC [95% CI] 0.85 [0.80–0.88], AP 0.11 [0.08–0.18], BWH Test: AUROC 0.86 [0.84–0.87], AP 0.19 [0.16–0.24]) and HFpEF (MGH Test: AUROC 0.84 [0.80–0.87], AP 0.20 [0.16–0.26], BWH Test: AUROC 0.82 [0.81–0.84], AP 0.17 [0.15–0.18]) (Table S4 and Figure S6).

Calibration of Echo2HF Model for Incident Heart Failure

Median Echo2HF risk was similar in the MGH and BWH test sets (10-year estimated HF risk, MGH: 5.2% [Q1 to Q3: 1.1% to 19.6%], BWH: 4.0% [Q1 to Q3: 0.9% to 15.8%, Figure S7). Echo2HF demonstrated good calibration, with ICI values consistent with low estimation error in both MGH and BWH Test sets (Table 2). Although overall calibration was favorable, Echo2HF tended to overestimate risk at the high end of the predicted risk distribution (Figure 4). Echo2HF was also well-calibrated for 10-year prediction of HF in both HFrEF and HFpEF, with similar estimation errors for HFrEF (ICI 0.08 [0.07–0.10]) and HFpEF (ICI 0.06 [0.05–0.06]) (Table S4 and Figure S8).

Figure 4. Echo2HF calibration of incident heart failure risk.

Figure 4.

Figure 4 depicts model calibration using Echo2HF (green) versus age and sex (gray), where the x-axis depicts predicted heart failure risk, and the y-axis depicts observed heart failure incidence in the test sets. Curves were fit using adaptive hazard regression. Perfect calibration is depicted by the hashed diagonal line. The integrated calibration index, an estimate of model error where lower values depict better calibration, is listed next to each curve.

Comparison to PCP-HF and PREVENT Scores

Echo2HF was compared with PCP-HF and PREVENT-HF in 1,116 individuals from MGH and 7,914 individuals from BWH with complete data for PCP-HF and PREVENT score calculations. In this subset, there were 66 incident HF events in the MGH Test (10-year cumulative risk 15.4%, 95% CI 9.3–21.2%) and 414 events in the BWH Test (11.1%, 95% CI 9.2–13.0%) over 10 years of follow-up.

Echo2HF demonstrated favorable performance relative to the PCP-HF and PREVENT scores. In BWH, Echo2HF had significantly higher AUROC than PCP-HF (difference 0.07, 95% CI 0.01–0.13, p =0.02). The difference relative to PREVENT was smaller and not statistically significant (difference 0.04, 95% CI −0.02–0.1, p=0.17). Echo2HF had higher AP than both PCP-HF and PREVENT (difference vs PCP-HF: 0.16, 95% CI 0.097–0.22, p<0.001; vs PREVENT: difference 0.16, 95% CI 0.11–0.22, p<0.001, Figure 5). Combined models incorporating Echo2HF with PCP-HF or PREVENT moderately higher discrimination than Echo2HF alone (Table S5).

Figure 5. Model discrimination comparison between Echo2HF, PCP-HF and PREVENT.

Figure 5.

Figure 5 displays comparisons between Echo2HF (green) and the PCP-HF (orange) and PREVENT (lavender) scores in the MGH (left) and BWH (right) test sets for the discrimination of 10-year incident heart failure (among n=1,116 individuals from MGH and n=7,914 individuals from BWH with complete data for PCP-HF and PREVENT score calculations). The top panels depict discrimination measured using the time-dependent area under the receiver operating characteristic curve (AUROC). The bottom panels depict discrimination using time-dependent average precision (AP). A model with no discriminative power (“No Skill”), or the performance observed when all individuals are classified as cases, is depicted by black triangles.

Echo2HF, PCP-HF, and PREVENT were all generally well-calibrated, although all tended to overestimate risk at higher predicted probabilities (Table 2, Figure 6).

Figure 6. Model calibration comparison between Echo2HF, PCP-HF and PREVENT.

Figure 6.

Figure 6 displays comparisons between Echo2HF (green) and the PCP-HF (orange) and PREVENT (lavender) scores in the test sets for the calibration of 10-year of incident heart failure (among n=1,116 individuals from MGH and n=7,914 individuals from BWH with complete data for PCP-HF and PREVENT score calculations). The integrated calibration index, an estimate of model error where lower values depict better calibration, is listed next to each curve.

Reclassification analyses favored Echo2HF over both PCP-HF and PREVENT in the BWH test set (PCP-HF: NRI 0.32, 95% CI 0.25–0.38; PREVENT-HF: NRI 0.29, 95% CI 0.25–0.36) which was driven by both increased sensitivity (case reclassification [NRI+], PCP-HF: 0.18, 95% CI 0.12–0.25; PREVENT: 0.05, 95% CI 0.006–0.12) and increased specificity (non-case reclassification [NRI−], PCP-HF: 0.13, 95% CI 0.12–0.15; PREVENT: 0.24, 95% CI 0.23–0.26) (Table S6). Continuous NRI was also favorable (PCP-HF: 0.69, 95% CI 0.51–0.80; PREVENT: 0.67, 95% CI 0.55–0.77).

Echo2HF Model Performance for 5-year Incident Heart Failure

Over 5 years of follow-up, there were 238 HF events in MGH (5-year cumulative risk 5.4%, 95% CI 4.7–6.0%) and 1,130 events in BWH (5-year cumulative risk 5.0%, 4.7–5.3%). Echo2HF-5 year demonstrated good discrimination (AUROC: MGH Test 0.86 [0.84–0.88], BWH 0.86 [0.85–0.87], p-value <0.001 versus age and sex for both; AP: MGH Test 0.19 [0.16–0.23], BWH 0.20 [0.18–0.22], p-value <0.001 versus age and sex for both) (Table 2 and Figure S9). The cumulative incidence of HF and hazard ratios for future HF were again higher among individuals with high versus low Echo2HF-5-year risk in BWH (HF incidence, high [above median] HF risk: 9.8% vs low [below median] HF risk: 0.6%, HR 19.3, 95% CI 14.9–24.9, p-value <0.001). Echo2HF-5-year also demonstrated excellent calibration (ICI [95% CI]: MGH Test 0.05 [0.04–0.06], BWH Test 0.03 [0.03–0.04], Table 2 and Figure S10).

Saliency Mapping

In exploratory analyses, we produced saliency maps depicting areas of the echocardiogram with the greatest influence on Echo2HF model prediction. Saliency maps demonstrated that the LV and RV myocardium and mitral and tricuspid valve annuli appeared to have the greatest influence on Echo2HF risk predictions (Figure S11).

Discussion

We developed Echo2HF, an echocardiogram-based deep learning 3D-CNN model that predicts 5- and 10-year risk of incident HF among patients undergoing clinically indicated echocardiography. Our findings are three-fold. First, Echo2HF predicts future HF in two ambulatory healthcare samples. Second, among individuals with available data, Echo2HF showed favorable performance relative to PCP-HF and PREVENT, two validated clinical risk tools for prediction of incident HF. Finally, Echo2HF offered reasonable discrimination among individuals with both HFrEF and HFpEF. Together, our findings suggest that deep learning analysis of echocardiograms may enhance HF risk prediction beyond currently available clinical tools, an important unmet need given the expanding availability of HF preventive therapies.

The feasibility and accuracy of deep learning echocardiographic interpretation have been established, but application to incident disease prediction remains limited. Prior studies have shown that deep learning of echocardiographic images can accurately and reproducibly measure cardiac structure and function, identify left ventricular systolic dysfunction, and detect prevalent diseases (including amyloidosis and hypertrophic and dilated cardiomyopathies). However, relatively few models have paired with longitudinal outcome data.811 By integrating deep learning echocardiographic interpretation with well-curated EHR cohorts and rigorously defined HF outcomes, we extend prior work showing that deep learning model-derived measures of left heart structure and function (e.g. model-estimated LVEF, LV end-diastolic dimension) are associated with incident HF.15 Echo2HF demonstrates that future HF can be estimated directly from echocardiographic images, without first extracting traditional measures of cardiac structure and function.

Our work suggests that the echocardiogram holds rich latent information relevant to prediction of HF beyond clinical risk variables. In the BWH external test set, Echo2HF achieved numerically higher discrimination than PCP-HF and PREVENT among individuals with complete data, and Echo2HF AP was significantly higher than both clinical scores. AP may be particularly relevant in this setting because incident HF is uncommon and effective risk stratification depends on identifying the subset of patients most likely to develop disease.5,23 Echo2HF also performed well across subgroups, including individuals without known cardiovascular risk factors and those with preserved LVEF, suggesting that Echo2HF leverages orthogonal information from traditional clinical risk models to improve risk estimation. This is important because many individuals who develop HF may not be identified as high risk using current approaches until structural heart disease, symptoms, or overt clinical risk factors are recognized. Combined models incorporating Echo2HF with PCP-HF and PREVENT probabilities demonstrated modest improvement in discrimination over each model alone, supporting the potential value of integrating imaging-derived and clinical risk estimates.

Echo2HF also discriminated reasonably well among patients with both future HFrEF and HFpEF in exploratory analyses. HF subtypes represent distinct entities despite similar clinical phenotypes, with unique risk factor profiles, etiologic drivers, patterns of structural remodeling, and response to therapy.2728 Prediction of HFpEF is particularly challenging because of its clinical and biological heterogeneity.29 Although Echo2HF performed slightly better for HFrEF than HFpEF, discrimination was favorable for both HF subtypes. These findings suggest that deep learning echocardiographic analysis can capture imaging features relevant to multiple HF pathways, including latent features not readily measured by conventional LVEF-based assessment. Because broadly validated HF risk prediction models incorporating conventional echocardiographic and biomarker data are limited, we could not determine whether deep learning analysis of echocardiographic videos adds incremental predictive value beyond these data, an important area for future investigation.

Saliency mapping provided exploratory biological context. The LV and RV myocardium were among the most influential regions for Echo2HF risk prediction, which is consistent with the central role of ventricular structure and function in HF development. The apparent importance of the mitral and tricuspid valve annuli is also notable, as annular motion reflects longitudinal systolic and diastolic function and has been shown to be associated with cardiometabolic risk factors including age, diabetes, and hypertension.30 Although saliency maps do not establish mechanism, their localization to biologically plausible cardiac structures supports the interpretation that Echo2HF is leveraging meaningful latent imaging features for HF risk prediction.

Our findings have several potential clinical implications. Echo2HF may be best positioned as a complementary tool for opportunistic HF risk stratification among patients already undergoing clinically indicated echocardiography, rather than a replacement for simpler clinical risk models. A plausible near-term implementation pathway would be opportunistic HF risk stratification from routinely acquired echocardiograms, with Echo2HF model output incorporated into the echocardiography report or downstream clinical workflows to prompt closer phenotyping, clinical review, or targeted preventive strategies for individuals at elevated predicted risk. Future work should define actionable thresholds, evaluate patient and provider responses to echocardiography-based HF risk estimates, assess whether deep-learning analysis of routine echocardiograms changes management, and determine whether the predictive value extends to point-of-care ultrasound, may be more scalable in some clinical settings. Ultimately, prospective validation, including selection of actionable thresholds, will be important to evaluate whether translation of models like Echo2HF will meaningfully improve HF-related outcomes through early identification of at-risk individuals for targeted provision of preventive therapies.

Limitations

Our study has several limitations. First, Echo2HF was trained on clinically indicated echocardiograms, which may introduce selection bias and confounding by indication. Although we used ambulatory cohorts designed to reduce ascertainment bias, TTEs are not routine screening tests, and the study population was older, more comorbid, and at higher HF risk than an unselected community population.16 Model performance should therefore be interpreted among patients selected for TTE. Second, the cohorts were drawn from two academic institutions within one integrated healthcare system from a single geographic location, limiting generalizability. Further validation in independent, diverse, and prospective cohorts is needed. Third, prevalent HF cases were included during model training, although all validation and test analyses excluded individuals with prevalent HF. This approach may have encouraged the model to learn features of established disease. Fourth, although overall calibration was favorable, Echo2HF tended to overestimate heart failure risk at higher predicted probabilities, suggesting that recalibration may be needed before implementation in new populations. Fifth, death was treated as a censoring event rather than as a competing risk, which may affect long-term absolute risk estimates and calibration, particularly over a 10-year horizon. However, strong 5-year performance provides reassurance that findings were not limited to long-term risk prediction. Sixth, comparisons with PCP-HF and PREVENT were subject to several practical limitations. PCP-HF was implemented in modified form because QRS duration was not consistently available, and fasting glucose was estimated from HbA1c when direct glucose values were unavailable. In addition, clinical variables were ascertained across a pragmatic EHR window around the index TTE, which may not fully reflect a strict baseline risk-score assessment. Seventh, although the HF definition used in this study was previously validated, incident HF events were ascertained from the EHR and may not capture events that occur outside the context of the EHR. Finally, Echo2HF used only PLAX, A4C, and A2C views; additional views, Doppler, and tissue doppler imaging may further improve prediction.

Conclusions

Echo2HF is an echocardiogram-based deep learning model that predicts incident HF directly from TTE images. Across two large ambulatory care samples comprising 105,565 individuals, Echo2HF discriminated future HF risk, showed favorable calibration, and demonstrated favorable predictive utility compared to two validated clinical HF risk scores, PCP-HF and PREVENT. Deep learning echocardiographic interpretation may support opportunistic HF risk stratification among patients undergoing echocardiography and may help prioritize selected at-risk individuals for preventive and therapeutic therapies. Prospective studies should define actionable thresholds, evaluate workflow integration, and determine whether model-guided care improves HF outcomes.

Supplementary Material

Supplementary Material

What is New?

  • A deep learning model applied directly to routine transthoracaic echocardiographic videos (Echo2HF) accurately predicts risk of incident heart failure without reliance on predefined echocardiographic measurements or clinical variables.

  • Echo2HF demonstrated consistent discrimination and calibration across two large ambulatory cohorts and outperformed established clinical HF risk scores for prediction of future heart failure.

What are the Clinical Implications?

  • Deep learning-enabled analysis of standard echocardiograms may enable opportunistic, scalable heart failure risk stratification among patients undergoing clinically indicated echocardiography, including individuals without known structural heart disease or traditional cardiovascular risk factors.

  • Deep learning echocardiographic models may help prioritize patients for heart failure preventive therapies.

Sources of Funding:

Dr. Lau is supported by the NIH (K23HL159243, R01HL180648), the American Heart Association, the Massachusetts Life Sciences Center, and the Tianqiao and Chrissy Chen Institute. Dr. Cunningham is supported by the American Heart Association (23CDA1052151) and the National Heart, Lung, and Blood Institute (1K23HL168163). Dr. Lubitz was supported by NIH grants 1R01HL139731, R01HL157635, and American Heart Association 18SFRN34250007 during this work. Dr. Khurshid is supported by the NIH (K23HL169839) and the AHA (23CDA1050571). Dr. Ellinor is supported by the NIH (1R01HL092577, K24HL105780), AHA (18SFRN34110082), Foundation Leducq (14CVD01), and by MAESTRIA (965286). Dr. Ho is supported by the NIH (R01 HL168889, R01HL160003, and K24 HL153669).

Relationships with Industry:

Dr. Lau consults or serves on advisory boards for SystoleHealth, Inc, Amissa Health, and Roon, unrelated to this work. Dr. Cunningham has consulted for Edgewise Therapeutics, us2.ai, and Maribel Health, and served on a Data Safety Monitoring Board for Cytokinetics. Dr. Lubitz is a full-time employee of Novartis as of July 2022. Dr. Lubitz has received sponsored research support from Bristol Myers Squibb, Pfizer, Boehringer Ingelheim, Fitbit, Medtronic, Premier, and IBM, and has consulted for Bristol Myers Squibb, Pfizer, Blackstone Life Sciences, and Invitae. Dr. Picard has previously consulted for Vertex. Dr. Ho has received past sponsored research support from Bayer AG. Dr. Ellinor receives sponsored research support from Bayer AG and IBM Health and has served on advisory boards or consulted for Bayer AG, Quest Diagnostics, MyoKardia, and Novartis. Dr. Ho has received consulting fees from Eli Lilly.

Biography

graphic file with name nihms-2196559-b0007.gif

Footnotes

Publisher's Disclaimer: This is a PDF of an article that has undergone enhancements after acceptance, such as the addition of a cover page and metadata, and formatting for readability. This version will undergo additional copyediting, typesetting and review before it is published in its final form. As such, this version is no longer the Accepted Manuscript, but it is not yet the definitive Version of Record; we are providing this early version to give early visibility of the article. Please note that Elsevier’s sharing policy for the Published Journal Article applies to this version, see: https://www.elsevier.com/about/policies-and-standards/sharing#4-published-journal-article. Please also note that, during the production process, errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

References

  • 1.Martin SS, Aday AW, Allen NB, et al. 2025 Heart Disease and Stroke Statistics: A Report of US and Global Data From the American Heart Association. Circulation. 2025;151:e41–e660. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Chen J, Normand S-LT, Wang Y, Krumholz HM. National and Regional Trends in Heart Failure Hospitalization and Mortality Rates for Medicare Beneficiaries, 1998–2008. JAMA. 2011;306:1669–1678. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.McMurray JJ, Petrie MC, Murdoch DR, Davie AP. Clinical epidemiology of heart failure: public and private health burden. Eur Heart J. 1998;19 Suppl P:P9–16. [PubMed] [Google Scholar]
  • 4.Heidenreich PA, Bozkurt B, Aguilar D, et al. 2022 AHA/ACC/HFSA Guideline for the Management of Heart Failure: A Report of the American College of Cardiology/American Heart Association Joint Committee on Clinical Practice Guidelines. Circulation. 2022;145:e895–e1032. [DOI] [PubMed] [Google Scholar]
  • 5.Khan SS, Ning H, Shah SJ, et al. 10-Year Risk Equations for Incident Heart Failure in the General Population. J Am Coll Cardiol. 2019;73:2388–2397. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Agarwal SK, Chambless LE, Ballantyne CM, et al. Prediction of Incident Heart Failure in General Practice: The ARIC Study. Circ Heart Fail. 2012;5:422–429. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Butler J, Kalogeropoulos A, Georgiopoulou V, Belue R, Rodondi N, Garcia M, Bauer DC, Satterfield S, Smith AL, Vaccarino V, Newman AB, Harris TB, Wilson PWF, Kritchevsky SB, Health ABC Study. Incident heart failure prediction in the elderly: the health ABC heart failure score. Circ Heart Fail 2008;1:125–133. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Ghorbani A, Ouyang D, Abid A, et al. Deep learning interpretation of echocardiograms. Npj Digit Med. 2020;3:1–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Ouyang D, He B, Ghorbani A, et al. Video-based AI for beat-to-beat assessment of cardiac function. Nature. 2020;580:252–256. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Sangha V, Nargesi AA, Dhingra LS, et al. Detection of Left Ventricular Systolic Dysfunction From Electrocardiographic Images. Circulation. 2023;148:765–777. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Zhang J, Gajjala S, Agrawal P, et al. Fully Automated Echocardiogram Interpretation in Clinical Practice. Circulation. 2018;138:1623–1635. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Kwon J-M, Kim K-H, Jeon K-H, Park J Deep learning for predicting in-hospital mortality among heart disease patients based on echocardiography. Echocardiogr Mt Kisco N. 2019;36:213–218. [DOI] [PubMed] [Google Scholar]
  • 13.Ulloa Cerna AE, Jing L, Good CW, et al. Deep-learning-assisted analysis of echocardiographic videos improves predictions of all-cause mortality. Nat Biomed Eng. 2021;5:546–554. [DOI] [PubMed] [Google Scholar]
  • 14.Valsaraj A, Kalmady SV, Sharma V, et al. Development and validation of echocardiography-based machine-learning models to predict mortality. eBioMedicine. 2023;90. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Lau ES, Di AP, Kopparapu K, et al. Deep Learning–Enabled Assessment of Left Heart Structure and Function Predicts Cardiovascular Outcomes. J Am Coll Cardiol. 2023;82:1936–1948. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Khurshid S, Reeder C, Harrington LX, et al. Cohort design and natural language processing to reduce bias in electronic health records research. NPJ Digit Med. 2022;5:47. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Haimovich JS, Diamant N, Khurshid S, et al. Artificial intelligence-enabled classification of hypertrophic heart diseases using electrocardiograms. Cardiovasc Digit Health J. 2023;4:48–59. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Cunningham JW, Singh P, Reeder C, et al. Natural Language Processing for Adjudication of Heart Failure in the Electronic Health Record. JACC Heart Fail. 2023;11:852–854. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Cunningham JW, Singh P, Reeder C, et al. Natural Language Processing for Adjudication of Heart Failure in a Multicenter Clinical Trial: A Secondary Analysis of a Randomized Clinical Trial. JAMA Cardiol. 2024;9:174–181. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Marti-Castellote PM, Reeder C, Claggett B, et al. Natural Language Processing to Adjudicate Heart Failure Hospitalizations in Global Clinical Trials. Circ Heart Fail. 2025;18:e012514. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.A scalable discrete-time survival model for neural networks [PeerJ] Accessed April 17, 2025. https://peerj.com/articles/6257/. [DOI] [PMC free article] [PubMed]
  • 22.Khurshid S, Friedman S, Reeder C, et al. ECG-Based Deep Learning and Clinical Risk Factors to Predict Atrial Fibrillation. Circulation. 2022;145:122–133. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Khan SS, Matsushita K, Sang Y, et al. Development and Validation of the American Heart Association’s PREVENT Equations. Circulation. 2024;149:430–449. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Bavishi A, Bruce M, Ning H, et al. Predictive Accuracy of Heart Failure-Specific Risk Equations in an Electronic Health Record-Based Cohort. Circ Heart Fail. 2020;13:e007462. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.eAG/A1C Conversion Calculator | American Diabetes Association; Accessed April 17, 2025. https://professional.diabetes.org/glucose_calc. [Google Scholar]
  • 26.Robins JM, Finkelstein DM. Correcting for noncompliance and dependent censoring in an AIDS Clinical Trial with inverse probability of censoring weighted (IPCW) log-rank tests. Biometrics. 2000;56:779–788. [DOI] [PubMed] [Google Scholar]
  • 27.Ho JE, Zern EK, Wooster L, et al. Differential Clinical Profiles, Exercise Responses, and Outcomes Associated With Existing HFpEF Definitions. Circulation. 2019;140:353–365. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.de Boer RA, Nayor M, deFilippi CR, et al. Association of Cardiovascular Biomarkers With Incident Heart Failure With Preserved and Reduced Ejection Fraction. JAMA Cardiol. 2018;3:215–224. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Choy M, Liang W, He J, et al. Phenotypes of heart failure with preserved ejection fraction and effect of spironolactone treatment. ESC Heart Fail. 2022;9:2567–2575. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Sohn D-W, Chai I-H, Lee D-J, et al. Assessment of Mitral Annulus Velocity by Doppler Tissue Imaging in the Evaluation of Left Ventricular Diastolic Function. J Am Coll Cardiol. 1997;30:474–480. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material

Data Availability Statement

MGH and BWH data contain protected health information and are not publicly available because individuals did not consent to data sharing. Source code and model weights for Echo2HF are available at: https://github.com/broadinstitute/ml4h/tree/master/model_zoo/Echo2HF

RESOURCES