Skip to main content
European Heart Journal. Digital Health logoLink to European Heart Journal. Digital Health
. 2026 Mar 25;7(4):ztag049. doi: 10.1093/ehjdh/ztag049

Wearable-Echo-FM: an ECG echo foundation model for 1-lead electrocardiography

Elizabeth Knight 1, Evangelos K Oikonomou 2, Arya Aminorroaya 3, Aline F Pedroso 4, Rohan Khera 5,6,7,8,✉,b
PMCID: PMC13131981  PMID: 42077383

Abstract

Aims

Artificial intelligence (AI) models can now detect patterns of structural heart diseases (SHDs) from electrocardiograms (ECGs), though scaling them requires the broader use of 1-lead ECGs that are now ubiquitous in wearable and portable devices. However, model development for these devices is limited by a lack of diagnostic labels for SHDs for wearable ECGs.

Methods and results

Here, we present Wearable-Echo-FM, a foundation model that encodes 1-lead ECGs with information from echocardiographic text reports. Using 194 551 1-lead ECG-echo pairs from 77 378 adults (2015–2018), we contrastively pre-trained ECG convolutional neural network (CNN) and RoBERTa text encoders. The ECG encoder was fine-tuned on a distinct progressively larger ECG set (250 to 250 260 ECGs) to detect different cardiac disorders: (i) left-ventricular systolic dysfunction (LVSD), (ii) diastolic dysfunction, and (iii) a composite SHD. This was compared with a randomly initialized CNN, with both approaches evaluated in an independent held-out test set. With the full training set, Wearable-Echo-FM matched the baseline CNN (AUROC 0.894 vs. 0.884 for LVSD; 0.849 vs. 0.843 diastolic dysfunction; 0.887 vs. 0.869 composite). With only 0.5% (∼1000 ECGs) of data, it markedly outperformed baseline (0.855 vs. 0.548; 0.819 vs. 0.582; 0.863 vs. 0.496, respectively).

Conclusion

Contrastive pre-training of 1-lead ECGs on echocardiographic text reduces label requirements for label-efficient development of SHD screening models on 1-lead ECGs, providing a foundation for future validation on wearable and portable devices.

Keywords: artificial intelligence, Contrastive learning, Foundation model, Structural heart disease, Wearable devices

Graphical Abstract

Graphical Abstract.

For image description, please refer to the figure legend and surrounding text.

Introduction

Structural heart diseases (SHD), including a range of myocardial and valvular abnormalities, are conditions with modifiable risk factors and evidence-based treatments that can alter the risk of heart failure (HF) and premature mortality.1,2 Early diagnosis is essential for the timely initiation of evidence-based therapies that can slow disease progression and reduce morbidity and mortality.3–7 Despite the availability of effective treatments, SHD remains frequently underdiagnosed, often only identified when patients present with overt symptoms.4,5,8,9 The reliance on resource-intensive cardiac imaging, specifically echocardiography, for diagnosing these conditions, limits early detection, particularly in low-resource settings.10

Recent advances in artificial intelligence (AI) have demonstrated the potential of electrocardiograms (ECGs) as screening tools for various cardiac conditions.1,11–14 Traditional 12-lead ECGs have been successfully leveraged for AI-based disease prediction;11,12,15 however, 1-lead ECGs, now integrated into widely available wearable and handheld devices, offer a more accessible alternative for large-scale screening.14–17 Early applications show the role of 1-lead ECGs obtained from wearable and portable devices in identifying left ventricular systolic dysfunction (LVSD), suggesting that the 1-lead data might have sufficient information to capture SHDs.1,14 However, the development of AI models for SHD screening from 1-lead ECGs remains constrained by the limited availability of labeled datasets, particularly for rare cardiac conditions.18,19 Moreover, development of models for individual diseases is resource-intensive, especially when limited labelled data are likely to be available for each condition.

To address these challenges, we propose an ECG-echo foundation model (Wearable-Echo-FM) that learns shared representations between 1-lead ECGs and structural cardiac phenotypes extracted from unstructured echocardiographic reports, enabling detection of clinically actionable SHDs, including LV systolic dysfunction and moderate or severe valvular disease, that are often present but undiagnosed in pre-symptomatic or minimally symptomatic patients. By leveraging contrastive pretraining and transfer learning, we sought to enhance the efficiency of AI models for SHD screening from 1-lead ECGs obtainable on wearable and portable devices. We hypothesize that our approach will enable more accurate detection of SHDs, even in settings with limited labeled data, thereby improving the accessibility and scalability of AI-based cardiovascular screening.

Methods

Study overview

Based on data from the electronic health records (EHR) of the Yale New Haven Health System (YNHHS), which encompasses five hospitals and affiliated sites across Connecticut and Rhode Island, we developed a foundation model using transthoracic echocardiography (TTE) reports and concurrent 1-lead ECGs isolated from standard 12-lead ECGs. The study was approved by the Yale Institutional Review Board, which waived the need for informed consent as the research involved secondary analysis of existing data. Initially, we identified a subset of patients who had undergone both TTE and ECG within 30 days of each other between 2015 and 2018 for pretraining and contrastive learning, which created a foundation model we named Wearable-Echo-FM. We then set aside a temporally distinct cohort of patient data recorded between 2019 and 2023 for model fine-tuning and evaluation. We divided these 2019–2023 ECG TTE pairs into three task-specific cohorts according to each structural heart disease (SHD) label described below. Each task’s cohort was used independently for training, validation, and testing to evaluate performance on that specific label. Subsequently, a temporally distinct and age and sex matched dataset from 2019 to 2023 was used to identify SHD conditions. We deliberately used a temporally distinct cohort (2019–2023) for fine-tuning and evaluation to better approximate real-world deployment, during which reporting templates, echo equipment, and clinical practice may evolve. This design introduces a temporal domain shift between pre-training and downstream tasks. We fine-tuned each model with varying amounts of data to compare the effectiveness of the Wearable-Echo-FM initialisation (weights trained from the contrastive pretraining step) vs. random initialisation.

Data source and study population

Patient-level EHR data were collected from the YNHHS, encompassing a diverse patient population where patients were either seeking inpatient or outpatient care. Patients aged 18 or older with both a TTE and an ECG within 30 days of each other between 2015 and 2023 were included in the study. If multiple ECGs were recorded within 30 days of the TTE, a maximum of five ECGs were included per patient in the training set. In cases with multiple TTEs within 30 days of an ECG, we chose the TTE closest in time to avoid duplicate labelling of the same ECG. We limited the training set to a maximum of five ECG-TTE pairs per individual to prevent patients with very frequent ECGs from disproportionately influencing model training. We excluded patients who, at the time of ECG recording, had a history of cardiac procedures such as coronary artery bypass grafting, aortic or mitral valve procedure, left ventricular assist device implantation, heart transplant, alcohol septal ablation, or ventricular myectomy. Exclusion of patients with prior cardiac surgery was intended to reduce the over-representation of individuals with very frequent ECGs and advanced disease, who are also those who were unlikely to be candidates for screening using wearable and portable devices. The dataset included demographics (age, gender, ethnic group), ECG data such as the date of the ECG recording, and corresponding cardiac diagnoses on the TTE, as noted below (Table 1).

Table 1.

Demographic and clinical characteristics of the model development population

Characteristic Pre-training (2015–2018) Fine-tuning Data (2019–2023)
LVEF ≤40% LVDD Composite SHD
Train Train Val Test Train Val Test Train Val Test
Number of ECG Echo pairs 194551 250260 11925 11923 123306 6727 6584 132310 6678 6540
Number of unique individuals 77378 95388 11925 11923 59831 6727 6584 58815 6678 6540
Positive ECG Echo pairs 29479 1042 1028 19572 763 768 59340 2447 2415
Gender
 Female 86309 (44.4%) 116014 (46.4%) 5904 (49.5) 5907 (49.5) 59683 (48.4) 3501 (52.0) 3389 (51.5) 62759 (47.4) 3394 (50.8) 3354 (51.3)
 Male 96463 (49.6%) 132821 (53.1%) 6021 (50.5) 6016 (50.5) 63621 (51.6) 3226 (48.0) 3195 (48.5) 69546 (52.6) 3284 (49.2) 3186 (48.7)
 Unknown 11779 (6.1%) 1425 (0.6%) 2 (0.0) 5 (0.0)
Age, median [Q1, Q3] 68.0 [57.0, 79.0] 69.6 [58.4, 79.5] 68.5 [56.5, 78.7] 68.5 [56.5, 78.7] 69.6 [58.4, 79.5] 68.5 [56.5, 78.7] 68.5 [56.5, 78.7] 69.6 [58.4, 79.5] 68.5 [56.5, 78.7] 68.5 [56.5, 78.7]
Ethnic Group, # (%)
 Caucasian 141346 (72.7%) 172190 (68.8%) 8106 (68.0%) 8211 (68.9%) 81667 (66.2%) 4376 (65.1%) 4330 (65.8%) 89258 (67.5%) 4373 (65.5%) 4348 (66.5%)
 Black 27478 (14.1%) 32921 (13.2%) 1425 (11.9%) 1385 (11.6%) 16464 (13.4%) 835 (12.4%) 805 (12.2%) 17701 (13.4%) 867 (13.0%) 783 (12.0%)
 Asian 2539 (1.3%) 3587 (1.4%) 178 (1.5%) 183 (1.5%) 1952 (1.6%) 106 (1.6%) 109 (1.7%) 2050 (1.5%) 114 (1.7%) 105 (1.6%)
 Hispanic 16421 (8.4%) 21423 (8.6%) 997 (8.4%) 992 (8.3%) 11570 (9.4%) 615 (9.1%) 598 (9.1%) 11971 (9.0%) 601 (9.0%) 563 (8.6%)
 Unknown 4997 (2.6%) 6091 (2.4%) 353 (3.0%) 324 (2.7%) 2934 (2.4%) 177 (2.6%) 165 (2.5%) 3109 (2.3%) 174 (2.6%) 170 (2.6%)
BMI, median [Q1, Q3] 27.73 [23.82, 32.75] 28.08 [24.15, 33.12] 28.05 [24.30, 32.83] 28.13 [24.23, 32.84] 28.12 [24.32, 33.00] 28.19 [24.39, 32.76] 28.10 [24.29, 32.71] 27.75 [23.97, 32.51] 27.89 [24.09, 32.37] 27.72 [24.04, 32.42]
Atrial Fibrillation History # (%) 39995 (20.56%) 43018 (17.19%) 1430 (11.99%) 1371 (11.50%) 11646 (9.44%) 388 (5.77%) 350 (5.32%) 20122 (15.21%) 678 (10.15%) 620 (9.48%)
Atrial Fibrillation Present on ECG # (%) 28780 (14.79%) 34419 (13.75%) 1346 (11.29%) 1344 (11.27%) 7077 (5.74%) 294 (4.37%) 251 (3.81%) 16297 (12.32%) 672 (10.06%) 617 (9.43%)

Abbreviations: Composite SHD, LVEF ≤40%, moderate or severe aortic stenosis, aortic regurgitation, mitral regurgitation and/or severe left ventricular hypertrophy; ECG, electrocardiogram; LVDD, left ventricular diastolic dysfunction; LVEF, left ventricular ejection fraction. Data are presented as median [interquartile range], or number (percentage). Percentages are column percentages and may not sum to 100% due to rounding and exclusion of small “Other” race/ethnicity categories.

We extracted lead I from standard 12-lead ECG recordings as our 1-lead input, as this lead approximates the vector used by many commercially available wrist-worn and handheld wearable devices. Our goal was to learn a representation that is adaptable to wearable and portable devices rather than specifically designed for any single consumer device. Of note, some wearables (e.g. chest strap or patch devices) have different electrode configurations, and we therefore chose lead I as a proxy for a common group of devices rather than an exact match for all wearable-derived ECG data.

Study outcome/outcome labels

Our study focused on three classification tasks using 1-lead ECGs, each addressing critical markers of cardiac health. First, we predicted left ventricular ejection fraction (LVEF) ≤40%, a well-established threshold for defining LVSD. Second, we classified patients with left ventricular diastolic dysfunction, as determined by echocardiographic imaging and followed the ASE 2016 criteria. We restricted labels to moderate or severe LV diastolic dysfunction. Lastly, we predicted a composite measure of SHDs. The composite SHD label was defined as the TTE confirmed presence of LVSD, moderate or severe left-sided valvular diseases, and/or standardized left ventricular hypertrophy (sLVH). Moderate or severe left-sided valvular diseases were identified by the presence of any moderate or severe aortic regurgitation (AR), aortic stenosis (AS), mitral regurgitation (MR), or mitral stenosis (MS). We characterized sLVH by an interventricular septal diameter at end diastole (IVSd) greater than 15 mm with concomitant moderate or severe left ventricular diastolic dysfunction. By predicting myocardial and valvular abnormalities, systolic and diastolic dysfunction, as well as functional and structural cardiac parameters, we sought to explore the breadth of phenotypic information encoded within 1-lead ECGs using the Wearable-Echo-FM model. For each classification task, non-case ECG TTE pairs were defined as those whose paired TTE did not meet the corresponding positive criteria. These non-cases may include individuals with other cardiovascular comorbidities or milder forms of structural disease below our thresholds. With the model finetuning data (2019–2023), we identified separate cohorts of ECG TTE pairs for each classification task: LVEF ≤40%, left ventricular diastolic dysfunction, and a composite of SHD. During the classification task, we dropped data with missing outcome labels (Table 1). We focused on echocardiographic states with clear guideline-directed treatment implications (LVEF ≤40%, moderate or severe valve disease, and sLVH with concomitant moderate/severe diastolic dysfunction), which are reliably captured in structured echocardiographic reports. These labels, therefore, represent actionable SHD, rather than the earliest subclinical remodelling. We used these task-specific datasets to fine-tune and evaluate the Wearable-Echo-FM model’s performance in predicting each clinical outcome through validation and testing. Details of the validation and test datasets are included in Supplementary Methods.

Model training

The study data were divided into two distinct subsets, which we split as pre 2019 data (2015 2018) for pretraining and contrastive learning purposes through a contrastive language image pre-training (CLIP) framework,20 and used the temporally distinct subset of ECGs done during 2019–2023 for finetuning (Figure 1A and B).

Figure 1.

For image description, please refer to the figure legend and surrounding text.

Overview of data flow and CLIP-style pre-training. Patient records from the YNHHS that contained paired TTE reports and ECGs acquired ≤30 days apart were split into 2015 2018 data (A) for contrastive pre-training and 2019–2023 data (B) for downstream fine-tuning and evaluation. Panel A shows pre-training; a CNN encodes 1-lead ECG signals while a RoBERTa-based encoder processes the corresponding TTE text. The two embeddings (ZECG, ZTTE) are pulled together for matching pairs (green arrows) and pushed apart for non-matching pairs (red arrows), aligning both modalities in a shared latent space. The weights learned in this stage serve as initialisation for subsequent fine-tuning on the temporally distinct cohort. Panel B shows fine-tuning and testing; pre-trained weights are transferred to a distinct 2019–2023 cohort, split 80/10/10% (train/validation/test) at the patient level (≤5 ECGs per subject in train; 1 in validation/test). Separate classifier heads are trained for three echo-derived labels (LVEF ≤40%, LV diastolic dysfunction, and composite SHD) using cross-entropy with class weighting and early stopping (patience = 6). To test label-efficiency, the training set is randomly down-sampled to 0.1%, 0.5%, 1%, 5%, 20%, 50%, and 100%. AUROC is used to measure classification ability. Abbreviations: YNHH, Yale New Haven Hospital. CNN, a convolutional neural network. LLM, large language model. FM, Foundation Model. TTE, Transthoracic Echocardiography. ECG, electrocardiogram.

Model architecture

We employed a seven-layer 1D convolutional neural network (CNN) to encode the 1-lead ECG waveform based on our previously published 1-lead LVSD architecture described in Khunte et al.,14 with modifications. This seven-layer CNN comprises repeated Conv—BatchNorm—ReLU- MaxPool blocks. Kernel sizes decrease from 7 to 3, and filters increase from 16 to 64, with pooling factors of 2 and 4 progressively reducing temporal resolution to a 64 by 10 feature map. This feature map is flattened and projected to a 512-dimensional, L2-normalized ECG embedding for contrastive pre-training. Echocardiography reports are encoded using a RoBERTa-based transformer (12 heads, 6 layers; max length 514) and projected into the same normalized 512-dimensional space (Figure 2), further architectural and training details are provided in Supplementary Methods. For processing textual data, we implement a transformer-based model inspired by Robustly Optimized BERT Pre-training Approach (RoBERTa), with 12 attention heads, 6 hidden layers, a maximum position embedding of 514 and a vocabulary size of 8192. We initially trained a ByteLevelBPETokenizer for this encoder and implemented the model with well-established hyperparameter values and architecture settings. Like the ECG encoder, the text encoder's final embedding size is also 512. Both the ECG and text encoders project their outputs into a shared 512-dimensional embedding space to enable contrastive learning.

Figure 2.

For image description, please refer to the figure legend and surrounding text.

Signal encoder architecture optimized for embedding used in a contrastive framework. The seven-layer 1D CNN model encoded the 10 s 1-lead ECGs. Each convolutional block consists of (Conv, BatchNorm, ReLU, MaxPool), with kernel sizes tapering from 7 to 3 and filters doubling every other layer (16, 32, 64) to capture both broad and fine-grained temporal features. Max-pooling strides of 2 and 4 progressively down-sample the 10 s input while expanding the receptive field. The final convolutional output is flattened and passed through two FC layers (64 and 32 units) with 0.5 dropout, producing a 512-dimensional ECG embedding that feeds the contrastive loss. Abbreviations: CNN, convolutional neural network; ECG, electrocardiogram. Conv, convolution. BatchNorm, Batch normalisation. ReLU, Rectified linear unit activation function. Max pool, maximum pooling layer. FC, fully connected (dense) layer; L2 normalisation, Euclidean (L2) vector norm.

The classification model was built upon the 1-lead encoder by including additional fully connected layers followed by a sigmoid activation with a dropout rate of 0.5 to prevent overfitting. Specifically, these fully connected layers consisted of two custom fully connected layer units, the first with 64 neurons and the second with 32 neurons. These were followed by a flattening operation and a final linear layer that mapped the 320 flattened features to two output neurons. A sigmoid activation was then applied to these two outputs, producing probability-like scores for binary classification.

Pretraining

We initialized the weights of our 1-lead ECG CNN using a CLIP-inspired approach that leveraged paired echocardiogram text data. We used masked language modelling to first pretrain the text encoder, a RoBERTa-based model, on TTE reports from the pre 2019 dataset to acclimate the model to better interpret Yale New Haven Health System (YNHHS) echocardiogram reports. To ensure sufficient textual content and quality, we filtered the data to include only reports with a string length of at least 100 characters, resulting in 87 342 reports used for pretraining. The median number of words per report was 256 (IQR, 198–310), indicating moderate variability in report lengths. We trained a Byte Pair Encoding (BPE) tokenizer using the ByteLevelBPETokenizer, with a vocabulary size of 8192 tokens and a minimum frequency threshold of 3 for token inclusion. The tokenizer was trained on the preprocessed reports, which were converted to lowercase and had dates and other identifying information removed to ensure anonymisation.

The RoBERTa model was pretrained on TTE reports alone for 150 epochs with a learning rate of 104 and a weight decay of 0.01, using a masked language modelling probability of 0.15 to predict masked tokens in the input text. We used a batch size of 164 and employed the Adam optimizer for training. On a patient level, the data was randomly split into training and evaluation sets with a 90:10 ratio. In the training dataset, the median number of tokens per report was 260 with an IQR of [202–314], while the evaluation dataset had a median of 258 tokens [IQR (200–312)]. This pretrained language model served as the textual encoder in our CLIP-inspired framework, facilitating the alignment of ECG signals with echocardiogram-derived textual features.

Next, a CNN was used in conjunction with the pre-trained RoBERTa language model to perform CLIP pretraining, where contrastive learning was employed to align the features extracted from the ECG data with the semantic information from the TTE reports. By training these components together, the goal was to develop pre-trained weights for the CNN that encapsulate rich, multimodal features. The model was trained with a batch size of 80, a cosine learning rate of 10,4 a weight decay of 10,8 for 200 epochs with 100 warm-up steps, and a maximum target length of 400. Early stopping was implemented with a patience of five epochs. As previously mentioned, the CLIP model used a projection dimension of 512 and a temperature of 1 and minimizes the contrastive loss (Figure 1A).

Classification

Following pretraining, we transferred the model weights to ECG report pairs acquired between 2019 and 2023 in a distinct set of patients from those included in the contrastive pretraining. We split the dataset of ECG TTE pairs randomly by patient ID into an 80-10-10 splits where 80% of the data was allocated for training, 10% was allocated for evaluation to enable early stopping, and 10% was used for the final test set (Table 1, Table 2 for counts of ECGs and individuals). We performed the 80-10-10 splits at the patient level, such that each patient contributed ECG TTE pairs exclusively to the training, validation, or test sets. Within the training set, we included up to five ECG TTE pairs per person; in the validation and test sets, we included only one randomly selected ECG TTE pair per patient. The training dataset was further filtered to include a maximum of five ECGs per individual, while the validation and testing datasets were filtered to include only one ECG per person. Echocardiography labels (LVEF ≤40%, left ventricular diastolic dysfunction, or composite SHD) were derived from paired TTEs (Figure 1B).

Table 2.

Performance metrics for detecting LVEF, LVDD and SHDs in the held-out test set

Label Fraction of Training Data (count) Test AUROC
Wearable-Echo-FM Standard Model
LVEF ≤40 0.001 (250) 0.841 [0.828, 0.853] 0.491 [0.473, 0.51]
0.005 (1251) 0.854 [0.842, 0.865] 0.548 [0.529, 0.567]
0.01 (2503) 0.863 [0.851, 0.874] 0.841 [0.828, 0.853]
0.05 (12 513) 0.876 [0.865, 0.886] 0.857 [0.845, 0.868]
0.2 (50 052) 0.883 [0.872, 0.893] 0.874 [0.862, 0.884]
0.5 (125 130) 0.887 [0.876, 0.896] 0.88 [0.869, 0.89]
1 (250 260) 0.894 [0.884, 0.903] 0.884 [0.874, 0.894]
LVDD 0.001 (123) 0.801 [0.775, 0.823] 0.51 [0.477, 0.542]
0.005 (617) 0.819 [0.795, 0.839] 0.582 [0.549, 0.615]
0.01 (1233) 0.806 [0.781, 0.827] 0.782 [0.758, 0.805]
0.05 (6165) 0.843 [0.822, 0.861] 0.804 [0.779, 0.827]
0.2 (24 661) 0.843 [0.821, 0.861] 0.838 [0.815, 0.856]
0.5 (61 653) 0.851 [0.83, 0.869] 0.851 [0.831, 0.869]
1 (123 306) 0.849 [0.828, 0.866] 0.843 [0.822, 0.862]
Composite SHD 0.001(132) 0.486 [0.466, 0.505] 0.503 [0.482, 0.524]
0.005 (662) 0.864 [0.851, 0.875] 0.496 [0.477, 0.516]
0.01(1323) 0.865 [0.853, 0.876] 0.831 [0.817, 0.843]
0.05 (6616) 0.862 [0.85, 0.873] 0.853 [0.84, 0.864]
0.2 (26 462) 0.878 [0.867, 0.889] 0.859 [0.846, 0.87]
0.5 (66 155) 0.873 [0.862, 0.883] 0.868 [0.856, 0.879]
1 (132 310) 0.887 [0.876, 0.896] 0.869 [0.858, 0.88]

Abbreviations: AUROC, area under the receiver operating characteristic curve; Composite SHD, LVEF ≤ 40%, moderate or severe aortic stenosis, aortic regurgitation, mitral regurgitation and/or severe left ventricular hypertrophy; LVDD, left ventricular diastolic dysfunction; LVEF, left ventricular ejection fraction; Standard Model, with random weights initialized; Wearable-Echo-FM, pretrained model with weights initialized from pretraining step. Data are presented as AUROC, 95% confidence interval.

To assess performance with varying amounts of labelled data, we fine-tuned separate models for each clinical label using progressively smaller random fractions from the model fine-tuning cohort. Down-sampling experiments at 0.1%, 0.5%, 1%, 5%, 20%, 50%, and 100% were performed by drawing a single random subset of the training set using a fixed random seed for each fraction to ensure reproducibility, while keeping the validation and test sets unchanged. We trained the model from the same initialized weights under each condition. For each fraction, we randomly selected the corresponding percentage of ECG report pairs while maintaining the same validation and test splits for consistency. Since some labels had fewer available examples, the exact sample size for each fraction varied across tasks. We trained each model with early stopping and patience of six epochs to prevent overfitting, using cross-entropy loss with class weights to account for label imbalance. All other hyperparameters (learning rate, batch size, optimizer) were the same as in the base fine-tuning setup. The final model performance was reported on the test set (Figure 1B).

Statistical analysis

We summarized demographic characteristics using medians with interquartile ranges (IQR) for continuous variables and counts with percentages for categorical variables. The model's testing performance for each training data subset was evaluated by calculating the area under the receiver operating characteristic curve (AUROC). To obtain the 95% confidence interval for the AUROC, we employ bootstrapping. All tests were two-sided, and statistical significance was set at an alpha level of 0.05. The analyses were performed using Python 3.9 with NumPy, Pandas and Scikit learn libraries.

Results

Study population

We included 524 317 ECG TTE pairs from patients who underwent both procedures within 30 days of each other between 2015 and 2023. These pairs were divided into distinct cohorts for model development and fine-tuning. The pretraining cohort, used for the development of Wearable-Echo-FM, included 194 551 ECG TTE pairs from 77 378 unique individuals. For task-specific model finetuning and testing, we included three post 2019 cohorts: LVEF ≤40% (250 260 ECGs from 95 388 individuals), left ventricular diastolic dysfunction (123 306 ECGs from 59 831 individuals), and a SHD (132 310 ECGs from 58 815 individuals). Additionally, we had separate validation and test sets for each condition.

In the pretraining cohort, patients had a median age of 68.0 [interquartile (IQR) 57.0, 79.0] years at the time of the ECG, and 86 309 (44.4%) were women. The cohort included 141 346 229 (72.7%) White, 27 478 358 (14.1%) Black, 2539 (1.3%) Asian, 16 421 (8.4%) Hispanic and 4997 (2.6%) unknown individuals. The model finetuning cohorts had a median age of 69.6 (IQR 58.4, 79.5) years, with 116 014 (46.4%) women. The population included 172 190 (68.8%) White, 32 921 (13.2%) Black, 3587 (1.4%) Asian, 21 423 (8.6%) Hispanic and 6091(2.4%) Unknown. The proportions of individual labels remained relatively consistent across the validation and test sets (Table 1).

Pretraining results

During contrastive pre-training, the ECG and text encoders were optimized jointly for 200 epochs with a batch size of 80, Adam (β1 = 0.9, β2 = 0.999), and a cosine learning-rate schedule. The cross-modal contrastive loss fell from 4.19 after the first epoch to 0.012 by epoch 200 when pretraining was set to stop.

Performance in held out test set

When trained on the full model finetuning cohort, our Wearable-Echo-FM model achieved comparable performance to the standard CNN with randomly initialized weights. For detecting LVEF ≤40%, the Wearable-Echo-FM model achieved an AUROC of 0.894 (95% CI: 0.884 0.903) compared with 0.884 (95% CI: 0.874 0.894) for the standard CNN. For LV diastolic dysfunction, the AUROCs were 0.849 (95% CI: 0.828 0.869) and 0.843 (95% CI: 0.822 0.862), respectively. For the composite of SHD outcome, the Wearable-Echo-FM model achieved an AUROC of 0.887 (95% CI: 0.876 0.896) vs. 0.869 (95% CI: 0.858 0.88) for the standard CNN (Table 2).

The model performance was dependent on the volume of training data for both standard and Wearable-Echo-FM. However, the Wearable-Echo-FM model consistently outperformed the standard CNN, with the performance gap widening as data became scarcer (Figure 3). In experiments using only 0.5% of the training data (0.005 of the training set, corresponding to 617–1251 ECGs across tasks as detailed in Table 2), the pretrained Wearable-Echo-FM model significantly outperformed the standard model across all tasks. For detecting LVEF ≤40%, the Wearable-Echo-FM model achieved an AUROC of 0.854 (95% CI: 0.842 0.865) compared with 0.548 (95% CI: 0.529 0.567) for the standard model. For LVDD, the AUROCs were 0.819 (95% CI: 0.795 0.839) compared with 0.582 (95% CI: 0.549 0.615) for the standard model and for composite SHD, the Wearable Echo FM model achieved an AUROC of 0.864 (95% CI: 0.851 0.875) vs. 0.496 (95% CI: 0.477 0.516) for the standard model (Table 2).

Figure 3.

For image description, please refer to the figure legend and surrounding text.

Model performance across different fractions of data for LVEF ≤40, LVDD, and the composite outcome. Test-set AUROC (mean ± 95% CI) for detecting LVEF ≤40%, LV diastolic dysfunction, and composite SHD as a function of the fraction of labelled training data used for fine-tuning (0.1% to 100%). The pre-trained Wearable-Echo-FM is shown in blue, and the randomly initialised CNN is shown in orange. Abbreviations: LVEF, left ventricular ejection fraction. LVDD, left ventricular diastolic dysfunction. Composite SHD, LVEF ≤40%, moderate or severe aortic stenosis, aortic regurgitation, mitral regurgitation and/or severe left ventricular hypertrophy. AUROC, area under the receiver operating characteristic curve. Wearable Echo FM, pretrained model with weights initialized from the pretraining step. Standard Model, with random weights initialized.

Discussion

In this study, we introduce Wearable-Echo-FM, a foundation model optimized for screening SHDs from 1-lead ECGs derived from lead I that leverages joint embeddings of ECG signals and paired echocardiographic reports. We demonstrate that the model enables efficient development of downstream applications, achieving superior performance relative to standard training techniques, especially when there is a limited number of cases. Notably, these findings were robust across analyses of various individual SHDs and a composite SHD measure, spanning features of systolic function, diastolic function, valvular abnormalities, and abnormal LV hypertrophy. These findings highlight Wearable Echo FM as a powerful model that can boost our ability to screen for SHD from 1-lead ECGs increasingly available from wearable and portable devices in the community.

The study builds upon existing literature in developing deep learning models to enhance the diagnostic capabilities of ubiquitous wearable devices.1,21 Prior work has focused on two key domains. The first of which represents a series of tools built to detect observable rhythm and conduction disorders.17 1-lead ECG sensors, integrated into consumer devices (such as watches or wristbands), can enable large scale community based screening and monitoring for conditions such as atrial fibrillation and bundle branch blocks.5,14 Advances in electrode placement and signal quality have allowed this technology to produce reliable 1-lead data and detect straightforward, clinically observable patterns that have been validated.22 The second represents tools that identify signatures of a common structural heart disorder, LVSD, a ‘hidden’ feature not identifiable by experts.1,14 In a prospective study in 2022, Attia et al. at the Mayo Clinic demonstrated that an AI ECG model using single-lead ECGs from smartwatches could detect LVSD with an AUROC of approximately 0.88 in real-world settings.23,24 Likewise, our group showed that deep-learning models can robustly identify LVSD from single-lead ECG signals adapted for portable and wearable devices, even under highly noisy conditions.1,21 Building on these advances, our foundation model approach further broadens single-lead ECG–based screening to a broader range of SHDs without requiring extensive labelled datasets.

By leveraging 1-lead ECGs readily obtainable from wearable and handheld devices, our framework enables the development of 1-lead ECG models that are adaptable to wearable and handheld devices, pending prospective validation, and which ultimately could be applied to community-level screening for SHDs.23 Early detection of these structural disorders, which often remain underdiagnosed until clinical symptoms arise, enables timely interventions that can substantially reduce morbidity, mortality, and healthcare costs.7,10 Through harnessing the deeper structural insights of echocardiographic findings, our approach provides an adaptable model that can be applied to underrepresented conditions where large, fully labelled datasets are scarce. As wearable ECG technology continues to be increasingly adopted, this foundation model can be fine-tuned for additional structural and functional cardiac abnormalities. The temporal separation between the pre-training (2015–2018) and fine-tuning (2019–2023) cohorts likely introduced modest shifts in equipment and reporting practices. This may have acted as a form of regularisation, encouraging the foundation model to learn representations that are stable across time rather than over-fitting to a single era. This adaptability broadens its clinical utility and offers a scalable, cost-effective solution for cardiac screening, particularly in resource-limited settings.

An important observation from our experiments is that the Wearable-Echo-FM model consistently outperformed the standard, randomly initialized CNN across a broad spectrum of training set sizes. For instance, when the model was fine-tuned with only 0.5% of the available data (roughly 1000 ECGs), the Wearable-Echo-FM still demonstrated robust detection capabilities for LVEF ≤40%, left ventricular diastolic dysfunction, and composite SHD, while the standard model’s performance degraded substantially. Even at higher fractions of the training data, the pre-trained model provided incremental gains, which suggests that its prior exposure to echocardiographic text and ECG pairings conferred a possible advantage. Because of its training on a broad array of echocardiographic findings, Wearable-Echo-FM captures SHD features even under conditions of extremely limited labeled data.

This study has certain limitations that merit consideration. First, while the model was trained on 1-lead signals extracted from clinical 12-lead ECGs rather than directly from consumer-grade wearables, prior work, including the prospective smartwatch study by Attia et al., has shown that 1-lead algorithms for detecting LVSD retain strong performance in real-world wearable ECGs. These findings suggest that our model may similarly generalize to real-world wearable ECG data, though direct validation is still needed, and extrapolation to consumer wearables must be considered hypothesis-generating until confirmed in prospective wearable datasets. Second, the model was evaluated based on SHD phenotypes defined by TTE, and its applicability to other cardiac conditions remains unexplored. Our ECGs were recorded on clinical Philips and GE systems with their native filter settings, which differ from those used in consumer wearables that often operate at lower sampling frequencies with stronger noise filtering. This device-level domain shift may affect transferability and will require dedicated validation. Our labels reflect established SHD (reduced LVEF and moderate/severe valve disease) rather than subtle or mild structural changes that precede the development of these structural and functional abnormalities. Extending the foundation model to earlier stages of remodelling will require alternative labelling strategies and was beyond the scope of this analysis. We did not repeat down-sampling with multiple random seeds or perform k-fold cross-validation due to the computational cost of retraining both models across all fractions; consequently, some residual sensitivity to the sample used at each fraction is possible. Therefore, some residual stochastic variability cannot be excluded, particularly in the full data regime where gains are modest. Nonetheless, the large and consistent improvements at very low data fractions, with non-overlapping confidence intervals, support a true benefit in label efficiency. In addition, our finetuning and test cohorts reflect patients undergoing clinically indicated echocardiography and therefore represent a higher risk referred population than typical community screening settings. Accordingly, while our primary goal was to assess discrimination based on AUROC, prevalence-dependent measures would be expected to differ in lower prevalence wearable screening populations and will require prospective validation and calibration. Lastly, the model was developed within a single health system, which further limits its generalisability. However, the development cohort from the YNHHS is one of the most racially and ethnically diverse in the US and closely mirrors the national demographic distribution, supporting the potential generalisability of our findings across diverse populations. The inherent variability of echocardiographic measurements in clinical reads could introduce label noise and may affect reported performance. Such non-differential misclassification would be expected to bias AUROC estimates toward the null, making our performance estimates conservative. Prospective studies in broader community settings, ideally with diverse device manufacturers, are essential to confirm the scalability and long-term utility of our approach for widespread SHD screening.

In this work, we present an ECG Echo foundation model that learns shared embeddings between 1-lead ECGs and complex structural cardiac phenotypes recorded on unstructured echocardiographic reports and enhances the efficiency of developing AI algorithms for SHD screening on wearable and portable devices.

Supplementary Material

ztag049_Supplementary_Data

Acknowledgements

Dr Khera was supported by the National Institutes of Health (under awards R01AG089981, R01HL167858, and K23HL153775) and the Doris Duke Charitable Foundation (under award 2022060). The funders had no role in the design and conduct of the study.

Contributor Information

Elizabeth Knight, Section of Cardiovascular Medicine, Department of Internal Medicine, Yale School of Medicine, 333 Cedar Street, PO Box 208017, New Haven, CT 06520-8017, USA.

Evangelos K Oikonomou, Section of Cardiovascular Medicine, Department of Internal Medicine, Yale School of Medicine, 333 Cedar Street, PO Box 208017, New Haven, CT 06520-8017, USA.

Arya Aminorroaya, Section of Cardiovascular Medicine, Department of Internal Medicine, Yale School of Medicine, 333 Cedar Street, PO Box 208017, New Haven, CT 06520-8017, USA.

Aline F Pedroso, Section of Cardiovascular Medicine, Department of Internal Medicine, Yale School of Medicine, 333 Cedar Street, PO Box 208017, New Haven, CT 06520-8017, USA.

Rohan Khera, Section of Cardiovascular Medicine, Department of Internal Medicine, Yale School of Medicine, 333 Cedar Street, PO Box 208017, New Haven, CT 06520-8017, USA; Center for Outcomes Research and Evaluation (CORE), Yale New Haven Hospital, 195 Church Street, 5th Floor, New Haven, CT 06510, USA; Department of Biomedical Informatics and Data Science, Yale School of Medicine, 101 College Street, Floor 10, New Haven, CT 06510, USA; Section of Health Informatics, Department of Biostatistics, Yale School of Public Health, 60 College Street, New Haven, CT 06520-8034, USA.

Supplementary material

Supplementary material is available at European Heart Journal – Digital Health.

Funding

Dr Khera is an Associate Editor of JAMA. He also receives support from the National Institute on Aging (R01AG089981) National Heart, Lung, and Blood Institute of the National Institutes of Health (R01HL167858 and K23HL153775) and the Doris Duke Charitable Foundation (under award 2022060). He receives support from the Blavatnik Family Foundation through the Blavatnik Fund for Innovation at Yale. He also receives research support, through Yale, from Bristol Myers Squibb, BridgeBio, and Novo Nordisk. In addition to 63/346,610, Dr Khera is a coinventor of U.S. Pending Patent Applications WO2023230345A1, US20220336048A1, 63/484,426, 63/508,315, 63/580,137, 63/606,203, 63/619,241, and 63/562,335. Dr Khera and Dr Oikonomou are co-founders of Evidence2Health, a precision health platform to improve evidence-based cardiovascular care. Dr Oikonomou is a co-inventor of the U.S. Patent Applications 63/508,315 & 63/177,117 and has been a consultant to Caristo Diagnostics Ltd (all outside the current work).

Author contributions

Elizabeth Knight (Formal analysis, Investigation [equal], Methodology, Writing—review & editing [supporting], Software, Validation, Visualisation, Writing—original draft [lead]), Evangelos K Oikonomou (Conceptualisation, Formal analysis, Investigation, Software, Validation, Visualisation, Writing—original draft [supporting], Project administration, Supervision [lead]), Arya Aminorroaya (Conceptualisation, Formal analysis, Software, Validation [supporting], Methodology [equal]), Aline F Pedroso [Conceptualisation, Methodology, Writing—original draft, Writing—review & editing (supporting)], and Rohan Khera (Conceptualisation, Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Supervision, Writing—review & editing [lead], Writing—original draft [supporting])

R.K. conceived and designed the study and accessed the data. E.K., A.A., E.K.O. and R.K. developed the model. E.K., E.K.O, A.A. and R.K. performed the statistical analyses. E.K.O., A.F.P., E.K. and R.K. drafted the manuscript. All authors reviewed the study design, provided critical feedback, and approved the final version of the manuscript. R.K. supervised the project, secured funding, and is the guarantor.

Data availability

The dataset cannot be made publicly available because it is an electronic health record and sharing this data externally without proper consent could compromise patient privacy and would violate the Institutional Review Board approval for the study.

Computer code

The code for the study is available here: https://github.com/CarDS-Yale/Wearable-Echo-FM. Upon request and appropriate agreements, we will also share the pre-trained single-lead ECG encoder weights to facilitate external replication and extension of this work.

References

  • 1. Aminorroaya  A, Dhingra  LS, Pedroso  AF, Shankar  SV, Coppi  A, Khunte  A, et al.  Development and multinational validation of an ensemble deep learning algorithm for detecting and predicting structural heart disease using noisy single-lead electrocardiograms. Eur Heart J Digit Health  2025;6:554–566. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Steinberg  DH, Staubach  S, Franke  J, Sievert  H. Defining structural heart disease in the adult patient: current scope, inherent challenges and future directions. Eur Heart J Suppl  2010;12:E2–E9. [Google Scholar]
  • 3. Bonow  RO, Leon  MB, Doshi  D, Moat  N. Management strategies and future challenges for aortic valve disease. Lancet Lond Engl  2016;387:1312–1323. [DOI] [PubMed] [Google Scholar]
  • 4. Wang  TJ, Evans  JC, Benjamin  EJ, Levy  D, LeRoy  EC, Vasan  RS. Natural history of asymptomatic left ventricular systolic dysfunction in the community. Circulation  2003;108:977–982. [DOI] [PubMed] [Google Scholar]
  • 5. d’Arcy  JL, Coffey  S, Loudon  MA, Kennedy  A, Pearson-Stuttard  J, Birks  J, et al.  Large-scale community echocardiographic screening reveals a major burden of undiagnosed valvular heart disease in older people: the OxVALVE population cohort study. Eur Heart J  2016;37:3515–3522. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. McDonagh  TA, Metra  M, Adamo  M, Gardner  RS, Baumbach  A, Böhm  M, et al.  2021 ESC guidelines for the diagnosis and treatment of acute and chronic heart failure. Eur Heart J  2021;42:3599–3726. [DOI] [PubMed] [Google Scholar]
  • 7. Heidenreich  PA, Bozkurt  B, Aguilar  D, Allen  LA, Byun  JJ, Colvin  MM, et al.  2022 AHA/ACC/HFSA guideline for the management of heart failure: a report of the American College of Cardiology/American Heart Association joint committee on clinical practice guidelines. Circulation  2024;145:e895–e1032. [DOI] [PubMed] [Google Scholar]
  • 8. Sara  JD, Toya  T, Taher  R, Lerman  A, Gersh  B, Anavekar  NS. Asymptomatic left ventricular systolic dysfunction. Eur Cardiol  2020;15:e13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Manning  WJ. Asymptomatic aortic stenosis in the elderly: a clinical review. JAMA  2013;310:1490–1497. [DOI] [PubMed] [Google Scholar]
  • 10. Galasko  GIW, Barnes  SC, Collinson  P, Lahiri  A, Senior  R. What is the most cost-effective strategy to screen for left ventricular systolic dysfunction: natriuretic peptides, the electrocardiogram, hand held echocardiography, traditional echocardiography, or their combination?  Eur Heart J  2006;27:193–200. [DOI] [PubMed] [Google Scholar]
  • 11. Dhingra  LS, Aminorroaya  A, Sangha  V, Pedroso  AF, Shankar  SV, Coppi  A, et al.  Ensemble deep learning algorithm for structural heart disease screening using electrocardiographic images: PRESENT SHD. JACC  2025;85:1302–1313. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Dhingra  LS, Aminorroaya  A, Sangha  V, Pedroso  AF, Asselbergs  FW, Brant  LCC, et al.  Heart failure risk stratification using artificial intelligence applied to electrocardiogram images: a multinational study. Eur Heart J  2025;46:1044–1054. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Dhingra  LS, Aminorroaya  A, Pedroso  AF, Khunte  A, Sangha  V, McIntyre  D, et al.  Artificial intelligence–enabled prediction of heart failure risk from single-lead electrocardiograms. JAMA Cardiol  2025;10:574–584. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Khunte  A, Sangha  V, Oikonomou  EK, Dhingra  LS, Aminorroaya  A, Mortazavi  BJ, et al.  Detection of left ventricular systolic dysfunction from single-lead electrocardiography adapted for portable and wearable devices. NPJ Digit Med  2023;6:1–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Sangha  V, Nargesi  AA, Dhingra  LS, Khunte  A, Mortazavi  BJ, Ribeiro  AH, et al.  Detection of left ventricular systolic dysfunction from electrocardiographic images. Circulation  2023;148:765–777. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Sangha  V, Mortazavi  BJ, Haimovich  AD, Ribeiro  AH, Brandt  CA, Jacoby  DL, et al.  Automated multilabel diagnosis on electrocardiographic images and signals. Nat Commun  2022;13:1583. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Kim  J, Lee  SJ, Ko  B, Lee  M, Lee  YS, Lee  KH. Identification of atrial fibrillation with single-lead Mobile ECG during normal Sinus rhythm using deep learning. J Korean Med Sci  2024;39:e56. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Vandenberk  B, Chew  DS, Prasana  D, Gupta  S, Exner  DV. Successes and challenges of artificial intelligence in cardiology. Front Digit Health  2023;5:1201392. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Lopes  RR, Bleijendaal  H, Ramos  LA, Verstraelen  TE, Amin  AS, Wilde  AAM, et al.  Improving electrocardiogram based detection of rare genetic heart disease using transfer learning: an application to phospholamban p.Arg14del mutation carriers. Comput Biol Med  2021;131:104262. [DOI] [PubMed] [Google Scholar]
  • 20. Radford  A, Kim  JW, Hallacy  C, Ramesh  A, Goh  G, Agarwal  S, et al.  Learning transferable visual models from natural language supervision. Int Conf Mach Learn  2021.
  • 21. Pedroso  AF, Khera  R. Leveraging AI-enhanced digital health with consumer devices for scalable cardiovascular screening, prediction, and monitoring. Npj Cardiovasc Health  2025;2:34. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. van der Zande  J, Strik  M, Dubois  R, Ploux  S, Alrub  SA, Caillol  T, et al.  Using a smartwatch to record precordial electrocardiograms: a validation study. Sensors (Basel)  2023;23:2555. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Attia  ZI, Harmon  DM, Dugan  J, Manka  L, Lopez-Jimenez  F, Lerman  A, et al.  Prospective evaluation of smartwatch-enabled detection of left ventricular dysfunction. Nat Med  2022;28:2497–2503. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Aminorroaya  A, Dhingra  LS, Camargos  AP, Shankar  SV, Khunte  A, Sangha  V, et al.  Study Protocol for the Artificial Intelligence-Driven Evaluation of Structural Heart Diseases Using Wearable Electrocardiogram (ID-SHD). medRxiv2024.03.18.24304477. 2024.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

ztag049_Supplementary_Data

Data Availability Statement

The dataset cannot be made publicly available because it is an electronic health record and sharing this data externally without proper consent could compromise patient privacy and would violate the Institutional Review Board approval for the study.


Articles from European Heart Journal. Digital Health are provided here courtesy of Oxford University Press on behalf of the European Society of Cardiology

RESOURCES