Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Jun 1.
Published in final edited form as: JACC Clin Electrophysiol. 2025 Mar 17;11(6):1308–1320. doi: 10.1016/j.jacep.2025.02.003

Expert-Level Automated Diagnosis of the Pediatric ECG Using a Deep Neural Network

Joshua Mayourian a, William G La Cava b, Sarah D de Ferranti a, Douglas Mah a, Mark Alexander a, Edward Walsh a, John K Triedman a
PMCID: PMC12197845  NIHMSID: NIHMS2062250  PMID: 40100196

Abstract

Background:

Disparate access to expert pediatric cardiologist care and interpretation of electrocardiograms (ECGs) persists worldwide. Artificial intelligence-enhanced ECG (AI-ECG) has shown promise for automated diagnosis of ECGs in adults, but has yet to be explored in pediatrics.

Objective:

To determine whether AI-ECG can accurately perform automated diagnosis of pediatric ECGs.

Methods:

This retrospective single-center cohort study included all patients with an ECG at Boston Children’s Hospital read by an experienced pediatric cardiologist (≥5,000 reads) between 2000–2022. A convolutional neural network was trained (75% of patients) and internally tested (25% of patients) on ECGs to predict ECG diagnoses. The primary outcome was a composite of any ECG abnormality (i.e., detecting normal vs. abnormal ECG). Secondary outcomes include Wolff Parkinson White (WPW) syndrome and prolonged QTc. Model performance was assessed with area under the receiver operating (AUROC) and precision recall (AUPRC) curves.

Results:

The main cohort consistent of 201,620 patients (49% male; 11% with known congenital heart disease) and 583,134 ECGs (median age 11.7 [IQR, 3.1–16.9] years; 56% any ECG abnormality; 1.0% WPW; 5.3% with prolonged QTc). AI-ECG outperformed commercial MUSE interpretations for detecting any abnormality (AUROC 0.94, AUPRC 0.96), WPW (AUROC 0.99, AUPRC 0.88), and prolonged QTc (AUROC 0.96, AUPRC 0.63). During readjudication of false negatives/positives, blinded expert readers were more likely to agree with AI-ECG than the original reader to detect any abnormality (p=0.001), WPW (p=0.01), and prolonged QTc (p=0.07).

Conclusions:

Our model provides expert-level automated diagnosis of the pediatric 12-lead ECG, which may improve access to care.

Keywords: Artificial intelligence, electrocardiogram, pediatric screening, long QTc, WPW

Tweet:

Artificial intelligence-enhanced ECG analysis provides expert-level interpretation of the pediatric ECG.

INTRODUCTION

The electrocardiogram (ECG) is a ubiquitous and inexpensive diagnostic test used in pediatric clinical practice worldwide for a wide range of indications: known acquired/congenital heart disease,1 suspicion for arrhythmia or structural heart disease,2 sports clearance,3 medication initiation/monitoring,4 and others. The growing volume of ECGs, along with considerations for universal ECG screening,5 underscores the need for rapid and reliable ECG interpretations.

Specialized expertise is required for pediatric ECG interpretation to account for the progressive anatomic/physiologic changes from the newborn stage to adulthood, which corresponds to significantly different ECG patterns and epidemiology. While many ECG vendors have implemented rule-based diagnostic algorithms, these are generally focused on adult literature with limited standalone value in the pediatric ECG setting. It would be of considerable value—especially in low-resource settings with limited pediatric cardiology expertise—to have access to a novel automated ECG diagnostic tool designed for pediatric use.

Deep learning-based approaches have proven successful for rapid and reliable automated interpretation of ECGs in the adult population,6 making it plausible that artificial intelligence-enhanced ECG (AI-ECG) may similarly aid in ECG interpretation for the pediatric population. However, given the unique considerations of pediatric ECGs, AI-ECG algorithms for adults are expected to have poor generalizability to pediatric cohorts. There are few available AI-ECG applications to pediatric and congenital cardiology,7,8 due in part the paucity of large, well-validated pediatric ECG databases which hinders similar research. Thus, there are currently to our knowledge no AI-ECG algorithm for pediatric ECG diagnosis.

In this manuscript, we address this gap by training and testing a convolutional neural network using >500,000 ECGs on >200,000 patients to reliably identify common and rare ECG findings in the pediatric population. Model performance was benchmarked to commercial software, and discordantly classified tracings were readjudicated to four expert pediatric electrophysiologists. Subgroup analysis defined model performance within a range of ages. Finally, model explainability analysis was performed.

METHODS

Study Population and Patient Assignment

Patient data was utilized from Boston Children’s Hospital. Inclusion criteria consisted of any patient with at least one ECG between 2000–2022 without missing metadata (e.g., missing medical record number, ECG event number, reading provider, or ECG diagnosis). Note that pediatric and adult congenital patients were included given: 1) adult congenital patients have unique ECG patterns/findings that are related to their underlying congenital heart lesion and their sequela of cardiac interventions; 2) pediatric cardiologists are frequently caring for these patients and interpreting their ECGs; 3) there is a global shortage of the adult congenital heart disease workforce.

To optimize training label accuracy, we removed ECGs from less experienced readers (i.e., providers with ≤ 5,000 ECGs). ECGs failing to pass quality control (for quality control details, see “Quality Control and Data Preprocessing” below) were also removed. Finally, ECGs with equivocal (e.g., “probably normal ECG variant”) or outdated (e.g., “counterclockwise rotation”) diagnoses were removed. The remaining ECGs comprised the main cohort, which was then partitioned at the patient level into training (75%) and testing (25%) cohorts (Online Figure 1).

Data Retrieval

ECG waveforms (I, II, and V1–V6) were obtained from the MUSE ECG data management system (GE Healthcare, Chicago, IL), with each lead corresponding to a one-dimensional vector sampled at 250 Hz for a 10 second duration (2500 samples). Leads III, aVF, aVL, and aVR were reconstructed using Einthoven’s law9 and the Goldberger equation.10

ECG diagnoses, age, sex, and congenital heart lesion diagnoses were identified based on the institutional Fyler coding system.11 During the study period, all ECGs were read using a custom internal software (ECG Reader) which requires physicians to select one or more pre-determined ECG diagnostic codes, providing a cleanly labeled dataset for training and testing purposes (Online Figure 2).

Quality Control and Data Preprocessing

Our quality control and preprocessing pipeline has been previously published.7 Briefly, ECGs without 2500 samples or missing lead information were removed. For each passing ECG, we then apply a high pass filter and trimming process to 2048 samples (~8 seconds) for ease of working with convolutional neural networks.

Definition of Primary and Secondary Outcomes

At a high level, we envision the following clinical workflow for an AI-ECG automated diagnosis algorithm: Is the ECG predicted by AI-ECG normal or abnormal? If normal, this may conceivably spare an expert from reviewing the ECG. If abnormal, the AI-ECG algorithm can provide a comprehensive list of pertinent ECG diagnoses.

Guided by this framework, the primary composite outcome was ECG diagnosis of any abnormality (defined as any diagnostic code other than “Normal for age” and “Sinus arrhythmia”) by the original cardiologist reader. Secondary outcomes included individual ECG diagnoses that are especially pertinent to pediatric ECG screening—namely WPW and prolonged QTc. A comprehensive list of all ECG diagnoses that were predicted by AI-ECG herein is shown in Online Table 1.

As secondary analyses, we also evaluated time-to-diagnosis of ECG abnormality, WPW, and prolonged QTc (see “Time-to-Event Analysis” below).

Model Selection, Architecture, and Training

Our model development and architecture mimics our previous work.7 As previously described, the convolutional neural network involves a residual block architecture that has been adapted for unidimensional signals (diagram shown elsewhere).7

Briefly, the model was developed exclusively on the training set, of which 5% was designated for validation and hyperparameter tuning. The input to the convolutional neural network was 12 × 2048 ECG samples. The final hyperparameters were obtained via a grid search: kernel size [3, 9, 17], batch size [8, 32, 64], and initial learning rate [0.01, 0.001, 0.0001]. The average cross-entropy was minimized using the Adam optimizer. Maximum 150 epochs were used with early stopping based on validation loss. Final hyperparameters for this model were kernel size 17, batch size 32, learning rate 0.001.

Given the known role of age and sex on ECG characteristics, we created a second model that incorporated age/sex as inputs along with ECG waveforms (AI-ECG+age+sex). The architecture is shown elsewhere.7 For the AI-ECG+age+sex model, final hyperparameters were kernel size 9, batch size 32, learning rate 0.001.

Performance Evaluation and Statistical Analyses

Model performance was evaluated exclusively on the testing cohort. Area under the receiver operating curve (AUROC) and area under the precision-recall (i.e., positive predictive value (PPV)-sensitivity) curve (AUPRC) were computed. Other performance metrics assessed include PPV, negative predictive value (NPV), sensitivity, specificity, F1, and accuracy. Given the imbalanced dataset and our objective to emulate human behavior when interpreting ECGs (i.e., balancing precision and recall), these metrics were calculated based on thresholds achieving the optimal F1 score. A similar cutoff strategy was implemented by prior AI-ECG works for automated ECG diagnosis in adults.6,12 The F1 score is defined as the harmonic mean of the precision and sensitivity, symmetrically representing both within one metric. 95% Confidence intervals were obtained via resampling with 1,000 bootstraps.

Benchmarking Model Performance

Model performance was benchmarked to commercial (GE MUSE) ECG interpretations.

Time-to-Event Analysis

Time-to-event analysis was performed for the following outcomes: diagnosis of any abnormality, WPW, and prolonged QTc. For each, time-to-diagnosis onset analysis was performed by: 1) including only patients with multiple ECGs; 2) stratifying patients into two groups based on the AI-ECG classification of the first ECGs: true negative or false positive; 3) assessing time-to-diagnosis after ECG within each group.

Cox proportional hazards regression was used to evaluate AI-ECG classification association with time from ECG until the diagnosis of interest. Hazard ratios were adjusted for age and sex. Statistical comparison between groups were based on log-rank testing. Patients who did not experience ECG diagnosis onset were censored at the time of last ECG.

Readjudication and Expert Agreement

Readjudication was performed on ECGs with discrepancies between the diagnostic classification assigned by the original reader and the AI-ECG classification (using the same cutoff as above). More specifically, for each outcome (any abnormality, WPW, prolonged QTc), 50 false positives (i.e., deemed positive by AI-ECG, but negative by the original reader) and 50 false negatives (i.e., deemed negative by AI-ECG, but positive by the original reader) were re-read by four senior pediatric electrophysiologists. These experts were blinded to the diagnoses by AI-ECG, the clinical indication, and the original reader. They were presented with the patient age, sex, and automated measurements of axes/intervals.

Agreement was assessed between the original reader and experts, AI-ECG and experts and all four experts using Cohen’s κ (for agreement between two entities) or Fleiss κ (for agreement between more than two entities).13 Values <0 indicate no agreement, with 0–0.20 as slight, 0.21–0.40 as fair, 0.41–0.60 as moderate, 0.61–0.80 as substantial, and 0.81–1 as near perfect agreement.

Model Explainability

Model behavior was investigated via median waveform analysis and saliency mapping.7 Briefly, median waveforms provide visual representations of high- and low-risk ECGs. The 100 highest predicted ECGs for WPW and prolonged QTc were used to create high-risk median waveforms. To contrast to normal ECGs, the 100 lowest predicted ECGs for any abnormality were used to create low-risk median waveforms.

Saliency mapping provides insight into important ECG patterns that contribute to model prediction. Using a Shapley Additive Explanations framework,14 saliency maps highlight ECG regions where a change in ECG voltage input corresponds to a change in output prediction. The 100 ECGs with highest predicted probability for each diagnosis were used to create saliency maps. For more details of median waveform analysis and saliency mapping, see elsewhere.7

Data Availability and Software

Requests for Boston Children’s Hospital data and related materials will be internally reviewed to clarify if the request is subject to intellectual property or confidentiality constraints. Shareable data and materials will be released under a material transfer agreement for non-commercial research purposes. Use of Boston Children’s Hospital data was approved by their respective Institutional Review Boards.

Programming code used to perform the analyses are available upon reasonable request. The convolutional neural network used the Keras framework with a Tensorflow (Google) backend using Python 3.9. Deep learning was executed on institutional graphics processing units. All other pre- and post-processing code was written in Python 3.9 and R 4.0, which was executed locally.

RESULTS

Patient Population Baseline Characteristics

There were 734,800 ECGs (238,072 patients) between 2000–2022; after removing ECGs with missing metadata (18,539 ECGs), less experienced readers (77,363 ECGs), failed quality control (10,683 ECGs), and equivocal diagnoses (45,081 ECGs), there were 583,134 ECGs (201,620 patients) comprising the main cohort (Online Figure 1). Within the main cohort, 11% of patients had known congenital heart disease.

The training and testing cohorts comprised of 437,350 ECGs (151,215 patients; 49% male; median age 11.6 [IQR 3.1–16.9] years) and 145,784 ECGs (50,405 patients; 49% male; median age 12.0 [3.3–17.0] years), respectively. In both cohorts, 33% had more than one diagnosis; 56% had any abnormality, 1.0% had WPW, and 5.3% had prolonged QTc (Table 1). The prevalence of congenital heart disease lesions and diagnoses of other ECG findings are highlighted in Table 1.

Table 1:

Baseline Characteristics of Training and Testing Cohorts

Training Cohort Testing Cohort
Patients N=151,215 N=50,405
Sex (male) 74,015 (49%) 24,742 (49%)
Known CHD Diagnosis 17,193 (11%) 5,605 (11%)
 Cardiomyopathy 2,597 (1.7%) 937 (1.9%)
 ASD 5,052 (3.3%) 1,592 (3.2%)
 CAVC 743 (0.5%) 225 (0.4%)
 CoA 3,076 (2.0%) 1,021 (2.0%)
 DORV 1,054 (0.7%) 359 (0.7%)
 D-loop TGA 1,677 (1.1%) 584 (1.2%)
 Ebstein 467 (0.3%) 148 (0.3%)
 HLHS 1,061 (0.7%) 321 (0.6%)
 L-loop TGA 642 (0.4%) 232 (0.5%)
 Pulmonary Atresia 1,094 (0.7%) 373 (0.7%)
 TAPVC 630 (0.4%) 209 (0.4%)
 Tricuspid Atresia 443 (0.3%) 134 (0.3%)
 Truncus Arteriosus 281 (0.2%) 100 (0.2%)
 VSD 7,671 (5.1%) 2,544 (5.0%)
 Dextrocardia 619 (0.4%) 202 (0.4%)
 ToF 2,486 (1.6%) 796 (1.6%)
ECGs N=437,350 N=145,784
Age at ECG in years 11.6 (3.1, 16.9) 12.0 (3.3, 17.0)
Number of Diagnoses
 1 292,517 (67%) 97,654 (67%)
 2 86,184 (20%) 28,735 (20%)
 3 36,957 (8.5%) 12,122 (8.3%)
 4 15,333 (3.5%) 5,121 (3.5%)
 5 5,002 (1.1%) 1,719 (1.2%)
 > 5 1,357 (0.3%) 433 (0.3%)
Composite Outcome
 Any Abnormality 244,522 (56%) 81,127 (56%)
Individual Diagnoses
 RSR’ 50,218 (11%) 17,115 (12%)
 NSSTT 46,084 (11%) 15,728 (11%)
 Atrial Fibrillation 1,680 (0.4%) 680 (0.5%)
 Atrial Flutter 1,618 (0.4%) 496 (0.3%)
 High-grade AV Block 1,423 (0.3%) 398 (0.3%)
 Pericardial ST/T Wave Changes 1,344 (0.3%) 420 (0.3%)
 Ischemic ST/T Wave Changes 1,341 (0.3%) 486 (0.3%)
 Prolonged QTc 23,242 (5.3%) 7,679 (5.3%)
 Sinus Bradycardia 14,180 (3.2%) 5,137 (3.5%)
 Sinus Tachycardia 21,618 (4.9%) 7,363 (5.1%)
 SVT 1,405 (0.3%) 510 (0.3%)
 WPW 4,441 (1.0%) 1,493 (1.0%)
 Technically Inadequate Study 4,021 (0.9%) 1,415 (1.0%)
 RVH 29,857 (6.8%) 9,092 (6.2%)
 LVH 10,664 (2.4%) 3,303 (2.3%)
 T-wave Inversions 11,051 (2.5%) 3,795 (2.6%)
 CRBBB 34,134 (7.8%) 10,773 (7.4%)

Abbreviations: Congenital heart disease (CHD); atrial septal defect (ASD); complete atrioventricular canal defect (CAVC); coarctation of the aorta (CoA); double outlet right ventricle (DORV); transposition of the great arteries (TGA); hypoplastic left heart syndrome (HLHS); total anamolous pulmonary venous connection (TAPVC); ventricular septal defect (VSD); tetralogy of Fallot (ToF); electrocardiogram (ECG); non-specific ST/T wave changes (NSSTT); supraventricular tachycardia (SVT); Wolff Parkinson White syndrome (WPW); right ventricular hypertrophy (RVH); left ventricular hypertrophy (LVH); complete right bundle branch block (CRBBB).

Model Performance

Performance of the AI-ECG and AI-ECG+age+sex models to detect any abnormality, WPW, and prolonged QTc is shown in Figure 1A. Excellent performance was achieved for AI-ECG to detect any abnormality (AUROC 0.94, AUPRC 0.96), WPW (AUROC 0.99, AUPRC 0.88), and prolonged QTc (AUROC 0.96, AUPRC 0.63). Model performance was nearly identical when adding age and sex as inputs (Figure 1 and Online Figure 3), outperforming the commercial MUSE interpretation benchmark for any abnormality with more pronounced differences for WPW and prolonged QTc (Figure 1A). Sensitivity, specificity, NPV, PPV, accuracy, and F1 scores for individual ECG diagnosis are shown in Online Table 1.

Figure 1: AI-ECG Model Performance to Detect ECG Abnormalities.

Figure 1:

(A) Performance of the artificial intelligence-enhanced ECG model alone (AI-ECG; blue) and with demographic data (AI-ECG+age+sex; orange) evaluated in the testing cohorts using receiver operating (AUROC; left) and precision recall (AUPRC; right) curves for the following outcomes: any abnormality, Wolff Parkinson White syndrome (WPW), and prolonged QTc. Model benchmarked to commercial MUSE GE interpretations (grey dot). AUROC and AUPRC metric values for each model and outcome are inset below with 95% confidence intervals in brackets. Dotted line represents chance. (B) Incidence of future abnormal ECG (top), WPW (middle), or prolonged QTc (bottom) for the test cohort stratified by initial network classification (true negative [TN] in green, false positive [FP] in red). Number of patients at risk over the 1-year period inset below. Hazard ratio and log-rank p-value inset.

Patients with initial false positive AI-ECG classification were more likely to have a future abnormal ECG (hazard ratio 2.0 [95% CI, 1.8–2.2]; p<0.001), WPW (HR 88 [95% CI, 35–218]; p<0.001), and prolonged QTc (HR 3.4 [95% CI 2.9–4.0]; p<0.001) (Figure 1B).

Subgroup Analysis

In a subgroup analysis (Figure 2), there was variation in performance by age, sex, and reading provider. In general, there was lower performance for ages <3 year old, most notably for age <1 week for any abnormality and prolonged QTc. Interestingly, performance increased from <1 week to 1 month and then 6 months, after which it plateaued. Performance for ages ≥ 18 years old were similar to the overall cohort. Performance did not vary by sex. Finally, there was slightly higher performance for a pediatric electrophysiology specialist, most notably for prolonged QTc. There was no appreciable difference in performance when testing on a provider’s first 5,000 ECG reads (any abnormality: AUROC 0.93, AUPRC 0.96; WPW: AUROC 0.98, AUPRC 0.88; prolonged QTc: AUROC 0.96, AUPRC 0.61) to their last 5,000 ECG reads (any abnormality: AUROC 0.94, AUPRC 0.96; WPW: AUROC 0.99, AUPRC 0.87; prolonged QTc: AUROC 0.96, AUPRC 0.62).

Figure 2: Model Performance Across Age and Sex Subgroups.

Figure 2:

Forest plots showing area under the receiver operating (AUROC; red) and precision recall (AUPRC; black) curve performance when stratifying by age, sex, and reader for diagnosing any abnormality (left), Wolff Parkinson White syndrome (WPW; middle) and prolonged QTc (right). Abbreviations: week (w); month (mo), year (y).

Readjudication

Readjudication was performed by four blinded expert electrophysiologists on ECGs with discrepancies between the original reader and the AI-ECG classification. As shown in Figure 3A heatmap, AI-ECG, but not the original reader, clustered with the experts. Interestingly, the inter-expert Fleiss κ agreement was modest across each outcome of interest, ranging from 0.21–0.49 (all p-values not significant). The expert readers on average were more likely to agree with AI-ECG than the original reader (Figure 3B).

Figure 3: Model Classification Readjudication.

Figure 3:

ECGs deemed as false positives (AI-ECG positive but original reader negative; FP) or false negatives (AI-ECG negative but original reader positive; FN) were readjudicated. (A) Heatmap of classifications (positive = dark; negative = light) by original reader (Original), AI-ECG, and four expert readers. Hierarchical clustering of AI-ECG and readers shown above. Fleiss κ (to assess agreement across experts) shown below. (B) Cohen’s κ for agreement between the original reader and expert readjudicator (Original-Expert Agreement; black) versus the AI-ECG and expert readjudicator (AI-ECG-Expert Agreement). The average from four experts is shown with student’s t-test comparisons between Original-Exert Agreement and AI-ECG-Expert Agreement.

There was a 2–2 tie among experts for 20% of the any abnormality readjudication ECGs, 17% of the WPW readjudication ECGs, and 19% of the prolonged QTc readjudication ECGs. Readjudication results are similar when excluding the ECGs without majority vote (Online Figure 4).

Model Explainability

Model behavior analysis was performed to compare salient features noted by AI-ECG in comparison to conventional rule-based approaches for diagnostics (Figure 4). For WPW, the most salient signals are within P waves, QRS complexes, and PR signals (limb leads I-II and precordial leads V1, V5–V6). High-risk waveforms unsurprisingly demonstrate pre-excitation (a “Delta wave”) with short PR intervals. For prolonged QTc, the most salient features were QRS complexes and T waves (precordial leads V1 and V6). High-risk waveforms unsurprisingly demonstrate a prolonged QTc interval.

Figure 4: Model Explainability.

Figure 4:

Visualization of high-risk (red) and low-risk (green) median waveforms for WPW (top) and prolonged QTc (bottom). Saliency mapping demarcates ECG regions with greatest (dark blue) and least (light blue) influence on each diagnosis.

DISCUSSION

Given the proliferation of ECGs for screening in children, the burgeoning population of young patients with congenital heart disease, and the unique considerations of the pediatric ECG, it is of great interest to create a pediatric-specific ECG diagnostic model. This work represents the first (to our knowledge) application of ECG-based deep learning to automatically diagnose ECG findings in the pediatric and adult congenital population. After training on >400,000 ECGs from nearly 150,000 patients with and without cardiac abnormalities, we demonstrate model performance is excellent (AUROC >0.9) for our main outcomes of interest, outperforming commercially available MUSE software. Model performance remains robust across a range of subgroups, and readjudication of misclassified ECGs demonstrated that four blinded senior electrophysiologists were more likely to agree with AI-ECG diagnoses than an experienced reader in these boundary cases. Finally, saliency mapping findings align with conventional rule-based methodologies implemented by humans, promoting clinician trust. Altogether, these findings demonstrate the promise of AI-ECG to assist clinicians with the rapid and reliable interpretation of ECGs, which may: 1) reduce missed diagnoses by less experienced clinicians; 2) promote screening programs; 3) facilitate improved access to expert care; and 4) decrease physician workload, which may help reduce burnout.15

Current State of ECG Interpretation and Automated Diagnosis

Deep learning approaches have been leveraged to enhance ECG diagnoses broadly in the general adult population. For example, Riberio et al. demonstrated a deep neural network can outperform cardiology trainees to recognize common ECG abnormalities in adults.6 Recently, AI-ECG has been used to detect myocardial infarction,16 with a pragmatic randomized controlled trial demonstrating AI-ECG–assisted triage of ST elevation myocardial infarction decreased the door-to-balloon time for patients presenting to the emergency department.17

Utilization of adult AI-ECG models for ECG interpretation are unlikely to be well-suited for pediatric ECG diagnoses for several reasons. First, the rapid and non-linear anatomic and physiologic changes from the newborn to adulthood make adult AI-ECG models not amenable to pediatric generalization, as: 1) a right axis deviation in newborns can be normal and reflect in utero physiology; 2) normal heart rate ranges evolve through infancy through adolescence; 3) T waves are upright in the first weeks of life, then invert, then gradually become upright during adolescence; 4) axes and intervals vary by age and sex.18 Second, there is a disproportionate prevalence of patients with congenital heart lesions with highly prevalent and specific ECG abnormalities that are rare in adult populations (e.g., superior axis deviation). Finally, the clinical questions commonly asked of the adult ECG (e.g., the presence of ischemia/infarction) are rarely encountered in pediatrics, with pediatric use cases gravitating toward screening applications for evaluation of non-specific symptoms and identification of rare ECG findings such as WPW and long QT syndrome.

Benchmarking to Commercial Software

To our knowledge, there are no published studies that validate the accuracy of MUSE interpretation for pediatric patients. This prompted us to include the MUSE interpretations as a benchmark (Figure 1A). As shown in Figure 1A, our model outperforms MUSE across all outcomes of interest with more prominent differences for highly relevant pediatric diagnoses (e.g., WPW, prolonged QTc).

Clinical Significance and Implications

From a clinical and translational perspective, we envision this rapid, reliable, and automated pediatric and congenital AI-ECG model may: 1) facilitate improved access to care; 2) promote screening programs; and 3) enhance physician workflow to reduce workload/burnout.

While pediatric and congenital heart disease is relatively common, unfortunately over 90% of children in low- and middle-income countries to not have access to cardiovascular care.19 Similarly, there are geographic and socioeconomic disparities in access to care from expert cardiologists in the United States.20 Delays in expert ECG reads may cause delays in care or missed care opportunities. Given that ECG is a ubiquitous and inexpensive diagnostic test, we envision this tool may help democratize pediatric cardiology care by providing ECG interpretation for all, and aid in prioritizing patients for referral. We envision this framework particularly for screening of WPW and prolonged QTc in areas with limited access to care. In addition, AI-ECG would inform the provider if the ECG is normal or abnormal, and for ECGs deemed abnormal, a list of predicted ECG diagnoses would be available.

Large-scale ECG screening has been considered for decades within the pediatric cardiology community for sudden cardiac arrest and young athlete sports clearance.3,18,21,22 However, the lack of experienced workforce and lack of inter-reader consensus/reproducibility has undermined these efforts. It is conceivable that AI-ECG may help address these limitations to enable such screening efforts, which may also enhance physician workflow to reduce workload/burnout. For example, since ~50% of ECGs are normal and our AI-ECG tool performs at an expert level to detect normal ECGs, then conceivably this tool may reduce the total number of ECGs reviewed by an expert in half. Of note, such promise to reduce workload has been shown for AI automated quantification of left ventricular ejection fraction from echocardiography.23

Model Insights into ECG Interpretation

Several insights were developed from model performance and behavior analysis. First, we note that AI-ECG performed similarly to AI-ECG+age+sex (Figure 1A) for primary/secondary outcomes and sinus tachycardia/bradycardia (Online Figure 3), suggesting the model may also learn this demographic data that is known to relate to changes in ECG axis and intervals.2,18,24,25 Second, we observed that false positive predictions by AI-ECG were more likely to be positive on ECGs soon thereafter (with a hazard ratio of 88 for WPW), suggesting either initial misdiagnoses by the reader or AI-ECG detecting subtle findings that become more pronounced on follow-up ECGs (e.g., intermittent pre-excitation; borderline prolonged QTc). This is supported by the low inter-expert κ achieved for each outcome (Figure 3A), which suggests discrepant reads between the original reader and AI-ECG are borderline or difficult boundary cases with significant disagreement even amongst expert readers. Finally, we note the lower performance of any abnormality and prolonged QTc in newborns (<1 week), which is consistent with the literature26 and may be attributed to the rapid evolution of the normal ECG in the first weeks of life. Indeed, performance increased from 1 week to 1 month and then 6 months, after which it plateaued. In contrast, performance for adults were comparable to the overall cohort, suggestive the model works across the lifespan.

Limitations

Several limitations are noted. First, we acknowledge the lack of external validation. Unfortunately, the sizeable network of outside pediatric institutions contacted do not have similar infrastructures available for coding ECG diagnoses, with most utilizing free text rather than coded diagnoses. While there are public ECG waveforms and diagnoses available,6 these datasets exclusively represent the general adult population, which is outside the scope of this study. Future work therefore includes multicenter collaboration (via federated learning27 to compile a larger and more heterogeneous training set) and external validation. Second, our model architecture requires access to digital waveform data. This limits translation to low resource settings where digitized data is less accessible and motivates future efforts to generate a model that uses ECG image inputs.28 Third, our threshold to optimize F1 was used to represent a human reader’s behavior when interpreting an ECG; however, if considering a screening tool, other thresholds may be considered (e.g., prioritizing negative predictive value or sensitivity). Further consideration is required to weigh the impact of resultant false positives/negatives. Fourth, other model architectures (e.g., foundational vision transformer model) must be considered for improving performance. Fifth, we acknowledge that a physician’s access to history of presenting illness may further enhance ECG interpretation. Finally, the limitations of saliency mapping and model explainability are acknowledged.29

Future Directions

Translation of AI-ECG to clinical practice will require several additional steps, including: 1) multicenter collaboration to improve the power and generalizability of this model; 2) pragmatic randomized clinical trials30 to study the accuracy and safety of AI-ECG and guide clinical implementation; 3) cost-benefit analyses to determine optimal application of pediatric and congenital AI-ECG models to various carefully crafted clinical use cases for screening and disease management; 4) generation of AI-ECG models that take ECG photo inputs (rather than digital waveform inputs) to broaden global accessibility.28

Conclusions

These findings demonstrate the promise of AI-ECG to interpret ECGs rapidly and reliably. This tool may facilitate larger screening program efforts, improved access to care, decrease physician workload, and potentially even improve accuracy of ECG reading.

Supplementary Material

1

Central Illustration: AI-ECG for Expert-Level Interpretation of the Pediatric ECG.

Central Illustration:

An artificial intelligence-enhanced ECG (AI-ECG) convolutional neural network (CNN) algorithm trained on a diverse pediatric cohort at Boston Children’s Hospital outperformed commercial software to predict any ECG abnormality, Wolff Parkinson White (WPW) syndrome, and prolonged QTc. Blinded human experts were more likely to agree with AI-ECG than the original reader.

CLINICAL PERSPECTIVES.

Competency in Patient Care:

Artificial intelligence-enabled ECG analysis shows promise to provide expert-level automated diagnosis of the pediatric 12-lead ECG. This model outperformed commercial software, and blinded experts were more likely to agree with the model than the original reader to detect any ECG abnormality, Wolff Parkinson White syndrome, and prolonged QTc.

Translational Outlook:

Our model provides expert-level automated diagnosis of the pediatric 12-lead ECG, which may promote screening programs, facilitate improved access to expert care, and decrease physician workload.

Acknowledgments:

The authors would like to acknowledge Boston Children’s Hospital’s High-Performance Computing Resources Clusters Enkefalos 2 (E2) made available for conducting the research reported in this publication.

Funding:

Funding support received from the Thrasher Research Fund Early Career Award (J.M.), Boston Children’s Hospital Electrophysiology Research Education Fund and Kostin Innovation Fund (J.M., J.K.T.) and NIH grant R00-LM012926 from the National Library of Medicine (W.G.L.).

Abbreviations:

AI-ECG

Artificial intelligence-enhanced electrocardiogram

AUROC

Area under the receiver operating curve

AUPRC

Area under the precision-recall curve

WPW

Wolff Parkinson White syndrome

PPV

Positive predictive value

NPV

Negative predictive value

ECG

Electrocardiogram

Footnotes

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

SUPPLEMENTAL MATERIAL:

Online Figures 1–4 and Online Table 1

Disclosures: None

REFERENCES

  • 1.Khairy P, Marelli AJ. Clinical use of electrocardiography in adults with congenital heart disease. Circulation. Dec 4 2007;116(23):2734–46. doi: 10.1161/CIRCULATIONAHA.107.691568 [DOI] [PubMed] [Google Scholar]
  • 2.Saarel EV, Granger S, Kaltman JR, et al. Electrocardiograms in Healthy North American Children in the Digital Age. Circ Arrhythm Electrophysiol. Jul 2018;11(7):e005808. doi: 10.1161/CIRCEP.117.005808 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Myerburg RJ, Vetter VL. Electrocardiograms should be included in preparticipation screening of athletes. Circulation. Nov 27 2007;116(22):2616–26; discussion 2626. doi: 10.1161/CIRCULATIONAHA.107.733519 [DOI] [PubMed] [Google Scholar]
  • 4.Vetter VL, Elia J, Erickson C, et al. Cardiovascular monitoring of children and adolescents with heart disease receiving medications for attention deficit/hyperactivity disorder [corrected]: a scientific statement from the American Heart Association Council on Cardiovascular Disease in the Young Congenital Cardiac Defects Committee and the Council on Cardiovascular Nursing. Circulation. May 6 2008;117(18):2407–23. doi: 10.1161/CIRCULATIONAHA.107.189473 [DOI] [PubMed] [Google Scholar]
  • 5.Vetter VL. Electrocardiographic screening of all infants, children, and teenagers should be performed. Circulation. Aug 19 2014;130(8):688–97; discussion 697. doi: 10.1161/CIRCULATIONAHA.114.009737 [DOI] [PubMed] [Google Scholar]
  • 6.Ribeiro AH, Ribeiro MH, Paixao GMM, et al. Automatic diagnosis of the 12-lead ECG using a deep neural network. Nat Commun. Apr 9 2020;11(1):1760. doi: 10.1038/s41467-020-15432-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Mayourian J, La Cava WG, Vaid A, et al. Pediatric ECG-Based Deep Learning to Predict Left Ventricular Dysfunction and Remodeling. Circulation. Mar 19 2024;149(12):917–931. doi: 10.1161/CIRCULATIONAHA.123.067750 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Mayourian J, Gearhart A, La Cava WG, et al. Deep Learning-Based Electrocardiogram Analysis Predicts Biventricular Dysfunction and Dilation in Congenital Heart Disease. J Am Coll Cardiol. Aug 27 2024;84(9):815–828. doi: 10.1016/j.jacc.2024.05.062 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Einthoven W Weiteres über das Elektrokardiogramm. Archiv für die gesamte Physiologie des Menschen und der Tiere. 1908/05/01 1908;122(12):517–584. doi: 10.1007/BF01677829 [DOI] [Google Scholar]
  • 10.Goldberger E A simple, indifferent, electrocardiographic electrode of zero potential and a technique of obtaining augmented, unipolar, extremity leads. American Heart Journal. 1942/04/01/ 1942;23(4):483–492. doi: 10.1016/S0002-8703(42)90293-X [DOI] [Google Scholar]
  • 11.Colan SD. Early Database Initiatives: The Fyler Codes. In: Barach PR, Jacobs JP, Lipshultz SE, Laussen PC, eds. Pediatric and Congenital Cardiac Care: Volume 1: Outcomes Analysis. Springer; London; 2015:163–169. [Google Scholar]
  • 12.Sangha V, Mortazavi BJ, Haimovich AD, et al. Automated multilabel diagnosis on electrocardiographic images and signals. Nat Commun. Mar 24 2022;13(1):1583. doi: 10.1038/s41467-022-29153-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.McHugh ML. Interrater reliability: the kappa statistic. Biochem Med (Zagreb). 2012;22(3):276–82. [PMC free article] [PubMed] [Google Scholar]
  • 14.Lundberg SM, Erion G, Chen H, et al. From Local Explanations to Global Understanding with Explainable AI for Trees. Nat Mach Intell. Jan 2020;2(1):56–67. doi: 10.1038/s42256-019-0138-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.McCormick AD, Lim HM, Strohacker CM, et al. Paediatric cardiology training: burnout, fulfilment, and fears. Cardiol Young. Nov 2023;33(11):2274–2281. doi: 10.1017/S1047951123000148 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Gustafsson S, Gedon D, Lampa E, et al. Development and validation of deep learning ECG-based prediction of myocardial infarction in emergency department patients. Sci Rep. Nov 15 2022;12(1):19615. doi: 10.1038/s41598-022-24254-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Lin C, Liu W-T, Chang C-H, et al. Artificial Intelligence–Powered Rapid Identification of ST-Elevation Myocardial Infarction via Electrocardiogram (ARISE) — A Pragmatic Randomized Controlled Trial. NEJM AI. 2024;1(7):AIoa2400190. doi:doi: 10.1056/AIoa2400190 [DOI] [Google Scholar]
  • 18.Bratincsak A, Kimata C, Limm-Chan BN, Vincent KP, Williams MR, Perry JC. Electrocardiogram Standards for Children and Young Adults Using Z-Scores. Circ Arrhythm Electrophysiol. Aug 2020;13(8):e008253. doi: 10.1161/CIRCEP.119.008253 [DOI] [PubMed] [Google Scholar]
  • 19.Zheleva B, Atwood JB. The invisible child: childhood heart disease in global health. Lancet. Jan 7 2017;389(10064):16–18. doi: 10.1016/S0140-6736(16)32185-7 [DOI] [PubMed] [Google Scholar]
  • 20.Kim JH, Cisneros T, Nguyen A, van Meijgaard J, Warraich HJ. Geographic Disparities in Access to Cardiologists in the United States. J Am Coll Cardiol. Jul 16 2024;84(3):315–316. doi: 10.1016/j.jacc.2024.04.054 [DOI] [PubMed] [Google Scholar]
  • 21.Chaitman BR. An electrocardiogram should not be included in routine preparticipation screening of young athletes. Circulation. Nov 27 2007;116(22):2610–4; discussion 2615. doi: 10.1161/CIRCULATIONAHA.107.711465 [DOI] [PubMed] [Google Scholar]
  • 22.Halkin A, Steinvil A, Rosso R, Adler A, Rozovski U, Viskin S. Preventing sudden death of athletes with electrocardiographic screening: what is the absolute benefit and how much will it cost? J Am Coll Cardiol. Dec 4 2012;60(22):2271–6. doi: 10.1016/j.jacc.2012.09.003 [DOI] [PubMed] [Google Scholar]
  • 23.He B, Kwan AC, Cho JH, et al. Blinded, randomized trial of sonographer versus AI cardiac function assessment. Nature. Apr 2023;616(7957):520–524. doi: 10.1038/s41586-023-05947-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.O’Sullivan D, Anjewierden S, Greason G, et al. Pediatric sex estimation using AI-enabled ECG analysis: influence of pubertal development. NPJ Digit Med. Jul 2 2024;7(1):176. doi: 10.1038/s41746-024-01165-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Dickinson DF. The normal ECG in childhood and adolescence. Heart. Dec 2005;91(12):1626–30. doi: 10.1136/hrt.2004.057307 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Stramba-Badiale M, Karnad DR, Goulene KM, et al. For neonatal ECG screening there is no reason to relinquish old Bazett’s correction. Eur Heart J. Aug 14 2018;39(31):2888–2895. doi: 10.1093/eurheartj/ehy284 [DOI] [PubMed] [Google Scholar]
  • 27.Goto S, Solanki D, John JE, et al. Multinational Federated Learning Approach to Train ECG and Echocardiogram Models for Hypertrophic Cardiomyopathy Detection. Circulation. Sep 6 2022;146(10):755–769. doi: 10.1161/CIRCULATIONAHA.121.058696 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Sangha V, Nargesi AA, Dhingra LS, et al. Detection of Left Ventricular Systolic Dysfunction From Electrocardiographic Images. Circulation. Aug 29 2023;148(9):765–777. doi: 10.1161/CIRCULATIONAHA.122.062646 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Ghassemi M, Oakden-Rayner L, Beam AL. The false hope of current approaches to explainable artificial intelligence in health care. Lancet Digit Health. Nov 2021;3(11):e745–e750. doi: 10.1016/S2589-7500(21)00208-9 [DOI] [PubMed] [Google Scholar]
  • 30.Yao X, Rushlow DR, Inselman JW, et al. Artificial intelligence-enabled electrocardiograms for identification of patients with low ejection fraction: a pragmatic, randomized clinical trial. Nat Med. May 2021;27(5):815–819. doi: 10.1038/s41591-021-01335-4 [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1

Data Availability Statement

Requests for Boston Children’s Hospital data and related materials will be internally reviewed to clarify if the request is subject to intellectual property or confidentiality constraints. Shareable data and materials will be released under a material transfer agreement for non-commercial research purposes. Use of Boston Children’s Hospital data was approved by their respective Institutional Review Boards.

Programming code used to perform the analyses are available upon reasonable request. The convolutional neural network used the Keras framework with a Tensorflow (Google) backend using Python 3.9. Deep learning was executed on institutional graphics processing units. All other pre- and post-processing code was written in Python 3.9 and R 4.0, which was executed locally.

RESOURCES