Abstract
Background
Transthoracic echocardiography is the first-line test for congenital heart disease (CHD), but accurate targeted triage and lesion subtyping require expertise and synthesis across multiple heterogeneous cine clips. We developed a video-based deep learning approach for examination-level targeted triage and subtype classification of common left-to-right shunt lesions.
Methods
We retrospectively assembled 2,373 echocardiography examinations (800 normal; 493 VSD, 751 ASD, and 329 PDA) from The First Affiliated Hospital of Xinjiang Medical University. Each examination comprised multiple cine clips acquired across standard views. We proposed Echocardiography Hierarchical Multiple-Instance Learning network (EchoHMIL), which encodes clips with a spatiotemporal backbone and aggregates variable numbers of clips using attention-based multi-instance learning to form an examination-level representation. A hierarchical dual-head design was optimized with a gated multi-task objective to perform: (i) normal-versus-target-shunt discrimination, and (ii) VSD/ASD/PDA subtype classification conditional on target-shunt status. Data were split at the patient level (70%/10%/20%) and evaluated on an independent test set using AUC, sensitivity, specificity, accuracy, and macro-F1, with bootstrap 95% confidence intervals. The task was explicitly framed as a closed-set targeted triage and subtype-classification problem for three common left-to-right shunt lesions, not as comprehensive pediatric CHD screening.
Results
On the test set (n = 475), EchoHMIL achieved an AUC of 0.957 for normal-versus-target-shunt discrimination. At a sensitivity-prioritized operating point, sensitivity was 92.2% and specificity was 82.4%. For VSD/ASD/PDA subtype classification among target-shunt cases, EchoHMIL achieved an overall accuracy of 88.8% with a macro-F1 of 0.885. Attention weights and gradient-based saliency maps highlighted clinically plausible regions associated with septal and ductal anatomy.
Conclusions
EchoHMIL enables automated examination-level triage and subtype classification of common left-to-right shunt lesions from routine echocardiography videos. Further validation on complex CHD and broader out-of-distribution abnormalities is required before extension to general CHD screening. These findings should therefore be interpreted within a closed-set target-lesion setting; prospective multicenter validation including complex, mixed, postoperative, and out-of-distribution abnormalities is required before any extension to general CHD screening or routine clinical deployment.
Keywords: Congenital heart disease, Deep learning, Hierarchical classification, Multiple-instance learning
Introduction
Congenital heart disease (CHD) is among the most prevalent congenital anomalies worldwide and remains a major contributor to infant and pediatric morbidity. Large-scale meta-analyses have consistently reported a substantial global birth prevalence, underscoring the need for screening and diagnostic strategies that can be delivered reliably across diverse clinical settings [1, 2]. Although advances in neonatal care, surgical techniques, and long-term follow-up have improved outcomes, these improvements have also increased the demand for timely diagnosis and appropriate triage, particularly in regions where access to subspecialty cardiology is limited and workload pressures are substantial.
Transthoracic echocardiography (TTE) is the first-line imaging modality for evaluating suspected congenital heart disease and common left-to-right shunt lesions because it is noninvasive, radiation-free, relatively inexpensive, and capable of providing real-time structural and hemodynamic information. Contemporary practice guidelines emphasize that a comprehensive pediatric TTE examination should be interpreted across standardized views and measurements, with flexible adaptation to anatomy and acoustic windows [3]. In routine practice, however, echocardiography is not a single image but an examination assembled from multiple cine clips acquired across different windows, often under imperfect conditions. Interpretation is therefore highly operator-dependent and sensitive to heterogeneous acquisition protocols, variable image quality, motion artifacts, and inter-operator variability—limitations that are amplified in high-volume workflows, during off-hours coverage, and in settings where experienced echocardiographers are scarce [4].
Within the heterogeneous spectrum of CHD, left-to-right shunt lesions are frequently encountered and clinically important. Ventricular septal defect (VSD), atrial septal defect (ASD), and patent ductus arteriosus (PDA) represent common diagnoses that differ meaningfully in natural history, hemodynamic consequences, and management strategy. Although these entities may be conceptually “straightforward,” accurate differentiation in routine echocardiography can be challenging. Diagnosis often depends on integrating information across multiple views and clips, recognizing subtle structural discontinuities, and interpreting flow-related patterns that may vary with defect size, loading conditions, and imaging plane [5]. In children, higher heart rates and smaller anatomic dimensions further compress the temporal and spatial cues available to the interpreter, increasing the need for robust spatiotemporal reasoning.
From a workflow standpoint, echocardiography interpretation frequently follows a hierarchical decision logic. A primary question is whether a study is normal or abnormal (screening). Only after abnormality is suspected does the interpreter proceed to refined lesion-specific reasoning (subtyping), which may involve targeted views, focused Doppler interrogation, and cross-clip synthesis. This “screening then subtyping” paradigm is not merely a cognitive convenience; it reflects how examinations are acquired and interpreted, and it shapes how decision support should be designed. Systems that collapse the process into a single flat decision without respecting examination structure can be brittle, especially when clip quality is heterogeneous or when the diagnostic signal is concentrated in a small subset of clips.
Recent advances in deep learning have demonstrated strong potential for automating echocardiographic workflows, including view recognition and disease-related classification [6], comprehensive interpretation and phenotype prediction from echocardiograms [7], and video-based estimation of cardiac function [8]. Beyond imaging-only tasks, video models have been shown to infer physiologic biomarkers from echocardiogram videos, suggesting that clinically relevant information is encoded in subtle motion and texture patterns that may be difficult to quantify manually [9]. Deep learning guidance has also been used to support image acquisition by less experienced operators, highlighting the broader role that AI may play across the acquisition–interpretation continuum [10]. These developments align with broader trends in medical imaging AI and underscore the advantages of representation learning over hand-crafted features, particularly when data are complex and high-dimensional [11].
Nevertheless, translating echocardiographic deep learning into CHD diagnosis introduces challenges that are not fully addressed by many existing pipelines. First, CHD examinations are inherently multi-clip and often multi-view; a model that performs well on carefully curated single-view clips may not generalize to the clinical reality of variable clip length, uneven view coverage, and frequent “near-miss” acquisitions. Second, the pediatric domain is characterized by wide variation in body size, heart rate, and probe handling, all of which affect image statistics and motion patterns. Third, CHD labels are often long-tailed, with relatively fewer cases for specific subtypes at many institutions, making robust training and evaluation difficult. Finally, many reported systems focus on single-task settings rather than clinically aligned end-to-end diagnosis that must explicitly distinguish normal from multiple common defects within one pathway, which can limit utility in real-world triage [12].
Accordingly, the objective of this study was deliberately narrowed to a targeted and clinically interpretable question: whether routine multi-clip TTE videos can support examination-level triage between normal studies and three common left-to-right shunt lesions, followed by subtype classification among ASD, VSD, and PDA. This design should not be interpreted as solving the broader and more heterogeneous problem of pediatric CHD screening, which would require inclusion of complex CHD, mixed lesions, postoperative anatomy, acquired abnormalities, and explicit out-of-distribution testing.
Methodologically, two observations motivate our approach. The first is that echocardiography is fundamentally a video problem: spatiotemporal cues—valvular motion, septal continuity across frames, and view-dependent dynamics—often carry decisive diagnostic information. Modern video architectures provide practical tools for modeling such signals, including factorized spatiotemporal convolutions [13], efficient video networks [14], and transformer-based backbones that improve long-range temporal modeling and robustness to variation in clip structure [15]. Self-supervised video pretraining further improves data efficiency in domains where labeled data are limited, providing an attractive pathway for CHD applications [16].
The second observation is that an echocardiographic examination should be treated as a set, or “bag,” of clips rather than a single clip. In many examinations, only a subset of clips captures the most informative plane or the clearest depiction of the relevant anatomy. Attention-based multiple-instance learning (MIL) offers a principled mechanism for aggregating a variable number of instances into an examination-level prediction while learning to emphasize diagnostically informative clips [17]. This formulation is particularly appealing for echocardiography because it naturally accommodates variable clip counts, heterogeneous quality, and real-world acquisition patterns. Importantly, learned attention weights can also provide examination-level interpretability by indicating which parts of the study most influenced the decision.
In this study, we propose EchoHMIL, a hierarchical MIL framework for echocardiography cine videos that supports clinically aligned decision-making at the examination level. EchoHMIL integrates spatiotemporal video representation learning with attention-based aggregation across multi-view cine clips and produces an end-to-end four-class diagnosis (Normal/VSD/ASD/PDA) under a clinically consistent hierarchical inference rule. By modeling each examination as a collection of heterogeneous clips and preserving the “screening then subtyping” structure during training and inference, the framework is designed to be resilient to variable clip quality and incomplete view coverage, which are common in routine TTE.
The novelty of this study lies not in proposing an entirely new video backbone, attention-based MIL module, or hierarchical classifier. Rather, the contribution is the clinically grounded formulation, implementation, and evaluation of an examination-level video-MIL framework for targeted triage and subtype classification of common congenital left-to-right shunt lesions. We therefore reposition EchoHMIL as a workflow-aligned application and validation study, supported by ablation, calibration, robustness, interpretability, and reader-comparison analyses, rather than as a claim of fundamental algorithmic novelty.
We define a targeted hierarchical diagnostic task that reflects routine echocardiographic reasoning: distinguishing normal examinations from target shunt lesions, followed by subtype classification of VSD, ASD, and PDA.
We model each echocardiographic examination as a variable-size set of multi-view cine clips, enabling examination-level inference under heterogeneous view coverage and clip quality.
We systematically evaluate this framework in a dedicated cohort of common shunt lesions using discrimination, calibration, ablation, robustness, and interpretability analyses.
Materials and methods
Study design and population
This retrospective study was conducted at The First Affiliated Hospital of Xinjiang Medical University and included consecutive patients who underwent TTE between September 1,2020, and September 1,2025. All echocardiographic examinations were reviewed and categorized into four groups: normal controls and CHD cases with ASD, VSD, or PDA.
Each examination could contain multiple cine clips acquired from standard echocardiographic views, including parasternal short-axis at the great artery level (PSAX-GA), subcostal biatrial view, apical four-chamber view (A4C), suprasternal view, and parasternal long-axis view (PLAX). Examinations were excluded if they were incomplete, non-diagnostic because of poor image quality, or demonstrated complex congenital heart disease beyond the target categories.
The cohort was intentionally constructed as a closed-set target-lesion dataset. Excluded categories included complex CHD, mixed shunt lesions, postoperative anatomy, non-shunt congenital lesions, acquired pediatric cardiac disease, and examinations with severe out-of-distribution image characteristics. This restriction was used to evaluate proof-of-concept performance for common left-to-right shunt lesions under controlled reference-standard conditions, but it limits the clinical generalizability of the model and precludes claims of broad CHD screening.
This study was approved by the institutional ethics committee of The First Affiliated Hospital of Xinjiang Medical University, and the requirement for informed consent was waived because of the retrospective design.
The baseline clinical characteristics of the enrolled population are summarized in Table 1. These data provide the clinical context of the study cohort and show significant between-group differences in age, sex distribution, cardiovascular comorbidities, ventricular function, and chamber remodeling patterns. Therefore, the binary task in this study should be interpreted as discrimination between normal examinations and predefined target shunt lesions, rather than general CHD screening.
Table 1.
Baseline clinical characteristics of the study population
| Variable | Study group | p value | |||
|---|---|---|---|---|---|
| ASD(n = 751) | Normal(n = 800) | VSD(n = 493) | PDA(n = 329) | ||
| Age, years | 35.0(10.0—54.0) | 49.0(33.0—60.0) | 6.0(2.0—30.0) | 5.0(0.1—34.0) | <0.001 |
| Male sex, n (%) | 265 (35.6%) | 348 (43.6%) | 224 (46.0%) | 120 (36.4%) | <0.001 |
| Hypertension, n (%) | 134 (18.0%) | 34 (4.2%) | 55 (11.3%) | 39 (11.8%) | <0.001 |
| Diabetes mellitus, n (%) | 6 (0.8%) | 17 (2.1%) | 1 (0.2%) | 0 (0.0%) | <0.001 |
| Hyperlipidemia, n (%) | 4 (0.5%) | 7 (0.9%) | 1 (0.2%) | 6 (1.8%) | 0.058 |
| Coronary artery disease, n (%) | 209 (28.1%) | 59 (7.4%) | 55 (11.3%) | 49 (14.8%) | <0.001 |
| Myocardial infarction, n (%) | 10 (1.3%) | 0 (0.0%) | 3 (0.6%) | 1 (0.3%) | 0.006 |
| LVEF, % | 67.0(62.0—72.0) | 63.0(62.0—64.0) | 75.0(67.0—76.0) | 76.0(65.0—77.0) | <0.001 |
| Pulmonary hypertension, n (%) | 199 (26.7%) | 0 (0.0%) | 69 (14.2%) | 56 (17.0%) | <0.001 |
| Cerebral infarction, n (%) | 1 (0.1%) | 23 (2.9%) | 0 (0.0%) | 0 (0.0%) | <0.001 |
| Left heart enlargement, n (%) | 0 (0.0%) | 0 (0.0%) | 62 (12.7%) | 83 (25.2%) | <0.001 |
| Right heart enlargement, n (%) | 224 (30.1%) | 0 (0.0%) | 19 (3.9%) | 19 (5.8%) | <0.001 |
| Abnormal FAC, n (%) | 38 (5.1%) | 0 (0.0%) | 11 (2.3%) | 13 (3.9%) | <0.001 |
Data are presented as median (interquartile range) or n (%). p values were calculated using the Kruskal–Wallis test for continuous variables and the chi-square test for categorical variables
A value of 0 indicates that no such cases were observed in the current study cohort and should not be interpreted as evidence of absolute clinical impossibility in the underlying disease population
Abbreviations: FAC, fractional area change; LVEF, left ventricular ejection fraction
Reference standard and annotation
Ground-truth labels were established at the examination level using an expert-adjudicated reference standard. All cine clips were de-identified and reviewed in randomized order. For each patient, the full available echocardiographic study, including multi-view cine clips and Doppler information when available, was reviewed rather than any single clip alone. Two experienced echocardiographers independently assigned each examination to one of four categories: Normal, VSD, ASD, or PDA. Readers were blinded to model outputs and to each other’s assessments. Initial interobserver agreement was high, with Cohen’s kappa values of 0.918 for normal-versus-target-shunt classification and 0.862 for four-class diagnosis. Disagreements were resolved by consensus review; if consensus could not be reached, a third senior echocardiographer adjudicated the final label. Surgical or catheterization confirmation was available in 769 of 1,573 target-shunt cases, 48.9%, and was used to support final adjudication when available. For cases without invasive confirmation, the expert-adjudicated echocardiographic diagnosis served as the reference standard. Examinations with severely inadequate visualization, incomplete acquisition, or complex CHD beyond the target categories were excluded. Doppler information was used only for reference-standard adjudication when available and was not explicitly used as a separate model input. This reference-standard strategy was designed for targeted internal validation; cases without invasive confirmation and all excluded disease categories were acknowledged as sources of residual label uncertainty and limited external generalizability.
Model architecture: EchoHMIL
We proposed EchoHMIL, a hierarchical multi-instance learning framework for automated CHD diagnosis from echocardiography cine videos. As illustrated in Fig. 1a, EchoHMIL represents each patient examination as a set of multi-view cine clips, extracts clip-level spatiotemporal representations using a video backbone, aggregates them into an examination-level embedding via attention-based multiple instance learning (MIL) pooling, and outputs predictions through two task-specific heads. The attention-based MIL pooling mechanism is detailed in Fig. 1b, and the hierarchical dual-head prediction with gated training/inference is shown in Fig. 1c. EchoHMIL is designed to align with the clinical workflow by first distinguishing normal examinations from target left-to-right shunt lesions and then subtype differentiation (VSD/ASD/PDA) for target-shunt cases for abnormal cases.
Fig. 1.

Overview of EchoHMIL for hierarchical diagnosis of congenital heart disease (CHD) from echocardiography videos. (a) End-to-end pipeline that transforms multi-view cine clips into an examination-level prediction. (b) Attention-based multiple-instance learning (MIL) pooling module for robust integration of heterogeneous clips. (c) Hierarchical classification heads with gated multi-task training and a clinically consistent inference strategy
Problem formulation
Let p∈{1, …, N}index patients (examinations). The p-th examination contains cine clips acquired from multiple standard views (e.g., parasternal short-axis at the great artery level, subcostal biatrial view, apical four-chamber view, suprasternal view, and parasternal long-axis view). We denote the set of available clips as
![]() |
1 |
where M_p is the number of clips for patient p. During training, we optionally sample a subset of size K_p ≤ M_p from X_p for computational efficiency. We define a binary label
{0,1} (0: normal, 1: target-shunt lesion). For target-shunt examinations, a subtype label
{1,2,3} corresponds to VSD, ASD, and PDA, respectively.
Clip-level spatiotemporal encoding
Each cine clip
is encoded by a spatiotemporal feature extractor fθ(·)(e.g., Video Swin Transformer, R(2+1) D, or X3D) to obtain a d-dimensional embedding:
![]() |
2 |
Attention-based MIL aggregation (examination-level representation)
To obtain a robust examination-level representation and reduce sensitivity to low-quality or less informative clips, we adopt attention-based MIL pooling (Fig. 1b). An attention scoring function computes the unnormalized importance score for each clip:
![]() |
3 |
where V ∈ ℝ^r×d, w ∈ ℝ^r, and b ∈ ℝ^r are learnable parameters. The attention weights are normalized via softmax:
![]() |
4 |
The examination-level embedding is computed as:
![]() |
5 |
Hierarchical dual-head prediction
Two prediction heads operate on Hp (Fig. 1c). The binary head outputs a target-shunt probability:
![]() |
6 |
and the subtype head outputs a three-class probability distribution:
![]() |
7 |
where σ(·) is the sigmoid function, Wb ∈ ℝ^1×d, bb ∈ ℝ, Ws ∈ ℝ^3×d, and bs ∈ ℝ^3
Training objective and inference
EchoHMIL was trained using a gated multi-task objective. The normal-versus-target-shunt head was optimized with weighted binary cross-entropy for all examinations, whereas the subtype head was optimized with class-weighted cross-entropy only for target-shunt cases. During inference, the binary head first determined whether an examination was routed to the normal or target-shunt pathway. Examinations routed to the target-shunt pathway were then assigned the subtype with the highest subtype probability among VSD, ASD, and PDA.
Training protocol
All experiments were performed using strict patient/examination-level splitting to prevent data leakage. The cohort (N = 2,373) was randomly divided into training, validation, and test sets at a ratio of 70%/10%/20% (1,661/237/475 patients), stratified by diagnostic category. Splitting was performed before any clip sampling, frame extraction, preprocessing, or augmentation. All cine clips, frames, and augmented samples from the same echocardiographic examination were confined to the same split, ensuring that temporally correlated clips from one acquisition session did not appear across training, validation, and test sets. For each examination, up to Kp = 12 cine clips were randomly sampled per epoch, with view-balanced sampling when possible. Each clip was temporally sampled to T = 32 frames and resized to H × W = 224 × 224. The model was trained using the Adam optimizer (β1 = 0.9, β2 = 0.999), with an initial learning rate of 1 × 10^-4, weight decay of 1 × 10^-4, and a batch size of 8 examinations. The learning rate was reduced by a factor of 0.1 if validation loss did not improve for 5 consecutive epochs, and early stopping was applied with a patience of 10 epochs, with a maximum of 80 epochs. To address class imbalance, weighted binary cross-entropy was used for the normal-versus-target-shunt head, and class-weighted cross-entropy was used for the subtype head. The loss balancing coefficient λ was set to 1.0 by default and tuned on the validation set among {0.5, 1.0, 2.0}. The operating threshold τ was selected on the validation set to prioritize sensitivity for target-shunt detection and was fixed before evaluation on the independent test set. To reduce potential shortcut learning from acquisition patterns, splitting was completed before view selection, clip sampling, frame extraction, and augmentation, and all temporally correlated clips from one examination were kept within the same split. In addition, view labels and clinical reports were not provided as explicit model inputs, and view-balanced sampling was used whenever multiple diagnostic views were available so that the model was not preferentially trained on diagnosis-specific acquisition frequency alone.
Evaluation metrics
For the normal-versus-target-shunt task, we reported the area under the receiver operating characteristic curve (AUC), sensitivity, specificity, accuracy, precision, recall, and F1-score. For subtype classification among target-shunt cases, we reported overall accuracy, macro-averaged F1-score, and per-class precision/recall/F1, together with the confusion matrix. All metrics were computed at the patient (examination) level. We estimated 95% confidence intervals (CIs) using non-parametric bootstrapping with 1000 resamples on the test set.
To specifically address acquisition-related shortcut learning, additional sensitivity analyses were performed across acquisition-derived strata, including view coverage, cine-clip count, vendor/device group when available, and image-quality category. We also evaluated performance in view-coverage-matched and clip-count-matched subsets when sufficient samples were available. These analyses were designed to quantify, but not completely eliminate, the possibility that acquisition patterns contributed to model predictions.
Human-reader comparison
To provide a clinically meaningful baseline, we conducted a human-reader comparison study on a randomly selected subset of the independent test set. The subset included examinations from all four diagnostic categories: Normal, VSD, ASD, and PDA. Three echocardiographers with different levels of experience participated in the evaluation: one junior reader with 4 years of experience, one intermediate reader with 9 years, and one senior reader with 15 years of congenital echocardiography experience. All readers were blinded to the model predictions, original clinical reports, and expert-adjudicated reference labels. Each reader reviewed the de-identified cine clips for each examination and assigned one examination-level diagnosis from the four categories. Reader performance was compared with EchoHMIL on the same subset using accuracy, sensitivity, specificity, macro-F1, and per-class precision, recall, and F1-score. Paired comparisons were performed using bootstrap-based testing or McNemar’s test where appropriate. This experiment was designed as an offline benchmark rather than a full simulation of routine echocardiographic interpretation; readers did not perform real-time scanning, did not interactively acquire additional views, and did not use complete clinical context beyond the de-identified cine clips.
Model interpretability
To improve interpretability, we generated gradient-based class activation maps (Grad-CAM) on representative frames and clips for both the screening and final diagnostic predictions. Specifically, Grad-CAM was computed for the screening head and for the four-class output (Normal, VSD, ASD, PDA), and the resulting heatmaps were overlaid on echocardiography frames to highlight regions that most strongly contributed to model decisions. For normal examinations, we expected attribution to be comparatively diffuse without a persistent focal hotspot, consistent with the absence of a discrete defect signature.
In addition, the attention weights from the MIL pooling module were visualized to indicate which cine clips within an examination were most influential for the final prediction, providing examination-level attribution complementary to pixel-level saliency. Together, these pixel-and clip-level explanations facilitate qualitative assessment of whether the model attends to anatomically plausible regions and clinically informative clips when distinguishing normal studies from defects and when differentiating defect subtypes.
Statistical analysis
Continuous variables were summarized as mean±standard deviation or median (interquartile range) as appropriate, and categorical variables were summarized as counts and percentages. When comparing screening AUCs between methods, we used the De Long test; for paired comparisons of other metrics, we applied bootstrap-based testing where appropriate. All tests were two-sided, and a p-value < 0.05 was considered statistically significant. Statistical analyses were performed using Python (version 3.10) with standard scientific computing libraries.
Results
Study population characteristics
The baseline clinical characteristics of the study population are shown in Table 1. Significant intergroup differences were observed for age, sex distribution, several cardiovascular comorbidities, ventricular systolic function, and chamber enlargement patterns. In general, patients with ASD were older than those with VSD and PDA, while VSD and PDA groups showed higher LVEF values and more frequent left-sided chamber enlargement. By contrast, right heart enlargement and pulmonary hypertension were more common in the ASD group, consistent with the underlying volume-loading pattern of atrial-level shunting. For variables with 0 observed events in a specific group, the table should be interpreted as reflecting the distribution within this retrospective cohort rather than implying that such findings are clinically impossible in that disease category. Table 2 summarizes the cohort composition and the distribution of echocardiographic cine-clips across views. The dataset included 2373 patients and 29,335 cine-clips, comprising 751 ASD, 493 VSD, 329 PDA, and 800 normal cases. At the clip level, ASD contributed the largest number of samples, whereas normal cases, despite being the largest group by patient count, had fewer clips per subject. Across all groups, PSAX-GA was the most frequently represented view, while suprasternal clips were the least common. View distribution also varied substantially by diagnosis, suggesting disease-specific acquisition patterns that may influence model training and evaluation. Because view distribution and clip availability differed across diagnostic categories, the results should be interpreted together with the acquisition-bias and robustness analyses rather than as evidence that the model relies exclusively on disease-specific anatomy.
Table 2.
Cohort composition and echocardiographic cine-clip distribution by view
| Group | Patients | PSAX-GA | Subcostal biatrial | A4C | Suprasternal | PLAX | Total clips |
|---|---|---|---|---|---|---|---|
| ASD | 751 | 2983 | 3781 | 3113 | 987 | 1663 | 12527 |
| VSD | 493 | 2424 | 666 | 1736 | 787 | 1182 | 6795 |
| PDA | 329 | 1897 | 407 | 634 | 649 | 644 | 4231 |
| Normal | 800 | 1155 | 812 | 907 | 800 | 2108 | 5782 |
| Total | 2373 | 8459 | 5666 | 6390 | 3223 | 5597 | 29335 |
Normal versus target-shunt discrimination
We first evaluated whether EchoHMIL can function as a clinically relevant targeted triage model to distinguish normal examinations from studies containing a target left-to-right shunt lesion. This analysis was performed on the independent patient-level test set comprising both normal controls and target-shunt cases, with target-shunt cases treated as the positive class for discrimination analyses. We report receiver operating characteristic (ROC) analysis to summarize performance across all operating points, and additionally report the precision-recall (PR) curve because the non-negligible prevalence of target-shunt cases in the test cohort influences positive predictive value and thus the perceived utility of a targeted triage tool in practice.
To avoid optimistic bias and to emulate a targeted triage setting in which missed target-shunt cases are undesirable, the decision threshold τ was pre-specified on the validation set using a sensitivity-oriented criterion (targeting high recall for target-shunt cases). Test-set sensitivity and specificity were then computed at this fixed threshold, thereby explicitly accounting for performance on both target-shunt cases (sensitivity) and normal controls (specificity).
As shown in Fig. 2, EchoHMIL demonstrated strong normal-versus-target-shunt discrimination, achieving an AUC of 0.957. The corresponding PR curve (Fig. 3) maintained favorable precision across a wide range of recall values, indicating that high sensitivity can be achieved without an excessive loss of precision. Figure 4 further summarizes the sensitivity-specificity trade-off as a function of the decision threshold and highlights the operating point τ=0.724 selected on the validation set. At this operating point, EchoHMIL achieved a sensitivity of 92.2% for target-shunt detection and a specificity of 82.4% on normal examinations (Table 3). Compared with a baseline single-head model trained as a flat four-class classifier, EchoHMIL achieved higher overall discrimination and a more favorable screening trade-off at matched sensitivity, consistent with the benefit of examination-level aggregation under heterogeneous clip quality and incomplete view coverage.
Fig. 2.

Receiver operating characteristic (ROC) curves for normal-versus-target-shunt discrimination on the independent test set. Curves are shown for EchoHMIL and a baseline flat 4-class single-head model
Fig. 3.

Precision-recall (PR) curves for normal-versus-target-shunt discrimination. PR analysis complements ROC by illustrating the precision-recall trade-off under the cohort class prevalence
Fig. 4.

Sensitivity and specificity as a function of the target-shunt triage threshold for EchoHMIL. The vertical dashed line indicates the prespecified operating point τ selected on the validation set to prioritize sensitivity
Table 3.
Performance of EchoHMIL for normal-versus-target-shunt discrimination at the operating threshold τ
| Method | AUC | Sensitivity | Specificity | F1-score |
|---|---|---|---|---|
| EchoHMIL | 0.957 | 92.2% | 82.4% | 0.918 |
| Baseline (4-class single head) | 0.894 | 92.2% | 66.1% | 0.883 |
Calibration of screening and subtype probabilities
In addition to discrimination metrics, we assessed model calibration using calibration curves, Brier score, and expected calibration error (ECE). As shown in Fig. 5, for the normal-versus-target-shunt task, EchoHMIL achieved a Brier score of 0.084 and an ECE of 0.041, indicating generally acceptable agreement between predicted probabilities and observed outcomes. For subtype probabilities, the Brier scores were 0.076, 0.064, and 0.083 for ASD, VSD, and PDA, respectively, with corresponding ECE values of 0.052, 0.045, and 0.061. Mild deviations from ideal calibration were observed in the high-probability bins, particularly for subtype classification, likely reflecting the smaller effective sample size within subtype-specific bins and residual inter-class overlap. These findings suggest that probability outputs may support triage and confidence interpretation, but additional calibration refinement and prospective validation are needed before clinical deployment.
Fig. 5.

Overview of calibration plots for EchoHMIL, including the normal-versus-target-shunt task and one-vs.-rest subtype probability plots for ASD, VSD, and PDA. The diagonal dashed line indicates ideal calibration
Subtype classification performance
We next evaluated end-to-end diagnostic performance under the clinically consistent hierarchical inference rule, where the final prediction is one of four categories: Normal, VSD, ASD, or PDA. In addition to reporting lesion-subtype metrics on reference-standard target-shunt examinations (where subtype discrimination is clinically interpreted after abnormality is established), we also report the four-class confusion matrix to explicitly characterize performance on normal controls and the resulting false-positive burden.
Among target-shunt cases, EchoHMIL achieved an overall subtype accuracy of 88.8% and a macro-F1 of 0.885 for VSD/ASD/PDA classification. As shown in Fig. 6, residual subtype errors occurred predominantly between ASD and PDA, whereas VSD was more distinctly recognized with fewer cross-class confusions. Importantly, normal controls were explicitly included in the end-to-end 4-way confusion matrix, enabling direct assessment of model behavior on normal examinations. The majority of normal studies were correctly routed to normal (Normal recall 82.2%; precision 80.4%; F1-score 0.812), while a minority were over-called as shunt lesions, consistent with the sensitivity-oriented screening operating point. Conversely, a small fraction of target-shunt examinations were routed to Normal (Fig. 6), representing missed abnormalities under the hierarchical rule.
Fig. 6.

End-to-end confusion matrix on the test set under the hierarchical decision rule
Human-reader comparison
To assess the clinical relevance of EchoHMIL, we compared its performance with that of echocardiographers of varying experience levels on a randomly selected subset of the independent test set. To ensure a paired comparison, EchoHMIL predictions were re-evaluated on the same reader-assessment subset. On this subset, EchoHMIL achieved a normal-versus-target-shunt sensitivity of 91.8% and specificity of 83.1%, compared with 84.6%/78.1% for the junior reader, 89.3%/83.6% for the intermediate reader, and 93.0%/87.2% for the senior reader. For end-to-end four-class diagnosis, EchoHMIL achieved an accuracy of 82.7% and macro-F1 of 0.824, compared with 73.9% and 0.716 for the junior reader, 80.4% and 0.792 for the intermediate reader, and 84.6% and 0.836 for the senior reader. These findings suggest that the model may provide useful decision-support information in an offline cine-review setting, but they should not be interpreted as definitive equivalence or superiority to experienced echocardiographers because the reader study included only three readers and did not reproduce real-time scanning, Doppler interrogation, or full clinical context.
Ablation and robustness analyses
We conducted ablation studies to quantify the contribution of (i) the spatiotemporal video backbone, (ii) attention-based multi-instance aggregation, and (iii) the clinically aligned hierarchical screening-subtyping design. Ablations were evaluated on the independent test set using normal-versus-target-shunt AUC for normal-versus-target-shunt discrimination and end-to-end 4-class metrics (Normal/VSD/ASD/PDA),which directly reflect the final diagnostic output (Table 4; Fig. 7).
Table 4.
End-to-end diagnostic performance of EchoHMIL for 4-class classification on the test set (normal, VSD, ASD, and PDA)
| Class | Precision | Recall (Sensitivity) | F1-score |
|---|---|---|---|
| VSD | 0.859 | 0.850 | 0.854 |
| ASD | 0.824 | 0.819 | 0.821 |
| PDA | 0.797 | 0.788 | 0.792 |
| Normal | 0.804 | 0.822 | 0.813 |
Macro average (4-class) - − 0.820
Overall accuracy (4-class) 82.1%
Note: On reference-standard target-shunt cases only (VSD/ASD/PDA), macro-F1 = 0.885 and accuracy = 88.8%
Fig. 7.

Ablation effects on end-to-end four-class macro-F1 relative to the full EchoHMIL model
Effect of multi-instance aggregation. Replacing attention-based MIL pooling with mean pooling reduced end-to-end 4-class macro-F1, and removing MIL aggregation entirely (single-clip inference) produced the largest degradation (Fig. 7). These findings indicate that examination-level evidence integration is critical in routine echocardiography, where view coverage and clip quality are heterogeneous, and diagnostically informative frames may be concentrated in a subset of clips.
Effect of hierarchical modeling. Substituting the hierarchical dual-head formulation with a flat 4-class single-head classifier led to inferior screening discrimination and reduced end-to-end macro-F1 (Table 5). This supports that explicitly modeling the clinical decision pathway (screening followed by subtype differentiation) improves robustness under real-world acquisition variability.
Table 5.
Ablation study on the independent test set. Normal-versus-target-shunt AUC is reported for normal-versus-target-shunt discrimination, and end-to-end metrics are reported for four-class diagnosis
| Variant | Normal-versus-target-shunt AUC | 4-class Accuracy (%) | 4-class Macro-F1 |
|---|---|---|---|
| EchoHMIL (Video Swin + Attn-MIL + Hier) | 0.957 | 82.1 | 0.820 |
| Backbone: R(2+1) D (keep Attn-MIL + Hier) | 0.949 | 81.3 | 0.812 |
| Backbone: X3D (keep Attn-MIL + Hier) | 0.944 | 80.6 | 0.804 |
| Mean pooling (replace Attn-MIL; keep Hier) | 0.950 | 80.0 | 0.798 |
| No MIL (single-clip; keep Hier) | 0.941 | 78.9 | 0.785 |
| Flat 4-class single head (no Hier) | 0.894 | 79.7 | 0.793 |
Backbone comparison. Among representative backbones, the transformer-based video encoder achieved the best overall trade-off, while R(2+1) D and X3D yielded modestly lower end-to-end performance, consistent with the benefit of stronger spatiotemporal representation learning for echocardiography cine interpretation.
Robustness across acquisition-related strata. Because patient-level clinical variables were not available in this de-identified video dataset, robustness analyses were performed across acquisition-derived subgroups, including vendor/device group when available, view coverage, cine-clip count, and image quality (Table 6; Fig. 8). Performance remained broadly favorable across major acquisition strata, but modest attenuation was observed in examinations with reduced view coverage, fewer clips, or lower image quality, consistent with the dependence of shunt recognition on adequate visualization of septal and ductal landmarks. In view-coverage-matched and clip-count-matched sensitivity analyses, the model retained acceptable discrimination, suggesting that the reported performance was not solely explained by gross differences in acquisition protocol. Nevertheless, these internal analyses cannot fully exclude residual shortcut learning; multicenter protocol-controlled validation remains necessary.
Table 6.
Robustness analysis across acquisition-related subgroups. Normal-versus-target-shunt AUC and end-to-end four-class metrics are reported with 95% CIs
| Subgroup | Normal-versus-target-shunt AUC(95% CI) | 4-class Accuracy (%) | 4-class Macro-F1 |
|---|---|---|---|
| Vendor A | 0.959 (0.949–0.968) | 82.6 | 0.823 |
| Vendor B | 0.954 (0.943–0.964) | 81.5 | 0.816 |
| Vendor C | 0.961 (0.951–0.970) | 82.3 | 0.821 |
| View coverage ≥ 4 views | 0.961 (0.952–0.969) | 82.8 | 0.825 |
| View coverage < 4 views | 0.948 (0.936–0.959) | 80.4 | 0.806 |
| Good image quality | 0.962 (0.953–0.970) | 83.0 | 0.827 |
| Moderate/Poor quality | 0.945 (0.932–0.957) | 80.7 | 0.808 |
| High clip count (e.g., ≥median) | 0.960 (0.951–0.968) | 82.7 | 0.824 |
| Low clip count (e.g., <median) | 0.949 (0.939–0.958) | 80.9 | 0.810 |
Fig. 8.

Robustness across acquisition-related subgroups. Left: normal-versus-target-shunt AUC. Right: end-to-end four-class accuracy. Error bars indicate 95% CIs. The analysis includes view-coverage, clip-count, and image-quality strata to assess sensitivity to acquisition-related bias
Interpretability and error analysis
To interrogate model decision-making, we visualized pixel-level attribution using Grad-CAM-style activation maps overlaid on representative frames (Fig. 9). Overall, the highlighted regions were spatially coherent within the echocardiographic sector and concentrated on anatomically plausible structures, suggesting that EchoHMIL predominantly leveraged clinically meaningful image evidence rather than background artifacts.
Fig. 9.

Qualitative interpretability via Grad-CAM-style activation maps. Activation heatmaps (jet colormap) are overlaid on representative echocardiography frames; warmer colors indicate higher model attention. (A) VSD: peak activations concentrate along the interventricular septum. (B) PDA: attention shifts toward the ductal/outflow vicinity. (C) Normal: activation is comparatively diffuse without a persistent focal hotspot. (D) ASD: peak activations localize around the interatrial septum and shunt-related motion cues
Across subtypes, EchoHMIL exhibited a consistent pattern of broad contextual attention with focal peaks (Fig. 9A–D): low-to-moderate activation covered the global cardiac silhouette to provide contextual grounding, whereas higher activation localized to lesion-relevant regions. Specifically, for VSD (Fig. 9A), the strongest responses clustered along the interventricular septum and adjacent endocardial boundaries, consistent with clinical interrogation of septal continuity and shunt-related motion cues. For PDA (Fig. 9B), attention shifted toward the ductal vicinity, emphasizing peri-great-vessel structures that commonly carry discriminative information for ductal patency. For the normal example (Fig. 9C), activations were weaker and more diffuse, largely tracking cardiac contours without a persistent focal hotspot, consistent with the absence of a discrete defect signature. For ASD (Fig. 9D), peak activations localized around the interatrial septum and nearby atrial landmarks, aligning with the expected anatomical locus for atrial-level shunting. Collectively, these qualitative results support that EchoHMIL learns subtype-specific attention patterns that are concordant with clinical reasoning.
We further performed a qualitative review of misclassified examinations. Errors were most frequently associated with: (i) suboptimal acoustic windows and motion artifacts that degrade septal delineation; (ii) small defects or weak/intermittent shunt signatures, where discriminative cues are subtle and temporally sparse; and (iii) incomplete or non-target view coverage, in which the most informative clips were absent or truncated. In such cases, attribution maps tended to become more spatially diffuse or to peak on secondary structures with shared appearance across shunt lesions, offering a plausible explanation for residual subtype confusions. Overall, EchoHMIL generally attends anatomically appropriate regions for both screening and subtyping, while failure modes are dominated by data-quality and view-availability constraints intrinsic to real-world echocardiography.
Discussion
In this study, we developed and validated EchoHMIL, a hierarchical video-based multi-instance learning (MIL) framework that supports end-to-end four-class diagnosis from TTE cine examinations, assigning each study to one of Normal, VSD, ASD, or PDA. Under a clinically consistent hierarchical inference rule, EchoHMIL achieved an end-to-end 4-class accuracy of 82.1% with a 4-class macro-F1 of 0.820 on the independent test set, while preserving strong normal-versus-target-shunt discrimination (AUC 0.957) at a sensitivity of 92.2% and specificity of 82.4% under a sensitivity-prioritized threshold. Among reference-standard target-shunt cases, subtype performance remained robust (accuracy 88.8%, macro-F1 0.885), indicating that the model retains lesion-specific discriminability after abnormality is established. Together, these results support the feasibility of clinically aligned, examination-level inference from heterogeneous TTE videos—consistent with contemporary echocardiography standards that emphasize comprehensive multi-view acquisition and interpretation rather than reliance on a single clip or view [3]. Importantly, these results should be interpreted as internal validation of a targeted closed-set triage/classification framework, not as evidence for comprehensive pediatric CHD screening.
Positioning relative to prior work
Deep learning for echocardiography has advanced rapidly, demonstrating that clinically meaningful signals can be extracted from both images and videos to support detection, triage, and workflow automation [18, 19]. However, many CHD-focused pipelines remain constrained by curated single-view inputs or narrowly defined tasks, which can limit robustness when deployed on routine examinations with variable view coverage, clip quality, and acquisition artifacts [20]. Recent efforts targeting specific shunt lesions (e.g., ASD) have reported encouraging results when models are trained on selected views or curated clips [21, 22], and parallel work in ultrasound has increasingly adopted attention/MIL-style aggregation to emulate expert selection of informative instances within multi-view studies [23]. Prior studies have shown the value of video-based deep learning, attention mechanisms, and MIL aggregation in echocardiography and ultrasound. Our study builds on these advances rather than claiming novelty for each individual component. The main distinction is the clinical formulation: EchoHMIL is designed for examination-level triage and subtype classification of common congenital left-to-right shunt lesions using multi-view cine examinations. Unlike approaches based on selected views or isolated clips, our framework aggregates heterogeneous clips from routine studies and applies a hierarchical decision rule that reflects clinical interpretation. Thus, the contribution is primarily the workflow-aligned application and validation of hierarchical MIL for targeted CHD shunt-lesion classification. We have therefore revised the manuscript to avoid overstating algorithmic novelty: the individual technical components are established, whereas the main contribution is their clinically motivated integration and evaluation at the examination level for a predefined shunt-lesion task.
Why the hierarchical MIL formulation matters
We attribute the observed performance to three complementary design choices. First, spatiotemporal video encoding enables representation learning from dynamic cues that are central to echocardiographic interpretation (e.g., septal continuity across frames, motion-related patterns, and view-dependent temporal signatures) and may be underutilized in static-image pipelines. Second, attention-based MIL pooling provides resilience to heterogeneous clip quality by allowing the model to up-weight diagnostically informative clips and down-weight non-diagnostic segments, mirroring how clinicians integrate evidence across a multi-clip examination. Third, the hierarchical dual-head design reduces negative transfer by gating lesion-specific learning to target-shunt cases during training, while the inference rule enforces clinical consistency (screen first, then subtype). This separation is particularly relevant when subtype cues are weak, temporally intermittent, or present in only a subset of views—conditions that are common in routine TTE.
Clinical implications and deployment considerations
From a translational perspective, EchoHMIL is best positioned as a targeted triage and decision-support tool for common left-to-right shunt lesions, rather than a standalone replacement for expert interpretation or a general CHD screening system. In practice, the model could help prioritize examinations suspected of ASD, VSD, or PDA for expert review, particularly in high-volume or resource-limited echocardiography settings. The end-to-end four-class output supports a workflow in which most normal studies are routed to Normal, while suspected abnormal cases receive a structured subtype suggestion. Because the operating threshold was tuned to prioritize sensitivity, false positives and overcalling of normal studies are expected; therefore, positive outputs should prompt careful clinician review rather than be treated as definitive diagnoses. For deployment, robust quality control, calibration, uncertainty estimation, and explicit handling of out-of-distribution inputs will be essential to avoid inappropriate overconfidence; these issues are increasingly recognized in recent work on echocardiographic video quality control and view recognition [24]. Future clinical evaluation should assess reader-model interaction, reporting time, inter-reader variability, and workflow impact in real-world echo laboratory settings. The present model should be interpreted as a targeted triage and classification framework for common left-to-right shunt lesions rather than a general CHD screening tool. Although the binary head distinguishes normal examinations from ASD, VSD, and PDA, complex CHD, rare lesions, mixed defects, and other out-of-distribution abnormalities were excluded from the current dataset. Therefore, the model’s ability to safely identify the broader spectrum of CHD remains unproven and requires prospective validation in unfiltered populations. In any future deployment, the model output should be interpreted as a triage signal that triggers expert review rather than as an autonomous diagnostic decision. A negative prediction would also require quality-control safeguards because complex CHD and non-target abnormalities were not represented in the current training or test cohorts.
Limitations and future work
This study has several limitations. First, it was retrospective and single-center; therefore, the reported performance should be interpreted as internal validation only. Because echocardiography is vendor-, operator-, and protocol-dependent, prospective multicenter external validation is required before clinical deployment. Second, EchoHMIL was developed for normal controls and three common left-to-right shunt lesions-ASD, VSD, and PDA. Complex CHD, mixed lesions, postoperative anatomy, rare congenital lesions, acquired abnormalities, and broader out-of-distribution cases were excluded; accordingly, the model should not be interpreted as a general pediatric CHD screening system. Third, acquisition-related bias and shortcut learning remain possible because view distribution, clip count, and acquisition patterns differed across diagnostic categories. Although we added view-balanced sampling, acquisition-stratified robustness analyses, and matched-subset sensitivity analyses, these internal analyses cannot fully rule out protocol-related confounding. Fourth, the human-reader comparison included only three echocardiographers and used offline cine review without real-time scanning, interactive Doppler interrogation, or full clinical context; therefore, it should be viewed as an exploratory benchmark rather than proof of equivalence to clinical echocardiographers. Fifth, the de-identified video dataset contained limited clinical and multimodal information, preventing analysis by symptoms, hemodynamic severity, comorbidities, Doppler findings, 3D TTE, or strain parameters. Finally, this study did not assess downstream clinical utility, including reporting time, inter-reader variability, referral decisions, or model-assisted workflow impact. Future work should focus on prospective multicenter validation, broader CHD coverage, protocol-controlled testing, explicit OOD detection, multimodal integration, and real-world reader-assistance studies.
Conclusions
EchoHMIL provides an examination-level hierarchical MIL framework for targeted triage and classification of common left-to-right shunt lesions, including ASD, VSD, and PDA. Prospective validation on broader CHD populations, including complex and out-of-distribution cases, is required before extension to general CHD screening. The current evidence supports only targeted closed-set use; broader screening claims require validation in unfiltered pediatric echocardiography populations.
Acknowledgements
Not applicable.
Abbreviations
- EchoHMIL
Echocardiography Hierarchical Multiple-Instance Learning network
- CHD
Congenital heart disease
- TTE
Transthoracic echocardiography
- VSD
Ventricular septal defect
- ASD
Atrial septal defect
- PDA
Patent ductus arteriosus
- PSAX-GA
Parasternal short-axis at the great artery level
- A4C
Apical four-chamber view
- PLAX
Parasternal long-axis view
- ROC
Receiver operating characteristic
- PR
Precision-recall
- AUC
Area under the receiver operating characteristic curve
- CIs
Confidence intervals
- FAC
Fractional area change
- LVEF
Left ventricular ejection fraction
- Grad-CAM
Gradient-weighted class activation mapping
- OOD
Out-of-distribution
Author contributions
Yuming Mu was responsible for the overall supervision of the project and critical revision of the manuscript. Peipei Zhang designed the study, analyzed the data, and wrote the first draft of the manuscript. Yang Wu contributed equally to this work and was involved in methodology development, data analysis, and manuscript revision. Xiaomei Hu assisted with data curation, literature review, and editing of the manuscript. All authors have read and approved the final version of the manuscript.
Funding
No funding.
Data availability
Data and code are available from the corresponding author upon reasonable request, subject to institutional regulations.
Declarations
Ethics approval and consent to participate
This study was conducted in accordance with the Declaration of Helsinki. This study was approved by the ethics committees of The First Affiliated Hospital of Xinjiang Medical University, with a waiver of informed consent.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Peipei Zhang and Yang Wu contributed equally to this work.
References
- 1.van der Linde D, Konings EEM, Slager MA, Witsenburg M, Helbing WA, Takkenberg JJM, et al. Birth prevalence of congenital heart disease worldwide: a systematic review and meta-analysis. J Am Coll Cardiol. 2011;58(21):2241–47. 10.1016/j.jacc.2011.08.025. [DOI] [PubMed] [Google Scholar]
- 2.Liu Y, Chen S, Zühlke L, Black GC, Choy MK, Li N, et al. Global birth prevalence of congenital heart defects 1970–2017: updated systematic review and meta-analysis of 260 studies. Int J Epidemiol. 2019;48(2):455–63. 10.1093/ije/dyz009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Lopez L, Saurers DL, Barker PCA, Cohen MS, Colan SD, Dwyer J, et al. Guidelines for performing a comprehensive pediatric transthoracic echocardiogram: recommendations from the American society of echocardiography. J Am Soc Echocardiogr. 2024;37(2):119–70. 10.1016/j.echo.2023.11.015. [DOI] [PubMed] [Google Scholar]
- 4.Myhre PL, Grenne B, Asch FM, Delgado V, Khera R, Lafitte S, et al. Artificial intelligence-enhanced echocardiography in cardiovascular disease management. Nat Rev Cardiol. 2026;23(3):164–82. 10.1038/s41569-025-01197-0. [DOI] [PubMed] [Google Scholar]
- 5.Hokanson JS, Ring K, Zhang X. A survey of pediatric cardiologists regarding non-emergent echocardiographic findings in asymptomatic newborns. Pediatr Cardiol. 2022;43(4):837–43. 10.1007/s00246-021-02795-8. [DOI] [PubMed] [Google Scholar]
- 6.Madani A, Ong JR, Tibrewal A, Mofrad MRK. Deep echocardiography: data-efficient supervised and semi-supervised deep learning towards automated diagnosis of cardiac disease. npj Digit Med. 2018;1(1):59. 10.1038/s41746-018-0065-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Ghorbani A, Ouyang D, Abid A, He B, Chen JH, Harrington RA, et al. Deep learning interpretation of echocardiograms. NPJ Digit Med. 2020;3(1):10. 10.1038/s41746-019-0216-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Ouyang D, He B, Ghorbani A, Yuan N, Ebinger J, Langlotz CP, et al. Video-based AI for beat-to-beat assessment of cardiac function. Nature. 2020;580(7802):252–56. 10.1038/s41586-020-2145-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Hughes JW, Yuan N, He B, Ouyang J, Ebinger J, Botting P, et al. Deep learning evaluation of biomarkers from echocardiogram videos. EBioMedicine. 2021;73:103613. 10.1016/j.ebiom.2021.103613. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Narang A, Bae R, Hong H, Thomas Y, Surette S, Cadieu C, et al. Utility of a deep-learning algorithm to guide novices to acquire echocardiograms for limited diagnostic use. JAMA Cardiol. 2021;6(6):624–32. 10.1001/jamacardio.2021.0185. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Litjens G, Kooi T, Bejnordi BE, Setio AAA, Ciompi F, Ghafoorian M, et al. A survey on deep learning in medical image analysis. Med Image Anal. 2017;42:60–88. 10.1016/j.media.2017.07.005. [DOI] [PubMed] [Google Scholar]
- 12.Chen F, Zhang HY, Wan YL, Jia JN, Wang RZ, Gao C, et al. Artificial intelligence-assisted organoid construction in congenital heart disease: current applications and future prospects. Front Bioeng Biotechnol. 2025;13:1691972. 10.3389/fbioe.2025.1691972. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Tran D, Wang H, Torresani L, Ray J, LeCun Y, Paluri M. A closer look at spatiotemporal convolutions for action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2018. 6450–59.
- 14.Feichtenhofer C. X3D: expanding architectures for efficient video recognition. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020. 203–13.
- 15.Liu Z, Ning J, Cao Y, Wei Y, Zhang Z, Lin S, et al. Video Swin transformer. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022. 3202–11.
- 16.Tong Z, Song Y, Wang J, Wang L. VideoMAE: masked autoencoders are data-efficient learners for self-supervised video pre-training. Adv Neural Inf Process Syst. 2022;35:10078–93. [Google Scholar]
- 17.Ilse M, Tomczak JM, Welling M. Attention-based deep multiple instance learning. Proceedings of the 35th International Conference on Machine Learning. Proc Mach Learn Res. 2018: 2127–36, 80.
- 18.Sahashi Y, Ouyang D, Okura H, Kagiyama N. AI-echocardiography: current status and future direction. J Cardiol. 2025;85(6):458–64. 10.1016/j.jjcc.2025.02.005. [DOI] [PubMed] [Google Scholar]
- 19.Holste G, Oikonomou EK, Mortazavi BJ, Wang Z, Khera R. Efficient deep learning-based automated diagnosis from echocardiography with contrastive self-supervised learning. Commun Med (Lond). 2024;4(1):133. 10.1038/s43856-024-00538-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Qayyum SN. A comprehensive review of applications of artificial intelligence in echocardiography. Curr Probl Cardiol. 2024;49(2):102250. 10.1016/j.cpcardiol.2023.102250. [DOI] [PubMed] [Google Scholar]
- 21.Liu Y, Hou S, Han X, Liang T, Hu M, Wang X, et al. Intelligent diagnosis of atrial septal defect in children using echocardiography with deep learning. Virtual Reality Intell Hardware. 2024;6(3):217–25. 10.1016/j.vrih.2023.05.002. [DOI] [Google Scholar]
- 22.Liu Y, Huang Q, Han X, Liang T, Zhang Z, Lu X, et al. Atrial septal defect detection in children based on ultrasound video using multiple instances learning. J Digit Imag. Inf. Med. 2024;37(3):965–75. 10.1007/s10278-024-00987-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Huang Z, Wessler BS, Hughes MC. Detecting heart disease from multi-view ultrasound images via supervised attention multiple instance learning. Proceedings of the 8th Machine Learning for Healthcare Conference. Proc Mach Learn Res. 2023: 285–307, 219. [PMC free article] [PubMed]
- 24.Song S, Qin Y, Yang H, Huang T, Fei H, Li X. EchoViewCLIP: advancing video quality control through high-performance view recognition of echocardiography. In: Jc G, et al. editors. Medical image computing and computer assisted intervention - MICCAI 2025. Lect notes comput sci. Vol. 15972. 2026. p. 181–91.
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
Data and code are available from the corresponding author upon reasonable request, subject to institutional regulations.







