Skip to main content
International Journal of Emergency Medicine logoLink to International Journal of Emergency Medicine
. 2025 Nov 25;18:247. doi: 10.1186/s12245-025-01065-1

Machine learning-based classification of carotid plaques via ultrasound: a systematic review and meta-analysis of diagnostic performance

Pooya Eini 1,✉, Peyman Eini 2, Homa Serpoush 3, Mohammad Rezayee 4, Jason Tremblay 4
PMCID: PMC12645702  PMID: 41291444

Abstract

Background

Machine learning (ML) models have gained traction for classifying carotid artery plaques via ultrasound imaging to differentiate high-risk (unstable) from low-risk (stable) plaques, a critical step for stroke risk prediction and guiding clinical interventions such as endarterectomy. However, prior studies report inconsistent diagnostic performance attributed to variations in algorithms, cohort diversity, and imaging protocols. This systematic review and meta-analysis aim to evaluate the pooled diagnostic accuracy of ML models for carotid plaque classification, addressing these inconsistencies to inform standardized clinical applications.

Methods

Five electronic databases were systematically searched up to February 28, 2025, for studies reporting diagnostic performance metrics of ML-based models in carotid plaque classification. Pooled performance metrics were analyzed using STATA, and the risk of bias was assessed using the PROBAST+AI tool.

Results

A total of 20 studies met the inclusion criteria, of which 13 provided sufficient data for quantitative synthesis. Sample sizes ranged from 15 to 413 patients, with 115– 81,000 images per study. Mean ages ranged from 27.5 to 75 years, mostly 60–70, and male representation ranged from 47% to 81%, except for one all-female cohort. The pooled sensitivity was 0.84 (95% CI: 0.74–0.90) and specificity was 0.96 (95% CI: 0.89–0.98), with a pooled AUC of 0.95 (95% CI: 0.93–0.97). Substantial heterogeneity was observed (I2 = 88.8% for sensitivity, 64.1% for specificity, and 68.1% overall). Meta-regression identified sample size and model architecture as significant sources of between-study heterogeneity. No evidence of publication bias was detected (p = 0.36). Quality assessment using PROBAST+AI indicated a low overall risk of bias in 70% of studies, moderate in 20%, and high in 10%. The GRADE approach rated the certainty of evidence as moderate, primarily due to inconsistency and study-level bias.

Conclusion

Machine learning models demonstrate promising diagnostic accuracy for carotid plaque classification, showing high pooled sensitivity and specificity. However, substantial heterogeneity and only moderate certainty of evidence suggest that these findings should be interpreted with caution.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12245-025-01065-1.

Keywords: Machine learning, Artificial intelligence, Ultrasound, Carotid artery plaque, Carotid atherosclerosis

Introduction

Stroke remains a major global health challenge, ranking as the second leading cause of death and the third leading cause of death and disability combined worldwide [1]. In 2021, it accounted for more than seven million deaths and over 150 million disability-adjusted life-years (DALYs) globally [2]. In the United States, stroke caused approximately 160,000 deaths in 2022 (about 17% of cardiovascular mortality), with incidence and mortality showing a modest upward trend over the past decade [3]. Global projections indicate that annual stroke deaths may increase by nearly 50% by 2050, driven by population aging, growth, and escalating risk factors such as obesity and hypertension [4]. These trends underscore the need for improved risk stratification and prevention strategies.

Carotid atherosclerotic plaques are a well-recognized contributor to ischemic stroke, particularly in the anterior circulation [5]. While plaques causing significant luminal stenosis are clearly implicated in stroke risk, emerging evidence indicates that non-stenotic plaques may also play a causal role [6]. The embolic potential of these plaques is thought to depend on high-risk features such as intraplaque hemorrhage, ulceration, and hypodensity [7]. Accurate identification and characterization of carotid plaque using advanced imaging modalities, including ultrasound, MRI, CTA, and PET, are therefore critical for stroke risk stratification and the development of targeted preventive strategies.

Artificial intelligence (AI), encompassing machine learning (ML) and deep learning (DL), has revolutionized medical imaging by automating feature extraction and improving diagnostic accuracy [8]. In carotid ultrasound imaging, machine learning models are used to interpret both longitudinal and transverse scans, enabling automated detection of occlusions, measurement of arterial diameter, and evaluation of plaque accumulation [9]. This approach reduces reliance on manual segmentation, minimizing human error and time requirements. Recent studies report CNN-based models achieving areas under the curve (AUC) exceeding 0.85 for detecting vulnerable plaques, surpassing traditional manual and semi-automated approaches [10, 11]. However, the lack of standardized imaging protocols, feature extraction methods, and regional differences across studies limits the generalizability and clinical translation of current findings.

This study aims to systematically evaluate the diagnostic performance of ML models for carotid plaque classification using ultrasound imaging, focusing on metrics such as sensitivity, specificity, and AUC to distinguish high-risk (unstable, symptomatic) from low-risk plaques. Through a comprehensive systematic review and meta-analysis, we seek to synthesize evidence on MLs accuracy, identify optimal algorithms, and assess generalizability across diverse cohorts, addressing standardization gaps to enhance clinical decision-making for stroke prevention.

Methods

Study design

This systematic review and meta-analysis was conducted in accordance with the PRISMA 2020 (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines [12]. The study protocol was prospectively registered on the Open Science Framework (OSF; registration ID: http://osf.io/duqsp). The PRISMA checklist is provided in the Supplementary Material.

Eligibility criteria

Inclusion criteria

  • Population and outcome: Studies that performed carotid plaque classification, defined as the assignment of atherosclerotic lesions into clinically meaningful categories (stable vs. unstable/vulnerable; calcified, fibrous, lipid-rich; with or without ulceration or intraplaque hemorrhage) based on imaging or histopathologic criteria.

  • Study design and language: Peer-reviewed publications in English, with no restriction on publication year or study design.

  • Definition requirement: Included studies clearly defined their plaque classification criteria or provided imaging/histopathology references. Studies lacking a clear definition were excluded or considered separately in sensitivity analyses.

  • Model type: Studies that employed machine learning–based prediction models applied to carotid plaque characterization.

  • Performance reporting: Studies that reported at least one diagnostic performance metric—accuracy, sensitivity, specificity, precision, or area under the curve (AUC)—or provided sufficient quantitative data for synthesis.

  • Imaging modality: All carotid ultrasound modalities were eligible, including B-mode, contrast-enhanced ultrasound (CEUS), shear-wave elastography (SWE), superb microvascular imaging (SMI), and super-resolution ultrasound. Details of the imaging modality and parameters were extracted, and subgroup analyses were performed to compare model performance across modalities.

Exclusion criteria

  • Language and subject: Non-English publications or non-human studies.

  • Model or focus: Studies that did not apply machine learning–based models or were not focused on carotid plaque classification.

  • Study type: Reviews, editorials, conference abstracts, letters, or case reports that did not contain original data.

  • Duplications: Studies using overlapping or duplicate datasets without clear differentiation were excluded to avoid duplication bias.

A summary of the inclusion criteria using the PICO framework is provided in Supplementary Table 1.

Search strategy

A comprehensive search of PubMed, Scopus, Web of Science, Embase, and ProQuest databases was conducted from inception to February 28, 2025, to identify eligible studies. To capture grey literature, the first 100 results from Google Scholar were also screened.

Search terms included combinations of keywords and MeSH terms related to “machine learning,” “artificial intelligence,” “carotid artery,” “plaque,” “atherosclerosis,” and “classification.” The whole search strategy is detailed in the Supplementary Table 2.

Study selection

Two independent reviewers performed the title and abstract screening using the Rayyan web platform, followed by full-text assessment against the predefined eligibility criteria. The reviewers were blinded to each other’s decisions during the initial screening phase to minimize selection bias. Discrepancies were resolved through consensus or by consultation with a third reviewer. Reference lists of included studies and relevant reviews were also hand-searched to identify additional eligible articles.

Data extraction

Two reviewers independently extracted data using a standardized pre-piloted form. Extracted information included:

  1. Study characteristics: author, publication year, country, sample size, gender distribution, and dataset type.

  2. Model characteristics: ML algorithm used, feature type, and validation approach.

  3. Performance metrics: accuracy, sensitivity, specificity, precision, and AUC, along with raw classification counts: True positives (TP), true negatives (TN), false positives (FP), and false negatives (FN).

When performance metrics were missing, they were calculated or imputed from available raw data whenever possible. When essential data were unavailable, no contact was made with the original authors. Discrepancies in extracted data were resolved through discussion and mutual agreement between the reviewers.

Quality assessment

The PROBAST+AI tool (Prediction model Risk Of Bias Assessment Tool) was used because it is designed explicitly for prediction model studies and includes AI-specific aspects such as algorithm development, validation, and handling of complex predictors [13]. This makes it more suitable than general tools like QUADAS-2 or NOS for assessing ML studies. The certainty of evidence for pooled outcomes was evaluated using the GRADE approach. This method considers factors such as risk of bias, inconsistency, imprecision, and publication bias, providing a structured and transparent evaluation of confidence in the results. Each study was rated for risk of bias as low, moderate, or high by two reviewers, with disagreements resolved by consensus or by a senior reviewer.

Statistical analysis

All statistical analyses were performed using Stata version 18.0, utilizing the MIDAS and METADATA modules. TP, TN, FP, and FN were extracted or calculated for each model to derive pooled estimates of sensitivity, specificity, accuracy, precision, and AUC. When multiple models were reported in a single study, the best-performing model, as defined by the original authors, was selected to avoid selection bias. None of the included studies were multi-arm or reported correlated outcomes. A random-effects model with restricted maximum likelihood (REML) estimation was used to pool diagnostic performance metrics. This approach accounts for both within- and between-study variability, making it more appropriate than a fixed-effects model when substantial heterogeneity is expected across studies, as it provides more conservative and generalizable estimates. Potential sources of heterogeneity were explored using meta-regression, leave-one-out sensitivity analyses, and subgroup analyses (based on model type, study sample size, region, imaging modality, and plaque type).

For subgroup analysis, the machine learning models were categorized into two main groups:

  • DL models included convolutional neural network (CNN)-based architectures such as MultiNet 2.0, VGG13, Inception_v3, ResNet18, SE-ResNext-50, CNN-BiLSTM, CANet, DL-DCCP, and other deep convolutional neural networks or ensemble schemes.

  • U-Net-based models comprised UNet, UNet+GDL+GC, HRU-Net with transfer learning, Attention-UNet, and other U-Net–derived deep learning variants, primarily used for segmentation and detailed plaque feature extraction.

Publication bias was evaluated using Deeks’ funnel plot asymmetry test, where a p-value < 0.05 indicates publication bias.

Results

Study selection

The study selection process followed the PRISMA 2020 framework (Fig. 1). A total of 816 records were identified through database searches: PubMed (n = 124), Scopus (n = 225), Web of Science (n = 273), ProQuest (n = 6), Embase (n = 88), and Google Scholar (n = 100). After removing duplicates, 383 records remained for screening. Based on title and abstract review, 330 records were excluded for irrelevance, non-ML approaches, or non-ultrasound imaging. Fifty-three full-text articles were reviewed for eligibility, and 33 were excluded for the following reasons: studies focused only on plaque detection (n = 14), lacked risk stratification (n = 10), were abstracts or editorials (n = 2), did not use ultrasound imaging (n = 4), or were review papers (n = 3) (Supplementary table 3). Ultimately, 20 studies met the inclusion criteria and were included in the qualitative synthesis, of which 13 provided sufficient data for the quantitative meta-analysis.

Fig. 1.

Fig. 1

Prisma diagram for study selection

Study characteristics

A total of 20 studies were analyzed, conducted across countries including Japan (n = 4), China (n = 4), the United States (n = 4), the United Kingdom (n = 3, including collaborations with Cyprus and Portugal), Japan combined with Hong Kong (n = 2), the Czech Republic (n = 1), and Greece (n = 1). Sample sizes varied significantly, ranging from 15 subjects to 413 patients; image counts ranged from 115 to 81,000 per study, and some studies reported large datasets (13,810 and 15,546 images) without patient counts. The mean age, reported in 11 studies, ranged from 27.5 ± 3.5 to 75 years, with most studies reporting mean ages between 60 and 69.9 years. Male representation, reported in 8 studies, ranged from 47% to 81%, with one Hong Kong cohort being entirely female. Validation strategies varied: 10-fold cross-validation in 7 studies, 5-fold in 5, 3-fold in 2, 4-fold in 1, and fixed splitting in 5. This reflects diverse approaches to model robustness assessment. Ultrasound imaging modalities predominantly included B-mode ultrasound (n = 15), with some studies incorporating Doppler ultrasound (n = 5) or contrast-enhanced ultrasound (CEUS; n = 1), and one study unspecified beyond ultrasound. Input characteristics for ML models included radiomics features (gray-level co-occurrence matrix [GLCM], gray-level run-length matrix [GLRLM], fractal dimension, texture features; n = 12), clinical features (hemoglobin A1c, lipid profiles, serum markers; n = 5), or a combination of clinical and radiomics features (n = 3). Radiomics methods were semi-automated in 11 studies, manual in 5 (sonographer-defined regions of interest [ROI]), and automated in 4, highlighting variability in feature-extraction workflows. Plaque classification grades varied, with 10 studies classifying plaques as stable versus unstable, four as high-risk versus low-risk, three as symptomatic versus asymptomatic, and 3 using multi-grade systems (normal/mild/moderate/severe or low/moderate/high-risk) (Table 1).

Table 1.

Studies characteristics

First author/Year Type of study Country Validation Ultrasound imaging Input characteristics Method of radiomics Sample size Plaque grade Mean age Male% Algorithms
Biswas et al. (2024) [10] Cohort study Japan 5-fold cross-validation Ultrasound (B-mode) Clinical features (Hemoglobin A1c, LDL cholesterol, HDL cholesterol, total cholesterol, creatinine, etc.) Automated 204 patients (407 images) Low/moderate/high-risk 69 ± 11 77 MultiNet 2.0
Chen et al. (2023) [11] Cohort study China 5-fold cross-validation Ultrasound (Doppler) Clinical features (carotid Doppler ultrasound images) Semi-automated 568 images Stable/unstable NR NR

Inception_v3

ResNet50

Singh et al. (2023) [14] Experimental study Cyprus, UK 5-fold cross-validation Ultrasound (B-mode) Radiomics (GLCM, GLRLM, Fractal dimension) Semi-automated 190 patients (223 images) High-risk/Low-risk 54 years in Database 1, 54 years in Database 2, and 27.5 ± 3.5 years in Database 3 58% in Database 1 and 60% in Database 2

SVM

ResNet-50

GoogLeNet

XceptionNet

Squeeze-Net

VGG16

Zhou et al. (2023) [15] Experimental study USA Fixed splitting Ultrasound (B-mode, Doppler) Clinical features (ultrasound video clips of carotid plaque) Semi-automated 169 patients (63225 images) Mild/severe NR NR

CSG-3DCT

I3D

SlowFast

TPN

TimeSformer

Vivit

MTV (B/2+S/8)

NL I3D

UniFormer

Kybic et al. (2025) [16] Experimental study USA 3-fold cross-validation Ultrasound (B-mode) Radiomic features Semi-automated 400 patients (6931 images) Stable/unstable 69 NR ResNet 18
Jain et al. (2023) [17] Experimental study USA 5-fold cross-validation Ultrasound (B-mode) Radiomic features Semi-automated 97 patients (970 images) Stable/unstable 75 47

UNet

SegNet-UNet

Kigka et al. (2021) [18] Experimental study USA 10-fold cross-validation Ultrasound (B-mode) Clinical Data - Risk Factors- Laboratory Data- Serum Markers- Imaging Semi-automated 208 patients Stable/unstable NR NR

SVM

RF

NB

J48

ANN

Lindsey et al. (2019) [19] Experimental study USA 10-fold cross-validation Ultrasound (Doppler) Radiomic features Semi-automated 13810 image Dataset 1/15546 image Dataset 2 Normal/mild/moderate/severe NR NR

ResNet-50

BN-Inception

SE-ResNext-50

SE-ResNext-50

BN-Inception

ResNet-50

Azzopardi et al. (2020) [20] Experimental study UK Fixed splitting Ultrasound (B-mode) Radiomic features Semi-automated 15 subjects (81000 images) Stable/unstable NR NR

UNET+GDL+GC

UNET+CE

UNET+GDL

UNET+CE+GC

Guang et al. (2021) [21] cohort study China 3-fold cross-validation Ultrasound (B-mode, CEUS image) Radiomic features Automated 205 patients Stable/unstable 61.6±8.4 81

DL-DCCP

Xception (Deep CNN)

RA-CEUS

Yoshidomi et al. (2024) [22] Experimental study Japan 5-fold cross-validation Ultrasound (B-mode) Radiomics (Plaque movement and surface images) Manual (Sonographer manually sets ROI) 200 patients High risk/low risk NR NR

CNN-BiLSTM

CNN-FC

CNN-LSTM

Kybic et al. (2024) [23] Experimental study Czech Fixed splitting Ultrasound (transversal images) Image-based (radiomics) Semi-automated 413 total patients Stable/unstable 69 NR CNN
Jain et al. (2021) [24] Cohort study Japan, Hong Kong 10-fold cross-validation Ultrasound (B-mode) Radiomics Semi-automated

Japanese Cohort: 165 patients (330 images)

Hong Kong Cohort: 50 patients (300 images)

Stable/unstable 68.25 years in the Japanese cohort and 60.2 years in the Hong Kong cohort 77% in the Japanese cohort and all female in the Hong Kong cohort UNet
Yuan et al. (2022) [25] Experimental study China 10-fold cross-validation Ultrasound (2D longitudinal) Radiomics Manual 90 patients (115 images) Moderate/severe 60 NR

HRU-Net with transfer learning

U-Net

FCN

Attention U-Net

DeepLabv3

M-Net

GCN

LinkNET

Jain et al. (2022) [26] Experimental study United Kingdom (DB1), Japan (DB2), Hong Kong (DB3) 5-fold cross-validation Ultrasound (B-mode) Radiomics Semi-automated DB1 (UK): 99 patients, DB2 (Japan): 190 patients, DB3 (Hong Kong): 50 patients, 679 images High-risk/Low-risk 68.78 62

Attention-UNet

UNet

UNet++

UNet3P

Fractal-UNet

Squeeze-UNet

Liu et al. (2024) [27] Experimental study China Fixed splitting Ultrasound Clinical features and radiomics Manual 491 patients (512 images) Stable/unstable 62.7 49

CANet

U-Net

U-Net++

DeepLabV3+

FPN

K. Jain et al. (2022) [28] Experimental study Japan 10-fold cross-validation Ultrasound (B-mode) Radiomics (texture features from plaque images) Manual 190 patients (379 images) Stable/unstable 68.78 ± 10.88 77.4

UNet

SegNet-UNet

Saba et al. (2020) [29] Cohort study UK 10-fold cross-validation Ultrasound (Doppler) Radiomics (texture features from plaque images) Automated 346 patients (2311 images) Symptomatic/asymptomatic 69.9 ± 7.8 61 Deep Learning (CNN)
Saba et al. (2021) [30] Experimental study UK, Portugal 10-fold cross-validation Ultrasound (Doppler) Radiomics (texture features from plaque images) Manual 506 patients (1260 images) Symptomatic/asymptomatic 69.9 ± 7.8 years in the London cohort and 67.5 ± 0.77 years in the Portugal cohort. 61% in the London cohort, and gender distribution was not specified for the Portugal cohort. Deep Learning (DCNN)
Ganitidis et al. (2021) [31] Experimental study Greece 4-fold cross-validation Ultrasound (B-mode) Radiomics (texture features from plaque images) Automated 53 patients Symptomatic/asymptomatic NR NR Simple averaging ensemble scheme

List of abbreviations: MultiNet 2.0: MultiNet 2.0 (Deep Learning-based model); VGG13: VGG13 architecture; Inception_v3: Inception version 3; ResNet50: Residual Network 50; SVM: Support Vector Machine; GoogLeNet: GoogLeNet architecture; XceptionNet: Xception Network; Squeeze-Net: SqueezeNet architecture; VGG16: VGG16 architecture; CSG-3DCT: CSG-3DCT model; I3D: Inflated 3D ConvNet; SlowFast: SlowFast network; TPN: Temporal Pyramid Network; TimeSformer: TimeSformer model; Vivit: Vision Transformer; MTV (B/2+S/8): Mixed Temporal Vision (B/2+S/8); NL I3D: Non-Local Inflated 3D ConvNet; UniFormer: UniFormer architecture; ResNet18: Residual Network 18; UNet: U-Net architecture; SegNet-UNet: SegNet-U-Net hybrid; RF: Random Forest; NB: Naive Bayes; J48: J48 Decision Tree; ANN: Artificial Neural Network; SE-ResNext-50: Squeeze-and-Excitation ResNext 50; BN-Inception: Batch Normalization Inception UNET+GDL+GC: U-Net with Generalized Dice Loss and Graph Cuts; DL-DCCP: Deep Learning with Dynamic Contrast-Enhanced CT Perfusion; RA-CEUS: Radiomics with Contrast-Enhanced Ultrasound; CNN-BiLSTM: Convolutional Neural Network with Bidirectional Long Short-Term Memory; CNN-FC: Convolutional Neural Network with Fully Connected layers; CNN-LSTM: Convolutional Neural Network with Long Short-Term Memory; HRU-Net: High-Resolution U-Net with transfer learning; FCN: Fully Convolutional Network; DeepLabv3: DeepLab version 3; M-Net: M-Net architecture; GCN: Graph Convolutional Network; LinkNET: LinkNet architecture; UNet3P: U-Net 3+ architecture; CANet: Context Aggregation Network; FPN: Feature Pyramid Network; Deep Learning (DCNN): Deep Learning with various Deep Convolutional Neural Network layer; B-mode: Brightness-mode ultrasound imaging; CEUS: Contrast-Enhanced Ultrasound; CT: Computed Tomography; DL: Deep Learning; GLCM: Gray-Level Co-occurrence Matrix; GLRLM: Gray-Level Run-Length Matrix; ROI: Region of Interest; US: Ultrasound; 2D: Two-Dimensional; 3D: Three-Dimensional; Fractal Dimension: Quantitative measure describing structural complexity of image texture; Radiomics: Quantitative extraction of high-dimensional imaging features from medical images; Texture Features: Quantitative metrics derived from image intensity and spatial patterns; DB: Database; LDL: Low-Density Lipoprotein; HDL: High-Density Lipoprotein; NR: Not Reported; Cross-Validation: Statistical technique for assessing model generalizability by partitioning data into folds; Fixed Splitting: Predetermined training/validation data separation method; Patients (n): Sample size indicator

The ML models were diverse, predominantly categorized into deep learning (DL) and U-Net-based architectures, with a few incorporating traditional ML approaches. Deep learning models included convolutional neural network (CNN)-based architectures such as MultiNet 2.0, VGG13, Inception_v3, ResNet18, SE-ResNext-50, CNN-BiLSTM, CANet, and various deep convolutional neural networks (DCNNs) with differing layer configurations, alongside specialized models like DL-DCCP (Deep Learning-Based Detection and Classification of Carotid Plaques) and a simple averaging ensemble scheme with a 0.465 threshold. U-Net-based models tailored for segmentation tasks include UNet, UNet+GDL+GC (integrating global context and guided deep learning), HRU-Net with transfer learning, Attention-UNet, and UNet-based deep learning variants, highlighting their utility for precise plaque feature extraction. Additionally, support vector machines (SVMs) were used in two studies, representing traditional ML approaches, while one study employed CSG-3DCT, a less common 3D convolutional neural network.

Of the 20 studies identified, 13 provided sufficient performance data for inclusion in the meta-analysis [11, 14, 16, 17, 21–24, 26–29].

Pooled diagnostic performance

This meta-analysis included 13 studies to evaluate the diagnostic performance of machine learning and deep learning models in ultrasound-based carotid plaque risk stratification. The best-performing model from each study, as reported by the original authors, was used for analysis to ensure comparability and represent the optimal diagnostic performance of each approach. Detailed model performance metrics are presented in Table 2. The pooled sensitivity was 0.84 (95% CI: 0.74–0.90) and specificity was 0.96 (95% CI: 0.89–0.98) (Fig. 2), with a pooled AUC of 0.95 (95% CI: 0.93–0.97) (Fig. 3). Substantial heterogeneity was observed across studies (I2 = 88.8% for sensitivity, 64.1% for specificity, and 68.1% overall).

Table 2.

Performance metrics of different ml models. The first row for each study indicates the best-performing model reported by the authors

First author/Year Machine learning algorithms Accuracy Sensitivity Specificity Precision F1score AUC
Biswas et al. (2024) [10] MultiNet 2.0 0.88 NR NR NR NR 0.88
Chen et al. (2023) [11] Inception_v3 0.9294 NR NR NR NR 0.915
ResNet50 0.8941 NR NR NR NR 0.853
Singh et al. (2023) [14] SVM 0.9685 0.9529 0.975 0.9732 0.975 NR
ResNet-50 0.8986 0.9714 0.775 0.8804 0.9224 NR
GoogLeNet 0.98 0.9857 0.9765 0.986 0.9856 NR
XceptionNet 0.9414 0.963 0.9044 0.9453 0.9539 NR
Squeeze-Net 0.9508 0.9714 0.9176 0.9518 0.9602 NR
VGG16 0.9777 0.9929 0.9529 0.9724 0.9823 NR
Zhou et al. (2023) [15] CSG-3DCT 0.831 0.824 NR 0.828 0.825 NR
I3D 0.788 0.808 NR 0.81 0.788 NR
SlowFast 0.703 0.708 NR 0.703 0.702 NR
TPN 0.78 0.785 NR 0.778 0.778 NR
TimeSformer 0.771 0.759 NR 0.768 0.762 NR
Vivit 0.703 0.708 NR 0.703 0.702 NR
MTV (B/2+S/8) 0.72 0.734 NR 0.73 0.72 NR
NL I3D 0.788 0.808 NR 0.81 0.788 NR
UniFormer 0.754 0.776 NR 0.781 0.754 NR
Kybic et al. (2025) [16] ResNet 18 0.48 0.932 0.098 0.466 0.621 0.614
Jain et al. (2023) [17] UNet 0.9856 0.903 0.9923 0.9923 NR 0.91
SegNet-UNet 0.9844 0.882 0.9927 0.9069 NR 0.905
Kigka et al. (2021) [18] SVM 0.76 NR NR NR NR 0.73
RF 0.68 NR NR NR NR 0.74
NB 0.58 NR NR NR NR 0.62
J48 0.62 NR NR NR NR 0.61
ANN 0.69 NR NR NR NR 0.72
Lindsey et al. (2019) [19] ResNet-50 0.9555 0.9448 NR 0.9884 0.9661 NR
BN-Inception 0.8761 0.88 NR 0.89 0.88 NR
SE-ResNext-50 0.8935 0.89 NR 0.9 0.89 NR
SE-ResNext-50 0.9781 0.9781 NR 0.9782 0.9781 NR
BN-Inception 0.9701 0.9701 NR 0.9704 0.9701 NR
ResNet-50 0.973 0.973 NR 0.973 0.973 NR
Azzopardi et al. (2020) [20] UNET+GDL+GC NR 0.94 0.969 NR NR NR
UNET+CE NR 0.937 0.966 NR NR NR
UNET+GDL NR 0.936 0.965 NR NR NR
UNET+CE+GC NR 0.939 0.969 NR NR NR
Guang et al. (2021) [21] DL-DCCP NR 0.792 0.844 NR NR 0.87 (0.82–0.92)
Xception (Deep CNN) NR 0.75 0.756 NR NR 0.77 (0.71–0.83)
RA-CEUS NR 0.875 0.444 NR NR 0.66 (0.56–0.76)
Yoshidomi et al. (2024) [22] CNN-BiLSTM 0.805 0.802 NR 0.81 NR NR
CNN-FC 0.775 0.771 NR 0.78 NR NR
CNN-LSTM 0.787 0.785 NR 0.8 NR NR
Kybic et al. (2024) [23] CNN 0.528 0.725 0.485 0.234 0.354 0.609
Jain et al. (2021) [24] UNet 0.9901 0.8637 0.9952 0.8855 0.8668 0.95
Yuan et al. (2022) [25] HRU-Net with transfer learning 0.977 NR NR NR NR NR
U-Net 0.969 NR NR NR NR NR
FCN 0.965 NR NR NR NR NR
Attention U-Net 0.968 NR NR NR NR NR
DeepLabv3 0.969 NR NR NR NR NR
M-Net 0.968 NR NR NR NR NR
GCN 0.965 NR NR NR NR NR
LinkNET 0.967 NR NR NR NR NR
Jain et al. (2022) [26] Attention-UNet 0.9858 0.8686 0.9952 0.9354 NR 0.99
UNet 0.9858 0.8743 0.9948 0.9311 NR 0.989
UNet++ 0.9851 0.8572 0.9953 0.937 NR 0.988
UNet3P 0.9851 0.8673 0.9947 0.929 NR 0.988
Fractal-UNet 0.9841 0.8568 0.9945 0.9261 NR 0.962
Squeeze-UNet 0.9853 0.8629 0.9951 0.934 NR 0.969
Liu et al. (2024) [27] CANet NR 0.961 NR 0.947 NR NR
U-Net NR 0.9055 NR 0.9123 NR NR
U-Net++ NR 0.8749 NR 0.8947 NR NR
DeepLabV3+ NR 0.9061 NR 0.9108 NR NR
FPN NR 0.9269 NR 0.9147 NR NR
K. Jain et al. (2022) [28] UNet 0.9908 0.886 0.9956 0.8905 NR 0.94
SegNet-UNet 0.9897 0.915 0.9926 0.8342 NR 0.93
Saba et al. (2020) [29] Deep Learning (CNN) 0.897 NR NR NR NR 0.91
Saba et al. (2021) [30] Deep Learning 0.8807 0.786 0.983 NR 0.872 0.88
Ganitidis et al. (2021) [31] Simple averaging ensemble 0.725 0.75 0.7 NR NR 0.73 (0.623–0.837)

Fig. 2.

Fig. 2

Forest plot for sensitivity and specificity

Fig. 3.

Fig. 3

Summary receiver operating characteristic (SROC) plot

Sensitivity (influence) analysis

A leave-one-out sensitivity analysis was performed using a random-effects model with restricted maximum likelihood (REML) estimation to assess the robustness of the pooled diagnostic odds ratio (lnDOR). Excluding each study sequentially produced minimal variation in the overall effect size (pooled lnDOR = 4.63; 95% CI = 3.17–6.09; p < 0.001), indicating that no single study disproportionately influenced the overall meta-analytic estimate. This consistency supports the stability and reliability of the pooled diagnostic performance across included studies (Fig. 4) (Supplementary table 4).

Fig. 4.

Fig. 4

Leave-one-out forest plot

Meta-regression analysis

Meta-regression identified sample size and model architecture as significant contributors to between-study heterogeneity, explaining a substantial proportion of variability (I2 = 72%). In contrast, factors such as study region showed a borderline effect, suggesting a moderate but non-significant contribution to heterogeneity. Plaque characteristics, validation strategy, and ultrasound modality did not significantly account for residual heterogeneity (Supplementary Table 5).

Subgroup analysis

Subgroup analysis by plaque type

For the plaque type subgroup, studies comparing stable versus unstable plaques demonstrated a pooled sensitivity of 0.82 (95% CI: 0.66–0.91) and specificity of 0.96 (95% CI: 0.85–0.99). Substantial heterogeneity was observed overall (I2 = 54.61%), particularly for sensitivity (I2 = 83.29%), with moderate heterogeneity for specificity (I2 = 59.28%). Meta-analysis could not be performed for high-risk vs. low-risk (n = 3) and symptomatic vs. asymptomatic (n = 2) subgroups due to the limited number of studies (Supplementary figure 1).

Subgroup analysis by imaging modality

For the imaging-based subgroup, studies using B-mode ultrasound (n = 8) reported a pooled sensitivity of 0.85 (95% CI: 0.73–0.92) and specificity of 0.96 (95% CI: 0.88–0.99). Overall heterogeneity was substantial (I2 = 65.31%), primarily driven by high variability in sensitivity (I2 = 84.72%), while heterogeneity in specificity was moderate (I2 = 54.80%). These findings suggest that differences in imaging parameters and interpretation criteria across studies may partly explain the variability in diagnostic performance (Supplementary Figure 2).

Subgroup analysis by model architecture

In the model-based subgroup, Deep Learning models (n = 5) demonstrated a pooled sensitivity of 0.79 (95% CI: 0.74–0.83) and specificity of 0.95 (95% CI: 0.79–0.99), with minimal overall heterogeneity (I2 = 0.02%), low variability in sensitivity (I2 = 13.69%), and moderate heterogeneity in specificity (I2 = 63.12%) (Supplementary figure 3).

In contrast, network-based models (n = 5) demonstrated a pooled sensitivity of 0.85 (95% CI: 0.68–0.94) and specificity of 0.98 (95% CI: 0.92–1.00). Overall heterogeneity was negligible (I2 = 0.01%), though sensitivity showed high variability (I2 = 79.58%) and specificity exhibited moderate heterogeneity (I2 = 40.03%). Meta-analysis could not be conducted for CNN (n = 2) and SVM (n = 1) subgroups due to the limited number of studies (Supplementary figure 4).

Subgroup analysis by region

In the regional subgroup analysis, studies conducted in East Asia demonstrated a pooled sensitivity of 0.84 (95% CI: 0.70–0.93) and specificity of 0.96 (95% CI: 0.79–0.99), indicating generally high diagnostic accuracy. However, substantial heterogeneity was observed (I2 = 76.13% for sensitivity and 59.63% for specificity and general heterogeneity of 48.18%) (supplementary figure 5). For studies conducted in Europe (n = 2), North America (n = 3), and multi-regional datasets (n = 3), meta-analysis was not performed due to the limited number of studies ( < 4), which precluded reliable pooled estimates.

Assessment of publication bias

Publication bias was investigated using Deeks’ funnel plot asymmetry test, which yielded a p-value of 0.36, indicating no significant asymmetry in the funnel plot (Fig. 5). This result suggests no risk of publication bias affecting the meta-analysis outcomes.

Fig. 5.

Fig. 5

Funnel plot

Quality assessment

In the model development phase, 90% of studies (n = 18) exhibited a low risk of bias across the Participants, Predictors, Outcomes, and Analyses domains, reflecting well-defined participant selection, consistent and objective predictor measurement, standardized outcome definitions, and robust analytical frameworks. Only two studies (10%) were rated as having a moderate risk due to incomplete reporting of predictor preprocessing steps or insufficient detail on analytical procedures. No studies were deemed at high risk of bias.

During model evaluation, the majority of studies (71%, n = 15) demonstrated low risk across all domains, supported by representative evaluation datasets, consistent predictor assessment, and rigorous analytical validation methods. Nonetheless, four studies (19%) showed moderate risk and two (10%) high risk, primarily due to limited transparency in validation protocols or potential overlap between training and evaluation datasets, which may have introduced bias in performance estimation.

Regarding applicability to carotid plaque classification, all studies (100%, n = 20) were judged to have low risk across all domains, indicating strong relevance and consistency with clinical practice. Overall, 70% of studies were classified as low risk, 20% as moderate, and 10% as high, underscoring generally sound methodological quality but highlighting residual concerns related to evaluation transparency and external generalizability (Supplementary figure 6).

Grade assessment

Evidence started as high certainty, reflecting observational diagnostic data. Risk of bias was rated low in 70% of studies per PROBAST+AI, supporting high initial certainty, though 20% moderate and 10% high ROB indicated some methodological concerns. Inconsistency was severe, with high I2 values, leading to a one-level downgrade due to substantial heterogeneity. Indirectness was minimal, as studies aligned with the review’s ultrasound-based focus. Imprecision varied: narrow confidence intervals for AUC and Specificity maintained certainty, while wider intervals for Sensitivity and likelihood ratios prompted a one-level downgrade. Overall certainty was rated moderate due to inconsistency and bias.

Discussion

Overall, ML models showed promising diagnostic performance for classifying carotid plaques from ultrasound images, although substantial variability was observed across studies. Meta-regression indicated that differences in sample size and model architecture contributed to much of this heterogeneity, primarily due to inconsistency and methodological differences. Moving forward, it is essential to compare different models to identify their respective strengths and weaknesses and explore strategies for integrating these approaches into clinical decision-making.

MRI vs ultrasound

MRI is a powerful tool for assessing carotid plaque vulnerability, visualizing key features such as intraplaque hemorrhage (IPH), lipid-rich necrotic core (LRNC), fibrous cap (FC) status, inflammation, and neovascularization, with high histologic validation [32]. IPH appears hyperintense on T1-weighted images, while LRNC is T2-hypointense and non-enhancing on contrast-enhanced scans; plaques with > 40% LRNC and thin FC are prone to rupture and major events [33]. MRI also tracks treatment effects like LRNC reduction with statins, and analyses show IPH is a stronger predictor of stroke than stenosis or clinical factors [32, 33]. The MATCH sequence shortens scan time to 7 minutes (vs. 40 minutes) with high sensitivity/specificity for IPH (≥89%/≥91%) and LRNC (≥81%/85%), though moderate agreement for calcification and fibrous tissue [34].

Ultrasound, while cost-effective and real-time for stenosis and neovascularization detection, is operator-dependent and limited in visualizing plaque characteristics [35, 36]. It measures smaller plaque areas (≈1.4-fold less than MRI) but excels in hemodynamic evaluation. MRI offers superior tissue characterization and stroke prediction, making ultrasound best for screening and MRI for detailed risk assessment. Emerging ultrasound strain imaging may help differentiate plaque composition, but lacks MRIs full validation [33].

Comparison of different ML models

Among the included studies, Kybic et al. (2025) [16] and Kybic et al. (2024) [23] exhibited lower performance (AUC 0.61 and 0.609, respectively) compared to the robust outcomes of other studies, potentially due to methodological limitations that hindered their efficacy. In Kybic et al. (2025), the use of a regression-based approach to predict plaque width (d) and increase (D) diverged from the classification focus of most studies, relying on a ResNet-18 network with a modest Dice score improvement (0.790 to 0.794) via self-supervised augmentation consistency (SAC). This modest gain suggests insufficient leverage of the large unlabeled dataset (6781 images), possibly due to the simplistic augmentation strategy (e.g., flips, rotations) and a Dice loss that, despite macro-averaging, may not adequately prioritize plaque-specific features over background classes. The regression network’s reliance on pre-segmented inputs, combined with channel dropout (probability 0.5), risks losing critical plaque texture data. At the same time, the small validation set (29 images) likely led to overfitting, undermining generalizability [16]. Similarly, Kybic et al. (2024) employed a U-Net with ResNet-34 and GANs for segmentation, followed by a shallow CNN for multicriterion regression of plaque attributes (e.g., width, echogenicity). The use of weak annotations (ellipsoidal masks) for 1399 images, while innovative, introduced noise from incomplete pixel labeling, and the GANs discriminator training instability (e.g., toggling at 40–60% accuracy) may have produced inconsistent segmentations. The regression network’s small architecture and heavy reliance on segmented masks—vulnerable to upstream errors—coupled with a limited training set (568 images for width increase), likely contributed to its poor predictive power (AUC 0.609). These methodological choices—regression over classification, suboptimal use of large datasets, and unstable training strategies—contrast with the classification-focused, well-validated deep learning models (e.g., GoogLeNet, UNet variants) in other studies, explaining their outsized contribution to heterogeneity (I2 > 90%) and the downward pull on pooled estimates like Accuracy (0.854) [23].

ML models vs conventional methods

Traditional carotid plaque assessment relied primarily on manual ultrasound interpretation and histological analysis, with quantitative metrics reported inconsistently across studies. Rafailidis et al. quantified plaque surface irregularity using the Surface Irregularity Index (SII) via color Doppler ultrasound (CDUS) and CEUS, demonstrating high interobserver agreement (ICC 0.979 for CDUS, 0.952 for CEUS), yet their classification of plaque surfaces showed no significant correlation with stroke occurrence and lacked sensitivity or specificity metrics [37]. Baldassarre et al. measured carotid intima-media thickness (cIMT) and internal carotid artery diameter (ICCAD), reporting hazard ratios of 1.27–1.47 per SD increase and a net reclassification improvement (NRI) of up to 20.1%, but did not provide diagnostic accuracy measures, reflecting the subjective nature of visual cIMT assessment [38]. Sonaglioni et al. found thicker cIMT in idiopathic pulmonary fibrosis patients (1.5 ± 0.3 mm vs. 1.1 ± 0.3 mm, p < 0.0001) with moderate intra-observer agreement (ICC 0.79–0.94), yet qualitative echogenicity assessment lacked quantitative accuracy metrics. CEUS improved diagnostic performance by visualizing plaque neovascularity [39]; Luo et al. reported Sensitivity 83–87.1%, Specificity 64.7–81.5%, and AUC 0.729–0.911 for plaque characterization, although manual plaque selection and modest repeatability limited scalability [40]. Histological validation demonstrated associations of ADAMTS4 and versican expression with plaque vulnerability (OR 1.144, 95% CI 1.007–1.300), but real-time application and diagnostic metrics were absent [41].

Non-machine learning meta-analyses on carotid imaging provide critical insights into stroke risk stratification using ultrasound, yet reveal limitations in prognostic accuracy and consistency. In 13 studies with 3,092 participants, ischemic stroke patients had higher plaque enhancement intensity (SMD = 0.71, 95% CI: 0.32–1.11) and a greater likelihood of plaque enhancement (OR = 3.25, 95% CI: 1.86–5.68), with pooled diagnostic sensitivity and specificity of 0.68 and 0.61 [42]. Increased carotid intima-media thickness (cIMT) per SD was associated with higher risks of coronary heart disease (HR = 1.10), stroke (HR = 1.08), and CVD (HR = 1.14), while MRI-derived wall thickness showed even stronger associations (HR = 1.27–1.58) and remained significant when modeled with cIMT [43]. Across 119 RCTs (100,667 patients), each 10 μm/year reduction in cIMT progression corresponded to a 9% lower CVD risk (RR = 0.91, 95% CrI: 0.87–0.94), consistent across interventions [44]. In 11 studies, echolucent carotid plaques predicted future cardiovascular events in asymptomatic (RR = 2.72) and recurrent symptoms in symptomatic patients (RR = 2.97), especially with severe stenosis and modern ultrasound [45]. However, analysis of 41 trials (18,307 participants) found no link between cIMT regression and outcomes (CHD OR = 0.82; stroke OR = 0.71), indicating cIMT regression may not reliably predict clinical endpoints [46]. Unlike traditional approaches, which often rely on subjective visual assessment or limited quantitative metrics, ML models can automatically extract high-dimensional features from large datasets, capturing subtle patterns associated with plaque vulnerability. This capability allows ML algorithms to generate personalized risk scores for individual patients, improve early detection of high-risk plaques, and potentially outperform conventional methods.

Clinical applicability

ML models represent a paradigm shift in carotid plaque classification, offering straightforward clinical utility when integrated appropriately. This automation enhances diagnostic accuracy and consistency in risk stratification. At the same time, there is a need for analysis of large-scale datasets, external validation, standardized datasets across diverse patient populations and imaging modalities, to improve generalizability. For high-risk patient identification, ML algorithms can combine multimodal data, including clinical variables and imaging features, to generate personalized vulnerability scores, which could guide thresholds for intervention such as endarterectomy or intensified medical therapy and facilitate early detection of unstable plaques prone to rupture. To support clinical adoption, these models should be validated on representative cohorts, with performance benchmarks guiding integration into the workflow. Additionally, the scalability of ML can streamline resource allocation, reduce radiologist workload through automated triage, and optimize follow-up strategies, fostering cost-effective, evidence-based care in vascular medicine. On the other hand, the reliance of ML models on large datasets and advanced computational resources may limit their practical implementation in routine clinical settings.

Limitations

This meta-analysis has several limitations. Substantial heterogeneity was observed, particularly for sensitivity and specificity, suggesting that pooled estimates may not directly translate to broader clinical practice. Most studies originated from Southeast Asia, creating geographic clustering that may limit applicability to other regions. Possible dataset overlap across studies could bias pooled estimates and overestimate model performance. Many studies did not validate ultrasound-based plaque classification against histopathology or longitudinal clinical outcomes, making it difficult to assess true diagnostic or prognostic value. Criteria for plaque classification varied across studies, reflecting a lack of standardization in defining plaque stability or vulnerability. Direct comparisons between machine learning models and conventional diagnostic methods were generally lacking, preventing head-to-head evaluation.

Future direction

Future research should prioritize model interpretability, integration of multimodal data, and validation across diverse populations. Emphasis should be placed on explainable AI to build clinician trust, combining imaging with genomic, proteomic, and clinical data for personalized risk assessment. Large, multi-center studies are needed to ensure generalizability and address biases, while real-time deployment in point-of-care ultrasound could improve efficiency and guide targeted interventions.

Conclusion

This meta-analysis suggests that machine learning models may have potential for carotid plaque risk classification, with generally high sensitivity, specificity, and overall accuracy. The overall low risk of bias, as assessed by PROBAST+AI, provides some support for these findings, but substantial heterogeneity and variability across studies limit confidence in the pooled estimates. Differences in sample size, model architecture, and methodology likely contribute to this variability. Further well-designed, standardized studies with external validation are needed to clarify the reliability and generalizability of ML-based diagnostic approaches.

Electronic supplementary material

Below is the link to the electronic supplementary material.

Supplementary Material 1 (769KB, docx)

Abbreviations

ANN

Artificial Neural Network

B-mode

Brightness-mode ultrasound imaging

CEUS

Contrast-Enhanced Ultrasound

CNN

Convolutional Neural Network

DL

Deep Learning

DCNN

Deep Convolutional Neural Network

FCN

Fully Convolutional Network

GLCM

Gray-Level Co-occurrence Matrix

GLRLM

Gray-Level Run-Length Matrix

HRU-Net

High-Resolution U-Net with transfer learning

Inception_v3

Inception version 3

ML

Machine Learning

ResNet18

Residual Network 18

ResNet50

Residual Network 50

RF

Random Forest

ROI

Region of Interest

SVM

Support Vector Machine

UNet

U-Net architecture

US

Ultrasound

VGG16

Visual Geometry Group 16 architecture

XceptionNet

Xception Network

Author contributions

P.E.1 was a primary contributor in the design, implementation, and writing of the manuscript. M.R., H.S., and P.E.2 independently assessed articles and extracted data. All authors read and approved the final manuscript. M.R. and J.T. performed statistical analysis.

Funding

None.

Data availability

All data generated or analyzed during this study are included in this published article [and its supplementary information files].

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Generative AI

In the preparation of this article, the authors utilized the Grammarly application to enhance linguistic accuracy and clarity. The manuscript underwent meticulous double-checking to ensure precision, and the authors assume full responsibility for the integrity and originality of the content presented herein.

Randomized controlled trial number

Not applicable.

Competing interests

TThe authors declare no competing interests.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Strilciuc S, Grad DA, Radu C, et al. The economic burden of stroke: a systematic review of cost of illness studies. J Med Life. 2021;14(5):606–19. 10.25122/jml-2021-0361. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Global, regional, and national burden of stroke and its risk factors, 1990-2021: a systematic analysis for the global burden of disease study 2021. Lancet Neurol. 2024;23(10):973–1003. 10.1016/s1474-4422(24)00369-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Martin SS, Aday AW, Allen NB, et al. Heart disease and stroke statistics: a report of US and global data from the American Heart Association. Circulation. 2025 2025;151(8):e41–660. 10.1161/cir.0000000000001303. [DOI] [PMC free article] [PubMed]
  • 4.Cheng Y, Lin Y, Shi H, et al. Projections of the stroke burden at the global, regional, and national levels up to 2050 based on the global burden of disease study 2021. J Am Heart Assoc. 2024;13(23):e036142. 10.1161/jaha.124.036142. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Diener HC, Easton JD, Hart RG, Kasner S, Kamel H, Ntaios G. Review and update of the concept of embolic stroke of undetermined source. Nat Rev Neurol. 2022;18(8):455–65. 10.1038/s41582-022-00663-4. [DOI] [PubMed] [Google Scholar]
  • 6.Kamtchum-Tatuene J, Wilman A, Saqqur M, Shuaib A, Jickling GC. Carotid plaque with high-risk features in embolic stroke of undetermined source: systematic review and meta-analysis. Stroke. 2020;51(1):311–14. 10.1161/strokeaha.119.027272. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Ospel JM, Marko M, Singh N, Goyal M, Almekhlafi MA. Prevalence of non-stenotic (<50%) carotid plaques in acute ischemic stroke and transient ischemic attack: a systematic review and meta-analysis. J Stroke Cerebrovasc Dis. 2020;29(10):105117. 10.1016/j.jstrokecerebrovasdis.2020.105117. [DOI] [PubMed] [Google Scholar]
  • 8.Sajjadi SM, Mohebbi A, Ehsani A, et al. Identifying abdominal aortic aneurysm size and presence using natural language processing of radiology reports: a systematic review and meta-analysis. Abdominal Radiol. 2025. 10.1007/s00261-025-04810-5. [DOI] [PubMed] [Google Scholar]
  • 9.Eini P, Eini P, Serpoush H, Rezayee M, Tremblay J. Machine learning models for carotid artery plaque detection: a systematic review of ultrasound-based diagnostic performance. J Stroke Cerebrovascular Dis. 2025;34(11):108446. 10.1016/j.jstrokecerebrovasdis.2025.108446. [DOI] [PubMed] [Google Scholar]
  • 10.Biswas M, Saba L, Kalra M, et al. MultiNet 2.0: a lightweight attention-based deep learning network for stenosis measurement in carotid ultrasound scans and cardiovascular risk assessment. Computerized Med Imag Graphics. 2024;117:102437. 10.1016/j.compmedimag.2024.102437. [DOI] [PubMed] [Google Scholar]
  • 11.Chen X-X, Kong Z-X, Wei S-F, et al. Ultrasound lmaging-vulnerable plaque diagnostics: automatic carotid plaque segmentation based on deep learning. J Radiat Res Appl Sci. 2023;16(3):100598. 10.1016/j.jrras.2023.100598. [Google Scholar]
  • 12.Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. 10.1136/bmj.n71. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Moons KGM, Damen JAA, Kaul T, et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ. 2025;388:e082505. 10.1136/bmj-2024-082505. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Singh S, Jain PK, Sharma N, Pohit M, Roy S. Atherosclerotic plaque classification in carotid ultrasound images using machine learning and explainable deep learning. Intell Med. 2024;4(2):83–95. 10.1016/j.imed.2023.05.003. [Google Scholar]
  • 15.Zhou X, Huang Y, Xue W, et al. Inflated 3D convolution-transformer for weakly-supervised carotid stenosis grading with ultrasound videos. In: Greenspan H, Madabhushi A, Mousavi P, et al., editors. Medical image computing and computer assisted intervention – MICCAI 2023. Cham: Springer Nature Switzerland; 2023. p. 511–20.
  • 16.Kybic J, Pakizer D, Kozel J, Michalčová P, Charvát F, Školoudík D. Atherosclerotic plaque stability prediction from longitudinal ultrasound images. In: Xu X, Cui Z, Rekik I, Ouyang X, Sun K, editors. Machine learning in medical imaging. Cham: Springer Nature Switzerland; 2025. p. 124–32. [Google Scholar]
  • 17.Jain PK, Sharma N, Roy S. Hybrid deep learning models for segmentation of atherosclerotic plaque in B-mode carotid ultrasound image. In: Sharma S, Subudhi B, Sahu U, editors. Intelligent control, robotics, and industrial automation. Singapore: Springer Nature Singapore; 2023. p. 807–19. [Google Scholar]
  • 18.Kigka VI, Sakellarios AI, Mantzaris MD, et al. A machine learning model for the identification of high-risk carotid atherosclerotic plaques. Annu Int Conf IEEE Eng Med Biol Soc. 2021;2021:2266–69. 10.1109/embc46164.2021.9630654. [DOI] [PubMed] [Google Scholar]
  • 19.Lindsey T, Garami Z. Automated stenosis classification of carotid artery sonography using deep neural networks. 2019 18th IEEE International Conference on Machine Learning and Applications (ICMLA). 2019:1880–84.
  • 20.Azzopardi C, Camilleri KP, Hicks YA. Bimodal automated carotid ultrasound segmentation using geometrically constrained deep neural networks. IEEE J Biomed Health Inf. 2020;24(4):1004–15. 10.1109/jbhi.2020.2965088. [DOI] [PubMed] [Google Scholar]
  • 21.Guang Y, He W, Ning B, et al. Deep learning-based carotid plaque vulnerability classification with multicentre contrast-enhanced ultrasound video: a comparative diagnostic study. BMJ Open. 2021;11(8):e047528. 10.1136/bmjopen-2020-047528. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Yoshidomi T, Kume S, Aizawa H, Furui A. Classification of carotid plaque with jellyfish sign through convolutional and recurrent neural networks utilizing plaque surface edges. 2024 46th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC):1–42024. [DOI] [PubMed]
  • 23.Kybic J, Pakizer D, Kozel J, Michalčová P, Charvát F, Školoudík D. Predikce stability aterosklerotického plátu z transverzálních ultrazvukových obrazů pomocí hlubokého učení. Čes Slov Neurol Neurochir. 2024;87(4).
  • 24.Jain PK, Sharma N, Saba L, et al. Unseen artificial intelligence—deep learning paradigm for segmentation of low atherosclerotic plaque in carotid ultrasound: a multicenter cardiovascular study. Diagnostics. [DOI] [PMC free article] [PubMed]
  • 25.Yuan Y, Li C, Zhang K, Hua Y, Zhang J. HRU-Net: a transfer learning method for carotid artery plaque segmentation in ultrasound images. Diagnostics. [DOI] [PMC free article] [PubMed]
  • 26.Jain PK, Dubey A, Saba L, et al. Attention-based UNet deep learning model for plaque segmentation in carotid ultrasound for stroke risk stratification: an artificial intelligence paradigm. J Cardiovasc Dev Dis. [DOI] [PMC free article] [PubMed]
  • 27.Liu M, Gao W, Song D, et al. A deep learning-based calculation system for plaque stenosis severity on common carotid artery of ultrasound images. Vascular. 2024;17085381241246312. 10.1177/17085381241246312. [DOI] [PubMed]
  • 28.Jain PK, Sharma N, Saba L, et al. Automated deep learning-based paradigm for high-risk plaque detection in B-mode common carotid ultrasound scans: an asymptomatic Japanese cohort study. Int Angiol. 2022;41(1):9–23. 10.23736/s0392-9590.21.04771-4. [DOI] [PubMed]
  • 29.Saba L, Sanagala SS, Gupta SK, et al. Ultrasound-based internal carotid artery plaque characterization using deep learning paradigm on a supercomputer: a cardiovascular disease/stroke risk assessment system. Int J Cardiovasc Imag. 2021;37(5):1511–28. 10.1007/s10554-020-02124-9. [DOI] [PubMed] [Google Scholar]
  • 30.Saba L, Sanagala SS, Gupta SK, et al. A multicenter study on carotid ultrasound plaque tissue characterization and classification using six deep artificial intelligence models: a stroke application. IEEE Trans Instrum Meas. 2021;70:1–12. 10.1109/TIM.2021.3052577.33776080 [Google Scholar]
  • 31.Ganitidis T, Athanasiou M, Dalakleidi K, Melanitis N, Golemati S, Nikita KS. Stratification of carotid atheromatous plaque using interpretable deep learning methods on B-mode ultrasound images. 2021 43rd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC):3902–052021. [DOI] [PubMed]
  • 32.Kassem M, Florea A, Mottaghy FM, van Oostenbrugge R, Kooi ME. Magnetic resonance imaging of carotid plaques: current status and clinical perspectives. Ann Transl Med. 2020;8(19):1266. 10.21037/atm-2020-cass-16. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Porambo ME, DeMarco JK. Mr imaging of vulnerable carotid plaque. Cardiovasc Diagn Ther. 2020;10(4):1019–31. 10.21037/cdt.2020.03.12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Kassem M, Nies KPH, Boswijk E, et al. Quantification of carotid plaque composition with a multi-contrast atherosclerosis characterization (match) MRI sequence. Front Cardiovasc Med. 2023;10:1227495. 10.3389/fcvm.2023.1227495. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Hou C, Li S, Zheng S, et al. Quality assessment of radiomics models in carotid plaque: a systematic review. Quant Imag Med Surg. 2023;14(1):1141–54. https://qims.amegroups.org/article/view/119159. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Hou C, Liu X-Y, Du Y, et al. Radiomics in carotid plaque: a systematic review and radiomics quality score assessment. Ultrasound Med Biol. 2023;49(12):2437–45. 10.1016/j.ultrasmedbio.2023.06.008. [DOI] [PubMed] [Google Scholar]
  • 37.Rafailidis V, Chryssogonidis I, Grisan E, et al. Does quantification of carotid plaque surface irregularities better detect symptomatic plaques compared to the subjective classification? J Ultrasound Med. 2019;38(12):3163–71. 10.1002/jum.15017. [DOI] [PubMed] [Google Scholar]
  • 38.Baldassarre D, Hamsten A, Veglia F, et al. Measurements of carotid intima-media thickness and of interadventitia common carotid diameter improve prediction of cardiovascular events: results of the improve (carotid intima media thickness [IMT] and IMT-progression as predictors of vascular events in a high risk European population) study. J Am Coll Cardiol. 2012;60(16):1489–99. 10.1016/j.jacc.2012.06.034. [DOI] [PubMed] [Google Scholar]
  • 39.Sonaglioni A, Caminati A, Lipsi R, Lombardo M, Harari S. Association between C-reactive protein and carotid plaque in mild-to-moderate idiopathic pulmonary fibrosis. Intern Emerg Med. 2021;16(6):1529–39. 10.1007/s11739-020-02607-6. [DOI] [PubMed] [Google Scholar]
  • 40.Luo X, Li W, Bai Y, Du L, Wu R, Li Z. Relation between carotid vulnerable plaques and peripheral leukocyte: a case-control study of comparison utilizing multi-parametric contrast-enhanced ultrasound. BMC Med Imag. 2019;19(1):74. 10.1186/s12880-019-0374-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Dong H, Du T, Premaratne S, et al. Relationship between ADAMTS4 and carotid atherosclerotic plaque vulnerability in humans. J Vasc Surg. 2018;67(4):1120–26. 10.1016/j.jvs.2017.08.075. [DOI] [PubMed] [Google Scholar]
  • 42.Costanzo P, Perrone-Filardi P, Vassallo E, et al. Does carotid intima-media thickness regression predict reduction of cardiovascular events? A meta-analysis of 41 randomized trials. J Am Coll Cardiol. 2010;56(24):2006–20. 10.1016/j.jacc.2010.05.059. [DOI] [PubMed] [Google Scholar]
  • 43.Willeit P, Tschiderer L, Allara E, et al. Carotid intima-media thickness progression as surrogate marker for cardiovascular risk: meta-analysis of 119 clinical trials involving 100 667 patients. Circulation. 2020;142(7):621–42. 10.1161/circulationaha.120.046361. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Zhang Y, Guallar E, Malhotra S, et al. Carotid artery wall thickness and incident cardiovascular events: a comparison between US and MRI in the multi-ethnic study of atherosclerosis (MESA). Radiology. 2018;289(3):649–57. 10.1148/radiol.2018173069. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Zhou S, Hui P. Predictive value of contrast-enhanced carotid ultrasound features for stroke risk: a systematic review and meta-analysis. Front Neurol. 2025;16:1487850. 10.3389/fneur.2025.1487850. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Jashari F, Ibrahimi P, Bajraktari G, Grönlund C, Wester P, Henein MY. Carotid plaque echogenicity predicts cerebrovascular symptoms: a systematic review and meta-analysis. Eur J Neurol. 2016;23(7):1241–47. 10.1111/ene.13017. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1 (769KB, docx)

Data Availability Statement

All data generated or analyzed during this study are included in this published article [and its supplementary information files].


Articles from International Journal of Emergency Medicine are provided here courtesy of Springer-Verlag

RESOURCES