Abstract
Background
Machine learning (ML) models have gained traction for classifying carotid artery plaques via ultrasound imaging to differentiate high-risk (unstable) from low-risk (stable) plaques, a critical step for stroke risk prediction and guiding clinical interventions such as endarterectomy. However, prior studies report inconsistent diagnostic performance attributed to variations in algorithms, cohort diversity, and imaging protocols. This systematic review and meta-analysis aim to evaluate the pooled diagnostic accuracy of ML models for carotid plaque classification, addressing these inconsistencies to inform standardized clinical applications.
Methods
Five electronic databases were systematically searched up to February 28, 2025, for studies reporting diagnostic performance metrics of ML-based models in carotid plaque classification. Pooled performance metrics were analyzed using STATA, and the risk of bias was assessed using the PROBAST+AI tool.
Results
A total of 20 studies met the inclusion criteria, of which 13 provided sufficient data for quantitative synthesis. Sample sizes ranged from 15 to 413 patients, with 115– 81,000 images per study. Mean ages ranged from 27.5 to 75 years, mostly 60–70, and male representation ranged from 47% to 81%, except for one all-female cohort. The pooled sensitivity was 0.84 (95% CI: 0.74–0.90) and specificity was 0.96 (95% CI: 0.89–0.98), with a pooled AUC of 0.95 (95% CI: 0.93–0.97). Substantial heterogeneity was observed (I2 = 88.8% for sensitivity, 64.1% for specificity, and 68.1% overall). Meta-regression identified sample size and model architecture as significant sources of between-study heterogeneity. No evidence of publication bias was detected (p = 0.36). Quality assessment using PROBAST+AI indicated a low overall risk of bias in 70% of studies, moderate in 20%, and high in 10%. The GRADE approach rated the certainty of evidence as moderate, primarily due to inconsistency and study-level bias.
Conclusion
Machine learning models demonstrate promising diagnostic accuracy for carotid plaque classification, showing high pooled sensitivity and specificity. However, substantial heterogeneity and only moderate certainty of evidence suggest that these findings should be interpreted with caution.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12245-025-01065-1.
Keywords: Machine learning, Artificial intelligence, Ultrasound, Carotid artery plaque, Carotid atherosclerosis
Introduction
Stroke remains a major global health challenge, ranking as the second leading cause of death and the third leading cause of death and disability combined worldwide [1]. In 2021, it accounted for more than seven million deaths and over 150 million disability-adjusted life-years (DALYs) globally [2]. In the United States, stroke caused approximately 160,000 deaths in 2022 (about 17% of cardiovascular mortality), with incidence and mortality showing a modest upward trend over the past decade [3]. Global projections indicate that annual stroke deaths may increase by nearly 50% by 2050, driven by population aging, growth, and escalating risk factors such as obesity and hypertension [4]. These trends underscore the need for improved risk stratification and prevention strategies.
Carotid atherosclerotic plaques are a well-recognized contributor to ischemic stroke, particularly in the anterior circulation [5]. While plaques causing significant luminal stenosis are clearly implicated in stroke risk, emerging evidence indicates that non-stenotic plaques may also play a causal role [6]. The embolic potential of these plaques is thought to depend on high-risk features such as intraplaque hemorrhage, ulceration, and hypodensity [7]. Accurate identification and characterization of carotid plaque using advanced imaging modalities, including ultrasound, MRI, CTA, and PET, are therefore critical for stroke risk stratification and the development of targeted preventive strategies.
Artificial intelligence (AI), encompassing machine learning (ML) and deep learning (DL), has revolutionized medical imaging by automating feature extraction and improving diagnostic accuracy [8]. In carotid ultrasound imaging, machine learning models are used to interpret both longitudinal and transverse scans, enabling automated detection of occlusions, measurement of arterial diameter, and evaluation of plaque accumulation [9]. This approach reduces reliance on manual segmentation, minimizing human error and time requirements. Recent studies report CNN-based models achieving areas under the curve (AUC) exceeding 0.85 for detecting vulnerable plaques, surpassing traditional manual and semi-automated approaches [10, 11]. However, the lack of standardized imaging protocols, feature extraction methods, and regional differences across studies limits the generalizability and clinical translation of current findings.
This study aims to systematically evaluate the diagnostic performance of ML models for carotid plaque classification using ultrasound imaging, focusing on metrics such as sensitivity, specificity, and AUC to distinguish high-risk (unstable, symptomatic) from low-risk plaques. Through a comprehensive systematic review and meta-analysis, we seek to synthesize evidence on MLs accuracy, identify optimal algorithms, and assess generalizability across diverse cohorts, addressing standardization gaps to enhance clinical decision-making for stroke prevention.
Methods
Study design
This systematic review and meta-analysis was conducted in accordance with the PRISMA 2020 (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines [12]. The study protocol was prospectively registered on the Open Science Framework (OSF; registration ID: http://osf.io/duqsp). The PRISMA checklist is provided in the Supplementary Material.
Eligibility criteria
Inclusion criteria
Population and outcome: Studies that performed carotid plaque classification, defined as the assignment of atherosclerotic lesions into clinically meaningful categories (stable vs. unstable/vulnerable; calcified, fibrous, lipid-rich; with or without ulceration or intraplaque hemorrhage) based on imaging or histopathologic criteria.
Study design and language: Peer-reviewed publications in English, with no restriction on publication year or study design.
Definition requirement: Included studies clearly defined their plaque classification criteria or provided imaging/histopathology references. Studies lacking a clear definition were excluded or considered separately in sensitivity analyses.
Model type: Studies that employed machine learning–based prediction models applied to carotid plaque characterization.
Performance reporting: Studies that reported at least one diagnostic performance metric—accuracy, sensitivity, specificity, precision, or area under the curve (AUC)—or provided sufficient quantitative data for synthesis.
Imaging modality: All carotid ultrasound modalities were eligible, including B-mode, contrast-enhanced ultrasound (CEUS), shear-wave elastography (SWE), superb microvascular imaging (SMI), and super-resolution ultrasound. Details of the imaging modality and parameters were extracted, and subgroup analyses were performed to compare model performance across modalities.
Exclusion criteria
Language and subject: Non-English publications or non-human studies.
Model or focus: Studies that did not apply machine learning–based models or were not focused on carotid plaque classification.
Study type: Reviews, editorials, conference abstracts, letters, or case reports that did not contain original data.
Duplications: Studies using overlapping or duplicate datasets without clear differentiation were excluded to avoid duplication bias.
A summary of the inclusion criteria using the PICO framework is provided in Supplementary Table 1.
Search strategy
A comprehensive search of PubMed, Scopus, Web of Science, Embase, and ProQuest databases was conducted from inception to February 28, 2025, to identify eligible studies. To capture grey literature, the first 100 results from Google Scholar were also screened.
Search terms included combinations of keywords and MeSH terms related to “machine learning,” “artificial intelligence,” “carotid artery,” “plaque,” “atherosclerosis,” and “classification.” The whole search strategy is detailed in the Supplementary Table 2.
Study selection
Two independent reviewers performed the title and abstract screening using the Rayyan web platform, followed by full-text assessment against the predefined eligibility criteria. The reviewers were blinded to each other’s decisions during the initial screening phase to minimize selection bias. Discrepancies were resolved through consensus or by consultation with a third reviewer. Reference lists of included studies and relevant reviews were also hand-searched to identify additional eligible articles.
Data extraction
Two reviewers independently extracted data using a standardized pre-piloted form. Extracted information included:
Study characteristics: author, publication year, country, sample size, gender distribution, and dataset type.
Model characteristics: ML algorithm used, feature type, and validation approach.
Performance metrics: accuracy, sensitivity, specificity, precision, and AUC, along with raw classification counts: True positives (TP), true negatives (TN), false positives (FP), and false negatives (FN).
When performance metrics were missing, they were calculated or imputed from available raw data whenever possible. When essential data were unavailable, no contact was made with the original authors. Discrepancies in extracted data were resolved through discussion and mutual agreement between the reviewers.
Quality assessment
The PROBAST+AI tool (Prediction model Risk Of Bias Assessment Tool) was used because it is designed explicitly for prediction model studies and includes AI-specific aspects such as algorithm development, validation, and handling of complex predictors [13]. This makes it more suitable than general tools like QUADAS-2 or NOS for assessing ML studies. The certainty of evidence for pooled outcomes was evaluated using the GRADE approach. This method considers factors such as risk of bias, inconsistency, imprecision, and publication bias, providing a structured and transparent evaluation of confidence in the results. Each study was rated for risk of bias as low, moderate, or high by two reviewers, with disagreements resolved by consensus or by a senior reviewer.
Statistical analysis
All statistical analyses were performed using Stata version 18.0, utilizing the MIDAS and METADATA modules. TP, TN, FP, and FN were extracted or calculated for each model to derive pooled estimates of sensitivity, specificity, accuracy, precision, and AUC. When multiple models were reported in a single study, the best-performing model, as defined by the original authors, was selected to avoid selection bias. None of the included studies were multi-arm or reported correlated outcomes. A random-effects model with restricted maximum likelihood (REML) estimation was used to pool diagnostic performance metrics. This approach accounts for both within- and between-study variability, making it more appropriate than a fixed-effects model when substantial heterogeneity is expected across studies, as it provides more conservative and generalizable estimates. Potential sources of heterogeneity were explored using meta-regression, leave-one-out sensitivity analyses, and subgroup analyses (based on model type, study sample size, region, imaging modality, and plaque type).
For subgroup analysis, the machine learning models were categorized into two main groups:
DL models included convolutional neural network (CNN)-based architectures such as MultiNet 2.0, VGG13, Inception_v3, ResNet18, SE-ResNext-50, CNN-BiLSTM, CANet, DL-DCCP, and other deep convolutional neural networks or ensemble schemes.
U-Net-based models comprised UNet, UNet+GDL+GC, HRU-Net with transfer learning, Attention-UNet, and other U-Net–derived deep learning variants, primarily used for segmentation and detailed plaque feature extraction.
Publication bias was evaluated using Deeks’ funnel plot asymmetry test, where a p-value < 0.05 indicates publication bias.
Results
Study selection
The study selection process followed the PRISMA 2020 framework (Fig. 1). A total of 816 records were identified through database searches: PubMed (n = 124), Scopus (n = 225), Web of Science (n = 273), ProQuest (n = 6), Embase (n = 88), and Google Scholar (n = 100). After removing duplicates, 383 records remained for screening. Based on title and abstract review, 330 records were excluded for irrelevance, non-ML approaches, or non-ultrasound imaging. Fifty-three full-text articles were reviewed for eligibility, and 33 were excluded for the following reasons: studies focused only on plaque detection (n = 14), lacked risk stratification (n = 10), were abstracts or editorials (n = 2), did not use ultrasound imaging (n = 4), or were review papers (n = 3) (Supplementary table 3). Ultimately, 20 studies met the inclusion criteria and were included in the qualitative synthesis, of which 13 provided sufficient data for the quantitative meta-analysis.
Fig. 1.
Prisma diagram for study selection
Study characteristics
A total of 20 studies were analyzed, conducted across countries including Japan (n = 4), China (n = 4), the United States (n = 4), the United Kingdom (n = 3, including collaborations with Cyprus and Portugal), Japan combined with Hong Kong (n = 2), the Czech Republic (n = 1), and Greece (n = 1). Sample sizes varied significantly, ranging from 15 subjects to 413 patients; image counts ranged from 115 to 81,000 per study, and some studies reported large datasets (13,810 and 15,546 images) without patient counts. The mean age, reported in 11 studies, ranged from 27.5 ± 3.5 to 75 years, with most studies reporting mean ages between 60 and 69.9 years. Male representation, reported in 8 studies, ranged from 47% to 81%, with one Hong Kong cohort being entirely female. Validation strategies varied: 10-fold cross-validation in 7 studies, 5-fold in 5, 3-fold in 2, 4-fold in 1, and fixed splitting in 5. This reflects diverse approaches to model robustness assessment. Ultrasound imaging modalities predominantly included B-mode ultrasound (n = 15), with some studies incorporating Doppler ultrasound (n = 5) or contrast-enhanced ultrasound (CEUS; n = 1), and one study unspecified beyond ultrasound. Input characteristics for ML models included radiomics features (gray-level co-occurrence matrix [GLCM], gray-level run-length matrix [GLRLM], fractal dimension, texture features; n = 12), clinical features (hemoglobin A1c, lipid profiles, serum markers; n = 5), or a combination of clinical and radiomics features (n = 3). Radiomics methods were semi-automated in 11 studies, manual in 5 (sonographer-defined regions of interest [ROI]), and automated in 4, highlighting variability in feature-extraction workflows. Plaque classification grades varied, with 10 studies classifying plaques as stable versus unstable, four as high-risk versus low-risk, three as symptomatic versus asymptomatic, and 3 using multi-grade systems (normal/mild/moderate/severe or low/moderate/high-risk) (Table 1).
Table 1.
Studies characteristics
| First author/Year | Type of study | Country | Validation | Ultrasound imaging | Input characteristics | Method of radiomics | Sample size | Plaque grade | Mean age | Male% | Algorithms |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Biswas et al. (2024) [10] | Cohort study | Japan | 5-fold cross-validation | Ultrasound (B-mode) | Clinical features (Hemoglobin A1c, LDL cholesterol, HDL cholesterol, total cholesterol, creatinine, etc.) | Automated | 204 patients (407 images) | Low/moderate/high-risk | 69 ± 11 | 77 | MultiNet 2.0 |
| Chen et al. (2023) [11] | Cohort study | China | 5-fold cross-validation | Ultrasound (Doppler) | Clinical features (carotid Doppler ultrasound images) | Semi-automated | 568 images | Stable/unstable | NR | NR |
Inception_v3 ResNet50 |
| Singh et al. (2023) [14] | Experimental study | Cyprus, UK | 5-fold cross-validation | Ultrasound (B-mode) | Radiomics (GLCM, GLRLM, Fractal dimension) | Semi-automated | 190 patients (223 images) | High-risk/Low-risk | 54 years in Database 1, 54 years in Database 2, and 27.5 ± 3.5 years in Database 3 | 58% in Database 1 and 60% in Database 2 |
SVM ResNet-50 GoogLeNet XceptionNet Squeeze-Net VGG16 |
| Zhou et al. (2023) [15] | Experimental study | USA | Fixed splitting | Ultrasound (B-mode, Doppler) | Clinical features (ultrasound video clips of carotid plaque) | Semi-automated | 169 patients (63225 images) | Mild/severe | NR | NR |
CSG-3DCT I3D SlowFast TPN TimeSformer Vivit MTV (B/2+S/8) NL I3D UniFormer |
| Kybic et al. (2025) [16] | Experimental study | USA | 3-fold cross-validation | Ultrasound (B-mode) | Radiomic features | Semi-automated | 400 patients (6931 images) | Stable/unstable | 69 | NR | ResNet 18 |
| Jain et al. (2023) [17] | Experimental study | USA | 5-fold cross-validation | Ultrasound (B-mode) | Radiomic features | Semi-automated | 97 patients (970 images) | Stable/unstable | 75 | 47 |
UNet SegNet-UNet |
| Kigka et al. (2021) [18] | Experimental study | USA | 10-fold cross-validation | Ultrasound (B-mode) | Clinical Data - Risk Factors- Laboratory Data- Serum Markers- Imaging | Semi-automated | 208 patients | Stable/unstable | NR | NR |
SVM RF NB J48 ANN |
| Lindsey et al. (2019) [19] | Experimental study | USA | 10-fold cross-validation | Ultrasound (Doppler) | Radiomic features | Semi-automated | 13810 image Dataset 1/15546 image Dataset 2 | Normal/mild/moderate/severe | NR | NR |
ResNet-50 BN-Inception SE-ResNext-50 SE-ResNext-50 BN-Inception ResNet-50 |
| Azzopardi et al. (2020) [20] | Experimental study | UK | Fixed splitting | Ultrasound (B-mode) | Radiomic features | Semi-automated | 15 subjects (81000 images) | Stable/unstable | NR | NR |
UNET+GDL+GC UNET+CE UNET+GDL UNET+CE+GC |
| Guang et al. (2021) [21] | cohort study | China | 3-fold cross-validation | Ultrasound (B-mode, CEUS image) | Radiomic features | Automated | 205 patients | Stable/unstable | 61.6±8.4 | 81 |
DL-DCCP Xception (Deep CNN) RA-CEUS |
| Yoshidomi et al. (2024) [22] | Experimental study | Japan | 5-fold cross-validation | Ultrasound (B-mode) | Radiomics (Plaque movement and surface images) | Manual (Sonographer manually sets ROI) | 200 patients | High risk/low risk | NR | NR |
CNN-BiLSTM CNN-FC CNN-LSTM |
| Kybic et al. (2024) [23] | Experimental study | Czech | Fixed splitting | Ultrasound (transversal images) | Image-based (radiomics) | Semi-automated | 413 total patients | Stable/unstable | 69 | NR | CNN |
| Jain et al. (2021) [24] | Cohort study | Japan, Hong Kong | 10-fold cross-validation | Ultrasound (B-mode) | Radiomics | Semi-automated |
Japanese Cohort: 165 patients (330 images) Hong Kong Cohort: 50 patients (300 images) |
Stable/unstable | 68.25 years in the Japanese cohort and 60.2 years in the Hong Kong cohort | 77% in the Japanese cohort and all female in the Hong Kong cohort | UNet |
| Yuan et al. (2022) [25] | Experimental study | China | 10-fold cross-validation | Ultrasound (2D longitudinal) | Radiomics | Manual | 90 patients (115 images) | Moderate/severe | 60 | NR |
HRU-Net with transfer learning U-Net FCN Attention U-Net DeepLabv3 M-Net GCN LinkNET |
| Jain et al. (2022) [26] | Experimental study | United Kingdom (DB1), Japan (DB2), Hong Kong (DB3) | 5-fold cross-validation | Ultrasound (B-mode) | Radiomics | Semi-automated | DB1 (UK): 99 patients, DB2 (Japan): 190 patients, DB3 (Hong Kong): 50 patients, 679 images | High-risk/Low-risk | 68.78 | 62 |
Attention-UNet UNet UNet++ UNet3P Fractal-UNet Squeeze-UNet |
| Liu et al. (2024) [27] | Experimental study | China | Fixed splitting | Ultrasound | Clinical features and radiomics | Manual | 491 patients (512 images) | Stable/unstable | 62.7 | 49 |
CANet U-Net U-Net++ DeepLabV3+ FPN |
| K. Jain et al. (2022) [28] | Experimental study | Japan | 10-fold cross-validation | Ultrasound (B-mode) | Radiomics (texture features from plaque images) | Manual | 190 patients (379 images) | Stable/unstable | 68.78 ± 10.88 | 77.4 |
UNet SegNet-UNet |
| Saba et al. (2020) [29] | Cohort study | UK | 10-fold cross-validation | Ultrasound (Doppler) | Radiomics (texture features from plaque images) | Automated | 346 patients (2311 images) | Symptomatic/asymptomatic | 69.9 ± 7.8 | 61 | Deep Learning (CNN) |
| Saba et al. (2021) [30] | Experimental study | UK, Portugal | 10-fold cross-validation | Ultrasound (Doppler) | Radiomics (texture features from plaque images) | Manual | 506 patients (1260 images) | Symptomatic/asymptomatic | 69.9 ± 7.8 years in the London cohort and 67.5 ± 0.77 years in the Portugal cohort. | 61% in the London cohort, and gender distribution was not specified for the Portugal cohort. | Deep Learning (DCNN) |
| Ganitidis et al. (2021) [31] | Experimental study | Greece | 4-fold cross-validation | Ultrasound (B-mode) | Radiomics (texture features from plaque images) | Automated | 53 patients | Symptomatic/asymptomatic | NR | NR | Simple averaging ensemble scheme |
List of abbreviations: MultiNet 2.0: MultiNet 2.0 (Deep Learning-based model); VGG13: VGG13 architecture; Inception_v3: Inception version 3; ResNet50: Residual Network 50; SVM: Support Vector Machine; GoogLeNet: GoogLeNet architecture; XceptionNet: Xception Network; Squeeze-Net: SqueezeNet architecture; VGG16: VGG16 architecture; CSG-3DCT: CSG-3DCT model; I3D: Inflated 3D ConvNet; SlowFast: SlowFast network; TPN: Temporal Pyramid Network; TimeSformer: TimeSformer model; Vivit: Vision Transformer; MTV (B/2+S/8): Mixed Temporal Vision (B/2+S/8); NL I3D: Non-Local Inflated 3D ConvNet; UniFormer: UniFormer architecture; ResNet18: Residual Network 18; UNet: U-Net architecture; SegNet-UNet: SegNet-U-Net hybrid; RF: Random Forest; NB: Naive Bayes; J48: J48 Decision Tree; ANN: Artificial Neural Network; SE-ResNext-50: Squeeze-and-Excitation ResNext 50; BN-Inception: Batch Normalization Inception UNET+GDL+GC: U-Net with Generalized Dice Loss and Graph Cuts; DL-DCCP: Deep Learning with Dynamic Contrast-Enhanced CT Perfusion; RA-CEUS: Radiomics with Contrast-Enhanced Ultrasound; CNN-BiLSTM: Convolutional Neural Network with Bidirectional Long Short-Term Memory; CNN-FC: Convolutional Neural Network with Fully Connected layers; CNN-LSTM: Convolutional Neural Network with Long Short-Term Memory; HRU-Net: High-Resolution U-Net with transfer learning; FCN: Fully Convolutional Network; DeepLabv3: DeepLab version 3; M-Net: M-Net architecture; GCN: Graph Convolutional Network; LinkNET: LinkNet architecture; UNet3P: U-Net 3+ architecture; CANet: Context Aggregation Network; FPN: Feature Pyramid Network; Deep Learning (DCNN): Deep Learning with various Deep Convolutional Neural Network layer; B-mode: Brightness-mode ultrasound imaging; CEUS: Contrast-Enhanced Ultrasound; CT: Computed Tomography; DL: Deep Learning; GLCM: Gray-Level Co-occurrence Matrix; GLRLM: Gray-Level Run-Length Matrix; ROI: Region of Interest; US: Ultrasound; 2D: Two-Dimensional; 3D: Three-Dimensional; Fractal Dimension: Quantitative measure describing structural complexity of image texture; Radiomics: Quantitative extraction of high-dimensional imaging features from medical images; Texture Features: Quantitative metrics derived from image intensity and spatial patterns; DB: Database; LDL: Low-Density Lipoprotein; HDL: High-Density Lipoprotein; NR: Not Reported; Cross-Validation: Statistical technique for assessing model generalizability by partitioning data into folds; Fixed Splitting: Predetermined training/validation data separation method; Patients (n): Sample size indicator
The ML models were diverse, predominantly categorized into deep learning (DL) and U-Net-based architectures, with a few incorporating traditional ML approaches. Deep learning models included convolutional neural network (CNN)-based architectures such as MultiNet 2.0, VGG13, Inception_v3, ResNet18, SE-ResNext-50, CNN-BiLSTM, CANet, and various deep convolutional neural networks (DCNNs) with differing layer configurations, alongside specialized models like DL-DCCP (Deep Learning-Based Detection and Classification of Carotid Plaques) and a simple averaging ensemble scheme with a 0.465 threshold. U-Net-based models tailored for segmentation tasks include UNet, UNet+GDL+GC (integrating global context and guided deep learning), HRU-Net with transfer learning, Attention-UNet, and UNet-based deep learning variants, highlighting their utility for precise plaque feature extraction. Additionally, support vector machines (SVMs) were used in two studies, representing traditional ML approaches, while one study employed CSG-3DCT, a less common 3D convolutional neural network.
Of the 20 studies identified, 13 provided sufficient performance data for inclusion in the meta-analysis [11, 14, 16, 17, 21–24, 26–29].
Pooled diagnostic performance
This meta-analysis included 13 studies to evaluate the diagnostic performance of machine learning and deep learning models in ultrasound-based carotid plaque risk stratification. The best-performing model from each study, as reported by the original authors, was used for analysis to ensure comparability and represent the optimal diagnostic performance of each approach. Detailed model performance metrics are presented in Table 2. The pooled sensitivity was 0.84 (95% CI: 0.74–0.90) and specificity was 0.96 (95% CI: 0.89–0.98) (Fig. 2), with a pooled AUC of 0.95 (95% CI: 0.93–0.97) (Fig. 3). Substantial heterogeneity was observed across studies (I2 = 88.8% for sensitivity, 64.1% for specificity, and 68.1% overall).
Table 2.
Performance metrics of different ml models. The first row for each study indicates the best-performing model reported by the authors
| First author/Year | Machine learning algorithms | Accuracy | Sensitivity | Specificity | Precision | F1score | AUC |
|---|---|---|---|---|---|---|---|
| Biswas et al. (2024) [10] | MultiNet 2.0 | 0.88 | NR | NR | NR | NR | 0.88 |
| Chen et al. (2023) [11] | Inception_v3 | 0.9294 | NR | NR | NR | NR | 0.915 |
| ResNet50 | 0.8941 | NR | NR | NR | NR | 0.853 | |
| Singh et al. (2023) [14] | SVM | 0.9685 | 0.9529 | 0.975 | 0.9732 | 0.975 | NR |
| ResNet-50 | 0.8986 | 0.9714 | 0.775 | 0.8804 | 0.9224 | NR | |
| GoogLeNet | 0.98 | 0.9857 | 0.9765 | 0.986 | 0.9856 | NR | |
| XceptionNet | 0.9414 | 0.963 | 0.9044 | 0.9453 | 0.9539 | NR | |
| Squeeze-Net | 0.9508 | 0.9714 | 0.9176 | 0.9518 | 0.9602 | NR | |
| VGG16 | 0.9777 | 0.9929 | 0.9529 | 0.9724 | 0.9823 | NR | |
| Zhou et al. (2023) [15] | CSG-3DCT | 0.831 | 0.824 | NR | 0.828 | 0.825 | NR |
| I3D | 0.788 | 0.808 | NR | 0.81 | 0.788 | NR | |
| SlowFast | 0.703 | 0.708 | NR | 0.703 | 0.702 | NR | |
| TPN | 0.78 | 0.785 | NR | 0.778 | 0.778 | NR | |
| TimeSformer | 0.771 | 0.759 | NR | 0.768 | 0.762 | NR | |
| Vivit | 0.703 | 0.708 | NR | 0.703 | 0.702 | NR | |
| MTV (B/2+S/8) | 0.72 | 0.734 | NR | 0.73 | 0.72 | NR | |
| NL I3D | 0.788 | 0.808 | NR | 0.81 | 0.788 | NR | |
| UniFormer | 0.754 | 0.776 | NR | 0.781 | 0.754 | NR | |
| Kybic et al. (2025) [16] | ResNet 18 | 0.48 | 0.932 | 0.098 | 0.466 | 0.621 | 0.614 |
| Jain et al. (2023) [17] | UNet | 0.9856 | 0.903 | 0.9923 | 0.9923 | NR | 0.91 |
| SegNet-UNet | 0.9844 | 0.882 | 0.9927 | 0.9069 | NR | 0.905 | |
| Kigka et al. (2021) [18] | SVM | 0.76 | NR | NR | NR | NR | 0.73 |
| RF | 0.68 | NR | NR | NR | NR | 0.74 | |
| NB | 0.58 | NR | NR | NR | NR | 0.62 | |
| J48 | 0.62 | NR | NR | NR | NR | 0.61 | |
| ANN | 0.69 | NR | NR | NR | NR | 0.72 | |
| Lindsey et al. (2019) [19] | ResNet-50 | 0.9555 | 0.9448 | NR | 0.9884 | 0.9661 | NR |
| BN-Inception | 0.8761 | 0.88 | NR | 0.89 | 0.88 | NR | |
| SE-ResNext-50 | 0.8935 | 0.89 | NR | 0.9 | 0.89 | NR | |
| SE-ResNext-50 | 0.9781 | 0.9781 | NR | 0.9782 | 0.9781 | NR | |
| BN-Inception | 0.9701 | 0.9701 | NR | 0.9704 | 0.9701 | NR | |
| ResNet-50 | 0.973 | 0.973 | NR | 0.973 | 0.973 | NR | |
| Azzopardi et al. (2020) [20] | UNET+GDL+GC | NR | 0.94 | 0.969 | NR | NR | NR |
| UNET+CE | NR | 0.937 | 0.966 | NR | NR | NR | |
| UNET+GDL | NR | 0.936 | 0.965 | NR | NR | NR | |
| UNET+CE+GC | NR | 0.939 | 0.969 | NR | NR | NR | |
| Guang et al. (2021) [21] | DL-DCCP | NR | 0.792 | 0.844 | NR | NR | 0.87 (0.82–0.92) |
| Xception (Deep CNN) | NR | 0.75 | 0.756 | NR | NR | 0.77 (0.71–0.83) | |
| RA-CEUS | NR | 0.875 | 0.444 | NR | NR | 0.66 (0.56–0.76) | |
| Yoshidomi et al. (2024) [22] | CNN-BiLSTM | 0.805 | 0.802 | NR | 0.81 | NR | NR |
| CNN-FC | 0.775 | 0.771 | NR | 0.78 | NR | NR | |
| CNN-LSTM | 0.787 | 0.785 | NR | 0.8 | NR | NR | |
| Kybic et al. (2024) [23] | CNN | 0.528 | 0.725 | 0.485 | 0.234 | 0.354 | 0.609 |
| Jain et al. (2021) [24] | UNet | 0.9901 | 0.8637 | 0.9952 | 0.8855 | 0.8668 | 0.95 |
| Yuan et al. (2022) [25] | HRU-Net with transfer learning | 0.977 | NR | NR | NR | NR | NR |
| U-Net | 0.969 | NR | NR | NR | NR | NR | |
| FCN | 0.965 | NR | NR | NR | NR | NR | |
| Attention U-Net | 0.968 | NR | NR | NR | NR | NR | |
| DeepLabv3 | 0.969 | NR | NR | NR | NR | NR | |
| M-Net | 0.968 | NR | NR | NR | NR | NR | |
| GCN | 0.965 | NR | NR | NR | NR | NR | |
| LinkNET | 0.967 | NR | NR | NR | NR | NR | |
| Jain et al. (2022) [26] | Attention-UNet | 0.9858 | 0.8686 | 0.9952 | 0.9354 | NR | 0.99 |
| UNet | 0.9858 | 0.8743 | 0.9948 | 0.9311 | NR | 0.989 | |
| UNet++ | 0.9851 | 0.8572 | 0.9953 | 0.937 | NR | 0.988 | |
| UNet3P | 0.9851 | 0.8673 | 0.9947 | 0.929 | NR | 0.988 | |
| Fractal-UNet | 0.9841 | 0.8568 | 0.9945 | 0.9261 | NR | 0.962 | |
| Squeeze-UNet | 0.9853 | 0.8629 | 0.9951 | 0.934 | NR | 0.969 | |
| Liu et al. (2024) [27] | CANet | NR | 0.961 | NR | 0.947 | NR | NR |
| U-Net | NR | 0.9055 | NR | 0.9123 | NR | NR | |
| U-Net++ | NR | 0.8749 | NR | 0.8947 | NR | NR | |
| DeepLabV3+ | NR | 0.9061 | NR | 0.9108 | NR | NR | |
| FPN | NR | 0.9269 | NR | 0.9147 | NR | NR | |
| K. Jain et al. (2022) [28] | UNet | 0.9908 | 0.886 | 0.9956 | 0.8905 | NR | 0.94 |
| SegNet-UNet | 0.9897 | 0.915 | 0.9926 | 0.8342 | NR | 0.93 | |
| Saba et al. (2020) [29] | Deep Learning (CNN) | 0.897 | NR | NR | NR | NR | 0.91 |
| Saba et al. (2021) [30] | Deep Learning | 0.8807 | 0.786 | 0.983 | NR | 0.872 | 0.88 |
| Ganitidis et al. (2021) [31] | Simple averaging ensemble | 0.725 | 0.75 | 0.7 | NR | NR | 0.73 (0.623–0.837) |
Fig. 2.
Forest plot for sensitivity and specificity
Fig. 3.
Summary receiver operating characteristic (SROC) plot
Sensitivity (influence) analysis
A leave-one-out sensitivity analysis was performed using a random-effects model with restricted maximum likelihood (REML) estimation to assess the robustness of the pooled diagnostic odds ratio (lnDOR). Excluding each study sequentially produced minimal variation in the overall effect size (pooled lnDOR = 4.63; 95% CI = 3.17–6.09; p < 0.001), indicating that no single study disproportionately influenced the overall meta-analytic estimate. This consistency supports the stability and reliability of the pooled diagnostic performance across included studies (Fig. 4) (Supplementary table 4).
Fig. 4.
Leave-one-out forest plot
Meta-regression analysis
Meta-regression identified sample size and model architecture as significant contributors to between-study heterogeneity, explaining a substantial proportion of variability (I2 = 72%). In contrast, factors such as study region showed a borderline effect, suggesting a moderate but non-significant contribution to heterogeneity. Plaque characteristics, validation strategy, and ultrasound modality did not significantly account for residual heterogeneity (Supplementary Table 5).
Subgroup analysis
Subgroup analysis by plaque type
For the plaque type subgroup, studies comparing stable versus unstable plaques demonstrated a pooled sensitivity of 0.82 (95% CI: 0.66–0.91) and specificity of 0.96 (95% CI: 0.85–0.99). Substantial heterogeneity was observed overall (I2 = 54.61%), particularly for sensitivity (I2 = 83.29%), with moderate heterogeneity for specificity (I2 = 59.28%). Meta-analysis could not be performed for high-risk vs. low-risk (n = 3) and symptomatic vs. asymptomatic (n = 2) subgroups due to the limited number of studies (Supplementary figure 1).
Subgroup analysis by imaging modality
For the imaging-based subgroup, studies using B-mode ultrasound (n = 8) reported a pooled sensitivity of 0.85 (95% CI: 0.73–0.92) and specificity of 0.96 (95% CI: 0.88–0.99). Overall heterogeneity was substantial (I2 = 65.31%), primarily driven by high variability in sensitivity (I2 = 84.72%), while heterogeneity in specificity was moderate (I2 = 54.80%). These findings suggest that differences in imaging parameters and interpretation criteria across studies may partly explain the variability in diagnostic performance (Supplementary Figure 2).
Subgroup analysis by model architecture
In the model-based subgroup, Deep Learning models (n = 5) demonstrated a pooled sensitivity of 0.79 (95% CI: 0.74–0.83) and specificity of 0.95 (95% CI: 0.79–0.99), with minimal overall heterogeneity (I2 = 0.02%), low variability in sensitivity (I2 = 13.69%), and moderate heterogeneity in specificity (I2 = 63.12%) (Supplementary figure 3).
In contrast, network-based models (n = 5) demonstrated a pooled sensitivity of 0.85 (95% CI: 0.68–0.94) and specificity of 0.98 (95% CI: 0.92–1.00). Overall heterogeneity was negligible (I2 = 0.01%), though sensitivity showed high variability (I2 = 79.58%) and specificity exhibited moderate heterogeneity (I2 = 40.03%). Meta-analysis could not be conducted for CNN (n = 2) and SVM (n = 1) subgroups due to the limited number of studies (Supplementary figure 4).
Subgroup analysis by region
In the regional subgroup analysis, studies conducted in East Asia demonstrated a pooled sensitivity of 0.84 (95% CI: 0.70–0.93) and specificity of 0.96 (95% CI: 0.79–0.99), indicating generally high diagnostic accuracy. However, substantial heterogeneity was observed (I2 = 76.13% for sensitivity and 59.63% for specificity and general heterogeneity of 48.18%) (supplementary figure 5). For studies conducted in Europe (n = 2), North America (n = 3), and multi-regional datasets (n = 3), meta-analysis was not performed due to the limited number of studies ( < 4), which precluded reliable pooled estimates.
Assessment of publication bias
Publication bias was investigated using Deeks’ funnel plot asymmetry test, which yielded a p-value of 0.36, indicating no significant asymmetry in the funnel plot (Fig. 5). This result suggests no risk of publication bias affecting the meta-analysis outcomes.
Fig. 5.
Funnel plot
Quality assessment
In the model development phase, 90% of studies (n = 18) exhibited a low risk of bias across the Participants, Predictors, Outcomes, and Analyses domains, reflecting well-defined participant selection, consistent and objective predictor measurement, standardized outcome definitions, and robust analytical frameworks. Only two studies (10%) were rated as having a moderate risk due to incomplete reporting of predictor preprocessing steps or insufficient detail on analytical procedures. No studies were deemed at high risk of bias.
During model evaluation, the majority of studies (71%, n = 15) demonstrated low risk across all domains, supported by representative evaluation datasets, consistent predictor assessment, and rigorous analytical validation methods. Nonetheless, four studies (19%) showed moderate risk and two (10%) high risk, primarily due to limited transparency in validation protocols or potential overlap between training and evaluation datasets, which may have introduced bias in performance estimation.
Regarding applicability to carotid plaque classification, all studies (100%, n = 20) were judged to have low risk across all domains, indicating strong relevance and consistency with clinical practice. Overall, 70% of studies were classified as low risk, 20% as moderate, and 10% as high, underscoring generally sound methodological quality but highlighting residual concerns related to evaluation transparency and external generalizability (Supplementary figure 6).
Grade assessment
Evidence started as high certainty, reflecting observational diagnostic data. Risk of bias was rated low in 70% of studies per PROBAST+AI, supporting high initial certainty, though 20% moderate and 10% high ROB indicated some methodological concerns. Inconsistency was severe, with high I2 values, leading to a one-level downgrade due to substantial heterogeneity. Indirectness was minimal, as studies aligned with the review’s ultrasound-based focus. Imprecision varied: narrow confidence intervals for AUC and Specificity maintained certainty, while wider intervals for Sensitivity and likelihood ratios prompted a one-level downgrade. Overall certainty was rated moderate due to inconsistency and bias.
Discussion
Overall, ML models showed promising diagnostic performance for classifying carotid plaques from ultrasound images, although substantial variability was observed across studies. Meta-regression indicated that differences in sample size and model architecture contributed to much of this heterogeneity, primarily due to inconsistency and methodological differences. Moving forward, it is essential to compare different models to identify their respective strengths and weaknesses and explore strategies for integrating these approaches into clinical decision-making.
MRI vs ultrasound
MRI is a powerful tool for assessing carotid plaque vulnerability, visualizing key features such as intraplaque hemorrhage (IPH), lipid-rich necrotic core (LRNC), fibrous cap (FC) status, inflammation, and neovascularization, with high histologic validation [32]. IPH appears hyperintense on T1-weighted images, while LRNC is T2-hypointense and non-enhancing on contrast-enhanced scans; plaques with > 40% LRNC and thin FC are prone to rupture and major events [33]. MRI also tracks treatment effects like LRNC reduction with statins, and analyses show IPH is a stronger predictor of stroke than stenosis or clinical factors [32, 33]. The MATCH sequence shortens scan time to 7 minutes (vs. 40 minutes) with high sensitivity/specificity for IPH (≥89%/≥91%) and LRNC (≥81%/85%), though moderate agreement for calcification and fibrous tissue [34].
Ultrasound, while cost-effective and real-time for stenosis and neovascularization detection, is operator-dependent and limited in visualizing plaque characteristics [35, 36]. It measures smaller plaque areas (≈1.4-fold less than MRI) but excels in hemodynamic evaluation. MRI offers superior tissue characterization and stroke prediction, making ultrasound best for screening and MRI for detailed risk assessment. Emerging ultrasound strain imaging may help differentiate plaque composition, but lacks MRIs full validation [33].
Comparison of different ML models
Among the included studies, Kybic et al. (2025) [16] and Kybic et al. (2024) [23] exhibited lower performance (AUC 0.61 and 0.609, respectively) compared to the robust outcomes of other studies, potentially due to methodological limitations that hindered their efficacy. In Kybic et al. (2025), the use of a regression-based approach to predict plaque width (d) and increase (D) diverged from the classification focus of most studies, relying on a ResNet-18 network with a modest Dice score improvement (0.790 to 0.794) via self-supervised augmentation consistency (SAC). This modest gain suggests insufficient leverage of the large unlabeled dataset (6781 images), possibly due to the simplistic augmentation strategy (e.g., flips, rotations) and a Dice loss that, despite macro-averaging, may not adequately prioritize plaque-specific features over background classes. The regression network’s reliance on pre-segmented inputs, combined with channel dropout (probability 0.5), risks losing critical plaque texture data. At the same time, the small validation set (29 images) likely led to overfitting, undermining generalizability [16]. Similarly, Kybic et al. (2024) employed a U-Net with ResNet-34 and GANs for segmentation, followed by a shallow CNN for multicriterion regression of plaque attributes (e.g., width, echogenicity). The use of weak annotations (ellipsoidal masks) for 1399 images, while innovative, introduced noise from incomplete pixel labeling, and the GANs discriminator training instability (e.g., toggling at 40–60% accuracy) may have produced inconsistent segmentations. The regression network’s small architecture and heavy reliance on segmented masks—vulnerable to upstream errors—coupled with a limited training set (568 images for width increase), likely contributed to its poor predictive power (AUC 0.609). These methodological choices—regression over classification, suboptimal use of large datasets, and unstable training strategies—contrast with the classification-focused, well-validated deep learning models (e.g., GoogLeNet, UNet variants) in other studies, explaining their outsized contribution to heterogeneity (I2 > 90%) and the downward pull on pooled estimates like Accuracy (0.854) [23].
ML models vs conventional methods
Traditional carotid plaque assessment relied primarily on manual ultrasound interpretation and histological analysis, with quantitative metrics reported inconsistently across studies. Rafailidis et al. quantified plaque surface irregularity using the Surface Irregularity Index (SII) via color Doppler ultrasound (CDUS) and CEUS, demonstrating high interobserver agreement (ICC 0.979 for CDUS, 0.952 for CEUS), yet their classification of plaque surfaces showed no significant correlation with stroke occurrence and lacked sensitivity or specificity metrics [37]. Baldassarre et al. measured carotid intima-media thickness (cIMT) and internal carotid artery diameter (ICCAD), reporting hazard ratios of 1.27–1.47 per SD increase and a net reclassification improvement (NRI) of up to 20.1%, but did not provide diagnostic accuracy measures, reflecting the subjective nature of visual cIMT assessment [38]. Sonaglioni et al. found thicker cIMT in idiopathic pulmonary fibrosis patients (1.5 ± 0.3 mm vs. 1.1 ± 0.3 mm, p < 0.0001) with moderate intra-observer agreement (ICC 0.79–0.94), yet qualitative echogenicity assessment lacked quantitative accuracy metrics. CEUS improved diagnostic performance by visualizing plaque neovascularity [39]; Luo et al. reported Sensitivity 83–87.1%, Specificity 64.7–81.5%, and AUC 0.729–0.911 for plaque characterization, although manual plaque selection and modest repeatability limited scalability [40]. Histological validation demonstrated associations of ADAMTS4 and versican expression with plaque vulnerability (OR 1.144, 95% CI 1.007–1.300), but real-time application and diagnostic metrics were absent [41].
Non-machine learning meta-analyses on carotid imaging provide critical insights into stroke risk stratification using ultrasound, yet reveal limitations in prognostic accuracy and consistency. In 13 studies with 3,092 participants, ischemic stroke patients had higher plaque enhancement intensity (SMD = 0.71, 95% CI: 0.32–1.11) and a greater likelihood of plaque enhancement (OR = 3.25, 95% CI: 1.86–5.68), with pooled diagnostic sensitivity and specificity of 0.68 and 0.61 [42]. Increased carotid intima-media thickness (cIMT) per SD was associated with higher risks of coronary heart disease (HR = 1.10), stroke (HR = 1.08), and CVD (HR = 1.14), while MRI-derived wall thickness showed even stronger associations (HR = 1.27–1.58) and remained significant when modeled with cIMT [43]. Across 119 RCTs (100,667 patients), each 10 μm/year reduction in cIMT progression corresponded to a 9% lower CVD risk (RR = 0.91, 95% CrI: 0.87–0.94), consistent across interventions [44]. In 11 studies, echolucent carotid plaques predicted future cardiovascular events in asymptomatic (RR = 2.72) and recurrent symptoms in symptomatic patients (RR = 2.97), especially with severe stenosis and modern ultrasound [45]. However, analysis of 41 trials (18,307 participants) found no link between cIMT regression and outcomes (CHD OR = 0.82; stroke OR = 0.71), indicating cIMT regression may not reliably predict clinical endpoints [46]. Unlike traditional approaches, which often rely on subjective visual assessment or limited quantitative metrics, ML models can automatically extract high-dimensional features from large datasets, capturing subtle patterns associated with plaque vulnerability. This capability allows ML algorithms to generate personalized risk scores for individual patients, improve early detection of high-risk plaques, and potentially outperform conventional methods.
Clinical applicability
ML models represent a paradigm shift in carotid plaque classification, offering straightforward clinical utility when integrated appropriately. This automation enhances diagnostic accuracy and consistency in risk stratification. At the same time, there is a need for analysis of large-scale datasets, external validation, standardized datasets across diverse patient populations and imaging modalities, to improve generalizability. For high-risk patient identification, ML algorithms can combine multimodal data, including clinical variables and imaging features, to generate personalized vulnerability scores, which could guide thresholds for intervention such as endarterectomy or intensified medical therapy and facilitate early detection of unstable plaques prone to rupture. To support clinical adoption, these models should be validated on representative cohorts, with performance benchmarks guiding integration into the workflow. Additionally, the scalability of ML can streamline resource allocation, reduce radiologist workload through automated triage, and optimize follow-up strategies, fostering cost-effective, evidence-based care in vascular medicine. On the other hand, the reliance of ML models on large datasets and advanced computational resources may limit their practical implementation in routine clinical settings.
Limitations
This meta-analysis has several limitations. Substantial heterogeneity was observed, particularly for sensitivity and specificity, suggesting that pooled estimates may not directly translate to broader clinical practice. Most studies originated from Southeast Asia, creating geographic clustering that may limit applicability to other regions. Possible dataset overlap across studies could bias pooled estimates and overestimate model performance. Many studies did not validate ultrasound-based plaque classification against histopathology or longitudinal clinical outcomes, making it difficult to assess true diagnostic or prognostic value. Criteria for plaque classification varied across studies, reflecting a lack of standardization in defining plaque stability or vulnerability. Direct comparisons between machine learning models and conventional diagnostic methods were generally lacking, preventing head-to-head evaluation.
Future direction
Future research should prioritize model interpretability, integration of multimodal data, and validation across diverse populations. Emphasis should be placed on explainable AI to build clinician trust, combining imaging with genomic, proteomic, and clinical data for personalized risk assessment. Large, multi-center studies are needed to ensure generalizability and address biases, while real-time deployment in point-of-care ultrasound could improve efficiency and guide targeted interventions.
Conclusion
This meta-analysis suggests that machine learning models may have potential for carotid plaque risk classification, with generally high sensitivity, specificity, and overall accuracy. The overall low risk of bias, as assessed by PROBAST+AI, provides some support for these findings, but substantial heterogeneity and variability across studies limit confidence in the pooled estimates. Differences in sample size, model architecture, and methodology likely contribute to this variability. Further well-designed, standardized studies with external validation are needed to clarify the reliability and generalizability of ML-based diagnostic approaches.
Electronic supplementary material
Below is the link to the electronic supplementary material.
Abbreviations
- ANN
Artificial Neural Network
- B-mode
Brightness-mode ultrasound imaging
- CEUS
Contrast-Enhanced Ultrasound
- CNN
Convolutional Neural Network
- DL
Deep Learning
- DCNN
Deep Convolutional Neural Network
- FCN
Fully Convolutional Network
- GLCM
Gray-Level Co-occurrence Matrix
- GLRLM
Gray-Level Run-Length Matrix
- HRU-Net
High-Resolution U-Net with transfer learning
- Inception_v3
Inception version 3
- ML
Machine Learning
- ResNet18
Residual Network 18
- ResNet50
Residual Network 50
- RF
Random Forest
- ROI
Region of Interest
- SVM
Support Vector Machine
- UNet
U-Net architecture
- US
Ultrasound
- VGG16
Visual Geometry Group 16 architecture
- XceptionNet
Xception Network
Author contributions
P.E.1 was a primary contributor in the design, implementation, and writing of the manuscript. M.R., H.S., and P.E.2 independently assessed articles and extracted data. All authors read and approved the final manuscript. M.R. and J.T. performed statistical analysis.
Funding
None.
Data availability
All data generated or analyzed during this study are included in this published article [and its supplementary information files].
Declarations
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Generative AI
In the preparation of this article, the authors utilized the Grammarly application to enhance linguistic accuracy and clarity. The manuscript underwent meticulous double-checking to ensure precision, and the authors assume full responsibility for the integrity and originality of the content presented herein.
Randomized controlled trial number
Not applicable.
Competing interests
TThe authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Strilciuc S, Grad DA, Radu C, et al. The economic burden of stroke: a systematic review of cost of illness studies. J Med Life. 2021;14(5):606–19. 10.25122/jml-2021-0361. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Global, regional, and national burden of stroke and its risk factors, 1990-2021: a systematic analysis for the global burden of disease study 2021. Lancet Neurol. 2024;23(10):973–1003. 10.1016/s1474-4422(24)00369-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Martin SS, Aday AW, Allen NB, et al. Heart disease and stroke statistics: a report of US and global data from the American Heart Association. Circulation. 2025 2025;151(8):e41–660. 10.1161/cir.0000000000001303. [DOI] [PMC free article] [PubMed]
- 4.Cheng Y, Lin Y, Shi H, et al. Projections of the stroke burden at the global, regional, and national levels up to 2050 based on the global burden of disease study 2021. J Am Heart Assoc. 2024;13(23):e036142. 10.1161/jaha.124.036142. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Diener HC, Easton JD, Hart RG, Kasner S, Kamel H, Ntaios G. Review and update of the concept of embolic stroke of undetermined source. Nat Rev Neurol. 2022;18(8):455–65. 10.1038/s41582-022-00663-4. [DOI] [PubMed] [Google Scholar]
- 6.Kamtchum-Tatuene J, Wilman A, Saqqur M, Shuaib A, Jickling GC. Carotid plaque with high-risk features in embolic stroke of undetermined source: systematic review and meta-analysis. Stroke. 2020;51(1):311–14. 10.1161/strokeaha.119.027272. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Ospel JM, Marko M, Singh N, Goyal M, Almekhlafi MA. Prevalence of non-stenotic (<50%) carotid plaques in acute ischemic stroke and transient ischemic attack: a systematic review and meta-analysis. J Stroke Cerebrovasc Dis. 2020;29(10):105117. 10.1016/j.jstrokecerebrovasdis.2020.105117. [DOI] [PubMed] [Google Scholar]
- 8.Sajjadi SM, Mohebbi A, Ehsani A, et al. Identifying abdominal aortic aneurysm size and presence using natural language processing of radiology reports: a systematic review and meta-analysis. Abdominal Radiol. 2025. 10.1007/s00261-025-04810-5. [DOI] [PubMed] [Google Scholar]
- 9.Eini P, Eini P, Serpoush H, Rezayee M, Tremblay J. Machine learning models for carotid artery plaque detection: a systematic review of ultrasound-based diagnostic performance. J Stroke Cerebrovascular Dis. 2025;34(11):108446. 10.1016/j.jstrokecerebrovasdis.2025.108446. [DOI] [PubMed] [Google Scholar]
- 10.Biswas M, Saba L, Kalra M, et al. MultiNet 2.0: a lightweight attention-based deep learning network for stenosis measurement in carotid ultrasound scans and cardiovascular risk assessment. Computerized Med Imag Graphics. 2024;117:102437. 10.1016/j.compmedimag.2024.102437. [DOI] [PubMed] [Google Scholar]
- 11.Chen X-X, Kong Z-X, Wei S-F, et al. Ultrasound lmaging-vulnerable plaque diagnostics: automatic carotid plaque segmentation based on deep learning. J Radiat Res Appl Sci. 2023;16(3):100598. 10.1016/j.jrras.2023.100598. [Google Scholar]
- 12.Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. 10.1136/bmj.n71. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Moons KGM, Damen JAA, Kaul T, et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ. 2025;388:e082505. 10.1136/bmj-2024-082505. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Singh S, Jain PK, Sharma N, Pohit M, Roy S. Atherosclerotic plaque classification in carotid ultrasound images using machine learning and explainable deep learning. Intell Med. 2024;4(2):83–95. 10.1016/j.imed.2023.05.003. [Google Scholar]
- 15.Zhou X, Huang Y, Xue W, et al. Inflated 3D convolution-transformer for weakly-supervised carotid stenosis grading with ultrasound videos. In: Greenspan H, Madabhushi A, Mousavi P, et al., editors. Medical image computing and computer assisted intervention – MICCAI 2023. Cham: Springer Nature Switzerland; 2023. p. 511–20.
- 16.Kybic J, Pakizer D, Kozel J, Michalčová P, Charvát F, Školoudík D. Atherosclerotic plaque stability prediction from longitudinal ultrasound images. In: Xu X, Cui Z, Rekik I, Ouyang X, Sun K, editors. Machine learning in medical imaging. Cham: Springer Nature Switzerland; 2025. p. 124–32. [Google Scholar]
- 17.Jain PK, Sharma N, Roy S. Hybrid deep learning models for segmentation of atherosclerotic plaque in B-mode carotid ultrasound image. In: Sharma S, Subudhi B, Sahu U, editors. Intelligent control, robotics, and industrial automation. Singapore: Springer Nature Singapore; 2023. p. 807–19. [Google Scholar]
- 18.Kigka VI, Sakellarios AI, Mantzaris MD, et al. A machine learning model for the identification of high-risk carotid atherosclerotic plaques. Annu Int Conf IEEE Eng Med Biol Soc. 2021;2021:2266–69. 10.1109/embc46164.2021.9630654. [DOI] [PubMed] [Google Scholar]
- 19.Lindsey T, Garami Z. Automated stenosis classification of carotid artery sonography using deep neural networks. 2019 18th IEEE International Conference on Machine Learning and Applications (ICMLA). 2019:1880–84.
- 20.Azzopardi C, Camilleri KP, Hicks YA. Bimodal automated carotid ultrasound segmentation using geometrically constrained deep neural networks. IEEE J Biomed Health Inf. 2020;24(4):1004–15. 10.1109/jbhi.2020.2965088. [DOI] [PubMed] [Google Scholar]
- 21.Guang Y, He W, Ning B, et al. Deep learning-based carotid plaque vulnerability classification with multicentre contrast-enhanced ultrasound video: a comparative diagnostic study. BMJ Open. 2021;11(8):e047528. 10.1136/bmjopen-2020-047528. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Yoshidomi T, Kume S, Aizawa H, Furui A. Classification of carotid plaque with jellyfish sign through convolutional and recurrent neural networks utilizing plaque surface edges. 2024 46th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC):1–42024. [DOI] [PubMed]
- 23.Kybic J, Pakizer D, Kozel J, Michalčová P, Charvát F, Školoudík D. Predikce stability aterosklerotického plátu z transverzálních ultrazvukových obrazů pomocí hlubokého učení. Čes Slov Neurol Neurochir. 2024;87(4).
- 24.Jain PK, Sharma N, Saba L, et al. Unseen artificial intelligence—deep learning paradigm for segmentation of low atherosclerotic plaque in carotid ultrasound: a multicenter cardiovascular study. Diagnostics. [DOI] [PMC free article] [PubMed]
- 25.Yuan Y, Li C, Zhang K, Hua Y, Zhang J. HRU-Net: a transfer learning method for carotid artery plaque segmentation in ultrasound images. Diagnostics. [DOI] [PMC free article] [PubMed]
- 26.Jain PK, Dubey A, Saba L, et al. Attention-based UNet deep learning model for plaque segmentation in carotid ultrasound for stroke risk stratification: an artificial intelligence paradigm. J Cardiovasc Dev Dis. [DOI] [PMC free article] [PubMed]
- 27.Liu M, Gao W, Song D, et al. A deep learning-based calculation system for plaque stenosis severity on common carotid artery of ultrasound images. Vascular. 2024;17085381241246312. 10.1177/17085381241246312. [DOI] [PubMed]
- 28.Jain PK, Sharma N, Saba L, et al. Automated deep learning-based paradigm for high-risk plaque detection in B-mode common carotid ultrasound scans: an asymptomatic Japanese cohort study. Int Angiol. 2022;41(1):9–23. 10.23736/s0392-9590.21.04771-4. [DOI] [PubMed]
- 29.Saba L, Sanagala SS, Gupta SK, et al. Ultrasound-based internal carotid artery plaque characterization using deep learning paradigm on a supercomputer: a cardiovascular disease/stroke risk assessment system. Int J Cardiovasc Imag. 2021;37(5):1511–28. 10.1007/s10554-020-02124-9. [DOI] [PubMed] [Google Scholar]
- 30.Saba L, Sanagala SS, Gupta SK, et al. A multicenter study on carotid ultrasound plaque tissue characterization and classification using six deep artificial intelligence models: a stroke application. IEEE Trans Instrum Meas. 2021;70:1–12. 10.1109/TIM.2021.3052577.33776080 [Google Scholar]
- 31.Ganitidis T, Athanasiou M, Dalakleidi K, Melanitis N, Golemati S, Nikita KS. Stratification of carotid atheromatous plaque using interpretable deep learning methods on B-mode ultrasound images. 2021 43rd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC):3902–052021. [DOI] [PubMed]
- 32.Kassem M, Florea A, Mottaghy FM, van Oostenbrugge R, Kooi ME. Magnetic resonance imaging of carotid plaques: current status and clinical perspectives. Ann Transl Med. 2020;8(19):1266. 10.21037/atm-2020-cass-16. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Porambo ME, DeMarco JK. Mr imaging of vulnerable carotid plaque. Cardiovasc Diagn Ther. 2020;10(4):1019–31. 10.21037/cdt.2020.03.12. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Kassem M, Nies KPH, Boswijk E, et al. Quantification of carotid plaque composition with a multi-contrast atherosclerosis characterization (match) MRI sequence. Front Cardiovasc Med. 2023;10:1227495. 10.3389/fcvm.2023.1227495. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Hou C, Li S, Zheng S, et al. Quality assessment of radiomics models in carotid plaque: a systematic review. Quant Imag Med Surg. 2023;14(1):1141–54. https://qims.amegroups.org/article/view/119159. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Hou C, Liu X-Y, Du Y, et al. Radiomics in carotid plaque: a systematic review and radiomics quality score assessment. Ultrasound Med Biol. 2023;49(12):2437–45. 10.1016/j.ultrasmedbio.2023.06.008. [DOI] [PubMed] [Google Scholar]
- 37.Rafailidis V, Chryssogonidis I, Grisan E, et al. Does quantification of carotid plaque surface irregularities better detect symptomatic plaques compared to the subjective classification? J Ultrasound Med. 2019;38(12):3163–71. 10.1002/jum.15017. [DOI] [PubMed] [Google Scholar]
- 38.Baldassarre D, Hamsten A, Veglia F, et al. Measurements of carotid intima-media thickness and of interadventitia common carotid diameter improve prediction of cardiovascular events: results of the improve (carotid intima media thickness [IMT] and IMT-progression as predictors of vascular events in a high risk European population) study. J Am Coll Cardiol. 2012;60(16):1489–99. 10.1016/j.jacc.2012.06.034. [DOI] [PubMed] [Google Scholar]
- 39.Sonaglioni A, Caminati A, Lipsi R, Lombardo M, Harari S. Association between C-reactive protein and carotid plaque in mild-to-moderate idiopathic pulmonary fibrosis. Intern Emerg Med. 2021;16(6):1529–39. 10.1007/s11739-020-02607-6. [DOI] [PubMed] [Google Scholar]
- 40.Luo X, Li W, Bai Y, Du L, Wu R, Li Z. Relation between carotid vulnerable plaques and peripheral leukocyte: a case-control study of comparison utilizing multi-parametric contrast-enhanced ultrasound. BMC Med Imag. 2019;19(1):74. 10.1186/s12880-019-0374-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Dong H, Du T, Premaratne S, et al. Relationship between ADAMTS4 and carotid atherosclerotic plaque vulnerability in humans. J Vasc Surg. 2018;67(4):1120–26. 10.1016/j.jvs.2017.08.075. [DOI] [PubMed] [Google Scholar]
- 42.Costanzo P, Perrone-Filardi P, Vassallo E, et al. Does carotid intima-media thickness regression predict reduction of cardiovascular events? A meta-analysis of 41 randomized trials. J Am Coll Cardiol. 2010;56(24):2006–20. 10.1016/j.jacc.2010.05.059. [DOI] [PubMed] [Google Scholar]
- 43.Willeit P, Tschiderer L, Allara E, et al. Carotid intima-media thickness progression as surrogate marker for cardiovascular risk: meta-analysis of 119 clinical trials involving 100 667 patients. Circulation. 2020;142(7):621–42. 10.1161/circulationaha.120.046361. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Zhang Y, Guallar E, Malhotra S, et al. Carotid artery wall thickness and incident cardiovascular events: a comparison between US and MRI in the multi-ethnic study of atherosclerosis (MESA). Radiology. 2018;289(3):649–57. 10.1148/radiol.2018173069. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Zhou S, Hui P. Predictive value of contrast-enhanced carotid ultrasound features for stroke risk: a systematic review and meta-analysis. Front Neurol. 2025;16:1487850. 10.3389/fneur.2025.1487850. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Jashari F, Ibrahimi P, Bajraktari G, Grönlund C, Wester P, Henein MY. Carotid plaque echogenicity predicts cerebrovascular symptoms: a systematic review and meta-analysis. Eur J Neurol. 2016;23(7):1241–47. 10.1111/ene.13017. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All data generated or analyzed during this study are included in this published article [and its supplementary information files].





