Skip to main content
Cardiology and Therapy logoLink to Cardiology and Therapy
. 2026 May 15;15(3):393–406. doi: 10.1007/s40119-026-00453-9

Concordance of Large Language Model Recommendations with Multidisciplinary Heart Team Decisions in Coronary Revascularization and Aortic Valve Intervention: A Systematic Review and Pooled Analysis

Armaun D Rouhi 1, Shreyas V Menon 2, Yazid K Ghanem 3, Jason J Han 4,✉
PMCID: PMC13542011  PMID: 42141253

Abstract

Introduction

The multidisciplinary heart team (HT) remains the cornerstone of decision-making for complex cardiovascular disease. Large language models (LLMs) and other generative artificial intelligence models have recently emerged as potential decision support tools across diverse clinical settings. We sought to synthesize current evidence and quantitatively estimate concordance between LLM recommendations and HT decisions.

Methods

A literature search was performed using PubMed, Scopus, and Web of Science for primary studies published between November 2022 and February 2026 that evaluated recommendations by LLMs against multidisciplinary HT decisions. Studies reporting overall agreement were included for quantitative pooling. Random-effects meta-analysis was performed to determine proportion of agreement.

Results

Four retrospective concordance studies were included regarding decision-making in coronary revascularization and aortic valve intervention. LLM–HT concordance ranged from 65% to 82% for coronary revascularization and was 77% for aortic valve intervention. In random-effects meta-analysis, the pooled agreement between LLM recommendations and HT decisions was 0.73 (95% CI 0.60–0.83) with substantial heterogeneity. Discordance stemmed from LLM reliance on outdated trial evidence and limited transparency regarding utilized data, with misclassifications observed in cases of octogenarians with aortic stenosis. Detailed prompts generally improved accuracy and reliability of LLM recommendations.

Conclusion

These preliminary findings suggest LLMs may have potential as adjunctive decision support tools for multidisciplinary HTs. There remains potential for misclassification when patient-specific factors and conflicting guidelines complicate decision-making. Further prospective evaluation across diverse LLMs is essential before clinical deployment can be recommended.

Keywords: Large language models, Artificial intelligence, ChatGPT, Heart team, Cardiology, Cardiothoracic surgery, Coronary revascularization, Aortic valve intervention, Aortic stenosis, Coronary artery disease

Key Summary Points

Across four included studies, LLM recommendations showed 65–82% concordance with heart team decisions in coronary revascularization and 77% for aortic valve intervention.
Pooled LLM–heart team agreement was 0.73 (95% CI 0.60–0.83) with substantial heterogeneity (I2 = 69.0%) and low-to-moderate risk of bias.
LLMs misclassified cases of patients aged 70–80 years old with severe aortic stenosis and multiple comorbidities.
Early results are largely exploratory; however, prospective evaluation is necessary to identify potential shortcomings of current LLMs before broader implementation as heart team decision support tools.

Introduction

The multidisciplinary heart team (HT), consisting of interventional cardiology, cardiac surgery, and clinical cardiology, is a contemporary model for collaborative decision-making in the management of complex coronary and valvular disease [1]. In patients with unclear treatment strategies, the American Heart Association (AHA) and American College of Cardiology (ACC) emphasize the HT approach to improve outcomes [2]. This approach affords both individualized care and consensus in particularly critical settings, such as multivessel coronary artery disease and aortic stenosis [3]. Multiple studies have associated the HT approach with reduced in-hospital mortality and enhanced long-term survival in mitral and aortic valve disease, although challenges persist regarding team composition and optimal workflow [4–6].

Despite its widespread adoption, the HT paradigm currently lacks standardized outcome metrics and randomized data supporting its efficacy [4, 5]. As such, there exists a pressing need for tools that support the function of HTs in synthesizing different evidence, facilitating multidisciplinary collaboration, and ultimately optimizing patient management.

Large language models (LLMs) and other generative artificial intelligence (AI) have recently emerged and been utilized in diverse clinical settings, including patient education, clinical documentation, and decision support [7–9]. These novel technologies have shown the unique ability to rapidly synthesize clinical data, established guidelines, and the current literature, potentially offering HTs an additional tool for evidence-based decision-making [7, 8]. The present review aims to describe current applications of LLMs for clinical decision support within HTs, appraise available evidence, and quantitatively estimate concordance between LLM recommendations and HT decisions.

Methods

Search Strategy

We conducted a comprehensive literature search using PubMed, Scopus, and Web of Science databases for English language articles published between November 2022 and February 2026. These dates were chosen to reflect the time period in between when LLMs were first introduced to the public and the time of this article’s writing to capture the most current literature. The following search terms were utilized: “Generative Artificial Intelligence”, “Generative AI”, “Large Language Model”, “ChatGPT”, “Artificial Intelligence”, “Heart Team,” “Multidisciplinary Team”, “Heart”, “Cardiac”, and “Cardiology”.

This article is based on previously conducted studies and does not contain any new studies with human participants or animals performed by any of the authors.

Eligibility Criteria

We included original studies that (1) evaluated clinical decision-making of at least one LLM in a cardiovascular pathology using structure clinical inputs or vignettes; (2) compared LLM recommendations against a multidisciplinary HT reference decision; and (3) reported percentage agreement or accuracy. Studies were excluded if they did not directly identify the specific LLMs used or lacked extractable concordance metrics necessary for quantitative pooling.

Data Extraction

Cardiovascular pathology, sample size (cases/patients), LLM version, decision structure (binary versus multiclass), HT composition, and overall agreement were extracted from each included study.

Outcomes

The primary outcome was overall agreement or accuracy, defined as the proportion of cases where the LLM recommendation and HT decision exactly matched. For management decisions that involved multiple classes (e.g., percutaneous coronary intervention [PCI] versus coronary artery bypass grafting [CABG] versus medical therapy), an exact category match was required for agreement. As the term “concordance” is used inconsistently across the literature, we predefined our outcome as exact agreement between LLM recommendations and HT decisions. This allowed quantitative pooling but necessarily collapsed richer performance information into a single metric.

Risk of Bias Assessment

We assessed the quality and risk of bias of all included studies using the Quality Assessment for Diagnostic Accuracy Studies (QUADAS-2) criteria, which examines patient or case selection, index test, reference standard, and flow and timing [10]. Quality assessment was performed independently by two reviewers (ADR and SVM), with disagreements resolved by discussion with a third author (YKG) designated as tiebreaker, although consensus was reached in all cases without adjudication.

Statistical Analysis

Agreement proportions were pooled using a random-effects meta-analysis on the logit scale to account for between-study heterogeneity. Between-study variance (τ2) was estimated using the restricted maximum likelihood (REML) estimator, and confidence intervals for the pooled estimate were computed with the Hartung–Knapp–Sidik–Jonkman (HKSJ) adjustment to account for uncertainty in τ2 given the small number of pooled studies [11, 12]. Pooled logit-scale estimates were back-transformed to proportions and reported with 95% confidence intervals (CI). As pre-specified sensitivity analyses, we re-estimated the pooled agreement using the DerSimonian–Laird and Paule–Mandel estimators and re-ran the primary model restricting each study to its oldest reported LLM [11]. Statistical heterogeneity was quantified using I2 and τ2. Pre-specified subgroup analyses compared (a) binary (CABG vs PCI) versus three-way (medical therapy vs PCI vs CABG; medical therapy vs TAVI vs SAVR) decision structures, and (b) coronary versus aortic pathologies. Statistical analyses were performed in Python 3.11.

Results

Study Selection and Characteristics

The study flow diagram is summarized in Fig. 1. Four studies contributed to the quantitative synthesis: Sudri et al., 2024 (n = 86, coronary, GPT-4), Migliaro et al., 2025 (n = 137, coronary, GPT-4o), Mola et al., 2025 (n = 128, coronary, GPT-o1), and Salihu et al., 2024 (n = 150, aortic, GPT-4), for a total of 501 patients (Table 1) [13–16].

Fig. 1.

Fig. 1

Study flow diagram

Table 1.

Included study characteristics

Study (first author, year) Cardiovascular pathology Sample size (cases) LLMs Decision structure HT composition LLM–HT agreement
Sudri, 2024 [13] CAD 86 ChatGPT-3.5; ChatGPT-4 Binary; PCI or CABG 1 interventional cardiologist; 1 cardiac surgeon ChatGPT-3.5: 67% overall; ChatGPT-4: 82% overall (ChatGPT-4 pooled; ChatGPT-3.5 in sensitivity)
Migliaro, 2025 [14] CAD 137 ChatGPT-4o Binary; PCI or CABG Unspecified number of interventional cardiologists, imaging specialists, cardiac surgeons; 1 vascular surgeon; 1 anesthesiologist; 1 geriatrician 65% overall; 82.4% (CABG), 44.4% (PCI) (ChatGPT-4o pooled)
Mola, 2025 [15] CAD 128 ChatGPT-4o; ChatGPT-o1 Multiclass; PCI or CABG or medical therapy ≥ 2 cardiac surgeons; ≥ 2 cardiologists

ChatGPT-4o: 43% (CABG), 68.7% (PCI), 0% (medical therapy); 42.2% overall (derived, 54/128)

ChatGPT-o1: 82% (CABG), 43.7% (PCI), 0% (medical therapy); 69.5% overall (derived, 89/128; ChatGPT-o1 pooled, ChatGPT-4o in sensitivity)

Salihu, 2024 [16] Severe aortic stenosis 150 ChatGPT-4 Multiclass; TAVI or SAVR or medical therapy Unspecified number of interventional cardiologists, imaging specialists, cardiac surgeons; 1 vascular surgeon; 1 anesthesiologist; 1 geriatrician 77% overall; 90% (TAVI), 65% (SAVR), 65% (medical therapy) (ChatGPT-4 pooled)

Mola et al. did not report an overall concordance proportion. Exact agreement was derived by class-weighting the reported per-class sensitivities against HT class counts (100 CABG, 16 PCI, 12 medical therapy). “Pooled” denotes the model carried forward to the primary random-effects meta-analysis. Alternative models were used in the oldest LLM sensitivity analysis

CABG coronary artery bypass grafting, CAD coronary artery disease, HT heart team, LLM large language model, PCI percutaneous coronary intervention, SAVR surgical aortic valve replacement, TAVI transcatheter aortic valve intervention

Coronary Revascularization Decision Support

Three recent studies compared AI-generated recommendations with HT decisions for PCI versus CABG, including cases with single and multivessel coronary artery disease (CAD). Sudri et al. analyzed 86 cases that included patients with triple vessel, left main CAD [13]. When given structured and comprehensive clinical and angiographic data, ChatGPT-4 achieved an overall agreement rate of approximately 82% with HT decisions. Performance was even stronger in high-risk groups, as agreement exceeded 90% for patients with concomitant left main CAD, triple vessel disease, and diabetes, highlighting that ChatGPT-4 aligned closely with HT clinical judgment in situations that were strongly guided by established revascularization pathways [2]. The authors also found that providing LLMs with detailed patient data was essential for optimal reliability while less structured inputs resulted in lower concordance with HT decisions.

Migliaro et al. examined 137 patients with multivessel CAD and compared ChatGPT-4o recommendations to HT decisions [14]. In this study, ChatGPT-4o demonstrated 65% concordance with HT management decisions. However, the level of agreement varied significantly depending on the chosen treatment by the HT. When the HT recommended CABG, ChatGPT-4o agreed in 82% of cases, demonstrating high concordance in scenarios where the disease burden was substantial and guidelines more clearly favored surgical revascularization. In contrast, when the HT selected PCI, ChatGPT-4o agreed in only 44% of cases. In cases with discordant recommendations between ChatGPT-4o and HT, the AI more often tended to recommend surgery. These instances of disagreement occurred in 48 (35%) cases and often reflected situations where clinical nuance beyond anatomic scoring systems influenced decision-making, such as patient comorbidities, operative risk profiles, or subtleties in lesion morphology and coronary physiology.

Mola et al. evaluated 128 cases for coronary revascularization and compared recommendations from ChatGPT-o1 and ChatGPT-4o with multidisciplinary HT decisions [15]. The HT in this study determined CABG in 78.1% of cases, PCI in 12.5%, and medical therapy in 9.4%. Compared to HT decisions, ChatGPT-o1 agreed with 82% of CABG recommendations, while ChatGPT-4o only demonstrated 43% agreement. ChatGPT-4o, however, showed greater agreement with the HT in 68.7% of PCI cases, ahead of ChatGPT-o1 at 43.7% agreement, suggesting that ChatGPT-4o was more likely to favor PCI in borderline or complex cases. Notably, both LLMs failed to correctly identify any patients assigned by the HT to medical therapy (0%), highlighting decision support limitations in cases where operative risk and broader clinical context outweighed indications for revascularization. The authors attributed this discordance to the present inability of LLMs to directly interpret imaging while integrating nuanced patient characteristics, such as vessel quality and perioperative risk. Because Mola et al. did not report an overall concordance proportion, exact agreement was derived by class-weighting the reported per-class sensitivities against HT class counts: ChatGPT-o1 concordance = (0.82 × 100 CABG) + (0.437 × 16 PCI) + (0 × 12 medical) = 89/128 = 0.695, and ChatGPT-4o concordance = (0.43 × 100) + (0.687 × 16) + (0 × 12) = 54/128 = 0.422. The most advanced model (ChatGPT-o1) was carried forward to the primary pooled analysis, with ChatGPT-4o used in the oldest-LLM sensitivity analysis.

Aortic Valve Intervention Decision Support

Salihu et al. evaluated how ChatGPT-4 could support HT decision-making in severe aortic stenosis, with each of 150 patients having three possible treatment options: TAVI, SAVR, or medical therapy [16]. Researchers created standardized vignettes that included 14 key variables derived from clinical, echocardiographic, cardiovascular imaging, and surgical risk assessments. These variables included patient age, New York Heart Association functional class, frailty evaluation, left ventricular ejection fraction, aortic valve area, mean valve gradient, coronary angiography findings, CT-based aortic calcification score, vascular access characteristics, carotid artery assessment, and Society of Thoracic Surgeons and EuroSCORE II risk scores.

Concordance between ChatGPT-4 and HT recommendations was 77% overall but varied by treatment, with ChatGPT-4 correctly matching HT decisions for 90% of patients selected for TAVI, 65% selected for SAVR, and 65% recommended for medical therapy [16]. In comparison, guideline-based classifiers performed worse. A European Society of Cardiology guideline decision tree showed 73% overall agreement with HT recommendations, while an AHA guideline-based classifier disagreed with the HT in 85 (57%) patients. Among the 20 (13%) patients for whom the HT recommended medical therapy, ChatGPT-4 did not suggest SAVR and instead recommended TAVI for seven cases. Review of these cases revealed serious comorbidities, including newly identified oncologic conditions and high perioperative risk not fully captured in the standardized assessments, suggesting that ChatGPT-4 acknowledged surgical risk but may have favored TAVI in borderline cases. Notably, for patients whom the HT decided on SAVR, ChatGPT-4 gave discordant recommendations most commonly in patients aged 70 to 80 years old, a population where ACC/AHA guidelines emphasize individualized choice between TAVI and SAVR [17]. Overall, ChatGPT-4 showed strong agreement with HT decisions and outperformed strict guideline-based decision trees, although accuracy was not uniform across treatment groups.

Pooled Agreement

For quantitative pooling, all studies were harmonized to a common outcome of exact agreement between LLM and HT decisions. The pooled LLM–HT concordance across four studies (REML estimator with HKSJ adjustment, k = 4, total n = 501 patients) was 0.73 (95% CI 0.60–0.83). Between-study heterogeneity was substantial (Cochran’s Q = 9.67, df = 3, I2 = 69.0%), with an estimated between-study variance τ2 = 0.060.

Sensitivity analyses supported the primary estimate, as re-running the model with the DerSimonian–Laird estimator yielded 0.73 (95% CI 0.66–0.80), and the Paule–Mandel estimator with HKSJ adjustment produced 0.73 (95% CI 0.60–0.83). In the pre-specified subgroup analysis by decision structure, the two binary coronary studies (Sudri, Migliaro; k = 2, n = 223) yielded a pooled concordance of 0.74 (95% CI 0.27–0.96), while the two three-way studies (Mola coronary, Salihu aortic; k = 2, n = 278) yielded 0.73 (95% CI 0.39–0.92). In the subgroup analysis by pathology, the three coronary studies (k = 3, n = 351) yielded a pooled concordance of 0.72 (95% CI 0.46–0.88); the single aortic study (Salihu, n = 150) is reported narratively (0.77). In a pre-specified sensitivity analysis restricting each study to its oldest reported LLM (GPT-3.5 for Sudri, GPT-4o for Migliaro and Mola, GPT-4 for Salihu), pooled concordance dropped to 0.63 (95% CI 0.39–0.83), consistent with a temporal improvement across LLM generations.

Quality Assessment

Risk of bias was assessed using the QUADAS-2 framework, which indicated variable levels of risk and applicability concerns across the included studies. Patient selection showed a higher risk of bias in most studies due to differences in study design and case selection. The index test was generally at a low risk of bias, indicating appropriate conduct of the tests in most studies. The reference standard was consistently at low risk of bias, as well as flow and timing, which indicated appropriate study execution. However, several studies showed high concerns regarding applicability, as real-world patient populations may vary from these study populations and imaging was largely not incorporated into LLMs (Table 2).

Table 2.

Quality assessment using QUADAS-2 framework

Study Risk of bias Applicability concerns
Patient selection Index test Reference standard Flow and timing Patient selection Index test Reference standard
Sudri et al. [13] High High Low High High Low Low
Migliaro et al. [14] High Low Low Low High Low Low
Mola et al. [15] High Low Low Low High High Low
Salihu et al. [16] Low Low Low Low Low Low Low

Discussion

The findings of this review highlight the emerging potential of LLMs as adjunctive tools for multidisciplinary HT decision-making. Across four exploratory studies encompassing both coronary and valvular disease, LLMs demonstrated moderate-to-high concordance with HT decisions, with a pooled agreement of 0.73 (95% CI 0.60–0.83). The width of this interval, which reflects both the small number of pooled studies and the HKSJ adjustment for uncertainty in the between-study variance, indicates that while the central estimate is consistent across estimators (REML, DerSimonian–Laird, and Paule–Mandel all produced 0.73), the true pooled concordance could plausibly range from approximately 60% to 83%. Performance also appeared to improve with more recent LLM generations. In the sensitivity analysis restricted to the oldest reported model in each study, the pooled estimate dropped to 0.63, consistent with a temporal trend in LLM clinical reasoning capability. This level of agreement was observed despite substantial heterogeneity across study designs and clinical domains. In individual studies, LLM–HT concordance reached 80% or higher in some subgroups with clearly guideline-driven indications [2, 13–17], but these subgroup observations have not been confirmed in pooled analysis.

These results suggest that LLMs are capable of playing an adjunctive support role in clinical cases, particularly when the case presentation is well structured and supported by clear guidelines. In the studies assessing coronary revascularization, LLM performance was highest when disease burden was substantial with clearly defined treatment pathways, such as in cases with left main or triple vessel CAD [13–15]. Similarly, in severe aortic stenosis cases, LLMs like ChatGPT-4 demonstrated strong overall agreement with HT decisions while outperforming strict guideline-based classifiers, meaning that LLMs were able to integrate multiple clinical variables in a more flexible manner than simple decision trees [16]. Discordance in aortic stenosis cases were limited to patient groups where current guidelines already emphasize individualized decision-making, such as octogenarians considered for TAVI or SAVR [16, 17].

Taken together, this current evidence suggests LLMs may achieve clinically meaningful agreement with HT decisions and potentially offer a scalable framework for supporting multidisciplinary HT decision-making. Contemporary HT workflows demand synthesis of large volumes of heterogenous information, involving clinical history, imaging, and risk scores, among other factors. The findings of the studies included in this review associated detailed and structured prompts with improved LLM–HT concordance, supporting LLMs as scalable tools to organize clinical data prior to multidisciplinary discussion [13–16]. To this end, while LLMs cannot replace expert analysis and consensus, LLMs can be proposed as adjunctive decision support tools to enhance efficiency in HT case review (Fig. 2). Incorporating LLMs into an iterative feedback loop with HT review may yield promising results in the near future.

Fig. 2.

Fig. 2

Conceptual workflow of LLM decision support within the HT framework. AS aortic stenosis, HT heart team, LLM large language model, SAVR surgical aortic valve replacement, TAVI transcatheter aortic valve intervention

However, a structural limitation of all included studies is that these general LLMs cannot directly interpret cardiac imaging, which is central to HT decision-making in both coronary and structural heart disease. Mola et al. explicitly attributed misclassifications to this present inability, which likely contributed to the observed 0% recall for medical therapy decisions [15]. Imaging findings may have been the decisive factor favoring conservative management over procedural intervention in this setting. Until multimodal models capable of native image interpretation are validated in HT decision-making, LLM decision support will remain dependent on the fidelity of text-based clinical summaries provided by clinicians.

The finding that both LLMs evaluated by Mola et al. achieved 0% agreement with HT decisions in cases determined to medical therapy warrants particular attention. This suggests that current GPT-based LLMs may carry a procedural bias in over-recommending intervention when operative risk and clinical factors better suit medical management [15]. This pattern is consistent with Salihu et al. who observed that ChatGPT-4 recommended TAVI for seven patients whom the HT assigned to medical therapy [12].

LLMs can pose several challenges to safe clinical deployment as HTs and health systems consider implementation. Recent studies in interventional cardiology have attributed misclassifications by LLMs to past or misapplied evidence. For example, ChatGPT-3.5, ChatGPT-4, and Bard have been reported to depend on outdated training data by citing the 2010 PARTNER trial when determining optimal aortic valve interventions [18]. Regarding coronary revascularization decision support, the included studies used standardized summaries containing clinical history, comorbidities, and detailed coronary lesion descriptions, finding that quality and completeness of the information provided to the LLMs strongly influenced its reliability [13–15]. Together, these results suggest that LLMs perform best in clearly defined CAD where revascularization guidelines strongly point toward surgical management. The consistently high concordance for CABG decisions across included studies suggests that LLMs are especially confident and reliable when disease burden is extensive and surgical therapy aligns with standard practice. However, the lower agreement for PCI decisions underscores that LLMs currently struggle with nuanced or borderline cases where expert clinicians weigh individual patient factors in addition to angiographic severity.

Issues regarding misclassification persist, especially among higher risk patient populations. LLMs commonly misclassify cardiac treatment recommendations for octogenarians, where European and US guidelines often conflict [2, 17]. It is essential that institutional expertise, patient comorbidities, and intraoperative factors be rigorously validated and cross-checked before input into AI decision support tools so as to avoid error in these high stakes management scenarios [19–21]. The incorporation of richer clinical information can reduce misclassification and discordance with HT recommendations, but doing so may raise concerns related to patient privacy. Transparent regulatory frameworks with robust consent processes and clear attribution of liability can help mitigate these risks to patient privacy and safety as AI is deployed in HTs [19–21].

Looking toward clinical implementation, the available evidence remains largely exploratory and currently positions LLMs as potential adjunctive decision support tools. Current medical and surgical applications of LLMs have been documented to exhibit shortcomings apart from misclassification, including hallucinations consisting of fabricated sources [7, 8, 19–21]. As such, rigorous regulatory oversight must be key to the implementation of any LLMs for clinical decision support, as well as transparency regarding data provenance and LLM adherence to established guidelines. Although the current literature does not yet directly examine data privacy, liability, reimbursement, or electronic health record (EHR) integration, previous calls for supervision implicitly acknowledge these items as essential to any future implementation [7, 8, 19–22]. Future research should embed LLMs into the EHR to evaluate their impact and identify appropriate structures for patient consent and physician oversight. Additionally, given the rapidly growing number of AI tools available to clinicians, future studies should evaluate whether LLM integration meaningfully saves clinicians time that can be redirected toward patient interactions. This is especially compounded by the increasing usage of free online LLMs by patients for second opinions, which necessitates the cooperation of professional societies, payers, and regulatory agencies to ensure safe deployment of LLMs for decision support [23].

Although the studies included in this review only evaluated GPT-based models, the broader environment of LLMs relevant to clinical decision support is rapidly expanding. GPT-4o is a multimodal GPT-family model, whereas GPT-o1 is specifically trained with reinforcement learning to perform complex reasoning using chain-of-thought [24, 25]. However, both remain proprietary systems, limiting independent verification of model weights, training-data provenance, update processes, and knowledge cutoffs [24]. Claude-3.5 Sonnet has demonstrated competitive performance in medical-agent benchmarking [26], and Claude-4 models have been introduced as hybrid reasoning models with strong general reasoning capabilities, although neither has been evaluated in a HT concordance study. Gemini-2.5 is a multimodal “thinking” model with long-context and tool-use capabilities [24], while domain-specific Med-PaLM 2 and Med-Gemini models have shown strong performance on US Medical Licensing Examination (USMLE)-style and broader medical benchmarks [27]. However, access to Med-PaLM 2 and Med-Gemini has remained limited to selected testing or research collaborations, and validation in HT decision-making is lacking. There is substantial variability across LLMs in prompt sensitivity, hallucination and omission rates, clinical-input handling, evaluation validity, and transparency [24–28]. As such, direct comparative evaluation of LLMs for HT decision-making is complicated yet remains an important direction for future research, and comparative claims should presently be treated as hypothesis-generating rather than definitive.

Limitations

This review should be considered in the context of several important limitations. Although the present work captured the best current evidence regarding LLMs as clinical decision support in the HT, updated and more advanced AI models are being rapidly made available to the public and thus may render the capabilities of the aforementioned LLMs as outdated. While different versions of ChatGPT were the only LLMs examined in this review as an indirect result of our selection criteria, other commercial, domain-specific, and open-source LLMs may yield different results than those in the present work. A comparative analysis of different LLM versions was not feasible given the relatively small number of available studies, substantial heterogeneity in clinical tasks, prompting strategies, and outcome definitions. This review also carries an inherent subjectivity in the selection and interpretation of studies. We attempted to minimize any bias by including a broad range of studies across different cardiovascular pathologies and conducting risk of bias assessment. Nevertheless, the scope of the review is limited by the availability and quality of the available literature.

Several additional methodological limitations deserve explicit acknowledgement. First, the pooled estimate is based on only four studies, which is at the lower bound for meta-analytic pooling. We addressed this by using the HKSJ adjustment, but the resulting confidence interval remains wide and should be interpreted as a provisional effect size rather than a precise population estimate. Second, the included studies differ in decision structure, pathology, LLM generation, and prompting strategy, and this clinical heterogeneity is the most likely driver of the observed statistical heterogeneity. Third, the primary outcome of exact agreement does not capture the clinical severity of disagreements. For example, an LLM recommending medical therapy when the HT chose CABG is a more consequential mismatch than CABG versus PCI, yet both count equally in our concordance metric. All included studies evaluated LLMs retrospectively on standardized vignettes. Prospective studies with real-time integration of imaging, laboratory data, and patient preferences remain needed before clinical deployment can be recommended. None of the included studies reported calibration, which will be an important consideration for safe deployment. Moreover, the pooled exact agreement proportion combines studies with binary decision structures and three-way decision structures. The pooled proportion is therefore a descriptive summary and does not correspond to a chance-adjusted performance metric. Lastly, no included study reported inter-rater reliability among individual HT members prior to consensus, and HT composition was variable. The pooled analysis treats these non-equivalent reference standards identically, which is a limitation of the currently available evidence. Future primary studies should report pre-consensus inter-rater reliability to permit assessment of reference-standard stability.

Conclusions

Overall, current LLMs are promising supplementary decision support tools for multidisciplinary HTs. In well-defined conditions such as left main or triple vessel disease and severe aortic stenosis, individual studies observed signals of higher agreement between LLM recommendations and HT decisions, but there is potential for misclassification when patient-specific factors and conflicting guidelines complicate decision making, and current LLMs show a systematic failure to recommend medical therapy when it is the appropriate choice. Contemporary LLMs continue to show areas of improvement regarding incorporation of current trial evidence, limited transparency regarding data or guidelines used, and misclassifications in high-risk populations. These findings underscore how LLMs can support clinician judgement yet require structured model inputs, ongoing audit and quality improvement, transparency about training data, and overall strong oversight before integration into HTs or other clinical workflows. In the near future, prospective trials assessing their efficacy in the EHR can help physicians and regulators address current challenges. HTs worldwide should take an active role in shaping the use of LLMs to enhance cardiovascular decision-making while safeguarding patient care.

Author Contribution

Armaun D. Rouhi—study conception, data curation, data analysis, manuscript writing, critical revision; Shreyas V. Menon—data curation, manuscript writing, critical revision; Yazid K. Ghanem—data curation, manuscript writing, critical revision; Jason J. Han—study conception, critical revision, supervision. All named authors meet the International Committee of Medical Journal Editors (ICMJE) criteria for authorship for this article, take responsibility for the integrity of the work as a whole, and have given their approval for this version to be published.

Funding

No funding or sponsorship was received for the publication of this article.

Data Availability

The datasets generated during and/or analyzed during the current study are available from the corresponding author on reasonable request.

Declarations

Conflict of Interest

Armaun D. Rouhi, Shreyas V. Menon, Yazid K. Ghanem, and Jason J. Han have nothing to disclose.

Ethical Approval

This article does not contain any new studies with human participants or animals performed by any of the authors.

References

  • 1.Han JJ, Brown CR. The heart team: a powerful paradigm for the future training of cardiovascular surgeons. J Am Coll Cardiol. 2018;71(23):2702–5. 10.1016/j.jacc.2018.05.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Lawton JS, Tamis-Holland JE, Bangalore S, et al. 2021 ACC/AHA/SCAI guideline for coronary artery revascularization: a report of the American College of Cardiology/American Heart Association Joint Committee on clinical practice guidelines. Circulation. 2022;145(3):e18–114. 10.1161/CIR.0000000000001038. [DOI] [PubMed] [Google Scholar]
  • 3.Khan S, Shi W, Kaneko T, Baron SJ. The evolving role of the multidisciplinary heart team in aortic stenosis. US Cardiol Rev. 2022;16:e19. 10.15420/usc.2022.04. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Arjomandi Rad A, Streukens S, Vainer J, Athanasiou T, Maessen J, Sardari Nia P. The current state of the multidisciplinary heart team approach: a systematic review. Eur J Cardiothorac Surg. 2024;67(1):ezae461. 10.1093/ejcts/ezae461. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Burlacu A, Covic A, Cinteza M, Lupu PM, Deac R, Tinica G. Exploring current evidence on the past, the present, and the future of the heart team: a narrative review. Cardiovasc Ther. 2020;2020:9241081. 10.1155/2020/9241081. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Marghitu T, Roberts SH, He J, et al. Impact of transcatheter aortic valve replacement use ratio on outcomes in patients with aortic valve disease. J Thorac Cardiovasc Surg. 2025;170(2):468-475.e2. 10.1016/j.jtcvs.2024.10.046. [DOI] [PubMed] [Google Scholar]
  • 7.Maddox TM, Embí P, Gerhart J, Goldsack J, Parikh RB, Sarich TC. Generative AI in medicine—evaluating progress and challenges. N Engl J Med. 2025. 10.1056/NEJMsb2503956. [DOI] [PubMed] [Google Scholar]
  • 8.Shah NH, Entwistle D, Pfeffer MA. Creation and adoption of large language models in medicine. JAMA. 2023;330(9):866–9. 10.1001/jama.2023.14217. [DOI] [PubMed] [Google Scholar]
  • 9.Rouhi AD, Ghanem YK, Yolchieva L, et al. Can artificial intelligence improve the readability of patient education materials on aortic stenosis? A pilot study. Cardiol Ther. 2024;13(1):137–47. 10.1007/s40119-023-00347-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Whiting PF, Rutjes AW, Westwood ME, et al. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. 2011;155(8):529–36. 10.7326/0003-4819-155-8-201110180-00009. [DOI] [PubMed] [Google Scholar]
  • 11.DerSimonian R, Laird N. Meta-analysis in clinical trials revisited. Contemp Clin Trials. 2015;45(Pt A):139–45. 10.1016/j.cct.2015.09.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Tanriver-Ayder E, Faes C, van de Casteele T, McCann SK, Macleod MR. Comparison of commonly used methods in random effects meta-analysis: application to preclinical data in drug discovery research. BMJ Open Sci. 2021;5(1):e100074. 10.1136/bmjos-2020-100074. . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Sudri K, Motro-Feingold I, Ramon-Gonen R, et al. Enhancing coronary revascularization decisions: the promising role of large language models as a decision-support tool for multidisciplinary heart team. Circ Cardiovasc Interv. 2024;17(11):e014201. 10.1161/CIRCINTERVENTIONS.124.014201. [DOI] [PubMed] [Google Scholar]
  • 14.Migliaro S, Celotto R, Teliti R, Mariani S, Altamura L, Tomai F. Comparing AI-driven and heart team decision-making in multivessel coronary artery disease. J Clin Med. 2025;14(13):4452. 10.3390/jcm14134452. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Mola S, Yıldırım A, Gül EB. Artificial intelligence in cardiac treatment decision-making: an evaluation of the performance of ChatGPT versus the heart team in coronary revascularization. Rev Cardiovasc Med. 2025;26(8):38705. 10.31083/RCM38705. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Salihu A, Meier D, Noirclerc N, et al. A study of ChatGPT in facilitating heart team decisions on severe aortic stenosis. EuroIntervention. 2024;20(8):e496–503. 10.4244/EIJ-D-23-00643. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Otto CM, Nishimura RA, Bonow RO, et al. 2020 ACC/AHA guideline for the management of patients with valvular heart disease: executive summary: a report of the American College of Cardiology/American Heart Association Joint Committee on clinical practice guidelines. Circulation. 2021;143(5):e35–71. 10.1161/CIR.0000000000000932. [DOI] [PubMed] [Google Scholar]
  • 18.Leon MB, Smith CR, Mack M, et al. Transcatheter aortic-valve implantation for aortic stenosis in patients who cannot undergo surgery. N Engl J Med. 2010;363(17):1597–607. 10.1056/NEJMoa1008232. [DOI] [PubMed] [Google Scholar]
  • 19.Itelman E, Witberg G, Kornowski R. AI-assisted clinical decision making in interventional cardiology: the potential of commercially available large language models. JACC Cardiovasc Interv. 2024;17(15):1858–60. 10.1016/j.jcin.2024.06.013. [DOI] [PubMed] [Google Scholar]
  • 20.Koch V. Can artificial intelligence help heart teams make decisions? EuroIntervention. 2024;20(8):e465–6. 10.4244/EIJ-E-24-00016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Zhang X, Wang J. Ensuring safety and reliability in ChatGPT-assisted pediatric cardiovascular surgical decision-making. J Thorac Cardiovasc Surg. 2025. 10.1016/j.jtcvs.2025.02.015. [DOI] [PubMed] [Google Scholar]
  • 22.Anyanwu EC, Fanaroff AC, Maddox TM. Large language models and revascularization decisions: the newest member of your multidisciplinary heart team? Circ Cardiovasc Interv. 2024;17(11):e014775. 10.1161/CIRCINTERVENTIONS.124.014775. [DOI] [PubMed] [Google Scholar]
  • 23.Lautrup AD, Hyrup T, Schneider-Kamp A, Dahl M, Lindholt JS, Schneider-Kamp P. Heart-to-heart with ChatGPT: the impact of patients consulting AI for cardiovascular health advice. Open Heart. 2023;10(2):e002455. 10.1136/openhrt-2023-002455. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Wang X, Xiong Z, Zou K, et al. Reasoning-driven large language models in medicine: opportunities, challenges, and the road ahead. Lancet Digit Health. 2026;8(1):100931. 10.1016/j.landig.2025.100931. [DOI] [PubMed] [Google Scholar]
  • 25.Li CP, Chu Y, Jia WW, et al. Punching above its weight: a head-to-head comparison of Deepseek-R1 and OpenAI-o1 on Pancreatic Adenocarcinoma-related questions. Int J Med Sci. 2025;22(15):3868–77. 10.7150/ijms.118887. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Bedi S, Cui H, Fuentes M, et al. Holistic evaluation of large language models for medical tasks with MedHELM. Nat Med. 2026;32(3):943–51. 10.1038/s41591-025-04151-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Singhal K, Tu T, Gottweis J, et al. Toward expert-level medical question answering with large language models. Nat Med. 2025;31(3):943–50. 10.1038/s41591-024-03423-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Suh PS, Shim WH, Suh CH, et al. Comparing large language model and human reader accuracy with New England Journal of Medicine image challenge case image inputs. Radiology. 2024;313(3):e241668. 10.1148/radiol.241668. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The datasets generated during and/or analyzed during the current study are available from the corresponding author on reasonable request.


Articles from Cardiology and Therapy are provided here courtesy of Springer

RESOURCES