Skip to main content
NPJ Digital Medicine logoLink to NPJ Digital Medicine
. 2026 Jun 2;9:672. doi: 10.1038/s41746-026-02837-6

Clinical outcomes and reporting quality of large language model interventions in practice: a systematic evidence map

Zixuan He 1,2, Lan Yang 1,2, Zitao Liang 1,2, Xiaofan Li 3, Jian Du 2,✉
PMCID: PMC13530203  PMID: 42230743

Abstract

Large language models (LLMs) are being deployed in clinical settings despite an underdeveloped evidence base regarding their real-world effectiveness. This study employed systematic evidence mapping to characterize outcome measures used in published studies and registered clinical trials (Jan 2022–Jun 2025) evaluating LLM performance. Analysis of 55 included studies revealed a predominance of human-AI collaborative designs (65.5%) for decision support and symptom management. LLM-only interventions focused on functional performance and operational or process impact outcomes (e.g., accuracy and time saving), whereas LLM-assisted interventions showed positive clinical effects, particularly in psychological health endpoints. Critical evidence gaps persist: diagnostic accuracy in randomized trials was notably lower and more variable (range 0.65–0.88) compared to non-randomized studies (typically ≥ 0.80); clinical efficiency impacts were inconsistent, and reporting quality was suboptimal (78.8% mean CONSORT-AI adherence), with critical omissions in handling data quality and performance errors. These findings indicate a heterogeneous and insufficient evidence landscape, necessitating standardized core outcome sets, mandatory use of specialized reporting guidelines, and robust clinical trials to ensure the safe integration of LLMs.

Subject terms: Diseases, Health care, Medical research

Introduction

Large language models (LLMs) are increasingly integrated into healthcare, with potential to enhance diagnostic accuracy, clinical decision-making, and operational efficiency. Empirical evaluations in controlled settings demonstrate that LLMs can achieve performance comparable to medical professionals, particularly in tasks such as diagnostic reasoning and managing complex clinical scenarios1–4. However, most evaluations benchmark model outputs against expert answers rather than real-world clinical endpoints, limiting their clinical relevance5–7. This evidence gap is especially concerning given the speed of adoption. By early 2025, for example, DeepSeek had reportedly been deployed across more than 300 hospitals in China, spanning diagnosis, patient education, and management systems—often without rigorous prospective validation. This gap between rapid implementation and limited clinical evaluation raises uncertainty regarding effectiveness, safety, and equity in real-world practice.

Performance variability further underscores these concerns. Unregulated implementation of LLMs will introduce substantial risks including diagnostic errors, ethical dilemmas, and potential sources of doctor-patient conflict. These risks manifest in two distinct patterns: uncritical adoption due to over-reliance on AI outputs, and increased cognitive burden for clinicians who attempt to verify AI-generated recommendations against established evidence8. Several clinical validation studies reveal concerning performance disparities and indicated that current LLMs perform significantly worse than clinicians on aggregate across all diseases. Furthermore, LLMs exhibit worrying tendencies toward ‘diagnostic anchoring’, forming premature conclusions without sufficient evidence and demonstrate limited ability to integrate sequential clinical information effectively9,10.

As standardized and consistent outcome measurement is fundamental to reliably evaluating the safety and efficacy of healthcare interventions11. The Core Outcome Measures in Effectiveness Trials (COMET) Initiative promotes the development and implementation of core outcome sets (COS) to address heterogeneity and reporting bias in clinical studies. These are agreed-upon sets of outcomes, established through consensus among key stakeholders (including patients, carers, and healthcare professionals). By mandating consistent measurement and reporting across trials, COMET enhances comparability, minimizes bias, and ensures that research assesses outcomes meaningful to those affected. Integrating such structured, patient-centered frameworks can align the evaluation of emerging tools like LLMs with clinically relevant endpoints, bridging the gap between technical performance and real-world health impact12,13.

Further, interpretation of existing evidence is hindered by substantial heterogeneity in study design and inconsistent reporting14. To address these challenges, specialized frameworks have been introduced, including the reporting guidelines for clinical trial reports for interventions involving artificial intelligence (the CONSORT-AI extension15) and guidelines for clinical trial protocols for interventions involving artificial intelligence (the SPIRIT-AI extension16) for randomized controlled trial (RCT) design and reporting of AI interventions. Most recently, the reporting guideline for chatbot health advice studies: the chatbot assessment reporting tool (CHART) statement17 was developed to further standardize the reporting for studies of generative AI health advice chatbots. Yet adherence to these standards remains limited, undermining transparency and reproducibility18,19.

In this study, we provide a systematic review and evidence map for studies evaluating clinical applications of LLMs. We synthesize published studies and registered trials, categorize interventions by mode (LLM-only or LLM-assisted), task (decision support, symptom intervention, or information queries), and outcome type (health endpoints, functional performance, operational or process impact), synthesized reported outcome measures and compare performance across categories. We also assess reporting quality against CONSORT-AI and CHART. This comprehensive evaluation aims to clarify the current evidence base, identify core domains and metrics for measuring LLMs’ clinical performance, and inform safe, ethical, and effective integration of LLMs into healthcare.

Results

In total, 318 publications were retrieved following deduplication. 303 articles that did not meet the inclusion criteria were excluded, and 15 studies containing 7 RCTs and 8 nonrandomized studies of intervention (NRSI) studies were included in the review. For registered trials, 456 trials were identified and screened for eligibility, of which 416 does not meet the inclusion criteria, a total of 40 trials comprising 33 RCT and 7 NRSI records were included, and a total of 320 outcomes were extracted. The study selection flow is detailed in the PRISMA diagram (Fig. 1). Table 1 presents key features of all 15 included published studies, with the detailed characteristics of all included publications and clinical trials listed in Supplementary Table 1. The Cohen’s Kappa index between researchers is 0.84, as detailed in Supplementary Note 2.

Fig. 1. Data identification and screening flowchart.

Fig. 1

This flowchart shows the identification and screening process. In total, 318 publications were retrieved following deduplication, 15 studies containing 7 RCTs and 8 nonrandomized studies of intervention (NRSI) studies were included in the review. For registered trials, 456 trials were identified and screened for eligibility. Ultimately, 33 RCT registration records and 7 NRSI records are included.

Table 1.

Key features of included publications

Study/Trial ID Intervention Task Sample size Primary Outcome Outcome Cate.
Publication- RCTs
Harari (2024)28 Supervised-ChatGPT Provide clinical guidance during cardiac arrest 54 clinicians Accuracy Functional performance
Gan (2025)29 ChatGPT assist Alleviate anxiety in total knee arthroplasty consent process 55 patients HADS-A Health endpoints
Goh (2023)30 ChatGPT assist Suggest chest pain clinical decisions 50 physicians Accuracy Functional performance
Goh (2024)1 Conventional resource plus ChatGPT Perform diagnostic decision reasoning 50 physicians Diagnostic performance Functional performance
Goh (2025)2 Conventional resource plus ChatGPT Perform management decision reasoning 92 physicians Performance score Functional performance
Heinz (2025)20 Gen-AI chatbots Provide mental disorders treatment 210 patients Disorder-specific symptom Health endpoints
Akdogan (2025)31 ChatGPT assist Reduce anxiety and depression in cancer patients 150 patients HADS-A Health endpoints
Publication- NRSIs
Vaira (2024)32 ChatGPT Answer questions and solve clinical scenarios of head and neck surgery 18 cases Accuracy Functional performance
Cheng (2024)33 ChatGPT Assess preoperative risk and manage day surgery anaesthesia 150 patients ASA-PS Health endpoints
Turan (2024)34 ChatGPT Predict ASA scores in preoperative risk assessment 2851 patients ASA score Health endpoints
Turan (2025)35 ChatGPT Predict postoperative ICU admission 406 patients Accuracy Functional performance
Sanduleanu (2024)36 ChatGPT Identify suspected Appendicitis 113 patients Accuracy Functional performance
Zhou (2024)37 Gemini Identify suspicious breast lesions 1062 patients Diagnostic performance Functional performance
Bayala (2025)38 ChatGPT, Gemini, Copilot, Claude Diagnose African rheumatic diseases 103 patients Accuracy Functional performance
Gurbuz (2025)39 ChatGPT, DeepSeek Address queries on gonarthrotic and total knee arthroplasty 100 patients Accuracy Functional performance

The majority of studies adopted the physician’s perspective, with only three focusing on patients. ChatGPT (including GPT - 4, GPT - 3.5, and other versions) was the most frequently evaluated LLM, followed by Gemini. The targeted medical conditions were diverse, including cardiac arrest, total knee arthroplasty, chest pain, mental disorders, pre-operative risk assessment, postoperative ICU admission, suspected appendicitis, and suspicious breast lesions.

Prevalence of human - AI collaborative design

Human-AI collaborative designs—where interventions utilized LLMs to assist rather than replace clinicians—were evaluated in 40.0% (6/15) of published studies, 75% (30/40) of registered trials, and 65.5% (36/55) when considering all 55 included records (published studies and registered trials combined). Among randomized controlled trials specifically, collaborative designs accounted for 6 of 7 published RCTs (85.7%) and 33 of 40 registered RCTs (82.5%). These patterns underscore the prevailing focus on augmenting rather than replacing clinical expertise.

In terms of comparative effectiveness, among the published studies that have reported results, 5 of 6 (83.3%) that compared LLM-based interventions against traditional methods or conventional resources reported superior performance. Findings from registered trials on comparative effectiveness are not yet available, as most studies are ongoing and results remain unpublished.

Categorization of outcome measures

We categorized the 320 extracted outcomes using the pre-specified four-tier framework designed to reflect the evaluation of LLM interventions along the translational pipeline. The Sankey diagram (Fig. 2) illustrates the structural relationships between user perspectives, LLM tasks, and outcome tiers. LLMs were primarily deployed for decision support (42.3%) and symptom intervention (36.8%), while pure information queries were less frequent (11.5%). Among the 320 documented outcomes, the majority (63.1%, n = 202) were health endpoints, dominated by scale-based patient-reported measures. Functional and operational outcomes comprised 21.6% (n = 69) of the total, driven mainly by diagnostic performance metrics (e.g., accuracy; n = 19) and process efficiency measures (e.g., time-to-decision; n = 9). Notably, only eight outcomes represented physiological or biomarker endpoints, underscoring a critical gap in objective clinical evaluation. The remaining 15.1% (n = 48) were classified as “Others” which included measure of physician trust, experience, satisfaction, etc.

Fig. 2. Sankey diagram between user perspectives, LLM tasks, and outcome categories.

Fig. 2

Fig. 2 illustrates the relationships between user perspectives, LLM tasks, and outcome tiers through a Sankey diagram. LLMs were predominantly deployed for clinical decision support (42.3%) and symptom intervention (36.8%), with information queries being less common (11.5%). Among 320 documented outcomes, most (63.1%) were health endpoints, largely patient-reported measures related to behavior and symptoms, whereas functional performance (17.3%) were dominated by diagnostic yield (e.g., accuracy) and operational/workflow impact (4.2%) dominated by efficiency.

Evidence map of LLM effects

We then synthesized findings from 15 published studies with available results, covering 48 distinct clinical outcomes (Fig. 3). The evidence map differentiates between LLM-only and LLM-assisted interventions. LLM-only interventions, where models operate independently, were primarily associated with functional performance such as accuracy, efficiency, and performance. By contrast, LLM-assisted interventions, where models support clinicians or patients, extended to a broader set of health endpoints, including psychological (e.g., anxiety, depression), cognitive, and physiological measures (e.g., heart rate variability, pain).

Fig. 3. Evidence map of LLM effects in clinical settings, by outcome type.

Fig. 3

Red edges represent positive effects, blue edges indicate negative effects, and purple edges reflect proportionally inconsistent evidence. Edges are not drawn where the comparison showed no statistically significant difference between LLM and comparator groups, or where the direction of effect was not reported or could not be determined. The full list of outcomes included in the evidence map are provided in Supplementary Table 2. Legend: Fig. 3 presents an evidence map synthesizing 48 clinical outcomes from 15 published studies, stratified by intervention type. LLM-only interventions were predominantly associated with functional performance or operational impact (e.g., accuracy and efficiency), while LLM-assisted interventions demonstrated broader impact across psychological, cognitive, and physiological health endpoints. Dense evidence clusters showed positive effects on accuracy, time saving, and psychological outcomes (e.g., HADS scores), whereas physiological biomarkers remained substantially under-represented.

The direction of effects varied across domains: dense evidence clusters were observed for improvements in accuracy, time saving, and psychological outcomes such as hospital anxiety and depression scale (HADS) scores. In contrast, physiological outcomes were infrequently assessed, reflecting a notable evidence gap due to the scarcity of biomarker endpoints. Overall, this evidence map highlights the differential impact of LLMs depending on their mode of integration into clinical workflows: while LLM-only applications address functional performance, LLM-assisted interventions are more likely to shape patient-centered and health-related outcomes. The observed heterogeneity also underscores the need for larger, higher-quality RCTs to resolve conflicting findings and expand evaluation beyond subjective and process-related endpoints.

We further explored the potential sources of heterogeneity in outcome effects, hypothesizing that variations may be substantially associated with differences in study design, particularly the user perspective. To investigate this, we categorized the 33 registered RCTs into two groups: those evaluating LLMs designed for use by healthcare professionals involved in direct patient care (including physicians, surgeons, and anesthetists), and those for patients. We counted and averaged the number of inclusion and exclusion criteria reported in each trial protocol or publication descriptively. Our analysis revealed that trials focusing on patient-facing LLMs employed more stringent eligibility criteria, with an average of 4.1 inclusion and 3.1 exclusion criteria, compared to an average of 2.5 inclusion and 2.1 exclusion criteria for physician-facing LLM trials. Furthermore, notable distinctions were observed between the two groups regarding restrictions on participant age, underlying health conditions, concurrent use of other treatments, digital accessibility, and clinical settings (Suppl. Table 3).

Accuracy and Time Efficiency of LLMs

As shown in Fig. 4, accuracy estimates varied widely across studies, ranging from 0.64 to 0.99. Non-randomized studies (NRSIs) reported consistently higher values, typically above 0.80, with the largest study (n = 2,851) achieving 0.92–0.99. By contrast, RCTs with smaller sample sizes (50–150) showed more moderate accuracy (0.65–0.88). When deployed as clinician support tools, LLMs demonstrated accuracies of 65–99%, with studies enrolling more than 100 participants consistently reporting values above 80%. Evidence for patient-side LLMs is limited and shows considerable variability in accuracy.

Fig. 4. Bubble chart of heterogeneity in the accuracy of LLM applications across clinical studies, by study design and user perspective.

Fig. 4

Bubble size is proportional to the study sample size. Fig. 4 illustrates the variability in diagnostic accuracy estimates across study designs. Non-randomized studies consistently reported higher accuracy values (≥ 0.80) compared to RCTs, which demonstrated more moderate performance (0.65–0.88). While physician-supported LLMs generally achieved accuracies above 80% in larger studies, evidence for patient-facing applications remains limited and highly variable.

Table 2 summarizes findings from four RCTs evaluating accuracy and time efficiency. Three RCTs reported significant accuracy gains of 20.4–38.3% with LLM-assisted interventions, while one found only a marginal, non-significant improvement (76% vs. 74%, p = 0.6). In terms of time efficiency, LLMs reduced decision-making time in two-third of study findings, although one RCT reported a non-significant reduction and another noted a significant increase in decision time in the LLM group (p = 0.022).

Table 2.

LLMs’ improvement in accuracy and efficiency compared to traditional methods in RCTs

graphic file with name 41746_2026_2837_Tab1_HTML.webp

Yellow denotes outcomes where the LLM intervention was superior to conventional care, while blue denotes inferiority. Darker shades indicate statistically significant differences, and lighter shades represent non-significant differences. This table includes only published studies that directly compared LLM-assisted interventions against traditional or conventional methods and reported quantifiable accuracy outcomes. Studies employing other comparator designs or not reporting accuracy as an outcome were excluded from this summary.

Reporting quality and risk of bias assessment of included studies

Overall concordance of RCTs with CONSORT-AI items was 78.8% (Supplementary Table 4 and Supplementary Table 5). Most studies clearly described the purpose of LLM intervention, intended users, and its place in the clinical workflows. However, several critical reporting gaps were identified. None of the studies addressed the management of poor-quality or missing input data (0% for item 5iii). Only two-thirds reported performance error analyses (66.7% for item 19). Additionally, only half of the studies specified how the AI intervention or its code could be accessed (50.0% for item 25), limiting transparency and reproducibility.

Notably, Heinz et al. (2025) conducted the first RCT of a generative AI therapy chatbot20, achieving relatively high reporting quality with 78.4% compliance to CONSORT-AI guidelines. However, evaluation against the specialized CHART guideline revealed substantial deficiencies. More than half of the required items were either absent or insufficiently reported. Key omissions included the identity and version of the underlying AI model, details of prompt engineering and disclosure, transparency of chatbot dialogues, clear benchmarking against clinical gold standards, and quantitative assessment of potential harms. Such gaps limit comprehensive appraisal, reproducibility, and robust safety evaluation of the intervention.

The risk-of-bias assessment was conducted for all included studies using the Cochrane RoB 2 tool for randomized controlled trials and the ROBINS-I tool for non-randomized studies. Detailed results of this assessment are presented in Supplementary Figs. 1, 2. Overall, the methodological quality of included studies varied considerably, with key limitations identified across multiple domains including selection of participants, selection of reported results, and outcome measurement. These findings underscore the need for improved trial conduct and reporting in future LLM clinical research.

Discussion

This systematic evidence map suggests that clinical evaluations of LLM interventions remain in their early stages. The published evidence base is limited and characterized by substantial heterogeneity across study designs, populations, and outcome measures. Human-AI collaborative designs predominate, indicating a clear emphasis on augmenting rather than replacing clinical expertise. However, critical evidence gaps persist: diagnostic accuracy in randomized controlled trials was notably lower and more variable than in non-randomized studies, effects on clinical efficiency were inconsistent, physiological or biomarker endpoints were scarcely represented, and reporting of AI-specific elements is frequently incomplete. Collectively, these patterns impede cross-study comparability and constrain our ability to infer clinical effectiveness beyond narrow tasks and short-term outcomes. Our risk‑of‑bias assessment further demonstrated that non‑randomized studies were consistently at higher risk of bias, and those with the highest accuracy estimates faced more methodological concerns, suggesting that the substantially higher accuracy may be driven in part by bias rather than by genuinely superior model performance. As recently argued in Nature Medicine, claims of clinical value for medical AI must be supported by proportionate evidence—stronger claims requiring stronger prospective validation—yet the field currently lacks consistent evidentiary standards, leading to scientific uncertainty and premature implementation21. Our findings align with this call by demonstrating that reliance on technical metrics alone, without corresponding evidence of clinical actionability or workflow benefit, risks adopting LLM tools whose real-world value remains uncertain and whose unintended consequences may be substantial.

At the efficacy level, several mechanisms may explain why psychological health endpoints appear more frequently—and sometimes demonstrate larger apparent effects—than physiological endpoints. Such outcomes are often proximal to the intervention (e.g., conversational support or education), are feasible to measure within short follow-up periods, and may be more sensitive to expectancy and context effects than objective biomarkers. In contrast, biomarker and “hard” clinical endpoints typically require longer follow-up, larger sample sizes, and more rigorous safety monitoring. These requirements are often difficult to reconcile with the rapid iterative cycles of LLM systems, where frequent model updates, prompt engineering changes, and complex workflow integrations introduce significant operational noise.

Our study interprets the observed performance heterogeneity with caution. Non-randomized evaluations may overestimate clinical benefits due to design artifacts—such as vignette-based testing and retrospective case selection—as well as limited blinding and selective reporting. Conversely, randomized controlled trials (RCTs) are more likely to encounter real-world operational constraints, thus yielding more modest and variable effects. While hybrid human-AI collectives have demonstrated the potential to outperform either party alone in complex diagnostics22, this synergy is not guaranteed. The effectiveness of human-AI collaboration likely depends on multiple factors, including task complexity, the clinical expertise of the user, the design of the AI interface, and the presence of clear escalation pathways for handling uncertain outputs. According to a synthesis of empirical studies on human-AI teaming, the realization of synergy depends heavily on collaboration design—specifically whether clinicians review AI outputs concurrently versus sequentially—and on clinician expertise, with junior practitioners often deriving significantly greater benefits than their senior counterparts23. These insights align with our findings that clinical effects vary by study context, underscoring the necessity for future trials to explicitly specify teaming modes, human oversight mechanisms, and user expertise to enable meaningful cross-intervention comparisons.

Relatedly, patient-facing trials in our sample tended to adopt stricter eligibility criteria than clinician-facing trials. This trend likely reflects heightened safety, ethical, and liability concerns regarding direct LLM-patient interactions. While such rigor improves internal validity, it may inadvertently reduce the generalizability of the findings. Future studies should transparently report the rationale and trade-offs of their eligibility decisions and, where feasible, complement early efficacy trials with pragmatic evaluations that better reflect routine clinical care.

For future research, our results propose three practical priorities. First, the development and adoption of a Core Outcome Set (COS) for AI-based clinical interventions that spans model performance, workflow impact, as well as health endpoints. Such standardization would enable meaningful cross-trial comparison and meta-analysis, addressing the heterogeneity that currently undermines evidence synthesis. Second, the explicit specification and evaluation of human-AI teaming configurations in future trial designs. Researchers should clearly define and systematically vary parameters including: the timing of AI input (concurrent vs. sequential use relative to clinician decision-making); boundaries of AI autonomy (e.g., required human oversight for specific tasks); escalation pathways for handling uncertain or high-risk cases; and the level of user expertise required to safely interpret AI outputs. These design choices fundamentally shape both the effectiveness and safety of LLM interventions and warrant rigorous empirical investigation. Third, enhanced methodological transparency through prospective registration and rigorous application of AI-specific reporting guidelines, such as CONSORT-AI, SPIRIT-AI, or CHART guidelines as applicable. We specifically recommend that trials: (a) pre-specify sample sizes adequate to detect clinically meaningful differences in patient-centred outcomes; (b) include diverse patient populations to assess generalizability and equity impacts; and (c) incorporate pre-planned analyses of performance errors and handling of poor-quality input data, which were commonly absent in the current evidence base.

This study is subject to several limitations. First, the relatively small number of published trials and the heterogeneity in study designs, populations, and outcome measures may limits the generalizability of our findings. The evidence map should therefore be viewed as a systematic overview of the current landscape rather than a definitive comparative assessment of LLM efficacy. Second, outcome metrics were reported inconsistently across studies, and critical domains—particularly physiological biomarkers and long-term clinical endpoints—were rarely assessed, constraining our ability to evaluate the full impact of LLM interventions. Third, despite our comprehensive search strategy, the potential for publication bias cannot be excluded, as studies reporting positive or statistically significant findings may be more likely to be published than those with null or negative results. Fourth, the rapid and continuous evolution of LLM technologies means that new evidence continues to emerge, and our findings represent a snapshot of the evidence up to June 2025. Finally, the scarcity of long-term follow-up data across included studies limits any conclusions regarding the sustained effects, safety, or cost-effectiveness of LLM-based interventions in routine clinical practice. Periodic updates to this evidence map will therefore be essential to track the accumulating evidence and refine our understanding as the field evolves.

In conclusion, the current evidence base for LLM interventions in clinical practice remains uneven. While early studies indicate potential benefits, the field urgently requires standardized outcome frameworks and explicit reporting of teaming modes. Transitioning from “AI-only” performance metrics to rigorous evaluations of human-AI synergy and clinically meaningful health endpoints will be critical for the safe and effective integration of LLMs into healthcare.

Methods

Data sources and search strategy

We conducted a systematic search to identify studies evaluating the clinical use of large language models (LLMs). We searched the following electronic databases for published studies: PubMed, Embase, and Web of Science. For registered trials, we searched ClinicalTrials.gov and the WHO International Clinical Trials Registry Platform (WHO ICTRP). The searching strategy employed a combination of terms related to LLMs and clinical trial restriction words (full searching strategy refers to Suppl. Note 1). We also screened reference lists of relevant reviews to identify additional studies. The search was restricted to studies published or registered between January 1, 2022, and June 30, 2025, to ensure inclusion of the most recent developments in this rapidly evolving field.

A prospective protocol for this systematic review was registered on the PROSPERO international prospective register of systematic reviews (CRD420251002623). The review was conducted and reported in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines24. The authors declare no patient or public involvement in this study. Human Ethics and Consent to Participate declarations is not applicable.

Eligibility criteria

Studies were included if they were clinical trials (randomized or non-randomized) conducted in healthcare settings and evaluated an LLM-based intervention for screening, diagnosis, treatment, or clinical management. Eligible studies encompassed participants of all ages and across various medical conditions, with no restrictions imposed on comparator types. Only publications in English or Chinese were considered for inclusion.

Studies were excluded if they were not conducted in a healthcare setting, such as those focused solely on LLM model development, training or testing. Additionally, studies employing LLMs for non-clinical purposes—including medical education, academic research assistance, or administrative support—were excluded. Trials in which LLMs served as comparators or ancillary tools rather than as the primary intervention under investigation were also excluded, along with non-interventional study designs such as reviews, commentaries, and conference abstracts without complete results. Note that any registered trial record that corresponded to an already included published study was excluded from the trial dataset to avoid duplicate counting of the same evidence.

Data extraction

Data extraction was performed by two independent reviewers using a standardized, piloted form. Any discrepancies between reviewers were resolved through consensus or, when necessary, by adjudication from a third reviewer. The extracted data include: (1) Study characteristics: study design, sample size, target medical condition, participant demographics, and details on randomization, blinding, and allocation concealment; (2) Intervention details: the specific LLM used (e.g., ChatGPT, Gemini) and the clinical task it was applied to; (3) Outcomes: all primary, secondary outcomes and adverse events recorded in the studies.

For all the above screening and extraction processes, two researchers independently screened clinical trials and publications eligible for inclusion and linked clinical trials with their corresponding publications. We calculated the Cohen’s Kappa index to evaluate the level of agreement between the two researchers during the screening process.

Definitions of categorization

Interventions were classified as either LLM-only or LLM-assisted based on the degree of human involvement in the clinical workflow. LLM-only interventions were defined as those in which the model operated autonomously, generating outputs directly for clinical use without human moderation or override. LLM-assisted interventions were defined as human-AI collaborative systems, where the LLM served as a decision support tool and its outputs were interpreted, verified, or integrated by a clinician before influencing patient care.

The primary clinical task addressed by each intervention was categorized into four mutually exclusive groups, borrowed ideas from the framework proposed by Han et al.25: (1) clinical decision support: interventions aiding diagnosis, treatment planning, or risk assessment; (2) symptom intervention: interventions targeting patient symptom management, psychological support, or behavioural modification; (3) information query: interventions designed to answer patient or clinician questions or provide clinical information; and (4) others: tasks not falling into the above categories (e.g., administrative task).

Drawing on prior work characterizing AI evaluation in healthcare, we classified outcomes using a four-tier framework designed to reflect the translational pipeline: (1) Functional performance: Outcomes assessing the technical capabilities of LLMs, including accuracy, sensitivity, specificity, calibration, and concordance with expert judgment; (2) Operational or process impact: Outcomes evaluating the effect of LLM interventions on clinical workflows and system efficiency, such as decision time, documentation efficiency, clinician workload, and cost-related measures; (3) Health endpoints: Outcomes directly reflecting patient health status, including symptoms, patient-reported outcomes, quality of life, adverse events, and physiological or biomarker measures; (4) Others: Outcomes that cannot be categorised as the previous three types of outcomes. Such as trust of AI, user experience, satisfaction, feasibility of LLM in practice, etc.

To ensure consistency, outcomes labelled heterogeneously across primary studies were mapped into these tiers using pre-specified rules, with any ambiguous cases resolved through consensus discussion. For outcomes related to accuracy and time efficiency, we extracted data only when studies provided explicit operational definitions. Accuracy required a clear reference standard against which model outputs were compared; time efficiency required objective measurement in standardized units. Studies using ambiguous or unverifiable outcome definitions were excluded from these specific analyses to maintain interpretability. Furthermore, the alignment of categorizations was independently checked across all studies by two reviewers, with any discrepancies resolved through team discussion.

Reporting quality assessment

The quality of reporting for included studies was assessed using established guidelines. For randomized controlled trials (RCTs), we used the CONSORT-AI checklist15. For studies focusing on AI-driven chatbots, we additionally applied the CHART guideline17. Two investigators independently reviewed each included study against the respective checklists. Disagreements were discussed until consensus was reached, with recourse to a third reviewer if needed. Adherence to the guidelines was quantified in two ways: per study (calculated as the percentage of applicable items fully addressed) and per checklist item (calculated as the percentage of studies that adequately reported that item).

Risk of bias assessment

We conducted a formal risk-of-bias assessment for all included studies. For randomized controlled trials, we used the Cochrane Risk of Bias 2 (RoB 2) tool26, which evaluates bias across five domains: the randomization process, deviations from intended interventions, missing outcome data, measurement of the outcome, and selection of the reported result. For non-randomized studies of interventions, we applied the Risk of Bias in Non-randomized Studies of Interventions (ROBINS-I) tool27, which assesses bias across seven domains: confounding, selection of participants, classification of interventions, deviations from intended interventions, missing data, measurement of outcomes, and selection of the reported result. Each domain was rated as “low risk,” “some concerns,” or “high risk” of bias, following the standardized algorithms and signalling questions provided in the respective guidance documents.

Data synthesis

Data synthesis was performed using Stata software (version 17.0). Descriptive statistics were used to summarize the distribution of studies and outcomes across the predefined categories. The evidence mapping results were visualized using Kumu (https://kumu.io/), a platform for organizing complex data into relationship maps.

Supplementary information

Supplementary Information (770.5KB, pdf)

Acknowledgements

This study was funded by the National Natural Science Foundation of China (Project number 72074006 to J.D.), the National Key R&D Program for Young Scientists (Project number 2022YFF0712000 to J.D.), and Michigan Medicine-PKUHSC Joint Institute (JI) Award (BMU2026JIUM001). The funders had no role in any aspect of this study.

Author contributions

The contributions made by the individual authors are as follows. Z.H. and J.D. conceptualised and designed the study. Z.H., L.Y., Z.L. and X.L. are responsible for data collection, data cleaning, and data extraction. Z.H. conducted the analyses, and J.D. contributed to the interpretation of results. The initial manuscript was drafted by Z.H., amendments are made by Z.H. and J.D. All authors have read and agreed to the published version of the manuscript.

Data availability

All data are publicly accessible via the PubMed (MEDLINE), Embase, Web of Science, ClinicalTrial.gov, and WHO ICTRP databases. This study contains no individual participant data. Clinical trial number is not applicable.

Code availability

The search strategy for this study is reported open access in Supplementary Note 1.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Supplementary information

The online version contains supplementary material available at https://doi.org/10.1038/s41746-026-02837-6.

References

  • 1.Goh, E. et al. Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Netw. Open7, e2440969 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Goh, E. et al. GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial. Nat. Med.31, 1233–1238 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Shi, X. et al. A Bibliographic Dataset of Health Artificial Intelligence Research. Health Data Sci.4, 0125 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Yang, K. et al. ECG-LM: Understanding Electrocardiogram with a Large Language Model. Health Data Sci.5, 0221 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Hager, P. et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat. Med.30, 2613–2622 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Bedi, S. et al. Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review. Jama333, 319–328 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Bean, A. M. et al. Reliability of LLMs as medical assistants for the general public: a randomized preregistered study. Nat. Med.32, 609–615 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.The Lancet Digital H. Rapid generative AI rollout in health care. Lancet Dig. Health7, 100909 (2025). [DOI] [PubMed]
  • 9.Moëll, B., Sand Aronsson, F. & Akbar, S. Medical reasoning in LLMs: an in-depth analysis of DeepSeek R1. Front Artif. Intell.8, 1616145 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Suri, G., Slater, L. R., Ziaee, A. & Nguyen, M. Do large language models show decision heuristics similar to humans? A case study using GPT-3.5. J. Exp. Psychol. Gen.153, 1066–1075 (2024). [DOI] [PubMed] [Google Scholar]
  • 11.The COMET Initiative. Core Outcome Measures in Effectiveness Trials. (2025).
  • 12.Cai, X., Wang, B., Cheng, L. & Ou, Q. Comments on “ChatGPT’s role in alleviating anxiety in total knee arthroplasty consent process: a randomized controlled trial pilot study. Int J. Surg.111, 4115–4116 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Zhang, C. & Chen, X. Letter to the editor, “Evaluating the accuracy of ChatGPT-4 in predicting ASA scores: A prospective multicentric study ChatGPT-4 in ASA score prediction. J. Clin. Anesth.98, 111571 (2024). [DOI] [PubMed] [Google Scholar]
  • 14.Zixuan, H., Lan, Y., Xiaofan, L. & Jian, D. Discrepancies in reported results between trial registries and journal articles for AI clinical research. eClinicalMedicine80, 103066 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Liu, X., Cruz Rivera, S., Moher, D., Calvert, M. J. & Denniston, A. K. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat. Med.26, 1364–1374 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Cruz Rivera, S., Liu, X., Chan, A. W., Denniston, A. K. & Calvert, M. J. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. Nat. Med.26, 1351–1363 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.CHART Collaborative. Reporting guideline for chatbot health advice studies: the Chatbot Assessment Reporting Tool (CHART) statement. BMJ Med.4, e001632 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Martindale, A. P. L. et al. Concordance of randomised controlled trials for artificial intelligence interventions with the CONSORT-AI reporting guidelines. Nat. Commun.15, 1619 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Pattathil, N., Zhao, J. Z. L., Sam-Oyerinde, O. & Felfeli, T. Adherence of randomised controlled trials using artificial intelligence in ophthalmology to CONSORT-AI guidelines: a systematic review and critical appraisal. BMJ Health Care Inform.30, e100757 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Heinz, M. V. et al. Randomized Trial of a Generative AI Chatbot for Mental Health Treatment. NEJM AI2, AIoa2400802 (2025). [Google Scholar]
  • 21.Show us the evidence for the value of medical AI. Nat. Med.32, 1163 (2026). [DOI] [PubMed]
  • 22.Zöller, N. et al. Human-AI collectives most accurately diagnose clinical vignettes. Proc. Natl. Acad. Sci. USA.122, e2426153122 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Liu, P., Zhang, J. & Chen, S. Human-AI teaming in healthcare: 1 + 1 > 2? npj Artif. Intell.1, 47 (2025).
  • 24.Page, M. J. et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. Bmj372, n71 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Han, R. et al. Randomised controlled trials evaluating artificial intelligence in clinical practice: a scoping review. Lancet Digit Health6, e367–e373 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Sterne, J. A. C. et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. Bmj366, l4898 (2019). [DOI] [PubMed] [Google Scholar]
  • 27.Sterne, J. A. et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. Bmj355, i4919 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Harari, R. E., Altaweel, A., Ahram, T., Keehner, M. & Shokoohi, H. A randomized controlled trial on evaluating clinician-supervised generative AI for decision support. Int. J. Med. Inform.195, 105701 (2025). [DOI] [PubMed] [Google Scholar]
  • 29.Gan, W. et al. Integrating ChatGPT in orthopedic education for medical undergraduates: randomized controlled trial. J. Med. Int. Res.26, e57037 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Goh E., et al. ChatGPT Influence on Medical Decision-Making, Bias, and Equity: A Randomized Study of Clinicians Evaluating Clinical Vignettes. (2023).
  • 31.Akdogan, O. et al. Effect of a ChatGPT-based digital counseling intervention on anxiety and depression in patients with cancer: A prospective, randomized trial. Eur. J. Cancer221, 115408 (2025). [DOI] [PubMed] [Google Scholar]
  • 32.Vaira, L. A. et al. Accuracy of ChatGPT-Generated Information on Head and Neck and Oromaxillofacial Surgery: A Multicenter Collaborative Analysis. Otolaryngol. - Head. Neck Surg. (U. S.)170, 1492–1503 (2024). [DOI] [PubMed] [Google Scholar]
  • 33.Cheng, T. et al. The performance of ChatGPT in day surgery and pre-anesthesia risk assessment: a case-control study of 150 simulated patient presentations. Perioperat. Med.13, 111 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Turan, E., Baydemir, A. E., Özcan, F. G. & Şahin, A. S. Evaluating the accuracy of ChatGPT-4 in predicting ASA scores: A prospective multicentric study ChatGPT-4 in ASA score prediction. J. Clin. Anesth.96, 111475 (2024). [DOI] [PubMed] [Google Scholar]
  • 35.Turan, E. I., Baydemir, A. E., Şahin, A. S. & Özcan, F. G. Effectiveness of ChatGPT-4 in predicting the human decision to send patients to the postoperative intensive care unit: a prospective multicentric study. Minerva Anestesiol (2025). [DOI] [PubMed]
  • 36.Sanduleanu, S. et al. Feasibility of GPT-3.5 versus Machine Learning for Automated Surgical Decision-Making Determination: A Multicenter Study on Suspected Appendicitis. Ai5, 1942–1954 (2024). [Google Scholar]
  • 37.Zhou, B. Y. et al. The Large Language Model Improves the Diagnostic Performance of Suspicious Breast Lesions by Radiologists Using Grayscale Ultrasound: A Multicenter Cohort Study. (2025).
  • 38.Bayala, Y. L. T. et al. Performance of the Large Language Models in African rheumatology: a diagnostic test accuracy study of ChatGPT-4, Gemini, Copilot, and Claude artificial intelligence. BMC Rheumatol.9, 54 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Gurbuz, S. et al. Comparative Efficacy of ChatGPT and DeepSeek in Addressing Patient Queries on Gonarthrosis and Total Knee Arthroplasty. Arthroplast Today33, 101730 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

All data are publicly accessible via the PubMed (MEDLINE), Embase, Web of Science, ClinicalTrial.gov, and WHO ICTRP databases. This study contains no individual participant data. Clinical trial number is not applicable.

The search strategy for this study is reported open access in Supplementary Note 1.


Articles from NPJ Digital Medicine are provided here courtesy of Nature Publishing Group

RESOURCES