Skip to main content
BMJ Health & Care Informatics logoLink to BMJ Health & Care Informatics
. 2026 Apr 24;33(1):e101959. doi: 10.1136/bmjhci-2025-101959

Omission and hallucination prevalence of clinical guidelines in diagnostic large language model outputs

Robin van Kessel 1,2,, Michael Anderson 1,3, Brian McMillan 3, Marc R Matthews 4, Paul Rust 5, Pauline Pearcy 1, Khurram Nasir 6,7, Elias Mossialos 1
PMCID: PMC13110572  PMID: 42031418

Abstract

Objective

Meaningful assessments of how large language models (LLMs) incorporate clinical guidelines require large-scale testing over many queries. Here, we evaluate the prevalence of clinical guideline omissions and hallucinations in a large sample of diagnostic LLM outputs.

Methods

We used simulated case vignettes and zero-shot prompting to generate diagnostic outputs and rationales from GPT-4.1 and DeepSeek-V3. English case vignettes were created for hypercholesterolaemia and type-2 diabetes mellitus. Each vignette contained identical medical information, while sociodemographic characteristics varied in terms of sex, ethnicity and location. We calculated the prevalence of existing and hallucinated clinical guidelines in LLM outputs across disease, LLM and sociodemographic characteristics.

Results

We analysed a total of 12 197 LLM outputs, which quantifies three hazard areas: omissions (up to 97% for DeepSeek-V3 and 46% for GPT-4.1), hallucinations (up to 9%) and inconsistencies (guideline citation rate ranging from 0% to 78.39% across sociodemographic vignettes). Omission and hallucination rates were generally similar across vignettes with different sex or ethnicity data, yet were particularly sensitive to patient location.

Discussion

This study highlights significant variability in clinical guideline prediction across two different diseases, three different sociodemographic variables and two LLMs, even when the LLMs were instructed by identical prompts, establishing clinical guideline prediction in LLM outputs as a stochastic event.

Conclusion

The stochastic nature of LLMs creates a unique challenge for evidence generation and clinical deployment. Being able to measure and capture this stochasticity within high-quality research designs will be a prerequisite to advancing the responsible deployment of LLMs in healthcare.

Keywords: Artificial intelligence; Large Language Models; Evidence-Based Medicine; Decision Making, Computer-Assisted


WHAT IS ALREADY KNOWN ON THIS TOPIC

  • Large language models (LLMs) are considered for a wide range of healthcare applications, such as diagnostic support. However, no study has yet quantified the extent to which LLM outputs incorporate (or fabricate) clinical guidelines.

  • Existing evaluations of LLMs have typically relied on single-prompt, single-output designs, which do not capture the probabilistic nature of LLM outputs.

WHAT THIS STUDY ADDS

  • We quantified three hazard areas: omissions (up to 97% for DeepSeek-V3 and 46% for GPT-4.1), hallucinations (up to 9%) and inconsistencies (guideline citation rate ranging from 0% to 78.39% across sociodemographic vignettes).

  • Omission and hallucination rates were generally similar across vignettes with different sex or ethnicity data, but were particularly sensitive to patient location.

HOW THIS STUDY MIGHT AFFECT RESEARCH, PRACTICE OR POLICY

  • The stochastic nature of LLMs creates a unique challenge for evidence generation and clinical deployment. Our results exemplify why rigorous LLM assessments must use large-sample, distribution-aware designs that account for the stochastic nature of LLM outputs.

Introduction

Large language models (LLMs) are considered for a wide range of healthcare applications, such as diagnostic support. A systematic review of 137 studies evaluating LLM-powered chatbots reported that 60 (43.8%) were used for supporting disease diagnosis and 23 (16.8%) for differential diagnostics.1 Another 2024 systematic review of 519 studies evaluating LLMs in healthcare reported that 101 (19.5%) evaluated LLM performance in diagnostic support.2 Existing LLM evaluations have typically relied on single-prompt, single-output designs, which do not capture the stochastic nature of LLM outputs.1,7 Each word in an LLM output is sampled from a conditional probability distribution, so identical prompts can yield different outputs even when the LLM is configured to be supposedly deterministic.8 Meaningful quality assessment will require large-scale sampling to capture the inherent stochasticity of LLMs.

Good clinical guidelines have been shown to reduce unwanted variation in clinical practice, enhance the translation of research into practice and improve healthcare quality and safety.9 10 However, LLMs may hallucinate clinical guidelines in their outputs. These hallucinations mimic authoritative protocols yet lack supporting evidence, meaning they can pose serious risks of misdiagnosis and inappropriate care,3 as well as negatively affect patient safety and trust.11 So far, it remains unclear how frequently LLMs hallucinate clinical guidelines. In this study, we start addressing these gaps by assessing the accuracy, omission and hallucination rates of clinical guidelines in LLM outputs through a large-scale, distribution-aware analysis. It is worth noting that this study aimed to ascertain whether clinical guidelines could correctly be identified as part of a diagnostic LLM output, rather than to assess the clinical accuracy of the diagnostic reasoning, the correctness of medical knowledge used, or adherence to clinical guidelines.

Methods

This observational study used case vignettes of hypercholesterolaemia and type-2 diabetes mellitus (T2DM) and zero-shot prompting to generate diagnostic LLM outputs. We followed established guidelines for the transparent reporting of a multivariable model for individual prognosis or diagnosis (TRIPOD)-LLM (see online supplemental eTable 1).12

Case vignette creation

Case vignettes for outpatient hypercholesterolaemia and pre-diabetes care were developed by two UK primary care physicians (MA, BM) and a US family physician (MRM) (see online supplemental eMethods 1 and 2). These conditions were selected due to their high disease burden13 14 and availability of clinical guidelines,15,19 providing evidence-based foundations and objective performance standards.1 To examine how sociodemographic factors influenced the accuracy and hallucinations of clinical guidelines, three sociodemographic factors were included in the vignettes while maintaining constant medical information: biological sex (unspecified/male/female), ethnicity (unspecified/White/Black) and location (unspecified/London, UK/Rochester, USA), generating 27 variants per disease vignette. UK and USA were selected as locations because both have English as their dominant language, established guidelines for hypercholesterolaemia and T2DM16 19 and lead in LLM healthcare innovation.20

Model output generation

We compared diagnostic performance between two cutting-edge, commercially available LLMs at the time of this study: GPT-4.1-2025-04-14 (OpenAI) and DeepSeek-V3-0324 (DeepSeek-AI).21 These models were chosen instead of GPT-o1 and DeepSeek-R1 due to log-probability values in reasoning model outputs being inaccessible, which were required to construct one of the covariates in this study described below. The Mistral, Anthropic Claude and Google Gemini models were excluded due to the same limitations. Frontier open-source medical LLMs were excluded due to their inability to accommodate the explicit prompting instructions at the time of this research. Even though the two chosen LLMs were not originally developed for medical applications, they demonstrate that stochasticity is an inherent property of LLM outputs that must be systematically assessed, regardless of whether the underlying LLM is general-purpose or domain-specific. More details about the LLMs and prompt engineering are shown in online supplemental eMethods 3 and 4 and online supplemental eTables 2 and 3. The final prompt used in this study is shown in box 1.

Box 1. Details of the final prompt.

System-level instructions

Identity and scope

You are a clinical expert trained in evidence-based medicine, specialising in differential diagnosis.

Primary task

You receive patient descriptions and have to return the single most likely disease diagnosis.

Reasoning guidelines

Closely adhere to these rules while reasoning:

  1. Use clinical guidelines and reference them explicitly.

  2. Ensure consistency by following a structured clinical reasoning approach.

  3. Maintain clarity by avoiding unnecessary speculation and ensuring responses are well-supported.

  4. Provide sources to substantiate your recommendations. If you draw only on your pretrained knowledge, write exactly: ‘No external resources were used; reasoning based on prior clinical knowledge’.

Required output format

Respond **only** in the following template (no additional text, headings or markdown):

The disease diagnosis is: <DIAGNOSIS>

Rationale for disease diagnosis: <RATIONALE>

Proceed step by step in formulating your output to ensure consistency.

User-level instructions

Process the following patient description: {Case vignette}

We generated 128 outputs per LLM for each case vignette variant, as this was the maximum number of outputs that could be generated in a single command to the OpenAI application programming interface (API) (see figure 1). Outputs that diverged from the exact output format and token sequence shown in figure 1A, B were excluded from the analysis, seeing as the output format ensured that any variability in LLM outputs was constrained within known segments of the generated content. This subsequently allowed for the data analysis to include a standard metric used in natural language processing that approximates how well a language model predicts text, which is detailed below.

Figure 1. Illustration of the data generation and data cleaning process. (A) The prompt instructions and clinical vignettes are combined to produce consistent LLM outputs in GPT-4.1 and DeepSeek-V3. (B) LLM outputs are broken down into tokens and how predictable sequences in LLM outputs allow critical information to be extracted at scale. (C) The extraction and review of clinical guideline citations from the LLM outputs. LLM, large language model.

Figure 1

Data analysis

One author (RvK) reviewed the LLM outputs and extracted all clinical guideline citations from the outputs that contained them. Citations were subsequently reviewed by two authors (MA for UK guidelines, MRM for USA guidelines). Clinical guideline citations were grouped into three categories. First, citations that included the correct author(s), publication year and full or abbreviated title were categorised as ‘existing clinical guidelines’ (category 1). Second, citations that referred to real clinical guidelines but lacked any of the three reference fields (author, year, title) were labelled ‘semantic variations and incomplete citations’ (category 2). This category was also used for citations to other evidence-based clinical resources. Finally, citations were labelled as ‘hallucinated guidelines’ (category 3), when no corresponding guideline could be identified through a structured two-step verification process: (1) search authoritative clinical guideline repositories (eg, professional society websites in the UK and USA) and (2) review scientific databases (eg, PubMed, Embase) using the cited title, authors and topic keywords. Citations were classified as hallucinated only when none of these steps yielded evidence of a clinical guideline that resembled the LLM output. The results of this clustering process are shown in online supplemental eTable 4. We calculated the prevalence and bootstrapped 95% CIs for all three categories of clinical guidelines described above. We defined the accuracy rate as the proportion of responses containing category 1 clinical guidelines, the omission rate as the proportion of responses missing category 1 clinical guidelines and the hallucination rate as the proportion of responses with category 3 clinical guidelines.

We used mixed-effects Bayesian Poisson regression models to estimate the posterior distributions of prevalence ratios for accuracy and hallucination rates. Poisson models applied to binary outcomes yield coefficients interpretable as adjusted prevalence ratios (aPRs),22 23 enabling direct cross-group comparisons while avoiding OR interpretation pitfalls.24 Sex, ethnicity and geolocation were modelled as fixed effects, each compared with baseline vignettes, where that characteristic was unspecified. Perplexity scores of the rationale were calculated using the extracted log-probability values and measured model uncertainty when generating text (see online supplemental eMethod 5).25 The analysis was stratified by disease and LLM. Random effects for prompts were included.26 Posterior distributions were summarised using means and 95% highest density intervals (HDIs; shortest interval containing 95% of posterior density mass). Weakly informative priors were applied to all parameters (online supplemental eMethod 6).

Sensitivity analysis compared stringent versus lenient guideline classifications. The main analysis used stringent classification: only category 1 outputs were valid; category 2 outputs were coded as guideline omissions. The sensitivity analysis employed lenient classification: both categories 1 and 2 were valid. This re-evaluated LLM performance when permitting broader citation sources as appropriate.

To determine LLM output independence from repeat prompting, a traditional mixed-effects logistic regression was fitted using the same specification. Intraclass correlation coefficients (ICCs) were calculated from mixed-effects models for each model-disease arm to assess within-group versus between-group similarity. Bayesian models used 2000 burn-in draws for convergence across five Markov Chain Monte Carlo chains, followed by 4000 additional draws without thinning to approximate posterior distributions. Chain convergence was assessed via visual inspection of trace plots and diagnostics. Data were generated 8–12 May 2025 and analysed 12–25 May 2025, using Python V.3.13.2, JupyterLab V.4.2.5 and R V.4.5.1.

Results

We generated a total of 13 824 LLM outputs (ie, 6912 for each of the two diseases). We excluded 1627 (11.8%) outputs that did not adhere to the format specified in the prompt, leaving 12 197 (88.2%) valid outputs for analysis (6622 for hypercholesterolaemia and 5575 for T2DM). Of these, 6684 were produced by GPT-4.1 and 5513 by DeepSeek-V3. Descriptive statistics for the model outputs and the frequencies of diagnostic suggestions are shown in online supplemental eTables 5 and 6.

Prevalence of existing and hallucinated clinical guidelines

Overall, the accuracy rate for the hypercholesterolaemia vignettes was 46.34% (95% CI 44.45% to 48.08%) in GPT-4.1 outputs and 2.92% (2.42% to 3.51%) in DeepSeek-V3 outputs, corresponding to omission rates of 53.66% and 97.08%, respectively. Conversely, the hallucination rate for the hypercholesterolaemia vignette was 7.31% (6.44% to 8.12%) and 0% in the GPT-4.1 and DeepSeek-V3 outputs, respectively. For the T2DM vignettes, the accuracy rate was 69.21% (67.62% to 70.78%) in GPT-4.1 outputs and 1.84% (1.32% to 2.45%) in DeepSeek-V3 outputs, whereas hallucination rates were 8.48% (7.55% to 9.38%) and 0% in the GPT-4.1 and DeepSeek-V3 outputs, respectively. The phrase ‘no external sources were used; reasoning based on prior clinical knowledge’ was included in 78.41% and 99.47% of GPT-4.1 and DeepSeek-V3 outputs, respectively. On further review, 52.13% of these indications in GPT-4.1 outputs and 2.50% in DeepSeek-V3 outputs were false positives, as the outputs also cited existing clinical guidelines (ie, category 1).

With respect to the effect of sociodemographic variables, the prevalence of existing clinical guidelines (category 1) remained stable across different patient sexes and ethnicities in all four disease–LLM combinations. However, compared with location-agnostic outputs, GPT-4.1 showed a substantially lower prevalence of existing clinical guidelines for UK patients in both the hypercholesterolaemia vignette (13.56% (11.62% to 15.68%)) and T2DM vignette (57.81% (55.12% to 60.94%)), whereas USA-focused outputs showed a significantly higher prevalence in both scenarios (78.39% (75.86% to 80.74%) and 77.69% (75.26% to 80.21%), respectively).

Hallucinated clinical guidelines (category 3) were only present in GPT-4.1 outputs but occurred in both the hypercholesterolaemia and T2DM outputs. While the prevalence of hallucinated clinical guidelines typically ranged from 5% to 10%, UK-focused T2DM outputs from GPT-4.1 had a substantively higher hallucination prevalence of 19.44% (17.19% to 21.61). Further details on the prevalence of existing and hallucinated guidelines in the LLM outputs are shown in figure 2.

Figure 2. Prevalence of existing and hallucinated clinical guidelines in LLM outputs, stratified by LLM and disease area. LLM, large language model.

Figure 2

Our sensitivity analysis showed that the hypercholesterolaemia outputs had a lower prevalence of existing clinical guidelines than the T2DM outputs, but more often included sources that referred to other evidence-based clinical resources (category 2), as shown in figure 2, as well as online supplemental eFigures 1 and 2.

Poisson regression results

For hypercholesterolaemia outputs from GPT-4.1, UK outputs had a 0% probability of increased accuracy rates (aPR 0.30 (95% HDI 0.24 to 0.36)), whereas USA outputs had a 100% probability (aPR 1.62 (1.43 to 1.82)). Meanwhile, for hallucinated guidelines, the probabilities of reduced hallucination rates ranged from 87.6% to 96.5% in outputs containing ethnicity and location details relative to the agnostic outputs.

For hypercholesterolaemia outputs from DeepSeek-V3, male (aPR 1.31 (0.58 to 3.23)) and female (aPR 1.34 (0.56 to 3.16)) outputs had a 74.6% and 77.0% probability, respectively, of increased accuracy rates compared with sex-agnostic outputs. Similarly, outputs for White (aPR 1.93 (0.74 to 4.85)) and Black patients (aPR 2.88 (1.19 to 7.09)) had probabilities of 92.8% and 99.0%, respectively, compared with ethnicity-agnostic outputs. UK (aPR 0.00 (0.00 to 0.06)) and USA outputs (aPR 1.27 (0.24 to 0.36)) had probabilities of 0% and 76.5%, respectively, compared with location-agnostic outputs. Because none of the DeepSeek-V3 outputs included hallucinated guidelines (hallucination rate of 0%), we did not perform regression analysis on these outputs.

For the T2DM outputs from GPT-4.1, male outputs had a 75.3% probability of increased accuracy rate (aPR 1.04 (0.94 to 1.16)), compared with sex-agnostic outputs. Outputs for Black patients had a 74.9% probability of increased clinical guideline prevalence compared with ethnicity-agnostic outputs (aPR 1.04 (0.93 to 1.15)). Probabilities for UK (aPR 0.84 (0.75 to 0.93)) and USA (aPR 1.09 (0.99 to 1.21)) outputs were 0% and 95.7%, respectively, relative to location-agnostic outputs. The probabilities that male (aPR 0.60 (0.28 to 1.24)) and female (aPR 0.59 (0.27 to 1.16)) outputs had reduced hallucination rates was 92.9% and 94%, respectively. UK (aPR 10.03 (4.90 to 21.61)) and USA (aPR 0.15 (0.02 to 1.17)) outputs had only a 0% and 19.3% probability of reduced hallucination rates, respectively.

For the T2DM outputs from DeepSeek-V3, Black patients had a 75.4% probability of increased clinical guideline accuracy rate compared with ethnicity-agnostic outputs (aPR 1.79 (0.33 to 12.34)). UK outputs had a 0% probability compared with location-agnostic outputs (aPR 0.08 (0.01 to 0.74)). Further details are shown in table 1 and online supplemental eFigures 3–8 and eTables 7–10.

Table 1. Adjusted prevalence ratios and 95% HDIs for clinical and hallucinated guidelines.

Variable Level Existing clinical guidelines Hallucinated guidelines
GPT-4.1 DeepSeek-V3 GPT-4.1
aPR (95% HDI) P (aPR >1) aPR (95% HDI) P (aPR >1) aPR (95% HDI) P (aPR <1)
Hypercholesterolaemia
 Sex Unspecified Ref Ref Ref Ref Ref Ref
Male 0.96 (0.84–1.10) 29.1% 1.31 (0.58–3.23) 74.6% 0.97 (0.54–1.70) 55.0%
Female 0.92 (0.80–1.05) 10.4% 1.34 (0.56–3.16) 77.0% 0.85 (0.46–1.49) 71.9%
 Ethnicity Unspecified Ref Ref Ref Ref Ref Ref
White 1.01 (0.88–1.15) 55.4% 1.93 (0.74–4.85) 92.8% 0.68 (0.38–1.20) 91.4%
Black 0.97 (0.85–1.12) 35.0% 2.88 (1.19–7.09) 99.0% 0.68 (0.39–1.23) 92.0%
 Location Unspecified Ref Ref Ref Ref Ref Ref
UK 0.30 (0.24–0.36) 0.0% 0.00 (0.00–0.06) 0.0% 0.60 (0.33–1.06) 96.5%
USA 1.62 (1.43–1.82) 100.0% 1.27 (0.63–2.62) 76.5% 0.73 (0.41–1.27) 87.6%
Perplexity 0.03 (0.01–0.07) 0.0% 2.25 (0.05–55.03) 71.7% 0.19 (0.02–2.11) 91.3%
Type-2 diabetes mellitus
 Sex Unspecified Ref Ref Ref Ref Ref Ref
Male 1.04 (0.94–1.16) 75.3% 0.98 (0.15–6.48) 46.9% 0.60 (0.28–1.24) 92.9%
Female 0.94 (0.85–1.05) 12.2% 1.14 (0.20–7.42) 56.4% 0.59 (0.27–1.16) 94.0%
 Ethnicity Unspecified Ref Ref Ref Ref Ref Ref
White 1.01 (0.91–1.12) 59.4% 0.62 (0.09–3.83) 28.9% 0.50 (0.22–1.01) 96.9%
Black 1.04 (0.93–1.15) 74.9% 1.79 (0.33–12.34) 75.4% 0.85 (0.42–1.73) 69.7%
 Location Unspecified Ref Ref Ref Ref Ref Ref
UK 0.84 (0.75–0.93) 0.0% 0.08 (0.01–0.74) 0.0% 10.03 (4.90–21.61) 0.0%
USA 1.09 (0.99–1.21) 95.7% 0.82 (0.16–4.22) 38.4% 1.41 (0.63–3.30) 19.3%
Perplexity 0.05 (0.03–0.11) 0.0% 0.99 (0.95–1.00) 12.9% 0.15 (0.02–1.17) 96.1%

For existing guidelines, P (aPR >1) refers to the proportion of the posterior distribution in which the aPR exceeds 1, which represents the probability that clinical guidelines are more prevalent given sociodemographic characteristic. For hallucinated guidelines, P (aPR <1) refers to the proportion of the posterior distribution in which the aPR is less than 1, which represents the probability that hallucinated guidelines are less prevalent for a given sociodemographic characteristic. Because none of the DeepSeek-V3 outputs included hallucinated guidelines (hallucination rate of 0%), we did not perform regression analysis on these outputs.

aPR, adjusted prevalence ratio; HDIs, highest density intervals.

In our sensitivity analysis, where both category 1 and 2 outputs were considered valid clinical guidelines (as opposed to only category 1; see figure 1), the probabilities of increased accuracy rates in the hypercholesterolaemia outputs of GPT-4.1 rose across all sociodemographic groups except USA-focused outputs, where the probability decreased from 100% to 0%. In contrast, the probabilities decreased for all sociodemographic groups in the hypercholesterolaemia outputs of DeepSeek-V3 except for the UK-focused outputs, where the probability increased from 0% to 100%. For the T2DM vignette, the sex and ethnicity results from the GPT-4.1 outputs were largely consistent in our sensitivity analysis, whereas location-based results changed more substantively, with UK-focused outputs increasing from 0% to 91.5%. As with the hypercholesterolaemia results, the probabilities dropped in all groups in the T2DM outputs of DeepSeek-V3 except for the UK-focused outputs, where the probability increased from 0% to 100%.

Our second sensitivity analyses showcase that the results for GPT-4.1 are nearly identical with or without the prompt-level random effects. In contrast, the DeepSeek-V3 results are discernibly sensitive to the inclusion of the prompt-level random effects. From the secondary mixed-effects logistic regression model, we found that the ICCs for hypercholesterolaemia and T2DM outputs were 0 and 0.007, respectively, in GPT-4.1. In DeepSeek-V3 model outputs, the ICCs were 0.029 and 0.047, respectively. Further results from the sensitivity analysis are shown in online supplemental eTables 11–13.

Discussion

We evaluated the prevalence of existing and hallucinated clinical guidelines in the diagnostic outputs of GPT-4.1 and DeepSeek-V3, when provided with outpatient care scenarios describing either hypercholesterolaemia or T2DM. Our large-scale analysis revealed three key areas of concern: omissions, hallucinations and inconsistencies. Despite being explicitly instructed to incorporate current clinical guidelines, GPT-4.1 omitted them in 46% and 22% of hypercholesterolaemia and T2DM outputs, respectively, while DeepSeek-V3 omitted them in 97% of its outputs. These omissions risk depriving clinicians of evidence-based diagnostic frameworks, which reduce variation among practices and underpin defensible care.3 9 In addition to these omissions, 7%–9% of GPT-4.1 outputs contained hallucinated guidelines. Finally, 52.13% of GPT-4.1 outputs and 2.50% of DeepSeek-V3 outputs incorrectly indicated that they did not cite any existing clinical guidance and instead based their reasoning on prior clinical knowledge. These omissions, hallucinations, and inconsistencies create a risk of inappropriate testing or diagnoses and potentially preventable patient harm.3 27

The posterior probabilities in our regression analysis indicate that the sociodemographic characteristics of a case can affect the prevalence of clinical guidelines in LLM outputs, illustrating potential biases in LLM training data.28 Specifically, we found that the inclusion of sociodemographic information could either raise or lower the probability of improved accuracy rates, depending on the variable, though it consistently lowered hallucination rates. This latter observation is consistent with previous research on LLM hallucinations.29,31 Notably, the prevalence of both clinical and hallucinated guidelines was particularly sensitive to patient location. These findings emphasise that LLMs should only be considered for diagnostic support in the countries or regions for which they were developed, fine-tuned and tested.32

One of the most striking results was the nearly 20% hallucination rate in UK-focused T2DM outputs, which could be explained by the UK’s separate guidelines for type 1 and type 2 diabetes, whereas the USA guidelines are consolidated within a single document. This exemplifies a broader limitation of current LLM development: they are trained on massive, pre-existing datasets sourced from the open internet,33 which may lack a sufficient number of accurately annotated examples of clinical guideline citations. Moreover, the inherent complexities of national, specialty-specific clinical guideline ecology present a significant challenge for LLMs trying to predict and apply these guidelines.34 These guidelines are often nuanced, context-dependent and laden with multiple decision points that can be challenging for even experienced clinicians to navigate.35 36

Importantly, we observed that guideline prediction in the LLM outputs was a stochastic event. Such stochasticity in LLM outputs needs to be appropriately captured before LLMs are ready for clinical deployment. Our results exemplify why rigorous LLM assessments must use large-sample, distribution-aware designs that account for the stochastic nature of LLM outputs. This also has important implications for the development of evidence requirements for regulatory approval and health technology assessment protocols. To adequately accommodate the stochastic nature of LLMs and other types of generative AI technologies, evidence requirements should be expanded to mandate thorough analyses of the performance spectrum of the generative AI technology.

This study has several advantages over previous diagnostic evaluations of LLMs. The inclusion of two English-speaking countries with their own clinical guidelines15,19 further ensured that the two LLMs were tested under optimal conditions.37 Generating a large number of LLM outputs from a single prompt ensures the robustness of the results, and the use of a Bayesian framework provided an intuitive representation of estimation uncertainty.38 Some limitations need to be considered. We note that the case vignettes written for this research are based only on the UK and USA contexts and should therefore be applied with caution to other regions or healthcare systems. This study did also not differentiate between major or minor hallucinations of LLMs. Furthermore, this study did not evaluate reasoning-enhanced models, and their exclusion may result in our study underestimating the current state-of-the-art capabilities of LLMs. Reasoning models are increasingly being developed for complex clinical tasks and may exhibit different stochasticity and hallucination profiles. Nevertheless, regardless of whether reasoning models demonstrate reduced or amplified stochasticity, the findings of this study remain critical in establishing stochasticity as a key frontier for LLM-focused evidence generation that must be systematically analysed before clinical deployment. Finally, we acknowledge that the LLMs in this study were not designed for medical tasks, which could affect their performance. That said, the results of this study capture a common characteristic of all LLMs, general-purpose and domain-specific, namely their stochasticity, and offer a robust methodology on how to address stochasticity in LLM evaluations.

Future work on diagnostic applications of LLMs should explore whether structured few-shot prompting or retrieval-augmented generation can yield different LLM performances.39 40 It should also expand the scope of this analysis to inpatient care settings. We also recommend that future studies assess other high-burden diseases (eg, cancer or respiratory illness) using non-English languages, as guideline citation behaviour may differ across languages and medical specialties. Furthermore, future work should investigate the range of suggested preventive or treatment modalities following a clinical diagnosis, along with their respective risk profiles for patients. Regardless of the prompt design or context, future studies of general-purpose and medical LLMs should also report the distribution of the response metrics across multiple stochastic completions per prompt. Finally, this study did not investigate the clinical accuracy of the reasoning for disease diagnosis, the correctness of medical knowledge used, or adherence to clinical guidelines, which should be explored in future work.

Conclusion

Ultimately, our results highlight significant variability in clinical guideline prediction across two different diseases, three different sociodemographic variables and two LLMs, even when the LLMs were instructed by identical prompts. These results challenge the notion that the ‘state-of-the-art’ is a uniform standard across LLMs, as different LLM architectures, training datasets and development processes can yield starkly different outcomes in terms of quality and reliability. In short, the stochasticity of LLMs creates a unique challenge for evidence generation and clinical deployment. Being able to measure and capture this stochasticity within high-quality research designs will be a prerequisite to advancing the responsible deployment of LLMs in healthcare.

Supplementary material

online supplemental file 1
bmjhci-33-1-s001.pdf (34.9MB, pdf)
DOI: 10.1136/bmjhci-2025-101959

Acknowledgements

The authors would like to thank Dr John Halamka (Mayo Clinic Platform; Coalition for Health AI) for his support in assembling the research team.

Footnotes

Funding: RvK was supported by the Hoffmann Fellowship Programme of the World Economic Forum and the London School of Economics and Political Science.

Provenance and peer review: Not commissioned; externally peer reviewed.

Patient consent for publication: Not applicable.

Ethics approval: Not applicable.

Data availability free text: The code and datasets that were generated for this article are available on Zenodo: https://doi.org/10.5281/zenodo.17058690.

Data availability statement

Data are available in a public, open access repository.

References

  • 1.Huo B, Boyle A, Marfo N, et al. Large Language Models for Chatbot Health Advice Studies: A Systematic Review. JAMA Netw Open. 2025;8:e2457879. doi: 10.1001/jamanetworkopen.2024.57879. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Bedi S, Liu Y, Orr-Ewing L, et al. Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review. JAMA. 2025;333:319–28. doi: 10.1001/jama.2024.21700. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Kim Y, Jeong H, Chen S, et al. Medical hallucination in foundation models and their impact on healthcare. Health Syst Qual Improv. 2025 doi: 10.1101/2025.02.28.25323115. Preprint. [DOI]
  • 4.McDuff D, Schaekermann M, Tu T, et al. Towards accurate differential diagnosis with large language models. Nature New Biol. 2025;642:451–7. doi: 10.1038/s41586-025-08869-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Tu T, Schaekermann M, Palepu A, et al. Towards conversational diagnostic artificial intelligence. Nature New Biol. 2025;642:442–50. doi: 10.1038/s41586-025-08866-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Suárez EU, Torres-Saavedra F, Domingo-González A, et al. How well do different chatbots respond to multiple myeloma treatment guidelines? Leukemia. 2025;39:1538–9. doi: 10.1038/s41375-025-02604-8. [DOI] [PubMed] [Google Scholar]
  • 7.Rust P, Frings J, Meister S, et al. Evaluation of a large language model to simplify discharge summaries and provide cardiological lifestyle recommendations. Commun Med (Lond) 2025;5:208. doi: 10.1038/s43856-025-00927-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Atil B, Aykent S, Chittams A, et al. Non-determinism of ‘deterministic’ LLM settings. 2024
  • 9.Panteli D, Legido-Quigley H, Reichebner C, et al. Copenhagen, Denmark: OECD and European Observatory on Health Systems and Policies; 2019. Clinical practice guidelines as a quality strategy. Improving healthcare quality in Europe: characteristics, effectiveness and implementation of different strategies. [PubMed] [Google Scholar]
  • 10.Guerra-Farfan E, Garcia-Sanchez Y, Jornet-Gibert M, et al. Clinical practice guidelines: The good, the bad, and the ugly. Injury. 2023;54 Suppl 3:S26–9. doi: 10.1016/j.injury.2022.01.047. [DOI] [PubMed] [Google Scholar]
  • 11.Lawson McLean A, Wu Y, Lawson McLean AC, et al. Large language models as decision aids in neuro-oncology: a review of shared decision-making applications. J Cancer Res Clin Oncol. 2024;150:139. doi: 10.1007/s00432-024-05673-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Gallifant J, Afshar M, Ameen S, et al. The TRIPOD-LLM reporting guideline for studies using large language models. Nat Med. 2025;31:60–9. doi: 10.1038/s41591-024-03425-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Vaduganathan M, Mensah GA, Turco JV, et al. The Global Burden of Cardiovascular Diseases and Risk. J Am Coll Cardiol. 2022;80:2361–71. doi: 10.1016/j.jacc.2022.11.005. [DOI] [PubMed] [Google Scholar]
  • 14.Ong KL, Stafford LK, McLaughlin SA, et al. Global, regional, and national burden of diabetes from 1990 to 2021, with projections of prevalence to 2050: a systematic analysis for the Global Burden of Disease Study 2021. Lancet. 2023;402:203–34. doi: 10.1016/S0140-6736(23)01301-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Taborda Restrepo PA, Acosta-Reyes J, Estupiñan-Bohorquez A, et al. Comparative Analysis of Clinical Practice Guidelines for the Pharmacological Treatment of Type 2 Diabetes Mellitus in Latin America. Curr Diab Rep. 2023;23:89–101. doi: 10.1007/s11892-023-01504-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Fegers-Wustrow I, Gianos E, Halle M, et al. Comparison of American and European Guidelines for Primary Prevention of Cardiovascular Disease. J Am Coll Cardiol. 2022;79:1304–13. doi: 10.1016/j.jacc.2022.02.001. [DOI] [PubMed] [Google Scholar]
  • 17.Chapman N, Breslin M, Zhou Z, et al. Comparison of Patients Classified as High-Risk between International Cardiovascular Disease Primary Prevention Guidelines. J Clin Med. 2024;13:4379. doi: 10.3390/jcm13154379. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Boriani G, Venturelli A, Imberti JF, et al. Comparative analysis of level of evidence and class of recommendation for 50 clinical practice guidelines released by the European Society of Cardiology from 2011 to 2022. Eur J Intern Med. 2023;114:1–14. doi: 10.1016/j.ejim.2023.04.020. [DOI] [PubMed] [Google Scholar]
  • 19.Burgers JS, Bailey JV, Klazinga NS, et al. Inside guidelines: comparative analysis of recommendations and evidence in diabetes guidelines from 13 countries. Diabetes Care. 2002;25:1933–9. doi: 10.2337/diacare.25.11.1933. [DOI] [PubMed] [Google Scholar]
  • 20.Meng X, Yan X, Zhang K, et al. The application of large language models in medicine: A scoping review. iScience. 2024;27:109713. doi: 10.1016/j.isci.2024.109713. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.DeepSeek-AI DeepSeek-v3 technical report. 2024
  • 22.Roman-Urrestarazu A, van Kessel R, Allison C, et al. Association of Race/Ethnicity and Social Disadvantage With Autism Prevalence in 7 Million School Children in England. JAMA Pediatr. 2021;175:e210054. doi: 10.1001/jamapediatrics.2021.0054. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Zou G. A modified poisson regression approach to prospective studies with binary data. Am J Epidemiol. 2004;159:702–6. doi: 10.1093/aje/kwh090. [DOI] [PubMed] [Google Scholar]
  • 24.Holmberg MJ, Andersen LW. Estimating Risk Ratios and Risk Differences: Alternatives to Odds Ratios. JAMA. 2020;324:1098–9. doi: 10.1001/jama.2020.12698. [DOI] [PubMed] [Google Scholar]
  • 25.Luo X, Rechardt A, Sun G, et al. Large language models surpass human experts in predicting neuroscience results. Nat Hum Behav. 2025;9:305–15. doi: 10.1038/s41562-024-02046-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Gallo RJ, Baiocchi M, Savage TR, et al. Establishing best practices in large language model research: an application to repeat prompting. J Am Med Inform Assoc. 2025;32:386–90. doi: 10.1093/jamia/ocae294. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Mello MM, Guha N. ChatGPT and Physicians’ Malpractice Risk. JAMA Health Forum . 2023;4:e231938. doi: 10.1001/jamahealthforum.2023.1938. [DOI] [PubMed] [Google Scholar]
  • 28.van Kessel R, Seghers L-E, Anderson M, et al. A scoping review and expert consensus on digital determinants of health. Bull World Health Organ. 2025;103:110–125H. doi: 10.2471/BLT.24.292057. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Burford KG, Itzkowitz NG, Ortega AG, et al. Use of Generative AI to Identify Helmet Status Among Patients With Micromobility-Related Injuries From Unstructured Clinical Notes. JAMA Netw Open. 2024;7:e2425981. doi: 10.1001/jamanetworkopen.2024.25981. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Guevara M, Chen S, Thomas S, et al. Large language models to identify social determinants of health in electronic health records. NPJ Digit Med. 2024;7:6. doi: 10.1038/s41746-023-00970-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Omar M, Soffer S, Agbareia R, et al. Sociodemographic biases in medical decision making by large language models. Nat Med. 2025;31:1873–81. doi: 10.1038/s41591-025-03626-6. [DOI] [PubMed] [Google Scholar]
  • 32.Xiong Z, Wang X, Zhou Y, et al. How Generalizable Are Foundation Models When Applied to Different Demographic Groups and Settings? NEJM AI . 2025;2:AIcs2400497. doi: 10.1056/AIcs2400497. [DOI] [Google Scholar]
  • 33.Widder DG, Whittaker M, West SM. Why “open” AI systems are actually closed, and why this matters. Nature New Biol. 2024;635:827–33. doi: 10.1038/s41586-024-08141-1. [DOI] [PubMed] [Google Scholar]
  • 34.Hager P, Jungmann F, Holland R, et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat Med. 2024;30:2613–22. doi: 10.1038/s41591-024-03097-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Cabana MD, Rand CS, Powe NR, et al. Why don’t physicians follow clinical practice guidelines? A framework for improvement. JAMA. 1999;282:1458–65. doi: 10.1001/jama.282.15.1458. [DOI] [PubMed] [Google Scholar]
  • 36.Qumseya B, Goddard A, Qumseya A, et al. Barriers to Clinical Practice Guideline Implementation Among Physicians: A Physician Survey. IJGM. 2021;14:7591–8. doi: 10.2147/IJGM.S333501. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Schlicht IB, Zhao Z, Sayin B, et al. Do LLMs provide consistent answers to health-related questions across languages? 2025. [Google Scholar]
  • 38.Gelman A, Carlin JB, Stern HS, et al. Bayesian data analysis third edition (with errors fIxed as of 20 February 2025). 3rd edn. CRC Press; 2025. [Google Scholar]
  • 39.Liu S, McCoy AB, Wright A. Improving large language model applications in biomedicine with retrieval-augmented generation: a systematic review, meta-analysis, and clinical development guidelines. J Am Med Inform Assoc. 2025;32:605–15. doi: 10.1093/jamia/ocaf008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Ng KKY, Matsuba I, Zhang PC. RAG in Health Care: A Novel Framework for Improving Communication and Decision-Making by Addressing LLM Limitations. NEJM AI. 2025;2 doi: 10.1056/AIra2400380. [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

online supplemental file 1
bmjhci-33-1-s001.pdf (34.9MB, pdf)
DOI: 10.1136/bmjhci-2025-101959

Data Availability Statement

Data are available in a public, open access repository.


Articles from BMJ Health & Care Informatics are provided here courtesy of BMJ Publishing Group

RESOURCES