Skip to main content
QJM: An International Journal of Medicine logoLink to QJM: An International Journal of Medicine
. 2025 May 19;118(11):802–804. doi: 10.1093/qjmed/hcaf114

Artificial intelligence for clinical reasoning: the reliability challenge and path to evidence-based practice

He Xu 1,2,3, Yueqing Wang 4, Yangqin Xun 5, Ruitai Shao 6,, Yang Jiao 7,
PMCID: PMC12778421  PMID: 40489895

Abstract

The integration of generative artificial intelligence (AI), particularly large language models (LLMs), into clinical reasoning heralds transformative potential for medical practice. However, their capacity to authentically replicate the complexity of human clinical decision-making remains uncertain—a challenge defined here as the reliability challenge. While studies demonstrate LLMs’ ability to pass medical licensing exams and achieve diagnostic accuracy comparable to physicians, critical limitations persist. Crucially, LLMs mimic reasoning patterns rather than executing genuine logical reasoning, and their reliance on outdated or non-regional data undermines clinical relevance. To bridge this gap, we advocate for a synergistic paradigm where physicians leverage advanced clinical expertise while AI evolves toward transparency and interpretability. This requires AI systems to integrate real-time, context-specific evidence, align with local healthcare constraints, and adopt explainable architectures (e.g. multi-step reasoning frameworks or clinical knowledge graphs) to demystify decision pathways. Ultimately, reliable AI for clinical reasoning hinges on harmonizing technological innovation with human oversight, ensuring ethical adherence to beneficence and non-maleficence while advancing evidence-based, patient-centered care.

Introduction

The rise of artificial intelligence (AI) marks a transformative era that looks set to reshape medical practice. With the advent of generative large language models (LLMs), especially those capable of producing chain-of-thought (CoT) narrative,1 AI’s role seems to be shifting from merely supporting to potentially executing clinical reasoning. However, this raises a critical question: can CoT AI authentically replicate the intricate clinical reasoning required of human clinical decision-making?

Certainly, the literature continues to expand with studies applying generative AI to different problems in healthcare. However, followers of the topic will also be aware that many AI studies exaggerate the benefits and potential and downplay the drawbacks of AI. Overall, the specific value and reliability of generative AI in clinical reasoning remain underexplored and poorly understood.2 Clinical settings are highly diverse, varying according to geography, resourcing, and the unique interpersonal dynamics of physician-patient relationships. Reliable clinical reasoning is fundamental for safe, high-quality care delivery in heterogeneous clinical settings, while reliable clinical AI for physician lies in the authentic clinical reasoning capability.

Here, we explore the current landscape of the reliability of generative AI for clinical reasoning, especially with respect to its potential to replicate and enhance human clinical decision-making. Our hope is to call for a reliable AI system towards future evidence-based practice.

What is the ‘reliability challenge’ of AI for clinical reasoning?

Clinical reasoning is a sophisticated process that integrates knowledge and experience to formulate accurate diagnoses and management plans. This dynamic and iterative process requires the synthesis of multifaceted and multimodal ‘data’ regarding individual symptoms, physical signs, imaging, laboratory test results, pathological examinations, risk factors, and medical histories. It takes years of rigorous training, consolidated with hands-on experience, for medical professionals to master this intricate skill. With the advent of LLMs, which appear capable of ‘reasoning’ based on giving answers to clinical questions entered as prompts into chatbots, some have speculated that AI looks poised to replace human clinicians by delivering faster and more accurate diagnoses. However, diagnostic inaccuracies from AI may constitute a clinical risk that is too high to be acceptable.

So what is the evidence base for AI being able to perform clinical reasoning? The existing literature highlights some impressive capabilities for LLMs in clinical reasoning tasks. For example, LLMs have been shown to pass the United States Medical Licensing Exam (USMLE).3 In further clinical case tests, LLMs were comparable to attending physicians and residents in terms of diagnostic accuracy and cannot-miss diagnoses. Indeed, LLMs achieved a median inclusion rate of 66.7% for cannot-miss diagnoses, similar to attending physicians (50.0%) and matching residents (66.7%).4

Despite this promise shown in clinical reasoning tests, LLMs also have notable limitations.5 LLMs commit more diagnostic errors than medical residents4 and misunderstand clinical terminology, such as illness scripts and problem lists, necessitating additional explanations.6 Furthermore, physician-AI chatbot (GPT-4) collaboration does not outperform physicians using conventional resources (e.g. UpToDate, Google).7 While AI may be more accurate in certain scenarios, it also generates a higher volume of errors, highlighting its practical shortcomings in clinical settings.

Critically, generating reasoned text or passing a medical exam does not equate to genuine logical reasoning. Researchers from Apple remind us that LLMs do not truly reason but instead mimic reasoning steps based on training data. A single additional clause in a question (prompt) can degrade performance by up to 65%,8 resulting in inconsistent and unreliable outputs. This sensitivity to prompts has also been observed in clinical reasoning examinations.6 Moreover, while LLMs can produce coherent reasoning texts, they lack access to high-quality regional medical data and up-to-date medical knowledge to ensure implementation of an optimal treatment plan. Retraining these models with the latest medical information is not only costly but also risks catastrophic loss or inadvertent assimilation of unreliable or misleading contents from the open web.

Call for AI toward evidence-based practice

Diagnostic inaccuracies from generative AI not only compromise patient safety in clinical settings but may also lead to inefficiency and misguided clinical decision-making.7 These errors violate the core ethnical principles of beneficence and non-maleficence, threatening patient safety and trust. Two critical factors have historically overcome these challenges: an improvement in digital tools and their effective implementation by humans.2 Building on this foundation, we, therefore, propose a transformative paradigm where evidence-based practice hinges on a synergistic collaboration between physicians and AI. In this paradigm, physicians must remain equipped with the latest medical knowledge and sophisticated clinical reasoning skills, while AI must evolve to be more interpretable and transparent. To achieve this, AI systems must not only develop authentic clinical reasoning capabilities that mirror human cognitive processes but also access and present the most recent evidence aligned with local clinical contexts, ensuring transparency through traceable sourcing and relevance to region-specific healthcare challenges and resource constraints.

It is worth remembering that interpretable AI models predate the emergence of LLM. For example, Liang et al.9 developed a hierarchical logistic regression classifier that systematically classified the electronic health record, starting with broad organ systems, followed by anatomic categories, and finally specific diagnosis groups. Similarly, reasoning-based AI, supported by expert clinical knowledge, has been designed to emulate the deliberate and methodical nature of human clinical reasoning. Causal graph knowledge, for example, explicitly delineates causal relationships with uncertainties and subsequently performs probabilistic reasoning.10

A promising trend, rapidly gaining momentum, is the integration of interpretable components into LLMs. Emerging initiatives aim to combine LLMs with multi-step reasoning11 or clinical reasoning graph knowledge to address the ‘black-box’ nature of purely diagnostic models.12 These advances are critical because clinicians need to first comprehend the rationale behind AI-generated recommendations to trust and integrate these tools in their decision-making processes.

Conclusions

The reliability of AI for clinical reasoning is a critical challenge that can be effectively resolved through a collaborative approach between technology and human clinicians. The rapid advancement of AI in clinical reasoning, when combined with critical clinical thinking, could unlock the transformative potential of AI for evidence-based practice.

Contributor Information

He Xu, Department of General Practice (General Internal Medicine), Peking Union Medical College Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China; School of Population Medicine and Public Health, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China; 4 + 4 MD Program, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China.

Yueqing Wang, School of Population Medicine and Public Health, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China.

Yangqin Xun, School of Population Medicine and Public Health, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China.

Ruitai Shao, School of Population Medicine and Public Health, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China.

Yang Jiao, Department of General Practice (General Internal Medicine), Peking Union Medical College Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China.

Author contributions

H.X. was responsible for the writing of the first draft of the thesis; H.X., Y.W., and Y.X. participated in the revision and text finishing of the first draft; Y.J. and R.S. were responsible for the selection of the thesis, structural design, data and information review, revision, quality control and review of the first draft of the article.

Conflict of interest: None declared.

Funding

National High Level Hospital Clinical Research Funding (2022-PUMCH-A-017); the non-profit Central Research Institute Fund of Chinese Academy of Medical Sciences (2022-ZHCH330-01); the non-profit Central Research Institute Fund of Chinese Academy of Medical Sciences (2021-RC330-004).

Patient consent for publication

Not applicable.

Data availability

Not applicable

References

  • 1. Wei J, Wang X, Schuurmans D, Bosma M, Ichter B, Xia F, et al. , arXiv.org. 2022. Chain-of-thought prompting elicits reasoning in large language models. https://arxiv.org/abs/2201.11903v6, (14 April 2025, date last accessed).
  • 2. Wachter RM, Brynjolfsson E.  Will generative artificial intelligence deliver on its promise in health care?  JAMA  2024; 331:65–9. [DOI] [PubMed] [Google Scholar]
  • 3. Shieh A, Tran B, He G, Kumar M, Freed JA, Majety P.  Assessing ChatGPT 4.0’s test performance and clinical diagnostic accuracy on USMLE STEP 2 CK and clinical case reports. Sci Rep  2024; 14:9330. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Cabral S, Restrepo D, Kanjee Z, Wilson P, Crowe B, Abdulnour RE, et al.  Clinical reasoning of a generative artificial intelligence model compared with physicians. JAMA Intern Med  2024; 184:581–3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Newman-Toker DE, Nassery N, Schaffer AC, Yu-Moe CW, Clemens GD, Wang Z, et al.  Burden of serious harms from diagnostic error in the USA. BMJ Qual Saf  2024; 33:109–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Strong E, DiGiammarino A, Weng Y, Kumar A, Hosamani P, Hom J, et al.  Chatbot vs medical student performance on free-response clinical reasoning examinations. JAMA Intern Med  2023; 183:1028–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Goh E, Gallo R, Hom J, Strong E, Weng Y, Kerman H, et al.  Large language model influence on diagnostic reasoning: a randomized clinical trial. JAMA Netw Open  2024; 7:e2440969. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Mirzadeh I, Alizadeh K, Shahrokhi H, Tuzel O, Bengio S, Farajtabar M. GSM-Symbolic: understanding the limitations of mathematical reasoning in large language models. arXiv; 2024. http://arxiv.org/abs/2410.05229 (3 November 2024, date last accessed)
  • 9. Liang H, Tsui BY, Ni H, Valentim CCS, Baxter SL, Liu G, et al.  Evaluation and accurate diagnoses of pediatric diseases using artificial intelligence. Nat Med  2019; 25:433–8. [DOI] [PubMed] [Google Scholar]
  • 10. Wu H, Shi W, Wang MD.  Developing a novel causal inference algorithm for personalized biomedical causal graph learning using meta machine learning. BMC Med Inform Decis Mak  2024; 24:137. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Gao S, Zhu R, Kong Z, Noori A, Su X, Ginder C  et al. TxAgent: an AI agent for therapeutic reasoning across a universe of tools. arXiv; 2025. http://arxiv.org/abs/2503.10970 (30 March 2025, date last accessed)
  • 12. Gao Y, Li R, Croxford E, Caskey J, Patterson BW, Churpek M, et al.  Leveraging medical knowledge graphs into large language models for diagnosis prediction: design and application study. JMIR.  2025; 4: e58670. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Not applicable


Articles from QJM: An International Journal of Medicine are provided here courtesy of Oxford University Press

RESOURCES