Skip to main content
BMJ Open Access logoLink to BMJ Open Access
editorial
. 2025 Apr 17;34(7):e018559. doi: 10.1136/bmjqs-2025-018559

Understanding the evidence for artificial intelligence in healthcare

Gretchen Purcell Jackson 1,2,, Edward H Shortliffe 3,4
PMCID: PMC12229047  PMID: 40246317

Scientific studies of artificial intelligence (AI) solutions in healthcare have been the subject of intense criticism—both in research publications and in the media.1,3 Early validations of predictive algorithms are criticised for not having meaningful clinical impact, and AI tools that make mistakes or fail to show immediate improvement in health outcomes are heralded as the first snowflakes in the next AI winter (a period of decreased interest in AI research and development). Scientific evidence is the language of trust in healthcare, and peer-reviewed studies evaluating AI solutions are key to fostering adoption. There are over two dozen reporting guidelines for AI in medicine,4 and many other consensus statements and standards that offer recommendations for the publication of research about medical AI.5 Despite such guidance, the average frontline clinician still struggles in interpreting the results of an AI study to determine whether the reported tool is safe and effective enough to use in practice. This editorial offers a pragmatic and clinically focused perspective on the evaluation of AI for healthcare.

As for most innovations in medicine, AI needs comprehensive and systematic evaluation. Comprehensive evaluation means that each element of an AI solution should be assessed throughout its development and after deployment to show that the system works, can be used effectively in relevant healthcare settings, and has an impact on health or healthcare delivery. The American Medical Informatics Association (AMIA), a leading professional society for biomedical informatics in the USA, developed the AI Evaluation Showcase as a scientific forum for robust AI research. The AMIA AI Showcase requires scientists to report evaluation of AI in three components: technical performance, usability and workflow, and impact6 (figure 1). Systematic evaluation means that these three phases of evaluation should be done in order. Clinicians often prioritise evidence of the effects on clinical outcomes, but if an AI tool is not accurate, and its usability has not been tested in a real clinical workflow, deployment is unlikely to produce the desired clinical result. The entire evaluation sequence—technical performance, usability and workflow, and impact—should also be repeated whenever conditions change, especially if the models may learn and change their performance over time.7 Model performance can vary for a wide variety of reasons such as changes in the underlying data used for prediction or behavioural changes from use of the model itself. For example, when a model is predicting mortality or suicide, the primary objective is not simply accurate forecasting. Instead, use of models to identify high-risk scenarios should also lead to actions to prevent the predicted adverse events and decrease their incidence. Over time, changes to clinician behaviours in response to using such AI models can influence model performance.8

Figure 1. A framework for the evaluation of artificial intelligence in healthcare.

Figure 1

In this issue of BMJ Quality and Safety, Tabuchi and colleagues present a study of an AI-based surgical safety system, which provides an excellent example of comprehensive and systematic AI evaluation.9 The AI-based tool was created to support ophthalmology teams in confirming the correct patient, eye and lens for implantation in cataract surgery using facial recognition and computer vision technologies. In ophthalmology, bilateral surgeries, short procedure times, and indistinguishable implants increase the likelihood of devastating medical errors. The technical performance of the AI algorithms had been reported previously,10 with successful authentication rates, on the first and with repeated attempts respectively, of 92.0% and 96.3% for patient identification, 82.5% and 98.2% for laterality confirmation, and 67.4% and 88.9% for lens verification. The team had also tested usability and elicited feedback to evolve the workflow for using the system.11 This most recent study then reassessed system performance, examined usability issues when deployed at scale, and focused on the clinical impact of the system by comparing medical errors and near-misses before and after adoption. In this observational study of 37 529 procedures (18 767 pre-adoption, 18 762 post-adoption), five medical errors occurred after implementation of the system compared with one beforehand. In four post-adoption cases where errors occurred, the AI system was not used, and in the fifth, the system correctly identified a laterality problem, but the information was not communicated to the surgeon by the operating room staff. Near misses increased significantly from 9 pre-implementation to 30 post-implementation, with growth in errors and near misses likely representing better detection of errors rather than unintended consequences of technology.

Evaluation of an AI tool’s technical performance may seem like an exercise best left to computer scientists, but a key part of this step is ensuring that algorithms are accurate enough for their clinical application. In general, most AI algorithms either predict or classify, so their performance is measured in a manner similar to the evaluation of diagnostic tests, using metrics such as sensitivity, specificity and area under the curve.12 13 Healthcare providers should pay particular attention to rates of false positives and false negatives, as well as their consequences, as clinical judgement is often needed in selecting performance thresholds. For example, the consequences of an AI tool that automatically interprets a radiographic image and misses a lung lesion (ie, a false negative) may be severe, whereas an AI tool searching the medical literature and finding an irrelevant study for your reading list (ie, a false positive) is annoying, but not likely to compromise life or limb. Studies of healthcare AI tools should explicitly report false positive and false negative rates rather than composite measures such as F1 scores (an evaluation metric that combines precision and recall) so that medical practitioners can determine their suitability for practice. In the Tabuchi study, technical performance improved over time, with authentication rates for all models exceeding 99% after 3 months. False positive or false negative rates were 0% if the system successfully completed all three steps of authentication, which is important for the critical surgical time-out tasks.10

It is important to measure the baseline performance for the task an AI algorithm is intended to replicate. AI is often designed to automate tasks that humans may do poorly or inconsistently. The Tabuchi study revealed a medical error rate that was higher than expected, most likely because the objective measurement of errors and near misses found some that were previously undetected. Baseline performance assessment in medicine can be a sensitive topic, as it may reveal that clinicians are not perfect. Creating the ground truth for performance evaluation must also take into account that medical experts do not always agree on the right answer when it comes to diagnostic or therapeutic decisions. Evaluation of MYCIN, one of the earliest computerised knowledge-based systems for the diagnosis and treatment of infectious diseases, illustrated this point—infectious disease specialists did not all agree on the best treatment.14 More recently, an evaluation of an AI-based oncology clinical decision support tool highlighted that both clinicians in practice and the AI tool made less preferred or unacceptable therapeutic recommendations when compared with the recommendation of a panel of experts.15

A commonly skipped step in evaluation is ensuring that AI solutions can be used by their intended users to accomplish the desired tasks in often complex healthcare workflows. It is important to envision the entire scenario in which information from an AI algorithm is used.8 For example, what happens when you are about to discharge a patient from the emergency department, and an AI algorithm predicts a high risk of suicide, or a popular consumer health application tells your mother that the mole on her arm is likely to be a lethal cancer? The phases of clinical research used for drugs and devices are useful for framing AI evaluation, but there are important nuances, as articulated by Park and colleagues.16 AI is similar to drugs and devices, as early-phase ‘laboratory’ studies are needed to demonstrate proof-of-concept technical performance, usability of prototypes, and potential for impact. AI differs from drugs and devices, which tend to function in a relatively predictable manner, whereas the output of an AI system often must be understood, trusted, contextualised, and used by humans, who can be highly unpredictable. AI evaluations should accordingly incorporate assessments of understanding, explainability, and ability to take appropriate action, and they should look for unexpected applications of the technology and unintended consequences. An active area of research in medical education is the determination of what competencies might be necessary for health professionals to learn to use AI safely in practice.17 For the ophthalmology surgical safety system, usability and workflow issues contributed to errors. Two of the five medical errors reported occurred in early implementation, when authentication rates were lower or required multiple attempts, leading frustrated staff to abandon using the system. Two implantation errors happened when the lens recognition model had not been trained on new implants. The system required three updates during the study period to stay current.

The ultimate goal of AI evaluation is demonstrating that the technology has a positive impact on health or healthcare delivery. Measuring clinical outcomes in comparative trials often takes years, and evaluation can pose a challenge for AI algorithms that learn and evolve rapidly over time. When AI tools are first introduced, it is appropriate to measure proximal or process outcomes—do clinical decision support tools change decisions or do population health management tools save time? Only with time can one look at long-term outcomes, such as whether a system improves outcomes such as survival, recurrence or costs of care. In designing clinical studies, it is also crucial to choose the right comparator. Everyone loves pitting humans against machines, but AI solutions that support rather than replace clinicians should be evaluated by comparing the performance of providers with versus without the tool, not human versus machine.

Several prominent healthcare leaders, including Curtis Langlotz and Jesse Ehrenfeld, have predicted that doctors may never be replaced by AI, but that those who use it will replace those who do not.18 19 Understanding AI and the scientific literature that evaluates it will be a key clinical competency for the next generation of healthcare providers.20 Technical performance studies, usability studies in real healthcare environments, and evaluations of short and long-term health impact are critical bricks in the foundation of evidence that is needed to assure safe and effective deployment of AI tools in healthcare environments.

Acknowledgements

The authors acknowledge Brent Hainsworth for his creation of Figure 1 and Jeff Williamson for his contributions to the development of the AMIA AI Evaluation Showcase.

Footnotes

Funding: The authors have not declared a specific grant for this research from any funding agency in the public, commercial or not-for-profit sectors.

Patient consent for publication: Not applicable.

Ethics approval: Not applicable.

Provenance and peer review: Commissioned; internally peer reviewed.

References

  • 1.Andaur Navarro CL, Damen JAA, Takada T, et al. Systematic review finds “spin” practices and poor reporting standards in studies on machine learning-based prediction models. J Clin Epidemiol. 2023;158:99–110. doi: 10.1016/j.jclinepi.2023.03.024. [DOI] [PubMed] [Google Scholar]
  • 2.Shahzad R, Ayub B, Siddiqui MAR. Quality of reporting of randomised controlled trials of artificial intelligence in healthcare: a systematic review. BMJ Open. 2022;12:e061519. doi: 10.1136/bmjopen-2022-061519. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Tahir D. CBS News; [23-Dec-2024]. Health care ai, intended to save money, turns out to require a lot of expensive humans.https://www.cbsnews.com/news/health-care-ai-cost-humans/ Available. Accessed. [Google Scholar]
  • 4.Kolbinger FR, Veldhuizen GP, Zhu J, et al. Reporting guidelines in medical artificial intelligence: a systematic review and meta-analysis. Commun Med (Lond) 2024;4:71. doi: 10.1038/s43856-024-00492-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Wang Y, Li N, Chen L, et al. Guidelines, Consensus Statements, and Standards for the Use of Artificial Intelligence in Medicine: Systematic Review. J Med Internet Res. 2023;25:e46089. doi: 10.2196/46089. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.AMIA Artificial intelligence evaluation showcase. [13-Jan-2025]. https://amia.org/education-events/amia-artificial-intelligence-evaluation-showcase Available. Accessed.
  • 7.Petersen C, Smith J, Freimuth RR, et al. Recommendations for the safe, effective use of adaptive CDS in the US healthcare system: an AMIA position paper. J Am Med Inform Assoc. 2021;28:677–84. doi: 10.1093/jamia/ocaa319. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Lenert MC, Matheny ME, Walsh CG. Prognostic models will be victims of their own success, unless…. J Am Med Inform Assoc. 2019;26:1645–50. doi: 10.1093/jamia/ocz145. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Tabuchi H, Ishitobi N, Deguchi H, et al. Large-scale observational study of AI-based patient and surgical material verification system in ophthalmology: real-world evaluation in 37 529 cases. BMJ Qual Saf. 2025;34:433–42. doi: 10.1136/bmjqs-2024-018018. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Kiuchi G, Tanabe M, Nagata K, et al. Deep Learning-Based System for Preoperative Safety Management in Cataract Surgery. J Clin Med. 2022;11:5397. doi: 10.3390/jcm11185397. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Tabuchi H, Masumoto H, Adachi S. Real-world testing of artificial intelligence system for surgical safety management. Invest Ophthalmol Vis Sci. 2020;61:2032–32. doi: 10.1167/iovs.61.13.32. [DOI] [Google Scholar]
  • 12.Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi: 10.1136/bmj-2023-078378. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Hernandez-Boussard T, Bozkurt S, Ioannidis JPA, et al. MINIMAR (MINimum Information for Medical AI Reporting): Developing reporting standards for artificial intelligence in health care. J Am Med Inform Assoc. 2020;27:2011–5. doi: 10.1093/jamia/ocaa088. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Yu VL, Fagan LM, Wraith SM, et al. Antimicrobial selection by a computer. A blinded evaluation by infectious diseases experts. JAMA. 1979;242:1279–82. [PubMed] [Google Scholar]
  • 15.Suwanvecho S, Suwanrusme H, Jirakulaporn T, et al. Comparison of an oncology clinical decision-support system’s recommendations with actual treatment decisions. J Am Med Inform Assoc. 2021;28:832–8. doi: 10.1093/jamia/ocaa334. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Park Y, Jackson GP, Foreman MA, et al. Evaluating artificial intelligence in medicine: phases of clinical research. JAMIA Open. 2020;3:326–31. doi: 10.1093/jamiaopen/ooaa033. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Russell RG, Lovett Novak L, Patel M, et al. Competencies for the Use of Artificial Intelligence-Based Tools by Health Care Professionals. Acad Med. 2023;98:348–56. doi: 10.1097/ACM.0000000000004963. [DOI] [PubMed] [Google Scholar]
  • 18.Langlotz CP. Will Artificial Intelligence Replace Radiologists? Radiol Artif Intell. 2019;1:e190058. doi: 10.1148/ryai.2019190058. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Robeznieks A. American Medical Association; [22-Feb-2025]. AI is already reshaping care. Here’s what it means for doctors.https://www.ama-assn.org/practice-management/digital/ai-already-reshaping-care-heres-what-it-means-doctors Available. Accessed. [Google Scholar]
  • 20.Matheny ME, Goldsack JC, Saria S, et al. Artificial Intelligence In Health And Health Care: Priorities For Action. Health Aff (Millwood) 2025;44:163–70. doi: 10.1377/hlthaff.2024.01003. [DOI] [PubMed] [Google Scholar]

Articles from BMJ Quality & Safety are provided here courtesy of BMJ Publishing Group

RESOURCES