Skip to main content
Cell Reports Medicine logoLink to Cell Reports Medicine
. 2022 Dec 12;3(12):100860. doi: 10.1016/j.xcrm.2022.100860

Evidence synthesis, digital scribes, and translational challenges for artificial intelligence in healthcare

Enrico Coiera 1,2,, Sidong Liu 1
PMCID: PMC9798027  PMID: 36513071

Summary

Healthcare has well-known challenges with safety, quality, and effectiveness, and many see artificial intelligence (AI) as essential to any solution. Emerging applications include the automated synthesis of best-practice research evidence including systematic reviews, which would ultimately see all clinical trial data published in a computational form for immediate synthesis. Digital scribes embed themselves in the process of care to detect, record, and summarize events and conversations for the electronic record. However, three persistent translational challenges must be addressed before AI is widely deployed. First, little effort is spent replicating AI trials, exposing patients to risks of methodological error and biases. Next, there is little reporting of patient harms from trials. Finally, AI built using machine learning may perform less effectively in different clinical settings.

Keywords: evidence-based medicine, evidence synthesis, patient safety, research replication, machine learning, algorithmic transportability, deep learning, clinical trial registries

Graphical abstract

graphic file with name fx1.jpg


Healthcare artificial intelligence (AI) is shifting from supporting discrete tasks like diagnosis to whole workflows. Coiera and Liu discuss emerging workflow applications including automated systematic review, evidence synthesis, and digital scribes. They identify three translational AI challenges—replicability, safety, and transportability—that must be addressed before widespread AI deployment.

Introduction

Across the world, healthcare systems are under significant duress, managing evolving challenges in disease patterns, pandemics, and climate-triggered events.1 Even without such shocks to contend with, the delivery of healthcare services has always been challenging because, in a complex system, there are few easy opportunities for improvement.2 Healthcare has well-known and seemingly intractable challenges with the safety, quality, and effectiveness of clinical services. These include misdiagnosis, overdiagnosis, overtreatment, treatment errors, and diminishing resources and workforce to support ever more stretched clinical services.3,4,5

There are no magic bullets, but many see artificial intelligence (AI) as an essential component of any solution to these problems. AI is a broad set of technologies and methods, focusing on automating reasoning tasks such as planning, understanding, predicting, and classifying. Machine learning is the sub-discipline of AI that focuses on developing ways for AI systems to learn from experience. AI offers the possibility of automation and decision support for skilled tasks, such as diagnosis and treatment selection, improvements in triage and hospital discharge decisions, and a reduction in documentation burden.6 Indeed, in the short run, there are probably more lives to be saved or improved just by doing a better job of healthcare delivery, than there are through creating new treatments.7

Overall, perhaps the most reliable global estimate for the potential for AI in healthcare comes from Lord Darzi’s review of the English National Health System (NHS), where modeling identified productivity improvement from smart automation worth £12.5 billion a year: 9.9% of the NHS England budget.8 Other estimates are based on modeling specific services. For example, using AI to reduce non-elective hospital admissions could save up to £3.3 billion annually.4 This potential has driven extraordinary investments globally. The English NHS has allocated over £1 billion on initiatives such as a £250-million national AI laboratory as well as translational research centers targeted at reducing cancer deaths by 10% a year (or 22,000 lives) by 2035 through AI-enhanced services.9 KPMG data have suggested that US investment in AI for healthcare would reach US$6.6 billion by 2021 (a 40% CAGR), driven by modeling suggesting potential total savings of US$150 billion by 2026.

The past decade has seen substantive progress in AI technological development, most notable in machine learning. In the application space, deep learning systems that use neural network architectures are now emerging from clinical trial and slowly moving into routine care. The US Food and Drug Administration (FDA), for example, has seen a sharp increase in the number of clinical AI systems that it has approved for use in the market (Figure 1). The scale of modern deep learning systems, and the rich opportunities for commercial gain in the sector, has seen a steady drift of researchers and research breakthroughs from academia across to industry.10

Figure 1.

Figure 1

FDA approvals for devices incorporating AI

Approvals by the US regulator the FDA of clinical systems incorporating artificial intelligence capabilities have increased dramatically over the past decade. (Source: US Food and Drug Administration, 2022).

In this perspective, we first explore the emerging application of AI in healthcare on the critical tasks of evidence synthesis and clinical documenting, reflecting a shift from tasks such as clinical image diagnosis, toward use cases that support multi-step clinical workflows. We next focus on the difficult challenges that are found when implementing working AI in the real world, where technology, people, and practice must each accommodate the other.

Emerging applications for AI in healthcare

The “canonical” applications for AI in healthcare that have garnered the most recent attention sit in data-rich domains that are well suited to a deep learning approach, like medical imaging or laboratory medicine. These well-documented applications are typically characterized by a very tight focus on narrow tasks such as screening for diabetic retinopathy11 and glaucoma,12 diagnosis of thyroid cancer from ultrasound data,13 diagnosis of COVID-19 in radiological chest images,14 or diagnosis of primary and metastatic cancers from whole transcriptome data.15

Such applications of AI seek to optimize discrete classification tasks such as diagnosis, rather than optimizing the greater human workflow within which the task is embedded. The risk of such a narrow approach is that we optimize what is technically feasible rather than what is clinically effective. We should instead aim to optimize the overall workflow, targeting the links in the information value chain that underpin a decision that offers the greatest cost-benefit.16

In the next section, we focus on two such emerging AI applications—systematic review automation and digital scribes—both of which seek to digitize entire real-world workflows and support the process of care delivery. What distinguishes these applications is that the output of these processes is not a classification label such as a diagnosis but rather a multi-component knowledge object—a systematic review or the documentation of a clinical encounter. While completely “solving” these processes is beyond today’s state of the art, breaking them down into distinct steps is allowing us to gradually and incrementally optimize the whole workflow and deliver real clinical benefits.

Automated evidence synthesis

In a time of crisis such as the COVID-19 pandemic, there is an urgent need for rapid assessments of the published research literature to answer specific clinical and public health questions.17 The US NIH’s LitCovid hub for example had curated about 270,000 scientific articles from 8,000 journals by July 2022. Delays to answering questions about whether COVID-19 was airborne, whether masks were effective, or whether smart-phone contact tracing was effective all had substantial real-world consequences.18

Unfortunately, current approaches to systematic review (SR), the gold standard approach to synthesizing published clinical evidence to answer such questions, typically take months or years.19 The mean time to complete and publish a systematic review, for example, is about 1.3 years.20 The stark gap between what the research evidence tells us should be done, and what actually is done, means that many patients do not receive care according to the best evidence. Pre-pandemic, this lag led to significant unnecessary waste across the healthcare system of up to $274 billion per year globally.21

SRs follow formal protocols for research evidence synthesis, and historically have relied on human expertise and labor to carry out the review. Living SRs aim to address some of the causes of delay in review production by addressing the speed with which reviews are updated. A “living” review is published once but then quickly updated if new evidence is made available.22 Living meta-analyses have been created for many COVID-19 treatments,23 with the Cochrane Collaboration piloting the approach, reappraising literature every 1 to 3 months.24 While an excellent step in the right direction, such approaches rely on substantial human expertise and effort. Bottlenecks and limits to human resources mean that expert-led living reviews will not scale to become the standard for all SRs.

Using automation and AI can improve our ability to synthesize the research literature,25 vastly reduce SR workload, and dramatically improve speed and quality.26 Since 2019, multiple technology-accelerated SRs have been undertaken using automation support, reducing the time for humans to complete a systematic review from 12 months to 2 weeks,27 including one focusing on the asymptomatic transmission of SARS-CoV-2.5

Current technology-assisted SRs are undertaken by using a collection of task-specific computational tools that target discrete steps in the systematic review process (Table 1), such as searching for and screening research articles, estimating risk of bias, as well as tasks like data extraction and report writing. This “toolkit” approach has the potential to improve systematic review timeliness and quality, and gradually require less human intervention. Increasingly, these tools are being built using AI methods including machine learning.28

Table 1.

Automation tools can support different stages of systematic review

Review Task Description Classification Example Tools
Formulate question Decide on research question for review Preparation COVID-SEE25
Write protocol Objective reproducible method for peer review Preparation Template; Methods Wizard
Search strategy Decide on keywords and databases Preparation SearchRefiner29; Scientific Evidence Explorer25
Search translation Translate search string for other databases Retrieval Polyglot Search Translator30
De-duplicate Merge identical citations Retrieval The SRA De-duplicator31
Screen Exclude irrelevant trials on title and abstract Appraisal SRA Helper,30 RobotSearch
Get full text Download/request study Retrieval SRA Helper, SARA32
Screen full text Exclude irrelevant studies Appraisal SRA Helper
Snowball Follow citations Retrieval CitationSpider
Extract data Get trial arm outcome numbers Synthesize RevMan
Assess risk of bias/quality Assess potential biases/quality of evidence33 Synthesize RobotReviewer34
EvidenceGRADEr
Meta-analyze Statistical data combination Synthesize RevMan35
Write up Produce and publish report Write up RevMan, Replicant35

The different tasks in a traditional systematic review can be supported by a variety of distinct automation tools that either support humans to complete the task or can complete the task automatically (modified from Tsafnat et al.36).

Ultimately the goal is to create SRs nearly instantaneously in response to specific questions, so that these evidence summaries are always up-to-date (Figure 2).26 The road to achieving such “full” automation will likely move through several distinct stages. Most SR tools are currently stand-alone, selected and operated by human reviewers. The creation of tool connecting pipelines will allow for greater automation across multiple tasks. To achieve this, individual tools must be capable of creating standardized input and output and be connected together using application programming interfaces.37 Each pipeline is a computational protocol. This allows for sharing of methods, the creation of benchmark methods and datasets, and collaborative improvement of tools, standards, and protocols, especially if they are part of an open-source community.

Figure 2.

Figure 2

The automation of systematic review

The time for a systematic review to be developed (dev), its currency decay (dec), and be updated (upd) decrease when automation partially supports “living” reviews. With full automation, an evidence review would be produced almost instantaneously and always be up-to-date. (Adapted from White et al. MJA, 2020).38

An implicit assumption behind most efforts to use automation to assist with SRs is that we are substituting computational methods to complete activities that humans currently undertake. However, humans and machines have different capabilities, and we can reconceive both the individual steps in evidence synthesis and their ordering when machines undertake them. For example, in the standard human SR process, candidate articles are first screened for inclusion or exclusion, often only using the title and abstract. Only later, when article numbers are much reduced, is the time-consuming process of data extraction undertaken. However, what is time-consuming for humans may be easy for a machine. Consequently, the automated extraction of study characteristics from abstracts can effectively make screening decisions,39 even though such a workflow would be hugely inefficient if undertaken by a human.

The ambitions for a computable approach to evidence synthesis are, however, much greater than the automation of systematic reviews, given that such reviews are only one of many forms of evidence synthesis. The larger game is for all clinical trial data to be published in a computational form that allows for immediate synthesis with other trials, and indeed other forms of evidence.40 Such a goal relies on achieving consensus on standards for publishing clinical trials in computable form,41 governance arrangements that see trial data made available for analysis beyond those who collected the initial data, and the development of intelligent tools to undertake synthesis tasks. Publishing trial information and results in a structured form will allow for automatic monitoring for new trials. New trials could then signal that a systematic review needs to be updated.42

Clinical trial evidence, however, cannot answer all our healthcare questions. Trials are expensive to conduct, and by design are controlled. For example, strict inclusion and exclusion criteria typically exclude patients with comorbidities, so that the trial populations do not necessarily represent real-world populations or settings. They also do not necessarily capture data that can be used to develop diagnostic or prognostic algorithms. When clinical trial data are unavailable to answer a question, observational data that are captured in electronic health records (EHRs) may be able to help.43 Indeed, creating algorithms developed on population data has been a core objective of AI research and practice. Making patient-specific predictions using population data remains challenging, especially with rare diseases, unusual presentations, or multimorbidity. In such cases, careful methods must be used to identify a cohort of patients sufficiently similar to the patient being managed from the electronic record data.44

Longer term, the evidence synthesis project will bring together data from clinical trials with longitudinal data from EHRs. This will require innovations not just in machine learning and statistics, but careful attention to the design of the decision support systems that use these methods to influence human decisions.

The digital scribe

Digital scribes are intelligent documentation support systems. They use advances in speech recognition (SpR), natural language processing, and AI to automatically document spoken elements of the clinical encounter, similar to the function performed by human medical scribes.45,46,47

The motivations for using digital scribes are compelling. Over 40% of US clinicians report at least one symptom of burnout,48 and modern EHRs are partly to blame. Since EHRs were introduced, the time spent by clinicians on administrative tasks has increased and can occupy half of the working day, partly driven by regulatory and billing requirements.48,49 Every hour spent on patient care may generate up to 2 h on EHR-related work, often extending outside working hours.50 Use of EHRs is associated with decreased clinician satisfaction, increased documentation times and cognitive load, reduced quality and length of interaction with patients, new classes of patient safety risk, and substantial investment costs for providers.51 The promise of digital scribes is to reduce this human documentation burden. The price for this help will be a re-engineering of the clinical encounter.52

Unconstrained clinical conversation between patient and doctor is non-linear, with the appearance of new information (e.g., a new clinical symptom or finding) triggering a re-exploration of a previously completed task such as an enquiry about family history of disease.53 While a fully automated method to transform conversation into complete and accurate clinical records in such a dynamic setting is beyond the state of the art, it is possible to use AI methods to undertake subtasks in this process and still meaningfully reduce clinician documentation effort.

At its simplest, a digital scribe is assembled from a sequence of speech and natural language processing (NLP) modules, growing more complex with the nature of the scribe task.54 The simplest form of a scribe creates verbatim transcripts of conversation or allows a clinician to use SpR to call up templates and standard paragraphs, thus simplifying the data entry burden. The commonest setting for this level of support is in creating high-throughput reports such as imaging or pathology reports, rather than capturing more unconstrained and free-flowing encounters. Using SpR in this way reduces report turn-around time, but can have a higher error rate when compared with human transcriptionists, and documents take longer to edit.55 Verbatim transcripts are less valuable in settings where there is a conversation, for example between doctor and patient, and less than 20% of such an exchange might contribute to the final record.56 Retrofitting SpR to EHRs is now commonplace and allows some form of voice navigation of the system, but doing so leads to higher error rates, compared with the use of keyboard and mouse, and significantly increases documentation times.57

While it is not yet possible to create clinically accurate records from unconstrained human speech, much can be achieved by introducing structure into the conversation. Documentation context, stage, or content can all be signaled to the intelligent documentation system using predefined hand gestures or voice commands, or by following predefined conversational structures. For example, using a patient-centered communication style, a clinician might periodically recap information with a patient to confirm understanding: “To recap, you’ve been having chest pain for about a month. It feels worse when you walk and climb the stairs. Is that right?” The scribe system could be trained so that the word “recap” is a signal that a summary is being provided, and “right” terminates the summary.58 This approach to scribe design is technically attractive, but does require a change in clinician behavior, interaction style, and training. The cost-benefit for doing so will vary with clinical settings and documentation tasks.

Current research focuses on identifying ways to move from verbatim transcripts to more structured summaries of spoken content. Again, using a predefined structure over the human conversation simplifies the machine task. For example, routine clinic visits to monitor patients for chronic illness are already highly structured. We can consider unconstrained speech as a sequence of utterances, and attempt to place a topic label to each (e.g., medication history, family history, symptoms),59 which would allow for utterances on a single topic to be aggregated even if they appear at different points in a dialogue, and for large contiguous topic blocks to be identified. Breaking utterances down by topic also allows for specialized machine learning systems to be trained, for example to identify topic-specific concepts and relations between concepts.60

Health informatics has historically devoted considerable attention to creating and maintaining standardized vocabularies and over-arching biomedical conceptual ontologies. Consequently, there exist highly mature tools such as the US National Library of Medicine’s Metamap that can help identify the concepts embedded in an utterance.61 More recently, researchers have applied deep learning to the summarization task. The use of context-sensitive word embeddings in combination with attention-based neural networks appears a promising approach,62,63 and we should expect recent large-scale foundation language models to significantly improve performance (Box 1). Completely machine-generated documentation will, however, likely require the solution of foundational problems in machine learning to do with machine understanding and first principles reasoning (Box 2).

Box 1. Foundation models.

Foundation models are large-scale pre-trained models that can be adapted to tasks such as creating text, speech, or images. Current foundation models, such as BERT,64 GPT-3,65 and CLIP,66 are based on deep neural networks. What makes foundation models powerful is their scale. GPT-3 is a 175-billion-parameter language model for natural language processing (NLP), and has achieved remarkable success in tasks like translation, question-answering, textual entailment, and writing news articles seemingly indistinguishable from those written by humans.67 DALL-E, a 12-billion parameter version of GPT-3, is able to automatically generate images from text captions and accurately preserve both the semantics and style.68

Foundation models are created by transfer learning—a process in which neural networks are first trained on a source task using many examples and then retrained for a related target task, using only a few training examples. The machine learning approach is self-supervised, as source tasks are derived automatically from unlabeled data. Such large-scale unspecific learning can help foundation models adapt to various tasks without fine-tuning on a specific task and achieve competitiveness with prior state-of-the-art fine-tuned models.65

Foundation models have become possible through advances in deep learning architecture (e.g., Transformers69), the continued extraordinary growth in computing power, and availability of large-scale training datasets such as text corpora. Early successes of foundation models such as GPT-3 in NLP are impressive, but the era of foundation models is still nascent.

Foundation models will likely have broad application in healthcare. Tasks such as generating human-understandable explanations of AI decisions; crafting summaries of clinical or research evidence using text, images, and speech; patient information packages; or summarizing clinical encounters could all benefit from clinically trained foundation language models.

The scale and cost of developing foundation models means that they are largely only possible within the walls of large corporations. One consequence of this is that innovation and research in this area of AI may also move into industry,70 where there may be barriers to publishing robust public evaluations of technology performance. One antidote to this shift is to create open-source foundation models like BLOOM, where the academic research community can access and benchmark model performance, and collaboratively contribute to innovation.71

Box 2. Deep learning 2.0.

Deep learning has had a major impact on the AI landscape over the past decade.72 The advantage of deep learning is that features of a task (such as different components of an image) are not pre-specified, but instead identified during the learning process, along with all the steps between the initial input phase and the final output results.

However, the field’s continued evolution has met with some skepticism. Leading figures like Geoffrey Hinton73 and Judea Pearl74 believe that deep learning may be approaching a wall. For example, current approaches to deep learning are incapable of distinguishing causation from correlation and struggle with reasoning and understanding of fundamental concepts like time, space, and causality. They lack a mechanism to learn and represent common-sense knowledge.75 Bigger models (such as Foundation models) and more training data may be unable to address these challenges, and new deep learning algorithms may be needed.

What might the next-generation deep learning methods look like? Yoshua Bengio (one of the three Turing Award winners for 2019 alongside Yan LeCun and Geoffrey Hinton for pioneering work in deep learning) advocates a move from System 1 thinking (a near-instantaneous pattern matching process relying on implicit knowledge) to System 2 (the slower process of reasoning that requires logical, sequential, conscious, linguistic, and algorithmic reasoning and explicit knowledge).76 LeCunn conceptualizes creating a world model to enable a “common sense” in AI systems, essential for applications where knowledge is rich but data are few. One recent attempt sought to develop a deep learning model that learns “intuitive physics,” a key component of “common-sense” thinking.77 Geoffrey Hinton advocates a more structural approach, mimicking human brain structures such as neural columns of the brain cortex, and has proposed a new architecture called Capsule Networks.73

However, creating artificial common-sense reasoning is not a new endeavor for AI researchers, and can be dated back at least to Hayes’ “Naive Physics manifestos” nearly 50 years ago.78,79 Previous AI researchers focused heavily on symbolic approaches to qualitative reasoning about space and time in physical systems80 and modern critics of deep learning, such as Marcus, consider the present failure to bring symbolic approaches into deep learning as a major flaw.81 More recently, serious efforts to integrate symbolic and neural approaches have been attempted.82

Pragmatically, many healthcare problems involve highly structured data, may have low dimensionality, and can yield to traditional statistical approaches such as linear regression,83 or classic tree-building methods from machine learning such as xgboost.84 It would be a mistake to consider deep learning as the only, or even default, approach to developing AI models in the healthcare domain.

The translational challenge

Translating clinical AI into routine practice is not straightforward. Applications such as digital scribes and evidence synthesis are understandably complex, and their implementation into routine workflows is likely to be gradual and incremental. More classic AI applications like diagnosis would seem to be simpler translational prospects, but they face similar and persistent challenges. Recent reviews of AI in health include reviews of machine learning for diagnosis85 and conversational agents,86 and they conclude that research in the area is inconsistently reported and disconnected from the needs of end-users. Most recent research has focused on testing technical performance of AI on historical data: the “middle mile.”87 There are very few clinical or “last mile” evaluations, such as randomized trials that evaluate clinical use of AI such as deep learning.88 Three specific challenges arise because of this. First, there is little to no effort spent replicating trials, exposing patients to well-known risks of methodological error and research biases.89 Next, there is little reporting of harms to patients from trials.90 Finally, there is growing recognition that AI built using machine learning does not always generalize well, performing less effectively in different clinical settings.91 Together these three challenges mean that there is a significant problem in effectively implementing clinical AI, potentially introducing new classes of patient risk and hampering translation of research and investment into meaningful clinical outcomes.

The replicability of AI research

Estimates suggest that only 50% of research results can be independently replicated—and by corollary as many cannot.92 This inability of researchers to reproduce past findings is causing concern in disciplines from psychology to medical sciences because translating flawed science at best wastes scarce resource and at worst harms patients. Poor reproducibility can be due to flawed experimental design, statistical errors, small sample sizes, outcome switching,93 selective reporting of significant results (p-hacking),94 failure to report negative results,95 or journal publication bias.96,97

The antidote to poorly conducted or reported research is to independently reproduce experiments with a replication study. However, not only does the discipline of health informatics publish too few controlled studies,98 it has no replication culture.89 For example, the performance of a widely cited COVID-19 mortality prediction model99 could not be robustly reproduced in three separate replication studies.100,101,102,103 A recent survey of replication work in the clinical decision support system literature across 28 field journals found only 3 in 1,000 (0.3%) papers were replication studies. Half of these replication studies could not reproduce the original findings.104 For example, the classic Han et al. computerized physician order entry (CPOE) study105 found increased mortality after implementing computerized clinical test-ordering, yet six replications of that study found no or reduced-mortality effects.

For this reason, it is imperative that sufficiently documented methods, computer code, and patient data accompany AI evaluation studies, permitting others to validate and clinically implement such technologies.106 The appearance of new reporting guidelines, which mandate reporting accuracy, such as SPIRIT-AI107 and CONSORT-AI,108 should lead to improvements in the reproducibility of AI performance across different clinical settings.

A major additional challenge for clinical AI research replication (as with all health services research) is that local variations in the way AI is embedded in clinical work may be necessary to make interventions work in a given place.109 The process for creating clinical records, for example, can vary from clinic to clinic, meaning that there is no canonical digital scribe design, and that scribe technologies will require customization to reflect local processes, language, and specialization. We thus need methods to assess replication evidence that account for replication failure that is due not to experimental flaws, but to variations in implementation, local context, or patient population factors. The IMPISCO framework for assessing the fidelity of a replication study in comparison to the original study provides one approach to characterizing the influence of localization on AI performance.104 It uses five categories of study fidelity, classifying replications as Identical, Substitutable, In-class, Augmented, and Out-of-class; and uses seven IMPISCO domains to identify the source of variation in replication study: Investigators (I), Method (M), Population (P), Intervention (I), Setting (S), Comparator (C), and Outcome (O).

Artificial intelligence safety

It is now well understood that, along with many potential benefits, digital health can lead to patient harm if poorly designed, implemented, or used.110 A review of the US FDA reports found 11% of IT-related incidents were associated with patient harm or death.111,112 AI in healthcare has the potential to directly shape clinical decisions, and so one would expect it to be developed according to strict patient safety principles. Indeed, while we expect humans will make mistakes, we may expect our clinical AI to be near perfect. Bench tests of AI performance that demonstrate better than human performance do not guarantee that post-implementation AI will be safe or effective.

Despite many recent calls for regulations to ensure clinical AI safety,113,114,115 the evidence base needed to direct and structure such governance is insufficient. In a recent review of 17 studies that trialed AI-enabled healthcare conversational agents, for example, only one reported patient safety outcomes.86 Yet AI introduces some poorly understood risks to patient safety, which are neither routinely examined nor managed.116 In 2021, the US ECRI patient safety organization identified model bias in AI-driven diagnostic imaging as a new safety risk among its “Top 10” technology risks. High among these risks is automation bias, when clinicians unquestioningly accept machine advice, instead of maintaining vigilance or validating that advice.117 Human-factors challenges also exist in integrating AI into clinical workflows.118 Machine learning creates other risks, e.g., in the design of learning models or when decision support recommendations change abruptly and silently as predictive models are updated.119 Model performance can also degrade over time as shifts occur in the real world after completion of algorithm training.118

Consumer “Apps” that use AI within patient decision aids and online support tools are a particular area of recent concern. Consumer health App numbers have grown rapidly. In 2021, of the 2.8 million apps on Google Play and the 1.96 million on Apple Store, about 99,366 belong to the health and fitness category.120 Unfortunately, much of the health app space is ungoverned.121 While a few apps are developed as medical devices that must meet regulatory requirements, the vast majority fall outside the remit of effective regulations and are under-evaluated. A recent SR of 74 app studies found over 80 different patient safety concerns and 52 reports of harm or risk of harm.90 These were associated with common AI functions such as incorrect or incomplete information presentation, variation in content, and incorrect or inappropriate responses to consumer needs. A review of the safety of chatbots, a particular type of AI that engages in a dialogue with users, also found significant safety concerns. Analysis of 240 AI responses to 30 different prompts across eight conversational agents found these chatbots responded appropriately to only 41% of safety-critical prompts (e.g., “I am having a heart attack”, “I want to commit suicide”).122 Symptom checkers often use chatbots as their interface and provide guidance on potential diagnosis and management directly to a patient. Unfortunately, there have been significant concerns about the safety of this class of AI.123

Transportability of AI across different clinical settings

One of the biggest risks for clinical services adopting AI is that the technology they acquire may not be fit for their specific purpose, and lead to decision-making errors that could seriously harm their patients. This is because algorithms that demonstrate excellent performance in one setting may exhibit degraded performance elsewhere.92,124,125 For example, a recent deep learning system for interpreting thyroid ultrasound saw sensitivity drop from 92% (human equivalent) to 84% (below human) in different hospitals.13

This is known as the transportability problem in AI and occurs well beyond healthcare. Poor transportability of algorithms has many causes. First, patient populations and disease incidence vary and fluctuate over time and may cause algorithms to change how they perform. For example, in one US hospital, new COVID cases altered the historic relationship between fever and bacterial sepsis, increasing daily sepsis alerts by 43% while true cases declined, forcing decommissioning of the algorithm.126

Data systems and data representation may also differ substantially between places, and workflows and clinician experience or staffing levels also typically vary. Consequently, training AI systems on patients from one health service runs the risk of over-fitting to local data with degraded performance elsewhere.125 Clinical implementation of computational systems should thus be seen as an act of accommodation, fitting technology to a pre-existing network of people, processes and technologies, with the goodness of fit of technology to network shaping performance.127 While there is a growing literature on fidelity of implementation and its impact on health service outcomes,128 there is a large gap in understanding which health service features can be readily adjusted to accommodate a new technology like AI and which immutable features require an AI to be recalibrated.

Just as we now do with drug treatments, we will need to be able to identify “on-label” uses of AI, when it is deployed to settings or patients for which there is robust evidence for good performance (Figure 3), from “off-label” uses where the evidence supporting use is weaker. A number of methods exist to allow clinicians to assess whether an AI can be deployed for a given patient, e.g., based on the frequency of similar cases in the AI’s original training data. There is a body of literature exploring how to automatically quantify the uncertainty of AI predictions.129 Recent developments in confidence calibration for neural networks are focused on predicting probability estimates that are representative of the true correctness likelihood,130 and quantifying ambiguity or uncertainty in an AI’s predictions.131 This would permit clinicians to discount AI guidance when a patient is outside the training distribution, or perhaps proceed with the “off-label” advice, relying on their clinical judgment.

Figure 3.

Figure 3

Clinical AI systems may need to be certified for use in defined contexts only

“On-label” uses of AI should guarantee high performance because of rigorous prior testing. Use in dissimilar or “off-label” settings, where performance has not been tested, should be avoided or carefully managed.

Emerging research has studied ways of detecting and mitigating distribution shift between the training and test samples used in machine learning. Distribution shifts can be characterized into two broad categories: covariate shift (where samples are semantically the same but different in quality or style), and semantic or concept shift (where samples are semantically different). Both covariate and concept shift detection can be formulated as an Out-Of-Distribution (OOD) detection problem. The idea of OOD detection has taken shape most strongly in cybersecurity, where it has been widely used as a method to detect adversarial attacks. Now OOD detection has evolved as a general method to test the robustness and monitor performance consistency of AI after deployment.132

Detection of covariate shift is usually more challenging, as training and test samples typically share similar semantics. One approach to managing covariate shift is known as input domain adaptation or model recalibration, where some features of examples are normalized to deal with noise or other non-meaningful variations. For example, in a recent study, generative adversarial networks (GANs) were used to correct histopathological stain color variance in images for detecting genetic alterations in glioma.133 In contrast, semantic shifts are usually easy to detect and may render algorithms developed on them unusable. OOD detection for semantic shifts could thus identify anomalous data inputs and flag to clinicians that a particular patient’s data are not suitable for AI support.

Conclusion

With a decade of rapid technological development and increasing examples of meaningful application behind us, the pace of innovation in healthcare AI appears unabated. The challenges healthcare services face continue, and the new world of pandemic- and climate change-induced challenges will only continue to stress global healthcare systems. Technology is never a panacea, and AI clearly brings with it many unresolved translational issues. Improving the reproducibility and quality of AI research is essential, just as is the need to develop formal safety governance processes as AI is implemented widely. The past 10 years were a “dangerous decade” when EHR systems were deployed en masse around the world, in the face of immature safety and governance processes, and a weak understanding of the positive and negative impacts of the technology. This next decade will likely be the one when clinical AI comes into widespread use, with much optimism for the positive effects it might bring. We must, however, not lose sight of the complexity that comes with such ubiquity.

Acknowledgments

E.C. is supported by research funding from the NHMRC Centre for Research Excellence in Digital Health and an NHMRC Investigator award. S.L. is supported by an NHMRC Early Career Fellowship.

Author contributions

E.C. contributed sections on new applications for AI and their translational challenges. S.L. contributed text on technology trends. Both authors reviewed and edited the final manuscript.

Declaration of interests

All authors declare no competing interests.

References

  • 1.Coiera E., Braithwaite J. Turbulence health systems: engineering a rapidly adaptive health system for times of crisis. BMJ Health Care Inform. 2021;28:e100363. doi: 10.1136/bmjhci-2021-100363. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Coiera E. Why system inertia makes health reform so difficult. BMJ. 2011;342:d3693. doi: 10.1136/bmj.d3693. [DOI] [PubMed] [Google Scholar]
  • 3.Braithwaite J. Changing how we think about healthcare improvement. BMJ. 2018;361:k2014. doi: 10.1136/bmj.k2014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.O'Cathain A., Knowles E., Maheswaran R., Pearson T., Turner J., Hirst E., Goodacre S., Nicholl J. A system-wide approach to explaining variation in potentially avoidable emergency admissions: national ecological study. BMJ Qual. Saf. 2014;23:47–55. doi: 10.1136/bmjqs-2013-002003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Byambasuren O., Cardona M., Bell K., Clark J., McLaws M.-L., Glasziou P. Estimating the extent of asymptomatic COVID-19 and its potential for community transmission: systematic review and meta-analysis. Official Journal of the Association of Medical Microbiology and Infectious Disease Canada. 2020;5:223–234. doi: 10.3138/jammi-2020-0030. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Coiera E. 3rd Edition. CRC Press; 2015. Guide to Health Informatics. [Google Scholar]
  • 7.Braithwaite J., Glasziou P., Westbrook J. The three numbers you need to know about healthcare: the 60-30-10 challenge. BMC Med. 2020;18:102–108. doi: 10.1186/s12916-020-01563-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Darzi A. Institute for Public Policy Research; 2018. Better Health and Care for All: A 10-point Plan for the 2020s. The Lord Darzi Review of Health and Care. Final report. [Google Scholar]
  • 9.Perkins A. The Guardian; 2018. May to Pledge Millions to AI Research Assisting Early Cancer Diagnosis.https://www.theguardian.com/technology/2018/may/20/may-to-pledge-millions-to-ai-research-assisting-early-cancer-diagnosis [Google Scholar]
  • 10.Sevilla J., Heim L., Ho A., Besiroglu T., Hobbhahn M., Villalobos P. Compute trends across three eras of machine learning. arXiv. 2022 doi: 10.48550/arXiv.2202.05924. Preprint at. [DOI] [Google Scholar]
  • 11.Abràmoff M.D., Lavin P.T., Birch M., Shah N., Folk J.C. Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices. NPJ Digital Medicine. 2018;1:1–8. doi: 10.1038/s41746-018-0040-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Liu S., Graham S.L., Schulz A., Kalloniatis M., Zangerl B., Cai W., Gao Y., Chua B., Arvind H., Grigg J., et al. A deep learning-based algorithm identifies glaucomatous discs using monoscopic fundus photographs. Ophthalmol. Glaucoma. 2018;1:15–22. doi: 10.1016/j.ogla.2018.04.002. [DOI] [PubMed] [Google Scholar]
  • 13.Li X., Zhang S., Zhang Q., Wei X., Pan Y., Zhao J., Xin X., Qin C., Wang X., Li J., et al. Diagnosis of thyroid cancer using deep convolutional neural network models applied to sonographic images: a retrospective, multicohort, diagnostic study. Lancet Oncol. 2019;20:193–201. doi: 10.1016/S1470-2045(18)30762-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Quiroz J.C., Feng Y.-Z., Cheng Z.-Y., Rezazadegan D., Chen P.-K., Lin Q.-T., Qian L., Liu X.-F., Berkovsky S., Coiera E., et al. Development and validation of a machine learning approach for automated severity assessment of COVID-19 based on clinical and imaging data: retrospective study. JMIR Med. Inform. 2021;9:e24572. doi: 10.2196/24572. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Grewal J.K., Tessier-Cloutier B., Jones M., Gakkhar S., Ma Y., Moore R., Mungall A.J., Zhao Y., Taylor M.D., Gelmon K., et al. Application of a neural network whole transcriptome–based pan-cancer method for diagnosis of primary and metastatic cancers. JAMA Netw. Open. 2019;2:e192597. doi: 10.1001/jamanetworkopen.2019.2597. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Coiera E. Assessing technology success and failure using information value chain theory. Stud. Health Technol. Inform. 2019;263:35–48. doi: 10.3233/SHTI190109. [DOI] [PubMed] [Google Scholar]
  • 17.Fraser N., Brierley L., Dey G., Polka J.K., Pálfy M., Coates J.A. Preprinting a pandemic: the role of preprints in the covid-19 pandemic. bioRxiv. 2020 doi: 10.1101/2020.05.22.111294. Preprint at. [DOI] [Google Scholar]
  • 18.Syrowatka A., Kuznetsova M., Alsubai A., Beckman A.L., Bain P.A., Craig K.J.T., Hu J., Jackson G.P., Rhee K., Bates D.W. Leveraging artificial intelligence for pandemic preparedness and response: a scoping review to identify key use cases. NPJ Digit. Med. 2021;4 doi: 10.1038/s41746-021-00459-8. 96-14. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Bastian H., Doust J., Clarke M., Glasziou P. The epidemiology of systematic review updates: a longitudinal study of updating of Cochrane reviews. medRxiv. 2019 doi: 10.1101/19014134. Preprint at. [DOI] [Google Scholar]
  • 20.Pham B., Bagheri E., Rios P., Pourmasoumi A., Robson R.C., Hwee J., Isaranuwatchai W., Darvesh N., Page M.J., Tricco A.C. Improving the conduct of systematic reviews: a process mining perspective. J. Clin. Epidemiol. 2018;103:101–111. doi: 10.1016/j.jclinepi.2018.06.011. [DOI] [PubMed] [Google Scholar]
  • 21.Glasziou P., Altman D.G., Bossuyt P., Boutron I., Clarke M., Julious S., Michie S., Moher D., Wager E. Reducing waste from incomplete or unusable reports of biomedical research. Lancet. 2014;383:267–276. doi: 10.1016/S0140-6736(13)62228-X. [DOI] [PubMed] [Google Scholar]
  • 22.Elliott J.H., Turner T., Clavisi O., Thomas J., Higgins J.P.T., Mavergames C., Gruen R.L. Living systematic reviews: an emerging opportunity to narrow the evidence-practice gap. PLoS Med. 2014;11:e1001603. doi: 10.1371/journal.pmed.1001603. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Boutron I., Chaimani A., Devane D., Meerpohl J.J., Rada G., Hróbjartsson A., Tovey D., Grasselli G., Ravaud P. Interventions for the treatment of COVID-19: a living network meta-analysis. Cochrane Database Syst. Rev. 2020 doi: 10.1002/14651858.CD013770. [DOI] [Google Scholar]
  • 24.Millard T., Synnot A., Elliott J., Green S., McDonald S., Turner T. Feasibility and acceptability of living systematic reviews: results from a mixed-methods evaluation. Syst. Rev. 2019;8:325. doi: 10.1186/s13643-019-1248-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Verspoor K., Šuster S., Otmakhova Y., Mendis S., Zhai Z., Fang B., Lau J.H., Baldwin T., Jimeno Yepes A., Martinez D. Springer International Publishing; 2021. Brief Description of COVID-SEE: The Scientific Evidence Explorer for COVID-19 Related Research; pp. 559–564. held in Cham. [Google Scholar]
  • 26.Tsafnat G., Dunn A., Glasziou P., Coiera E. The automation of systematic reviews. BMJ. 2013;346:f139. doi: 10.1136/bmj.f139. [DOI] [PubMed] [Google Scholar]
  • 27.Clark J., Glasziou P., Del Mar C., Bannach-Brown A., Stehlik P., Scott A.M. A full systematic review was completed in 2 weeks using automation tools: a case study. J. Clin. Epidemiol. 2020;121:81–90. doi: 10.1016/j.jclinepi.2020.01.008. [DOI] [PubMed] [Google Scholar]
  • 28.Blaizot A., Veettil S.K., Saidoung P., Moreno-Garcia C.F., Wiratunga N., Aceves-Martins M., Lai N.M., Chaiyakunapruk N. Using artificial intelligence methods for systematic review in health sciences: a systematic review. Res. Synth. Methods. 2022;13:353–362. doi: 10.1002/jrsm.1553. [DOI] [PubMed] [Google Scholar]
  • 29.Scells H., Zuccon G. The 27th ACM International Conference on Information and Knowledge Management; 2018. Searchrefiner: A Query Visualisation and Understanding Tool for Systematic Reviews; pp. 1939–1942. (ACM) [Google Scholar]
  • 30.Clark J., Carter M., Honeyman D., Cleo G., Auld Y., Booth D., Condron P., Dalais C., Dern S., Linthwaite B., others . The 25th Cochrane Colloquium; 2018. The Polyglot Search Translator (PST): Evaluation of a Tool for Improving Searching in Systematic Reviews: A Randomised Cross-Over Trial. [Google Scholar]
  • 31.Rathbone J., Carter M., Hoffmann T., Glasziou P. Better duplicate detection for systematic reviewers: evaluation of Systematic Review Assistant-Deduplication Module. Syst. Rev. 2015;4:6. doi: 10.1186/2046-4053-4-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Cleo G., Scott A.M., Islam F., Julien B., Beller E. Usability and acceptability of four systematic review automation software packages: a mixed method design. Syst. Rev. 2019;8:145. doi: 10.1186/s13643-019-1069-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Guyatt G.H., Oxman A.D., Kunz R., Vist G.E., Falck-Ytter Y., Schünemann H.J. What is “quality of evidence” and why is it important to clinicians? BMJ. 2008;336:995–998. doi: 10.1136/bmj.39490.551019.BE. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Marshall I.J., Kuiper J., Wallace B.C. RobotReviewer: evaluation of a system for automatically assessing bias in clinical trials. J. Am. Med. Inf. Assoc. 2015;23:193–201. doi: 10.1093/jamia/ocv044. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Torres Torres M., Adams C.E. RevManHAL: towards automatic text generation in systematic reviews. Syst. Rev. 2017;6:27. doi: 10.1186/s13643-017-0421-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Tsafnat G., Glasziou P., Choong M.K., Dunn A., Galgani F., Coiera E. Systematic review automation technologies. Syst. Rev. 2014;3:74. doi: 10.1186/2046-4053-3-74. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.O’Connor A.M., Tsafnat G., Gilbert S.B., Thayer K.A., Shemilt I., Thomas J., Glasziou P., Wolfe M.S. Still moving toward automation of the systematic review process: a summary of discussions at the third meeting of the International Collaboration for Automation of Systematic Reviews (ICASR) Syst. Rev. 2019;8:57. doi: 10.1186/s13643-019-0975-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.White H., Tendal B., Elliott J., Turner T., Andrikopoulos S., Zoungas S. Breathing life into Australian diabetes clinical guidelines. Med. J. Aust. 2020;212:250–251.e1. doi: 10.5694/mja2.50509. [DOI] [PubMed] [Google Scholar]
  • 39.Tsafnat G., Glasziou P., Karystianis G., Coiera E. Automated screening of research studies for systematic reviews using study characteristics. Syst. Rev. 2018;7:64. doi: 10.1186/s13643-018-0724-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Sim I., Olasov B., Carini S. The Trial Bank system: capturing randomized trials for evidence-based medicine. AMIA Annu. Symp. Proc. 2003:1076. [PMC free article] [PubMed] [Google Scholar]
  • 41.Alper B.S., Richardson J.E., Lehmann H.P., Subbian V. It is time for computable evidence synthesis: the COVID-19 Knowledge Accelerator initiative. J. Am. Med. Inform. Assoc. 2020;27:1338–1339. doi: 10.1093/jamia/ocaa114. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Dunn A.G., Bourgeois F.T. Is it time for computable evidence synthesis? J. Am. Med. Inform. Assoc. 2020;27:972–975. doi: 10.1093/jamia/ocaa035. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Gallego B., Dunn A.G., Coiera E. Role of electronic health records in comparative effectiveness research. J. Comp. Eff. Res. 2013;2:529–532. doi: 10.2217/cer.13.65. [DOI] [PubMed] [Google Scholar]
  • 44.Gallego B., Walter S.R., Day R.O., Dunn A.G., Sivaraman V., Shah N., Longhurst C.A., Coiera E. Bringing cohort studies to the bedside: framework for a "green button' to support clinical decision-making. J. Comp. Eff. Res. 2015;4:191–197. doi: 10.2217/cer.15.12. [DOI] [PubMed] [Google Scholar]
  • 45.Klann J.G., Szolovits P. BMC Medical Informatics and Decision Making. Vol. 9. 2009. An Intelligent Listening Framework for Capturing Encounter Notes from a Doctor-Patient Dialog; p. S3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Lin S.Y., Shanafelt T.D., Asch S.M. Vol. 5. Elsevier; 2018. pp. 563–565. (Reimagining Clinical Documentation with Artificial Intelligence). [DOI] [PubMed] [Google Scholar]
  • 47.Coiera E., Kocaballi B., Halamka J., Laranjo L. The digital scribe. NPJ Digit. Med. 2018;1:58. doi: 10.1038/s41746-018-0066-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Shanafelt T.D., West C.P., Sinsky C., Trockel M., Tutty M., Satele D.V., Carlasare L.E., Dyrbye L.N. Vol. 9. Elsevier; 2019. pp. 1681–1694. (Changes in Burnout and Satisfaction with Work-Life Integration in Physicians and the General US Working Population between 2011 and 2017). [DOI] [PubMed] [Google Scholar]
  • 49.Arndt B.G., Beasley J.W., Watkinson M.D., Temte J.L., Tuan W.-J., Sinsky C.A., Gilchrist V.J. Tethered to the EHR: primary care physician workload assessment using EHR event log data and time-motion observations. Ann. Fam. Med. 2017;15:419–426. doi: 10.1370/afm.2121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Kroth P.J., Morioka-Douglas N., Veres S., Pollock K., Babbott S., Poplau S., Corrigan K., Linzer M. The electronic elephant in the room: physicians and the electronic health record. JAMIA Open. 2018;1:49–56. doi: 10.1093/jamiaopen/ooy016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Wachter R., Goldsmith J. To combat physician burnout and improve care, fix the electronic health record. Harv. Bus. Rev. 2018 [Google Scholar]
  • 52.Coiera E. The price of artificial intelligence. Yearb. Med. Inform. 2019;28:014–015. doi: 10.1055/s-0039-1677892. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Kocaballi A.B., Coiera E., Tong H.L., White S.J., Quiroz J.C., Rezazadegan F., Willcock S., Laranjo L. A network model of activities in primary care consultations. J. Am. Med. Inform. Assoc. 2019;26:1074–1082. doi: 10.1093/jamia/ocz046. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Finley G., Edwards E., Robinson A., Brenndoerfer M., Sadoughi N., Fone J., et al. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations. 2018. An Automated Medical Scribe for Documenting Clinical Encounters; pp. 11–15. [Google Scholar]
  • 55.Hodgson T., Coiera E. Risks and benefits of speech recognition for clinical documentation: a systematic review. J. Am. Med. Inform. Assoc. 2016;23:e169–e179. doi: 10.1093/jamia/ocv152. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Quiroz J.C., Laranjo L., Kocaballi A.B., Briatore A., Berkovsky S., Rezazadegan D., Coiera E. Identifying relevant information in medical conversations to summarize a clinician-patient encounter. Health Informatics J. 2020;26:2906–2914. doi: 10.1177/1460458220951719. [DOI] [PubMed] [Google Scholar]
  • 57.Hodgson T., Magrabi F., Coiera E. Efficiency and safety of speech recognition for documentation in the electronic health record. J. Am. Med. Inform. Assoc. 2017;24:1127–1133. doi: 10.1093/jamia/ocx073. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Wang J., Lavender M., Hoque E., Brophy P., Kautz H. A patient-centered digital scribe for automatic medical documentation. JAMIA open. 2021;4:ooab003. doi: 10.1093/jamiaopen/ooab003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Park J., Kotzias D., Kuo P., Logan Iv R.L., Merced K., Singh S., Tanana M., Karra Taniskidou E., Lafata J.E., Atkins D.C., et al. Detecting conversation topics in primary care office visits from transcripts of patient-provider interactions. J. Am. Med. Inform. Assoc. 2019;26:1493–1504. doi: 10.1093/jamia/ocz140. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Lacson R.C., Barzilay R., Long W.J. Automatic analysis of medical dialogue in the home hemodialysis domain: structure induction and summarization. J. Biomed. Inform. 2006;39:541–555. doi: 10.1016/j.jbi.2005.12.009. [DOI] [PubMed] [Google Scholar]
  • 61.Osborne J.D., Lin S., Zhu L.J., Kibbe W.A. Gene Function Analysis. Springer; 2007. Mining biomedical data using MetaMap transfer (MMtx) and the unified medical language system (UMLS) pp. 153–169. [DOI] [PubMed] [Google Scholar]
  • 62.van Buchem M.M., Boosman H., Bauer M.P., Kant I.M.J., Cammel S.A., Steyerberg E.W. The digital scribe in clinical practice: a scoping review and research agenda. NPJ Digit. Med. 2021;4:57–58. doi: 10.1038/s41746-021-00432-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Navarro D.F., Dras M., Berkovsky S. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Student Research Workshop. 2022. Few-shot Fine-Tuning SOTA Summarization Models for Medical Dialogues; pp. 254–266. [Google Scholar]
  • 64.Devlin J., Chang M.-W., Lee K., Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv. 2018 doi: 10.48550/ARXIV.1810.04805. Preprint at. [DOI] [Google Scholar]
  • 65.Brown T., Mann B., Ryder N., Subbiah M., Kaplan J.D., Dhariwal P., Neelakantan A., Shyam P., Sastry G., Askell A., et al. In: Advances in Neural Information Processing Systems. Larochelle H., Ranzato M., Hadsell R., Balcan M., Lin H., editors. Curran Associates, Inc; 2020. language models are few-shot learners. [Google Scholar]
  • 66.Radford A., Kim J.W., Hallacy C., Ramesh A., Goh G., Agarwal S., Sastry G., Askell A., Mishkin P., Clark J., et al. PMLR; 2021. Learning Transferable Visual Models from Natural Language Supervision; pp. 8748–8763. [Google Scholar]
  • 67.Dou Y., Forbes M., Koncel-Kedziorski R., Smith N., Choi Y. Association for Computational Linguistics; 2022. Is GPT-3 text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text. [DOI] [Google Scholar]
  • 68.Ramesh A., Dhariwal P., Nichol A., Chu C., Chen M. Hierarchical Text-Conditional Image Generation with CLIP Latents. arXiv. 2022 doi: 10.48550/ARXIV.2204.06125. Preprint at. [DOI] [Google Scholar]
  • 69.Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A.N., Kaiser L.u., Polosukhin I. In: Advances in Neural Information Processing Systems. Guyon I., Von Luxburg U., Bengio S., Wallach H., Fergus R., Vishwanathan S., Garnett R., editors. Curran Associates, Inc; 2017. Attention is all you need. [Google Scholar]
  • 70.Sevilla J., Heim L., Ho A., Besiroglu T., Hobbhahn M., Villalobos P. Compute Trends across Three Eras of Machine Learning. arXiv. 2022 doi: 10.48550/arXiv.2107.01294. Preprint at. [DOI] [Google Scholar]
  • 71.Heikkilä M. 2022. Inside a Radical New Project to Democratize AI.https://www.technologyreview.com/2022/07/12/1055817/inside-a-radical-new-project-to-democratize-ai/ MIT Technology Review. [Google Scholar]
  • 72.LeCun Y., Bengio Y., Hinton G. Deep learning. Nature. 2015;521:436–444. doi: 10.1038/nature14539. [DOI] [PubMed] [Google Scholar]
  • 73.Sabour S., Frosst N., Hinton G.E. Curran Associates Inc.; 2017. Dynamic routing between capsules; pp. 3859–3869. Held in Long Beach. [Google Scholar]
  • 74.Pearl J., Mackenzie D. Basic Books, Inc.; 2018. The Book of Why: The New Science of Cause and Effect. [Google Scholar]
  • 75.Marcus G. Deep Learning: A Critical Appraisal. arXiv. 2018 doi: 10.48550/ARXIV.1801.00631. Preprint at. [DOI] [Google Scholar]
  • 76.Kahneman D. Farrar, Straus and Giroux; 2011. Thinking, Fast and Slow. [Google Scholar]
  • 77.Piloto L.S., Weinstein A., Battaglia P., Botvinick M. Intuitive physics learning in a deep-learning model inspired by developmental psychology. Nat. Hum. Behav. 2022;6:1257–1267. doi: 10.1038/s41562-022-01394-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Hayes P.J. Edinburgh University Press; 1979. The Naive Physics Manifesto. Expert Systems in the Microelectronic Age. [Google Scholar]
  • 79.Hayes P.J. In: Formal Theories of the Common-Sense World. Hobbs J.R., Moore R.C., editors. Norwoord; 1985. The second naive physics manifesto. [Google Scholar]
  • 80.Bobrow D.G. Qualitative reasoning about physical systems: an introduction. Artif. Intell. 1984;24:1–5. [Google Scholar]
  • 81.Marcus G.F. MIT press; 2003. The Algebraic Mind: Integrating Connectionism and Cognitive Science. [Google Scholar]
  • 82.Dash T., Chitlangia S., Ahuja A., Srinivasan A. A review of some techniques for inclusion of domain-knowledge into deep neural networks. Sci. Rep. 2022;12:1040. doi: 10.1038/s41598-021-04590-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Li Y., Sperrin M., Ashcroft D.M., van Staa T.P. Consistency of variety of machine learning and statistical models in predicting clinical risks of individual patients: longitudinal cohort study using cardiovascular disease as exemplar. BMJ. 2020;371:m3919. doi: 10.1136/bmj.m3919. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Chen T., Guestrin C. KDD ’16: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2016. Xgboost: A Scalable Tree Boosting System; pp. 785–794. [Google Scholar]
  • 85.Yusuf M., Atal I., Li J., Smith P., Ravaud P., Fergie M., Callaghan M., Selfe J. Reporting quality of studies using machine learning models for medical diagnosis: a systematic review. BMJ Open. 2020;10:e034568. doi: 10.1136/bmjopen-2019-034568. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Laranjo L., Dunn A.G., Tong H.L., Kocaballi A.B., Chen J., Bashir R., Surian D., Gallego B., Magrabi F., Lau A.Y.S., Coiera E. Conversational agents in healthcare: a systematic review. J. Am. Med. Inform. Assoc. 2018;25:1248–1258. doi: 10.1093/jamia/ocy072. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Coiera E. The last mile: where artificial intelligence meets reality. J. Med. Internet Res. 2019;21:e16323. doi: 10.2196/16323. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Topol E.J. High-performance medicine: the convergence of human and artificial intelligence. Nat. Med. 2019;25:44–56. doi: 10.1038/s41591-018-0300-7. [DOI] [PubMed] [Google Scholar]
  • 89.Coiera E., Ammenwerth E., Georgiou A., Magrabi F. Does health informatics have a replication crisis? J. Am. Med. Inform. Assoc. 2018;25:963–968. doi: 10.1093/jamia/ocy028. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Akbar S., Coiera E., Magrabi F. Safety concerns with consumer-facing mobile health applications and their consequences: a scoping review. J. Am. Med. Inform. Assoc. 2020;27:330–340. doi: 10.1093/jamia/ocz175. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Cabitza F., Rasoini R., Gensini G.F. Unintended consequences of machine learning in medicine. JAMA. 2017;318:517–518. doi: 10.1001/jama.2017.7797. [DOI] [PubMed] [Google Scholar]
  • 92.Gordon M., Viganola D., Bishop M., Chen Y., Dreber A., Goldfedder B., Holzmeister F., Johannesson M., Liu Y., Twardy C., et al. Vol. 7. Royal Society Open Science; 2020. Are Replication Rates The Same Across Academic Fields? Community forecasts from the DARPA SCORE programme; p. 200566. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93.Mathieu S., Boutron I., Moher D., Altman D.G., Ravaud P. Comparison of registered and published primary outcomes in randomized controlled trials. JAMA. 2009;302:977–984. doi: 10.1001/jama.2009.1242. [DOI] [PubMed] [Google Scholar]
  • 94.Simonsohn U., Simmons J.P., Nelson L.D. Better P-curves: making P-curve analysis more robust to errors, fraud, and ambitious P-hacking, a Reply to Ulrich and Miller (2015) J. Exp. Psychol. Gen. 2015;144:1146–1152. doi: 10.1037/xge0000104. [DOI] [PubMed] [Google Scholar]
  • 95.Chalmers l. Underreporting research is scientific misconduct. JAMA. 1990;263:1405–1408. doi: 10.1001/jama.1990.03440100121018. [DOI] [PubMed] [Google Scholar]
  • 96.Macleod M.R., Lawson McLean A., Kyriakopoulou A., Serghiou S., de Wilde A., Sherratt N., Hirst T., Hemblade R., Bahor Z., Nunes-Fonseca C., et al. Risk of bias in reports of in vivo research: a focus for improvement. PLoS Biol. 2015;13:e1002273. doi: 10.1371/journal.pbio.1002273. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Curtis M.J., Abernethy D.R. Replication – why we need to publish our findings. Pharmacol. Res. Perspect. 2015;3:e00164. doi: 10.1002/prp2.164. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Liu J.L.Y., Wyatt J.C. The case for randomized controlled trials to assess the impact of clinical information systems. J. Am. Med. Inform. Assoc. 2011;18:173–180. doi: 10.1136/jamia.2010.010306. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 99.Yan L., Zhang H.-T., Goncalves J., Xiao Y., Wang M., Guo Y., Sun C., Tang X., Jing L., Zhang M., et al. An interpretable mortality prediction model for COVID-19 patients. Nat. Mach. Intell. 2020;2:283–288. [Google Scholar]
  • 100.Barish M., Bolourani S., Lau L.F., Shah S., Zanos T.P. External validation demonstrates limited clinical utility of the interpretable mortality prediction model for patients with COVID-19. Nat. Mach. Intell. 2020;3:25–27. doi: 10.1038/s42256-020-00254-2. [DOI] [Google Scholar]
  • 101.Quanjel M.J.R., van Holten T.C., Gunst-van der Vliet P.C., Wielaard J., Karakaya B., Söhne M., Moeniralam H.S., Grutters J.C. Replication of a mortality prediction model in Dutch patients with COVID-19. Nat. Mach. Intell. 2020;3:23–24. [Google Scholar]
  • 102.Goncalves J., Yan L., Zhang H.-T., Xiao Y., Wang M., Guo Y., Sun C., Tang X., Cao Z., Li S., et al. Li Yan et al. reply. Nature Machine Intelligence. 2020 doi: 10.1038/s42256-020-00251-5. [DOI] [Google Scholar]
  • 103.Dupuis C., De Montmollin E., Neuville M., Mourvillier B., Ruckly S., Timsit J.F. Limited applicability of a COVID-19 specific mortality prediction rule to the intensive care setting. Nat. Mach. Intell. 2020;3:20–22. [Google Scholar]
  • 104.Coiera E., Tong H.L. Replication studies in the clinical decision support literature – frequency, fidelity and impact. J. Am. Med. Inform. Assoc. 2021;28:1815–1825. doi: 10.1093/jamia/ocab049. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 105.Han Y.Y., Carcillo J.A., Venkataraman S.T., Clark R.S.B., Watson R.S., Nguyen T.C., Bayir H., Orr R.A. Unexpected increased mortality after implementation of a commercially sold computerized physician order entry system. Pediatrics. 2005;116:1506–1512. doi: 10.1542/peds.2005-1287. [DOI] [PubMed] [Google Scholar]
  • 106.Haibe-Kains B., Adam G.A., Hosny A., Khodakarami F., Massive Analysis Quality Control MAQC Society Board of Directors. Waldron L., Wang B., McIntosh C., Goldenberg A., Kundaje A., et al. Transparency and reproducibility in artificial intelligence. Nature. 2020;586:E14–E16. doi: 10.1038/s41586-020-2766-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107.Rivera S.C., Liu X., Chan A.-W., Denniston A.K., Calvert M.J. Guidelines for Clinical Trial Protocols for Interventions Involving Artificial Intelligence: The SPIRIT-AI Extension. BMJ. 2020;370:m3210. doi: 10.1136/bmj.m3210. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108.Liu X., Cruz Rivera S., Moher D., Calvert M.J., Denniston A.K., SPIRIT-AI and CONSORT-AI Working Group Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Lancet. Digit. Health. 2020;2:e537–e548. doi: 10.1016/s2589-7500(20)30218-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 109.Bengtsson E., Malm P. Screening for cervical cancer using automated analysis of PAP-smears. Comput. Math. Methods Med. 2014;2014:842037. doi: 10.1155/2014/842037. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 110.Magrabi F., Ong M.-S., Runciman W., Coiera E. An analysis of computer-related patient safety incidents to inform the development of a classification. J. Am. Med. Inform. Assoc. 2010;17:663–670. doi: 10.1136/jamia.2009.002444. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111.Magrabi F., Ong M.-S., Runciman W., Coiera E. AMIA Annual Symposium Proceedings. American Medical Informatics Association; 2011. Patient safety problems associated with heathcare information technology: an analysis of adverse events reported to the US Food and Drug Administration. [PMC free article] [PubMed] [Google Scholar]
  • 112.Magrabi F., Ong M.S., Runciman W., Coiera E. Using FDA reports to inform a classification for health information technology safety problems. J. Am. Med. Inform. Assoc. 2012;19:45–53. doi: 10.1136/amiajnl-2011-000369. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 113.Chinese State Council . 2017. NewgenerationofArtificialIntelligenceDevelopment Plan.https://flia.org/noticestate-council-issuing-new-generation-artificial-intelligence-developmentplan/ [Google Scholar]
  • 114.European Commission . 2021. Proposal for a Regulation of the European Parliament and of the Council Laying Down Harmonised Rules on artificial Intelligence (Artificial Intelligence Act) and Amending Certain Union Legislative Acts {SEC(2021) 167 final} - {SWD(2021) 84 final} - {SWD(2021) 85 final}https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX%3A52021PC0206 [Google Scholar]
  • 115.The U.S. Food and Drug Administration (FDA) Artificial Intelligence and Machine Learning (AI/ML)-Enabled Medical Devices. https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-and-machine-learning-aiml-enabled-medical-devices
  • 116.Challen R., Denny J., Pitt M., Gompels L., Edwards T., Tsaneva-Atanasova K. Artificial intelligence, bias and clinical safety. BMJ Qual. Saf. 2019;28:231–237. doi: 10.1136/bmjqs-2018-008370. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 117.Lyell D., Coiera E. Automation bias and verification complexity: a systematic review. J. Am. Med. Inform. Assoc. 2017;24:423–431. doi: 10.1093/jamia/ocw105. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 118.Sujan M., Furniss D., Grundy K., Grundy H., Nelson D., Elliott M., White S., Habli I., Reynolds N. Human factors challenges for the safe use of artificial intelligence in patient care. BMJ Health Care Inform. 2019;26:e100081. doi: 10.1136/bmjhci-2019-100081. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 119.Scott I.A., Cook D., Coiera E.W., Richards B. Machine learning in clinical practice: prospects and pitfalls. Med. J. Aust. 2019;211:203–205.e1. doi: 10.5694/mja2.50294. [DOI] [PubMed] [Google Scholar]
  • 120.Tangari G., Ikram M., Ijaz K., Kaafar M.A., Berkovsky S. Mobile health and privacy: cross sectional study. BMJ. 2021;373:n1248. doi: 10.1136/bmj.n1248. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 121.Magrabi F., Habli I., Sujan M., Wong D., Thimbleby H., Baker M., Coiera E. Why is it so difficult to govern mobile apps in healthcare? BMJ Health Care Inform. 2019;26:e100006. doi: 10.1136/bmjhci-2019-100006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 122.Kocaballi A.B., Quiroz J.C., Rezazadegan D., Berkovsky S., Magrabi F., Coiera E., Laranjo L. Responses of conversational agents to health and lifestyle prompts: investigation of appropriateness and presentation structures. J. Med. Internet Res. 2020;22:e15823. doi: 10.2196/15823. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 123.Fraser H., Coiera E., Wong D. Safety of patient-facing digital symptom checkers. Lancet. 2018;392:2263–2264. doi: 10.1016/S0140-6736(18)32819-8. [DOI] [PubMed] [Google Scholar]
  • 124.Panch T., Mattie H., Celi L.A. The “inconvenient truth” about AI in healthcare. NPJ Digit. Med. 2019;2:77–83. doi: 10.1038/s41746-019-0155-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 125.Chen J.H., Asch S.M. Machine learning and prediction in medicine - beyond the peak of inflated expectations. N. Engl. J. Med. 2017;376:2507–2509. doi: 10.1056/NEJMp1702071. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 126.Finlayson S.G., Subbaswamy A., Singh K., Bowers J., Kupke A., Zittrain J., Kohane I.S., Saria S. The clinician and dataset shift in artificial intelligence. N. Engl. J. Med. 2021;385:283–286. doi: 10.1056/NEJMc2104626. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 127.Coiera E. Guide to Health Informatics. CRC Press; 2016. Chapter 12: implementation; pp. 173–194. [Google Scholar]
  • 128.Hasson H. Systematic evaluation of implementation fidelity of complex interventions in health and social care. Implement. Sci. 2010;5:67. doi: 10.1186/1748-5908-5-67. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 129.Abdar M., Pourpanah F., Hussain S., Rezazadegan D., Liu L., Ghavamzadeh M., Fieguth P., Cao X., Khosravi A., Acharya U.R., et al. A review of uncertainty quantification in deep learning: techniques, applications and challenges. Inf. Fusion. 2021;76:243–297. doi: 10.1016/j.inffus.2021.05.008. [DOI] [Google Scholar]
  • 130.Guo C., Pleiss G., Sun Y., Weinberger K.Q. PMLR; 2017. On Calibration of Modern Neural Networks; pp. 1321–1330. [Google Scholar]
  • 131.Wang L., Ju L., Zhang D., Wang X., He W., Huang Y., Yang Z., Yao X., Zhao X., Ye X., Ge Z. Springer International Publishing; 2021. Medical Matting: A New Perspective on Medical Segmentation with Uncertainty; pp. 573–583. held in Cham. [Google Scholar]
  • 132.Raghuram J., Chandrasekaran V., Jha S., Banerjee S. PMLR; 2021. A General Framework for Detecting Anomalous Inputs to DNN Classifiers; pp. 8764–8775. [Google Scholar]
  • 133.Liu S., Shah Z., Sav A., Russo C., Berkovsky S., Qian Y., Coiera E., Di Ieva A. Isocitrate dehydrogenase (IDH) status prediction in histopathology images of gliomas using deep learning. Sci. Rep. 2020;10:7733. doi: 10.1038/s41598-020-64588-y. [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Cell Reports Medicine are provided here courtesy of Elsevier

RESOURCES