1. Introduction
Over the last decade, we have seen increasing development and adoption of new digital technologies in healthcare and medicine. Among these technologies, artificial intelligence (AI) has gripped the public attention in ways that few others have. In the years ahead, AI and Generative AI (GenAI) is poised to broadly reshape healthcare. At the Health Systems journal, we continue to see an increasing volume of manuscript submissions. As the journal positions itself to attract and publish top quality novel papers in AI/GenAI for health, we, the editors felt that it would be helpful to discuss new frontiers and pathways for healthcare systems research that uses AI/GenAI in order to guide those manuscripts to be ready for publication.
We broadly define Artificial Intelligence, or AI, as technology that enables computers and machines to simulate human intelligence and problem-solving capabilities (Stryker & Kavlakoglu, 2024). We can further classify AI into two broad categories. Predictive AI uses data and algorithms to predict a variable of interest (e.g., diagnosis, prognosis, hospital readmission, etc.) (Khare, 2023). Generative AI refers to a class of artificial intelligence systems designed to create new content, such as text, images, music, or videos, by learning patterns from large datasets (Eysenbach, 2023). In this article, we use the broad term AI to include both types. A large part of the recent success of AI in medicine and health has been due to advances in machine learning – computer programs that learn without being explicitly programmed – and deep learning systems that are based on many-layered neural networks that have shown remarkable accuracy in their tasks (Beede et al., 2020; Hinton et al., 2006). GenAI-driven chatbots and virtual health assistants can answer patient queries, provide health education, and offer reminders for medication or appointments. These AI tools can generate responses tailored to individual patient needs, making communication more relevant and accessible, and are being deployed in many medical settings, with internet connections and access to ChatGPT (Eysenbach, 2023). Simultaneously, many medical institutions are evaluating AI to assist with tasks that humans find tedious or time-consuming. Current research underscores the productivity enhancements that such GenAI tools are bringing to medicine, potentially improving accuracy of medical tasks, and opening up new opportunities (Rajpurkar et al., 2022; Wachter & Brynjolfsson, 2024).
In this editorial, we begin by taking a quick look at the recent past of AI and its use in health. We then present the current landscape of AI research in health. We further discuss promising avenues for novel innovations in health systems research. We present the challenges posed by AI models when they are implemented from the lab to clinics, practice and the point-of-care. All throughout these discussions, we will point out opportunities for different types of research that our authors can be working on and submit for consideration at Health Systems journal.
2. A look at the history of AI in medicine
The history of AI in medicine dates back to the 1960s. During this time, a small group of researchers conducted foundational research which led to the creation of DENDRAL, an expert system designed for chemical analysis, that laid the groundwork for future medical applications (Lederberg, 1990). Simultaneously, the MYCIN system was developed at Stanford University (Shortliffe, 2012). MYCIN was an early AI program designed to assist physicians by recommending treatments for certain bacterial infections, showcasing the potential of AI in clinical decision support. These systems were built on the foundations of Turing test in which Alan Turing posed the questions: “Can machines think?” The 1970s and 1980s saw a surge in the development of expert systems, which aimed to mimic the decision-making abilities of human experts. Systems like INTERNIST-I (Miller et al., 1985) were created to assist in diagnosing complex medical cases. These systems showed the potential of rule-based programs to handle large databases of medical knowledge and make informed decisions. However, they also revealed significant limitations, such as the difficulty in generalizing knowledge and non-flexibility of rule-based systems.
Bayes’ Theorem, a fundamental concept in probability theory and Florence Nightingale’s pioneering work in statistical visualization, was adopted during the 1970s to modern machine learning. Bayes’ Theorem provides a mathematical framework for updating the probability of a hypothesis based on new evidence. It underpins many machine learning algorithms, especially in fields like classification, spam filtering, and medical diagnosis, through Bayesian inference. Florence Nightingale, known for her groundbreaking use of data visualization to communicate the impact of sanitation on mortality rates, introduced the importance of using statistics to inform decisions. Her innovative visualizations inspired new methods for presenting complex data, which is now a key element in modern machine learning for understanding and visualizing patterns.
The 2000s saw an explosion of electronic health records (EHRs) and the availability of large datasets, which led to the rise of data mining using machine learning and predictive AI (Rajpurkar et al., 2022). AI systems could now learn from vast amounts of data. Several machine learning methods such as support vector machines, neural networks, random-forest classification, and natural language processing started to be applied in various medical domains, including genomics, proteomics, and drug discovery. IBM’s Watson, which gained fame by winning Jeopardy, was adapted for healthcare to assist in diagnosing and recommending treatments for cancer patients (Jie et al., 2021). However, certain limitations were observed with Watson.
The 2010s witnessed the rise of deep learning, which revolutionized many areas of AI, including medicine. Convolutional neural networks (CNNs) for image classification and recurrent neural networks (RNNs) for speech recognition tasks showed remarkable success. In healthcare, deep learning models achieved significant milestones in radiology (Ardila et al., 2019), dermatology (Courtiol et al., 2019), and ophthalmology (Liu et al., 2019), often outperforming human experts in tasks like detecting tumours in medical images (Hollon et al., 2020). The decade also saw the proliferation of big data analytics, analysing vast amounts of patient data to identify patterns and trends, leading to more accurate and timely decisions.
In the early 2020s, the COVID-19 pandemic accelerated the adoption of AI technologies for disease tracking, predictive modeling, and vaccine development (Lyon et al., 2021). AI since then has become more integrated into clinical practice and public health. AI-driven telemedicine platforms were quickly adopted to provide remote patient care. Natural language processing was integrated with EHRs, enhancing patient management and research capabilities. AI’s role in genomics continued to expand, contributing to advancements in precision medicine and understanding genetic diseases.
Today AI/GenAI continues to make strides as a critical tool in healthcare. The focus has shifted towards ethical concerns, ensuring AI systems are transparent, fair, and free from bias. Today AI researchers, clinicians, and policy-makers are developing robust frameworks for AI integration. AI’s journey in medicine from the 1960s to now reflects a progression from theoretical research to practical, widespread application, continually transforming the healthcare landscape. The recent past paves the way for a bright present and future of AI in medicine and health research.
3. Current landscape of medical AI research
The current landscape of AI in medicine and healthcare research is characterized by rapid advancements and widespread exploration of AI-enabled tools in various healthcare domains. We can broadly categorize this research into three key areas of focus: diagnostics and prognosis, precision medicine and operational efficiency. Figure 1 shows how AI in medicine research is growing across these three areas. Below we discuss them and provide examples of each.
Figure 1.

A snapshot of current landscape of AI in medicine research.
3.1. Diagnostics & prognosis
The first focus area has been diagnostics in which AI-powered tools are being extensively developed and validated for their ability to analyse medical images, pathology slides, and genetic data with high accuracy, often surpassing human capabilities. AI is extensively used in the analysis of medical images (e.g., X-rays, MRIs, CT scans) Ardila et al. (2019), and pathology slides (Jackson et al., 2020). Research shows that pathologists can diagnose breast cancer faster and more accurately with the help of AI software (Jackson et al., 2020). Deep learning models can detect abnormalities, tumours, and other conditions with higher accuracy and can outperform human radiologists in speed (Zhou et al., 2020). For example, AI tools are used to identify lung nodules in chest scans Ardila et al. (2019 or detect diabetic retinopathy in eye images (Wolf et al., 2020).
As EHR data proliferated in the last few years, we now have AI models that can analyse vast amount of patient data to predict disease onset, progression, and patient outcomes. Predictive analytics is helping to identify patients at high risk for conditions such as diabetes, heart disease, or sepsis, enabling early intervention and preventive care (Adler et al., 2020). Some hospitals are also using AI for cancer diagnosis and prognosis to predict the patient’s chances of recovery, potential complications, and overall response to treatment (Huang et al., 2020).
3.2. Personalized medicine
AI has accelerated the field of personalized medicine. New AI tools have the ability to analyse large datasets, identify patterns, and predict individual patient responses to treatments. This has led to more targeted therapies and better patient outcomes (Esteva et al., 2017).
With AI, it is now possible to tailor treatments to individual patients by analysing genetic information, electronic health records, and other data sources. This personalized approach enhances the effectiveness of treatments and minimizes side effects (Gainza et al., 2020). Oncology is benefitting tremendously from AI-driven precision medicine, impacting the selection of better therapies based on a patient’s genetic profile (Huang et al., 2020).
AI is proving to be a game changer in drug discovery and development by significantly accelerating the identification and validation of new drug candidates. Machine learning algorithms can predict which compounds are most likely to succeed by analysing vast datasets, including chemical properties, biological activities, and clinical trial outcomes (Gainza et al., 2020; Gussow et al., 2020). AI models can identify potential drug candidates by screening millions of molecules quickly, optimizing lead compounds, and predicting their efficacy and safety profiles. Furthermore, AI can easily analyse genomic and proteomic data to identify novel drug targets (Stokes et al., 2020; Zhavoronkov et al., 2019). AI has simplified clinical trials by assisting in patient recruitment, monitoring adverse effects, and predicting patient responses (Harrer et al., 2019). This has led to trial efficiency and higher success rates. Deep learning models for molecular analysis have been shown to accelerate the discovery of novel drugs by reducing the need for slower, more costly physical experiments (Stokes et al., 2020). AI can also simulate drug interactions with biological targets, speeding up the identification of promising candidates for repurposing. Overall, AI reduces the time and cost associated with bringing new drugs to market, ultimately facilitating the development of more effective and personalized treatments.
3.3. Operational efficiency
The third focus area of current research in AI for health is the gains made in operational efficiency. AI tools are being used to optimize hospital operations, from streamlining administrative tasks to predicting patient admissions and managing resources more effectively. Predictive AI is assisting in scheduling surgeries (Martinez et al., 2021), reducing wait times, and managing emergency department resources more effectively (Hijry & Olawoyin, 2021).
Hospital robots can either aid a human surgeon or in some cases execute operations by themselves. These systems assist surgeons in performing minimally invasive surgeries with greater accuracy and less variability. Studies show that robotic surgery reduces recovery times and provides better patient outcomes (Zemmar et al., 2020).
Operational efficiencies have been gained by using AI enhanced telemedicine systems (Shaik et al., 2023). They provide decision support tools and automate routine tasks. AI-driven remote monitoring technologies track patients’ vital signs and other health metrics, alerting healthcare providers to potential issues before they become critical. Such systems are now being used to monitor chronic disease patients such as heart failure and diabetes (Alnosayan et al., 2017).
Clinical decision support systems (CDSS) provide healthcare professionals with tools for more accurate and timely decision-making. AI-enabled CDSS are helping both with operational efficiencies as well as with personalized medicine. They can process EHRs to identify potential diagnoses, recommend personalized treatment plans, and flag early warning signs of deteriorating health conditions (Ferrão et al., 2021). Natural language processing (NLP) enables these systems to extract critical information from unstructured data, such as clinical notes and research articles, ensuring that clinicians have access to the latest medical knowledge (Sheikhalishahi et al., 2019). AI-integrated CDSS not only reduces the cognitive burden on healthcare providers but also enhances patient safety and outcomes by supporting more informed and evidence-based medical decisions.
4. Opportunities and promising avenues for novel AI healthcare systems research
Most AI-related manuscripts that we receive fall into two categories: either they are a systematic review paper or follow the traditional approach of supervised learning on labelled data to train an AI system (often using multiple ML models) and then evaluating the system by comparing against human experts. In this section, we try to provide novel ideas and innovative areas of AI in health research that the editorial team are interested in.
4.1. AI and Human-in-the-loop collaboration systems
Instead of a head-to-head comparison of AI with humans, we are now beginning to see study setups where many medical practices are embracing to involve human-in-the-loop setups, where clinicians and practitioners actively collaborate with AI systems and provide oversight (Rajpurkar et al., 2022). Within such setups, we are seeing two forms of collaboration: i) humans and AI cooperating with predictions and ii) humans using GenAI to enhance collaboration. The former approach makes disease predictions a joint effort, while the latter brings enormous knowledge of GenAI into the discussions.
4.2. AI and human setup
We are just beginning to see collaborative setups between AI and humans in health. Such setups enhance patient care and improve clinical outcomes. These collaborations are designed to combine the computational power and data-processing capabilities of AI with the critical thinking, empathy, and nuanced understanding of human clinicians (Rajpurkar et al., 2022). For example, a recent study has shown that clinical experts and AI in combination achieve better performance than experts alone (Patel et al., 2019). Another study found that AI-assisted clinical experts surpassed both humans and AI alone in detecting malignant nodules on chest X-rays (Sim et al., 2020).
4.3. Use of generative AI models in medicine
Generative AI (GenAI) is increasingly being adopted in practice and inside clinics as an assistive tool. GenAI refers to a class of artificial intelligence models designed to create new content that mimics human-like creativity. These models, which include techniques such as Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs), are trained on large datasets and can generate a wide range of outputs, from text and images to music and videos (Sai et al., 2024). Unlike traditional AI, which typically focuses on recognizing patterns and making predictions based on existing data, GenAI learns the underlying structure of the data to produce original content that is novel yet coherent and contextually relevant.
Introduction of ChatGPT by OpenAI brought the world’s attention to the power of generative AI and Large Language Models (LLMs). LLM models, such as OpenAI’s GPT-4 or Google’s GEMINI, excel at a wide range of language-related tasks, including text completion, translation, summarization, and answering questions. Their ability to understand context and generate coherent, contextually appropriate responses makes them valuable tools in applications like chatbots, virtual assistants, and content creation. GenAI models and tools are continually advancing every day, and we now can generate text (ChatGPT, Llama 3.1), audio (Murf AI Voice Generator, OpenAI Musenet), video (Google Imagen Video, OpenAI Sora), images (Dalle-2, Midjourney), 3D visual (Microsoft Rodin Diffusion), and code (DeepMind AlphaCode). Despite their impressive capabilities, GenAI and LLMs also pose challenges, such as potential biases in generated content and the need for substantial computational resources for training and deployment.
The scope and range of GenAI applications in medicine is vast. Figure 2 shows a snapshot of the applications of GenAI in medicine and health. An essential element in disease diagnosis using medical imaging is the quality of scans such as X-rays or CT scan. When images are distorted due to noise, missing data, or low resolution, it can lead to misinterpretations. GenAI can be used to remove noise from images, and this tends to improve the accuracy of quantitative analysis. An enhanced generative adversarial network for the super-resolution reconstruction of images led to improved diagnosis results (Bing et al., 2019). GenAI methods are now being used to create synthetic datasets that resemble patient data, increasing the models’ robustness. GenAI models are also capable of generating data on rare conditions such as Aquagenic Urticaria (Rothbaum & McGee, 2016). It is observed that GenAI is playing a pivotal role in addressing the limited datasets available to build AI predictive models (Gootjes-Dreesbach et al., 2020).
Figure 2.

Applications of GenAI on medicine and health (adapted from Sai et al., 2024).
Generative AI is revolutionizing medical education and simulation by creating highly realistic, interactive, and personalized learning experiences for students and professionals. These AI-driven tools can generate 3D representations of organs, bones, or blood vessels, allowing learners to practice diagnosing and treating a wide range of conditions in a safe, controlled environment. For instance, AI can simulate patient interactions, complete with realistic symptoms and responses, enabling medical trainees to develop and refine their clinical skills. Medical residents are practicing surgical skills using robotics on computer generated 3D replicas of various human anatomy. Additionally, generative AI is being used to produce tailored educational content, such as personalized quizzes and adaptive learning modules, to meet individual learning needs (Eysenbach, 2023; Sai et al., 2024).
Drug development by pharma brings new drugs and therapies to market. However, it is a costly and time-consuming process with often low success rates. Compounds and molecules play a crucial role in drug discovery, serving as the fundamental building blocks for developing new therapeutic agents. The process of drug discovery involves identifying and optimizing these chemical entities to interact with specific biological targets, such as proteins or enzymes, to produce a desired therapeutic effect (Senior et al., 2020). GenAI can generate novel compounds, optimize drug candidates, and predict properties of drugs as well as identify candidates for possible drug repurposing. They are particularly good with 3D models of molecules (Li et al., 2021).
Clinical trials and Randomized Control Trials have been the gold standard to evaluate new interventions and translate research findings into clinical practice. However, the complexity and challenges of recruiting a diverse patient pool while ensuring safety make clinical trials tedious. AI integrated inside EHRs can now analyze vast amounts of patient data to quickly identify eligible candidates who meet specific trial criteria. GenAI in particular can also personalize recruitment strategies by generating tailored communication and engagement plans, thereby improving patient retention rates throughout the trial (Mueller et al., 2023). GenAI is optimizing trial protocols by simulating various trial designs and predicting their outcomes. AI can suggest modifications to protocols that maximize statistical power and minimize costs and duration, ensuring a more efficient trial process (Gootjes-Dreesbach et al., 2020).
New and exciting research is emerging on using GenAI to address personalized treatments (Wang et al., 2019), text summarization of medical records (Li et al., 2020), public epidemiology, and prosthetic organ design. Recent research suggests that AI holds enormous potential in solid organ transplants, particularly in areas such as organ allocation, donor-recipient pairing and precision transplant pathology (Peloso et al., 2022).
4.4. Moving beyond images in medical AI research
While the initial application of AI in medicine has been on imaging tasks such as image classification, now we are seeing interesting new research in which deep learning models can learn from many other kinds of input data. Recent work has used rich data sources such a genomic data, molecular information, natural language, medical signals such as electrocardiogram (ECG) and multimodal data (Rajpurkar et al., 2022). AlphaFold represented a breakthrough in the key task of protein folding (Senior et al., 2020), while Alley et al. made strides in the area of protein analysis, creating statistical summaries that help neural networks learn with less data (Alley et al., 2019).
AI and CRISPR gene editing are enabling scientists with more precise, efficient, and innovative approaches to understanding and treating genetic diseases. CRISPR, a groundbreaking gene-editing technology, allows scientists to make precise modifications to DNA, potentially correcting genetic mutations that might have caused the disease. However, designing effective CRISPR interventions requires intricate knowledge of gene interactions and potential off-target effects. Here, AI plays a crucial role by analyzing large genomic datasets to predict the outcomes of CRISPR edits, identifying optimal target sites, and minimizing unintended consequences (Wang et al., 2019). Machine learning algorithms can process large genomic sequencing data to identify patterns and predict which genes are likely to be implicated in particular diseases.
4.5. Shift towards new problem formulation with unlabeled data
Our editorial team is particularly interested in research with unsupervised learning, which involves learning from data without any labels. Labelling in healthcare has proven to be costly and time-consuming and obtaining such accurate datasets can be difficult. Papers that find novel patterns and categories such as those using clustering algorithms (Haraty et al., 2015) that group similar data points together have been applied to conditions such as sepsis and cancer. These categories can reveal novel patterns in disease manifestation which leads to better diagnosis and treatment.
Autoencoders and other unsupervised learning models can be trained to detect anomalies in medical images, such as MRIs, CT scans, and X-rays. These models learn to represent normal anatomy and can subsequently highlight areas that deviate from the norm, assisting radiologists in identifying potential abnormalities like tumours, lesions, or fractures. Unsupervised learning methods are employed to analyse EHRs to identify patterns and correlations that might not be immediately apparent (Haraty et al., 2015). Unsupervised learning is applied to wearable device data and health app data to identify patterns in physical activity, sleep, and other health behaviors. Clustering and pattern recognition techniques can provide insights into different lifestyle patterns and their impacts on health, enabling personalized recommendations for improving overall well-being.
5. Research addressing the next big challenges
There is no doubt that AI will continue to play an important role in medicine and health over the next few decades. A lot of promising novel discoveries in diagnosis, prognosis, and treatments will be fuelled by the use of AI models and tools. However, there are few major challenges and barriers that AI and GenAI faces. Below, we highlight a few such challenges and seek papers that can address them in novel ways specially from a systems perspective (Brailsford et al., 2012).
5.1. Bias, hallucination, and fabrication with GenAI
In the context of large language models (LLMs), “hallucination” refers to the phenomenon where the model generates information that seems plausible but is not grounded in the input data or reality. These hallucinations can manifest as fabrications, inaccuracies, or statements that are not supported by any factual basis. For instance, an LLM might generate a detailed but entirely fictional account of a medical case or provide incorrect answers to factual questions (Walters & Wilder, 2023). There have been several recent papers in which 55% of GPT-3.5 and 18% of GPT-4 citations were fabricated (Walters & Wilder, 2023).
Bias in training data is a universal problem in AI and can be reflected by LLMs. LLMs trained on historical medical literature may inherit gender biases present in the data. For example, they might downplay or misdiagnose conditions that predominantly affect women, such as endometriosis or fibromyalgia, because historical medical texts often dismissed or underrepresented these conditions (Schaul et al., 2023). This can lead to the model providing less accurate or less empathetic advice to female patients. Two growing concerns with LLMS emerge in healthcare applications: 1) to what extent do LLMs exhibit social bias based on patients’ protected attributes (like race) and 2) how do design choices (like architecture design and prompting strategies) influence the observed biases? Significant racial bias has been found in dermatology applications of LLMS. Datasets used by researchers may contain cultural bias, age bias, and bias in mental health stigma that can affect the results. Addressing these biases requires careful attention to the training data used, paying careful scrutiny to prompt design on bias patterns, and incorporating diverse perspectives and sources to create more equitable and accurate LLMs in health and medicine (Omiye et al., 2024).
5.2. Translational AI – implementations from lab to clinic
As with drugs and devices, we need to move from “basic science” to “clinical science”. Unfortunately, today many researched AI models are not implemented, and the few implemented models are not researched for effectiveness. We know that the best evidence arises from clinical interventions which is often represented by randomized controlled trials (RCTs) or systematic reviews of RCTs. Conducting RCT studies of AI or GenAI is not trivial, actually hard. There are very few RCTs with AI use that demonstrate value from real-world use (Plana et al., 2022). A reason for that might be as the recent AMA poll found only 38% of physicians are using AI tools (Plana et al., 2022).
5.3. AI model performance drifts
It has been noted that often AI models performance degrade over time. This phenomenon, called performance drift, occurs when models trained on certain datasets start to produce less accurate or reliable results as the underlying data distribution changes. Several factors can cause such performance drifts. The characteristics of the input data may change over time due to new trends, behaviors, or external factors that were not present in the original training data. Data can exhibit seasonal patterns that the model may not have been trained to handle. For instance, certain medical conditions might be more common in specific seasons, and a model trained on data from other times might not perform well during those periods. There are two main types of data drift. “Covariate drift” occurs when the distribution of input features (independent variables) changes over time, while the relationship between inputs and outputs remains constant. The other type “Concept drift” happens when the underlying relationship between inputs and outputs changes, such as target variable behaves differently due to external factors. Other causes of drift can be due to changes in clinical practices and guidelines in which model trained on outdated practices might not perform well when new treatments or diagnostic criteria are adopted (Bayram et al., 2022). We need to build evidence of efficacy by implementing more AI models in the clinic and check for model drifts.
5.4. AI and Analytics in claims fraud in value-based care models
As we move from a fee-for-service to value-based care models, detecting and preventing claim fraud in healthcare is becoming important. AI algorithms can help to identify behavioural patterns from historical claims data. AI can detect unusually high billing amounts, repeated procedures, or patterns in patient visits that could suggest fraud (Agarwal, 2023). Risk scores assigned to claims by AI can be based on various factors, such as the provider’s history, claim amount, treatment type, and patient characteristics. Claims with high-risk scores are flagged for further investigation by human analysts. AI can quickly map relationships between different healthcare entities (e.g., doctors, patients, clinics) to detect unusual connections that might indicate organized fraud schemes (e.g., collusion between a doctor and a patient).
5.5. AI for optimizing operations in hospitals
Column generation is an effective method primarily used in solving large-scale linear programming problems, particularly in areas like operations research (e.g., vehicle routing, cutting stock problems). It involves iteratively optimizing a subproblem containing a subset of columns (variables) and obtains a set of dual values to generate new columns with negative reduced costs. ML can guide the column generation process by predicting which columns (or variables) are more likely to improve the dual-value objective function. Instead of relying on purely mathematical rules to generate new columns, ML models can predict which columns are most promising based on previous optimization outcomes and problem characteristics, speeding up convergence. Machine learning models can learn from past problem instances and solutions, enabling the optimization process to anticipate future needs.
AI is also transforming supply chain management systems in healthcare. New demand forecasting systems utilize AI to analyse historical sales data, market trends, and external factors to predict future demand (Kumar et al., 2023). Inventory management integrated with AI is being used for predicting when stock needs to be replenished and optimizing reorder points. This minimizes excess inventory in hospitals while ensuring sufficient stock to meet demand. Machine learning models can adjust these predictions based on real-time sales data and changing conditions. AI is also used to optimize transportation routes and hospital-related logistics.
5.6. Accountability, regulation, and privacy with data use
The regulatory and governance challenges with AI in healthcare and medicine revolve around a myriad of factors such as ensuring safety, efficacy, fairness, transparency, and accountability (Reddy et al., 2020). We created a framework that can help to understand the various core governance and regulatory issues as shown below in Figure 3. A full detailed discussion of each of these is beyond the scope of this editorial, but we elaborate on a few below.
Figure 3.

A regulatory and governance core issues framework with ai in health.
As AI models begin to be deployed, regulatory bodies must develop and ensure standards for validation and testing, fairness and bias, and protect patient privacy. Of critical concern is determining who is accountable for AI-driven decisions, especially in situations where AI recommendations lead to adverse outcomes. As AI systems take on more responsibility in clinical settings, a concern facing the system is that clinicians can become over reliant on AI, perhaps seeing a gradual decline in their own skills (Rajpurkar et al., 2022). This in turn gives AI companies and developers more power. But at the heart of all this remains patient protection and privacy, ensuring that AI systems comply with data privacy regulations, such as GDPR in Europe or HIPAA in the United States, to protect patient confidentiality (Forcier et al., 2019). Current research needs to ensure that patients are adequately informed about the use of AI in their care and have given consent. We must address ethical concerns related to any potential misuse of AI such as surveillance or discrimination. Clinical facilities must boost their security measures to prevent data breaches and unauthorized access to sensitive health information.
The proliferation of AI also raises concerns around accountability, as it is currently unclear whether developers, regulators, sellers, or healthcare providers should be held accountable if a model makes mistakes even after being thoroughly clinically validated. Currently, doctors are held liable when they deviate from the standard of care and patient injury occurs (Rajpurkar et al., 2022). It becomes critically important to establish clear legal and ethical guidelines to govern the use of AI in health, ensuring that healthcare providers, developers, and other stakeholders understand their responsibilities and liabilities. This should typically involve a diverse range of stakeholders, including clinicians, patients, ethicists, technologists, and policymakers, in the governance of AI in health to ensure comprehensive and balanced oversight. Within the context of accountability lies the issue of AI explainability. It is important to develop methods to make AI decisions interpretable and understandable to doctors and patients, enabling informed decision-making and trust.
6. Conclusions
Artificial Intelligence (AI) has emerged as a transformative force in healthcare, driving significant advancements across various domains, including diagnostics, personalized medicine, operational efficiency, and patient care. In this editorial, we have provided prospective authors with topics, insights, and potential novel research that we seek for the journal. AI-powered tools enhance the accuracy and speed of medical imaging, predict patient outcomes through sophisticated analytics, and enable personalized treatment plans by analyzing vast datasets. These capabilities not only improve patient outcomes but also streamline healthcare operations, making care delivery more efficient and cost-effective. Despite these advancements, AI in healthcare faces several challenges, including addressing biases in algorithms, ensuring data privacy and security, and achieving transparency and explainability in AI-driven decisions. Regulatory and ethical considerations are paramount to ensure that AI technologies are safe, fair, and equitable.
In conclusion, AI holds immense potential to revolutionize healthcare, offering solutions that enhance patient outcomes, operational efficiency, and overall quality of care. Continued research and innovation, guided by ethical considerations and regulatory oversight, will be crucial in fully realizing the benefits of AI in healthcare and addressing the challenges that lie ahead.
Acknowledgments
The authors are grateful for very insightful feedback received on an earlier version of this editorial from a few highly accomplished area editors of the Health Systems board. We would like to thank and acknowledge Christina Bartenschlager, Arin Brahma, and Rashida F. Parks for their help.
References
- Adler, E. D., Voors, A. A., Klein, L., Macheret, F., Braun, O. O., Urey, M. A., Zhu, W., Sama, I., Tadel, M., Campagnari, C., Greenberg, B., & Yagil, A. (2020). Improving risk prediction in heart failure using machine learning. European Journal of Heart Failure, 22(1), 139–147. 10.1002/ejhf.1628 [DOI] [PubMed] [Google Scholar]
- Agarwal, S. (2023). An intelligent machine learning approach for fraud detection in medical claim insurance: A comprehensive study. Scholars Journal of Engineering and Technology, 11(9), 191–200. 10.36347/sjet.2023.v11i09.003 [DOI] [Google Scholar]
- Alley, E. C., Khimulya, G., Biswas, S., AlQuraishi, M., & Church, G. M. (2019). Unified rational protein engineering with sequence-based deep representation learning. Nature Methods, 16(12), 1315–1322. 10.1038/s41592-019-0598-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Alnosayan, N., Chatterjee, S., Alluhaidan, A., Lee, E., & Feenstra, L. H. (2017). Design and usability of a heart failure mHealth system: A pilot study. JMIR Human Factors, 4(1), e6481. 10.2196/humanfactors.6481 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ardila, D., Kiraly, A. P., Bharadwaj, S., Choi, B., Reicher, J. J., Peng, L., Tse, D., Etemadi, M., Ye, W., Corrado, G., Naidich, D. P., & Shetty, S. (2019). End-to-end lung cancer screening with three-dimensional deep learning on low-dose chest computed tomography. Nature Medicine, 25(6), 954–961. 10.1038/s41591-019-0447-x [DOI] [PubMed] [Google Scholar]
- Bayram, F., Ahmed, B. S., & Kassler, A. (2022). From concept drift to model degradation: An overview on performance-aware drift detectors. Knowledge-Based Systems, 245, 108632. 10.1016/j.knosys.2022.108632 [DOI] [Google Scholar]
- Beede, E., Baylor, E., Hersch, F., Iurchenko, A., Wilcox, L., Ruamviboonsuk, P., & Vardoulakis, L. M. (2020, April). A human-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy. Proceedings of the 2020 CHI conference on human factors in computing systems (pp. 1–12). [Google Scholar]
- Bing, X., Zhang, W., Zheng, L., & Zhang, Y. (2019). ‘Medical image super resolution using improved generative adversarial networks. Institute of Electrical and Electronics Engineers Access, 7, 145030–145038. 10.1109/ACCESS.2019.2944862 [DOI] [Google Scholar]
- Brailsford, S., Harper, P., LeRouge, C., & Payton, F. C. (2012). Editorial. Health Systems, 1(1), 1–6. 10.1057/hs.2012.9 [DOI] [Google Scholar]
- Courtiol, P., Maussion, C., Moarii, M., Pronier, E., Pilcer, S., Sefta, M., Manceron, P., Toldo, S., Zaslavskiy, M., Le Stang, N., Girard, N., Elemento, O., Nicholson, A. G., Blay, J.-Y., Galateau-Sallé, F., Wainrib, G., & Clozel, T. (2019). Deep learning-based classification of mesothelioma improves prediction of patient outcome. Nature Medicine, 25(10), 1519–1525. 10.1038/s41591-019-0583-3 [DOI] [PubMed] [Google Scholar]
- Esteva, A., Kuprel, B., Novoa, R. A., Ko, J., Swetter, S. M., Blau, H. M., & Thrun, S. (2017). Dermatologist-level classification of skin cancer with deep neural networks. Nature, 542(7639), 115–118. 10.1038/nature21056 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Eysenbach, G. (2023). The role of ChatGPT, generative language models, and artificial intelligence in medical education: A conversation with ChatGPT and a call for papers. JMIR Medical Education, 9(1), e46885. 10.2196/46885 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ferrão, J. C., Oliveira, M. D., Janela, F., Martins, H. M., & Gartner, D. (2021). Can structured EHR data support clinical coding? A data mining approach. Health Systems, 10(2), 138–161. 10.1080/20476965.2020.1729666 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Forcier, M. B., Gallois, H., Mullan, S., & Joly, Y. (2019). Integrating artificial intelligence into health care through data access: Can the GDPR act as a beacon for policymakers? Journal of Law & the Biosciences, 6(1), 317–335. 10.1093/jlb/lsz013 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gainza, P., Sverrisson, F., Monti, F., Rodolà, E., Boscaini, D., Bronstein, M. M., & Correia, B. E. (2020). Deciphering interaction fingerprints from protein molecular surfaces using geometric deep learning. Nature Methods, 17(2), 184–192. 10.1038/s41592-019-0666-6 [DOI] [PubMed] [Google Scholar]
- Gootjes-Dreesbach, L., Sood, M., Sahay, A., Hofmann-Apitius, M., & Fröhlich, H. (2020, May). ‘Variational autoencoder modular bayesian networks for simulation of heterogeneous clinical study data. Frontiers in Big Data, 3, 16. 10.3389/fdata.2020.00016 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gussow, A. B., Park, A. E., Borges, A. L., Shmakov, S. A., Makarova, K. S., Wolf, Y. I., Bondy-Denomy, J., & Koonin, E. V. (2020). Machine-learning approach expands the repertoire of anti-crispr protein families. Nature Communications, 11(1), 3784. 10.1038/s41467-020-17652-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Haraty, R. A., Dimishkieh, M., & Masud, M. (2015). An enhanced k-means clustering algorithm for pattern discovery in healthcare data. International Journal of Distributed Sensor Networks, 11(6), 615740. 10.1155/2015/615740 [DOI] [Google Scholar]
- Harrer, S., Shah, P., Antony, B., & Hu, J. (2019). Artificial intelligence for clinical trial design. Trends in Pharmacological Sciences, 40(8), 577–591. 10.1016/j.tips.2019.05.005 [DOI] [PubMed] [Google Scholar]
- Hijry, H., & Olawoyin, R. (2021). Predicting patient waiting time in the queue system using deep learning algorithms in the emergency room. International Journal of Industrial Engineering and Operations Management, 3(1), 33–45. 10.46254/j.ieom.20210103 [DOI] [Google Scholar]
- Hinton, G. E., Osindero, S., & Teh, Y. W. (2006). A fast learning algorithm for deep belief nets. Neural Computation, 18(7), 1527–1554. 10.1162/neco.2006.18.7.1527 [DOI] [PubMed] [Google Scholar]
- Hollon, T. C., Pandian, B., Adapa, A. R., Urias, E., Save, A. V., Khalsa, S. S. S., Eichberg, D. G., D’Amico, R. S., Farooq, Z. U., Lewis, S., Petridis, P. D., Marie, T., Shah, A. H., Garton, H. J. L., Maher, C. O., Heth, J. A., McKean, E. L., Sullivan, S. E., Hervey-Jumper, S. L., Patil, P. G., & Camelo-Piragua, S. (2020). Near real-time intraoperative brain tumor diagnosis using stimulated Raman histology and deep neural networks. Nature Medicine, 26(1), 52–58. 10.1038/s41591-019-0715-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Huang, S., Yang, J., Fong, S., & Zhao, Q. (2020). Artificial intelligence in cancer diagnosis and prognosis: Opportunities and challenges. Cancer Letters, 471, 61–71. 10.1016/j.canlet.2019.12.007 [DOI] [PubMed] [Google Scholar]
- Jackson, H. W., Fischer, J. R., Zanotelli, V. R. T., Ali, H. R., Mechera, R., Soysal, S. D., Moch, H., Muenst, S., Varga, Z., Weber, W. P., & Bodenmiller, B. (2020). The single-cell pathology landscape of breast cancer. Nature, 578(7796), 615–620. 10.1038/s41586-019-1876-x [DOI] [PubMed] [Google Scholar]
- Jie, Z., Zhiying, Z., & Li, L. (2021, March). ‘A meta-analysis of Watson for oncology in clinical application. Scientific Reports, 11(1), 5792. 10.1038/s41598-021-84973-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Khare, Y. (2023). Generative AI vs predictive AI: What is the difference? Analytics Vidhya. Retrieved December 12, 2023, from https://www.analyticsvidhya.com/blog/2023/09/generative-ai-vs-predictive-ai/
- Kumar, A., Mani, V., Jain, V., Gupta, H., & Venkatesh, V. G. (2023). Managing healthcare supply chain through artificial intelligence (AI): A study of critical success factors. Computers & Industrial Engineering, 175, 108815. 10.1016/j.cie.2022.108815 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lederberg, J. (1990). How DENDRAL was conceived and born. In Bruce I. Blum & Karen Duncan (Eds.), A history of medical informatics (pp. 14–44). New York: Association for Computing Machinery. [Google Scholar]
- Li, Y., Nair, P., Lu, X. H., Wen, Z., Wang, Y., Dehaghi, A. A. K., Miao, Y., Liu, W., Ordog, T., Biernacka, J. M., Ryu, E., Olson, J. E., Frye, M. A., Liu, A., Guo, L., Marelli, A., Ahuja, Y., Davila-Velderrain, J., & Kellis, M. (2020). Inferring multimodal latent topics from electronic health records. Nature Communications, 11(1), 2536. 10.1038/s41467-020-16378-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li, Y., Pei, J., & Lai, L. (2021). ‘Structure-based de novo drug design using 3D deep generative models. Chemical Science, 12(41), 13664–13675. 10.1039/D1SC04444C [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liu, H., Li, L., Wormstone, I. M., Qiao, C., Zhang, C., Liu, P., Li, S., Wang, H., Mou, D., Pang, R., Yang, D., Zangwill, L. M., Moghimi, S., Hou, H., Bowd, C., Jiang, L., Chen, Y., Hu, M., Xu, Y., Kang, H., & Xu, M. (2019). Development and validation of a deep learning system to detect glaucomatous optic neuropathy using fundus photographs. JAMA Ophthalmology, 137(12), 1353–1360. 10.1001/jamaophthalmol.2019.3501 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lyon, V., LeRouge, C., Fruhling, A., & Thompson, M. (2021). Home testing for COVID-19 and other virus outbreaks: The complex system of translating to communities. Health Systems, 10(4), 298–317. 10.1080/20476965.2021.1952905 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Martinez, O., Martinez, C., Parra, C. A., Rugeles, S., & Suarez, D. R. (2021). Machine learning for surgical time prediction. Computer Methods and Programs in Biomedicine, 208, 106220. 10.1016/j.cmpb.2021.106220 [DOI] [PubMed] [Google Scholar]
- Miller, R. A., Pople, H. E., Jr., & Myers, J. D. (1985). Internist-I, an experimental computer-based diagnostic consultant for general internal medicine. In James A. Reggia, Stanley Tuhrim (Eds.), Computer-assisted medical decision making (pp. 139–158). Springer New York. [DOI] [PubMed] [Google Scholar]
- Mueller, J., Parikh, R. B., & Noble, A. (2023). Evaluating clinical trial inclusion/exclusion criteria from claims using generative artificial intelligence. Technical Report.
- Omiye, J. A., Gui, H., Rezaei, S. J., Zou, J., & Daneshjou, R. (2024). Large language models in medicine: The potentials and pitfalls: A narrative review. Annals of Internal Medicine, 177(2), 210–220. 10.7326/M23-2772 [DOI] [PubMed] [Google Scholar]
- Patel, B. N., Rosenberg, L., Willcox, G., Baltaxe, D., Lyons, M., Irvin, J., Rajpurkar, P., Amrhein, T., Gupta, R., Halabi, S., Langlotz, C., Lo, E., Mammarappallil, J., Mariano, A. J., Riley, G., Seekins, J., Shen, L., Zucker, E., & Lungren, M. P. (2019). Human–machine partnership with artificial intelligence for chest radiograph diagnosis. NPJ Digital Medicine, 2(1), 111. 10.1038/s41746-019-0189-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Peloso, A., Moeckli, B., Delaune, V., Oldani, G., Andres, A., & Compagnon, P. (2022). Artificial intelligence: Present and future potential for solid organ transplantation. Transplant International, 35, 10640. 10.3389/ti.2022.10640 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Plana, D., Shung, D. L., Grimshaw, A. A., Saraf, A., Sung, J. J., & Kann, B. H. (2022). Randomized clinical trials of machine learning interventions in health care: A systematic review. JAMA Network Open, 5(9), e2233946–e2233946. 10.1001/jamanetworkopen.2022.33946 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rajpurkar, P., Chen, E., Banerjee, O., & Topol, E. J. (2022). AI in health and medicine. Nature Medicine, 28(1), 31–38. 10.1038/s41591-021-01614-0 [DOI] [PubMed] [Google Scholar]
- Reddy, S., Allan, S., Coghlan, S., & Cooper, P. (2020). A governance model for the application of AI in health care. Journal of the American Medical Informatics Association, 27(3), 491–497. 10.1093/jamia/ocz192 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rothbaum, R., & McGee, J. (2016, November). ‘Aquagenic urticaria: Diagnostic and management challenges. Journal of Asthma and Allergy, 9, 209–213. 10.2147/JAA.S91505 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sai, S., Gaur, A., Sai, R., Chamola, V., Guizani, M., & Rodrigues, J. J. (2024). Generative AI for transformative healthcare: A comprehensive study of emerging models, applications, case studies, and limitations. Institute of Electrical and Electronics Engineers Access, 12, 31078–31106. 10.1109/ACCESS.2024.3367715 [DOI] [Google Scholar]
- Schaul, K., Chen, S. Y., & Tiku, N. (2023). Inside the secret list of websites that make AI like ChatGPT sound smart [WWW document]. Retrieved December 20, 2023, from https://www.washingtonpost.com/technology/interactive/2023/ai-chatbot-learning/.
- Senior, A. W., Evans, R., Jumper, J., Kirkpatrick, J., Sifre, L., Green, T., Qin, C., Žídek, A., Nelson, A. W. R., Bridgland, A., Penedones, H., Petersen, S., Simonyan, K., Crossan, S., Kohli, P., Jones, D. T., Silver, D., Kavukcuoglu, K., & Hassabis, D. (2020). Improved protein structure prediction using potentials from deep learning. Nature, 577(7792), 706–710. 10.1038/s41586-019-1923-7 [DOI] [PubMed] [Google Scholar]
- Shaik, T., Tao, X., Higgins, N., Li, L., Gururajan, R., Zhou, X., & Acharya, U. R. (2023). Remote patient monitoring using artificial intelligence: Current state, applications, and challenges. Wiley Interdisciplinary Reviews Data Mining and Knowledge Discovery, 13(2), e1485. 10.1002/widm.1485 [DOI] [Google Scholar]
- Sheikhalishahi, S., Miotto, R., Dudley, J. T., Lavelli, A., Rinaldi, F., & Osmani, V. (2019). Natural language processing of clinical notes on chronic diseases: Systematic review. JMIR Medical Informatics, 7(2), e12239. 10.2196/12239 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shortliffe, E. (Ed.). (2012). Computer-based medical consultations: MYCIN (Vol. 2). Elsevier. [Google Scholar]
- Sim, Y., Chung, M. J., Kotter, E., Yune, S., Kim, M., Do, S., Han, K., Kim, H., Yang, S., Lee, D.-J., & Choi, B. W. (2020). Deep convolutional neural network–based software improves radiologist detection of malignant lung nodules on chest radiographs. Radiology, 294(1), 199–209. 10.1148/radiol.2019182465 [DOI] [PubMed] [Google Scholar]
- Stokes, J. M., Yang, K., Swanson, K., Jin, W., Cubillos-Ruiz, A., Donghia, N. M., MacNair, C. R., French, S., Carfrae, L. A., Bloom-Ackermann, Z., Tran, V. M., Chiappino-Pepe, A., Badran, A. H., Andrews, I. W., Chory, E. J., Church, G. M., Brown, E. D., Jaakkola, T. S., Barzilay, R., & Collins, J. J. (2020). A deep learning approach to antibiotic discovery. Cell, 181(2), 475–483. 10.1016/j.cell.2020.04.001 [DOI] [PubMed] [Google Scholar]
- Stryker, C., & Kavlakoglu, E.. 2024. What is artificial intelligence (AI)?. Accessed 15 September 2024. https://www.ibm.com/topics/artificial-intelligence
- Wachter, R. M., & Brynjolfsson, E. (2024). Will generative artificial intelligence deliver on its promise in health care? JAMA, 331(1), 65–69. 10.1001/jama.2023.25054 [DOI] [PubMed] [Google Scholar]
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13(1), 14045. 10.1038/s41598-023-41032-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang, D., Zhang, C., Wang, B., Li, B., Wang, Q., Liu, D., Wang, H., Zhou, Y., Shi, L., Lan, F., & Wang, Y. (2019). Optimized CRISPR guide RNA design for two high-fidelity Cas9 variants by deep learning. Nature Communications, 10(1), 4284. 10.1038/s41467-019-12281-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wolf, R. M., Channa, R., Abramoff, M. D., & Lehmann, H. P. (2020). Cost-effectiveness of autonomous point-of-care diabetic retinopathy screening for pediatric patients with diabetes. JAMA Ophthalmology, 138(10), 1063–1069. 10.1001/jamaophthalmol.2020.3190 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zemmar, A., Lozano, A. M., & Nelson, B. J. (2020). The rise of robots in surgical environments during COVID-19. Nature Machine Intelligence, 2(10), 566–572. 10.1038/s42256-020-00238-2 [DOI] [Google Scholar]
- Zhavoronkov, A., Ivanenkov, Y. A., Aliper, A., Veselov, M. S., Aladinskiy, V. A., Aladinskaya, A. V., Terentiev, V. A., Polykovskiy, D. A., Kuznetsov, M. D., Asadulaev, A., Volkov, Y., Zholus, A., Shayakhmetov, R. R., Zhebrak, A., Minaeva, L. I., Zagribelnyy, B. A., Lee, L. H., Soll, R., Madge, D., Xing, L., & Guo, T. (2019). Deep learning enables rapid identification of potent DDR1 kinase inhibitors. Nature Biotechnology, 37(9), 1038–1040. 10.1038/s41587-019-0224-x [DOI] [PubMed] [Google Scholar]
- Zhou, D., Tian, F., Tian, X., Sun, L., Huang, X., Zhao, F., Zhou, N., Chen, Z., Zhang, Q., Yang, M., Yang, Y., Guo, X., Li, Z., Liu, J., Wang, J., Wang, J., Wang, B., Zhang, G., Sun, B., Zhang, W., & Chen, K. (2020). Diagnostic evaluation of a deep learning model for optical diagnosis of colorectal cancer. Nature Communications, 11(1), 2961. 10.1038/s41467-020-16777-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
