Abstract
The success of AI solutions in health systems depends on governance from use case inception through deployment and auditing. This proposed early pipeline governance framework for vendor AI solutions highlights a four-pronged approach: strategic alignment, executive sponsorship, impact and value case assessment, and risk assessment. Each component can be scaled to health systems of any size and the risk and impact assessments can take place simultaneously or sequentially.
Subject terms: Health policy, Health services
Introduction to the AI governance challenge
AI adoption in healthcare is growing rapidly and needs oversight and governance to ensure that its potential clinical and operational impact is realized. However, there are no paradigm frameworks on which health systems can rely to build their own AI governance structure. There is also uncertainty around the guidance that regulatory bodies will provide, compelling health systems to rely on academic and conceptual models that often lack structured, data-driven outcome assessments1. Most governance frameworks remain largely theoretical and categorization-based, leaving significant gaps for practical implementation2,3. The current health system AI governance recommendations may fall short in either being overly broad and focused on theoretical ethical and compliance concerns rather than real-world issues2,4,5, or excessively narrow and not applicable and adaptable to various health systems6.
Governance of the AI pipeline: the selection of vendor solutions and return on investment
The governance of AI in healthcare requires oversight of AI solutions from their inception as use cases through design and training, deployment, and ongoing monitoring and auditing. In the AI governance pipeline, model accuracy, while essential, has been emphasized, whereas relatively less attention has been paid to model selection, particularly for vendor solutions, which may be opaque7.
Model selection, which is the entrance to the AI solution pipeline, is a key step that, when well governed, helps ensure an appropriate return on investment (ROI) for the solution. Helping ensure that an AI solution delivers an ROI involves a projection of the impact that it will have and also an assessment of the risk it poses. While the appropriate reimbursement for AI solutions is still being explored8 and there exist retrospective calculations of ROI generated by specific AI solutions ROI9, the importance of considering ROI during solution selection has not been emphasized in the literature. However, it is essential to a genuinely responsible AI governance structure. Without an appropriate ROI, the implementation of AI to improve the accessibility and quality of healthcare, patient and clinician experience, and operational improvements will not be realized. Any innovation, such as health AI, must be sustainable and a large part of sustainability is achieving an ROI.
A proposed framework for the selection of vendor solutions
A real-world, adaptable, and translatable approach to selecting vendor AI solutions should include four components: alignment with a health system’s strategic goals, assignment of executive sponsorship and accountability, projected impact, and risk assessment with multi-stakeholder input. This four-pronged approach has several benefits. First, it can be readily adopted and adapted to the existing governance structure of various hospitals and health systems. It is also easily understood and accessible, requiring little to no additional personnel to implement. While this approach involves the input of multidisciplinary stakeholders, it also holds executive leaders responsible for model integration and adoption post-deployment. Finally, from inception, it facilitates ROI through alignment with strategic goals, demonstration of impact, and thorough risk assessment.
The ROI for all solutions is best realized if there is alignment with strategic goals, executive sponsorship, and a defined and measurable impact methodology. Because the development, training, and performance of solutions acquired from vendors may not be transparent, assessing impact and risk before acquiring these solutions can present a challenge. However, such assessments represent an essential (although often overlooked) step in the governance lifecycle10. Compared with the other domains, a thorough pre-acquisition risk assessment of vendor solutions is more dependent on the solution itself and is difficult to make independently. Close collaboration with the vendor, transparency, and due diligence in assessing the risk of a solution before acquiring it not only helps ensure a positive ROI but is also essential for responsible governance3,11.
Strategic alignment and executive sponsorship
Governing a vendor AI solution involves a significant commitment and investment. Alignment of AI governance overall with the strategic priorities of the system ensures consistency, removes friction, and also facilitates the realization of ROI. Misalignment risks diminishing focus on responsible AI and value, the AI capability is built to deliver12. The executive sponsor should be a senior leader who champions the proposed AI solution, is accountable for securing financial support, and proactively identifies mitigation strategies. Executive sponsorship is critical and must lead efforts to integrate the AI solution into existing workflows, securing stakeholder adoption, and promoting organizational awareness and education regarding its use. (Fig. 1).
Fig. 1. The four components of the proposed governance framework.
A real-world, adaptable framework for AI governance centers around four key components to guide health systems in evaluating and adopting vendor AI solutions. These four components are designed to ensure strategic alignment, shared accountability, measurable value, and minimized risk.
Impact and value case assessment
The Impact and Value Case Assessment, along with the Risk Assessment tools, are systematic ways of evaluating an AI solution’s potential for ROI. The Impact and Value Case has 6 domains: Executive Summary, Background and Problem Statement, Solution Landscape, Proposed AI Solution, Value Proposition and Quantifiable Objectives, and Cost Analysis. Overall, the assessment considers the current problem that the AI solution is intended to solve, whether an AI solution is necessary to solve the problem, whether a vendor-built solution is available, what is the quantifiable value that the solution will provide, and what the cost of the solution will be, broadly defined.
Figure 2 provides details of the Impact and Value Assessment and applies the assessment to a hypothetical Ambient AI scribe use case (Scribr).
Fig. 2. Vendor assessment framework: impact and value case assessment and application to ambient AI scribe use case.
Early assessment of the potential impact and value of a vendor AI solution is an important step in the governance process. This figure provides a detailed description of each step in the process, including an executive summary, a background and problem statement, evaluation of the solution landscape, a description of the proposed AI solution, the value proposition, and quantifiable objectives that the solution proposes to achieve, and the estimated cost of the solution. The application of each step to a hypothetical AI scribe use case (Scribr) illustrates the real-world utility of this framework.
Description, objective, benefits
The Impact and Value Case Assessment begins by identifying the problem or issue that the solution is proposed to address and the benefit it aims to achieve. For instance, consider efforts to improve postoperative care following heart surgery and a proposed AI solution that would better predict postoperative complications in this population. The primary benefit gained from the AI solution is the ability to preoperatively identify patients at high risk of complications, so that the facility can institute a strategy to regularly and consistently reduce the risk in those patients. Secondary benefits may include reductions in overall costs associated with complications, shorter lengths of stay (LOS) following heart surgery, and fewer reportable post-operative complications and/or readmissions, not to mention allowing the facility to deliver higher quality care and increase patient satisfaction.
Current situation and need for AI/Automation
To better understand the problem and provide context, the current workflow and process for addressing the problem are considered, emphasizing whether or not an AI-enabled solution is needed. In the example of postoperative complications discussed previously, the problem may not be the inability to accurately identify high-risk patients but may instead be poor intraoperative technique or an ineffective workflow to deal with complications postoperatively. In these examples, where the issue is one other than more accurately identifying patients who are at high risk of having complications after heart surgery, the proposed AI solution would not be effective in improving postoperative care.
In other instances, non-AI solutions may be as good, if not better, than an AI-enabled solution for addressing a problem. For example, perhaps a more robust risk calculator that considers multiple variables (but doesn’t rely on AI) would suffice. Not every problem requires an AI solution, and carefully choosing which problems benefit from AI the most helps predict the impact and ensure ROI13.
Existing vendors, solution description, and likelihood of success
The next step in creating an Impact and Value Case is identifying potential vendors, describing their solution in detail, and estimating the likelihood of its success in addressing the identified problem. Identifying potential vendors applies to those with “plug and play” solutions as well as those who will customize their solutions. For the surgical complication predictive model described previously, a niche vendor may have already developed, trained, and validated a solution that would only require fine-tuning in the local ecosystem, or a vendor may have a model that would require local data for training and validation. In either instance, the process of identifying potential vendors requires due diligence and familiarity with the AI solution landscape, which often caters to specific specialties or types of solutions. After surveying the landscape, a solution or a slate of solutions would be identified that best addresses the problem. The details of the solution are described, specifically how the solution will address the problem for which it is proposed. For the postoperative complication model, this entails describing the model’s output and how the output will be used to reduce post-operative complications. It is important at this step to also evaluate the solution’s “real-world” success at other health systems, the solution’s product roadmap and the vendor’s relationship with their existing customers.
Value proposition and quantifiable objectives
Whether the proposed value added by the AI solution is clinical, operational, or both, it is crucial to accurately estimate the solution’s impact in quantifiable units that can be measured. Even if a solution is accurate, that does not necessarily translate into a measurable impact14. For instance, in the case of the postoperative complication predictive model, accurately predicting an individual’s risk of complications after heart surgery doesn’t ensure that complications, overall costs, LOS, or reportable post-operative complications will be reduced or that patients will be more satisfied with higher quality care. An AI solution alone will not lead to a measurable impact or produce a meaningful ROI unless it is thoughtfully integrated into a workflow that can translate the model’s output into an operational and/or clinical impact. Predicting a solution’s impact begins with measuring the status quo and determining how much realistic benefit the solution will provide. Accurate and realistic predictions of the expected benefits that are quantifiable – like the average decrease in LOS, for example – can then be converted into a projected financial benefit. Not every benefit derived from an AI solution translates directly into a clear economic benefit: decreased pain and suffering associated that result from reduced complications may not confer monetary savings to the hospital or health system. In such instances, other types of financial or tangible benefits must be demonstrated to justify the monetary investment in the solution.
Cost analysis
Projected costs of an AI solution include the costs of acquiring it and also validating, piloting, deploying, and maintaining it. Part of this calculation is considering the timeframe over which deployment will occur and its scalability. Costs should also consider potential disruptions to existing workflows that may occur before the solution is fully functional.
Risk assessment
While terms like “uncertainty” and “safety” appear in AI governance policies4,5 and tools exist to identify potential legal risks of AI solutions15, formal risk assessments of AI solutions under consideration are rare. However, we believe the development of an objective, robust risk assessment tool to be an invaluable part of an AI Governance structure. To develop such a tool, we identified 12 domains from the perspective of multiple enterprise stakeholders and established guidelines. [Cybersecurity and Infrastructure Security Agency (CISA) and the National Institute of Standards and Technology (NIST)—particularly the NIST AI Risk Management Framework - (AI RMF)] Each domain was weighted according to an estimate of its contribution to the overall risk of harm compared to other domains. The average risk weight was 6.1%, whereas the high risk weight was 15%. All domains were considered average risk except for three: clinical documentation and decision support, quality and patient safety, and enterprise risk, which were high-risk domains. Within each domain, variables that contributed to risk assessment were identified.
Table 1 provides details of the Risk Assessment and applies the assessment to a hypothetical Ambient AI scribe use case (Scribr).
Table 1.
Vendor assessment framework: risk assessment and application to ambient AI scribe use case
| Domain | Description | Risk Weight (%) | Application to Ambient AI Scribe Use Case | Ambient AI Scribe Use Case Risk Assessment |
|---|---|---|---|---|
| Model Type | Assesses the type of AI/ML technology used (e.g., generative, predictive), considering potential opacity and complexity. | 6.1% | Speech-to-text; Generative AI (LLM) | High risk |
| Data Acquisition & Integrity | Evaluates how data is collected, labeled, validated, and governed to ensure reliability and minimize error. | 6.1% | Source data are live audio streams containing PHI. There is a risk of transcription errors and/ or mislabeled clinical information. | High risk |
| Model Performance | Examines the accuracy, precision, recall, and robustness of the model in clinical or operational environments. | 6.1% | Scribr has a < 5% transcription error rate; >95% overall accuracy (summary and coding) | Low risk |
| Interpretability & Transparency | Measures how explainable and understandable the AI outputs are to human stakeholders, especially clinicians and patients. | 6.1% | The inner summarization chain is opaque, but clinicians can compare the visit note output to their recollection of the visit. | Low risk |
| Scalability & Maintenance | Assesses whether the AI system can be updated, maintained, and scaled across the enterprise while preserving performance and safety. | 6.1% | Easy enterprise scalability. Low maintenance as updates are automatically pushed out to devices. | Low risk |
| Legal, Research & IP | Evaluates compliance with legal and regulatory frameworks, including ownership of models and data and use in research contexts. | 6.1% | ITSA executed between Scribr vendor and health system; data will not be used by the vendor for model training; EMR is HIPAA compliant. | Low risk |
| Enterprise Risk | Reviews how the AI model could impact organizational reputation, liability, compliance, and financial exposure. | 15% | Clinician training and periodic audits to assure accurate documentation and billing; audits during phased rollout; reporting process for identified errors that could impact patient care | High risk |
| Security | Considers cybersecurity risks, including data breaches, model inversion, and adversarial attacks. | 6.1% | FIPS 140-3 is used for data encryption | Low risk |
| Ethics & Access Fairness | Assesses the model’s alignment with principles of fairness, justice, and non-discrimination. | 6.1% | High accuracy across patient and clinician groups. | Low risk |
| Physician Engagement | Evaluates the extent to which physicians were engaged in the AI development and implementation processes. | 6.1% | Extensive education; clinical champions (early adopters) to provide in-person support; clinical sponsor with model oversight | Low risk |
| Clinical Doc & Decision Support | Reviews the AI model’s role in influencing clinical documentation, decision-making, and care delivery pathways. | 15% | Clinician education and periodic audits to assure accuracy; clinicians must attest to visit note’s accuracy before signing. | High risk |
| Quality & Patient Safety | Measures potential impact on patient outcomes, clinical safety, and healthcare quality indicators. | 15% | Required patient validation of high-risk medical information (e.g., medications and dosages, laterality for procedures, allergies). | High risk |
Model type, training, performance, and maintenance
For instance, within the Model Type domain, a generative AI model was considered higher risk than a model that relied on structured patient data (low risk) or audio and text data (intermediate risk). This is because generative AI models are more prone to inaccuracies and hallucinations than the other model types and thus pose a higher risk of harm16,17. Similarly, audio and text data require machine interpretation and data extraction, which introduces an opportunity for inaccuracies. Compared with other data sources, structured patient data has the lowest risk of introducing errors and is thus considered the lowest risk18.
The Data Risk and Integrity domain assesses the transparency of a vendor-built model’s training data. Risk is gauged based on the degree of transparency of the model’s training and testing datasets. If there is complete access to training and testing data, the model is considered low risk from the Data Risk and Integrity domain. However, the model is considered intermediate risk if there is only access to testing and training data statistics, or if there is only access to features used in the model or a data dictionary for features used in the model.
The Model Performance domain is assessed based on how the model’s performance has been assessed and the results of the assessment. A model would be deemed high risk if its performance had not been evaluated on an independent and representative data set or if its performance report or validation study was not available for review. If any potential concerns with a model’s performance had been reported, the model would be considered to be high risk.
The Interpretability and Transparency domain assesses the extent to which the logic behind the model’s output can be interpreted. The assessment requests measures of interpretability through tools such as SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations). If a solution lacks transparency, clinicians may struggle to trust or appropriately act on its recommendations, undermining effective clinical decision-making and potentially endangering patient safety. By contrast, interpretable AI solutions that clearly show how results are generated can foster trust among clinicians and patients, facilitating streamlined integration into care processes19,20. Transparency also enables health professionals to detect errors in AI suggestions, preventing unintended harms19. Interpretability and transparency mean understanding why a model produced a specific output. Tools like SHAP and LIME help explain which patient data points (like age or lab results) influenced a prediction, so clinicians can trust, validate, and act on the results21.
The Scalability and Maintenance domain assesses the vendor’s relationship with the solution post-deployment. It emphasizes the need for an established process for notification and support when a model is updated (intermediate risk), for establishing thresholds for model retraining (intermediate risk), and for monitoring and maintaining the data input pipeline (high risk).
Model and data security and compliance
The degree of data encryption and the execution of an Information Technology Service Addendum (ITSA) have the most significant impact on the Security domain’s risk assessment. FIPS 140-3 (Federal Information Processing Standard 140-3) is the gold standard for cryptographic security, mandated by NIST and required for federal agencies and healthcare organizations handling sensitive data22. Encryption in transit protects data from interception during transmission, while encryption at rest secures stored data against breaches. Given the risk of cyberattacks and regulatory requirements such as the Health Insurance Portability and Accountability Act (HIPAA), FIPS 140-3 compliance reduces vulnerabilities, safeguards patient data, and ensures regulatory adherence. Executing an ITSA between the health system and the vendor outlines the data security and privacy measures that will be enacted to ensure compliance with HIPAA.
The Legal, Research, and Intellectual Property domain emphasizes vendor access to sensitive patient-level data, whether an ITSA is needed, and whether the data will be used by the vendor for model training purposes. Sensitive patient-level data refers to a subset of individually identifiable health information that is afforded heightened privacy protections under U.S. federal law due to the nature of the information. This includes data which falls under the broader regulatory framework of Protected Health Information (PHI) as defined by HIPAA, as such this sensitive classification imposes stricter requirements for access, use, disclosure, and the use of sensitive data, and subsequent data use for training is considered high risk.
Ethics, access fairness, and clinical engagement
The Ethics and Access Fairness domain assesses whether either the use case or the model’s performance fairly distributes burdens and benefits. This domain analyzes the use case to assess whether the model aligns with principles of fairness, justice, and non-discrimination.
Assessing the degree of clinical oversight for outputs produced by AI solutions is the focus of the Clinician Engagement domain. There are two main considerations in this domain. The first is the extent to which clinicians are informed about the limitations of the model they are using, such as the risk of hallucinations, inaccurate predictions, and incomplete summaries from the EMR. The second is the degree of control that clinicians maintain over the model’s output. This applies both to determining the appropriateness of interventions that are triggered from the output and to assessing output for accuracy before sharing it with patients and other clinicians.
Enterprise risk, decision support, quality and safety
The Enterprise Risk domain is one of the three highest-weighted domains when assessing overall risk (15%). This is because the harm caused by AI solutions can have far-reaching financial, reputational, and operational implications for the entire health system. A failure in one AI system can also damage trust and cause cascading effects on other services. Within the Enterprise Risk domain, the consideration that contributes most to the overall assessment is whether the model’s output is retained in the electronic medical record (EMR). There are at least four reasons for the emphasis on whether AI outputs are recorded in the EMR. First, clinicians rely on accurate information in the EMR to deliver patient care. AI-enabled predictions are just that: predictions, and not necessarily clinical truths. Placing an AI-enabled prediction in the EMR risks having a clinician misunderstand how to integrate a prediction rather than factual clinical information into their clinical reasoning. Second, not only might AI-enabled predictions be misinterpreted, they may also be inaccurate. Even if the clinician correctly integrated the information, the information itself may have been wrong. Third, AI-generated outputs in the LMR are subject to HIPAA, FDA, and CMS regulations, raising compliance concerns. The legal implications of using AI to assist with diagnosis or treatment recommendations, as well as the potential for privacy concerns, have received attention in the literature23,24. Lastly, storing AI outputs also raises costly eDiscovery and data storage concerns25.
Clinical Documentation and Decision Support is another high-risk domain and is assessed based on how autonomously the solution functions without clinical oversight. The Clinician Engagement domain assesses the solution’s role in influencing clinical documentation, decision-making, and care delivery pathways, particularly the limitations of the output. This domain assesses the solution itself and its potential to directly impact clinical decisions, including diagnoses and treatments. A predictive model may provide a clinician with information without prompting an associated intervention. Particularly if the model is predicting a future event, such as sepsis, this can lead to confusion, frustration, and decreased clinical adoption26–28. A solution would be considered high risk if it is either overly autonomous or if it provides an output without recommending an appropriate clinical response.
Quality and Patient Safety is the third high-risk domain and is assessed based on previous safety events associated with the solution, whether there is a process for reporting safety concerns and events associated with the solution, and the extent to which unintended harms can be identified pre-deployment. In healthcare settings, AI systems always pose a risk of inadvertently contributing to diagnostic errors, treatment delays, or inappropriate interventions that can exacerbate existing vulnerabilities in patient care29. Without rigorous quality assurance and continuous safety monitoring, the integration of AI into clinical workflows may lead to adverse outcomes and undermine trust in healthcare delivery. Therefore, if a model poses a risk of unintended harm to patients - it is weighted as one of the highest risks in the AI risk assessment. By incorporating robust quality and patient safety measures into AI risk assessments, it enables the health system to establish systematic protocols for incident reporting, error analysis, risk mitigation, and real-world performance monitoring. Proactive strategies are essential not only for identifying latent safety threats before deployment but also for ensuring that AI tools continuously meet established healthcare quality indicators and reduce the likelihood of patient harm, and improve overall clinical outcomes.
A scalable implementation process for the assessment of vendor solutions
The first step in the early governance of use cases proposed by clinical and operational leaders includes executive sponsorship and strategic alignment with the health system’s goals. Every health system, regardless of its size, will have executive leadership (CEO, COO, CNO, CMO, CIO, CFO, etc.) as well as defined strategic goals. After the proposed AI solution has an executive sponsor and strategic alignment, the next step is to begin the process of estimating its impact and its potential for harm. These individual assessments can take place simultaneously or in sequence based on both the solution’s characteristics and the health system’s priorities and resources. For instance, even before completing the entire risk assessment, an intelligent supply-chain sensing operational solution for hospital consumable goods may be placed at a lower risk since it neither interfaces with patients, supports clinical decisions, nor has extensive exposure to PHI. In such a case, the assessment might begin with the Impact and Value Case and be followed by a formal risk assessment. Allowing flexibility as to whether the Impact and Value Assessment or the Risk Assessment occurs simultaneously or in sequence makes the framework scalable for smaller health systems that may not have the resources to complete the assessments simultaneously. Smaller health systems can also decentralize the use of these tools such that the effort is spread across several departments and is not the responsibility of any one.
Conclusion
This proposed four-pronged approach to the governance of vendor solutions early in the pipeline addresses critical gaps left by existing frameworks, which often lack practical, broad implementation strategies that focus more on accuracy than impact. The approach to early governance emphasizes the need for a thorough assessment of vendor solutions, which may be opaque. Vendors who are hesitant to share the measurable impact of their solutions, and not just the solution’s accuracy on training data, might be best approached with a degree of suspicion. The same process of due diligence applied to other large-scale investments by health systems should also be applied to AI solutions, perhaps even more so. Balancing projected impact with a comprehensive risk assessment provides the executive sponsor with the necessary information prior to investing in the solution. Furthermore, while not automatic, such stringency early in the pipeline helps ensure that the solution will realize an ROI that is aligned with the health system’s strategic goals. For AI solutions to realize their potential in improved health care access, quality, safety, and satisfaction, they must be sustainable.
In addition, the approach itself is adaptable to any health system, regardless of its size, complexity, or existing governance structure. It is readily scalable without adding additional layers of administrative burden. It is also systematic, thorough, and applicable to any AI solution. It is also intended to be iterative and adaptable, with aspects of impact prediction and risk assessment adjusted as necessary to both align with health systems’ needs and strategic goals. For instance, for many clinical models, nurse leaders and practicing nurses must be included in impact predictions and risk assessments. The same is true for other members of the healthcare team who have valuable insights to offer to the process.
Just as it is essential to measure and evaluate the ROI of AI solutions, it is also important that the process by which vendor solutions are chosen is also evaluated. Outcome metrics such as the accuracy with which the model’s impact was predicted or its risk assessed must be tracked, and the information used to further refine the process. When AI solutions either underperform or are associated with an adverse event, a thorough review must include whether there were gaps in the early assessment of impact and risk. Just as all of healthcare aims to continuously improve in quality and safety, so too must the governance of AI solutions. Thorough, early assessment of vendor solutions situates health systems best to realize an ROI, have sustainable AI programs, and fulfill their social contract to deliver the best possible healthcare.
Acknowledgements
We acknowledge Craig Solid, PhD, for his assistance in developing this manuscript.
Author contributions
C.B., D.B., A.Z., and L.K. drafted the manuscript, and C.B., L.K., R.W., S.S., and J.A. formulated the process described in the manuscript, including developing the tools described. All authors reviewed the manuscript.
Data availability
No datasets were generated or analyzed during the current study.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Trang, B. FDA commissioner: Health systems have to “step up” on AI regulation. https://www.statnews.com/2024/09/11/fda-health-ai-regulation-robert-califf-hospitals-role/ (2024).
- 2.Borkowski, A. A., Jakey, C. E., Thomas, L. B., Viswanadhan, N. & Mastorides, S. M. Establishing a Hospital Artificial Intelligence Committee to improve patient care. Fed. Pr.39, 334–336 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Parker, V. J., Economou, N. J. & Silcox, C. AI governance in health systems: aligning innovation, accountability, and trust. Preprint at https://healthpolicy.duke.edu/sites/default/files/2024-10/AI%20Governance%20in%20Health%20Systems.pdf (2024).
- 4.Liao, F., Adelaine, S., Afshar, M. & Patterson, B. W. Governance of Clinical AI applications to facilitate safe and equitable deployment in a large health system: Key elements and early successes. Front. Digit. Heal.4, 931439 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Nong, P., Hamasha, R., Singh, K., Adler-Milstein, J. & Platt, J. How academic medical centers govern AI prediction tools in the context of uncertainty and evolving regulation. NEJM AI1, 10.1056/AIp2300048 (2024).
- 6.Whittaker, R. et al. An example of governance for AI in health services from Aotearoa New Zealand. npj Digit. Med.6, 164 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Fehr, J., Citro, B., Malpani, R., Lippert, C. & Madai, V. I. A trustworthy AI reality-check: the lack of transparency of artificial intelligence products in healthcare. Front. Digit. Heal.6, 1267290 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Abràmoff, M. D. et al. A reimbursement framework for artificial intelligence in healthcare. npj Digit. Med.5, 72 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Bharadwaj, P. et al. Unlocking the value: quantifying the return on investment of hospital artificial intelligence. J. Am. Coll. Radiol.21, 1677–1685 (2024). [DOI] [PubMed] [Google Scholar]
- 10.Hussein, R. et al. Advancing Healthcare AI Governance: A Comprehensive Maturity Model Based on Systematic Review. medRxiv 2024.12.30.24319785 10.1101/2024.12.30.24319785 (2024).
- 11.MIT Technology Review Insights. Reframing digital transformation through the lens of generative AI | MIT Technology Review. https://www.technologyreview.com/2025/02/06/1111007/reframing-digital-transformation-through-the-lens-of-generative-ai/ (2025).
- 12.Center for Connected Medicine at UPMC. How Health Systems Are Navigating the Complexities of AI | CCM Reports. https://enterprises.upmc.com/app/uploads/2025/03/How_Health_Systems_Are_Navigating_The_Complexities_Of_AI_CCM_Reports.pdf.
- 13.World Health Organization. Ethics and governance of artificial intelligence for health. https://www.who.int/publications/i/item/9789240029200 (2021).
- 14.Binkley, C. & Loftus, T. Encoding Bioethics: AI in Clinical Decision-Making. (University of California Press, 2024).
- 15.Meeus, S. et al. AI in Healthcare: Navigating Legal Risk Assessment with JusticeBot. in Frontiers in Artificial Intelligence and Applications, Volume395: Legal Knowledge and Information Systems 384–386.
- 16.Barassi, V. Toward a Theory of AI Errors: Making Sense of Hallucinations, Catastrophic Failures, and the Fallacy of Generative AI. Harv. Data Sci. Rev. 10.1162/99608f92.ad8ebbd4 (2024).
- 17.Varghese, J. & Chapiro, J. ChatGPT: The transformative influence of generative AI on science and healthcare. J. Hepatol.80, 977–980 (2024). [DOI] [PubMed] [Google Scholar]
- 18.Meystre, S. M. et al. Clinical data reuse or secondary use: current status and potential future progress. Yearb. Méd. Inform.26, 38–52 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Abgrall, G., Holder, A. L., Dagdia, Z. C., Zeitouni, K. & Monnet, X. Should AI models be explainable to clinicians?. Crit. Care28, 301 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Tighe, P., Mossburg, S. & Gale, B. Artificial intelligence and patient safety: promise and challenges. PSNet [internet] (2024).
- 21.Knapič, S., Malhi, A., Saluja, R. & Främling, K. Explainable Artificial Intelligence for human decision-support system in medical domain. arXiv10.48550/arxiv.2105.02357 (2021).
- 22.National Institute of Standards and Technology (US). Security requirements for cryptographic modules. 10.6028/nist.fips.140-3 (2019).
- 23.Duffourc, M. & Gerke, S. Generative AI in health care and liability risks for physicians and safety concerns for patients. JAMA330, 313–314 (2023). [DOI] [PubMed] [Google Scholar]
- 24.Rosic, A. Legal implications of artificial intelligence in health care. Clin. Dermatol.42, 451–459 (2024). [DOI] [PubMed] [Google Scholar]
- 25.Fitzmaurice, G. AI is causing a data storage crisis for enterprises | ITPro. https://www.itpro.com/hardware/storage/ai-is-causing-a-data-storage-crisis-for-enterprises (2024).
- 26.Habib, A. R., Lin, A. L. & Grant, R. W. The Epic Sepsis model falls short—the importance of external validation. JAMA Intern. Med.181, 1040–1041 (2021). [DOI] [PubMed] [Google Scholar]
- 27.Wong, A. et al. External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Intern. Med.181, 1065–1070 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.NIH. Widely used sepsis prediction tool is less effective than Michigan doctors thought | NHLBI, NIH. https://www.nhlbi.nih.gov/news/2021/widely-used-sepsis-prediction-tool-less-effective-michigan-doctors-thought (2021).
- 29.Rajkomar, A., Dean, J. & Kohane, I. Machine learning in medicine. N. Engl. J. Med.380, 1347–1358 (2019). [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
No datasets were generated or analyzed during the current study.


