Abstract
Background
While generative artificial intelligence (AI) is rapidly proliferating in healthcare research and clinical settings, there is a lack of actionable research ethics standards that reflect the unique technical features of generative AI, such as hallucination, agentic autonomy, and decontextualization. Existing AI ethics studies primarily focus on presenting universal principles, and Delphi studies report difficulties in deriving expert consensus owing to the overly broad scope of AI concepts. We aim to systematically identify ethical issues in generative AI research within the healthcare domain and to develop practical ethical guidelines and a checklist that researchers can use across the entire research and development lifecycle.
Methods
We applied a three-stage modified Delphi method in accordance with the Conducting and REporting DElphi Studies (CREDES) guidelines. Thirty-five experts spanning the fields of medicine, law/ethics/policy, and AI technology were invited. Round 1 was conducted as an in-person workshop involving 18 experts, while Rounds 2 (n = 32) and 3 (n = 27) were conducted via online surveys rating 56 and 43 items, respectively, using a 7-point Likert scale. Consensus criteria were set at interquartile range (IQR) ≤ 1.5 and coefficient of variation (CV) < 0.5 (high consensus) for Round 2, and a stricter criterion of IQR ≤ 1.0 (strengthened consensus) for Round 3.
Results
In Round 2, 97.7% of Likert-scaled items (42/43) entered the acceptable range for guideline adoption. In Round 3, 60.5% of items reached consensus under the strengthened criterion (IQR ≤ 1.0), and 46.5% achieved strong consensus (IQR ≤ 1.0 and CV ≤ 0.2). By domain, documentation standards (mean 6.06), safety measures (mean 5.95), and evaluation methods (mean 5.86) recorded the highest importance. For individual items, explainable AI (mean 6.48), ensuring the diversity of training data (mean 6.44), and human-in-the-loop (mean 6.33) were derived as core items and top-priority strategies. Based on these findings, an ethical framework comprising three domains (data, governance, and design-by-value) and eight value dimensions, alongside a lifecycle checklist categorized into pre-development, development, and post-deployment stages, was developed.
Conclusions
We developed a differentiated ethical framework and practical checklist that reflect the technical features of generative AI. These outputs can serve as sector-specific guidance for healthcare under the Framework Act on AI and as criteria for institutional review board reviews.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12910-026-01513-4.
Keywords: generative AI, healthcare ethics, Delphi study, research ethics guidelines, explainable AI, value alignment, checklist
Background
Since 2023, the rapid growth of generative artificial intelligence (AI) sparked by large language models (LLMs) has begun to transform the healthcare system [1, 2]. While early generative AI primarily focused on natural language processing and information summarization, by late 2024, agentic AI capable of autonomous planning, execution, and self-evaluation, alongside multimodal AI that integrates multiple data types such as text, images, and voice, emerged as core drivers of healthcare innovation [2]. As agentic and multimodal AI proliferate in clinical settings [3], ethical dilemmas qualitatively distinct from those of traditional predictive AI are surfacing. In particular, the autonomy of generative AI, hallucination based on probabilistic reasoning [4], and decontextualization—where the source and context of information are lost—create serious ethical challenges that obscure the locus of responsibility in medical practice.
Regulatory approaches to healthcare AI vary significantly across nations, complicating the establishment of an integrated governance framework [5]. The European Union, through the AI Act, categorizes medical AI as high-risk systems, with core high-risk obligations scheduled from August 2026 [6]. Meanwhile, South Korea’s Framework Act on AI, promulgated on January 21, 2025 and in force since January 22, 2026, classifies healthcare AI that poses risks to life or safety as high-impact AI, but adopts an approach centered on self-verification and best-efforts obligations rather than prior approval [7]. Despite these policy discussions, however, there remains insufficient provision of concrete, actionable ethical standards that researchers and clinicians who actually develop and validate generative AI can reference for routine decision-making [8].
The United States established a risk-based regulatory framework under the Biden administration’s Executive Order on the Safe, Secure, and Trustworthy Development and Use of AI (14110, 2023). However, the Trump administration revoked this and replaced it with an innovation-driven Executive Order (14179, 2025), shifting the direction toward securing technological competitiveness through deregulation [9, 10]. Such disparities in regulatory approaches impact international joint research and multinational clinical trials, with the lack of practical implementation guidelines being particularly problematic.
In clinical settings, the agentic nature of generative AI reshapes fundamental changes in healthcare professionals’ decision-making processes. As AI autonomously generates and proposes alternatives, automation bias (the uncritical acceptance of AI recommendations by medical staff) and deskilling (the weakening of professional judgment) are emerging as major ethical issues [11]. These risks cannot be resolved solely through improvements in technical accuracy; they can only be mitigated when explainability and human-in-the-loop mechanisms are integrated throughout the lifecycle [12]. Therefore, research ethics for generative AI demands a new responsibility-sharing model that goes beyond mere harm prevention to operate in a zone where the agency of healthcare professionals and patient safety are mutually complementary.
Furthermore, the creative outputs of generative AI pose new questions regarding the integrity and reliability of medical data. Generating synthetic medical images or virtual patient cases to augment research data can be an innovative means to resolve data scarcity but carries the risk of distorting clinical evidence by blurring the boundaries between reality and fiction [2]. This reinforces the argument that trust should be viewed not merely as an indicator of performance but as a relational value combining technical robustness and data veracity [13]. Consequently, the ethical governance of generative AI must be designed as a dynamic process that transcends static legal regulations to monitor bias and execute value alignment throughout the technology’s lifecycle. Specific, standardized reporting and verification systems are needed to support that process.
Studies have highlighted the tension between the technological potential offered by LLMs in healthcare and the accompanying risks from multiple angles. For example, Denecke et al. [2], in a three-round Delphi study of health informatics and medical natural language processing experts, identified benefits such as improved clinical documentation efficiency and communication in multilingual environments, while simultaneously pointing out the risks of misinformation and biased decision-making. Starke et al. [13] proposed that trust in medical AI is a social construct extending beyond simple performance metrics, suggesting that health equity and researcher diversity are prerequisites for building trust.
However, reaching a consensus on AI ethics is inherently difficult; even when achieved, the dual nature of the technology is often pronounced. Stahl et al. [14], in a large-scale Delphi study, attempted to prioritize ethical issues and mitigation strategies but failed to reach a clear consensus (with top strategies heavily skewed toward non-specific measures such as education, investigative journalism, and the exchange of best practices, while legal and regulatory instruments were excluded from the priorities). They identified the overly broad scope of the AI concept and systemic complexity of the ecosystem as the causes, calling for a sophisticated conceptual reconfiguration centered on specific technologies and application domains. Similarly, Mahl et al. [15], in a public health Delphi study, observed divergent attitudes depending on the application area; despite high consensus levels, experts acknowledged serious algorithmic bias and privacy risks while remaining skeptical about using AI for community engagement. These results suggest that AI is not an isolated tool but a system that interacts across the healthcare ecosystem. This necessitates a systemic perspective that considers the balance among interdependent values, moving beyond reductionism that evaluates values in isolation. To achieve this, there is a growing demand for methodological rigor in Delphi techniques, such as ensuring panel diversity, defining clear statistical consensus criteria, and providing transparent feedback between rounds [16].
Accordingly, building on the need for a “more refined and specific concept” proposed by Stahl et al. [14], we sought to address the consensus difficulties reported in previous studies by narrowing the scope to a clear technological subject, “generative AI,” and a specific application context, “healthcare research and development (R&D).” Specifically, this study addresses the following research questions: (1) What are the ethical values and principles that experts should prioritize in healthcare generative AI research? (2) How can a practical ethical checklist be structured for researchers to apply in practice across the entire R&D lifecycle? We used a three-stage modified Delphi method in compliance with the Conducting and REporting DElphi Studies (CREDES) reporting guidance [17], and the CREDES checklist is included in Multimedia Appendix 1.
Methods
Study design
We conducted a three-stage modified Delphi study following the CREDES guidelines [17]. The design combined (i) an initial qualitative Delphi focus group (Round 1) to refine scope and generate content, with (ii) two subsequent quantitative Delphi survey rounds (Rounds 2–3) to rate the importance of candidate items and strengthen consensus [18–20]. This modification was selected because generative AI in healthcare involves rapidly evolving technical concepts and regulatory contexts; a structured qualitative stage helps establish shared conceptual grounding and improves item clarity prior to quantitative consensus rating [16, 20].
Expert panel recruitment and composition
Purposive sampling was used to assemble a multidisciplinary expert panel spanning three domains: (1) medicine and clinical research, (2) law/ethics/policy (including institutional review board (IRB)-related expertise), and (3) AI and health informatics. A total of 35 experts were invited based on professional roles, domain expertise, and relevant scholarly or practical experience [19]. The target panel size was informed by prior Delphi studies in healthcare AI and common methodological recommendations, while allowing for attrition across rounds [2, 13, 15, 19].
To enhance both feasibility and breadth, we adopted a two-tier panel approach typical of modified Delphi designs. First, in the core working group (Round 1), 18 experts participated in a qualitative Delphi focus group. Second, in the expanded Delphi panel (Rounds 2–3), we recruited additional experts (“outer members”) for the quantitative rating rounds through relevant academic groups and professional associations across the three domains, expanding the panel beyond the core working group. These additional participants were not involved in the Round 1 focus group discussion and participated only in the anonymized survey rounds.
In Round 2, 32 experts participated (core + additional recruits). In Round 3, 27 experts participated (five non-respondents from Round 2; attrition rate 15.6%), which is comparable to attrition commonly reported in multi-round Delphi studies. Participant characteristics are summarized in Table 1, and a detailed discipline breakdown is provided in Multimedia Appendix 4.
Table 1.
Demographic characteristics of expert panel
| Category | Sub-category | Round 2 (N = 32) | Round 3 (N = 27) |
|---|---|---|---|
| Gender | Male | 15 (46.9%) | 12 (44.4%) |
| Female | 17 (53.1%) | 15 (55.6%) | |
| Work Experience | Less than 5 years | 3 (9.4%) | 3 (11.1%) |
| 5–10 years | 10 (31.3%) | 9 (33.3%) | |
| 11–20 years | 14 (43.8%) | 12 (44.4%) | |
| 20 + years | 5 (15.6%) | 3 (11.1%) | |
| GenAI Utilization | Limited | 6 (18.8%) | 5 (18.5%) |
| Moderate | 18 (56.3%) | 11 (40.7%) | |
| Professional | 8 (25.0%) | 11 (40.7%) | |
| IRB Experience | Yes | 15 (46.9%) | 15 (55.6%) |
| No | 17 (53.1%) | 12 (44.4%) |
Survey instrument development
The survey instrument was developed based on a systematic literature review and the results of the Round 1 expert workshop. First, foundational items were derived by reviewing domestic and international literature on healthcare AI ethics [5, 12, 21], international regulatory trends (the EU AI Act [6] and the Framework Act on AI [7]), and existing ethical frameworks [22–24]. In the Round 1 workshop, experts reviewed these foundational items and proposed additional issues reflecting the technical features of generative AI. Based on this, a survey instrument consisting of 56 items across nine thematic areas was developed: data protection, evaluation methods, safety measures, documentation standards, attribution of responsibility, monitoring, explainability, bias management, and value alignment (see Multimedia Appendix 2 for the full questionnaire). Each item was measured using a 7-point Likert scale (1 = not important at all, 7 = very important), and ranking scales were concurrently applied for the bias effects and value-alignment domains. Prior to conducting the main survey, the draft questionnaire underwent a pilot review by three experts with experience in healthcare AI ethics research to evaluate item clarity, content validity, and the appropriateness of the response burden. Based on their feedback, item phrasing was revised and overlapping items were consolidated (Figure 1).
Fig. 1.
Participant flow diagram
Delphi procedure
The study was conducted from August to November 2025 in three stages.
Round 1 (qualitative Delphi focus group; August 8, 2025)
18 experts participated in an in-person, structured discussion focused on (i) clarifying the scope of “healthcare generative AI research” for this guideline, (ii) identifying and refining ethically salient issues specific to generative AI (e.g., hallucination, autonomy/agentic behavior, decontextualization), and (iii) generating candidate items and practical implementation considerations for subsequent surveys. Outputs from Round 1 were used to structure and refine the Round 2 questionnaire (item pool and domain structure), rather than to quantify consensus.
Round 2 (quantitative Delphi survey; September 29–October 2, 2025)
The expanded Delphi panel completed a first online survey rating the importance of 56 items using a 7-point Likert scale (1 = not important at all, 7 = very important). Two sections (bias effects and value-alignment mechanisms) used ranking formats to elicit priorities (Methods).
Round 3 (quantitative Delphi survey; November 5–12, 2025)
Round 3 was conducted with (i) controlled feedback from Round 2 (aggregated statistics and anonymized comments) and (ii) a refined item set totaling 43 items. The Round 3 instrument consisted of: (1) re-rating of items with unresolved agreement in Round 2 (12 items) and (2) operational sub-items (31 items) developed to strengthen and specify issues already raised in Round 1 discussions and Round 2 open-ended responses. These were not treated as conceptually “new themes” but as granular, checklist-ready specifications (e.g., clarifying actors, stages, verification steps, and implementable safeguards) intended to improve actionability and reduce ambiguity when translating principles into lifecycle checklist items.
Data analysis and consensus criteria
We summarized Likert-type items using mean, median, interquartile range (IQR), and coefficient of variation (CV). Analyses were conducted in R (version 4.5; R Foundation for Statistical Computing, Vienna, Austria) using tidyverse, psych, and ggplot2. Missing responses were handled by item-level exclusion for the relevant analysis.
Consensus definitions (Likert items)
Consistent with contemporary Delphi reporting recommendations emphasizing transparent, pre-specified thresholds [17, 20], we applied a priori consensus rules. For Round 2, these were: high consensus: IQR ≤ 1.5 and CV < 0.5, moderate consensus: IQR ≤ 2.0 and CV < 0.8, and low consensus: IQR > 2.0 or CV ≥ 0.8. For Round 3 (strengthened consensus), we pre-specified a stricter dispersion criterion of IQR ≤ 1.0 to identify items showing tightened convergence after feedback and refinement. Items simultaneously satisfying IQR ≤ 1.0 and CV ≤ 0.2 were classified as having achieved strong consensus.
Carry-forward and refinement rules
Round 3 included (a) all Round 2 items categorized as moderate or low consensus for re-rating and (b) operational sub-items created by decomposing broader Round 1/2 discussion points into implementable statements (e.g., specifying verification stages, IRB review elements, or bias-management actions). These sub-items were explicitly anchored to previously discussed domains and were used to support downstream checklist construction; they were not introduced as unrelated “new” constructs late in the Delphi. Specifically, the 31 operational sub-items were derived as follows: 10 items on emerging ethical issues (e.g., human-in-the-loop, explainable AI, hypersuasion prevention) were distilled from Round 1 panel discussions on the agentic and persuasive properties of generative AI; 4 items on expert verification qualifications and 5 items on IRB review elements were operationalized from the Round 1 proposal for a standardized expert verification procedure; and 12 items on bias type prioritization and mitigation strategies were developed from Round 2 open-ended responses identifying specific bias categories and management approaches (see Multimedia Appendix 2 for the full Round 3 questionnaire).
Handling items not meeting consensus
Items that did not meet dispersion-based consensus thresholds under the strengthened Round 3 criteria were not discarded from the narrative; instead, they were handled under a conservative negative rule. They were not promoted to required checklist elements on the basis of Delphi results alone. Depending on their practical relevance and residual support, they were retained as recommended, conditional, or contextual considerations, or excluded from the general checklist. No item with low or no consensus on the same narrow obligation was elevated to REQ; when a related safeguard was retained as REQ, the audit mapping identifies the independent consensus-supported or baseline procedural rationale.
Ranking-format items (priority elicitation)
Two sections used ranking formats to elicit relative priorities: bias-effect sources (3-point ranking) and value-alignment mechanisms (6-point ranking). As ranking responses are ordinal and, in the 3-point format, may involve ties or coarse categorization, we summarized ranking results descriptively using (i) the distribution of ranks (e.g., proportion assigned highest priority) and (ii) median rank and IQR of ranks for each item. Lower ranks indicate higher perceived severity/priority. Ranking results were used to identify relative priorities within the domain rather than to compute Likert-style means or CV-based consensus classifications.
Separation of empirical and interpretive layers
Mean scores, medians, IQRs, CVs, ranking distributions, and consensus labels constitute the empirical Delphi findings. The ethical framework, value dimensions, lifecycle placement, and checklist status categories represent structured post-Delphi synthesis. To make this distinction explicit, the checklist tables report both the empirical evidence basis (e.g., high consensus, strong consensus, ranked priority, no consensus) and the final checklist status (REQ, REC, CON, CTX, or EXC). Conceptual links to bioethics traditions and policy frameworks are presented in the Discussion as normative interpretation, not as additional Delphi findings.
Framework synthesis and checklist development
After completion of Rounds 2–3, we synthesized the final ethical framework and lifecycle checklist through a structured mapping process consisting of four steps. First, all Delphi items were organized under the nine thematic areas used in the questionnaire. Second, we consolidated these themes into three overarching domains (data, governance, and design-by-value) and eight value dimensions: privacy, trust, safety, forward-looking responsibility, backward-looking responsibility, explainability, debiasing, and alignment. Third, each Delphi item (and each Round 3 operational sub-item) was mapped to at least one value dimension and to a research-lifecycle stage (pre-development, development, post-deployment, or cross-cutting). Fourth, items were assigned a final checklist status based on the following decision rules, with analytic elevations permitted only under the four conditions specified below:
REQ (required): Items meeting the strengthened Round 3 consensus criterion (IQR ≤ 1.0), strong Round 3 consensus (IQR ≤ 1.0 and CV ≤ 0.2), or high Round 2 consensus (IQR ≤ 1.5, CV < 0.5), when the item was a minimum safeguard for patient safety, protected health data, accountable documentation, IRB/regulatory compliance, or another baseline governance control. High- or strong-consensus items were not automatically assigned REQ if they functioned primarily as secondary safeguards, public-facing transparency aids, or context-dependent technologies.
REC (recommended): Items with high or moderate support that functioned as secondary safeguards, implementation aids, or useful practices without sufficient grounds for mandatory status. This category also included items that showed narrowed disagreement in Round 3 but did not meet strengthened consensus.
CON (conditional): Empirically supported items whose appropriateness depends on the data environment, deployment architecture, clinical specialty, institutional resources, or legal context (e.g., privacy-enhancing technologies whose necessity depends on the data environment).
CTX (contextual): Items retained as setting-specific considerations or institutional options, but not generalized as checklist requirements.
EXC (excluded): Items that did not reach consensus by the end of Round 3 (IQR > 2.0 or classified as low consensus) and were therefore not promoted to the checklist as general requirements.
For reporting clarity, these status codes were used throughout the checklist and audit tables. Items could be elevated to REQ above their raw Delphi convergence label only under one of four pre-specified conditions: (E-A) baseline infrastructure or governance-compliance controls in the high-sensitivity health-data context; (E-B) synthesized items operationalizing Round 1 generative-AI-specific themes with Round 3 sub-item support, used in this study exclusively for agentic-AI concerns; (E-C) external-oversight requirements where the panel endorsed oversight necessity without converging on a single macro-governance model; and (E-D) top-ranked priorities from Round 2 ranking-format sections, where Likert-style thresholds did not apply. CV-only Round 3 convergence was treated as supporting evidence for synthesis or elevation, not as consensus. Each elevation is flagged in Multimedia Appendix 3 Table B with its corresponding code (E-A through E-D) and rationale; the Table B header further specifies the pre-specified conditions under which Delphi items were combined, transformed in wording, or synthesized into checklist items.
Items from the ranking-format sections (bias effects and value alignment) were translated into checklist priority using ordinal rank position, median rank, and the consistency of accompanying panel comments rather than Likert-style consensus categories. In Table 2, these rows are therefore labeled as ranked priorities rather than as high/moderate/low consensus items.
Table 2.
Condensed audit trail: Mapping from Delphi items to ethical framework and checklist elements
| Value Dimension | Delphi Item (Round) | Thematic Area | Lifecycle Stage | Checklist Element | Status | Consensus Status |
|---|---|---|---|---|---|---|
| Domain 1: Data | ||||||
| Privacy | Q1.1.5 (R2) | Data Protection | Pre-Dev | Strict data access protocols | REQ | High (R2) |
| Privacy | Q1.1.6 (R2) | Data Protection | Pre-Dev | Risk-based network separation/segmentation plan | REQ | Moderate (R2); Elevation E-A |
| Privacy | Q1.1.7 (R2) | Data Protection | Pre-Dev | Contractual liability for 3rd party sharing | REQ | High (R2) |
| Privacy | Q1.1.3 (R2) | Data Protection | During | Secure multi-party computation | CON | High (R2) |
| Trust | Q1.2.3 (R2) | Evaluation | During | Expert verification of outputs | REQ | High (R2) |
| Trust | Q1.2.5 (R2) | Evaluation | During | Prompt robustness review | REQ | High (R2) |
| Trust | Q1.2.6 (R2) | Evaluation | During | Reproducibility check | REC | High (R2) |
| Trust | Q1.2.2 (R2); Q1-1 (R3) | Evaluation | Post-Dep | Verification benchmarks | REC | Moderate (R2); no strengthened consensus (R3) |
| Safety | Q1.3.6 (R2) | Safety | Pre-Dev | Use category limitation | REQ | High (R2) |
| Safety | Q2-2 (R3) | Safety | During | Decision boundary setting | REQ | Strong (R3) |
| Safety | Q1.3.2 (R2) | Safety | During | Uncertainty quantification standards | REQ | High (R2) |
| Safety | Q7-4 (R3) | Emerging Issues | During | Human-in-the-loop for high-risk decisions | REQ | Strong (R3) |
| Safety | Q2-1 (R3) | Safety | During | AI output reliance score | REC | Strong (R3) |
| Domain 2: Governance | ||||||
| Fwd Responsibility | Q2.2.2 (R2) | Documentation | Pre-Dev | Data source documentation | REQ | High (R2) |
| Fwd Responsibility | Q2.2.4 (R2); Q3-1 (R3) | Documentation | During | Data security plan | REQ | Moderate (R2); below strengthened threshold (R3); Elevation E-A |
| Fwd Responsibility | Q2.2.1 (R2) | Documentation | Post-Dep | Model card | REC | High (R2) |
| Fwd Responsibility | Q2.2.3 (R2) | Documentation | Post-Dep | Bias audit report | REC | High (R2) |
| Fwd Responsibility | Q3-2 (R3) | Documentation | Post-Dep | Continuous monitoring plan | REQ | Strong (R3) |
| Bwd Responsibility | Q2.3.6 (R2) | Liability | Pre-Dev | Shared liability protocols | REQ | High (R2) |
| Bwd Responsibility | Q4-1 (R3) | Liability | During | Healthcare professional responsibility | REQ | Strong (R3) |
| Bwd Responsibility | Q4-2 (R3) | Liability | During | Institution responsibility | CTX | IQR consensus (R3) |
| Bwd Responsibility | Q4-3 (R3) | Liability | — | Developer responsibility | EXC | No consensus (R3) |
| Domain 3: Design-by-Value | ||||||
| Explainability | Q2.5.1 (R2) | Explainability | Field-specific | Clinical decision support | REC | High (R2) |
| Explainability | Q7-3 (R3) | Emerging Issues | All stages | Explainable AI requirement | REQ | Strong (R3) |
| Explainability | Q6-3 (R3) | Explainability | — | Hospital administration | EXC | No consensus (R3) |
| Debiasing | Q10-2-1 (R3) | Bias Mgmt | Pre-Dev | Training data diversity | REQ | Strong (R3) |
| Debiasing | Q10-2-5 (R3) | Bias Mgmt | During | Transparency & explainability of bias risks | REQ | Strong (R3) |
| Debiasing | Q10-2-3 (R3) | Bias Mgmt | During | Bias mitigation algorithms | REC | Strong (R3) |
| Debiasing | Q10-2-2 (R3) | Bias Mgmt | Post-Dep | Regular bias monitoring | REQ | Strong (R3) |
| Alignment | Q3.3.2 (R2) | Value Alignment | 1st priority | Medical accuracy | REQ | Ranked priority 1 (R2); Elevation E-D |
| Alignment | Q3.3.1 (R2) | Value Alignment | 2nd priority | Patient-centered values | REQ | Ranked priority 2 (R2); Elevation E-D |
| Alignment | Q3.3.4 (R2) | Value Alignment | 3rd priority | Social values | REC | Ranked priority 3 (R2) |
| Alignment | Q3.3.6 (R2) | Value Alignment | — | Value balancing | EXC | Low (R2) |
REQ = required; REC = recommended; CON = conditional; CTX = retained as a context-dependent consideration; EXC = not promoted to the checklist as a general requirement (see Methods for decision rules). Strong (R3) = IQR ≤ 1.0 and CV ≤ 0.2; IQR consensus (R3) = IQR ≤ 1.0 only; Below strengthened threshold (R3) = disagreement narrowed in Round 3 but the prespecified IQR threshold was not fully met; High (R2) = IQR ≤ 1.5 and CV < 0.5; Moderate (R2) = IQR ≤ 2.0 and CV < 0.8; Ranked priority (R2) = ordinal priority derived from the ranking-format section rather than Likert-style consensus; No consensus / Low = thresholds not met. Elevation E-A to E-D indicates that final checklist status is stronger than the raw convergence label under the elevation rules in Methods; the rationale is documented in Multimedia Appendix 3 Table B. CV-only Round 3 convergence is treated as supporting evidence for mapping or elevation, not as consensus. R2 = Round 2; R3 = Round 3; Fwd = Forward-looking; Bwd = Backward-looking. For complete item-level statistics, see Multimedia Appendices 5 and 6
This approach preserved a clear link from empirical consensus results to the final framework structure and checklist wording. Table 2 provides a condensed mapping illustrating how key Delphi items were translated into framework domains, lifecycle stages, and checklist elements; this table should be read together with the full checklist text (Table A), the full audit mapping and elevation-code rationale (Table B) in Multimedia Appendix 3, and the item-level statistics in Multimedia Appendices 5 and 6.
Ethical considerations
This study was approved by the IRB of Yonsei University (approval number: 4-2025-0777). Written informed consent was obtained from all expert panel members after they received explanations regarding the purpose and methods of the study, voluntary nature of their participation, and guarantee of anonymity. The collected data were anonymized for analysis, and personally identifiable information was stored separately to be destroyed upon the completion of the study. This study was conducted in accordance with the Bioethics and Safety Act [25] and the Declaration of Helsinki.
Results
Round 1: expert forum and establishment of the research framework
Round 1 was conducted in the form of an expert forum on August 8, 2025, with 18 participating experts from the fields of medical ethics (n = 9), law (n = 1), philosophy (n = 2), data science (n = 2), policy research (n = 2), and the AI industry (n = 1). The remaining one expert had a combined background spanning clinical practice and health informatics. The objective was to review the foundational items derived through the literature review and to structure the Delphi survey instrument. The research team and panel categorized the scope of AI use into its use as a research tool and as a research subject, confirming the need for comprehensive guidelines applicable to both domains. The panel agreed that the ethical principles would be based on the three principles of the Belmont Report (autonomy, beneficence, and justice) [26], while incorporating additional principles such as transparency, accountability, and inclusiveness to reflect the features of generative AI. Through discussion, four core review domains were established: autonomy (the potential for excessive AI intervention and manipulation), safety (hallucination and information errors), explainability (the black-box problem), and fairness/bias (lack of representativeness in training data).
Three key issues raised during the panel discussion were directly reflected in the design of the Round 2 survey. First, the issue of data sovereignty, where domestic medical data are transmitted to overseas clouds when using foreign LLMs, was translated into the privacy-domain items “liability for compensation under third-party utilization contracts” and “network separation.” Second, the industry panel criticized that “abstract ethical principles do not work in the field” and expressed the burden on development sites owing to overlapping legal and ethical regulations. This became the rationale for designing the final output as a dual structure comprising declarative principles and an actionable checklist. Third, the introduction of a verification procedure based on a standardized evaluation dataset that can be used during the IRB review stage was proposed. This led to the development of the “verification benchmark” item in Round 2 and the “expert verification stage” item in Round 3.
Round 2: derivation of core ethical principles
The Round 2 Delphi survey, consisting of 56 items, was conducted with 32 experts. Among 43 Likert-scaled items, 28 (65.1%) reached high consensus and 97.7% entered the acceptable range; the two ranking-format domains are reported separately below using rank-agreement metrics rather than importance-consensus classifications. These findings indicate a robust consensus among the expert group regarding the necessity of research ethics for healthcare generative AI.
According to the item-level analysis results (Table 3), the items demonstrating the highest importance were explainability in clinical decision support (mean 6.68, IQR 0.25, CV 0.100) and explainability in diagnostic support (mean 6.57, IQR 1.00, CV 0.087). Both items exhibited a very low CV (< 0.10), highlighting a prominent convergence of views among the experts. The top 10 items were distributed across six domains: explainability (n = 2), evaluation (n = 3), documentation (n = 2), safety (n = 1), data protection (n = 1), and responsibility (n = 1), suggesting the presence of multidimensional ethical demands rather than the prioritization of a single specific domain. Notably, the inclusion of three items from the evaluation domain—expert verification (mean 6.29), prompt robustness (mean 6.21), and result reproducibility (mean 6.11)—within the top 10 indicates that review by human experts and system robustness are prioritized over automated metrics.
Table 3.
Top 10 highest-rated items and low-consensus items (Round 2, n = 32)
| Rank | Domain | Item | Mean | Median | SD | IQR | CV |
|---|---|---|---|---|---|---|---|
| 1 | Explainability | Clinical decision support (Q2.5.1) | 6.68 | 7.0 | 0.67 | 0.25 | 0.100 |
| 2 | Explainability | Diagnosis support (Q2.5.2) | 6.57 | 7.0 | 0.57 | 1.00 | 0.087 |
| 3 | Evaluation | Expert verification (Q1.2.3) | 6.29 | 6.0 | 0.90 | 1.00 | 0.143 |
| 4 | Documentation | Data source documentation (Q2.2.2) | 6.29 | 7.0 | 0.94 | 1.00 | 0.149 |
| 5 | Safety | Explicit explanation of AI reasoning (Q1.3.4) | 6.25 | 6.0 | 0.70 | 1.00 | 0.112 |
| 6 | Data protection | Strict data access protocol (Q1.1.5) | 6.21 | 6.5 | 1.23 | 1.00 | 0.198 |
| 7 | Evaluation | Prompt robustness (Q1.2.5) | 6.21 | 6.0 | 0.74 | 1.00 | 0.119 |
| 8 | Responsibility | Shared responsibility model (Q2.3.6) | 6.18 | 7.0 | 1.28 | 1.00 | 0.207 |
| 9 | Documentation | Model card (Q2.2.1) | 6.11 | 6.0 | 0.69 | 1.00 | 0.112 |
| 10 | Evaluation | Result reproducibility (Q1.2.6) | 6.11 | 6.0 | 0.99 | 1.00 | 0.163 |
| 42 | Explainability | Hospital administration (Q2.5.7) | 4.46 | 5.0 | 1.73 | 3.00 | 0.388 |
| — | Value alignment | Value balancing (Q3.3.6) | 3.14† | 3.0 | 1.80 | 4.00 | 0.573 |
† Value Balancing (Q3.3.6) was measured on a 6-point forced-ranking scale (1 = highest priority, 6 = lowest priority), unlike other items measured on a 7-point Likert scale. The mean of 3.14 represents the average rank position, not a Likert importance score. Lower values indicate higher perceived priority
SD, standard deviation, AI, artificial intelligence, IQR, interquartile range, CV, coefficient of variation
In contrast, explainability in hospital administrative processing (mean 4.46, IQR 3.0) and value balancing (IQR 4.0, CV 0.573) showed severe disagreement among experts, making them the only items out of 56 classified as having low consensus (bottom of Table 3). The demand for explainability in hospital administrative AI showed a gap of 2.22 points compared with the highest-scoring item (clinical decision support, mean 6.68), demonstrating that the required level of explainability is clearly differentiated according to clinical risk.
In the importance evaluation results for the nine thematic areas (Table 3), documentation standards recorded the highest score with a mean of 6.06, while safety measures (mean 5.95) and evaluation methods (mean 5.86) formed the top tier. This indicates that among the nine rated thematic areas, transparent documentation of the research process and prevention of potential harm to patients received the highest endorsement.
Among the individual items, explainability in clinical decision support demonstrated the highest importance with a mean of 6.68, followed by explainability in diagnostic support (mean 6.57), expert verification (mean 6.29), and documentation of standardized data sources (mean 6.29).
Regarding the governance structure, a hybrid model of an IRB and a central regulatory agency received the highest plurality of support (25.0%), followed by the standalone IRB model (21.4%). However, no single governance model achieved majority endorsement, indicating limited agreement on the optimal oversight structure. The distribution of responses suggests that experts recognized the potential value of combining national standards with institution-level reviews, though the lack of clear consensus underscores the need for further deliberation on governance design. In contrast, distinct disagreements were confirmed regarding the balancing of value alignment (IQR 4.0, CV 0.57) and the explainability of AI for hospital administrative processing (IQR 3.0, CV 0.39).
Round 3: strengthening consensus and deriving detailed strategies
Round 3 was conducted with 27 participating experts on a total of 43 items, including 12 re-evaluation items from Round 2 and 31 operational sub-items derived from Round 1–2 discussions. As a result of strengthening the consensus criterion to IQR ≤ 1.0, 26 out of the 43 items (60.5%) reached consensus. Among these, 20 items (46.5%) achieved strong consensus, simultaneously satisfying both IQR ≤ 1.0 and CV ≤ 0.2 (see Multimedia Appendix 6 for detailed item-level statistics).
Among the 12 items re-evaluated from Round 2, setting decision-making boundaries (mean 6.26, IQR 1.0) and AI output reliability scores (mean 5.85, IQR 0.5) reached consensus even under the strengthened criteria, and were thus confirmed as core safety requirements of the guideline. In contrast, developer responsibility (mean 5.11, IQR 2.0) and hospital administration explainability (mean 4.74, IQR 2.0) failed to reach consensus in Round 3 and were therefore not promoted into the condensed core checklist.
Among the operational sub-items, explainable AI (mean 6.48, IQR 1.0) and human-in-the-loop (mean 6.33, IQR 1.0) recorded the highest scores (Multimedia Appendix 6). The prevention of high persuasion, where human autonomous decision-making is compromised by the excessive persuasiveness of AI, and the mitigation of clinical deskilling were also selected as major management targets under high consensus. Regarding bias management, ensuring the diversity of training data received the highest support with a mean of 6.44, followed by strengthening transparency and explainability (mean 6.22). This indicates that experts rated securing representativeness at the data collection stage and transparently disclosing the process as higher priorities than post hoc algorithmic correction (mean 5.70). While demographic bias (gender, race, etc.) showed high consensus (IQR 0.5), socioeconomic bias (income, insurance type, etc.) and algorithmic bias exhibited a large divergence of opinions among experts, with IQRs ranging from 2.0 to 3.0.
In terms of the attribution of responsibility, institutional responsibility (mean 5.59) and healthcare professional responsibility (mean 5.56) reached consensus, whereas developer responsibility (mean 5.11, IQR 2.0) did not. This indicates that experts converged on assigning primary responsibility to the institution and to healthcare professionals, whereas their assessment of developer responsibility remained divergent.
Expert verification system and derivation of the final guideline
In the survey regarding the appropriate timing for expert intervention, performance evaluation (77.8%) and algorithm design (74.1%) were identified as the most important intervention points. Final approval (59.3%), post-deployment monitoring (59.3%), and the model training stage (44.4%) showed relatively lower response rates. This suggests that beyond the passive role of simply granting final approval to a completed model or monitoring it post-deployment, “ethics by design,” which reflects ethical standards from the planning and design stages of the AI model, is central.
Across Rounds 1–3, the empirical Delphi findings clustered most consistently around clinical safety, transparent documentation, and explainability. Several items, such as explainability requirements for clinical decision support, documentation of data provenance, and safeguards addressing hallucination, showed tighter convergence as rounds progressed and refinement increased. Items that did not meet the strengthened Round 3 dispersion criterion (IQR above the pre-specified threshold) were not promoted to REQ status by empirical consensus alone. When such items were nevertheless assigned REQ status, they were justified under one of the four elevation codes in Multimedia Appendix 3 Table B; otherwise, they were retained as REC, CON, or CTX items in the final guideline package. The final guideline was structured dually, consisting of declarative ethical principles and an actionable checklist (Fig. 2). The checklist aligns with the framework’s three domains (data, governance, and design-by-value) and eight value dimensions. It deploys 53 inspection items across the pre-development (eleven items), development (twenty-two items), post-deployment (twelve items), cross-cutting (two items), application-specific explainability (one item), and value alignment (five items). Furthermore, the generative AI-specific issue checklist includes 23 inspection items across six domains: hallucination, uncertainty quantification, persuasive AI, prompt vulnerability, data source transparency, and agentic AI. Items were promoted into the core checklist as REQ when they both met the consensus-driven REQ rule and functioned as minimum safeguards; high- or strong-consensus items that were secondary or context-dependent were retained as REC or CON. A small number of baseline safeguards, agentic-AI safeguards, oversight requirements, and top-ranked alignment priorities were retained as REQ despite lower raw convergence; these cases are flagged explicitly in the audit mapping using E-A through E-D elevation codes. Lower-convergence items were otherwise retained as REC, CON, or CTX items, and low-consensus items were EXC from the checklist as general requirements. Excluded items included high-intensity regulation of hospital administrative AI, developer-led attribution of responsibility, and the prioritized resolution of socioeconomic bias. For ease of practical use, Table 4 regroups selected REQ items by stage and operational focus rather than reproducing the full eight-value structure; the full checklist, full audit mapping, and application examples by clinical specialty are provided in Multimedia Appendix 3 Tables A and B.
Fig. 2.
Ethical framework for healthcare generative AI research ethics
Table 4.
Condensed core checklist by development stage (selected required items)
| Stage | Operational focus | Checklist Item |
|---|---|---|
| Pre-development | Privacy | Establish strict data access protocols |
| Privacy | Define contractual liability for third-party data sharing | |
| Safety | Clarify and restrict AI use categories | |
| Forward-looking responsibility | Document all data sources | |
| Backward-looking responsibility | Review and agree on shared liability protocols | |
| Debiasing | Assess training data representativeness | |
| During development | Trust | Conduct expert validation of generated outputs |
| Trust | Test prompt robustness (adversarial prompts, jailbreak) | |
| Safety | Set decision boundaries and user approval procedures | |
| Safety | Document AI system design explicitly | |
| Safety | Establish uncertainty quantification standards | |
| Forward-looking responsibility | Ensure responsible R&D conduct with role-based accountability | |
| Debiasing | Review labeling processes for bias | |
| Debiasing | Conduct data bias auditing and correction | |
| Post-deployment | Forward-looking responsibility | Develop continuous monitoring plan |
| Trust | Establish user feedback collection and analysis process | |
| Backward-looking responsibility | Implement post-deployment continuous monitoring plan | |
| Debiasing | Monitor for real-world bias emergence | |
| All stages | Forward-looking responsibility | Comply with government agency ethical oversight |
| Forward-looking responsibility | Comply with IRB review and oversight |
Figure 2legend. The figure presents the ethical framework organized into three hierarchical levels. The outermost layer represents the three overarching domains: Data (left), Governance (center), and Design-by-Value (right). Within each domain, the constituent value dimensions are shown: Privacy, Trust, and Safety under Data; Forward-looking Responsibility and Backward-looking Responsibility under Governance; and Explainability, Debiasing, and Alignment under Design-by-Value. Arrows indicate the mapping relationships between value dimensions and the research lifecycle stages (pre-development, development, post-deployment) shown in the inner ring. Bidirectional arrows between domains indicate interdependencies (e.g., documentation standards under Governance inform and are informed by Trust under Data). The center of the diagram represents the lifecycle checklist, which integrates items from all eight value dimensions across the three lifecycle stages. The shading of each value dimension reflects the overall consensus strength achieved in the Delphi process (darker shading = stronger consensus).
Discussion
Principal results
This study developed research ethics guidelines for healthcare generative AI through a three-stage modified Delphi method. The principal results are summarized in three main points. First, in Round 2, 97.7% of Likert-scaled items entered the acceptable range for guideline adoption. Among the nine rated thematic areas, documentation standards (mean 6.06), safety measures (mean 5.95), and evaluation methods (mean 5.86) received the highest endorsement. Second, in Round 3, under the strengthened consensus criterion (IQR ≤ 1.0), 60.5% of items reached consensus and 46.5% achieved strong consensus (IQR ≤ 1.0 and CV ≤ 0.2). Explainable AI and human-in-the-loop were derived as core items among the operational sub-items, and ensuring the diversity of training data was agreed upon as the top-priority strategy for bias mitigation. Third, an ethical framework comprising three domains (data, governance, and design-by-value) and eight value dimensions, alongside a lifecycle checklist categorized by development stage, was developed as the primary output.
Comparison with prior work
Compared with prior Delphi studies, the main distinction of this study is its narrower scope and operational output. Stahl et al. [14] reported difficulty reaching consensus when AI was treated as a broad cross-sector concept; by limiting the target to healthcare generative AI R&D, this study achieved a high acceptable consensus rate. Whereas Denecke et al. [2] mapped LLM opportunities and risks, Starke et al. [13] developed a conceptual account of trust, and Mahl et al. [15] showed application-specific disagreement in public health AI, our study produced a checklist directly tied to expert-rated research ethics priorities.
In relation to existing ethical frameworks, this study shares privacy, safety, explainability, and debiasing concerns with Jobin et al. [5], the WHO [22], and Ning et al. [8] (Table 5). Its main structural contribution is threefold: separating responsibility into forward-looking and backward-looking obligations, treating value alignment as an independent construct, and deriving checklist items from Delphi consensus rather than from scoping review alone. Compared with Hendricks-Sturrup et al. [27], it also extends beyond health equity and researcher diversity to cover the full development process.
Table 5.
Comparison of the proposed framework with existing AI ethics frameworks
| Ethical dimension | Jobin et al. [5] | WHO [22] | Ning et al. [8] | This study (N = 32) |
|---|---|---|---|---|
| Domain 1. Data | ||||
| Privacy | Privacy | Data protection & consent | Privacy + Security | Shared: Privacy as technical controls (differential privacy, federated learning) |
| Trust | Trust | (Implicit in data governance) | Trust (user confidence) | Shared: Trust as output verification, prompt robustness, reproducibility |
| Safety | Non-maleficence | Well-being, safety | Non-maleficence | Shared: Safety as decision boundaries, uncertainty quantification, use-category limits |
| Domain 2. Governance | ||||
| Forward-looking responsibility | Responsibility | Responsibility & accountability | Accountability | Unique: governance body designation plus documentation standards (model card, bias audit report) |
| Backward-looking responsibility | (Not separated) | (Not separated) | (Not separated) | Unique: liability attribution plus post-deployment monitoring; developer responsibility did not reach consensus under the strengthened Round 3 criterion (IQR 2.0) |
| Domain 3. Design-by-value | ||||
| Explainability | Transparency | Transparency, explainability, & intelligibility | Transparency (disclosure, documentation) | Shared: explainability in clinical decision support (6.68) and diagnosis (6.57) rated highest across all items |
| Debiasing | Justice & fairness | Inclusiveness & equity | Equity (bias, disparity) | Shared: debiasing, with training data diversity (6.44) prioritized over algorithmic correction (5.70) |
| Alignment | — | — | — | Unique: value alignment with ranked priorities of medical accuracy > patient-centered value > social value > stakeholder integration |
| Principles present in other frameworks but not defined as separate values in this study | ||||
| Autonomy | Freedom & autonomy (34/84) | Protecting human autonomy | Autonomy (self-determination) | Addressed within alignment and emerging-issue consensus on HITL (6.33); grounded in Kantian respect for persons [30] and principlist autonomy [29] |
| Beneficence | Beneficence (41/84) | Promoting human well-being | (Noted as under-discussed) | Addressed implicitly through safety and alignment values |
| Integrity | — | — | Integrity (plagiarism, intellectual property) | Not included; beyond the scope of this study (research ethics) |
| Sustainability | Sustainability (14/84) | Responsive & sustainable AI | — | Not included; beyond the scope of this study (research ethics) |
| Methodological characteristics | ||||
| Approach | Scoping review of gray literature (soft-law) | Expert advisory panel + policy analysis | Scoping review of 193 peer-reviewed articles | 3-round modified Delphi with quantitative consensus criteria (IQR, CV) |
| Scope | General AI, cross-sector | AI for health (all types) | Generative AI for healthcare | Generative AI for healthcare research ethics |
| Actionable output | Principle mapping | 6 guiding principles + governance recommendations | TREGAI reporting checklist | Lifecycle checklist (pre/during/post development) + ethical framework |
Unique = structural contribution of this study not present as a separate construct in the compared frameworks; Shared = principle also present in prior frameworks but operationalized here with Delphi-derived consensus scores (7-point Likert, shown in parentheses where applicable)
AI, artificial intelligence, CV, coefficient of variation, IQR, interquartile range, TREGAI, Transparent Reporting of Ethics for Generative Artificial Intelligence, WHO, World Health Organization
Proposal for a generative AI ethical framework
The framework proposed below is an interpretive synthesis of the Delphi findings rather than a separate empirical result. Unlike existing healthcare AI ethics studies that focus on presenting universal principles such as accountability and fairness [5, 23], the ethical framework of this study takes a differentiated approach by focusing on the probabilistic generative capabilities and agentic nature of generative AI. Beyond traditional passive ethical principles, it established active values such as trust and value alignment [28] as core values, reflecting the possibility of AI intervening as a communication agent between medical staff and patients rather than merely acting as a tool.
A particularly notable aspect of the framework’s structure is the clear distinction between issues exclusively considered in generative AI (reliability, forward-looking responsibility, alignment) and issues requiring re-evaluation in generative AI (backward-looking responsibility, explainability). This distinction helps set priorities for ethical review. To clarify the research scope, this study defined broader socio-structural issues—such as the replacement of medical personnel, deskilling, Big Tech monopolies, and cost/sustainability—as problems beyond the individual researcher’s control, strictly limiting the scope to actionable guidelines researchers must comply with during the R&D stage. Furthermore, general big data ethics issues such as data collection, consent, and storage, problems of paper fabrication using AI, patient disempowerment, and sustainability were intentionally excluded, as they were deemed matters to be addressed within the realms for universal ethics and traditional research ethics.
Situating the framework within bioethics traditions
The following analysis shifts from the empirical Delphi findings reported above to a normative reading that situates the framework within established ethical traditions. The three-domain framework developed in this study is broadly consonant with established traditions in biomedical ethics and moral philosophy. The principlist framework of Beauchamp and Childress [29] provides a useful orienting structure, while Kantian respect for persons [30] and O’Neill’s account of autonomy and trust in institutional contexts [31] offer complementary normative resources for interpreting the framework’s emphasis on agency, accountability, and institutional trust. Trust in AI-mediated healthcare cannot rest on abstract principles alone but must be sustained through transparent and verifiable practices.
The three domains map onto this principlist architecture in a differentiated form. The data domain—encompassing consent, training data diversity, and documentation standards—engages autonomy, beneficence, and justice, operationalizing the obligation to respect participant agency while ensuring that data practices serve patient welfare and do not systematically disadvantage underrepresented groups. The governance domain—covering human-in-the-loop mechanisms, accountability structures, and post-deployment monitoring—reflects the principles of non-maleficence and justice, translating the duty to prevent harm and distribute technological risk equitably into concrete oversight requirements. Autonomy and participatory justice are engaged through the design-by-value domain, which requires that affected stakeholders have meaningful opportunities to shape how AI systems are designed, evaluated, and constrained.
However, the framework departs from classical principlist formulations in three respects. First, privacy and trust cannot be adequately captured by the conventional medical concept of confidentiality. The synthetic generation of patient-like data, hallucination risks, and the opacity of large language models introduce structural vulnerabilities that no one-time disclosure or consent transaction can address. As O’Neill argues, trust in institutional contexts requires not merely the absence of deception but ongoing transparency and verifiable accountability [31]. In this framework, trust is accordingly treated as a relational and dynamic property—built through documentation standards, accuracy verification, and expert oversight—rather than a static attribute of data handling.
Second, the framework’s distinction between forward-looking and backward-looking responsibility cannot be fully accommodated by principlist analysis alone. Forward-looking responsibility—the obligation to anticipate, document, and mitigate harms before deployment—is grounded in the Kantian duty to act on universalizable maxims and to foresee the consequences of one’s choices [30]. Backward-looking responsibility—accountability for harms occurring after deployment—is better approached through a contractualist lens: Scanlon’s account of principles that no affected party could reasonably reject [32] provides a normative basis for determining when post hoc accountability obligations arise and who bears them, particularly where causal responsibility is distributed across developers, institutions, and clinicians.
Third, the design-by-value domain extends the requirement that technical choices remain answerable to the rational agency of affected stakeholders into a more operationally specified form. Van Wynsberghe’s Care-Centered Value-Sensitive Design (CCVSD) framework, which systematically incorporates relational and contextual care values into the design process of healthcare technologies, provides a methodological precedent for this approach [33]. Recent work on patient and public involvement (PPI) in healthcare AI further suggests that involving patients and publics can strengthen trust, legitimacy, and adoption [34], while narrative review evidence indicates that value-sensitive approaches remain under-operationalized in applied health AI practice [35]. Together, these extensions reflect the framework’s position at the intersection of established bioethics traditions and the novel normative demands posed by generative AI in healthcare.
Response strategies to intensified risk factors
Core risk factors derived through the Delphi survey were prominent in three domains: privacy, safety, and bias. For privacy, generative AI poses a risk of re-identifying the training data of LLMs through mechanisms such as training data extraction attacks [36]. In response, this guideline proposed strict data access protocols, risk-based network separation or segmentation planning, and contractual safeguards for third-party data use as core protections, while treating privacy-enhancing technologies such as secure multi-party computation and homomorphic encryption as conditional measures whose necessity depends on the data environment. These measures address the unique re-identification risks of generative AI, going beyond traditional data protection measures.
In the safety domain, hallucination and the deceptive persuasiveness of AI were raised as major concerns. Experts derived uncertainty quantification and the setting of decision-making boundaries as required ethical safeguards, reflecting the “safety by design” principle, which keeps AI outputs within clinically safe boundaries. The high importance scores for explainability in the clinical decision support and diagnostic support domains (means of 6.68 and 6.57, respectively) indicate that experts viewed explainability as essential in areas directly related to patients’ lives [22, 24].
Regarding bias management, a stage-based bias review process covering the pre-development, development, and post-deployment stages was established to prevent the social biases inherent in large-scale training data from leading to inequalities in medical outcomes [8, 37]. Experts rated ensuring the diversity of training data (mean 6.44) more highly than post hoc algorithmic correction (mean 5.70) as a bias mitigation strategy, indicating that the panel prioritized upstream representativeness at the data collection stage over downstream technical correction.
Bifurcation of responsibility structure and governance system
To resolve the role ambiguity between the Bioethics and Safety Act [25] and the Framework Act on AI [7], this study bifurcated the concept of responsibility into forward- and backward-looking responsibility. Forward-looking responsibility refers to managerial and supervisory duties during the research stage; it complements institutional blind spots by specifying the oversight functions (e.g., standardizing documentation, publishing model cards [38]) that IRBs and data review committees should perform. Backward-looking responsibility refers to the liability for compensation in the event of clinical harm caused by AI malfunction, assigning post-deployment monitoring obligations to researchers and developers.
Experts showed higher consensus on institutional responsibility (mean 5.59) and healthcare professional responsibility (mean 5.56) than on developer responsibility (mean 5.11, IQR 2.0). This indicates that the panel converged on assigning primary responsibility to the institution and to healthcare professionals, whereas their assessment of developer responsibility remained divergent. This result contrasts with the uncertainty regarding the locus of responsibility reported in Stahl et al. [14] (i.e., the finding that consensus was not reached on the executing entity even for clearly defined mitigation strategies). The ability to reach a consensus on an institution- and healthcare professional-centered responsibility structure in this study likely reflects the specific application context of healthcare, which already has an established legal and institutional responsibility framework.
Regarding governance structure, as noted above, no single model achieved majority endorsement. While the hybrid model combining an IRB with a central regulatory agency received the highest share of support (25.0%), the distribution across multiple models—including the standalone IRB (21.4%)—reflects limited agreement rather than a clear preference. This finding suggests that governance design for healthcare generative AI remains an open question requiring further empirical investigation and stakeholder deliberation.
Structuring the checklist as a practical tool
The final output of this study was structured dually, comprising declarative ethical principles and an actionable checklist. This dual structure directly responds to the practitioner demand identified in Round 1 and aligns with Stahl et al.‘s [14] finding that domain experts consistently prioritize practical tools—frameworks, guidelines, and toolkits—over abstract regulatory instruments.
Unlike publication-stage reporting instruments, this checklist functions as a prospective, lifecycle self-inspection tool grounded in empirical expert consensus rather than scoping review [8].
From guideline to practice: implementation considerations
The development of a consensus-based guideline does not, by itself, guarantee uptake in research practice. The clearest institutional pathway runs through IRBs. Under South Korea’s Bioethics and Safety Act, IRBs currently review healthcare AI research without explicit criteria for evaluating generative AI-specific risks such as hallucination, prompt vulnerability, or post-deployment monitoring obligations. Recent scoping work argues that IRBs should address safety and accuracy evaluation, bias assessment, autonomy protection, and explainability requirements, supported by amendments to the Bioethics and Safety act [39]; related work on AI governance of medical research databases likewise emphasized data access, accountability, and public-interest justification [40]. The present checklist can help fill this gap by providing evidence-linked review criteria that distinguish directly consensus-supported requirements from analytically elevated procedural safeguards, while its staged structure supports continuous monitoring rather than one-time protocol approval.
A second pathway lies in academic publication standards. Reporting guidelines such as CONSORT [41] and CREDES [17] show that checklists are most influential when journals incorporate them into submission requirements and peer-review expectations, although uptake remains uneven [42]. The present checklist, organized by development stage and linked to expert consensus, could serve a similar function for healthcare generative AI research by requiring authors to document training data provenance, explainability measures, and post-deployment monitoring plans before publication.
From a governance perspective, the Korean regulatory context illustrates both opportunities and challenges. South Korea’s Framework Act on AI [7], which entered into force on January 22, 2026, classifies healthcare AI as high-impact and establishes obligations related to risk assessment, documentation, human oversight, and safety verification. The present guideline can function as sector-specific guidance under this framework by supplying operational criteria that the Act does not fully specify. This architecture is broadly consistent with the EU AI Act [6] and WHO regulatory guidance [43], but the Korean framework is distinctive in its self-verification orientation, in which primary accountability rests with developers rather than prior regulatory approval. A publicly available, consensus-derived checklist is therefore especially useful as a benchmark for developer documentation and institutional review.
The Delphi findings map directly onto these implementation pathways. Items achieving strong Round 3 consensus—including explainable AI, human-in-the-loop, training data diversity, and continuous monitoring—are precisely those that IRBs and journals can most readily codify as review criteria. Analytically elevated REQ items should not be interpreted as having the same empirical strength as items achieving strong Round 3 consensus; their rationale is procedural and normative, with the evidential basis documented in Table 6 and Multimedia Appendix 3 Table B.
Table 6.
Importance evaluation ranking by domain (Round 2, n = 32)
| Rank | Domain | Mean score (± SD) | High-consensus items | Mean IQR |
|---|---|---|---|---|
| 1 | Documentation standards | 6.06 (± 0.86) | 3/5 (60%) | 1.40 |
| 2 | Safety measures | 5.95 (± 0.93) | 4/6 (66.7%) | 1.21 |
| 3 | Evaluation methods | 5.86 (± 1.00) | 5/6 (83.3%) | 1.21 |
| 4 | Explainability requirements | 5.70 (± 1.13) | 4/7 (57.1%) | 1.46 |
| 5 | Data protection | 5.64 (± 1.18) | 6/7 (85.7%) | 1.21 |
| 6 | Monitoring initiative | 5.49 (± 1.35) | 4/6 (66.7%) | 1.29 |
| 7 | Attribution of responsibility | 5.24 (± 1.38) | 2/6 (33.3%) | 1.67 |
The mean IQR represents the average interquartile range of the items within each domain, with lower values indicating higher consensus. The bias management and value-alignment domains used ranking scales and were, therefore, excluded from this table. For detailed item-level statistics, see Multimedia Appendix 5
SD, standard deviation, IQR, interquartile range
Items that did not reach consensus—such as developer-led responsibility attribution and hospital administrative explainability—are retained only as REC or CTX items, preserving flexibility while remaining transparent about the current boundaries of expert agreement.
Realizing the potential of this guideline nonetheless requires deliberate institutional effort. Voluntary adoption by researchers and developers cannot be assumed: prior literature consistently documents the gap between guideline creation and implementation across healthcare domains [39, 42]. Conditions most conducive to adoption include IRB curriculum integration, reference in government guidance under the Framework Act on AI, and journal endorsement. Pilot testing in IRB review processes and iterative refinement remain priorities for future work, as discussed further in the Limitations section.
Limitations and future work
Several limitations should be noted. First, the participant pool expanded after the Round 1 qualitative focus group. This was an intentional feature of the modified Delphi design rather than a post hoc change: Round 1 served as a structured content-generation and conceptual-grounding stage, whereas Rounds 2–3 served as formal consensus-rating rounds. Additional “outer” experts were recruited through academic groups and professional associations to broaden disciplinary representation. Although panel expansion may introduce heterogeneity and reduce apparent convergence compared with a fixed-panel Delphi, we mitigated this risk by using standardized survey instruments derived from Round 1, providing controlled feedback between rounds, and applying pre-specified dispersion-based thresholds. Future work could replicate these findings using a fully closed panel or compare closed- and expanded-panel Delphi designs.
Second, final checklist status assignment was not subjected to a separate consensus vote. Although explicit decision rules and elevation codes distinguish consensus-driven requirements from analytically elevated requirements, this translation step still involves normative judgment and should be tested in implementation settings. Third, the rapid advancement of generative AI may give rise to additional ethical issues not explored in this study. Fourth, field verification of the checklist’s practical applicability and effectiveness has not yet been conducted; pilot testing in real IRB review processes is therefore a priority.
Fifth, the reliance on a South Korean expert panel may introduce cultural and legal contextual influences, including differences between Korea’s self-verification-oriented Framework Act on AI and the EU AI Act’s conformity-assessment model. Nevertheless, core areas of consensus, such as explainability, human-in-the-loop, training data diversity, and documentation standards, reflect technical properties of generative AI that are relevant across jurisdictions. Future cross-cultural Delphi studies using multinational panels are needed to assess international applicability. Sixth, the absence of patient and public representatives may have affected the prioritization of patient-facing concerns; future iterations should incorporate patient perspectives.
Conclusions
As healthcare generative AI adoption accelerates, this study developed an actionable research ethics guideline through expert consensus. Its main contribution is a differentiated ethical framework that reflects the probabilistic generative capabilities and agentic nature of generative AI (Fig. 2; Table 5), while remaining grounded in established bioethics traditions. The framework extends existing AI ethics work by specifying value alignment, distinguishing forward- and backward-looking responsibility, and linking abstract principles to an actionable checklist with a transparent, traceable link from Delphi consensus to checklist elements.
From a policy perspective, this guideline functions as a sub-guideline capable of specifically applying the high-risk AI management standards of the current Framework Act on AI [7] to the healthcare sector. Simultaneously, it provides practical review criteria that IRB members can reference when evaluating generative AI research proposals within the framework of the existing Bioethics and Safety Act [25]. Realizing the potential of this guideline, however, requires active engagement from IRBs, journals, and regulatory bodies to create the institutional conditions for adoption, as well as ongoing pilot testing and iterative refinement. Future work should operate this guideline as a “living document” that is continuously updated in accordance with technological advancements. Subsequent development of specific guidelines reflecting the needs of individual clinical specialties, such as psychiatry and radiology, effectiveness evaluation through use in actual IRB review processes, and cross-cultural verification studies through multinational expert panels are needed.
Supplementary Information
Acknowledgements
We thank all the experts who participated in the Delphi panel, including the members of the Healthcare Artificial Intelligence Ethics Research Group.
Generative artificial intelligence disclosure
Generative artificial intelligence (Google Gemini and Anthropic Claude) was used for the initial English translation and the drafting of figures.
Abbreviations
- AI
artificial intelligence
- CREDES
Conducting and REporting DElphi Studies
- CV
coefficient of variation
- HITL
human-in-the-loop
- IQR
interquartile range
- IRB
institutional review board
- LLM
large language model
- R&D
research and development
- TREGAI
Transparent Reporting of Ethics for Generative Artificial Intelligence
Authors’ contributions
HJC: Study design, data collection, analysis, and manuscript drafting. SJS: Expert consultation and critical revision of the manuscript. JAS: Data analysis, interpretation of results, and critical revision of the manuscript. JYY: Data collection, Delphi survey administration, and critical revision of the manuscript. HWC: Data collection and critical revision of the manuscript. JHK: Study design, expert panel composition, supervision, and manuscript review. All authors reviewed and approved the final manuscript.
Funding
This work was supported by the Division of Healthcare Artificial Intelligence Research, National Institute of Health, Korea Disease Control and Prevention Agency (Project No.: 11-1790399-100091-01).
Data availability
The Delphi survey data used in this study are not publicly available to protect the anonymity of the participants but are available from the corresponding author upon reasonable request. The full text of the guidelines and checklist is included in Multimedia Appendix 3.
Declarations
Ethics approval and consent to participate
This study was approved by the Institutional Review Board of Yonsei University (approval number: 4-2025-0777). Written informed consent was obtained from all expert panel members. All procedures involving human participants were performed in accordance with relevant guidelines and regulations, including the Declaration of Helsinki and its later amendments.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW. Large language models in medicine. Nat Med. 2023;29:1930–40. 10.1038/s41591-023-02448-8. [DOI] [PubMed] [Google Scholar]
- 2.Denecke K, May R, LLMHealthGroup, Rivera Romero O. Potential of large language models in health care: Delphi study. J Med Internet Res. 2024;26:e52399. 10.2196/52399. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Meskó B, Topol EJ. The imperative for regulatory oversight of large language models (or generative AI) in healthcare. npj Digit Med. 2023;6:120. 10.1038/s41746-023-00873-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Ji Z, Lee N, Frieske R, Yu T, Su D, Xu Y, et al. Survey of hallucination in natural language generation. ACM Comput Surv. 2023;55:1–38. 10.1145/3571730. [Google Scholar]
- 5.Jobin A, Ienca M, Vayena E. The global landscape of AI ethics guidelines. Nat Mach Intell. 2019;1:389–99. 10.1038/s42256-019-0088-2. [Google Scholar]
- 6.European Parliament and Council. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Off J Eur Union. 2024:L2024/1689.
- 7.Ministry of Government Legislation, Republic of Korea. Framework Act on the Development of Artificial Intelligence and the Creation of a Foundation for Trust. Act 20676; promulgated January 21, 2025; in force January 22, 2026.
- 8.Ning Y, Teixayavong S, Shang Y, Savulescu J, Nagaraj V, Miao D, et al. Generative artificial intelligence and ethical considerations in health care: a scoping review and ethics checklist. Lancet Digit Health. 2024;6:e848–56. 10.1016/S2589-7500(24)00143-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Executive Office of the President. Executive Order 14110: safe, secure, and trustworthy development and use of artificial intelligence. Fed Reg 75191. (November 1, 2023). https://www.federalregister.gov/documents/2023/11/01/2023-24283/safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence;88.
- 10.Executive Office of the President. Executive Order 14179: removing barriers to American leadership in artificial intelligence. Fed Reg 8741. (January 31, 2025). https://www.federalregister.gov/documents/2025/01/31/2025-02172/removing-barriers-to-american-leadership-in-artificial-intelligence;90.
- 11.Budzyń K, Romańczyk M, Kitala D, Kołodziej P, Bugajski M, Adami HO, et al. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. Lancet Gastroenterol Hepatol. 2025;10:896–903. 10.1016/S2468-1253(25)00133-5. [DOI] [PubMed] [Google Scholar]
- 12.Kumar S, Datta S, Singh V, Datta D, Kumar Singh S, Sharma R. Applications, challenges, and future directions of human-in-the-loop learning. IEEE Access. 2024;12:75735–60. 10.1109/ACCESS.2024.3401547. [Google Scholar]
- 13.Starke G, Gille F, Termine A, Aquino YSJ, Chavarriaga R, Ferrario A, et al. Finding consensus on trust in AI in health care: recommendations from a panel of international experts. J Med Internet Res. 2025;27:e56306. 10.2196/56306. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Stahl BC, Brooks L, Hatzakis T, Santiago N, Wright D. Exploring ethics and human rights in artificial intelligence – a Delphi study. Technol Forecast Soc Change. 2023;191:122502. 10.1016/j.techfore.2023.122502. [Google Scholar]
- 15.Mahl D, Schäfer MS, Voinea SA, Adib K, Duncan B, Salvi C, et al. Responsible artificial intelligence in public health: a Delphi study on risk communication, community engagement and infodemic management. BMJ Glob Health. 2025;10:e018545. 10.1136/bmjgh-2024-018545. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Alon I, Haidar H, Haidar A, Guimón J. The future of artificial intelligence: insights from recent Delphi studies. Futures. 2025;165:103514. 10.1016/j.futures.2024.103514. [Google Scholar]
- 17.Jünger S, Payne SA, Brine J, Radbruch L, Brearley SG. Guidance on Conducting and REporting DElphi Studies (CREDES) in palliative care: recommendations based on a methodological systematic review. Palliat Med. 2017;31:684–706. 10.1177/0269216317690685. [DOI] [PubMed] [Google Scholar]
- 18.Linstone HA, Turoff M, editors. The Delphi Method: Techniques and Applications. Reading, MA: Addison-Wesley; 1975. [Google Scholar]
- 19.Hasson F, Keeney S, McKenna H. Research guidelines for the Delphi survey technique. J Adv Nurs. 2000;32:1008–15. 10.1046/j.1365-2648.2000.t01-1-01567.x. [PubMed] [Google Scholar]
- 20.Diamond IR, Grant RC, Feldman BM, Pencharz PB, Ling SC, Moore AM, et al. Defining consensus: a systematic review recommends methodologic criteria for reporting of Delphi studies. J Clin Epidemiol. 2014;67:401–9. 10.1016/j.jclinepi.2013.12.002. [DOI] [PubMed] [Google Scholar]
- 21.Floridi L, Cowls J, Beltrametti M, Chatila R, Chazerand P, Dignum V, et al. AI4People-An ethical framework for a good AI society: opportunities, risks, principles, and recommendations. Minds Mach (Dordr). 2018;28:689–707. 10.1007/s11023-018-9482-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.World Health Organization. Ethics and Governance of Artificial Intelligence for Health: WHO Guidance. Geneva: WHO; 2021. [Google Scholar]
- 23.Vayena E, Blasimme A, Cohen IG. Machine learning in medicine: addressing ethical challenges. PLOS Med. 2018;15:e1002689. 10.1371/journal.pmed.1002689. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Char DS, Shah NH, Magnus D. Implementing machine learning in health care—addressing ethical challenges. N Engl J Med. 2018;378:981–3. 10.1056/NEJMp1714229. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.National Assembly of The Republic of Korea. Bioeth Saf Act Act No 19887; 2023.
- 26.National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research. The Belmont Report: Ethical Principles and Guidelines for the Protection of Human Subjects of Research. Washington, DC: United States Department of Health, Education and Welfare; 1979. [PubMed] [Google Scholar]
- 27.Hendricks-Sturrup R, Simmons M, Anders S, Aneni K, Wright Clayton E, Coco J, et al. Developing ethics and equity principles, terms, and engagement tools to advance health equity and researcher diversity in AI and machine learning: modified Delphi approach. JMIR Ai. 2023;2:e52888. 10.2196/52888. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Gabriel I. Artificial intelligence, values, and alignment. Minds Mach. 2020;30:411–37. 10.1007/s11023-020-09539-2. [Google Scholar]
- 29.Beauchamp TL, Childress JF. Principles of Biomedical Ethics. 8th ed. New York: Oxford University Press; 2019.
- 30.Kant I. In: Gregor M, Timmermann J, editors. Groundwork of the metaphysics of morals. 2nd ed. Cambridge: Cambridge University Press; 2012. 10.1017/cbo9780511919978.
- 31.O’Neill O. Autonomy and trust in bioethics. Cambridge: Cambridge University Press; 2002. 10.1017/s095382080800304x.
- 32.Scanlon TM. What we owe to each other. Cambridge: Harvard University Press; 1998. 10.2307/j.ctv134vmrn.
- 33.van Wynsberghe A. Designing robots for care: Care centered value-sensitive design. Sci Eng Ethics. 2013;19(2):407–33. 10.1007/s11948-011-9343-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Banerjee S, Alsop P, Jones L, Cardinal RN. Patient and public involvement to build trust in artificial intelligence: a framework, tools, and case studies. Patterns (N Y). 2022;3(6):100506. 10.1016/j.patter.2022.100506. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Long Y, Novak L, Walsh CG. Searching for Value Sensitive Design in Applied Health AI: a narrative review. Yearb Med Inf. 2024;33(1):75–82. 10.1055/s-0044-1800723. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Carlini N, Tramer F, Wallace E, Jagielski M, Herbert-Voss A, Lee K et al. Extracting training data from large language models. In: 30th USENIX Security Symposium (USENIX Security 21); 2021.
- 37.Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366:447–53. 10.1126/science.aax2342. [DOI] [PubMed] [Google Scholar]
- 38.Mitchell M, Wu S, Zaldivar A, Barnes P, Vasserman L, Hutchinson B et al. Model cards for model reporting. In: Proceedings of the Conference on Fairness, Accountability, and Transparency. New York, NY, USA: ACM; 2019. pp. 220-9. 10.1145/3287560.3287596.
- 39.Kim J. The role of IRBs in healthcare generative AI research and directions for revising the Bioethics Act, based on a scoping review. Bio Ethics Policy. 2025;9(1):87–119. 10.23183/konibp.2025.9.1.004. [Google Scholar]
- 40.McKay F, Williams BJ, Prestwich G, Biller-Andorno N, Aro AR, Schicktanz S, et al. Artificial intelligence and medical research databases: ethical review by data access committees. BMC Med Ethics. 2023;24:49. 10.1186/s12910-023-00927-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Schulz KF, Altman DG, Moher D, CONSORT Group. CONSORT 2010 statement: updated guidelines for reporting parallel group randomised trials. BMJ. 2010;340:c332. 10.1136/bmj.c332. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Turner L, Shamseer L, Altman DG, Schulz KF, Moher D. Does use of the CONSORT Statement impact the completeness of reporting of randomised controlled trials published in medical journals? A Cochrane review. Syst Rev. 2012;1:60. 10.1186/2046-4053-1-60. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.World Health Organization. Regulatory Considerations on Artificial Intelligence for Health. Geneva: WHO; 2023. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The Delphi survey data used in this study are not publicly available to protect the anonymity of the participants but are available from the corresponding author upon reasonable request. The full text of the guidelines and checklist is included in Multimedia Appendix 3.


