Abstract
Artificial Intelligence (AI)-powered autonomous systems are increasingly entering healthcare, yet concerns about their reliability, safety, and responsible use present significant barriers to adoption. Building on prior conceptual work, this study introduces a refined and empirically validated framework designed to support the safe and responsible integration of AI in clinical and regulatory contexts. The original framework was developed from semi-structured interviews with 15 experts across clinical, technical, ethical, and regulatory domains, and was subsequently validated through a structured process involving 10 newly recruited participants. Validation combined quantitative ratings and qualitative feedback, yielding consistently high scores for relevance, clarity, and usability, alongside strong endorsement of practical utility. The resulting framework consists of ten dimensions spanning technical, ethical, and operational categories, and is aligned with international standards such as ISO 21448 and the NIST AI Risk Management Framework. By addressing critical issues including data quality, explainability, fairness, and human–AI collaboration, the framework moves beyond abstract principles to provide actionable guidance. It offers clinicians, developers, regulators, and procurement bodies a structured tool to evaluate, monitor, and guide the responsible adoption of autonomous AI systems in healthcare ecosystems.
Keywords: Artificial Intelligence, Autonomous Systems, Trustworthiness, Healthcare, Framework
Subject terms: Health policy, Machine learning
Introduction
Artificial Intelligence (AI) is revolutionizing modern healthcare by enabling increasingly autonomous systems (AS) to support or replace human judgment in clinical, operational, and administrative settings1,2. These systems range from decision-support tools to fully autonomous agents capable of analyzing large volumes of medical data, drawing conclusions, and initiating actions with minimal human intervention3,4. AI is already demonstrating potential in fields such as radiology, pathology, surgery, and chronic disease management, where it enhances diagnostic accuracy, reduces clinician workload, and expands access to care. However, these capabilities also raise complex challenges related to safety, ethics, accountability, and trust5.
Trust is a foundational element for the successful deployment of AI-powered systems in healthcare. It extends beyond technical performance to encompass human-centered concerns such as fairness, transparency, reliability, and ethical alignment6,7. As AI systems increasingly operate in high-stakes, life-critical environments, stakeholders–including clinicians, patients, regulators, and developers–demand confidence not only in the system’s outputs, but also in how those outputs are derived and contextualized within care delivery. These concerns have become central to the emerging discourse on Responsible AI, which emphasizes the need to embed ethical principles, accountability, and human values throughout the AI lifecycle.
Global standards, such as ISO 21448 (Safety of the Intended Functionality, SOTIF) and the NIST AI Risk Management Framework, aim to foster the development of trustworthy AI8,9. However, these frameworks are often abstract and lack actionable guidance suitable for healthcare’s dynamic and interdisciplinary contexts10,11. Trust assessments in AI remain fragmented, particularly in settings where alignment with medical ethics, regulatory requirements, and complex institutional workflows is critical. The emergence of generative AI further complicates this landscape, introducing new concerns related to hallucination, data provenance, explainability, and prompt manipulation12–14.
Although several reports and academic studies have defined trustworthy or responsible AI15,16, they often converge on principles–such as fairness, transparency, robustness, and accountability–without addressing the practical implementation of these concepts in healthcare. Most frameworks are developed from a top-down perspective, emphasizing ethical ideals and technical constraints, yet they lack sufficient engagement with frontline stakeholders. As a result, real-world applications often struggle to operationalize responsible AI in ways that account for both system-level reliability and the lived realities of clinical practice17,18.
The challenges of establishing trust and responsibility in healthcare AI are amplified by the diversity of users and contexts. Clinicians require systems that align with evidence-based guidelines and support shared decision-making. Patients need assurance that AI respects their autonomy and privacy. Regulators demand accountability and transparency, while healthcare institutions seek tools that integrate seamlessly into existing workflows without introducing safety risks or liability concerns19,20. A one-size-fits-all approach is insufficient; instead, a multidimensional and context-sensitive framework is essential.
Despite the growing momentum around responsible AI, few frameworks have been empirically validated or tailored to healthcare-specific environments11,15. Existing models tend to emphasize high-level principles but lack operational tools to guide developers, clinicians, and regulators in evaluating real-world responsibility and trustworthiness. Moreover, there is limited integration between technical performance criteria and sociotechnical factors such as clinician usability, patient autonomy, and institutional readiness.
This study introduces a practical, expert-informed framework for assessing the responsibility and trustworthiness of AI-powered autonomous systems in healthcare. Developed through qualitative interviews with a diverse panel of stakeholders–including clinicians, AI researchers, ethicists, engineers, and regulators21–and validated in the present study through structured expert feedback, the framework identifies ten critical dimensions that collectively define responsible AI in healthcare. These include data integrity, interpretability, clinical relevance, fairness, regulatory compliance, human-AI collaboration, and user experience.
The contributions of this paper are threefold:
A validated assessment framework is presented, grounded in interdisciplinary expertise and aligned with international standards, to operationalize responsible AI in healthcare.
Ten core dimensions of responsibility and trustworthiness are identified, reflecting both technical and human-centered considerations, and offering a holistic lens for evaluation.
Actionable guidance is provided for applying the framework in real-world settings, including procurement, deployment, regulation, and oversight.
The remainder of the paper is structured as follows: Section "Background and related work" provides background and reviews related work on responsible and trustworthy AI frameworks and their limitations in healthcare. Section "Methodology" describes the study methodology and participant selection. Section "Results: A validated framework for responsible and trustworthy healthcare AI systems" presents the results, including the validated responsible AI framework. Section "Discussion" discusses the broader implications of the findings. Section "Practical implications" outlines practical implications for implementation and stakeholder engagement. Section "Limitations" discusses the study’s limitations. Section "Future work" proposes directions for future research. Section "Conclusion" concludes the study with reflections on the framework’s potential to support safer, more ethical, and context-aware AI integration in clinical environments.
Background and related work
The deployment of Artificial Intelligence (AI)-powered autonomous systems (AS) in healthcare marks a paradigm shift in how medical services are delivered, monitored, and managed. As these systems begin to perform tasks once reserved exclusively for trained professionals–such as diagnosing diseases, predicting patient deterioration, recommending treatments, and automating administrative decisions–they raise new questions about performance, safety, and trust. This section provides an in-depth review of the current state of autonomous systems in healthcare, conceptual foundations of trust and responsible AI, healthcare-specific challenges, and existing frameworks, highlighting the critical gaps that this study seeks to address.
The rise of autonomous systems in healthcare
AI-powered autonomous systems have been implemented across a wide spectrum of healthcare applications. In diagnostic imaging, convolutional neural networks (CNNs) have demonstrated expert-level accuracy in identifying pathologies from X-rays, CT scans, and MRIs4,22,23. In surgical settings, robotic systems such as the Da Vinci Surgical System assist in minimally invasive procedures, providing precision and reducing recovery time24. In hospital wards and intensive care units (ICUs), AI systems monitor patient vitals to predict sepsis, cardiac arrest, and other adverse events25,26. Administrative tasks, including triage, billing, and resource scheduling, are increasingly being optimized using AI algorithms27,28.
Despite their promise, these systems are not uniformly adopted. Barriers include concerns about explainability, legal liability, data privacy, and ethical fairness29,30. Clinicians are often reluctant to rely on black-box models that lack intuitive interpretability, while administrators remain cautious due to unclear regulatory pathways. Moreover, disparities in data representation can lead to biased predictions, raising equity concerns31.
Defining trust and responsible AI in healthcare systems
Trust in AI is a multidimensional concept that extends beyond accuracy or performance. Researchers and policy makers have proposed overlapping but distinct definitions of trustworthiness, usually involving dimensions such as safety, robustness, transparency, fairness, explainability, privacy, human agency and accountability6,7,10. In recent years, this discussion has expanded into the broader paradigm of Responsible AI, which integrates these dimensions with additional emphasis on social impact, governance, and accountability throughout the AI lifecycle. Responsible AI highlights not only whether systems can be trusted, but also whether they are designed, deployed, and governed in ways that are safe, ethical, and aligned with human values.
The European Commission’s High-Level Expert Group on AI outlines seven key requirements for trustworthy and responsible AI: human agency and oversight, technical robustness and safety, privacy and data governance, transparency, diversity and fairness, societal and environmental well-being, and accountability32. Similarly, the OECD AI Principles promote inclusive growth, human-centered values, transparency, robustness, and accountability16. NIST’s AI Risk Management Framework provides a flexible, outcomes-based structure to identify and mitigate AI risks33, while ISO 23894 and ISO 42001 offer formal structures for risk management and AI management systems respectively34,35.
Although comprehensive, these standards are often high-level and domain-agnostic. They serve as ethical north stars but leave stakeholders–especially in healthcare–without concrete guidance for real-world implementation.
Healthcare-specific challenges for responsible AI
Healthcare is a high-stakes domain where the consequences of AI failure can be fatal. Trust and responsibility in this context are influenced not only by technical factors but also by social, emotional, and ethical considerations19,36. Clinical environments are shaped by unique features such as interdisciplinary workflows, regulatory scrutiny, patient vulnerability, and the expectation of compassionate care17,37.
One of the primary challenges is ensuring transparency and explainability. Clinicians require justifiable reasoning behind AI recommendations to incorporate them into shared decision-making processes38. Furthermore, patient trust is influenced by perceived autonomy, data control, and ethical alignment. Concerns about surveillance, consent, and the use of personal health data must be proactively addressed39,40.
Bias and fairness are also heightened concerns. Numerous studies have shown that AI models trained on imbalanced datasets can perpetuate health disparities, particularly along racial, gender, and socioeconomic lines31,41. In response, scholars call for frameworks of responsible AI that account for sociotechnical realities and promote inclusivity.
Gaps in current frameworks
Despite growing consensus on the principles of trustworthy and responsible AI, there remains a disconnect between conceptual models and practical implementation, especially in healthcare11,18. Many existing frameworks lack empirical validation and fail to consider the diverse perspectives of key stakeholders. Few studies offer operationalized metrics or context-aware criteria tailored to healthcare workflows, regulatory requirements, and ethical expectations15,42.
Notably, recent frameworks like AMLAS (Assurance of Machine Learning for Autonomous Systems)43 and AIAAIC (Assessment of AI Applications in Clinical Settings)44 have attempted to address these concerns. AMLAS provides a structured approach for safety assurance in machine learning-based systems, primarily in automotive contexts, while AIAAIC offers criteria for the safe implementation of AI in clinical workflows. However, these efforts remain limited in scope and generalizability.
This study responds to this critical gap by developing a healthcare-specific, expert-informed, and validated framework for responsible AI. Unlike prior efforts, this framework is grounded in real-world insights from clinicians, regulators, ethicists, and AI developers. It synthesizes interdisciplinary knowledge into ten actionable dimensions that capture both technical robustness and sociotechnical relevance, offering a path toward operationalizing responsible AI in clinical deployments.
Comparative overview of AI trust and responsible AI frameworks in healthcare
To better understand the limitations of existing standards in addressing the unique demands of healthcare, Table 1 compares widely recognized frameworks across key responsible AI dimensions. This comparative analysis reveals important disparities between high-level policy guidance and the practical realities of deploying AI-powered autonomous systems in clinical settings.
Table 1.
Comparison of Existing AI Trust Frameworks Across Key Dimensions.
| Trust Dimension | NIST RMF | ISO 42001 | ISO 23894 | AMLAS | AIAAIC |
|---|---|---|---|---|---|
| Data Quality & Bias | ![]() |
![]() |
![]() |
— | ![]() |
| Explainability & Interpretability | ![]() |
— | ![]() |
![]() |
![]() |
| Fairness | ![]() |
![]() |
![]() |
— | ![]() |
| Human-AI Collaboration | — | — | — | ![]() |
![]() |
| Clinical Relevance | — | — | — | — | ![]() |
| Usability / UX | — | — | — | — | ![]() |
| Regulatory Compliance | ![]() |
![]() |
![]() |
— | ![]() |
| Real-World Validation | — | — | — | ![]() |
![]() |
| Continuous Monitoring | ![]() |
— | — | ![]() |
![]() |
| Ethical Alignment | ![]() |
![]() |
— | — | ![]() |
As seen in Table 1, most existing frameworks cover general purpose attributes such as fairness, explainability, and risk management. For example, the NIST AI Risk Management Framework (RMF) emphasizes governance and risk mitigation but lacks prescriptive guidance for clinical usability or patient-centered metrics33. ISO 42001 and ISO 23894 focus on organizational processes and technical robustness, yet provide little on sociotechnical interaction or ethical nuance34,35.
AMLAS, though developed for safety assurance in autonomous vehicles, offers promising elements like structured validation and continuous monitoring43, but it does not address patient trust, usability, or clinical applicability. AIAAIC, a more healthcare-focused initiative, addresses these dimensions more directly, particularly in terms of clinical relevance and real-world evaluation44. However, its adoption remains limited and lacks empirical grounding.
This landscape illustrates the need for a unified, operational, and validated framework that bridges technical, ethical, and clinical considerations. The framework presented in this study responds to that need, offering a multidimensional tool developed from real-world stakeholder input and contextualized to the healthcare environment.
Methodology
This study builds upon previously published work21, which proposed an initial conceptual framework for assessing the trustworthiness of AI-powered autonomous systems (AS) in healthcare based on semi-structured expert interviews. In the current study, a multi-phase, qualitative research design grounded in interpretivist paradigms was employed to refine, operationalize, and validate that framework. Given the complex, value-laden, and interdisciplinary nature of trust and responsibility in clinical AI deployment, qualitative methods were particularly well-suited to elicit rich, context-dependent insights from diverse expert stakeholders. The methodology was structured around two core phases: (1) framework co-construction through additional semi-structured expert interviews, and (2) framework validation through structured feedback and thematic evaluation.
By explicitly situating this process within the discourse on Responsible AI, the study ensured that the refined framework addressed not only technical robustness but also ethical, regulatory, and human-centered considerations. This methodological approach therefore links theoretical principles of Responsible AI with practical, stakeholder-informed mechanisms for evaluation in real-world healthcare contexts.
Ethical approval and research oversight
Prior to data collection, ethical approval was obtained from the Institutional Review Board (IRB) at Najran University (Protocol Code: 012979-029337-DS). All participants provided informed consent and were assured of confidentiality, data protection, and the right to withdraw at any point. Interview transcripts and feedback data were anonymized and stored securely in compliance with institutional and GDPR-aligned ethical guidelines. All methods were carried out in accordance with relevant guidelines and regulations, and in compliance with the principles of the Declaration of Helsinki.
Participant recruitment and sampling strategy
In a previous study21, a purposive sampling strategy was employed to ensure diverse representation across disciplines and geographies. Fifteen experts were recruited based on inclusion criteria requiring a minimum of five years of experience in one or more of the following areas: AI system development, healthcare technology implementation, regulatory oversight, clinical decision-making, health informatics, ethical AI governance, or human-computer interaction (HCI).
These participants, drawn from academic, clinical, regulatory, and industry sectors across North America, Europe, Asia, and the Middle East, contributed to the co-construction of the initial framework. In the current study, to ensure an independent evaluation, ten new experts were recruited for the validation phase. These experts represented diverse expertise across clinical, ethical, technical, and regulatory domains (see Table 2). All participants provided informed consent prior to participation, and recruitment was conducted in accordance with the approved protocol, the principles of the Declaration of Helsinki, and relevant ethical guidelines and regulations. This recruitment strategy not only ensured fresh perspectives and avoided overlap with the original cohort, but also reflected the principles of Responsible AI by prioritizing inclusivity, interdisciplinarity, and accountability in the validation process.
Table 2.
Participants in the Follow-Up Validation Phase.
| Expert | Job Title | Experience (Years) | Industry/Domain | Country | Race/Ethnicity |
|---|---|---|---|---|---|
| Expert 1 | Clinical Oncologist | 15 | Healthcare | USA | Asian |
| Expert 2 | Ethical AI Consultant | 12 | AI Ethics | Canada | Caucasian |
| Expert 3 | AI Systems Engineer | 9 | Medical Device Manufacturing | Austria | African |
| Expert 4 | Regulatory Affairs Specialist | 10 | Healthcare Regulations | Canada | Middle Eastern |
| Expert 5 | Radiologist | 8 | Healthcare | India | Indian |
| Expert 6 | Human-Computer Interaction Researcher | 7 | Academia | Germany | Caucasian |
| Expert 7 | AI Product Manager | 7 | AI Technology | UK | Asian |
| Expert 8 | AI Researcher | 10 | Healthcare | USA | Caucasian |
| Expert 9 | Biomedical Data Scientist | 11 | Health Informatics | Saudi Arabia | Middle Eastern |
| Expert 10 | Healthcare Policy Advisor | 14 | Public Health / Regulation | Australia | Caucasian |
Phase I: Semi-structured interviews for framework co-construction
The initial framework used in this study was developed in earlier work21, which involved semi-structured interviews with 15 experts. These interviews elicited expert views on what constitutes trustworthiness in AI-powered autonomous healthcare systems. The interview protocol was informed by key dimensions identified in the literature and international standards such as the NIST AI RMF, ISO 21448 (SOTIF), ISO 23894, and the European Commission’s Ethical Guidelines for Trustworthy AI.
Interview themes included:
Definitions of trust and perceived risk in clinical AI
The relevance and applicability of existing AI standards to healthcare
Real-world barriers to implementation (e.g., explainability, bias, liability)
Suggested attributes for a trustworthiness assessment framework
This body of data formed the foundation for the framework refined and validated in the current study.
Phase II: Framework validation and refinement
A preliminary version of the refined framework was developed based on thematic synthesis of interview data collected in this study and insights from the earlier work21. To assess the framework’s clarity, completeness, and practical utility, a follow-up validation phase was conducted with 10 newly recruited experts who had not participated in the initial framework development. This ensured fresh perspectives and reduced bias, while maintaining diversity across clinical, technical, ethical, and regulatory domains (see Table 2). Each participant was provided with a visual framework diagram and a structured feedback instrument comprising:
Likert-scale ratings (1 to 5) across five dimensions: clarity, completeness, novelty, relevance, and usability
Open-ended prompts (e.g., “What is missing?”, “How would you improve this?”)
The feedback data were analyzed using a hybrid approach combining descriptive statistics for quantitative items and thematic coding for open responses. Framework revisions were made iteratively in response to the findings, ensuring conceptual clarity, contextual fit, and alignment with real-world constraints and needs.
Data analysis approach
A hybrid deductive–inductive thematic analysis was employed. Initial codes were drawn from pre-defined dimensions based on literature (e.g., transparency, safety, accountability), while emergent codes (e.g., workflow fit, emotional trust, user adaptation) were generated through close reading of transcripts. Coding was conducted using NVivo 12 software, and intercoder reliability was ensured through collaborative review and memoing. Final themes were clustered into ten dimensions, which form the basis of the validated trust assessment framework presented in the next section.
Trust dimensions mapped to literature and practice
The final framework integrates both normative dimensions derived from global AI standards and empirical themes grounded in practice. Table 3 illustrates how each framework dimension was informed by existing standards, participant insights, and thematic analysis.
Table 3.
Mapping of Framework Dimensions to Trust and Responsible AI Sources.
| Dimension | Literature/Standard | Expert-Informed Themes |
|---|---|---|
| Data Quality | ISO 23894, NIST RMF | Representation, noise, clinical validity |
| Interpretability | ISO 42001, AIAAIC | Clinician trust, explainability, feedback loops |
| Fairness | EC Ethics Guidelines, NIST RMF | Bias mitigation, equity in outcomes |
| Clinical Relevance | AIAAIC | Outcome alignment, evidence-based design |
| Privacy/Security | GDPR, ISO 23894 | Data control, trust in data flows |
| Human-AI Collaboration | AMLAS, AIAAIC | Shared agency, clinician oversight |
| User Experience | HCI Literature | Usability, workflow integration |
| Regulatory Compliance | FDA, ISO 42001 | Approval readiness, liability boundaries |
| Robustness | SOTIF, NIST RMF | Edge case handling, system failure risks |
| Monitoring/Feedback | AMLAS | Continuous learning, performance decay |
Justification for validation metrics
To ensure methodological rigor in evaluating the proposed framework, five established validation criteria were used: Clarity, Completeness, Relevance, Novelty, and Usability. These dimensions are rooted in interdisciplinary evaluation theories from fields such as health informatics, human-computer interaction (HCI), and implementation science.
Clarity was used to determine whether the framework components were well-defined, logically structured, and presented in a manner accessible to both technical and non-technical audiences. It also assessed whether the language used was unambiguous, the terminology consistent, and the relationships between dimensions clearly articulated. Drawing from principles in cognitive interviewing and survey design validation, this metric ensures that users can easily comprehend and navigate the framework without misinterpretation or confusion45,46.
Completeness evaluated whether the framework comprehensively captured the multidimensional aspects of trust relevant to AI deployment in healthcare. This included assessing whether it addressed technical, ethical, organizational, and experiential dimensions, without omitting key constructs commonly cited in the literature or raised by stakeholders. Informed by construct completeness theory and conceptual modeling in information systems, this metric aimed to prevent under-specification and ensure holistic coverage47,48.
Relevance measured the extent to which the framework addressed real-world concerns faced by stakeholders such as clinicians, developers, procurement officers, and regulators. This included consideration of context-specific challenges (e.g., workflow integration, clinical uncertainty) and alignment with the daily practices and decisions of healthcare actors. The metric reflects evaluation standards from qualitative research and implementation science, emphasizing stakeholder-centeredness and contextual adaptability49,50.
Novelty assessed the contribution of the framework in offering new insights, configurations, or integrations not previously seen in existing trust models. This was particularly focused on how the framework synthesized known dimensions with emergent themes such as emotional trust, shared decision-making, and dynamic adaptability in healthcare AI. It drew from theoretical innovation criteria in organizational studies, evaluating both content originality and the distinctiveness of its structure51.
Usability examined whether the framework could be practically adopted by intended users, including its intuitiveness, ease of navigation, and alignment with institutional processes such as procurement, auditing, and ethical review. This dimension also considered the cognitive load of using the framework and the potential for tool adaptation (e.g., checklists, digital dashboards). Grounded in ISO 9241-11 and usability engineering principles, it ensures that the framework moves beyond theory to support everyday decision-making52,53.
Results: A validated framework for responsible and trustworthy healthcare AI systems
The results of this study present a validated, stakeholder-informed framework for operationalizing responsible and trustworthy AI in healthcare. Drawing on in-depth qualitative interviews with 15 experts and follow-up validation from 10 newly recruited experts, the framework was refined through iterative thematic analysis and structured evaluation. To improve comprehension and usability, the ten dimensions have been grouped into three overarching categories: Technical, Ethical, and Operational, as visualized in Fig. 1.
Fig. 1.
Validated framework dimensions for trust in AI-powered healthcare systems, grouped into technical, ethical, and operational categories.
Framework dimensions and structure
The validation phase confirmed and strengthened the ten dimensions originally identified during the development stage. Ten newly recruited experts consistently affirmed that trust in healthcare AI requires attention across technical, ethical, and operational domains. Their evaluations not only reinforced the relevance of these dimensions but also contextualized them with new practical insights and applications.
Technical Dimensions. This category includes Data Quality, Model Validation, and Robustness. These dimensions underpin both technical robustness and responsible AI assurance. Experts emphasized these as foundational for algorithmic trust. Data quality was cited as essential for fairness and reliability, with one clinical oncologist noting, “We’ve had models trained on data that didn’t reflect our population. The outputs were completely off.” Another participant added, “Data shifts between institutions or populations can completely erode performance–and trust.” Model validation, according to participants, should extend beyond internal testing to include external audits and ongoing scenario-based evaluation. An AI engineer summarized, “Validation isn’t a one-time event–it’s an ongoing responsibility.” Robustness was also highlighted, particularly in high-pressure environments. A systems engineer noted, “Trust collapses the moment the system fails under pressure. Resilience is not optional in clinical AI.”
Ethical Dimensions. This category includes Fairness, Privacy & Security, and Regulatory Compliance. These dimensions capture the ethical and governance pillars of responsible AI, ensuring fairness, privacy, and compliance. Fairness was emphasized as a real and pressing concern. An AI ethics consultant shared, “Bias in healthcare isn’t theoretical–it causes real harm.” Experts called for routine demographic audits. Privacy and security were seen as foundational for trust, particularly in light of regulatory frameworks like GDPR and HIPAA. A regulatory specialist emphasized, “Without clear data governance, we cannot talk about trust.” Finally, regulatory compliance was viewed as a non-negotiable requirement. One policymaker stated, “Compliance is foundational. Without it, you don’t even get in the door.”
Operational Dimensions. These dimensions–Human-AI Collaboration, Clinical Relevance, Explainability, and User Experience–reflect the practical integration of responsible AI into clinical workflows, usability, and human–AI collaboration. Experts emphasized that trust is enhanced when AI augments human decision-making. A radiologist shared, “I want to work with the AI, not against it. There needs to be a sense of partnership.” Clinical relevance was seen as essential to adoption, with one CMO remarking, “We’ve seen great models that never make it past the pilot stage because they don’t fit real workflows.” Explainability was another key factor: “Doctors need to understand AI reasoning, especially in borderline cases,” said one participant. User experience, including emotional trust, was also highlighted. A patient advocate noted, “Trust is emotional as much as logical. People need to feel like the AI is on their side.”
Expert validation and refinement
Ten newly recruited experts participated in a structured validation process to assess the framework’s clarity, completeness, relevance, novelty, and usability. This independent cohort was selected to ensure fresh perspectives and reduce potential bias from participants involved in the original framework development. The validation process included Likert-scale ratings, open-ended feedback, and contextual suggestions. The strong scores across these metrics indicate that the framework not only demonstrates conceptual validity but also meets the practical expectations of Responsible AI principles in healthcare contexts.
Quantitative Ratings. Table 4 presents individual expert scores across the five validation metrics, and Fig. 2 visualizes the mean scores. Clarity received a mean score of 4.6, completeness 4.4, relevance 4.9, novelty 4.1, and usability 4.5. These results affirm that the framework is perceived as conceptually coherent, substantively comprehensive, and highly applicable in real-world settings. Notably, relevance achieved the highest agreement, with nearly all participants assigning it a perfect score of 5.
Table 4.
Expert Survey Ratings and Qualitative Feedback Summary.
| Expert Role | Clarity | Completeness | Relevance | Novelty | Usability | Top Use Case | Feedback Summary |
|---|---|---|---|---|---|---|---|
| Clinical Oncologist | 5 | 5 | 5 | 4 | 5 | Clinical workflow integration | Add real examples |
| Ethical AI Consultant | 4 | 4 | 5 | 5 | 4 | Ethical board review | Include role-specific views |
| AI Systems Engineer | 5 | 4 | 5 | 4 | 5 | Procurement checklist | Visual tools appreciated |
| Regulatory Affairs Specialist | 5 | 5 | 5 | 4 | 4 | Regulatory compliance | Compliance needs highlighting |
| Radiologist | 4 | 4 | 5 | 3 | 4 | Radiology model evaluation | Good foundation |
| HCI Researcher | 5 | 4 | 5 | 4 | 5 | UX improvement | Usability is a strength |
| AI Product Manager | 4 | 4 | 4 | 5 | 4 | AI tool onboarding | Workflow fit is critical |
| AI Researcher | 5 | 5 | 5 | 4 | 5 | Research grant review | Supports multiple roles |
| Biomedical Data Scientist | 5 | 4 | 5 | 4 | 5 | Data pipeline validation | Transparency in preprocessing critical; recommend auditability |
| Healthcare Policy Advisor | 4 | 5 | 5 | 4 | 4 | Public health policy review | Patient data protection central; regulation must be exceeded, not just met |
Fig. 2.
Mean expert ratings across the five framework validation metrics.
Role-Based Interpretations. Experts identified diverse use cases aligned with their roles. Clinical experts (e.g., oncologists, radiologists) emphasized workflow integration and radiology model evaluation. A radiologist noted, “This framework can guide which AI tools actually improve care rather than just adding noise.” Technical stakeholders–such as AI engineers, product managers, and researchers–noted relevance for tool onboarding, grant reviews, and procurement checklists. One AI product manager explained, “It provides a checklist for what we should demand before adopting a tool.” Regulatory and ethical professionals highlighted its applicability for audit preparedness and ethical board reviews. As one regulatory specialist stated, “This framework could serve directly as an audit readiness template.” Across all roles, participants framed the framework as a practical mechanism for implementing Responsible AI, moving beyond abstract principles into role-sensitive decision-making.
Thematic Feedback Insights. Three themes emerged from qualitative feedback:
Role-specific customization. Several experts highlighted the importance of tailoring the framework to different stakeholder needs. As one product manager noted, “A procurement officer and a clinician will not ask the same questions. The framework should allow role-sensitive checklists.”
Real-world examples. Participants emphasized that the framework would gain practical traction if it were accompanied by concrete use cases. A clinical oncologist explained, “If I can see how this was applied to a real-world AI tool, it becomes far easier to trust and use.”
Intuitive visual structure. Experts consistently valued the diagrammatic representation. One HCI researcher observed, “The visual layout makes the concepts stick–without it, the framework risks feeling abstract.”
Framework Refinement Actions. In response to this feedback, the visual diagram was refined (Fig. 1), a stakeholder-focused question matrix was introduced (Table 5), and a future development plan was outlined to include digital tools and dashboards based on the validated framework. These refinements ensure that the framework is not only validated by diverse experts but also positioned as a practical instrument for advancing Responsible AI in healthcare.
Table 5.
Assessment Questions by Stakeholder Role. This table operationalizes the validated framework, providing clinicians, developers, and regulators with role-specific questions that can be directly applied during system design, procurement, or audit processes.
| Dimension | Clinician / Operator Perspective | Developer / Regulatory Perspective |
|---|---|---|
| Technical | ||
| Data Quality | Does the system produce consistent results across patient types and sites? | Are data sources representative, current, documented, and validated end-to-end? |
| Model Validation | Has performance been verified on external datasets reflective of our population? | Is there evidence of external validation, calibration, and post-deployment monitoring? |
| Robustness | Does the tool remain reliable under emergencies, shift changes, or rare cases? | Have stress tests and failure modes (incl. outliers and drift) been specified and mitigated? |
| Ethical | ||
| Fairness | Are outcomes comparable across key demographic groups we serve? | Were fairness metrics selected, evaluated, and reported with mitigation plans? |
| Privacy & Security | Are consent, access control, and audit trails clear for my patients? | Do data handling and model operations comply with GDPR/HIPAA and security best practices? |
| Regulatory Compliance | Does use of this tool align with our institutional policies and clinical governance? | Are documentation and evidence mapped to relevant regulations/standards for audit readiness? |
| Operational | ||
|
Human–AI Collaboration |
Can I easily override, contest, or add notes to the AI’s recommendation? | Are fallback mechanisms, escalation paths, and human-in-the-loop points defined? |
| Clinical Relevance | Does the tool address a real workflow pain point and fit our care pathways? | Has context of use been specified (indications, users, settings) with success criteria? |
| Explainability | Do I understand why this recommendation was made, especially in borderline cases? | Are explanations faithful, traceable, and exposed via APIs/UX for audit and users? |
| User Experience | Is the interface intuitive and low-friction in my daily practice? | Has usability (e.g., ISO 9241-11) been tested with target users, with issues tracked and resolved? |
Discussion
This study contributes to both the theoretical and applied discourse on responsible and trustworthy AI in healthcare by translating abstract ethical principles into a validated, stakeholder-informed framework. Through iterative design and evaluation involving 15 experts in the development phase and 10 newly recruited experts in the validation phase, the framework bridges the gap between normative guidelines and real-world decision-making, offering actionable criteria for assessing and advancing Responsible AI in healthcare.
Theoretical Contributions. The framework extends traditional trust models by integrating underrepresented dimensions such as Human-AI collaboration, emotional trust, and clinical workflow relevance. These additions reflect an expanded understanding of AI not merely as a technical artifact, but as a socio-technical actor embedded in complex healthcare systems. For instance, the inclusion of explainability and interpretability directly addresses the cognitive needs of clinicians, while robustness and monitoring highlight the adaptive nature of AI in live, high-stakes environments. By synthesizing technical, ethical, and operational considerations, the framework reinforces the idea that trust–and by extension Responsible AI–is a multidimensional construct rather than a singular quality.
Contextual Relevance. Unlike generic AI risk management guidelines, this framework is grounded in the unique demands of healthcare. It incorporates domain-specific considerations such as diagnostic uncertainty, liability, and patient autonomy. Feedback from regulatory and clinical experts reinforced the necessity of these elements. As one participant noted, “Compliance is foundational. Without it, you don’t even get in the door.” This insight informed the strong positioning of Regulatory Compliance as a core dimension. More broadly, this grounding in clinical and regulatory contexts distinguishes the framework from generic AI governance principles and ensures its alignment with frontline healthcare realities, thereby embodying the practical implementation of Responsible AI.
Comparison with Existing Models. While global standards such as ISO 23894, the NIST AI RMF, and the EC Ethics Guidelines provide foundational principles, they often lack operational guidance specific to healthcare contexts. This framework distinguishes itself by translating those abstract values into role-sensitive evaluation tools–such as checklists and domain-aligned dimensions. It also builds upon earlier work21 by incorporating structured validation with new participants, real-world use cases, and visual aids to support stakeholder engagement. The alignment of these findings with previous frameworks reinforces well-known trust priorities such as fairness and safety. However, several dimensions–particularly emotional trust, clinical integration, and role-specific usability–emerged strongly in the results and are underrepresented in existing models. These findings suggest a shift toward more contextualized, implementation-ready frameworks tailored to frontline needs and essential for advancing Responsible AI in practice.
Validation of Practical Utility. The strong expert ratings, particularly in relevance and usability, provide empirical validation of the framework’s practical value. These results support prior observations in the literature that emphasize the need for context-sensitive, human-centered evaluation tools, but also advance the field by offering a structured, validated approach. Qualitative feedback confirmed that the framework not only aligns with theoretical priorities but is also perceived as actionable in daily clinical, technical, and regulatory workflows. The refinement actions taken in response to validation feedback–including enhancing the visual framework diagram, introducing a stakeholder-focused question matrix, and outlining digital tool development–demonstrate adaptability and readiness for integration into practice as a Responsible AI tool.
Generative AI Considerations. The emergence of generative AI introduces new challenges for trust and responsibility, including hallucination, prompt manipulation, and traceability of outputs. This framework remains adaptable to such developments by foregrounding explainability, data provenance, and continuous monitoring as dynamic, evaluable dimensions. It encourages healthcare stakeholders to apply the framework not only to rule-based and supervised models but also to emergent systems where decision accountability and contextual alignment are increasingly critical. By doing so, the framework positions itself as future-proof, capable of evolving alongside technological advances and serving as a foundation for Responsible AI in healthcare.
Practical implications
The validated framework provides actionable guidance for a wide range of stakeholders involved in the development, evaluation, and governance of AI-powered healthcare systems. Clinicians can use the framework to assess whether a system aligns with clinical standards, improves workflow integration, and maintains transparency in its decision-making processes. For AI developers, the framework serves as a benchmark for incorporating responsible and trust-centric design principles, including performance monitoring, user-centered explainability, and ethical safeguards, early in the development lifecycle.
Procurement officers and hospital IT committees may apply the framework to evaluate vendor proposals, ensuring that adopted systems are safe, explainable, and compliant with regulatory expectations. Similarly, regulators and policymakers can adapt the framework into auditing protocols, compliance checklists, or certification pathways to support oversight and accountability. In this way, the framework directly operationalizes the principles of Responsible AI into healthcare practice.
Importantly, Table 5 translates the framework into role-specific assessment questions, enabling direct use in procurement audits, clinical product evaluations, and regulatory reviews. For example, clinicians can use it to determine workflow fit and override mechanisms, developers can apply it as a design and testing checklist, and regulators can use it to structure audit and certification processes. This structured orientation ensures that Responsible AI is not only an abstract principle but a practical tool for decision-making across stakeholder groups.
The strong validation ratings, particularly in relevance (4.90) and usability (4.60), underscore the framework’s immediate utility across procurement audits, clinical product evaluations, grant proposal assessments, and medical board or ethics committee reviews. By offering structured, stakeholder-specific utility, the framework bridges the gap between high-level AI principles and operational decision-making in healthcare environments, thereby advancing the responsible adoption of AI-powered autonomous systems.
Limitations
This study has several limitations. First, although the qualitative sample was intentionally diverse, it was limited to 15 experts in the initial framework development phase and 10 newly recruited participants in the validation phase. While qualitative research does not seek statistical generalizability, this relatively small cohort may constrain the range of perspectives represented. Nevertheless, the number of validators is consistent with established norms in human-centered evaluation and usability testing, where 5–10 expert reviews are typically sufficient to identify major issues49,52. Future studies could benefit from engaging larger and more heterogeneous samples across varied healthcare systems, thereby broadening the Responsible AI considerations that inform framework refinement.
Second, the framework has not yet been evaluated through real-time deployment in clinical environments. While expert feedback offers strong early validation, future work should examine how the framework performs when integrated into actual workflows, decision-making processes, and technology procurement scenarios. Embedding the framework into practice will be a critical step toward demonstrating its value as an operational tool for Responsible AI in healthcare.
Third, the study relied on self-reported insights from participants. Although thematic saturation was reached, the results may be subject to recall bias, social desirability effects, or other forms of response bias inherent in qualitative interviews. Addressing these potential biases in future studies through triangulation with observational or quantitative data would strengthen the robustness of validation efforts.
Finally, although the framework was designed to be broadly applicable across healthcare domains, additional value may be gained through domain-specific customization. Future research could explore tailoring the framework to specific clinical specialties–such as oncology, radiology, or primary care–to improve contextual relevance and adoption. Such adaptations would further align the framework with the principles of Responsible AI, ensuring inclusivity, fairness, and contextual sensitivity in diverse healthcare settings.
Future work
Building on this study, several directions are envisioned to enhance the framework’s usability, applicability, and impact. One key avenue involves transforming the validated framework into an interactive digital toolkit, including scorecards and visual dashboards to support decision-making in healthcare institutions. Such tools could streamline evaluations by clinicians, administrators, and developers, while also embedding the principles of Responsible AI into daily practice.
Future research should also explore longitudinal deployment of the framework within real-world clinical environments to assess its influence on trust dynamics, technology adoption, and cross-stakeholder collaboration over time. In parallel, expanding the validation process to include patients and caregivers will help ensure that assessments of Responsible AI reflect not only technical and institutional needs, but also lay perspectives and lived experiences. This inclusion would strengthen the social accountability and human-centered orientation of the framework.
Additionally, future validation should test the framework quantitatively (e.g., inter-rater reliability, test–retest) across larger samples to complement the qualitative findings reported here. Such empirical rigor would reinforce the framework’s standing as a robust tool for operationalizing Responsible AI in healthcare.
Strategic collaboration with regulatory bodies also presents a promising opportunity. By integrating the framework into existing approval workflows, audit protocols, and procurement guidelines, regulatory pilots could further validate its utility while shaping policy standards for Responsible AI in healthcare. This regulatory engagement would help bridge the gap between international principles and enforceable, context-specific governance mechanisms.
Conclusion
This study presents a validated, expert-informed, and context-sensitive framework for assessing trust and responsibility in AI-powered autonomous healthcare systems. Building on prior work with 15 development experts, the framework was independently validated with 10 newly recruited experts across clinical, technical, ethical, and regulatory domains. This dual-phase approach strengthens confidence in its robustness, inclusivity, and applicability across diverse healthcare contexts.
By integrating dimensions such as interpretability, fairness, patient outcomes, and system adaptability, the framework advances efforts to operationalize Responsible AI in healthcare. It supports structured evaluation, fosters a shared understanding of both trust and accountability requirements, and promotes safer, more equitable, and transparent AI integration into clinical workflows.
Future deployment in real-world settings, along with domain-specific adaptations, will be essential to further validate and refine its impact. As AI technologies continue to evolve–particularly in the era of generative systems–this framework offers a practical foundation for Responsible AI, aligning technical integrity with human-centered values and providing actionable guidance for developers, clinicians, regulators, and policymakers alike.
Author contributions
The author confirms sole responsibility for study design, data collection, analysis, and manuscript preparation.
Funding
The author is thankful to the Deanship of Graduate Studies and Scientific Research at Najran University for funding this work, under the Easy Research Funding program grant code NU/EFP/SERC/13/217
Data availability
The dataset used in this study is publicly available at: https://doi.org/10.5281/zenodo.10445881.
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Topol, E. Deep medicine: how artificial intelligence can make healthcare human again. Basic Books (2019).
- 2.Jiang, F. et al. Artificial intelligence in healthcare: past, present and future. Stroke Vasc. Neurol.2, 230–243 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Yu, K.-H., Beam, A. L. & Kohane, I. S. Artificial intelligence in healthcare. Nat. Biomed. Eng.2, 719–731 (2018). [DOI] [PubMed] [Google Scholar]
- 4.Esteva, A. et al. Dermatologist-level classification of skin cancer with deep neural networks. Nature542, 115–118 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.He, J. et al. Practical lessons for the responsible development of ai technologies. Nat. Mach. Intell.1, 12–21 (2019). [Google Scholar]
- 6.Mittelstadt, B. D., Allo, P., Taddeo, M., Wachter, S. & Floridi, L. The ethics of algorithms: Mapping the debate. Big Data Soc.3, (2016).
- 7.Jobin, A., Ienca, M. & Vayena, E. The global landscape of ai ethics guidelines. Nat. Mach. Intell.1, 389–399 (2019). [Google Scholar]
- 8.International Organization for Standardization. Iso 21448: Road vehicles–safety of the intended functionality (2019).
- 9.NIST. Artificial intelligence risk management framework (ai rmf 1.0). Available at: https://www.nist.gov/itl/ai-risk-management-framework (2023).
- 10.Floridi, L. et al. Ai4people-an ethical framework for a good ai society: Opportunities, risks, principles, and recommendations. Minds Mach.28, 689–707 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Kaur, P. et al. Trustworthy ai: A review. ACM Comput. Surv. (CSUR)55, 1–38 (2022). [Google Scholar]
- 12.Bommasani, R. et al. On the opportunities and risks of foundation models. arXiv:2108.07258 (2021).
- 13.Salhab, W. & Ameyed, D. & Jaafar, F & Mcheick, H. A Systematic Literature Review on AI Safety: Identifying Trends, Challenges, and Future Directions. IEEE Access12, 131762–131784 (2024).
- 14.Rao, A. et al. Hallucinations in generative ai: A survey. arXiv:2303.15535 (2023).
- 15.Organization, W. H. Ethics and governance of artificial intelligence for health (2021).
- 16.OECD. Oecd principles on artificial intelligence. Available at: https://www.oecd.org/going-digital/ai/principles/ (2019).
- 17.Rajkomar, A., Dean, J. & Kohane, I. Machine learning in medicine. N. Engl. J. Med.380, 1347–1358 (2019). [DOI] [PubMed] [Google Scholar]
- 18.Asan, O. & Bayrak, Alparslan E. & Choudhury, A. Artificial intelligence and human trust in healthcare: focus on clinicians. Journal of medical Internet research22(6), e15154 JMIR Publications Inc., Toronto, Canada (2020). [DOI] [PMC free article] [PubMed]
- 19.Leveson, N. Engineering a Safer World: Systems Thinking Applied to Safety (MIT Press, 2016). [Google Scholar]
- 20.Alelyani, Turki. Decoding trust in large language models for healthcare in Saudi Arabia. Scientific Reports15(1), 35276 Nature Publishing Group UK London (2025). [DOI] [PMC free article] [PubMed]
- 21.Alelyani, T. Establishing trust in artificial intelligence-driven autonomous healthcare systems: an expert-guided framework. Front. Digit. Health6, 1474692 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Ardila, D. et al. End-to-end lung cancer screening with three-dimensional deep learning on low-dose chest computed tomography. Nat. Med.25, 954–961 (2019). [DOI] [PubMed] [Google Scholar]
- 23.Ozturk, T. et al. Automated detection of covid-19 cases using deep neural networks with x-ray images. Comput. Biol. Med.121, 103792 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Yang, G.-Z. et al. Medical robotics-regulatory, ethical, and legal considerations for increasing levels of autonomy. Sci. Robot.2, eaam8638 (2017). [DOI] [PubMed] [Google Scholar]
- 25.Shickel, B. et al. Deep ehr: A survey of recent advances in deep learning techniques for electronic health record (ehr) analysis. IEEE J. Biomed. Health Inform.22, 1589–1604 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Lamsal, R. & Kumar, TV Vijay. Artificial intelligence and early warning systems. AI and Robotics in Disaster Studies, 13–32 Springer (2020).
- 27.Bates, D. W. et al. Big data in health care: using analytics to identify and manage high-risk and high-cost patients. Health Affairs33, 1123–1131 (2018). [DOI] [PubMed] [Google Scholar]
- 28.Obermeyer, Z. & Emanuel, E. J. Predicting the future-big data, machine learning, and clinical medicine. N. Engl. J. Med.375, 1216–1219 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Holzinger, A. et al. Causability and explainability of ai in medicine. Wiley Interdiscip. Rev.: Data Min. Knowl. Discov.9, e1312 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Miotto, R. & Wang, F & Wang, S. & Jiang, X. & Dudley, J.T. Deep learning for healthcare: review, opportunities and challenges. Briefings in bioinformatics19(6), 1236–1246 Oxford University Press (2018). [DOI] [PMC free article] [PubMed]
- 31.Obermeyer, Z. et al. Dissecting racial bias in an algorithm used to manage the health of populations. Science366, 447–453 (2019). [DOI] [PubMed] [Google Scholar]
- 32.Smuha, N. A. The eu approach to ethics guidelines for trustworthy artificial intelligence. Computer Law Review International20, 97–106 (2019). [Google Scholar]
- 33.National Institute of Standards and Technology (NIST). Artificial intelligence risk management framework (ai rmf 1.0). https://www.nist.gov/itl/ai-risk-management-framework (2023). Accessed April 2025.
- 34.for Standardization, I. O. Iso/iec 23894:2023–artificial intelligence risk management (2023).
- 35.for Standardization, I. O. Iso/iec 42001: Artificial intelligence management system (2023).
- 36.He, J. et al. Human-centered artificial intelligence: Reliable, safe & trustworthy. arXiv:2006.11458 (2020).
- 37.Berner, E. S. Clinical decision support systems. 233, Springer (2007).
- 38.Holzinger, A. et al. Explainable ai and multi-modal causability in medicine. Nat. Mach. Intell.1, 418–426 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Morley, J. et al. The ethics of AI in health care: a mapping review. Social science \& medicine260, 113172 (2020). [DOI] [PubMed]
- 40.Alelyani, T. Decoding trust in large language models for healthcare in Saudi Arabia. Scientific Reports15(1), 35276 Nature Publishing Group UK London (2025). [DOI] [PMC free article] [PubMed]
- 41.Mittermaier, M. & Raza, M.M. & Kvedar, J.C. Bias in AI-based models for medical applications: challenges and mitigation strategies. NPJ Digital Medicine6(1), 113 Nature Publishing Group UK London (2023). [DOI] [PMC free article] [PubMed]
- 42.Madiega, T. Regulating artificial intelligence in the european union. Tech. Rep. PE 698.792, European Parliamentary Research Service (EPRS), Brussels (2022). Accessed April 2025.
- 43.Hawkins, R. et al. Assurance of machine learning for autonomous systems (amlas) (White Paper, University of York, 2021). [Google Scholar]
- 44.Consortium, A. Assessment of ai applications in clinical settings (aiaaic). Available at: https://www.aiaaic.org (2023).
- 45.Willis, G. B. Cognitive interviewing: A tool for improving questionnaire design (SAGE Publications, 2005).
- 46.Davis, B. E. Instrument review: Getting the most from your panel of experts. Appl. Nurs. Res.5, 194–197 (1992). [Google Scholar]
- 47.Gregor, S. The nature of theory in information systems. MIS quarterly30, 611–642 (2006). [Google Scholar]
- 48.MacKenzie, S. B., Podsakoff, P. M. & Podsakoff, N. P. Construct measurement and validation procedures in mis and behavioral research: Integrating new and existing techniques. MIS quarterly35, 293–334 (2011). [Google Scholar]
- 49.Patton, M. Q. Qualitative Research and Evaluation Methods: Integrating Theory and Practice 4th edn. (SAGE Publications, 2015). [Google Scholar]
- 50.Fixsen, D. L., Naoom, S. F., Blase, K. A., Friedman, R. M. & Wallace, F. Implementation research: A synthesis of the literature. Tampa, FL:University of South Florida, Louis de la Parte Florida Mental Health Institute, The National Implementation Research Network (FMHI Publication)231, 1–119 (2005). [Google Scholar]
- 51.Corley, K. G. & Gioia, D. A. From the editors: Publishing in amj-part 7: What’s different about qualitative research?. Acad. Manag. J.54, 432–435 (2011). [Google Scholar]
- 52.Nielsen, J. Usability Engineering (Morgan Kaufmann, 1994). [Google Scholar]
- 53.International Organization for Standardization. Iso 9241-11:2018 - ergonomics of human-system interaction – part 11: Usability: Definitions and concepts (2018). Available at: https://www.iso.org/standard/63500.html.
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The dataset used in this study is publicly available at: https://doi.org/10.5281/zenodo.10445881.






























