Abstract
Objectives
To identify the lowest sensitivity and specificity that physicians and the general population consider acceptable for medical artificial intelligence (AI), relative to current human performance.
Methods
In a nationwide, cross-sectional survey in Sweden, 2025, random samples of 500 physicians and 500 adults from the general population were mailed a questionnaire presenting three vignettes (chest pain triage, sore throat triage, ECG myocardial infarction detection) with the corresponding human performance. Participants reported the maximum number of cases an AI should be allowed to miss or over-refer.
Results
Response rates were 45% among physicians and 31% in the general population. Both groups demanded higher AI accuracy than the human benchmark for all cases. In the chest pain triage vignette, the nurse correctly referred 84 of 100 true emergencies; physicians required the AI to correctly refer 11 additional patients (95% sensitivity) and the general population demanded referral of 16 additional patients (100% sensitivity) (p<0.001 for both groups). Among 100 patients not requiring referral, the nurse would mistakenly refer 66. Both groups required the AI to reduce unnecessary referrals by 16 (50% specificity) (p<0.001). A similar pattern was observed in the other vignettes.
Discussion
The accuracy thresholds required by the respondents exceed the performance of many existing systems, although emerging AI research shows promise in narrowing the gap.
Conclusion
Physicians and the general population require medical AI systems to outperform human clinicians. When implementing AI in healthcare settings, early engagement with both groups may be necessary to align expectations with real-world system performance.
Keywords: Artificial intelligence; Decision Support Systems, Clinical
WHAT IS ALREADY KNOWN ON THIS TOPIC
Existing research shows growing artificial intelligence (AI) use in healthcare, but lacks data on user expectations for AI performance.
WHAT THIS STUDY ADDS
This research demonstrates that both physicians and the general population demand AI to outperform human clinicians, prioritising high sensitivity.
It also reveals polarised views on specificity and moderate levels of trust in AI.
HOW THIS STUDY MIGHT AFFECT RESEARCH, PRACTICE OR POLICY
These findings emphasise the need for transparent communication about AI capabilities, early stakeholder engagement in development and careful consideration of sensitivity-specificity trade-offs when implementing AI in healthcare.
Introduction
As artificial intelligence (AI) tools enter clinics and patient smartphones at unprecedented speed, a fundamental question emerges: what level of performance is ‘good enough’ for AI to guide medical decisions? Early applications—such as rules-based expert systems for ECG interpretation—have been in use for decades, but the emergence of large language models has considerably broadened the scope and accessibility of AI tools.1
Adoption is already outpacing regulation: self-reported use of non-medical-grade generative AI rose among UK general practitioners from 20% in 2024 to 25% in 2025,2 3 while 17% of the US general population reported using such tools for health-related queries at least monthly.4 This rapid uptake contrasts with the absence of proper validation and consensus on performance thresholds for safe and trusted deployment.
Trust is central to adoption in healthcare.5 6 For both patients and physicians, performance—typically expressed as sensitivity and specificity—is a cornerstone of trust.7,9 AI classification systems require balancing sensitivity and specificity. Prioritising sensitivity risks increasing false positives, leading to unnecessary examinations, anxiety and resource strain,10,12 whereas prioritising specificity raises the risk of false negatives, missed diagnosis and patient harm.13 Currently, performance targets are largely set by developers, and it is often not known to what extent the preferences of end-users are taken into account.
Importantly, it is not self-evident that all stakeholders must demand AI to outperform humans; even modest accuracy can add value by easing workload, saving resources and expanding access to care.14
To date, studies exploring the use of AI and opinions about AI use in healthcare have been limited in scope, often relying on convenience samples from prerecruited online panels, raising concerns about representativeness.15 16 To our knowledge, none have inquired about minimum acceptable performance levels for AI in medicine using nationally representative random samples.
This study aimed to assess and compare the views of physicians and the general population on the minimum acceptable sensitivity and specificity for medical AI systems across different clinical vignettes.
Objectives
The primary outcomes were physicians’ and the general population’s acceptable sensitivity levels for AI (median), compared with estimated human performance.
Secondary outcomes included the same measure but for specificity, as well as reported use of AI for medical purposes and trust in medical advice provided by AI.
Methods
Study design and setting
We conducted a nationwide, cross-sectional survey in Sweden between February and May 2025. Eligible physicians and adult residents were invited by postal mail to complete a questionnaire on AI use, trust and the minimum sensitivity and specificity they would consider acceptable for AI systems in three clinical vignettes. We used the STROBE (Strengthening the Reporting of Observational Studies in Epidemiology) reporting checklist17 when editing.
Study population and sampling frame
For physicians, the sampling frame was the National Board of Health and Welfare register, which by law is required to include all licensed physicians practicing in Sweden. Eligible physicians were those with a valid personal identification number, a Swedish address and under 70 years of age. For the general population, the sampling frame was the Swedish population registry, which covers an estimated 98–99% of residents; individuals aged 18 years or older were eligible. From each frame, 500 individuals were selected using simple random sampling without stratification.
Survey design
Consultation with experts in medical ethics, digitalisation and clinical AI informed the survey design. From several suggested themes, we focused on two: (1) accuracy requirements and (2) AI use and trust.
Clinical vignettes were developed to test expectations for AI sensitivity and specificity across two dimensions: high-stakes versus low-stakes cases, and autonomous operation versus decision support.
Questionnaire items were piloted and refined through cognitive interviews with physicians and members of the public to ensure clarity.18 19
Accuracy requirements for medical AI were measured as sensitivity (the proportion of patients who need referral and are referred) and specificity (the proportion of patients who do not need referral and are not referred). Cognitive interviews showed that respondents struggled to provide a numeric sensitivity or specificity when no reference was provided. Without any reference frame, many respondents resorted to ‘round numbers’ (50%, 100%) or left the item blank. Therefore, in the final version, respondents were shown an estimated human performance for comparison, and then asked how many cases an AI should be allowed to miss or refer unnecessarily (see online supplemental file 1 for details on the sources used to define human performance benchmarks). For example, in one vignette, patients with chest pain were considered for referral to emergency care. The questions were phrased as “the nurse misses 12 of 100 who need emergency care, how many should an AI be allowed to miss?”. Responses were converted to sensitivity and specificity by subtracting the numbers from 100. The three vignettes were:
A nurse deciding by phone whether a patient with chest pain needs emergency care.
A nurse deciding by phone whether a patient with a sore throat needs a doctor’s appointment.
A doctor detecting signs of myocardial infarction on an ECG.
Additional items included personal experiences with AI, such as medical use, and trust in AI systems. A detailed list of variables examined is found in online supplemental table 1.
Data collection
Separate surveys were sent to physicians (online supplemental file 1) and members of the general population (online supplemental file 2), with several overlapping questions. Surveys—including study details—were mailed with options to reply by postal mail or online. Two reminders were issued, the second offering a lottery-ticket incentive.
To protect privacy, the survey was anonymous. Postal respondents could optionally register their participation non-anonymously to stop reminders. Online access required a personal code, but it was not linked to responses, and no demographic questions were collected in the survey itself. Consequently, demographic analyses were only possible for online respondents and for postal respondents who explicitly registered their participation.
Sample size calculation
The study was originally powered to estimate the mean absolute sensitivity with a 95% CI width of 2%, which required 190 participants. After pilot testing, the questionnaire was refined, and the primary endpoint was changed to a median comparison of respondents’ required AI performance against a fixed human-performance threshold. Detecting that at least 60% of participants would place AI above the human threshold required 199 participants (α=0.05, 80% power). This is essentially the same magnitude as the initial target, so the planned invitation of 500 individuals (anticipating ≥40% response, based on comparable studies20 21) remained adequate for the revised primary analysis.
Analytical methods
Age and sex were obtained from the national registers for all respondents who opted to declare their participation, as were the medical specialties of the physicians. The median age was compared between each sample and its corresponding sampling frame using a one-sample Wilcoxon signed-rank test. A binomial test was used to compare sex proportions in the samples with the true proportions in the sampling frames.
We report medians with IQRs for AI versus human sensitivity and specificity. Sign tests evaluated differences to human benchmarks; Mann-Whitney U and Fisher’s exact tests compared groups on continuous and categorical outcomes, respectively.
For categorical secondary outcome variables, Wilson CIs without continuity correction were calculated for proportions. Comparisons of proportions of categorical outcomes between groups were tested for significance using the χ2 test for multiple categories and Fisher’s exact test in case of a single proportion. For ordinal variables, the Mann-Whitney U was used to assess significant differences between groups. These analyses were exploratory, and p values for secondary outcomes should be interpreted with caution. For each question, any blank answers were excluded from the analysis.
Result
Response rate and representativeness
Of the 500 physicians invited, 223 responded (response rate 45%). Among these, 171 (77%) submitted the optional, separate, non-anonymous declaration confirming their participation. Among the 500 individuals from the general population, 155 responded (response rate 31%), of whom 93 (60%) provided a declaration of participation (flowchart available in online supplemental figure 1).
Age, sex and medical specialty distribution
The median age was 53 years for the general population sample, modestly higher than the sampling frame median of 49 years (p=0.045). Otherwise, no statistically significant differences in age or sex were found between the samples and their respective sampling frames (table 1, online supplemental figure 2). Likewise, there was no statistically significant difference in the distribution of specialties among the physicians between the sample and the sampling frame (online supplemental figure 3).
Table 1. Summary statistics about the respondents.
| Physicians | General population | |||
|---|---|---|---|---|
| Sample | Sampling frame | Sample | Sampling frame | |
| Total respondents | 223 | — | 155 | — |
| Respondents with known demographics | 171 | — | 93 | — |
| Mean age (±SD) | 47±11 | 48±11 | 53±17 | 51±20 |
| Median age (IQR) | 45 (37–55) | 46 (38–57) | 53 (38–65)* | 49 (34–66) |
| Female proportion (%) | 46 | 51 | 48 | 50 |
Demographics are only known for respondents who declared separately that they had submitted the anonymous survey. Demographic data is based on this portion of respondents.
p<0.05 compared with the sampling frame.
Sensitivity and specificity
Sensitivity
Both physicians and the general population proposed higher median sensitivity requirements for AI than for humans in all three vignettes. In the chest pain triage vignette, the nurse correctly referred 84 of 100 true emergencies; physicians required the AI to correctly refer 11 additional patients (95% sensitivity) and the general population 16 additional patients (100% sensitivity) (p<0.001 for both groups). A similar pattern was observed in the other two cases (table 2, figure 1).
Table 2. Proposed sensitivity requirements for AI compared with human performance.
| Sore throat triage | Chest pain triage | ECG interpretation | ||||
|---|---|---|---|---|---|---|
| Provided human level: 88% | Provided human level: 84% | Provided human level: 80% | ||||
| Physicians | General population | Physicians | General population | Physicians | General population | |
| Median answer (Q1–Q3) (%) | 92 (88–95) | 95 (90–100) | 95 (90–100) | 100 (90–100) | 98 (90–100) | 100 (90–100) |
| Median difference from human level (Q1–Q3) (pp) | 4.0 (0.0–7.0)* | 7.0 (1.8–12)* | 11 (6.0–16)* | 16 (6.0–16)* | 18 (10–20)* | 20 (10–20)* |
| Proportion proposing exactly 100% (%) | 9.6 | 37 | 27 | 50 | 39 | 53 |
| Proportion above human level (%) | 74 | 77 | 87 | 83 | 91 | 86 |
| Proportion equal to human level (%) | 20 | 17 | 12 | 11 | 7.3 | 10 |
p<0.001 compared with the provided human level.
AI, artificial intelligence; pp, percentage points; Q1, first quartile; Q3, third quartile.
Figure 1. Distribution of required AI performance levels compared with human performance. Survey responses from physicians and the general population regarding the required sensitivity (top) and specificity (bottom) of AI systems across three clinical vignettes (sore throat triage, chest pain triage and ECG interpretation for myocardial infarction). Values on the y-axis represent the required performance level expressed as a percentage. Dashed red lines indicate the provided human performance level for each vignette. AI, artificial intelligence.
Across vignettes, 7–20% of respondents in each group suggested the same requirement as for humans, while 74–91% wanted stricter requirements for AI. In all vignettes, a greater proportion of the general population than of physicians demanded 100% sensitivity (p<0.01 for all).
Specificity
In the chest pain vignette, of 100 patients not requiring emergency care, the nurse mistakenly referred 66. Both groups required, in median, that the AI reduce unnecessary referrals by 16 (50% specificity) (p<0.001). A similar pattern was seen for the sore throat vignette. In the ECG interpretation vignette, however, where the doctor mistakenly classified only 1 of 100 non-myocardial infarction ECGs as showing myocardial infarction, both groups set the same requirement for the AI as for humans (99% specificity) (table 3, figure 1).
Table 3. Proposed specificity requirements for AI compared with human performance.
| Sore throat triage | Chest pain triage | ECG myocardial infarction detection | ||||
|---|---|---|---|---|---|---|
| Provided human level: 30% | Provided human level: 34% | Provided human level: 99% | ||||
| Physicians | General population | Physicians | General population | Physicians | General population | |
| Median answer (Q1–Q3) (%) | 50 (30–70) | 50 (30–90) | 50 (34–70) | 50 (34–90) | 99 (95–100) | 99 (98–100) |
| Median difference from human level (Q1–Q3) (pp) | 20 (0.0–40)* | 20 (0.0–60)* | 16 (0.0–36)* | 16 (0.0–56)* | 0.0 (−4.0–1.0) | 0.0 (−1.5–1.0) |
| Proportion proposing exactly 100% (%) | 6.4 | 14 | 4.6 | 15 | 27 | 38 |
| Proportion above human level (%) | 71 | 67 | 63 | 63 | 28 | 38 |
| Proportion equal to human level (%) | 22 | 19 | 20 | 16 | 32 | 31 |
| Proportion proposing exactly 0% (%) | 0.46 | 8.2 | 1.8 | 14 | 0.91 | 0.4 |
p<0.001 compared with the provided human level. Non-significant results tested at α=0.05.
pp, percentage points; Q1, first quartile; Q3, third quartile.
The proportion demanding the same specificity as humans ranged from 16% to 32%, while 63–71% wanted more stringent criteria for AI in the triage cases. Among the general population, more respondents than physicians proposed both 100% and 0% specificity requirements (p<0.05 for all vignettes). Response distributions showed spikes at the human level, at 50% and at 100% in both groups for the triage vignettes, with an additional spike at 0% among the general population.
AI usage
The use of AI chatbots for any purpose was more common among physicians than among the general population: 72% of physicians had tried a chatbot, compared with 53% of the general population, a difference of 20 percentage points (95% CI 5.7 to 33). ChatGPT was the most frequently used chatbot in both groups (online supplemental figure 4 and online supplemental table 2).
Among physicians who had used a chatbot, 33% (95% CI 26% to 41%) reported asking questions related to real patient cases. Of these, 37% (95% CI 25% to 50%) said they had incorporated the chatbot-generated answers into their clinical work. Overall, this corresponds to 8.5% (95% CI 5.5% to 13%) of all physicians using chatbots for clinical decision-making. Common use cases included suggesting differential diagnosis, evaluation and treatment options (online supplemental figure 5).
Additionally, 6.3% (95% CI 3.8% to 10%) of physicians reported using chatbots for administration (eg, drafting referral or patient letters), bringing the total proportion using chatbots for clinical or administrative tasks to 13% (95% CI 9.6% to 19%). Furthermore, 73% (95% CI 66% to 78%) had used some form of AI other than large language models, most commonly ECG interpretation software and speech-to-text transcription. Less frequently reported uses included radiology image interpretation, drafting clinical notes using digital scribes, clinical decision support systems such as the Swedish Alma platform, and automated polyp detection during colonoscopy (details in online supplemental figure 6).
Among the general population, 12% (95% CI 8.0% to 18%) had used chatbots for advice on health concerns for themselves or family members.
Levels of trust in chatbots’ answers to medical questions were similar between physicians and the general population. Most respondents reported moderate trust, with smaller proportions expressing either low or high trust, and none reporting complete trust (figure 2). Physicians’ trust in chatbots was also very similar to their trust in software ECG interpretation.
Figure 2. Comparison of trust in chatbots and ECG interpretation software. Chatbot trust responses include only physicians who had asked chatbots about real patient cases and members of the general population who had asked about their own or family members’ health. ECG interpretation software trust includes only physicians who had received an ECG software result within the past 6 months.
Discussion
Principal findings
Both physicians and the public required, in median, medical AI to achieve higher sensitivity than human clinicians across all vignettes. For high-stakes cases, expectations were 11–20 percentage points above human levels (80–85%), while for sore throat triage, expectations were 4–7 points higher. These results suggest respondents set stricter thresholds in high-stakes than in low-stakes settings. Respondents also expected higher specificity, but with greater variability.
Notably, the public was polarised on specificity in triage: many demanded either 100% specificity (no false positives) or 0% (refer all). These extremes are impractical for gatekeeping systems and may reflect either unawareness that higher specificity typically reduces sensitivity, or divergent attitudes toward risk.
A considerable proportion of both physicians and the general population already reported using AI for medical purposes, underscoring the importance of aligning user requirements with actual system capabilities in future implementation. Respondents demanded high accuracy but reported only moderate trust. Strikingly, physicians expressed similar levels of trust in unvalidated chatbots as in established ECG interpretation software—revealing a gap between desired standards and actual expectations in practice.
Strengths and limitations
Random sampling from national registries provided greater representativeness than convenience panels, and cognitive interviews ensured comprehension. Offering both postal and digital formats likely reduced bias related to digital literacy.
Nevertheless, the response rates, although consistent with recent healthcare survey research,22 23 leave open the possibility of non-response bias. Non-respondents may systematically differ in socioeconomic status or country of birth.24 The anonymous design protected respondents’ privacy but prevented post-hoc weighting or demographic subgroup analyses. As the study was conducted in Sweden, generalisability to other healthcare systems and cultural contexts may be limited.
Providing human reference values likely anchored expectations, limiting the validity of the absolute accuracy levels. Restricting to three vignettes reduced the burden for respondents, but limited generalisability. Future work should examine a broader range of clinical tasks, track whether expectations evolve as AI performance improves, and assess the views of other healthcare professionals.
Comparison with previous studies
Physicians’ clinical chatbot use (13%) was lower than UK (25%) and US (21%) surveys.3 25 Methodological differences may help explain variation in reported use. Prior studies used broader wording of questions and recruitment strategies that relied heavily on digital engagement, while our survey employed probability-based random sampling with both postal and digital response options. The observed differences may also reflect true differences between physicians in Sweden and those in the UK and US.
As in prior work, respondents favoured sensitivity over specificity.26 Consistent with our findings that AI is expected to outperform human clinicians, a randomised, blinded factorial survey showed that a primary care physician’s endorsement of a diagnostic AI tool as outperforming specialist physicians was the strongest predictor of patients choosing AI-based diagnosis.27
We found polarised views on specificity, likely reflecting different personal attitudes toward risk, particularly fear of false negatives versus false positives. Prior work has proposed that medical AI systems could be calibrated along this spectrum in line with individual patient preferences.13 However, reasoning about overdiagnosis and accuracy trade-offs is challenging.28 Taken together, these findings suggest that the choice of sensitivity–specificity thresholds is not a straightforward, one-size-fits-all matter.
The accuracy thresholds required by our respondents exceed the performance of many existing systems—for instance, symptom checkers with sensitivities around 50% for emergency vignettes29 or ECG interpretation tools with sensitivities in the range of 65% for myocardial infarction detection.30 By contrast, recent experimental conversational diagnostic AI systems have reported performance gains that, when roughly translated to sensitivity, exceed physicians by about 10 percentage points on average1 and by up to 60 percentage points in particularly challenging cases, as reported in a preprint.31 These performance levels are broadly consistent with those expressed by our respondents, suggesting that such systems might be well-received by the general population. However, no specificity levels are reported, and these systems remain at the research stage and are not yet deployed in clinical care.
Kostick-Quenet et al argue that trust in AI should be appropriately aligned with its actual capabilities and limitations, to avoid both over-reliance and underutilisation.32 Previous studies suggest that trust in AI is shaped more by confidence in healthcare institutions than by individual AI literacy.33 Institutional endorsement, transparent validation processes and clear governance may therefore be pivotal in aligning user trust with the real capabilities and limitations of AI systems. Recent work also underscores that successful AI adoption requires more than model optimisation and highlights the need to integrate implementation science into the development process.34 35 A hybrid approach has been proposed, in which clinicians are engaged from the outset, including in problem definition, co-design of user interfaces and workflow integration and evaluation of the fully realised tool in clinical practice.
Conclusion
Both physicians and the general population require medical AI to outperform human clinicians, yet these expectations often exceed the capabilities of current systems. Along with the observed polarisation in specificity preferences, this underscores the need for early engagement with clinicians and the general population when implementing AI in healthcare settings. Such dialogue should focus on realistic accuracy thresholds and the trade-offs between false negatives and false positives. Transparent communication and clear documentation of system-tuning decisions may help preserve trust when errors inevitably occur. Future research should explore how best to structure these discussions and extend the scope to a broader range of clinical vignettes.
Supplementary material
Footnotes
Funding: The research was funded by two sources: the Local Research Council in Södra Älvsborg and the Regional Research Council in Västra Götaland region. The Local Research Council provided two grants, the first one (VGFOUSA-P-983748) providing support for research/supervision (postdoctoral/docent level) with 20% work time allocated and the second with a project-related grant (Grant no. VGFOUSA-997111). The Regional Research Council in Västra Götalands region provided a project-related grant (VGFOUREG-995674) (project funding – regional collaborative projects, new application). These funding sources had no role in the study design, data collection, analysis, decision to publish or preparation of the manuscript. The author PN holds a 50% postdoctoral position, funded by the University Healthcare Clinical Research (USVE) programme in the Region of Skåne.
Provenance and peer review: Not commissioned; externally peer reviewed.
Patient consent for publication: Not applicable.
Ethics approval: This study involves human participants and was approved by Swedish Ethical Review Authority. Reference number 2024-02513-01. Participants gave informed consent to participate in the study before taking part.
Data availability statement
Data are available upon reasonable request.
References
- 1.Tu T, Palepu A, Schaekermann M, et al. Towards Conversational Diagnostic AI. arXiv. 2024 doi: 10.48550/arXiv.2401.05654. [DOI] [Google Scholar]
- 2.Blease CR, Locher C, Gaab J, et al. Generative artificial intelligence in primary care: an online survey of UK general practitioners. BMJ Health Care Inform. 2024;31:e101102. doi: 10.1136/bmjhci-2024-101102. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Blease C, Hagström J, Sanchez CG, et al. General practitioners’ adoption of generative artificial intelligence in clinical practice in the UK: An updated online survey. Digit Health. 2025;11:20552076251394287. doi: 10.1177/20552076251394287. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.kffalexm . KFF; 2024. KFF Health Misinformation Tracking Poll: Artificial Intelligence and Health Information.https://www.kff.org/public-opinion/kff-health-misinformation-tracking-poll-artificial-intelligence-and-health-information/ Available. [Google Scholar]
- 5.Birkhäuer J, Gaab J, Kossowsky J, et al. Trust in the health care professional and health outcome: A meta-analysis. PLoS One. 2017;12:e0170988. doi: 10.1371/journal.pone.0170988. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Liu JYW, Sorwar G, Rahman MS, et al. The role of trust and habit in the adoption of mHealth by older adults in Hong Kong: a healthcare technology service acceptance (HTSA) model. BMC Geriatr. 2023;23:73. doi: 10.1186/s12877-023-03779-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Nundy S, Montgomery T, Wachter RM. Promoting Trust Between Patients and Physicians in the Era of Artificial Intelligence. JAMA. 2019;322:497–8. doi: 10.1001/jama.2018.20563. [DOI] [PubMed] [Google Scholar]
- 8.Sandhu S, Lin AL, Brajer N, et al. Integrating a Machine Learning System Into Clinical Workflows: Qualitative Study. J Med Internet Res. 2020;22:e22421. doi: 10.2196/22421. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Tan BTN, Khan MI, Saleh MA, et al. Empowering Healthcare through Precision Medicine: Unveiling the Nexus of Social Factors and Trust. Healthcare (Basel) 2023;11:3177. doi: 10.3390/healthcare11243177. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Madhavan S, Hackshaw A, Hubbell E, et al. Estimating the Burden of False Positives and Implementation Costs From Adding Multiple Single Cancer Tests or a Single Multi-Cancer Test to Standard-Of-Care Screening. Cancer Med. 2025;14:e70776. doi: 10.1002/cam4.70776. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Wiens J, Saria S, Sendak M, et al. Do no harm: a roadmap for responsible machine learning for health care. Nat Med. 2019;25:1337–40. doi: 10.1038/s41591-019-0548-6. [DOI] [PubMed] [Google Scholar]
- 12.Brodersen J, Siersma VD. Long-term psychosocial consequences of false-positive screening mammography. Ann Fam Med. 2013;11:106–15. doi: 10.1370/afm.1466. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Birch J, Creel KA, Jha AK, et al. Clinical decisions using AI must consider patient values. Nat Med. 2022;28:229–32. doi: 10.1038/s41591-021-01624-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Wang Y, Liu C, Hu W, et al. Economic evaluation for medical artificial intelligence: accuracy vs. cost-effectiveness in a diabetic retinopathy screening case. NPJ Digit Med. 2024;7:43. doi: 10.1038/s41746-024-01032-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Sumner J, Wang Y, Tan SY, et al. Perspectives and Experiences With Large Language Models in Health Care: Survey Study. J Med Internet Res. 2025;27:e67383. doi: 10.2196/67383. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Witkowski K, Dougherty RB, Neely SR. Public perceptions of artificial intelligence in healthcare: ethical concerns and opportunities for patient-centered care. BMC Med Ethics. 2024;25:74. doi: 10.1186/s12910-024-01066-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.von Elm E, Altman D, Egger M, et al. The EQUATOR network reporting guideline platform; 2025. The STROBE reporting checklist.https://resources.equator-network.org/reporting-guidelines/strobe/strobe-checklist.docx Available. [Google Scholar]
- 18.Groves RM, Fowler FJ, Couper MP, et al. Survey methodology. 2nd. Wiley (Wiley Series in Survey Methodology); 2009. edn. [Google Scholar]
- 19.Fowler FJ. Survey research methods. 5th. SAGE Publications; 2014. edn. [Google Scholar]
- 20.Alexanderson K, Arrelöv B, Friberg E, et al. Läkares erfarenheter av arbete med sjukskrivning av patienter. 2018 doi: 10.13140/RG.2.2.21827.91682. [DOI]
- 21.Ohlsson-Nevo E, Hiyoshi A, Norén P, et al. The Swedish RAND-36: psychometric characteristics and reference data from the Mid-Swed Health Survey. J Patient Rep Outcomes. 2021;5:66. doi: 10.1186/s41687-021-00331-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Booker QS, Austin JD, Balasubramanian BA. Survey strategies to increase participant response rates in primary care research studies. Fam Pract. 2021;38:699–702. doi: 10.1093/fampra/cmab070. [DOI] [PubMed] [Google Scholar]
- 23.Brtnikova M, Crane LA, Allison MA, et al. A method for achieving high response rates in national surveys of U.S. primary care physicians. PLoS One. 2018;13:e0202755. doi: 10.1371/journal.pone.0202755. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Vo CQ, Samuelsen P-J, Sommerseth HL, et al. Comparing the sociodemographic characteristics of participants and non-participants in the population-based Tromsø Study. BMC Public Health. 2023;23:994. doi: 10.1186/s12889-023-15928-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Solmonovich RL, Kouba I, Lee JY, et al. Physician awareness of, interest in, and current use of artificial intelligence large language model-based virtual assistants. PLoS One. 2025;20:e0320749. doi: 10.1371/journal.pone.0320749. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Busch F, Hoffmann L, Xu L, et al. Multinational Attitudes Toward AI in Health Care and Diagnostics Among Hospital Patients. JAMA Netw Open. 2025;8:e2514452. doi: 10.1001/jamanetworkopen.2025.14452. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Robertson C, Woods A, Bergstrand K, et al. Diverse patients’ attitudes towards Artificial Intelligence (AI) in diagnosis. PLOS Digit Health. 2023;2:e0000237. doi: 10.1371/journal.pdig.0000237. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Nagler RH, Franklin Fowler E, Gollust SE. Women’s Awareness of and Responses to Messages About Breast Cancer Overdiagnosis and Overtreatment: Results From a 2016 National Survey. Med Care. 2017;55:879–85. doi: 10.1097/MLR.0000000000000798. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Schmieding ML, Kopka M, Schmidt K, et al. Triage Accuracy of Symptom Checker Apps: 5-Year Follow-up Evaluation. J Med Internet Res. 2022;24:e31810. doi: 10.2196/31810. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Garvey JL, Zegre-Hemsey J, Gregg R, et al. Electrocardiographic diagnosis of ST segment elevation myocardial infarction: An evaluation of three automated interpretation algorithms. J Electrocardiol. 2016;49:728–32. doi: 10.1016/j.jelectrocard.2016.04.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Nori H, Daswani M, Kelly C, et al. Sequential diagnosis with language models. arXiv. 2025 Preprint.
- 32.Kostick-Quenet K, Lang BH, Smith J, et al. Trust criteria for artificial intelligence in health: normative and epistemic considerations. J Med Ethics. 2024;50:544–51. doi: 10.1136/jme-2023-109338. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Nong P, Platt J. Patients’ Trust in Health Systems to Use Artificial Intelligence. JAMA Netw Open. 2025;8:e2460628. doi: 10.1001/jamanetworkopen.2024.60628. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Li RC, Asch SM, Shah NH. Developing a delivery science for artificial intelligence in healthcare. NPJ Digit Med. 2020;3:107. doi: 10.1038/s41746-020-00318-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Fayaz-Bakhsh A, Tania J, Lutfi SL, et al. What Is Implementation Science: And Why It Matters for Bridging the Artificial Intelligence Innovation-to-Application Gap in Medical Imaging. PET Clin. 2026;21:1–16. doi: 10.1016/j.cpet.2025.09.002. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Data are available upon reasonable request.


