Abstract
Introduction
Artificial intelligence (AI) is reshaping healthcare, enabled by advances in computing, affordable data storage, and the widespread adoption of electronic health records (EHRs). Machine learning (ML), deep learning (DL), and natural language processing (NLP) are increasingly used for disease diagnosis, risk prediction, and treatment planning.
Objective
This systematic review aimed to examine AI applications across clinical domains from 2020 to 2025, assess their diagnostic accuracy and clinical performance relative to standard practice, identify key implementation barriers including regulatory compliance, algorithmic fairness, and transparency challenges, and compare validation practices and methodological quality with earlier systematic reviews.
Methods
This systematic review followed PRISMA 2020 guidelines. We searched five databases (PubMed, IEEE Xplore, Web of Science, Springer, and Semantic Scholar) for studies published from January 2020 to September 2025. We included original clinical AI studies that reported prospective validation and/or external validation.
Results
Twenty studies met the inclusion criteria. Publication volume peaked in 2024 (n = 7, 35.0%). DL approaches were most common (n = 12, 60.0%), with convolutional neural networks (CNNs) frequently applied to medical imaging tasks. By clinical domain, 30.0% of studies focused on radiology (n = 6), 20.0% on oncology (n = 4), and 15.0% on cardiology (n = 3). For imaging-based diagnostic models, the descriptive median performance across individual studies was 0.91 AUC (no formal meta-analysis was conducted due to heterogeneity in study designs, populations, and outcome metrics). The most frequently reported challenges were regulatory compliance (55.0%, n = 11), limited algorithmic transparency (40.0%, n = 8), data quality limitations (35.0%, n = 7), and barriers to clinical integration (30.0%, n = 6).
Conclusions
AI demonstrates strong potential to improve the effectiveness, safety, and quality of healthcare. However, broader clinical adoption remains constrained by regulatory requirements, interpretability gaps, data quality issues, and workflow integration challenges, underscoring the need for stronger validation practices and more implementation-focused research.
Keywords: artificial intelligence, machine learning, deep learning, healthcare, clinical decision support, clinical validation, health services
Introduction
One of the biggest changes in modern medical practices is the application of AI in healthcare. AI refers to computational systems that can perform tasks typically requiring human intelligence, such as pattern recognition, decision-making, and language understanding.1,2 Its key subfields relevant to healthcare include ML, which enables systems to learn from data without being explicitly programmed; DL, which uses multi-layered neural networks to analyze complex data such as medical images; and natural language processing (NLP), which allows computers to interpret clinical text such as physician notes and medical literature.3,4
AI systems are currently used to detect cancerous lesions in mammograms, 5 predict patient deterioration from vital signs in intensive care units, 6 and extract diagnostic information from unstructured clinical notes. 7 AI has progressively evolved from early experimental models between 2020 and 2025 into clinically validated systems that help doctors diagnose, treat, and care for their patients. Advances in computing power, the creation of more sophisticated algorithms, lower data storage costs, and the widespread use of electronic health records (EHRs) have all contributed to this quick progress.2,3,8
The amount of medical data is growing at an accelerated rate, with healthcare data doubling roughly every 73 days and accounting for more than 30% of global data volume. 9 This enormous data flow has brought about both benefits and difficulties for healthcare systems. Traditional analytic methods sometimes cannot manage datasets this size, necessitating the use of more complex computer approaches. ML and DL can offer assistance by spotting patterns in data, bolstering clinical judgments, and predicting patient outcomes.7,10
EHRs have radically altered the collection, organization, and utilization of healthcare data. These digital archives contain extensive patient data, such as demographics, medical history, imaging tests, lab results, prescription records, and clinical documentation. 11 EHRs contain both ordered data, such test results and diagnoses, and unstructured data, like medical notes. This combination gives AI a strong base on which to operate. When healthcare teams combine EHR data with big data technologies and ML, they can forecast potential outcomes and suggest the best course of action. 12
The COVID-19 pandemic from 2020 to 2023 demonstrated a significant turning point for AI in healthcare. It exposed serious flaws in AI while also showcasing its potential. Researchers swiftly created AI models during this time to help with diagnosis, forecast patient outcomes, and manage hospital resources, often utilizing CT scans and chest X-rays. 13 Due to biased datasets and methodological limitations, systematic evaluations revealed that the majority of COVID-19 AI models were not suitable for clinical deployment. 14 The significance of extremely comprehensive validation, open reporting, and cautious application of AI systems in dynamic healthcare contexts was amply demonstrated by our experience.
Recent advances in technology have significantly expanded the possible uses of AI in medicine. Transformer models were first created for natural language processing (NLP), but they have shown promise in clinical text analysis and the creation of medical documentation. 4 Large language models (LLM) like GPT-4 and Med-PaLM have proven to have reasoning skills that clearly approach expert level performance on medical licensing tests.15,16 Multimodal AI systems those that combine multiple types of data within a single analytical framework now integrate clinical language, genomic sequences, medical images, and structured electronic records. For example, a multimodal system might analyze a patient’s chest CT scan alongside their laboratory results and clinical history to achieve more accurate diagnoses than any single data source could provide alone. 17
A standard framework shift in medical AI had been generated by foundation models trained on large-scale datasets, allowing for natural language interaction, multimodal integration, and zeroshot learning (i.e., the ability to perform new clinical tasks without requiring task-specific training examples). 18 Federated learning (FL) a technique that enables multiple healthcare institutions to collaboratively train AI models on their local data without transferring sensitive patient records to a central server has also addressed important privacy concerns.19,20 The translation of AI research into therapeutic applications such as AI-guided intraoperative monitoring that reduces adverse events during surgery 6 and AI-assisted screening systems that improve early cancer detection 5 has grown more rapid as a result of these advancements.
Despite these advances, significant gaps remain. Many healthcare AI studies focus predominantly on algorithm development and retrospective evaluation, with limited evidence regarding real-world implementation, clinical impact, and patient outcomes.21-23 Understanding of factors contributing to effective adoption including workflow integration, organizational readiness, and clinician training has grown but remains incomplete. 24 Additionally, ethical concerns regarding algorithmic bias, accountability, transparency, and equitable access require stronger frameworks and practical mitigation strategies. 25 Earlier systematic reviews covering the 2010–2019 period identified low rates of external validation, poor reporting standards, and limited attention to fairness,26,27 raising the question of whether more recent research has addressed these shortcomings.
To address these gaps, this systematic review pursues a single overarching aim: to evaluate the current state of clinical AI validation and implementation readiness across healthcare domains from 2020 to 2025. Specifically, we examine four interconnected dimensions: (1) the current landscape of AI applications in clinical diagnosis, treatment, and healthcare delivery; (2) the diagnostic accuracy and clinical performance of AI systems compared to standard clinical practice; (3) implementation barriers, including regulatory compliance, algorithmic fairness, and transparency challenges; and (4) changes in validation practices and methodological quality compared to earlier systematic reviews.26,27 Based on 20 rigorously selected studies, the findings provide evidencebased recommendations for clinicians, researchers, administrators, and policymakers seeking to implement AI in healthcare safely and effectively.
Figure 1 presents the conceptual framework guiding this review, organized around four interconnected components. On the input side, five categories of healthcare data serve as the foundation for AI systems: electronic health records (EHRs) containing both structured and unstructured clinical data, 11 medical imaging modalities such as CT, MRI, and X-ray, genomic and biomarker data, wearable and Internet of Medical Things (IoMT) devices generating continuous physiological signals, and clinical notes and published literature. These data sources feed into the AI technologies and methods employed across the included studies, ranging from established approaches such as deep learning, classical ML, and neural networks to emerging paradigms including NLP, large language models, and foundation models. 18 On the output side, the framework maps the clinical application domains where these technologies are deployed, including radiology and diagnostic imaging, oncology and pathology, cardiology and risk prediction, critical care and ICU monitoring, and primary care and clinical support. 2 Finally, the implementation challenges identified across studieS regulatory compliance (55%), algorithmic transparency (40%), data quality (35%), and clinical integration (30%) represent the barriers that determine whether technically successful AI systems achieve real-world clinical adoption. This framework guided our data extraction and synthesis by ensuring that each included study was analyzed not only for its technical approach and clinical performance, but also for the implementation context that influences translational success.
Figure 1.

Conceptual framework: AI in healthcare (2020–2025). This figure illustrates the key components of AI implementation in clinical settings, including data sources (left), AI technologies and methods (top), clinical applications (right), and implementation challenges (bottom). Corner panels show review statistics, key performance metrics, publication trends, and specialty distribution from the 20 included studies
Methods
This systematic review was carried out following the Preferred Reporting Items for Systematic Reviews and Meta Analyses (PRISMA) 2020 guidelines. 28 The complete study selection process is illustrated in Figure 2.
Figure 2.

PRISMA 2020 flow diagram showing the systematic identification, screening, eligibility assessment, and inclusion of studies
Research Questions
This review addresses the following research questions:
1. What are the current applications of AI and ML in clinical treatment, diagnosis, and healthcare delivery from 2020 to 2025?
2. What is the diagnostic accuracy and clinical performance of AI systems compared to standard practice methods and clinician performance?
3. What putting into practice barriers, ethical considerations, and regulatory requirements affect clinical AI deployment?
4. How do recent AI healthcare studies (2020–2025) compare to earlier systematic reviews in terms of validation practices and methodological quality?
Protocol and Registration
The review protocol was in advance developed and documented prior to commencing the systematic search. The protocol established inclusion and search strategies, exclusion criteria, data extraction procedures, and quality assessment methods to minimize systematic error. All protocol were documented transparently.
Search Strategy
A search strategy was developed in consultation with medical librarian to ensure inclusion of relevant literature from clinical medicine, computer science, and health informatics.
Databases: Five academic databases were searched: PubMed, IEEE Xplore Digital Library, Web of Science Core Collection, Springer Link, and Semantic Scholar.
Search Terms: The search combined two concept groups using Boolean operators:
(“artificial intelligence” OR “machine learning” OR “deep learning” OR “neural network” OR “natural language processing” OR “decision tree*” OR “predictive model” OR “foundation model”)
AND
(“healthcare” OR “clinical” OR “diagnosis” OR “clinical decision” OR “treatment” OR “electronic health record” OR “patient care” OR “medical imaging” OR “prognosis”)
Timeframe: January 2020 to September 2025, capturing the period of rapid AI advancement including pandemic-era developments and emergence of foundation models.
Additional Sources: Reference lists of included studies and key journals (Nature Medicine, The Lancet Digital Health, JAMA Network Open, npj Digital Medicine) were manually searched.
Study Selection Process
Initial Search Results
• PubMed: n = 187
• IEEE Xplore: n = 134
• Web of Science: n = 142
• Springer: n = 78
• Semantic Scholar: n = 60
• Manual search: n = 8
Total records identified: n = 601.
Duplicate Removal: After automated and manual duplicate identification, 118 duplicates were removed, leaving 483 unique records for screening.
Title and Abstract Screening: Two independent reviewers assessed all 483 records against inclusion criteria. Records were categorized as “include,” “exclude,” or “uncertain.” Studies marked “uncertain” by either reviewer proceeded to full-text review.
During this phase, 378 records were excluded:
• Not healthcare-related: n = 156
• No AI/ML application: n = 124
• Review/commentary/editorial: n = 98
Records proceeding to full-text review: n = 105.
Full-Text Assessment: Two reviewers independently evaluated 105 full-text articles. Disagreements were resolved through discussion or third-reviewer consultation.
85 articles were excluded:
• No external or prospective validation: n = 34
• Insufficient methodological detail: n = 22
• Inappropriate study design: n = 16
• Non-English language: n = 8
• Duplicate data/retracted: n = 5
Final Inclusion: n = 20 studies.
The final inclusion rate was 19.0% of full-text articles assessed (20/105) and 4.1% of unique records screened (20/483).
Eligibility Criteria
Inclusion Criteria
1. AI, ML, DL, or related computational methods applied to clinical healthcare
2. Clinical diagnosis, treatment, prognosis, risk prediction, or healthcare delivery
3. Peer-reviewed journal article or conference proceeding with full paper
4. English language publication
5. Published January 2020 to September 2025
6. External validation or prospective clinical validation
7. Clear description of algorithms, datasets, and performance metrics operationally defined as explicit reporting of (a) the algorithm type or architecture used (e.g., CNN, random forest, transformer), (b) the source and size of training and testing datasets, and (c) at least one standard quantitative performance metric such as AUC, sensitivity, specificity, accuracy, or F1 score
8. Appropriate statistical analyses with reported confidence intervals or other measures of uncertainty (e.g., p-values, bootstrapped confidence intervals), enabling assessment of the reliability and precision of reported results
Exclusion Criteria
1. Non-peer-reviewed articles, editorials, or gray literature
2. Internal validation only without external testing
3. Insufficient methodological detail defined as studies that did not clearly specify the AI algorithm used, the data sources for model training and evaluation, or any quantitative performance outcomes, thereby preventing adequate quality assessment
4. Preclinical or animal studies only
5. Duplicate publications or retracted articles
Data Extraction
A standardized extraction form was developed and pilot-tested on 5 studies. Two reviewers independently extracted:
Study Characteristics: Authors, year, journal, country, study design, setting, sample size, funding sources.
AI Technology: Algorithm type, architecture, input data types, training procedures, frameworks used, interpretability methods.
Clinical Application: Specialty, clinical task, target condition, workflow integration.
Performance: Metrics (AUC, sensitivity, specificity, accuracy), comparator benchmarks, external validation results.
Implementation: Barriers identified, regulatory considerations, adoption factors.
Quality Assessment
Study quality was assessed using a modified Critical Appraisal Skills Programme (CASP) checklist incorporating AI-specific criteria from TRIPOD-AI and CONSORT-AI guidelines.29-31 Key criteria included:
• Clear research objectives and clinical relevance
• Appropriate algorithm selection and architecture description
• Adequate sample size with power calculation
• External validation on independent dataset
• Performance across demographic subgroups
• Comparison to appropriate clinical benchmarks
• Consideration of implementation factors
Results
Study Selection and Characteristics
Following comprehensive systematic screening according to PRISMA 2020 guidelines, twenty studies in Table 1 met all inclusion criteria and were included in the final qualitative synthesis. The systematic search strategy found out a total of 601 potentially relevant records across five major electronic databases. PubMed, serving as the primary biomedical literature database, contributed the largest proportion with 187 records identified through our comprehensive search terms combining AI, ML, and healthcare related terminology. IEEE Xplore Digital Library, representing the computer science and engineering literature where many AI methodological advances are first published, contributed 134 records. Web of Science Core Collection yielded 142 records spanning multidisciplinary journals covering both clinical medicine and computational sciences. Springer Link provided 78 records from its extensive collection of work of medical informatics and healthcare technology journals. Semantic Scholar, leveraging its AI-powered academic search capabilities, contributed an additional 60 records that complemented traditional database searches by identifying relevant preprints and conference proceedings that might be missed by conventional bibliographic databases.
Table 1.
Characteristics of 20 Included Studies (2020–2025). Included Studies Grouped by Publication Year With Study Design, Clinical Specialty, AI Method, Clinical Task, and Validation Strategy
| Ref | Study design | Specialty | AI method | Clinical task | Validation |
|---|---|---|---|---|---|
| 5 | Retrospective cohort | Radiology | DL | Breast cancer screening | External (UK, US) |
| 32 | Diagnostic study | Pathology | CNN, DL | Gleason grading | External (multi-site) |
| 13 | Systematic review | Radiology | DL | COVID-19 detection | Evidence synthesis |
| 33 | Retrospective analysis | Multiple | ML | Dataset shift analysis | Multi-site retrospective |
| 34 | Retrospective analysis | Dermatology | DL | Skin lesion diagnosis | External (multi-site) |
| 31 | Reporting guideline | Multiple | Various | DECIDE-AI guidelines | Expert consensus |
| 35 | Framework/comm | eMntuarltyiple | ML | Ethical ML framework | Conceptual framework |
| 18 | Narrative review | Multiple | Foundation | Medical AI foundation | Literature synthesis |
| 36 | Prospective cohort | Cardiology | DL | CVD risk prediction | Prospective external |
| 37 | Narrative review | Oncology | ML | Cancer biomarkers | Literature synthesis |
| 6 | Randomized clinical trial | Critical Care | ML | Hypotension prediction | Prospective RCT |
| 20 | Systematic review | Multiple | Federated | Privacy-preserving AI | Evidence synthesis |
| 38 | Prospective survey | Radiology | ML | AI user experience | Prospective survey |
| 30 | Systematic review | Multiple | Various | CONSORT-AI adherence | Evidence synthesis |
| 39 | Systematic review | LLM, DL | Various | ChatGPT healthcare | Evidence synthesis |
| 40 | Narrative review | Cardiology | ML | QRISK | Literature synthesis |
| 17 | Comprehensive review | Radiology | ML, DL | Multi-modal diagnosis | Literature synthesis |
| 22 | Systematic review | Multiple | Various | AI impact framework | Evidence synthesis |
| 15 | Empirical evaluation | Primary Care | LLM, DL | Medical QA | Internal benchmark |
| 23 | Systematic review | Multiple | Various | AI capabilities review | Evidence synthesis |
The duplicate removal process ruled out 118 records that appeared across multiple databases, indicating that the inter disciplinary nature of healthcare AI research that results in publication in journals indexed by multiple bibliographic services. Following automated deduplication using reference management software with subsequent manual verification to ensure accuracy, 483 unique records remained for initial eligibility screening. Two independent reviewers conducted title and abstract screening against pre defined eligibility criteria, with disagreements resolved through structured discussion and consultation with a third senior reviewer when consensus could not be achieved. 378 records that patently failed to satisfy the basic inclusion criteria based on information accessible in titles and abstracts were eliminated from this initial screening phase, which was intended to efficiently identify obviously not qualified records while maintaining all poten-tially relevant research for detailed evaluation. At this point, non-peer-reviewed publications like editorials, commentary, and news articles, studies on AI applications outside of clinical healthcare contexts, publications describing algorithm development without any of the clinical validation, and studies published outside of the designated 2020–2025 timeframe were the most frequent causes of exclusion.
The remaining 105 full-text papers completed a thorough eligibility evaluation in accordance with all inclusion and exclusion criteria. In order to evaluate methodological rigor, validation strategies, clinical relevance, and reporting completeness, complete articles needs to be obtained and carefully examined throughout this thorough review step. 85 publications were excluded from the full-text review for particular methodological and scope-related reasons. 28 articles with inadequate methodological detail that hindered a proper quality assessment made up the largest exclusion category. These included studies that did not provide enough specificity in their descriptions of algorithm architecture, training procedures, or validation methodology to assess the quality of the research. Two articles have been formally retracted by journals due to discovered errors or misbehavior, while three articles were found to be duplicate publications claiming overlapping datasets with previously published work. Sixteen articles—including mainly technical algorithm papers without clinical application, simulation investigations without patient data, and review articles without original information were eliminated due to inappropriate study designs that did not meet our review criteria. Due to resource limitations that prevented accurate translation and evaluation, eight non-English language articles were excluded.
The demanding methodological criteria used to guarantee the inclusion of only high-quality, rigorously validated studies appropriate for guiding clinical implementation decisions are reflected in the final inclusion rate of 19.0% of full text articles assessed (20 of 105) and 4.1% of unique records initially screened (20 of 483). By eliminating research with methodology problems that would reduce confidence in reported results, this selective technique significantly improves the validity of combined findings while restricting the number of included studies. The full list of the twenty included research is included in 1, along with precise details about each one, such as publication year, journal location, clinical specialty focus, AI methodology used, particular clinical application addressed, and validation strategy.
The time-based distribution of publications throughout the course of the five-year evaluation period showed distinct patterns that demonstrated the dynamic evolution of healthcare AI research, developments in regulations, and worries about clinical adoption. Twenty percent of the total included sample came from four foundational studies conducted in the first two years of the evalua tion period (2020 and 2021). These early studies not only established critical performance benchmarks proving AI capabilities in clinical contexts, but they also uncovered important methodological and generalizability concerns that would have a substantial impact on later research orientations and evaluation standards. When validated on a cohort of over 25,000 women from healthcare systems in the United States and the United Kingdom, the AI system showed absolute reductions of 5.7% in false positive rates and 9.4% in false negative rates compared to human readers. 5 In 2020, McKinney and colleagues published their seminal study in Nature presenting a DL system for breast cancer screening that achieved diagnostic accuracy surpassing the average performance of board-certified radiologists.
With external validation across geographically and demographically diverse patient populations, this study provided strong evidence for generalizability that had been lacking in previous healthcare AI research, making it one of the first extensive, rigorously validated demonstrations that AI could achieve genuinely expert-level performance in a real-world screening context. The clinical implications were significant: fewer false positives would spare women the psychological anguish, inconvenience, and possible complications of needless biopsies and follow-up procedures, while fewer false negatives would allow for early cancer detection when prognosis is best and treatment is most effective. The development of an automated system for Gleason grading of prostate cancer biopsy specimens was also published by Bulten and colleagues in The Lancet Oncology in 2020. This system achieved agreement with expert pathologists that significantly exceeded the agreement observed between pathologists themselves. 32 AI may be able to lessen the diagnostic variability in cancer grading that has been shown to impact treatment choices and patient outcomes, as seen by the system’s kappa agreement of 0.918 compared to inter-pathologist agreement of 0.819. This result was particularly significant because grading variability directly translates to treatment variability, which has implications for both patient survival and quality of life. This is because the Gleason grade directly determines the percentage of patients who receive radiation therapy, radical prostatectomy, active surveillance, or other treatments. Significant methodological advancements were implemented in 2021, however they were complemented by passion and insightful observations on the disconnect between algorithm development and clinical deployment preparedness. In a thorough systematic analysis published in Nature Machine Intelligence, Roberts and colleagues looked at 314 ML models created during the COVID-19 pandemic for the diagnosis of SARS-CoV-2 infection from chest imaging. They determined that, despite frequently impressive reported performance metrics, the great majority of these models had basic methodological flaws that made them unsuitable for clinical deployment.citeroberts2021common. The identified problems included training on datasets containing systematic biases where COVID-19 positive and negative cases differed in ways unrelated to disease (such as patient positioning, image acquisition protocols, or institutional imaging practices), lack of external validation on independent patient populations, inappropriate comparisons with weak baselines, and overfitting to training data characteristics that would not generalize to new clinical contexts. This analysis had profound influence on subsequent healthcare AI research by establishing clear quality criteria for clinical AI development and demonstrating that impressive performance on internal test sets provides insufficient supporting proof for clinical utility. Finlayson and colleagues contributed complementary methodological insights through their New England Journal of Medicine publication examining the phenomenon of dataset shift the situation where the statistical properties of deployment data differ from training data and demonstrating how AI models trained on historical healthcare data can fail when rolled out in changing clinical environments where patient populations, documentation practices, treatment patterns, or disease prevalence evolve over time. 33
The intermediate period of 2022 and 2023 showed a distinctly accelerating research momentum, with seven investigations making up 35.0% of the sample. Increasingly sophisticated methods for addressing algorithmic fairness, adherence to new reporting standards, and investigation of revolutionary new paradigms, such as foundation models, were the hallmarks of this era. show through thorough external validation that commercial dermatology AI systems considerably and clinically meaningfully decreased diagnostic accuracy when examining skin lesions on people with darker skin tones as opposed to those with lighter skin tones, thereby radically changing conversations about AI fairness in healthcare. 34
The performance differences were attributed to the training datasets’ extreme underrepresentation of various skin kinds, which were mostly composed of images of individuals with fair skin. This illustrates how past biases in clinical data gathering and medical research immediately translate into algorithmic biases that could worsen healthcare inequities that already afflict minority populations rather than ameliorate them. In following healthcare AI research, regulatory guidelines, 35 and clinical deployment criteria, this study significantly enhanced the focus on demographic subgroup analysis and fairness considerations. A significant gap in methodological standardization that had impeded research quality and comparability across studies had been addressed by the DECIDE-AI guidelines published in BMJ, which established expert consensus-based recommendations for the design, conduct, and reporting of early-stage clinical evaluation studies for AI-based diagnostic systems. 31
Foundation models emerged in 2023 as a revolutionary paradigm with potentially significant consequences for the development and application of AI in healthcare. Moor and colleagues published a seminal analysis in Nature establishing conceptual frameworks for understanding how large neural networks pre trained on massive datasets could be adapted to diverse medical applications with minimal task specific training, potentially accelerating AI development across clinical domains while raising new questions about validation, interpretability, and appropriate clinical use. 18 The first prospective randomized controlled trials evaluating DL for cardiovascular risk prediction, demonstrating in their JAMIA publication that AI analysis of retinal photographs achieved superior discrimination compared to the Pooled Cohort Equations currently recommended for cardiovascular risk assessment, with C-index of 0.81 versus 0.72 representing a clinically meaningful 12.5% relative improvement in discriminating individuals who would versus would not experience cardiovascular events. 36 JAMA results from a randomized controlled trial demonstrating that ML based intraoperative hypotension prediction could reduce hypotensive episodes during surgery and improve patient outcomes, providing highest level evidence that AI systems could translate to meaningful clinical benefit when as required integrated into care processes. 6
The number of publications reached its highest level in 2024, with seven studies accounting for 35.0% of the sample. This may have been due to a number of factors coming together, such as the maturation of research projects started during the COVID-19 pandemic, the growing adoption of rigorous validation methodologies in response to earlier criticisms, the increased regulatory attention creating clearer pathways for the approval of AI medical devices, and the expansion of AI applications into new clinical domains and modalities. By presenting FL approaches that allow collaborative model training across multiple institutions without requiring the centralization of sensitive patient data, advanced privacy-preserving AI development is demonstrated. This addresses important concerns regarding data governance, patient privacy, and the ethical and legal constraints that have historically limited the amount of multi-institutional AI development. 20
Clinical acceptance depends on factors far beyond raw diagnostic accuracy, such as interface usability, confidence calibration, explanation quality, workflow integration, and trust calibration insights essential for translating technical AI advances into adopted clinical tools, as determined by real-world human factors dimensions of radiology AI adoption through prospective surveys of practicing Korean radiologists. 37 With significant ramifications for developers looking for international deployment, the increasingly complicated global regulatory regulatory environment for AI medical devices documents significant variation in requirements, timelines, and approval pathways across jurisdictions, 38 among which are the United States, United Kingdom, European Union, Canada, Japan, China, and emerging markets. 39 AI-assisted clinical decision support, medical education, and patient communication could all be greatly impacted by LLMs’ ability to achieve medical reasoning capabilities which are comparable to physician-level performance on clinical questions and medical licensing examination content, according to two additional studies conducted in 2025 through September (10.0.
AI Technologies and Clinical Applications
The technological approaches employed throughout all of the included research demonstrated the ongoing dominance of significantly-established DL methodologies as well as the emergence of revolutionary new paradigms that could fundamentally alter the development and application of AI in healthcare. Because individual studies commonly employ numerous complimentary AI approaches to maximize clinical performance, there is a significant overlap between technology categories. The collected literature was dominated by DL techniques, with twelve investigations (60.0%) utilizing deep neural network architectures. This is particularly true for applications that use unstructured data types, like physiological time-series signals, clinical narratives, and medical images, where DL algorithms have shown noticeably superior pattern recognition abilities over conventional feature-engineered ML techniques. CNNs are the most widely utilized specialized architecture, with eight studies (40.0%) using CNN-based techniques for image processing in radiology, pathology, dermatology, and ophthalmology applications. Classical ML algorithms continued to be highly relevant with ten studies (50.0%) employing traditional approaches such as random forests, gradient boosting techniques, and SVM. This is particularly true for structured clinical data, where simpler models are preferred based on interpretability requirements. Five research (25.0%) used NLP techniques made possible by transformer architecture advancements, indicating a significant expansion of NLP applications. With four research (20.0%) examining big pre-trained models that are capable of handling a variety of tasks relatively little extra training, foundation models emerged as a revolutionary advance.
The distribution of AI applications across clinical specialties can be observed in Table 2 and Figure 3, which show a concentration in imaging-intensive domains along with new applications in primary care and critical care settings. With six research that comprised 30% of the sample, radiology applications were most prevalent in the included literature, indicating a natural connection between DL image analysis capabilities and radiological practice that concentrated on image interpretation.
Table 2.
Distribution of AI Applications by Clinical Specialty (n = 20). Summary of Included Studies Grouped by Clinical Specialty, Highlighting Primary AI Technologies and Key Clinical Applications
| Clinical specialty | n | % | Primary AI technologies | Key applications |
|---|---|---|---|---|
| Radiology | 6 | 30.0 | CNN, DL, Transfer learning | Chest X-ray, CT, MRI, mammography |
| Oncology | 4 | 20.0 | DL, ML, Genomic analysis | Cancer detection, treatment response |
| Cardiology | 3 | 15.0 | Neural networks, ML | ECG analysis, cardiovascular risk prediction |
| Ophthalmology | 2 | 10.0 | CNN, DL | Diabetic retinopathy, glaucoma |
| Critical Care | 2 | 10.0 | ML, Time-series models | Sepsis prediction, mortality risk |
| Pathology | 1 | 5.0 | CNN, Whole-slide imaging | Tumor grading, cancer classification |
| Dermatology | 1 | 5.0 | CNN, Classification | Skin lesion diagnosis, fairness |
| Primary Care | 1 | 5.0 | LLM, Foundation models | Clinical QA, documentation |
Figure 3.

Distribution of AI applications by clinical specialty among included studies (n=20). Radiology represents the largest proportion (30%), followed by oncology (20%) and cardiology (15%). Imaging-intensive specialties collectively account for 65% of applications, reflecting the natural alignment between DL capabilities and image-based diagnostic tasks
With six research that make up 30% of the complete sample, radiology applications predominated in the included literature, indicating a natural connection between DL image analysis capabilities and radiological practice focused on picture interpretation. In a variety of imaging modalities and clinical tasks, AI systems demonstrated expert-level diagnostic ability.
In comparison to typical radiologist performance, breast cancer screening AI achieved an AUC of 0.889 with 5.7% fewer false positives and 9.4% less false negatives. 5 AUC of 0.91 for multi-modal diagnostic imaging AI that integrates CT, MRI, and clinical data, as compared with 0.87 for single modality specialist evaluation. 17 In accordance with Roberts and colleagues’ critical methodological review, the majority of COVID-19 imaging AI models found unsuitable for deployment due to dataset biases and validation issues. 13 Adoption of AI by radiologists has been affected by human factors, and acceptability varies depending on workflow integration, confidence calibration, and interface design in along with raw accuracy. 37
Oncology represented the second largest application domain with four studies (20.0%) addressing cancer detection, histopathological grading, and prognostic prediction. Bulten and colleagues achieved AI Gleason grading with kappa agreement of 0.918 exceeding inter pathologist agreement of 0.819, demonstrating potential to reduce diagnostic variability in high stakes cancer grading. 32 Muthamilselvan and colleagues achieved cancer prognosis prediction with AUC 0.85 versus 0.78 for standard clinical models. 40 Cardiology applications (15.0%, n=3) included cardiovascular risk prediction, with Lee and colleagues demonstrating retinal imaging-based AI achieving C-index 0.81 versus 0.72 for Pooled Cohort Equations in a prospective randomized trial, 36 and Shishehbori and colleagues achieving AUC 0.79 versus 0.71 for Framingham using gradient boosting. 41 Critical care applications demonstrated AI capability for time-sensitive prediction, ML hypotension prediction achieving AUROC 0.84 versus 0.75 for standard monitoring and reducing adverse events in a randomized trial. 6 Dermatology AI achieved AUC 0.91 versus 0.88 for expert dermatologists, though with documented performance disparities across skin tones highlighting fairness concerns. 34 LLM research showing medical reasoning accuracy comparable to the physician level provided the development of primary care applications. 15
Performance and Validation
With a descriptive median AUC of 0.91 across individual imaging studies (not a pooled metaanalytic estimate) and consistent outperformance of established clinical risk calculators for prediction tasks, AI systems demonstrated diagnostic and prognostic performance across clinical domains that were comparable to or exceeded human specialist benchmarks in the individual studies reviewed. Due to heterogeneity in clinical tasks, patient populations, and evaluation metrics across included studies, formal meta-analytic pooling was not performed. AI breast cancer screening in radiology achieved an AUC of 0.889 compared to radiologists’ AUC of 0.871, resulting in significant decreases in both missed benign tumors and needless biopsies. 5 AI Gleason grading consistency (kappa 0.918) in pathology is significantly greater than expert pathologist agreement (kappa 0.819). 32 Cardiovascular risk models achieved a 12.5% relative improvement in discrimination, consistently outperforming clinical calculators that were decades old. 36 Given that external validation, which was carried out in only 45.0% of studies, usually revealed 5 to 15% performance degradation compared to internal validation, performance metrics must be interpreted carefully. This demonstrates the degree to which AI models learn institution-specific patterns the fact that do not generalize.33,42
Implementation Barriers
Several implementation limitations have been identified in all of the included studies, which reflects the complexity of the sociotechnical difficulties of clinical AI deployment, where technical performance has become only one aspect of effective implementation. The frequency distribution of the barriers found can be observed in Table 3.
Table 3.
Implementation Barriers Identified (n = 20)
| Implementation barrier | n | % |
|---|---|---|
| Regulatory compliance | 11 | 55.0 |
| Algorithm transparency | 8 | 40.0 |
| Data quality/bias | 7 | 35.0 |
| Clinical integration | 6 | 30.0 |
| Clinician training | 5 | 25.0 |
| Cost/resources | 4 | 20.0 |
| Patient acceptance | 3 | 15.0 |
Eleven research (55.0%) identified regulatory problems as major hurdles, making regulatory compliance the most often mentioned barrier. Adaptive AI systems that may change through learning or updates are difficult to integrate into traditional medical device frameworks built for static technologies. While the EU Medical Device Regulation included extensive regulations that many developers find difficult to manage, the FDA completed advice on Predetermined Change Control Plans for AI/ML devices in December 2024.43,44 substantial regulatory variation across global jurisdictions complicating international deployment. 39 Eight studies (40.0%) have been influenced by algorithm transparency issues; in clinical settings where professional responsibility for choices is required, the ”black box” character of high-performing DL models presents accountability issues.45,46 Seven studies (35.0%) were impacted by data quality and algorithmic bias issues; training data constraints directly result in performance disparities that impact disadvantaged populations.34,47 Six studies (30.0%) suffered from by clinical workflow integration issues, and regardless of technical performance, AI tools that disrupt existing procedures encounter resistance.48,49 Concerns about patient acceptance (15.0%), cost and resource limitations (20.0%), and physician training requirements (25.0%) identified more challenges.
Comparison With Earlier Reviews
When compared to earlier systematic investigations conducted between 2015 and 2019, notable improvements in implementation awareness and validation rigor were discovered. Table 4 displays this comparison.
Table 4.
Comparison of Current Review With Earlier Systematic Reviews. Key Differences in Scope, Validation Practices, Methodological Quality, Performance Reporting, and Implementation Considerations Between Prior Reviews and the Current Review (2020–2025)
| Characteristic | Liu et al (2019) 26 | Nagendran et al (2020) 27 | Current review (2020–2025) |
|---|---|---|---|
| Review period | 2012–2019 | 2010–2019 | 2020–2025 |
| Studies included | 82 | 81 | 20 |
| Focus | Medical imaging | Deep learning claims | Clinical AI validation |
| Validation Practices | |||
| External validation rate | 6% | 23% | 45% |
| Prospective validation | 8% | 10% | 25% |
| Multi-site testing | 12% | 15% | 35% |
| Methodological Quality | |||
| Sample size justification | 15% | 18% | 40% |
| Subgroup analysis | 8% | 12% | 35% |
| Comparator: clinician | 58% | 42% | 65% |
| Reporting: TRIPOD/CONSORTAI | N/A | 5% | 45% |
| Performance (Imaging Tasks) | |||
| Median AUC | 0.94 | 0.93 | 0.91 |
| AI ≥ Human expert | 94% | 70% | 75% |
| Implementation Considerations | |||
| Regulatory discussion | 12% | 18% | 55% |
| Bias/fairness analysis | 5% | 8% | 35% |
| Clinical workflow discussed | 8% | 15% | 45% |
Note. All performance values (median AUC, AI ≥ Human expert percentages) represent descriptive summaries across individual included studies. No formal meta-analytic pooling was performed due to heterogeneity in study designs, clinical tasks, and outcome metrics.
External validation rates quickly increased from 6 to 23% to 45% as a result of lessons learnt from high-profile AI failures, such as COVID-19 imaging models. 13 Prospective validation went from 8 to 10% to 25%, while multi-site testing went from 12 to 15% to 35%. Sample size jus-tification grew from 15 to 18% to 40%, subgroup analysis increased from 8 to 12% to 35%, and reporting of standard adherence increased from nearly zero to 45%. With regulatory discussion rising from 12–18% to 55%, bias analysis from 5–8% to 35%, and workflow discussion from 8–15% to 45%, implementation awareness significantly improved. The modest decline in median.
AUC from 0.94 to 0.91 likely reflects more rigorous evaluation methodology rather than decreased capability, as external validation on challenging datasets produces more realistic estimates than internal testing on favorable data subsets.
Discussion
Principal Findings
Twenty AI healthcare research that were published between 2020 and 2025 a time of considerable advancement in the field were included in this systematic analysis. For unstructured data, especially medical imaging, DL continued to be the most popular methodology, whereas traditional ML continued to be used for organized clinical data. NLP and foundation models become significant developments, especially after 2023.
Thirty percent of applications were in radiology, where AI systems routinely performed on par with or better than human experts in controlled environments. Strong technical performance can be demonstrated by the median diagnostic AUC of 0.91, but external validation showed 5–15% performance degradation, indicating continued difficulties with generalizability.26,27
Significant advancements in validation procedures have been demonstrated as compared to previous systematic studies. External validation rates increased from 6 to 23% between 2015 and 2019, and from 2020 to 2025, they reached 45%. These gains reflect a stronger emphasis on generalizability and lessons learned from high-profile AI failures. A move toward practical application was evident in the percentage of prospective validation studies, which rose from 8 to 10% to 25%. Subgroup analyses were present in 35% of recent studies, compared to 8–12% in previous research, suggesting a notable increase in the focus on algorithmic fairness.
Comparison With Existing Literature
Our findings are consistent with earlier systematic assessments that demonstrate AI’s excellent technical performance under controlled circumstances.50,51 AI either matched or surpassed physician performance in 14 out of 20 imaging comparisons. 26 Inappropriate comparators and a l lack of external validation are significant methodological mistakes. 27
Our findings enhance on this research by showing advancements in scientific rigor and validation procedures between 2020 and 2025. The field appears to be reacting to previous criticisms based on the rise in external validation, prospective studies, and regulatory awareness. There are still significant gaps: just 45% of research involved external validation, and there have been still significant implementation issues.
Geographical concentration in high-income nations (90% from North America, Europe, and East Asia) is consistent with more general trends in health informatics research. 39 Lowand middle-income nations generally underrepresented, which restricts generalizability and creates fairness issues. Due to training data biases, dermatology AI performs poorly on darker skin tones, demonstrating how demographic and regional constraints result in clinical performance disparities. 34
Clinical Implications
The quality, safety, efficiency, and accessibility of healthcare could be substantially enhanced by AI technologies. Improved diagnostic accuracy has the potential to reduce the estimated millions of diagnostic errors occurring annually. 52 Enhanced risk stratification enables targeted preventive interventions and intensified monitoring for high-risk patients. The integration of multiple data modalities supports precision medicine approaches to individualized therapy. 53
However, realizing these benefits requires addressing several interconnected implementation challenges. First, clinician training programs must be developed to ensure that healthcare professionals can effectively interpret and appropriately act upon AI-generated recommendations while maintaining independent clinical judgment. 54 Second, healthcare organizations must invest not only in technical infrastructure but also in change management processes and ongoing support systems that facilitate sustainable AI adoption. Third, regulatory frameworks must balance patient safety protections with sufficient flexibility to accommodate the iterative, continuously learning nature of AI systems, as highlighted by the FDA’s guidance on Predetermined Change Control Plans for AI/ML devices. 43 Fourth, payment and reimbursement models should be aligned with quality-improvement objectives to incentivize appropriate AI adoption rather than penalizing innovation. 55 Fifth, policymakers should pursue greater international harmonization of AI medical device regulations, as our findings confirm significant variation across jurisdictions that complicates cross-border deployment. 39 Finally, requirements for demographic subgroup performance reporting should be embedded in regulatory approval processes to ensure that AI systems do not exacerbate existing health disparities, as demonstrated by the documented performance gaps in dermatology AI.34,56
Future Directions
Real-World Validation: There are still few prospective clinical trials contrasting AI-assisted care with conventional care. 6 To evaluate efficacy, cost-effectiveness, and implementation aspects, pragmatic trials integrated into standard clinical settings has to be given top priority.
Health Equity: Throughout the development of AI, algorithmic bias and fairness must be carefully considered. Future research should report performance across demographic groupings, assess mitigation techniques, and investigate causes of bias. 56 Diverse stakeholder input can be ensured through community-engaged research methodologies.
Implementation Science: It is necessary to conduct methodical study utilizing implementation science frameworks in order to comprehend the elements that contribute to the successful adoption of AI. Research should look at training efficacy, workflow integration, organizational preparedness, and sustainability. 57
Emerging Technologies: Foundation models, FL, and explainable AI techniques hold promise for addressing current limitations. Multi modal models integrating diverse data types may improve diagnostic accuracy. FL allows collaborative development while preserving data privacy.19,20 Explainable AI approaches promise greater transparency through causal inference and neural symbolic integration.58,59
Limitations
This review is limited in a number of ways. First, we only included peer-reviewed Englishlanguage papers, which may have left out pertinent gray literature or research in other languages. Because AI is developing so quickly, new advances might not yet be published in peer-reviewed journals. In September 2025, our hunt came to an end.
Second, as studies with favorable results are published more frequently than those with negative results, publication bias probably has an impact on the results. 60 A lot of failed AI projects never make it to publication.
Third, for the majority of applications, formal meta-analysis was not possible due to study heterogeneity. Direct comparison was difficult due to differences in datasets, algorithms, evaluation metrics, and populations.
Fourth, the majority of research came from affluent countries with robust healthcare systems. There is also uncertainty over generalizability to lowand middle-income nations with distinct therapeutic practices, legal systems, and data infrastructure.
Conclusions
AI has substantial potential to transform healthcare by processing vast data volumes, identifying complex patterns, and supporting clinical decision making. This systematic review of 20 studies from 2020 to 2025 demonstrates continued technical advancement alongside meaningful improvements in validation practices compared to the 2015 to 2019 period.DL and ML achieve excellent technical performance across clinical applications, with medical imaging, oncology, and cardiology as leading research areas. In controlled settings, AI systems frequently match or exceed human expert performance. External validation rates have increased from 6 to 23% to 45%, and attention to algorithmic fairness has grown substantially. Implementation challenges persist. Regulatory frameworks, algorithm transparency, data quality, and clinical integration remain significant barriers. Only 45% of studies included external validation, and prospective clinical trials remain scarce. Successful AI implementation requires investments in clinician training, organizational infrastructure, and change management. Policymakers must develop regulatory frameworks balancing innovation with safety. Patient values, safety, and equitable access must remain central priorities. Future research should prioritize real-world validation through pragmatic clinical trials, health equity research addressing algorithmic bias, and implementation science understanding adoption determinants. Foundation models, FL, and explainable AI offer promising directions for addressing current limitations.
Footnotes
Funding: The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported by the Ministry of Education, Ministry of Science & ICT, and Ministry of Health and Welfare, Republic of Korea (grant numbers: NRF [2021-R1-I1A2 (059735)], RS [2024-0040 (5650)], RS [2019-II19 (0421)], RS [2025-2544 (3209)], RS [2020-II20 (1821)], RS [2026-2561 (0094)], RS [2026-2561 (5334)], and RS [2026-2561 (3012)].
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
ORCID iD
Ghulam Hussain Noori https://orcid.org/0009-0005-4780-8765
References
- 1.Korteling JE, Boer-Visschedijk GC, Blankendaal RAM, Boonekamp RC, Eikelboom AR. Human-versus artificial intelligence. Front Artif Intell. 2021;4:622364. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25:44-56. [DOI] [PubMed] [Google Scholar]
- 3.Rajkomar A, Dean J, Kohane I. Machine learning in medicine. N Engl J Med. 2019;380:1347-1358. [DOI] [PubMed] [Google Scholar]
- 4.Davenport T, Kalakota R. The potential for artificial intelligence in healthcare. Future Healthc J. 2019;6:94-98. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.McKinney SM, Sieniek M, Godbole V, et al. International evaluation of an AI system for breast cancer screening. Nature. 2020;577:89-94. [DOI] [PubMed] [Google Scholar]
- 6.Wijnberge M, Geerts BF, Hol L, et al. Effect of a machine learning-derived early warning system for intraoperative hypotension vs standard care on depth and duration of intraoperative hypotension during elective noncardiac surgery: the HYPE randomized clinical trial. JAMA. 2020;323:1052-1060. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.He J, Baxter SL, Xu J, Xu J, Zhou X, Zhang K. The practical implementation of artificial intelligence technologies in medicine. Nat Med. 2019;25:30-36. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Beam AL, Kohane IS. Big data and machine learning in health care. JAMA. 2018;319:1317-1318. [DOI] [PubMed] [Google Scholar]
- 9.Capital Markets . Healthcare data RBC designed to inform. Published 2016. https://www.rbccm.com/en/gib/healthcare/episode/the_healthcare_data_explosion, Accessed December 7, 2025.
- 10.Abbas SR, Abbas Z, Zahir A, Lee SW. Advancing genome-based precision medicine: a review on machine learning applications for rare genetic disorders. Brief Bioinform. 2025;26:bbaf329. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Ebad SA. Healthcare software design and implementation: a project failure case. Softw Pract Exp. 2020;50:1258-1276. [Google Scholar]
- 12.Grand View Research . Artificial Intelligence (AI) in Healthcare Market Size, Share and Trends Analysis Report. Published 2025. Available at. https://www.grandviewresearch.com/industry-analysis/artificial-intelligence-ai-healthcare-market, Accessed December 7, 2025.
- 13.Roberts M, Driggs D, Thorpe M, et al. Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans. Nat Mach Intell. 2021;3:199-217. [Google Scholar]
- 14.Wynants L, Van Calster B, Collins GS, et al. Prediction models for diagnosis and prognosis of COVID-19: systematic review and critical appraisal. BMJ. 2020;369:m1328. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Singhal K, Tu T, Gottweis J, et al. Toward expert-level medical question answering with large language models. Nat Med. 2025;31:943-950. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Ali S, Shahab O, Al Shabeeb R, et al. General purpose large language models match human performance on gastroenterology board exam self-assessments. medRxiv. 2023:2023-2029 [Google Scholar]
- 17.Xu X, Li J, Zhu Z, et al. A comprehensive review on synergy of multi-modal data and AI technologies in medical diagnosis. Bioengineering. 2024;11:219. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Moor M, Banerjee O, Abad ZSH, et al. Foundation models for generalist medical artificial intelligence. Nature. 2023;616:259-265. [DOI] [PubMed] [Google Scholar]
- 19.Rieke N, Hancox J, Li W, et al. The future of digital health with federated learning. NPJ Digit Med. 2020;3:119. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Abbas SR, Abbas Z, Zahir A, Lee SW. Federated learning in smart healthcare: a comprehensive review on privacy, security, and predictive analytics with IoT integration. Healthcare. 2024;12:2587. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Parikh RB, Teeple S, Navathe AS. Addressing bias in artificial intelligence in health care. JAMA. 2019;322:2377-2378. [DOI] [PubMed] [Google Scholar]
- 22.Jacob C, Brasier N, Laurenzi E, et al. AI for IMPACTS framework for evaluating the long-term real-world impacts of AI-powered clinician tools: systematic review and narrative synthesis. J Med Internet Res. 2025;27:e67485. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Abbas SR, Seol H, Abbas Z, Lee SW. Exploring the role of artificial intelligence in smart healthcare: a capability and function-oriented review. Healthcare. 2025;13:1642. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Gkegkas C. Artificial intelligence in mental healthcare: adoption, perception, and strategic implementation: a mixed-methods study of clinician experiences and organizational readiness. Published. 2025 [Google Scholar]
- 25.Panarese P, Grasso MM, Solinas C. Algorithmic bias, fairness, and inclusivity: a multilevel framework for justice-oriented AI. AI Soc. 2025;41:1-23. [Google Scholar]
- 26.Liu X, Faes L, Kale AU, et al. A comparison of deep learning performance against healthcare professionals in detecting diseases from medical imaging: a systematic review and metaanalysis. Lancet Digit Health. 2019;1:e271-e297. [DOI] [PubMed] [Google Scholar]
- 27.Nagendran M, Chen Y, Lovejoy CA, et al. Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims of deep learning studies. BMJ. 2020;368:m689. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Liu X, Rivera SC, Moher D, et al. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Lancet Digit Health. 2020;2:e537-e548. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Martindale APL, Llewellyn CD, De Visser RO, et al. Concordance of randomised controlled trials for artificial intelligence interventions with the CONSORT-AI reporting guidelines. Nat Commun. 2024;15:1619. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Vasey B, Nagendran M, Campbell B, et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. BMJ. 2022:377. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Bulten W, Pinckaers H, Van Boven H, et al. Automated deep-learning system for Gleason grading of prostate cancer using biopsies: a diagnostic study. Lancet Oncol. 2020;21:233-241. [DOI] [PubMed] [Google Scholar]
- 33.Finlayson SG, Subbaswamy A, Singh K, et al. The clinician and dataset shift in artificial intelligence. N Engl J Med. 2021;385:283-286. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Daneshjou R, Vodrahalli K, Novoa RA, et al. Disparities in dermatology AI performance on a diverse, curated clinical image set. Sci Adv. 2022;8:eabq6147. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Chen IY, Pierson E, Rose S, Joshi S, Ferryman K, Ghassemi M. Ethical machine learning in healthcare. Annu Rev Biomed Data Sci. 2021;4:123-144. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Lee CJ, Rim TH, Kang HG, et al. Pivotal trial of a deep-learning-based retinal biomarker (Reti-CVD) in the prediction of cardiovascular disease: data from CMERC-HI. J Am Med Inform Assoc. 2023;31:130-138. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Hwang EJ, Park JE, Song KD, et al. 2023 survey on user experience of artificial intelligence software in radiology by the Korean Society of Radiology. Korean J Radiol. 2024;25:613. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Neha F, Bhati D, Shukla DK, Amiruzzaman M. ChatGPT: transforming healthcare with AI. AI. 2024;5:2618-2650. [Google Scholar]
- 39.Palaniappan K, Lin EYT, Vogel S. Global regulatory frameworks for the use of artificial intelligence (AI) in the healthcare services sector. Healthcare. 2024;12:562. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Muthamilselvan S, Ramasami Sundhar Baabu P, Palaniappan A. Ramasami Sundhar Baabu P, Palaniappan A. Microfluidics for profiling miRNA biomarker panels in AI-assisted cancer diagnosis and prognosis. Technol Cancer Res Treat. 2023;22:15330338231185284. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Shishehbori F, Awan Z. Enhancing cardiovascular disease risk prediction with machine learning models. arXiv. 2024 [Google Scholar]
- 42.Zech JR, Badgeley MA, Liu M, Costa AB, Titano JJ, Oermann EK. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a crosssectional study. PLoS Med. 2018;15:e1002683. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.U.S. Food and Drug Administration . Marketing submission recommendations for a predetermined change control plan for artificial intelligence/machine learning (AI/ML)-enabled device software functions. Published. 2024. [Google Scholar]
- 44.European Parliament and Council . Regulation (EU) 2017/745 on medical devices. Published. 2017. [Google Scholar]
- 45.Char DS, Shah NH, Magnus D. Implementing machine learning in health care: addressing ethical challenges. N Engl J Med. 2018;378:981-983. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Lundberg SM, Lee SI. A unified approach to interpreting model predictions. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. 2017;30:47654774 [Google Scholar]
- 47.Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366:447-453. [DOI] [PubMed] [Google Scholar]
- 48.Sendak MP, Gao M, Brajer N, Balu S. Presenting machine learning model information to clinical end users with model facts labels. NPJ Digit Med. 2020;3:41. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Ancker JS, Edwards A, Nosal S, Hauser D, Mauer E, Kaushal R. Effects of workload, work complexity, and repeated alerts on alert fatigue in a clinical decision support system. BMC Med Inform Decis Mak. 2017;17:36. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Jiang F, Jiang Y, Zhi H, et al. Artificial intelligence in healthcare: past, present and future. Stroke Vasc Neurol. 2017;2:2-243. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Haenssle HA, Fink C, Schneiderbauer R, et al. Man against machine: diagnostic performance of a deep learning convolutional neural network for dermoscopic melanoma recognition in comparison to 58 dermatologists. Ann Oncol. 2018;29:1836-1842. [DOI] [PubMed] [Google Scholar]
- 52.Faiyazuddin M, Rahman SJQ, Anand G, et al. The impact of artificial intelligence on healthcare: a comprehensive review of advancements in diagnostics, treatment, and operational efficiency. Health Sci Rep. 2025;8:e70312. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Ashley EA. Towards precision medicine. Nat Rev Genet. 2016;17:507-522. [DOI] [PubMed] [Google Scholar]
- 54.Chan KS, Zary N. Applications and challenges of implementing artificial intelligence in medical education: integrative review. JMIR Med Educ. 2019;5:e13930.e13930. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Reddy S, Allan S, Coghlan S, Cooper P. A governance model for the application of AI in health care. J Am Med Inform Assoc. 2020;27:491-497. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Vyas DA, Eisenstein LG, Jones DS. Hidden in plain sight: reconsidering the use of race correction in clinical algorithms. N Engl J Med. 2020;383:874-882. [DOI] [PubMed] [Google Scholar]
- 57.Proctor E, Silmere H, Raghavan R, et al. Outcomes for implementation research: conceptual distinctions, measurement challenges, and research agenda. Adm Policy Ment Health. 2011;38:65-76. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Holzinger A, Langs G, Denk H, Zatloukal K, Mu¨ller H. Causability and explainability of artificial intelligence in medicine. Wiley Interdiscip Rev Data Min Knowl Discov. 2019;9:e1312. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Houssein EH, Gamal AM, Younis EMG, Mohamed E. Explainable artificial intelligence for medical imaging systems using deep learning: a comprehensive review. Cluster Comput. 2025;28:469. [Google Scholar]
- 60.Easterbrook PJ, Gopalan R, Berlin JA, Matthews DR. Publication bias in clinical research. Lancet. 1991;337:867-872. [DOI] [PubMed] [Google Scholar]
