Abstract
Objectives
There is emerging use of artificial intelligence (AI) models to aid diagnostic imaging. This review examined and critically appraised the application of AI models to identify surgical pathology from radiological images of the abdominopelvic cavity, to identify current limitations and inform future research.
Design
Systematic review.
Data sources
Systematic database searches (Medline, EMBASE, Cochrane Central Register of Controlled Trials) were performed. Date limitations (January 2012 to July 2021) were applied.
Eligibility criteria
Primary research studies were considered for eligibility using the PIRT (participants, index test(s), reference standard and target condition) framework. Only publications in the English language were eligible for inclusion in the review.
Data extraction and synthesis
Study characteristics, descriptions of AI models and outcomes assessing diagnostic performance were extracted by independent reviewers. A narrative synthesis was performed in accordance with the Synthesis Without Meta-analysis guidelines. Risk of bias was assessed (Quality Assessment of Diagnostic Accuracy Studies-2 (QUADAS-2)).
Results
Fifteen retrospective studies were included. Studies were diverse in surgical specialty, the intention of the AI applications and the models used. AI training and test sets comprised a median of 130 (range: 5–2440) and 37 (range: 10–1045) patients, respectively. Diagnostic performance of models varied (range: 70%–95% sensitivity, 53%–98% specificity). Only four studies compared the AI model with human performance. Reporting of studies was unstandardised and often lacking in detail. Most studies (n=14) were judged as having overall high risk of bias with concerns regarding applicability.
Conclusions
AI application in this field is diverse. Adherence to reporting guidelines is warranted. With finite healthcare resources, future endeavours may benefit from targeting areas where radiological expertise is in high demand to provide greater efficiency in clinical care. Translation to clinical practice and adoption of a multidisciplinary approach should be of high priority.
PROSPERO registration number
CRD42021237249.
Keywords: SURGERY, Diagnostic radiology, Adult surgery
Strengths and limitations of this study.
This systematic review examined and critically appraised the application of artificial intelligence models to identify surgical pathology from cross-sectional radiological images, including CT, MR, CT-positron emission tomography and bone scans, and identified current limitations and evidence gaps, and, thereby, focusing future research efforts.
Robust methodology was undertaken including screening, data extraction and of risk of bias assessment by two independent reviewers.
Findings may be limited by including English language publications only and incomplete reporting in some of the included studies.
Introduction
The widespread adoption of digital healthcare provides vast data to enable the application of artificial intelligence (AI) in pattern recognition.1 This can alleviate the burden of, or replace, tasks traditionally dependent on clinicians. Examples include the interpretation of medical images for diagnostic, prognostic, surveillance and management decisions, which otherwise rely on a limited number of interpreters and human resources.2 There has been a surge of research into the use of AI in diagnostic imaging, exploring how it can support clinicians and provide greater efficacy and efficiency in clinical care.3
Systematic reviews on the diagnostic accuracy of AI in medical imaging (including respiratory medicine, ophthalmology and breast cancer),4 gastroenterology,5 neurosurgery6 and vascular surgery7 have demonstrated the diverse application of AI models to detect many pathologies. A variety of imaging modalities have been explored (eg, CT, MR and positron emission tomography (PET)). AI models can demonstrate diagnostic performances equivalent to that of experts8 and with greater efficiency, for example, in the time taken to diagnose childhood cataracts (5.7 min quicker than senior consultants).9 However, while AI technologies offer to markedly reduce the clinical workload, ‘black box’ techniques (AI algorithms) can be difficult or impossible to interpret, which can be a barrier to adopting these techniques into clinical practice. Furthermore, many AI studies are proof-of-concept10 and poorly reported,11 12 including limited details on participants, making it difficult to replicate or interpret the study findings.12
A systematic review of the diagnostic accuracy of AI models in cross-sectional radiological imaging of the abdominopelvic cavity is lacking.13 Synthesis of the current AI research in this area could benefit several different surgical specialities which image this region, such as endocrine surgery, gastrointestinal surgery, obstetrics and gynaecology, urology and vascular surgery to guide their clinical decision-making. This study aimed to conduct a systematic review to examine and critically appraise the application of AI models to identify surgical pathology from cross-sectional radiological images, including CT, MR, CT-PET and bone scans, of the abdominopelvic cavity, to identify current limitations and inform future research efforts.
Methods
Protocol and registration
This systematic review was registered with the International Prospective Register of Systematic Reviews. A study protocol has previously been published.13 It is reported in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-analyses of Diagnostic Test Accuracy Studies.14
Information sources
Electronic searches of OVID SP versions of Medline, EMBASE and the Cochrane Central Register of Controlled Trials databases were conducted to identify all potentially relevant studies. Date limitations between 1 January 2012 and 31 July 2021 were applied, to account for advancements in machine learning performance and the development of deep learning approaches since 2012, in line with existing reviews.15 Reference lists of included articles were screened to identify further relevant studies.
Search strategy and study selection
A comprehensive search syntax was developed with adaptation from three existing search strategies6 7 16 and guidance from an information specialist using text words and medical subject headings related to three domains: ‘artificial intelligence’, ‘diagnostic imaging’ and the ‘abdominopelvic cavity’ (online supplemental appendix S1). Database search results were imported into reference management software (EndNote V.X9, Clarivate Analytics, USA) and duplicates were removed.
bmjopen-2022-064739supp001.pdf (40.8KB, pdf)
Assessment of study eligibility was performed in two stages. First, titles and abstracts were screened for inclusion by two independent reviewers (GEF and CH). Any conflicts were resolved through discussion, referring to the wider study team if required. Final eligibility was assessed by a full-text review of potentially eligible studies by the same process. Management of the screening process was aided by Rayyan software (Rayyan Systems, Cambridge, Massachusetts).17
Eligibility criteria
Primary research studies were considered for eligibility using the PIRT (participants, index test(s), reference standard and target condition) framework.18 Participants were adults with pathology within the abdominopelvic cavity diagnosed using the following radiological modalities: CT, MR, CT-PET or bone scans. Diagnostic endoscopy was excluded as existing reviews have explored the performance of AI models in this area.19 20 The index test was studies considering AI models as an intervention with the aim to provide a diagnosis. The reference standard was ‘standard practice’ to allow for variation across the included studies. The target condition was abdominopelvic cavity pathology which has had, or may warrant, an invasive procedure21 for therapeutic intent.
Excluded were secondary research studies (eg, systematic reviews), case reports and case series, absence of full text (eg, conference abstracts), animal studies and non-English articles.
Data extraction and management
Data extraction from the included articles was independently performed by two reviewers (GEF and CH). Data management software (REDCap V.9.5.23, Vanderbilt University, USA)22 and a predesigned standardised form were used. Data were extracted under the following three subheadings:
Study characteristics
Extracted data included the name of the first author, their affiliated country, composition of the study team (eg, software engineers, radiologists and surgeons who routinely operate on surgical pathology within the abdominopelvic cavity (eg, gastrointestinal surgeons, urologists and gynaecologists)), year of publication, study aim and design (ie, ‘prospective’ or ‘retrospective’), surgical subspeciality and pathology studied (ie, benign, malignant tumours, multiple or other). Information on the reporting of ethics and/or regulatory approval (eg, Medicines and Healthcare products Regulatory Agency), patient and publication involvement and authors’ mention of using a reporting guideline (eg, Standards for Reporting of Diagnostic Accuracy Studies) was recorded.
Training data
Extracted data on the input features (data used to develop the AI model) included the modality of cross-sectional radiological imaging (CT, MR, CT-PET and bone scans), the AI model used, the reference standard and the reporting and size of the training and test sets. Information on whether the training data came from the studies dataet or publicly available datasets was recorded.
Outcomes
The performance of the AI models and human comparator (where applicable) was extracted. Diagnostic measures of accuracy included reported sensitivity, specificity, positive predictive values and the area under the receiver operating characteristic curve (AUC). The interpretation time (seconds) for studies comparing the performance of the AI model with a human comparator (where reported) was extracted.
Risk of bias and applicability
Risk of bias was assessed independently by two reviewers (GEF and RM) using the Quality Assessment of Diagnostic Accuracy Studies-2 (QUADAS-2) tool.23 A version of the QUADAS-2 tool for AI studies24 was still in development and not available at the time of conducting the current review. The generic QUADAS-2 tool with the pre-existing modified signalling questions was used to assess four domains including patient selection, index test, reference standard and flow and timing.23 An overall judgement of ‘at risk of bias’ or ‘concerns regarding applicability’ was assigned if more one or more domains were judged as ‘high’ or ‘unclear’.23 Judgements of applicability assessed whether the study matched the review question.23
Data synthesis
A narrative synthesis was conducted according to the Synthesis Without Meta-analysis guidelines.25 The synthesis was planned to focus on the primary outcome with studies grouped by the modality of radiological imaging, surgical subspeciality and pathology studied, as outlined in the protocol.13 A broader approach was, however, adopted due to the small number of included studies and their heterogeneity. A meta-analysis was not performed due to the broad nature of the included studies.
Patient and public involvement
As part of the wider programme of work (Bristol Biomedical Research Centre, National Institute for Health Research Bristol BRC), patients and the public were consulted on their views of AI being used to guide doctors to make decisions about treatment. Overall, it was perceived positively, and they were supportive of its adoption in healthcare.
Results
Database searching identified 628 records, with a further five studies identified through reference lists of included articles. After the removal of duplicates, 580 were screened and, 52 full-text articles were assessed for eligibility. Fifteen studies were finally included (figure 1).26–40
Figure 1.
PRISMA flow diagram. PRISMA, Preferred Reporting Items for Systematic Reviews and Meta-Analysis.
Study characteristics
Characteristics of the included studies and details of the AI models are summarised (tables 1 and 2). All were retrospective studies conducted in six different countries: Japan (n=4), USA (n=4), China (n=3), India (n=2), Turkey (n=1) and South Korea (n=1). Studies were all proof-of-concept (ie, not applied in a clinical setting), from four surgical specialities: urology (n=6), gastrointestinal surgery (n=6), endocrine surgery (n=1) and gynaecology (n=1). One study was not specific to a single specialty and involved whole-body CT-PET across four anatomical regions (head-and-neck, chest, abdomen and pelvis).29 Most studies focused on malignant tumours (n=11) using CT as the imaging modality (n=8). Most studies (n=14) included an ethical approval statement. No studies, however, mentioned patient and public involvement or the use of a reporting guideline.
Table 1.
Characteristics of the included studies
| Characteristics | Number of studies n=15 |
| Country of origin | |
| Japan | 4 |
| USA | 4 |
| China | 3 |
| India | 2 |
| Turkey | 1 |
| South Korea | 1 |
| Year of publication | |
| 2020 | 5 |
| 2019 | 5 |
| 2018 | 2 |
| 2017 | 1 |
| 2016 | 1 |
| 2013 | 1 |
| Surgical subspecialty | |
| Urology | 6 |
| Gastrointestinal | 6 |
| Endocrine | 1 |
| Gynaecology | 1 |
| Not reported | 1 |
| Pathology studied | |
| Malignant tumours | 11 |
| Multiple pathologies | 2 |
| Other | 2 |
| Modality of radiological imaging | |
| CT | 8 |
| CT-PET | 2 |
| MR | 2 |
| PET | 1 |
| MR and CT-PET | 1 |
| Bone scans | 1 |
PET, positron emission tomography.
Table 2.
Study details and AI models
| First author (Reference) |
Country of origin | Surgical specialty | Pathology studied | Modality of radiological imaging | AI models | Reference standard | Size of training set (patients) | Size of test set (patients) |
| Acar et al26 | Turkey | Urology | Malignancy | CT-PET | KNN | Two nuclear medicine physicians | 90% (total 75 patients) | 10% (total 75 patients) |
| Coy et al27 | USA | Urology | Malignancy | CT | CNN | Pathologist (number not specified) | 90%(total 179 patients) | 10%(total 179 patients) |
| Han et al28 | South Korea | Urology | Malignancy | CT | CNN | Pathologist (number not specified) | 135 | 34 |
| Koizumi et al30 | Japan | Urology | Malignancy | Bone scans | ANN | Two radiology nuclear physicians | Not reported | 226 |
| Lee et al31 | USA | Urology | Malignancy | PET | CNN | One imaging physician | Not reported | Not reported |
| Oberai et al35 | USA | Urology | Malignancy | CT | CNN | One genitourinary pathologist | 120 | 23 |
| Lu et al32 | China | Gastrointestinal | Malignancy | MR | CNN | Radiologists (number not specified) | Not reported | Not reported (+ 414 validation test set) |
| Nayak et al34 | India | Gastrointestinal | Multiple (cirrhosis, hepatocellular carcinoma) | CT | SVM | One radiologist | 40 | Not reported |
| Sethi et al37 | India | Gastrointestinal | Multiple (normal, cystic, calculus or tumour tissues) | CT | ANN, SVM | Unclear | Not reported | Not reported |
| Yasaka et al38 | Japan | Gastrointestinal | Malignancy | CT | CNN | Radiologists (number not specified) | 460 | 100 |
| Yuan et al39 | China | Gastrointestinal | Malignancy | CT | SVM, CNN | One pathologist | 130 | 40 |
| Zhao et al40 | China | Gastrointestinal | Malignancy | MR | CNN | Three radiologists | 293 | 81 |
| Saiprasad et al36 | USA | Endocrine surgery | Other (normal vs abnormal) | CT | RFC | Two radiologists | 5 | 10 |
| Nakagawa et al33 | Japan | Gynaecology | Malignancy | MR and CT-PET | Logistic regression (univariate and multivariate) | Two radiologists and pathologists (number not specified) | 66 | Unclear |
| Kawauchi et al29 | Japan | Uncharacterised | Other (benign/malignant/equivocal) | CT-PET | CNN | One nuclear medicine physician | 2440 | 1045 |
AI, artificial intelligence; ANN, artificial neural network; CNN, convolutional neural network; KNN, k-nearest neighbours algorithm; PET, positron emission tomography; RFC, random forest classification; SVM, support vector machine.
The composition and expertise within each study team varied. Three studies (n=3) comprised teams including software engineers, radiologists and surgeons. A small number of studies were comprised of only radiologists (n=2) or software engineers alone (n=1). Four study teams were comprised software engineers and radiologists together. Of the remaining five studies, three had no apparent radiological team and two were comprised of radiologists with either a team of software engineers, surgeons or physicians.
Training data
AI training and test sets within the included studies comprised a median of 130 (range: 5–2440) and 37 (range: 10–1045) patients, respectively (table 2). This information, however, was not always available (n=6 studies). All training data came from the studies collected data and not from pre-existing readily available data sets. There was variability in the reference standard, including the number of clinicians (range: 1–3) for either a radiological (n=9) or histological (n=4) diagnosis or both (n=1). Only one study had an unclear reference standard (table 2).
Outcomes
The intent of the AI applications in the included studies varied, with the majority focusing on diagnosing advanced or recurrent cancer (n=6 studies) and four studies classifying the pathology (ie, normal or abnormal, or benign or malignant) (online supplemental table 1). Diagnostic performance of the AI models ranged between 70.0% and 95.0% sensitivity, and 52.9% and 98.0% specificity (online supplemental table 1). The reporting of the diagnostic measures of accuracy was unstandardised, for example, there were different output measures across the studies, and three studies did not report all of their outcome measures (online supplemental table 1).
bmjopen-2022-064739supp002.pdf (55.5KB, pdf)
Following model development with a training and tuning set, four studies had an external validation test set and compared the performance of the AI model with a human comparator (a radiologist) (online supplemental table 1).32 33 39 40 The diagnostic performance of the radiologists ranged between 57.4% and 62.8% sensitivity and 0.89 and 0.97 AUC (online supplemental table 1). In these studies, there was variability in both the number of patients (range 50–414; from one to six different centres) and number of radiologists (range 2–4) and their years of experience. These studies reported the diagnostic performance of the AI model as either superior (n=2; in rectal and advanced gastrointestinal cancer)39 40 or comparable (n=2; in metastatic gastrointestinal pathology and a gynaecology study distinguishing between malignant tumours and benign pathology) to radiologists.32 33 Two of these studies included a comparison of the interpretation time of the AI model with the radiologists.32 40 AI models outperformed the radiologists in both studies (1–2 s vs 200 s per case, p<0.0540 and 20 s vs 600 s for an average of 100 MRIs32).
Risk of bias and applicability
With the exception of one study,40 all the included studies had an overall judgement of ‘at risk of bias’ and ‘concerns regarding applicability’ (table 3). This was predominantly due to comparisons based on either a single clinician’s assessment (n=10) or an unclear number of clinicians for the assessment (n=5), from either a small (eg, 10 patients in one study) or unclear size of test set, and most models were developed only from internal validations (n=11) (table 2 and online supplemental table 1).
Table 3.
Study quality assessment (QUADAS-2 tool)
| First author (reference) | Risk of bias | Applicability concerns | |||||
| Patient selection | Index test | Reference standard | Flow and timing | Patient selection | Index test | Reference standard | |
| Acar et al26 | Low | Low | High | Low | Low | Low | High |
| Coy et al27 | Low | Low | High | Low | Low | Low | High |
| Han et al28 | Unclear | Low | High | Low | Low | Low | High |
| Kawauchi et al29 | Low | Low | High | Low | Low | Low | High |
| Lu et al32 | Unclear | Low | High | Low | Low | Low | High |
| Nakagawa et al33 | Low | Low | High | Low | Low | Low | High |
| Koizumi et al30 | Low | Low | High | Low | Low | Low | High |
| Lee et al31 | Low | Low | High | Low | Low | Low | High |
| Nayak et al34 | Unclear | Low | Unclear | Low | Low | Low | High |
| Oberai et al35 | Low | Low | High | Low | Low | Low | High |
| Yasaka et al38 | Unclear | Low | High | Low | Low | Low | High |
| Saiprasad et al36 | Unclear | Low | High | Low | Low | Low | High |
| Sethi et al37 | High | Low | Unclear | Low | Low | Low | Unclear |
| Yuan et al39 | Low | Low | High | Low | Low | Low | High |
| Zhao et al40 | Low | Low | Low | Low | Low | Low | Low |
QUADAS-2, Quality Assessment of Diagnostic Accuracy Studies-2.
Discussion
The major finding of this review was the heterogeneity in the AI applications across the included studies regardless of the surgical specialty or pathology. Early phase studies of AI innovation, particularly focusing on advanced or recurrent malignancy, were identified with promising diagnostic accuracies to support clinical decision-making. Future AI research could benefit from targeting areas where radiological expertise is in high demand or the data are complex to interpret; for example, adrenal incidentalomas41 and images from virtual colonoscopy.2 Attention should also be directed to the governance of AI, particularly on where the responsibility lies if the AI model misses a lesion.
In this review, several reporting issues were identified, including for the reference standard and training data. Poor adherence to reporting guidelines is a common finding in the existing literature for diagnostic accuracy studies assessing AI interventions.4 11 The Standards for Reporting of Diagnostic Accuracy-AI (STARD-AI) Steering Group are developing an AI-specific extension to the STARD statement, which aims to improve reporting of AI diagnostic accuracy studies.42 This steering group highlighted three pitfalls, which are also reflected in this review: (1) unclear methodological interpretation (eg, methods of validation and comparison to human performance), (2) unstandardised nomenclature (eg, varying definition of the term ‘validation’) and (3) heterogeneity of the outcome measures (eg, sensitivity, specificity, predictive values and AUC).43 Endeavours to address this include the development of specific reporting guidelines for authors of AI studies, including protocols (SPIRIT-AI),44 reports (CONSORT-AI)45 and proposals (MINimum Information for Medical AI Reporting (MINIMAR)).46 These efforts should improve both the reporting quality and make it easier to interpret and compare AI studies.
A minority of the included studies compared the diagnostic performance of the AI model with a clinician’s diagnosis. These studies reported a faster and superior or equivalent diagnostic performance with the AI model. A recent review found only 51 studies worldwide reporting the implementation and evaluation of AI applications in clinical practice.47 While many AI studies are currently retrospective and proof-of-concept,10 11 47 which may be appropriate for early phase surgical research, future efforts should evaluate the role of AI in a clinical setting. This should adopt a multidisciplinary team of all relevant stakeholders (eg, software engineers, radiologists and surgeons) to ensure that the diverse and relevant skill sets can work together to produce both high-quality and clinically relevant AI research.
This review included a robust methodology, comprehensive search strategy and a multidisciplinary team. Some limitations, however, are acknowledged. Despite having a broad search strategy, relevant studies may have been missed by excluding articles that were not published in the English language. It did not encompass all diagnostic applications of AI in this region, such as diagnostic imaging for prognostic, surveillance and management decisions meaning findings are not generalisable to these wider contexts. However, recommendations for prioritising future endeavours on clinical need, adhering to reporting guidelines and standardised and transparent reporting can be considered appropriate for all studies assessing AI interventions in healthcare.
This review identified a diverse application of AI innovation in this field. Most studies were proof-of-concept and more ‘comparator’ studies in the clinical setting are needed. Future AI research could build on existing studies with translation to clinical practice, adopting a multidisciplinary approach, including patient and public involvement, which was lacking in the studies of this review. This could target areas of clinical need. Adherence to existing and developing guidelines for reporting AI studies, such as SPIRIT-AI,44 CONSORT-AI,45 STARD-AI43 and DECIDE-AI,48 is warranted.
Supplementary Material
Acknowledgments
The authors would like to thank Ms Catherine Borwick, Information Specialist, University of Bristol, for her input and expertise to develop the search strategy. The authors are grateful to Dr Christin Hoffmann (University of Bristol) for leading patient and public feedback sessions as part of the wider programme of this research and Dr Xiaoxuan Liu (University of Birmingham and University Hospitals Birmingham NHS Foundation Trust) for providing expert comments on the manuscript.
Footnotes
Twitter: @George_Fowler1, @NatalieBlencowe, @Neil_J_Smart, @CSR_Bris
Contributors: GEF and NSB conceived the idea for this systematic review. All authors (GEF, NSB, CH, MPC, NS and RM) contributed to the design of the study. GEF and CH performed screening and data extraction. GEF and RCM conducted the QUADAS-2 assessments. All authors contributed towards the data analysis and interpretation of the data. GEF drafted the manuscript (guarantor of review) and all authors were involved in critiquing the manuscript. All authors approved the final manuscript before submission.
Funding: NSB is funded by an MRC Clinical Scientist Award (grant number: MR/S001751/1). This study was supported by the National Institute for Health and Care Research Bristol Biomedical Research Centre. The views expressed are those of the authors and not necessarily those of the NIHR or the Department of Health and Social Care.
Competing interests: None declared.
Patient and public involvement: Patients and/or the public were not involved in the design, or conduct, or reporting, or dissemination plans of this research.
Provenance and peer review: Not commissioned; externally peer reviewed.
Supplemental material: This content has been supplied by the author(s). It has not been vetted by BMJ Publishing Group Limited (BMJ) and may not have been peer-reviewed. Any opinions or recommendations discussed are solely those of the author(s) and are not endorsed by BMJ. BMJ disclaims all liability and responsibility arising from any reliance placed on the content. Where the content includes any translated material, BMJ does not warrant the accuracy and reliability of the translations (including but not limited to local regulations, clinical guidelines, terminology, drug names and drug dosages), and is not responsible for any error and/or omissions arising from translation and adaptation or otherwise.
Data availability statement
All data relevant to the study are included in the article or uploaded as supplementary information.
Ethics statements
Patient consent for publication
Not applicable.
Ethics approval
Ethical approval was not required as no primary data were collected.
References
- 1.Oren O, Gersh BJ, Bhatt DL. Artificial intelligence in medical imaging: switching from radiographic pathological data to clinically meaningful endpoints. Lancet Digit Health 2020;2:e486–8. 10.1016/S2589-7500(20)30160-6 [DOI] [PubMed] [Google Scholar]
- 2.Hosny A, Parmar C, Quackenbush J, et al. Artificial intelligence in radiology. Nat Rev Cancer 2018;18:500–10. 10.1038/s41568-018-0016-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.He J, Baxter SL, Xu J, et al. The practical implementation of artificial intelligence technologies in medicine. Nat Med 2019;25:30–6. 10.1038/s41591-018-0307-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Aggarwal R, Sounderajah V, Martin G, et al. Diagnostic accuracy of deep learning in medical imaging: a systematic review and meta-analysis. NPJ Digit Med 2021;4:65. 10.1038/s41746-021-00438-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Jin P, Ji X, Kang W, et al. Artificial intelligence in gastric cancer: a systematic review. J Cancer Res Clin Oncol 2020;146:2339–50. 10.1007/s00432-020-03304-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Senders JT, Arnaout O, Karhade AV, et al. Natural and artificial intelligence in neurosurgery: a systematic review. Neurosurgery 2018;83:181–92. 10.1093/neuros/nyx384 [DOI] [PubMed] [Google Scholar]
- 7.Raffort J, Adam C, Carrier M, et al. Artificial intelligence in abdominal aortic aneurysm. J Vasc Surg 2020;72:321–33. 10.1016/j.jvs.2019.12.026 [DOI] [PubMed] [Google Scholar]
- 8.Esteva A, Kuprel B, Novoa RA, et al. Dermatologist-level classification of skin cancer with deep neural networks. Nature 2017;542:115–8. 10.1038/nature21056 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Lin H, Li R, Liu Z, et al. Diagnostic efficacy and therapeutic decision-making capacity of an artificial intelligence platform for childhood cataracts in eye clinics: a multicentre randomized controlled trial. EClinicalMedicine 2019;9:52–9. 10.1016/j.eclinm.2019.03.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Ben-Israel D, Jacobs WB, Casha S, et al. The impact of machine learning on patient care: a systematic review. Artif Intell Med 2020;103:101785. 10.1016/j.artmed.2019.101785 [DOI] [PubMed] [Google Scholar]
- 11.Nagendran M, Chen Y, Lovejoy CA, et al. Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims of deep learning studies. BMJ 2020;368:m689. 10.1136/bmj.m689 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Yusuf M, Atal I, Li J, et al. Reporting quality of studies using machine learning models for medical diagnosis: a systematic review. BMJ Open 2020;10:e034568. 10.1136/bmjopen-2019-034568 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Fowler GE, Macefield RC, Hardacre C, et al. Artificial intelligence as a diagnostic aid in cross-sectional radiological imaging of the abdominopelvic cavity: a protocol for a systematic review. BMJ Open 2021;11:e054411. 10.1136/bmjopen-2021-054411 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.McInnes MDF, Moher D, Thombs BD, et al. Preferred reporting items for a systematic review and meta-analysis of diagnostic test accuracy studies: the PRISMA-DTA statement. JAMA 2018;319:388–96. 10.1001/jama.2017.19163 [DOI] [PubMed] [Google Scholar]
- 15.Liu X, Faes L, Kale AU, et al. A comparison of deep learning performance against health-care professionals in detecting diseases from medical imaging: a systematic review and meta-analysis. Lancet Digit Health 2019;1:e271–97. 10.1016/S2589-7500(19)30123-2 [DOI] [PubMed] [Google Scholar]
- 16.Yang Y, Jin G, Pang Y, et al. The diagnostic accuracy of artificial intelligence in thoracic diseases. Medicine (Baltimore) 2020;99:e19114. 10.1097/MD.0000000000019114 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Ouzzani M, Hammady H, Fedorowicz Z, et al. Rayyan-a web and mobile APP for systematic reviews. Syst Rev 2016;5:210. 10.1186/s13643-016-0384-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Deeks JJ, Wisniewski S, Davenport C, et al. Guide to the contents of a cochrane diagnostic test accuracy protocol. In: Cochrane Handbook for Systematic Reviews of Diagnostic Test Accuracy Version 100 The Cochrane Collaboration. The Cochrane Collaboration, 2013: 1–15. [Google Scholar]
- 19.Hassan C, Spadaccini M, Iannone A, et al. Performance of artificial intelligence in colonoscopy for adenoma and polyp detection: a systematic review and meta-analysis. Gastrointest Endosc 2021;93:77–85. 10.1016/j.gie.2020.06.059 [DOI] [PubMed] [Google Scholar]
- 20.Lui TKL, Tsui VWM, Leung WK. Accuracy of artificial intelligence-assisted detection of upper Gi lesions: a systematic review and meta-analysis. Gastrointest Endosc 2020;92:821–30. 10.1016/j.gie.2020.06.034 [DOI] [PubMed] [Google Scholar]
- 21.Cousins S, Blencowe NS, Blazeby JM. What is an invasive procedure? A definition to inform study design, evidence synthesis and research tracking. BMJ Open 2019;9:e028576. 10.1136/bmjopen-2018-028576 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Harris PA, Taylor R, Minor BL, et al. The redcap consortium: building an international community of software platform partners. J Biomed Inform 2019;95:103208. 10.1016/j.jbi.2019.103208 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Whiting PF, Rutjes AWS, Westwood ME, et al. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med 2011;155:529–36. 10.7326/0003-4819-155-8-201110180-00009 [DOI] [PubMed] [Google Scholar]
- 24.Sounderajah V, Ashrafian H, Rose S, et al. A quality assessment tool for artificial intelligence-centered diagnostic test accuracy studies: QUADAS-AI. Nat Med 2021;27:1663–5. 10.1038/s41591-021-01517-0 [DOI] [PubMed] [Google Scholar]
- 25.Campbell M, McKenzie JE, Sowden A, et al. Synthesis without meta-analysis (SWiM) in systematic reviews: reporting guideline. BMJ 2020;368:l6890. 10.1136/bmj.l6890 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Acar E, Leblebici A, Ellidokuz BE, et al. Machine learning for differentiating metastatic and completely responded sclerotic bone lesion in prostate cancer: a retrospective radiomics study. Br J Radiol 2019;92:20190286. 10.1259/bjr.20190286 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Coy H, Hsieh K, Wu W, et al. Deep learning and radiomics: the utility of google tensorflow. Abdom Radiol (NY) 2019;44:2009–20. 10.1007/s00261-019-01929-0 [DOI] [PubMed] [Google Scholar]
- 28.Han S, Hwang SI, Lee HJ. The classification of renal cancer in 3-phase CT images using a deep learning method. J Digit Imaging 2019;32:638–43. 10.1007/s10278-019-00230-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Kawauchi K, Furuya S, Hirata K, et al. A convolutional neural network-based system to classify patients using FDG PET/CT examinations. BMC Cancer 2020;20:227. 10.1186/s12885-020-6694-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Koizumi M, Motegi K, Koyama M, et al. Diagnostic performance of a computer-assisted diagnosis system for bone scintigraphy of newly developed skeletal metastasis in prostate cancer patients: search for low-sensitivity subgroups. Ann Nucl Med 2017;31:521–8. 10.1007/s12149-017-1175-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Lee JJ, Yang H, Franc BL, et al. Deep learning detection of prostate cancer recurrence with 18f-FACBC (fluciclovine, axumin®) positron emission tomography. Eur J Nucl Med Mol Imaging 2020;47:2992–7. 10.1007/s00259-020-04912-w [DOI] [PubMed] [Google Scholar]
- 32.Lu Y, Yu Q, Gao Y, et al. Identification of metastatic lymph nodes in MR imaging with faster region-based convolutional neural networks. Cancer Res 2018;78:5135–43. 10.1158/0008-5472.CAN-18-0494 [DOI] [PubMed] [Google Scholar]
- 33.Nakagawa M, Nakaura T, Namimoto T, et al. A multiparametric MRI-based machine learning to distinguish between uterine sarcoma and benign leiomyoma: comparison with 18F-FDG PET/CT. Clinical Radiology 2019;74:167. 10.1016/j.crad.2018.10.010 [DOI] [PubMed] [Google Scholar]
- 34.Nayak A, Baidya Kayal E, Arya M, et al. Computer-aided diagnosis of cirrhosis and hepatocellular carcinoma using multi-phase abdomen CT. Int J Comput Assist Radiol Surg 2019;14:1341–52. 10.1007/s11548-019-01991-5 [DOI] [PubMed] [Google Scholar]
- 35.Oberai A, Varghese B, Cen S, et al. Deep learning based classification of solid lipid-poor contrast enhancing renal masses using contrast enhanced CT. Br J Radiol 2020;93:20200002. 10.1259/bjr.20200002 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Saiprasad G, Chang C-I, Safdar N, et al. Adrenal gland abnormality detection using random forest classification. J Digit Imaging 2013;26:891–7. 10.1007/s10278-012-9554-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Sethi G, Saini BS. Computer aided diagnosis system for abdomen diseases in computed tomography images. Biocybern Biomed Eng 2016;36:42–55. 10.1016/j.bbe.2015.10.008 [DOI] [Google Scholar]
- 38.Yasaka K, Akai H, Abe O, et al. Deep learning with CNN showed high diagnostic performance in differentiation of liver masses at dynamic CT. Radiology 2018;286:887–96. 10.1148/radiol.2017170706 [DOI] [PubMed] [Google Scholar]
- 39.Yuan Z, Xu T, Cai J, et al. Development and validation of an image-based deep learning algorithm for detection of synchronous peritoneal carcinomatosis in colorectal cancer. Ann Surg 2022;275:e645–51. 10.1097/SLA.0000000000004229 [DOI] [PubMed] [Google Scholar]
- 40.Zhao X, Xie P, Wang M, et al. Deep learning-based fully automated detection and segmentation of lymph nodes on multiparametric-mri for rectal cancer: a multicentre study. EBioMedicine 2020;56:102780. 10.1016/j.ebiom.2020.102780 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Garrett RW, Nepute JC, Hayek ME, et al. Adrenal incidentalomas: clinical controversies and modified recommendations. AJR Am J Roentgenol 2016;206:1170–8. 10.2214/AJR.15.15475 [DOI] [PubMed] [Google Scholar]
- 42.Sounderajah V, Ashrafian H, Karthikesalingam A, et al. Developing specific reporting standards in artificial intelligence centred research. Ann Surg 2022;275:e547–8. 10.1097/SLA.0000000000005294 [DOI] [PubMed] [Google Scholar]
- 43.Sounderajah V, Ashrafian H, Aggarwal R, et al. Developing specific reporting guidelines for diagnostic accuracy studies assessing AI interventions: the STARD-AI steering group. Nat Med 2020;26:807–8. 10.1038/s41591-020-0941-1 [DOI] [PubMed] [Google Scholar]
- 44.Cruz Rivera S, Liu X, Chan A-W, et al. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. Nat Med 2020;26:1351–63. 10.1038/s41591-020-1037-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Liu X, Cruz Rivera S, Moher D, et al. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat Med 2020;26:1364–74. 10.1038/s41591-020-1034-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Hernandez-Boussard T, Bozkurt S, Ioannidis JPA, et al. MINIMAR (minimum information for medical AI reporting): developing reporting standards for artificial intelligence in health care. J Am Med Inform Assoc 2020;27:2011–5. 10.1093/jamia/ocaa088 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Yin J, Ngiam KY, Teo HH. Role of artificial intelligence applications in real-life clinical practice: systematic review. J Med Internet Res 2021;23:e25759. 10.2196/25759 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Vasey B, Clifton DA, Collins GS. DECIDE-AI: new reporting guidelines to bridge the development-to-implementation gap in clinical artificial intelligence. Nat Med 2021;27:186–7. 10.1038/s41591-021-01229-5 [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
bmjopen-2022-064739supp001.pdf (40.8KB, pdf)
bmjopen-2022-064739supp002.pdf (55.5KB, pdf)
Data Availability Statement
All data relevant to the study are included in the article or uploaded as supplementary information.

