Skip to main content
Indian Journal of Orthopaedics logoLink to Indian Journal of Orthopaedics
. 2024 Apr 8;58(5):457–469. doi: 10.1007/s43465-024-01130-6

Diagnostic Accuracy of Artificial Intelligence-Based Algorithms in Automated Detection of Neck of Femur Fracture on a Plain Radiograph: A Systematic Review and Meta-analysis

Manish Raj 1,, Arshad Ayub 2, Arup Kumar Pal 3, Jitesh Pradhan 4, Naushad Varish 5, Sumit Kumar 6, Seshadri Reddy Varikasuvu 7
PMCID: PMC11058182  PMID: 38694696

Abstract

Objectives

To evaluate the diagnostic accuracy of artificial intelligence-based algorithms in identifying neck of femur fracture on a plain radiograph.

Design

Systematic review and meta-analysis.

Data sources

PubMed, Web of science, Scopus, IEEE, and the Science direct databases were searched from inception to 30 July 2023.

Eligibility criteria for study selection

Eligible article types were descriptive, analytical, or trial studies published in the English language providing data on the utility of artificial intelligence (AI) based algorithms in the detection of the neck of the femur (NOF) fracture on plain X-ray.

Main outcome measures

The prespecified primary outcome was to calculate the sensitivity, specificity, accuracy, Youden index, and positive and negative likelihood ratios. Two teams of reviewers (each consisting of two members) extracted the data from available information in each study. The risk of bias was assessed using a mix of the CLAIM (the Checklist for AI in Medical Imaging) and QUADAS-2 (A Revised Tool for the Quality Assessment of Diagnostic Accuracy Studies) criteria.

Results

Of the 437 articles retrieved, five were eligible for inclusion, and the pooled sensitivity of AIs in diagnosing the fracture NOF was 85%, with a specificity of 87%. For all studies, the pooled Youden index (YI) was 0.73. The average positive likelihood ratio (PLR) was 19.88, whereas the negative likelihood ratio (NLR) was 0.17. The random effects model showed an overall odds of 1.16 (0.84–1.61) in the forest plot, comparing the AI system with those of human diagnosis. The overall heterogeneity of the studies was marginal (I2 = 51%). The CLAIM criteria for risk of bias assessment had an overall >70% score.

Conclusion

Artificial intelligence (AI)-based algorithms can be used as a diagnostic adjunct, benefiting clinicians by taking less time and effort in neck of the femur (NOF) fracture diagnosis.

Study registration

PROSPERO CRD42022375449.

Supplementary Information

The online version contains supplementary material available at 10.1007/s43465-024-01130-6.

Keywords: Artificial intelligence, Fracture detection, Neck of femur, Plain radiograph

Introduction

The junctional configuration of the neck of the femur between the femoral head and shaft makes it the weakest and most fracture-prone portion of the bone, especially in older people whose osteoporotic bones are more brittle [1]. In the next 30 years, it is anticipated that the incidence of neck of femur fractures will double due to the global increase in the number and proportion of older people in the population, also known as population aging [2]. Rapid and immediate diagnosis is required to minimize mortality and morbidity related to neck of femur (NOF) fracture management. The busy emergency room commonly uses a plain radiograph (X-ray) to quickly and affordably discover and diagnose fractures. The misdiagnosis rate in identifying fractures ranges from 4 to 9%, and delay in the diagnosis and initiation of treatment causes adverse effects on long-term outcomes in these patients [35]. A plain radiograph’s interpretation or diagnosis error might have several causes, from the reader’s experience level to perception errors in reading the radiograph [68]. Front-line residents and inexperienced junior doctors are more prone to make diagnostic mistakes while treating many trauma patients in an emergency, because misdiagnosis can occur in a busy emergency room under pressure. Peripheral health systems in developing countries lack specialized medical personnel like radiologists. In situations of a neck of femur fracture, the impact of this diagnostic error ranges from a minor condition to a life-threatening one, and prompt identification and treatment can avoid unfavorable outcomes in these fractures. CT scans and MRI can prevent the diagnostic error in these fractures. However, limiting factors are the higher cost associated with it, the more extended duration requirement, and the availability of these diagnostic tests to limited specialized health centers in developing countries. Literature suggests that medical image analysis benefits significantly from an algorithm with artificial intelligence (AI) influences [913]. It may be able to quickly identify all trauma-related abnormalities on pelvic radiographs, avoiding false positives. It has been reported in the past literature that the artificial intelligence-based (AI) algorithm produces more precise diagnosis findings in the clinical setting and makes it easier to handle trauma patients quickly and effectively [1418]. As misdiagnosis can happen in a busy emergency department, real-time advice from the AI-based system can assist front-line residents and junior doctors when treating several trauma patients. It will also be very helpful in developing nations’ peripheral health systems, where there are insufficient radiologists or doctors with specialized training. Various studies done in the past reported different diagnostic accuracy of AI-based algorithms in automated detection of neck of the femur fracture on plain radiographs [1923]. However, to the best of our knowledge, we have not found any systematic review on the role of artificial intelligence in detecting neck of femur (NOF) fractures only on plain radiographs. This systematic review aims to collect and evaluate relevant clinical research from the literature, examine the diagnostic accuracy of AI-based algorithms in identifying neck of femur (NOF) fractures on plain radiographs, and provide some recommendations on the dependability of diagnostic accuracy in this area.

Methods

A systematic review and meta-analysis of the available literature on artificial intelligence in the neck of femur (NOF) fracture detection on a plain radiograph was performed. We registered the review protocol in PROSPERO International prospective register of systematic reviews (registration number CRD42022375449).

Literature Search and Inclusion Criteria

English-language articles with any publication date were taken into consideration. All original studies—descriptive, analytical, or clinical—that provided sufficient and calculable data to assess the usefulness or diagnostic accuracy of an AI-based algorithm in detecting the neck of the femur on a plain X-ray were eligible for inclusion in this study. Studies with comparable study objectives were included; however, there were no constraints on the dataset size, training method, or artificial intelligence type. This study excluded all review articles, systematic reviews, photographic reviews, narrative reviews, scoping reviews, multimedia files (online videos, podcasts), case reports, editorials, commentaries, and conference abstracts on the use of artificial intelligence in detecting fractures in the neck of the femur on plain radiographs. Articles without a clear reference standard, detailed subgroup reporting, or those referring to robotics or natural language processing (NLP) rather than image analysis were also removed to assess if an appropriate cohort of fracture neck femurs was analyzed. It also excluded articles on the veterinary population and related topics that included confounding factors like malignancy or diseases related to the bones.

PubMed, Web of science, Scopus, IEEE, and the Science direct databases were searched for eligible articles published on any date in the English language only. The search in this systematic review was carried out in July 2022 and included all the articles published up to 30 July 2023. Using a search technique that included phrases and word variations related to “Artificial Intelligence” [Mesh] AND “Femoral neck fractures” [Mesh] AND “Radiography” [Mesh] articles were obtained from various databases (“Artificial Intelligence” “Femoral neck fractures” “Radiography” keywords were used to search Science Direct, Scopus, and similar keywords were used to search Web of science). Afterward, those articles that had the neck of femur fracture exclusively in the study were selected manually to avoid missing any relevant study.

In addition, a manual search for references that might be connected was done. As advised by the cochrane collaboration, two of us independently assessed all possibly relevant studies’ titles, abstracts, and complete texts. The third and fourth reviewers decided on any disputes. According to the previously established inclusion and exclusion criteria, we evaluated the full-text articles of the remaining research and then chose the publications that qualified. The authors, organizations, and publications were not hidden from the reviewers. Six reviewers— two from the medical field and four with backgrounds in engineering—had finished the screening and selecting studies based on their titles and abstracts. Examining acceptable studies’ abstracts led to retrieving their complete texts from online databases and other relevant journals. Any disagreements among the reviewers were resolved through meetings by consensus (offline or online whatever is possible).

Data Extraction and Outcome Assessments

The following information was extracted from the included articles: authors, publication year, study duration, image type, fracture type, number of patients or images used, sociodemographic data, fracture classification (if any), diagnostic reference/ground truth if mentioned, and augments of each study. Furthermore, the names of the AIs, CNN architecture types, and data input proportions in training, validation, testing, and diagnosis and classification in the investigations were all derived from the published articles. A brief description of the different convolutional neural network (CNN) and dense neural network (DNN) used by the researcher for detecting neck of the femur fracture on plain radiographs is presented in Fig. 1.

Fig. 1.

Fig. 1

Two common learning methods have been adopted by the various researchers for detecting Neck of Femur (NOF) fracture are transfer learning and learning from the scratch. In transfer learning-based schemes, researchers have used existing and well known CNN and DNN architectures. In learning from scratch scheme, researchers have built their own CNN and DNN models for NOF fracture detection. a Describes the architecture of AlexNet that consist of eight layers. AlexNet is a sequential model. b Describes the basic architecture of a Generative Adversarial Network (GAN). GAN works in a bi-folded manner, initially GAN uses a feature extractor part of extract features and subsequently it uses a discriminator part for the NOF classification task. c Describes the inception module of the GoogLeNet which consist of 22 layers. The first inception model is a naïve version whereas the second module is the dimension reduction–based inception module which narrow downs the quantity of trainable parameters. d Describes an architecture of a custom built sequential model that consist of five blocks followed by a fully connected (FC) layer, SoftMax layer, and a Classification layer. In this custom model, each block consists 1 Max Pooling layer which follows 1 Convolutional layer with Batch Normalization and ReLU activation function. e Describes the basic architecture of a ResNet-18 model which consist of 18 layers. The ResNet-18 uses residual block in place of sequential layers to solve the degradation problem

Sensitivity, specificity, AUC (area Under the ROC curve), and accuracy were decided to be the primary outcomes along with Odds (from meta-analysis) and were extracted from every study. Youden index, positive and negative likelihood ratios were also calculated from the available information from each study. A random effects model was generated using the forest plot in the meta-analysis, providing the individual and overall Odds of AI detecting the correct diagnosis and the I2 value to assess the heterogeneity among the selected studies for meta-analysis.

Data Analysis

The software R version 4.2.1 was used to perform the meta-analysis, and the spider/radar chart showing CLAIM checklist scores was made using the Wondershare Edrawmax app.

Risk-of-Bias Assessment

The quality of each study was evaluated using a mix of the CLAIM (the Checklist for AI in Medical Imaging) and QUADAS-2 (A Revised Tool for the Quality Assessment of Diagnostic Accuracy Studies) criteria [24, 25]. The CLAIM tool "Checklist for Artificial Intelligence in Medical Imaging Reporting Adherence," was used to assess the reporting quality and adherence of studies related to artificial intelligence in medical imaging. While QUADAS is a tool specifically designed for assessing the risk a bias and methodological quality of studies that evaluate the accuracy of diagnostic tests. It helps reviewers evaluate the validity and applicability of diagnostic accuracy studies by assessing factors like patient selection, index test, reference standard, and flow and timing. Two of the reviewers/researchers did the ROB assessment independently and their common consensus was presented in the manuscript graphically using MS Excel and Word charts.

Result:

Study Selection, Characteristics, and Quality Assessment

We conducted a systematic review with meta-analysis of the available literature on artificial intelligence in the neck of femur (NOF) fracture on a plain radiograph following the PRISMA (preferred reported items for systematic reviews and meta-analysis) statement guidelines.

We obtained 437 studies from five databases, 344 of which were deleted due to duplication (confirmed manually by another two reviewers) and veterinary, review papers and dentistry investigations. Again, 44 studies were removed, because they were based on CT, MRI, malignancy, osteoporosis, or pathological diseases other than fracture. For eligibility, the remaining 39 studies were examined, and 34 with fractures other than the neck femur were removed, resulting in a final total of five (Fig. 2). Therefore, a total of five studies were used in the systematic review, while 3 out of 5 were included in the meta-analysis to determine the random effects model (Fig. 3). The comparisons of different CNN/DNN model used for NOF fracture detection from all the studies included in the paper is shown in Table 1.

Fig. 2.

Fig. 2

The PRISMA flow chart of the study selection process

Fig. 3.

Fig. 3

Forest plot depicting OR of different AI systems in the included studies (Comparing residents/physicians with AI)

Table 1.

Comparisons of CNN/DNN Model used for NOF fracture detection from studies that were included (N = 5)

S. no. References Model Layers Model Type Filter Sizes Advantage Model Complexity
1 Bae et al. [23] ResNet- 18 18 Residual 3 × 3 Convolutional Most efficient deep CNN with very low error rate Moderately complex
2 Mu et al. [20] RCNN + ResNet- 50 50 +  Ensemble

1 × 1 Convolutional

3 × 3 Convolutional

Region-based feature extractor, deep network, multi-scale and multi-level feature extraction for ensemble learning Very complex
3 Mutasa et al. [22] GAN 4 Sequential None Simple model with efficient in feature extraction and good for handling missing data Simple
Custom Residual model 21 Residual + Convolutional + Inception 3 × 3 Convolutional Sequential model with convolutional, residual, and inception blocks with spatial transform layer which stabilizes the gradients Moderately complex
4 Beyaz et al. [21] Custom model 8 Sequential 3 × 3 Convolutional Simple sequential model with ReLU and Batch Normalization, easy to train Very simple
5 Adams et al. [19] AlexNet 8 Sequential

11 × 11 Convolutional

5 × 5 Convolutional

3 × 3 Convolutional

Simple sequential model, easy to train Very simple
GoogLe Net 22 Inception

1 × 1 Convolutional

5 × 5 Convolutional

3 × 3 Convolutional

Less parameters so faster on training Moderately complex

The researchers have used sequential, convolutional, residual, and inception module-based networks in NOF detection. A total of seven CNN models have been used for NOF detection, of which five are existing pre-trained models and two are custom-built models trained from scratch. The brief objectives and conclusions of all the selected five studies that use AI for NOF detection are presented in Table 2. A detailed comparison of the statistical parameters of different studies is presented in Table 3. The details of dataset size, data input partition, and features are presented in Table 4. It was found that the pooled accuracy of AI systems for all the included studies in diagnosing the fracture neck femur was 87.28%, with the highest accuracy of 98.3% in the Bae et al. study [23]. A pooled Youden index (YI) was 0.73 for all the studies, excluding the Adams et al. study. [19] It was highest (0.96) in Bae et al. merged data set performance. The average positive likelihood ratio (PLR) for all the studies, excluding Adams et al. [19], was 19.88, whereas the negative likelihood ratio (NLR) was 0.17. The highest PLR (74.84) was in Bae et al. merged data set [23]. The model-assisted chief physician's performance was also high (33.13). A higher positive likelihood ratio tells us that the likelihood of fracture is higher. Regarding sensitivity, our study estimates a pooled sensitivity of AIs in diagnosing the fracture NOF to be 85% and a specificity of 87%. It was highest in the study by Bae et al. [23] (97% sensitivity and 98% specificity). Figure 3 shows a Forest plot depicting OR of different AI systems in the included studies (comparing residents/physicians with AI), including three studies, with Mu et al. (2021) having two components where they compared residents and chief physicians with AI. [20] It was found that three out of the four studies had an odds ratio of > 1 (Mu et al. residents having maximum odds of 1.81(0.87–3.73) in diagnosing the fractures when compared to those unaided). The random effects model showed an odds of 1.16 (0.84–1.61). Mutasa (2020) 22 and Bae et al. studies were not considered here as they were not comparing the AI system with those of human diagnosis. The overall heterogeneity of the studies was marginal (I2 = 51%), because there was a different methodology in all the individual studies, and none were a proper randomized controlled trial.

Table 2.

Study characteristics, objectives and conclusions drawn (N = 5)

Study Year Centers Objectives Conclusion
Bae et al. 2021 Multicentric 2 centers To develop a convolutional neural network (CNN) for detection of NOF fracture with datasets including non- displacement and displacement cases and to perform external validation for this CNN technique It can be said that a CNN algorithm for detection of both displaced and non-displaced NOF fracture using plain X-rays can be used
Mu et al. 2021 Multicentric 3 centers To detect and locate femoral neck fractures on radiographs by using a DL algorithm Femoral neck fractures on pelvic radiographs can be detected and located with high sensitivity and specificity by the object detection DCNN that has been trained with a limited dataset
Mutasa et al. 2020 Single To use deep learning with advanced data augmentation techniques to diagnose and classify femoral neck fractures A deep learning CNN method on AP radiographs, may reasonably predict the existence of femoral neck fractures. It can also learn to distinguish between displaced Garden III/IV fractures and non-displaced Garden I/II fractures
Beyaz et al. 2020 Single To detect femoral neck fracture on pelvic x rays using deep learning techniques (using GA) Accuracy and specificity increased
Adams et al. 2018 Single The purpose of the study is to evaluate the accuracy of DCNNs for detecting NOF fractures on radiographs, in comparison with perceptual training in medically-naïve individuals when trained on the same set of radiographic images Top-performing medically naive individuals can learn to recognise fractures just as well as a DCNN can with less than an hour of perceptual training

Table 3.

Diagnostic accuracy of test results from studies that were included (N = 5)

References Scheme Sensitivity (95% CI) Specificity (95% CI) Accuracy AUC YI Youden’s Index Positive likelihood ratio Negative likelihood ratio
Bae et al. [23] Merged data set 0.973 (0.96–0.98) 0.98 (0.98–0.99) 0.98 0.98 0.96 74.84 0.02
Mu et al. [20] Residents 0.82 (0.81–0.83) 0.89 (0.88–0.89) 0.85 NA 0.71 7.54 0.19
Model assisted 0.92 (0.92–0.93) 0.91 (0.90–0.91) 0.91 0.83 10.39 0.07
Chief physician 0.93 (0.93–0.94) 0.96 (0.96–0.97) 0.95 0.90 27.20 0.06
Model assisted 0.94 (0.94–0.95) 0.97 (0.96–0.97) 0.95 0.91 33.13 0.05
Mutasa et al. [22] Garden I & II 0.54 (0.45–0.62) 0.93 (0.90–0.95) 0.80 NA 0.47 7.71 0.49
Garden III & IV 0.91 (0.88–0.93) 0.83 (0.78–0.87) 0.86 0.74 5.35 0.10
Beyaz et al. [21] Without GA 0.82 (0.80–0.84) 0.69 (0.65–0.72) 0.77 NA 0.51 2.68 0.21
With GA 0.82 (0.80–0.84) 0.72 (0.69–0.76) 0.79 0.55 3.05 0.23
Adams et al. [19] GoogLeNet NA NA 0.90 0.98 NA NA NA
AlexNet NA NA 0.85 0.94 NA NA NA
Pooled variables 0.85 0.87 0.87 0.95 0.73 19.88 0.17

Youden Index is a measure of the overall effectiveness of a diagnostic test. It considers both sensitivity and specificity to determine how well a test can discriminate between positive and negative cases

Positive Likelihood Ratio is a measure of how much the odds of a positive test result increase the odds of having the condition of interest

Negative Likelihood Ratio measures how much the odds of a negative test result decrease the odds of having the condition

Table 4.

The features, characteristics and dataset size of the selected studies (N = 5)

References No. of images Fracture & non- fracture Data input partition for training, validation, and testing of AI Model Reference/ground truth Type of AI model used Pre-training of AI Model Used Final comparison done between
Bae et al. [23] 4189 images

1109 Fracture

3080 No Fracture

Training = 80%

Validation = 10%

Test = 10%

Two emergency medicine specialists ResNet-18 with convolutional block attention module (CBAM) +  +  Yes

External validation of deep learning module

No comparison with physician

Mu et al. [20] 1491 images from 1491 patients

1008 Fracture

483 No Fracture

Training = 41%

Validation = 6.7%

Internal Test 1 = 12.7%

Internal Test 2 = 15.8%

External Test 1 = 12.7%

External Test 2 = 11.2%

Two radiologists having 13 years and 15 years of experiences respectively Digital radiography fracture detection system [DR-FDS] with RCNN and ResNet-50 Yes 8 Residents (experience 1–3 years) and 8 attending physicians (experience 5–10 years) with model assisted and unassisted
Mutasa et al. [22] 1063 images from 550 patients

127 Garden I/II Fracture

610 Garden III/IV Fracture

326 No fracture

Training + validation = 90%

Test = 10%

Musculoskeletal fellowship trained radiologists

Generative Adversarial Network (GAN) and Custom Build Residual Network with 21

Layers

Yes Testing AI Based Algorithm No comparison with physician
Beyaz et al. [21] 234 images from 65patients

149 Fracture

85 Non fracture

Not Stated Not Stated

CNN with Genetic

Algorithm-based Optimization

No Deep learning assisted with GA and without GA
Adams et al. [19] 805

403 Fracture

402 Non fracture

Training = 57%

Validation = 29%

Test = 14%

142 undergraduate students for the detection Two DCNNs (AlexNet and GoogLeNet) Yes

Comparisons of two DCNNs (AlexNet and GoogLeNet) with top-performing

medically-naive humans

Risk-of-Bias Assessment

The risk of bias and quality assessment of the selected studies was done using QUADAS and CLAIM tools [24, 25]. The QUADAS tool for risk-of-bias assessment shows that 4 out of 5 selected studies had a ‘low’ risk of bias in ‘patient selection’ and ‘flow & timing domains’ (Fig. 4A).

Fig. 4.

Fig. 4

A Methodological quality assessment of the included studies using the QUADAS-2 tool for Risk of bias. B Methodological quality assessment of the included studies using the QUADAS-2 tool for applicability concerns. C CLAIM checklist scores (percentage)

Three out of 5 studies had a low risk of bias in the domain of ‘reference standard,’ while all the studies were unclear in the ‘index test’ domain. When’ concerns regarding applicability’ was assessed, all the studies had low risks in patient selection and reference standard. In contrast, only one study, i.e., Mutasa et al. had high risk in the domain of the ‘index test’ (Fig. 4B).

The spider/radar chart shows the percentage score obtained by the studies according to the (CLAIM) (Fig. 4C). All the included studies were observed to have an overall > 70% score, with Beyaz et al. finding a maximum of 90.4% score, followed by Bae et al. (88.09%) [21, 23]. Mutasa et al. study secured least scores, i.e., 73.8% [22].

Discussion

Principal Findings

Publications relating to AI-based decision support algorithms or computer-aided diagnosis for detecting NOF fractures have been described in literature. Five studies included in the present systematic review evaluated the accuracy of different AI algorithms which were trained and tested with images. Mu et al. reported a deep convolutional network-based femoral neck fracture detection system on radiographs [20]. They found a statistically significant increase in accuracy 0.95%, sensitivity 0.94 (0.94–0.95), and specificity 0.97 (0.96–0.97) in cases of model-assisted performance of the chief physician comparable to the chief physician alone diagnostic accuracy of 0.95%, sensitivity 0.93 (0.93–0.94) and specificity 0.96 (0.96–0.97). This study found a similar improvement in performance metrics in the resident group. This finding suggests AI model-assisted clinicians were significantly more sensitive and specific in detecting fractures than were unassisted clinicians, and the clinicians’ diagnostic ability improved with the assistance of the AI-based fracture detection system. Adams et al. evaluates the accuracy of 2 deep convolutional neural networks (AlexNet and GoogLeNet) to perceptual training in medically naive individuals for detecting neck of femur (NOF) fractures on radiographs [19]. CNN GoogLeNet outperformed CNN AlexNet with higher accuracy. This study was a follow-up to one by Chen et al. [26] that trained 142 medically naive undergraduate students in perceptual training for the detection of NOF fractures and reported that with 48 h and 52 min of training, respectively, top-performing medically illiterate students were able to detect neck of femur fractures as accurately as radiology residents and board-certified radiology consultants [26]. Adams et al., utilizing the same database as used by a study by Chen et al., found that top-performing medically-naıve humans can achieve similar learning with less than an hour of perceptual training with pre-trained DCNN [19].This study proposes a novel application of AI-based tools in human training in the medical sector to improve doctors’ diagnostic accuracy. Mutasa et al. reported a deep learning algorithm for accurately diagnosing and classifying femoral neck fractures [22]. They found in their study that the average AUC was 0.96, and displaced fractures were more accurately detected than undisplaced or minimally displaced NOFs. The difficulty in distinguishing undisplaced/minimally displaced NOF fracture from normal hip radiographs was expected, as this diagnosis can be difficult even for a radiologist. The study’s relevance was developing an AI-based automated tool that would prioritize positive femoral neck fracture cases on X-ray and provide a second opinion to the interpreting clinician, resulting in a valuable addition to the emergency radiology workflow. Beyaz et al. used deep learning and a genetic algorithm (GA) to detect NOF fractures in radiographs and discovered an overall accuracy of 77.7%, which increased by 1.6% when the GA was included [21]. The rate of fractured bone detection was 83%. Bae et al., in their study on 4189 images with an X-ray as the first screening tool, found a deep learning system that improved the capability to detect NOF fractures [23]. They conducted external validation of a deep learning algorithm for detecting femoral neck fractures. Although the sensitivity of easily displaced fractures (1.00) was higher than that of difficult non-displaced fractures (0.93), the latter had a lower sensitivity. In addition, they stated that the finished model developed using a dataset from one hospital may be exported to other hospitals and utilized as a screening tool for NOF fractures.

Comparison with Other Studies

Cha et al. [27], in their systematic review, observed an overall AUC of 0.96, while it was similar in our study, too (0.95) [27]. Accuracy was found to be 0.91 in their study, while it was 0.87 in our systematic review. The accuracy of fracture detection reported in the systematic review by David et al. (2019) ranged from 83 to 98% involving X-rays as well as CT scans which are comparable to our study [17]. Regarding sensitivity, our study estimates a pooled sensitivity of AIs in diagnosing the fracture NOF to be 85% and a specificity of 88%. It was highest in the study by Bae et al. (97% Sensitivity and 98% Specificity) [23]. In a Systematic review of “Artificial intelligence for radiological paediatric fracture assessment” by Shelmerdine et al. (2022), the algorithms tested on the test dataset achieved sensitivities of 88.9–90.7%, with a specificity of 90.9–100%, which is higher than our study, probably because our study focuses on a specific (NOF) fracture thus making it more challenging to diagnose [28]. Although in a systematic review of the diagnostic accuracy of deep learning in orthopaedic fractures by Yang et al. (2020), the pooled sensitivity and specificity, including 17 trials, were 0.87 and 0.91, respectively, which is also quite similar to our study [29]. This study’s AUC was comparable to our study’s (0.95). Another systematic review by Kuo et al. (2022) on AI in fracture detection shows similar results where pooled sensitivity and specificity for AI in the diagnosis of fracture ranged from 91 to 92%. [14].

Strength of Study

To the best of our knowledge, this is the first systematic review and meta-analysis to assess the diagnostic accuracy of artificial intelligence-based algorithms in identifying neck of femur (NOF) fractures on radiographs. The present study shows that AI-based algorithms had a high pooled accuracy of 87.28%, sensitivity of 85%, and specificity of 87% when detecting neck of femur fractures on X-rays. According to the random effects model, the AI system had odds of 1.16 (0.84–1.61) compared to human diagnosis. This systematic review and meta-analysis illustrate the clinical applicability and diagnostic accuracy of these algorithms, which has the potential to detect the NOF fractures on plain X rays. The findings of this study are a valuable contribution to the current literature and can help inform clinical decision-making and improve patient outcomes.

Limitation of this Study

Like every other study, this study has some limitations as we have not involved the quality and degree of AI-based training for the radiograph images. Since only radiographs are the inputs, it lacks patient history and clinical covariates. Further, the inclusion of these parameters will progress toward the improvement of substantial diagnostic accuracy. All the considered publications included in this study used a doctor, radiologist, or sophisticated imaging (such as CT images) to derive the ground truth labels (the AI reference standard) for training AI algorithms. Therefore, there is a possibility that these ground truth labels are vulnerable to human mistakes, which can lead to errors and incorrect dataset interpretations. It is always desirable that a better ground truth label (containing clinical co-variates and operational findings) may be generated for devising effective AI systems in real-time applications. The publications in this review should have included information on economic cost analyses or functional effect assessments, because the widespread adoption of AI applications as a diagnostic adjunct in healthcare systems would require significant funding and investment.

Implications for Practice and Future Research

We anticipated that our findings would help make decisions about using AI algorithms to diagnose neck of femur fractures on radiographs. To frame an efficient AI-based system that can achieve accuracy comparable to an experienced medical practitioner requires a large volume of training database. With this extensive training and emerging AI-based tools, one might achieve the desired accuracy. However, the rapid advancement of AI technology and its ability to learn quickly can assist it in reaching the level of a specialist. More research, however, is required to determine whether AI technology can accurately detect fractures in real-time clinical settings.

Conclusion

This review summarizes the current evidence on using AI applications to detect fractures in the femoral neck region on X-rays. AI-based algorithms can detect fractures in the neck of the femur, suggesting their application as a diagnostic adjunct benefitting clinicians since they will take less time and effort in fracture diagnosis. These AI-based tools can also train clinicians to make diagnoses with greater accuracy.

Supplementary Information

Below is the link to the electronic supplementary material.

Funding

None.

Declarations

Conflict of Interest

All authors have completed the ICMJE uniform disclosure form at www.icmje.org/disclosure-of-interest/ and declare: no support from any organization for the submitted work; no financial relationships with any organizations that might have an interest in the submitted work in the previous three years; no other relationships or activities that could appear to have influenced the submitted work.

Ethical Approval

Not required.

Informed Consent

For this type of study informed consent is not required.

Footnotes

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Augat P, Bliven E, Hackl S. Biomechanics of femoral neck fractures and implications for fixation. Journal of Orthopaedic Trauma. 2019;33(Suppl 1):S27–S32. doi: 10.1097/BOT.0000000000001365. [DOI] [PubMed] [Google Scholar]
  • 2.Gullberg B, Johnell O, Kanis JA. World-wide projections for hip fracture. Osteoporosis International. 1997;7(5):407–413. doi: 10.1007/pl00004148. [DOI] [PubMed] [Google Scholar]
  • 3.Pinto A, Berritto D, Russo A, Riccitiello F, Caruso M, Belfiore MP, Papapietro VR, Carotti M, Pinto F, Giovagnoni A, Romano L. Traumatic fractures in adults: missed diagnosis on plain radiographs in the emergency department. Acta Biomed. 2018;89(1-S):111–123. doi: 10.23750/abm.v89i1-S.7015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Guly HR. Diagnostic errors in an accident and emergency department. Emergency Medicine Journal. 2001;18(4):263–269. doi: 10.1136/emj.18.4.263. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Miele V, Galluzzo M, Trinci M. Missed fractures in the emergency department. InErrors in radiology. Milano: Springer; 2012. pp. 39–50. [Google Scholar]
  • 6.Brady AP. Error and discrepancy in radiology: inevitable or avoidable? Insights Imaging. 2017;8(1):171–182. doi: 10.1007/s13244-016-0534-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Waite S, Scott J, Gale B, Fuchs T, Kolla S, Reede D. Interpretive error in radiology. AJR. American Journal of Roentgenology. 2017;208(4):739–749. doi: 10.2214/AJR.16.16963. [DOI] [PubMed] [Google Scholar]
  • 8.Degnan AJ, Ghobadi EH, Hardy P, Krupinski E, Scali EP, Stratchko L, Ulano A, Walker E, Wasnik AP, Auffermann WF. Perceptual and interpretive error in diagnostic radiology-causes and potential solutions. Academic Radiology. 2019;26(6):833–845. doi: 10.1016/j.acra.2018.11.006. [DOI] [PubMed] [Google Scholar]
  • 9.Tang X. The role of artificial intelligence in medical imaging research. BJR Open. 2019;2(1):20190031. doi: 10.1259/bjro.20190031. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Aggarwal R, Sounderajah V, Martin G, Ting DS, Karthikesalingam A, King D, Ashrafian H, Darzi A. Diagnostic accuracy of deep learning in medical imaging: A systematic review and meta-analysis. NPJ Digital Medicine. 2021;4(1):1–23. doi: 10.1038/s41746-021-00438-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Pesapane F, Codari M, Sardanelli F. Artificial intelligence in medical imaging: Threat or opportunity? Radiologists again at the forefront of innovation in medicine. European Radiology Experimental. 2018;2:35. doi: 10.1186/s41747-018-0061-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Puttagunta M, Ravi S. Medical image analysis based on deep learning approach. Multimed Tools Appl. 2021;80(16):24365–24398. doi: 10.1007/s11042-021-10707-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Fourcade A, Khonsari RH. Deep learning in medical image analysis: A third eye for doctors. J Stomatol Oral Maxillofac Surg. 2019;120(4):279–288. doi: 10.1016/j.jormas.2019.06.002. [DOI] [PubMed] [Google Scholar]
  • 14.Kuo RYL, Harrison C, Curran TA, Jones B, Freethy A, Cussons D, Stewart M, Collins GS, Furniss D. Artificial intelligence in fracture detection: A systematic review and meta-analysis. Radiology. 2022;304(1):50–62. doi: 10.1148/radiol.211785. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Groot OQ, Bongers MER, Ogink PT, Senders JT, Karhade AV, Bramer JAM, Verlaan JJ, Schwab JH. Does artificial intelligence outperform natural intelligence in interpreting musculoskeletal radiological studies? A systematic review. Clinical Orthopaedics and Related Research. 2020;478(12):2751–2764. doi: 10.1097/CORR.0000000000001360. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Zhang X, Yang Y, Shen YW, Zhang KR, Jiang ZK, Ma LT, Ding C, Wang BY, Meng Y, Liu H. Diagnostic accuracy and potential covariates of artificial intelligence for diagnosing orthopedic fractures: A systematic literature review and meta-analysis. European Radiology. 2022;32(10):7196–7216. doi: 10.1007/s00330-022-08956-4. [DOI] [PubMed] [Google Scholar]
  • 17.Langerhuizen DWG, Janssen SJ, Mallee WH, van den Bekerom MPJ, Ring D, Kerkhoffs GMMJ, Jaarsma RL, Doornberg JN. What are the applications and limitations of artificial intelligence for fracture detection and classification in orthopaedic trauma imaging? A systematic review. Clinical Orthopaedics and Related Research. 2019;477(11):2482–2491. doi: 10.1097/CORR.0000000000000848. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Olczak J, Fahlberg N, Maki A, Razavian AS, Jilert A, Stark A, Sköldenberg O, Gordon M. Artificial intelligence for analyzing orthopedic trauma radiographs. Acta Orthop. 2017;88(6):581–586. doi: 10.1080/17453674.2017.1344459. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Adams M, Chen W, Holcdorf D, McCusker MW, Howe PD, Gaillard F. Computer vs human: Deep learning versus perceptual training for the detection of neck of femur fractures. Journal of medical imaging and radiation oncology. 2019;63(1):27–32. doi: 10.1111/1754-9485.12828. [DOI] [PubMed] [Google Scholar]
  • 20.Mu L, Qu T, Dong D, Li X, Pei Y, Wang Y, Shi G, Li Y, He F, Zhang H. Fine-tuned deep convolutional networks for the detection of femoral neck fractures on pelvic radiographs: A multicenter dataset validation. IEEE Access. 2021;24(9):78495–78503. doi: 10.1109/ACCESS.2021.3082952. [DOI] [Google Scholar]
  • 21.Beyaz S, Açıcı K, Sümer E. Femoral neck fracture detection in X-ray images using deep learning and genetic algorithm approaches. Jt Dis Relat Surg. 2020;31(2):175–183. doi: 10.5606/ehc.2020.72163. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Mutasa S, Varada S, Goel A, Wong TT, Rasiej MJ. Advanced deep learning techniques applied to automated femoral neck fracture detection and classification. Journal of Digital Imaging. 2020;33(5):1209–1217. doi: 10.1007/s10278-020-00364-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Bae J, Yu S, Oh J, Kim TH, Chung JH, Byun H, Yoon MS, Ahn C, Lee DK. External validation of deep learning algorithm for detecting and visualizing femoral neck fracture including displaced and non-displaced fracture on plain X-ray. J Digit Imaging. 2021;34(5):1099–1109. doi: 10.1007/s10278-021-00499-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.QUADAS-C tool | Cochrane Methods [Internet]. [cited 2023 Jan 1]. Available from: https://methods.cochrane.org/methods-cochrane/quadas-c-tool
  • 25.Mongan J, Moy L, Kahn CE., Jr Checklist for artificial intelligence in medical imaging (CLAIM): A guide for authors and reviewers. Radiol Artif Intell. 2020;2(2):e200029. doi: 10.1148/ryai.2020200029. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Chen W, HolcDorf D, McCusker MW, Gaillard F, Howe PDL. Perceptual training to improve hip fracture identification in conventional radiographs. PLoS ONE. 2017;12(12):e0189192. doi: 10.1371/journal.pone.0189192. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Cha Y, Kim JT, Park CH, Kim JW, Lee SY, Yoo JI. Artificial intelligence and machine learning on diagnosis and classification of hip fracture: Systematic review. Journal of Orthopaedic Surgery and Research. 2022;17(1):520. doi: 10.1186/s13018-022-03408-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Shelmerdine SC, White RD, Liu H, Arthurs OJ, Sebire NJ. Artificial intelligence for radiological paediatric fracture assessment: A systematic review. Insights Into Imaging. 2022;13(1):94. doi: 10.1186/s13244-022-01234-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Yang S, Yin B, Cao W, Feng C, Fan G, He S. Diagnostic accuracy of deep learning in orthopaedic fractures: A systematic review and meta-analysis. Clinical Radiology. 2020;75(9):713.e17–713.e28. doi: 10.1016/j.crad.2020.05.021. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials


Articles from Indian Journal of Orthopaedics are provided here courtesy of Indian Orthopaedic Association

RESOURCES