Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Jul 11;44(7):e70251. doi: 10.1002/jor.70251

Artificial Intelligence‐Based Automatic Screening System for Hip Dysplasia

Kazuki Miyama 1, Kenji Kitamura 1,, Masanori Fujii 2, Goro Motomura 1, Satoshi Hamai 1, Shinya Kawahara 1, Ryosuke Yamaguchi 1, Takeshi Utsunomiya 1, Soichiro Yoshino 1, Makoto Endo 1, Yasuharu Nakashima 1
PMCID: PMC13356558  PMID: 42436605

ABSTRACT

To develop an artificial intelligence‐based radiographic screening system for hip dysplasia and to evaluate its diagnostic performance and clinical utility. Seventy‐three patients with hip dysplasia who underwent unilateral periacetabular osteotomy and 29 asymptomatic volunteers were included. Radiographs from patients with hip dysplasia were divided into training (n = 40), validation (n = 7), and testing (n = 26) datasets, whereas radiographs from volunteers were used exclusively for testing. The artificial intelligence system detects and segments anatomical landmarks on pelvic radiographs, quantifies parameters reflecting anterior, posterior, and lateral acetabular coverage, and determines whether a patient has hip dysplasia based on these measurements. It comprises two deep learning models: DeepLabCut and U‐Net, which can effectively learn anatomical structures from relatively small training data. Ground‐truth measurements were obtained from repeated assessments by a hip surgeon and a non‐hip surgeon, and hip dysplasia status was independently determined by two hip surgeons in blinded assessments. Agreement with the ground‐truth was assessed using intraclass correlation coefficients, and screening performance was compared with that of general orthopaedic residents using the test images. Agreement with the ground‐truth was good to excellent [intraclass correlation coefficient: 0.70–0.95]. The system showed the best overall screening performance (sensitivity, 0.96 vs. 0.44–0.97; specificity, 0.92 vs. 0.16–0.92; area under the curve, 0.94 vs. 0.57–0.76; F‐measure, 0.96 vs. 0.60–0.81) compared with the residents. Our artificial intelligence‐based screening system may serve as a primary care‐oriented radiographic screening and referral‐support tool for clinicians, including non‐specialists, despite being trained on a limited dataset.

Keywords: artificial intelligence, deep learning, hip dysplasia, periacetabular osteotomy

1. Introduction

Hip dysplasia (HD) is typically characterized by a shallow acetabular concavity and reduced acetabular coverage of the femoral head [1, 2]. These morphological abnormalities increase joint contact pressure within the acetabulum and predispose patients to early‐onset hip osteoarthritis [3, 4]. Early diagnosis of HD allows timely intervention, such as periacetabular osteotomy (PAO), which can improve the prognosis [5], whereas delayed diagnosis may result in irreversible progression of osteoarthritis and the need for total hip arthroplasty (THA) at a relatively young age [6].

Radiographic assessment is essential for the early diagnosis of HD [1, 2, 3, 6, 7], as it is less expensive and more widely available than three‐dimensional (3D) computed tomography or magnetic resonance imaging evaluation. Traditionally, HD has been diagnosed based on lateral acetabular coverage measured on supine anteroposterior pelvic radiographs. Commonly used parameters to quantify lateral coverage include the lateral center‐edge angle (LCEA), Sharp angle, acetabular roof obliquity (ARO), and extrusion index (EI) [1, 8, 9]. Recent studies, however, have suggested that radiographs taken in the standing position are more appropriate for assessing hip deformities, because supine radiographs may not accurately represent the functional relationship between the acetabulum and the femur [10, 11]. Furthermore, anterior and posterior acetabular coverage are now recognized as key pathological features of HD [12, 13]. Thus, several studies have recommended the use of metrics that reflect both anterior [anterior wall index (AWI), anterior coverage (AC)] and posterior acetabular coverage [posterior wall index (PWI), posterior coverage (PC)] [1, 9, 13, 14, 15]. Incorporating these parameters into radiographic assessments has the potential to enhance the detection of HD.

Despite their value, radiographic assessments for HD in daily clinical practice have several limitations. First, they are time‐consuming. Evaluating only the lateral coverage may be quicker but is insufficient for accurate diagnosis. A more comprehensive radiographic evaluation that includes anterior and posterior coverage improves diagnostic accuracy but substantially increases the complexity of the assessment. Second, the clinicians who initially assess patients with suspected HD are often not hip specialists, which may lead to uncertainty regarding diagnostic thresholds and criteria. Third, radiographic measurements are technically demanding and subject to considerable intra‐ and inter‐observer variability, making reproducibility difficult to ensure [7, 8, 16]. These issues render early and accurate diagnosis challenging. Artificial intelligence (AI) offers a potential solution to these problems and AI‐based systems are already being introduced into clinical practice [17]. AI provides consistent measurements without intra‐observer variability and can rapidly process complex and large‐scale data. AI‐based tools could therefore be used to accurately and instantaneously measure radiographic parameters and perform comprehensive screening by integrating multiple measurements.

The purpose of this study was to develop an AI‐based automated screening system that assesses the presence of HD on standing pelvic radiographs, considering the lateral, anterior, and posterior acetabular coverage. We aimed to evaluate the reproducibility of this system and to clarify its clinical utility.

2. Methods

2.1. Study Design

This study was approved by the institutional review board (approval number 22089), and written informed consent was obtained from all participants. We included 82 patients with frank or borderline HD (LCEA < 25°, Sharp angle > 45°, or ARO > 15°) [18, 19] who had standing anteroposterior pelvic radiographs and underwent unilateral PAO at our institution between September 2016 and July 2022 and 30 asymptomatic volunteers (Figure 1) [20, 21]. Patients with advanced osteoarthritis (Tönnis grade ≥ 2) (n = 1) [22] or poor quality images (n = 8) were excluded. One volunteer was excluded because of a previous hip surgery (Figure 1). Ultimately, 73 patients with HD and 29 volunteers were included in this study (Table 1). The 73 radiographs from patients with HD were divided into training (n = 40), validation (n = 7), and testing (n = 26) datasets for development and evaluation of the AI system (Figure 1). The 29 radiographs from volunteers were used exclusively for testing the AI system.

Figure 1.

Figure 1

Overview of patient selection. HD, Hip dysplasia.

Table 1.

Demographic data.

Training (n = 40)
Age (Mean ± SD) 36.9 ± 10.2
Gender (M: F) 3:37
Validation (n  = 7)
Age (Mean ± SD) 36.0 ± 11.3
Gender (M: F) 0:7
Test (n  = 26)
Age (Mean ± SD) 40.2 ± 10.2
Gender (M: F) 0:26
HD total (n  = 73)
Age (Mean ± SD) 38.0 ± 10.3
Gender (M: F) 3:70
Volunteer (n  = 29)
Age (Mean ± SD) 34.6 ± 6.2
Gender (M: F) 3:26

Abbreviations: F; female, HD; hip dysplasia, M; male, SD; standard deviation.

2.2. Screening Pipeline of the AI System

Anatomical landmarks were annotated and the radiographic parameters were measured as previously described (Figure 2) [1, 14, 18, 23].

Figure 2.

Figure 2

Anatomical landmarks used to measure the radiographic parameters. Each symbol represents the following: C1,2, circle approximating the femoral head, AAW, anterior acetabular wall, PAW, posterior acetabular wall, P1,10, center of the femoral head, P2,3, lateral and medial edge of the acetabular sourcil, P4: lateral edge of C1, P5,6: teardrop, P7,8, medial and lateral margin of the intersection of the femoral neck axis and the AAW, P9, lateral margin of the intersection of the femoral neck axis and the PAW, P11, center of the femoral neck, R1,2, radii of C1 and C2, X, horizontal length of femoral head uncovered by the acetabulum, L1, line connecting P5 and P6, L2, line passing through the P1 and perpendicular to L1. Lateral center‐edge angle is the angle between lines L2 and P1P2. The sharp angle is the angle between lines L1 and P2P5. Acetabular roof obliquity is the angle between lines L1 and P2P3. Extrusion index is defined as X divided by 2×R1 (%). Anterior wall index is defined as P7P8 divided by R2. Anterior coverage is defined as the area of the anterior acetabular wall divided by the area of circle C2 (blue shaded area). Posterior wall index is defined as P7P9 divided by R2. Posterior coverage is defined as the area of the posterior acetabular wall divided by the area of circle C2 (yellow shaded area).

The screening pipeline of the AI system is illustrated in Figure 3. The system comprises three components: two detection models that identify anatomical landmarks, a segmentation model that delineates the anterior and posterior acetabular walls, and a measurement module that calculates eight radiographic parameters and outputs a screening result indicating whether the hip is dysplastic.

Figure 3.

Figure 3

Screening pipeline of the AI system. The AI system comprises three parts: two detection models that detect anatomical landmarks, a segmentation model that segments the anterior and posterior acetabular wall, and a measurement module that calculates eight radiographic parameters (LCEA, Sharp angle, ARO, EI, AWI, AC, PWI, and PC) based on the detected landmarks and segmented regions and outputs a screening result indicating whether the hip is dysplastic. AAW, anterior acetabular wall; AC, anterior coverage; AI, acetabular index; AI, artificial intelligence; ARO, acetabular roof obliquity; AWI, anterior wall index; EI, extrusion index; HD, hip dysplasia; LCEA, lateral center‐edge angle; P'1‐P'8, the surface of the upper and lower femoral head in contact with the acetabulum; P'11, the center of the femoral head; P'12, the center of the femoral neck; P'13,14, the lateral and medial edge of the acetabular sourcil; P'15, teardrop; P'9,10, lateral and medial cortices of the femoral neck; PAW, posterior acetabular wall; PC, posterior coverage; PWI, posterior wall index.

For image preprocessing (Figure 3 : Step 1), the original radiographs were first resized to 3000 × 2000 pixels and then divided into left and right halves (1500 × 2000 pixels). Images of left hips were horizontally flipped so that all hips were represented as right hips. This preprocessing pipeline was applied consistently to the training, validation, and test datasets. After preprocessing, the first detection model detected 10 anatomical landmarks (Figure 3 : Step 2). The centers of the femoral head (P'11) and neck (P'12) were then derived based on these landmarks (Figure 3 : Step 3). The image was subsequently cropped (894 × 894 pixels) and centered on P'11 (Figure 3 : Step 4). In this cropped image, the second detection model identified three anatomical landmarks (P'13‐15), and the segmentation model delineated the anterior and posterior acetabular walls (Figure 3 : Steps 5 and 6).

The measurement module automatically computed eight radiographic parameters (LCEA, Sharp angle, EI, ARO, AWI, AC, PWI, and PC) from the detected landmarks and segmented regions (Figure 3 : Step 7). Based on previous reporting thresholds [14, 15], the measurement module categorized each parameter as indicative of HD, normal morphology, or femoroacetabular impingement (Table 2 ). If the LCEA fell within the HD range, 3 points were assigned; for each of the other parameters classified as HD, 1 point (out of a possible 10 points) was assigned. When the total score exceeded 4 points, the measurement module classified the hip as dysplastic (Figure 3 : Step 7).

Table 2.

Cutoff values of each radiographic parameter.

LCEA Sharp ARO EI AWI AC PWI PC
HD < 25° > 43° > 14° > 27% < 0.3 < 14% < 0.81 < 35%
Normal 25°−33° 38°−43° 3°−14° 17%–27% 0.3–0.51 14%–26% 0.81–1.14 35%–47%
FAI > 33° < 38° < 3° < 17% > 0.51 > 26% > 1.14 > 47%

Abbreviations: AC, anterior coverage; ARO, acetabular roof obliquity; AWI, anterior wall index; EI, extrusion index; FAI, femoroacetabular impingement; HD, hip dysplasia; LCEA, lateral center‐edge angle; PC, posterior coverage; PWI, posterior wall index.

2.3. Training the Detection and Segmentation Models

The DeepLabCut deep learning framework [24] was used as the detection model in its standard transfer‐learning configuration with an ImageNet‐pretrained ResNet backbone. It can be trained to detect various target structures with relatively limited training data and has already been applied in human medicine [25, 26, 27]. Two orthopaedic surgeons (KM, a non‐hip surgeon, and KK, a hip surgeon) annotated all the landmarks on each image through consensus before training. Two DeepLabCut models were trained using these annotated images: the first to detect 10 landmarks (P'1‐P'10 in Figure 3 : Step 2), and the second to detect three landmarks (P'13‐P'15 in Figure 3 : Step 5).

The U‐Net deep learning architecture was used as the segmentation model [28]. Two orthopaedic surgeons annotated the anterior and posterior acetabular walls for each training image using software ( Labelme , MIT Computer Science and Artificial Intelligence Laboratory, Cambridge, MA) [29]. The U‐Net was trained for 1000 epochs to segment the anterior and posterior acetabular walls. The mean squared error loss was used as the loss function and the Adaptive Momentum (Adam) algorithm [30] was used as the optimizer. Early stopping [31] was employed to improve generalizability: training was halted when the validation loss failed to improve for 20 consecutive epochs.

2.4. Ground‐Truth for Radiographic Parameters and Screening

Two orthopaedic surgeons (KM, a non‐hip surgeon, and KK, a hip surgeon) annotated the five landmarks (the centers of the femoral head and femoral neck, the lateral and medial edges of the acetabular sourcil, and teardrop) and two regions (the anterior and posterior acetabular wall) twice on the 26 radiographs in the HD test set using Labelme software. Eight radiographic parameters were then calculated from these annotated landmarks and regions. Each surgeon (KM and KK) independently measured each radiographic parameter twice in a blinded manner, yielding four measurements per parameter. The mean of these four measurements was used as the ground‐truth (GT) because averaging repeated expert measurements was expected to reduce random measurement error while incorporating both observers rather than privileging a single reading. The intraclass correlation coefficients (ICCs) for the eight radiographic parameters between the two orthopaedic surgeons were excellent or good [32] (Table 3).

Table 3.

The measurement results of our system and ground‐truth values.

LCEA Sharp ARO EI AWI AC PWI PC
Our system 13.3 ± 6.7° 46.8 ± 3.1° 17.9 ± 5.0° 35.5 ± 7.4% 0.26 ± 0.10 10.6 ± 4.1% 0.82 ± 0.18 40.8 ± 7.7%
GT 13.5 ± 5.9° 46.5 ± 3.0° 18.9 ± 4.8° 35.4 ± 5.1% 0.26 ± 0.09 11.0 ± 4.4% 0.82 ± 0.18 40.9 ± 7.8%
Absolute error 1.6 ± 1.7° 0.7 ± 0.6° 1.5 ± 1.5° 1.8 ± 1.7% 0.05 ± 0.06 1.8 ± 1.8% 0.04 ± 0.06 2.2 ± 1.4%
Intraclass correlation coefficients between the system and GT 0.93 0.95 0.91 0.94 0.70 0.82 0.95 0.94
Intraclass correlation coefficients between non‐hip surgeon and hip surgeon 0.92 0.89 0.93 0.91 0.76 0.82 0.91 0.92
Intraclass correlation coefficients within the system 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00
Intra‐rater reliability of non‐hip surgeon 0.92 0.88 0.91 0.93 0.74 0.83 0.86 0.89
Intra‐rater reliability of hip surgeon 0.93 0.96 0.95 0.92 0.84 0.75 0.95 0.93

Values are presented as mean ± standard deviation.

Abbreviations: AC, anterior coverage; ARO, acetabular roof obliquity; AWI, anterior wall index; EI, extrusion index; GT, ground‐truth; LCEA, lateral center‐edge angle; PC, posterior coverage; PWI, posterior wall index.

Two hip surgeons (KK and MF) independently and blindly classified each hip as dysplastic or normal, using criteria based on previous reports [1, 14, 18, 23]. They evaluated 55 radiographs of 110 hips (26 patients with HD and 29 volunteers; Figure 1). In cases of disagreement, consensus was reached through discussion. These consensus classifications constituted the GT for HD status.

2.5. Evaluation of the Accuracy of Measuring Radiographic Measurements and Screening

The accuracy of the AI system in measuring radiographic parameters was evaluated using 26 test images from patients with HD. Performance was evaluated in four ways: the absolute error between the AI system and the GT, Bland‐Altman analysis of agreement, the ICC between the AI system and the GT, and the intra‐rater reliability of the AI system when each parameter was measured twice. ICCs were classified as follows [32]: ≥ 0.9, excellent; ≥ 0.75–0.89, good; ≥ 0.5–0.74, moderate; and < 0.50, poor. To explore error characteristics across the spectrum of acetabular undercoverage, absolute errors were also descriptively summarized among hips with GT LCEA < 25°, based on published LCEA ranges for borderline (20–25°), mild (15–20°), and moderate‐to‐severe (< 15°) dysplasia, and were grouped into greater undercoverage (LCEA < 15°) and lesser undercoverage (LCEA 15° to < 25°) [33].

The screening performance of the AI system was evaluated using 110 hips on 55 radiographs (26 patients with HD and 29 volunteers) to determine whether each hip was dysplastic or not. The screening performance was quantified using sensitivity, specificity, area under the curve (AUC), and F‐measure.

Three general orthopaedic residents (orthopaedic resident 1 with 4 years of experience, and orthopaedic residents 2 and 3 with 5 years of experience) performed screening using the same images used in the AI system. We compared the screening performance of the AI system with that of the residents. For formal comparison, 95% confidence intervals (CIs) for sensitivity and specificity were calculated with the Wilson method [34], and 95% CIs for AUC and F‐measure were estimated by patient‐level bootstrap resampling (2000 iterations) [35]. Because the AI system and each resident provided paired dichotomous classifications on the same hips, paired comparisons were performed using McNemar's test [36].

3. Results

3.1. Radiographic Measurement Performance

For the angular parameters (LCEA, Sharp angle, and ARO), all mean absolute errors between the AI system and GT were within 1.8° (0.7°−1.8°), and the ICCs between the AI system and GT were excellent (0.91‐0.95). For ratio‐based parameters (EI, AWI, AC, PWI, and PC), mean absolute errors were within 5.0%, and the ICCs between the AI system and GT were good to excellent (0.70‐0.95). The intra‐observer ICCs of the AI were 1.00 for all parameters, indicating perfect test–retest reliability. Bland–Altman analysis for LCEA demonstrated a small mean difference between the AI system and the GT (− 0.19°), with 95% limits of agreement ranging from −4.80° to 4.41°, suggesting minimal systematic bias (Supplementary Figure S1). Among hips with GT LCEA < 25°, the mean absolute error remained low in those with greater undercoverage (LCEA < 15°; n = 31, mean absolute error 1.81°) and those with lesser undercoverage (LCEA 15° to < 25°; n = 20, mean absolute error 1.41°), with no clear increase in error observed in hips with greater undercoverage.

3.2. Screening Performance

The AI system receives a pelvic radiograph as input and outputs a screening result (Figure 4A : patient with HD, Figure 4B : volunteer). The AI system accurately distinguished patients with HD from individuals with normal hips, with only a few false‐negative and false‐positive results (Figure 5A). In contrast, residents 1 and 2 failed to detect HD in multiple cases, whereas resident 3 classified almost all hips as dysplastic, resulting in very low specificity (Figure 5B–D). The AI system demonstrated the best overall screening performance (sensitivity 0.96 [95% CI, 0.88–0.99], specificity 0.92 [95% CI, 0.79–0.97], AUC 0.94 [95% CI, 0.87–0.99], F‐measure 0.96 [95% CI, 0.91–0.99]) compared with the residents (sensitivity 0.44–0.97, specificity 0.16–0.92, AUC 0.57–0.76, F‐measure 0.60–0.81) (Table 4). Paired McNemar tests showed significantly fewer misclassifications by the AI system than by each resident (all p < 0.001).

Figure 4.

Figure 4

Representative screening results of the AI system in (A) a patient with HD and (B) a volunteer with a normal hip. Each red point on the radiograph represents a detected landmark, and all landmarks were detected accurately. Measured radiographic parameters are displayed in the upper left of each radiograph. The classification of each parameter (HD, normal, or FAI) is summarized in the accompanying table. The final screening result is shown above each radiograph. AC, anterior coverage; AI, artificial intelligence; ARO, acetabular roof obliquity; AWI, anterior wall index; EI, extrusion index; FAI, femoroacetabular impingement; HD, hip dysplasia; LCEA, lateral center‐edge angle; PC, posterior coverage; PWI, posterior wall index.

Figure 5.

Figure 5

Confusion matrix of the screening results of (A) the AI system and (B–D) the three residents. In each matrix, the vertical axis indicates the GT status, and the horizontal axis indicates the predicted status. AI, artificial intelligence; GT, ground‐truth; HD, hip dysplasia.

Table 4.

Screening performance.

Rater Sensitivity Specificity AUC F‐measure
AI system 0.96 (0.88–0.99) 0.92 (0.79–0.97) 0.94 (0.87–0.99) 0.96 (0.91–0.99)
Resident 1 0.76 (0.65–0.85) 0.76 (0.61–0.87) 0.76 (0.66–0.86) 0.81 (0.71–0.89)
Resident 2 0.44 (0.34–0.56) 0.92 (0.79–0.97) 0.68 (0.61–0.76) 0.60 (0.47–0.71)
Resident 3 0.97 (0.90–0.99) 0.16 (0.07–0.30) 0.57 (0.49–0.66) 0.80 (0.71–0.88)

Values are presented as point estimates (95% confidence intervals).

Abbreviations: AI, artificial intelligence; AUC, area under the curve; CI, confidence interval.

4. Discussion

In this study, we proposed a novel AI‐based screening system for HD. Our AI system accurately and instantaneously evaluated acetabular coverage not only laterally but also anteriorly and posteriorly on two‐dimensional standing pelvic radiographs, providing a comprehensive radiographic assessment. By integrating these measurements, the system screened for HD with markedly better performance than general orthopaedic residents.

The principal advantage of our AI‐based system is its ability to automatically assess lateral, anterior, and posterior acetabular coverage on standing pelvic radiographs. Historically, HD has been diagnosed by assessing only lateral acetabular coverage on supine radiographs [1, 8, 9]. However, recent studies have indicated that supine radiographs may not reflect the functional orientation of the acetabulum relative to the femur, and that standing radiographs are more appropriate for evaluating hip deformities in the abnormal mechanical environment of dysplastic hips, such as increased joint contact pressure and overload on the anterosuperior acetabulum [10, 37]. Moreover, anterior and posterior coverage are crucial determinants of HD [12, 13]. In particular, anterior coverage has been identified as one of the most influential factors affecting the natural history after PAO [12]. Therefore, our comprehensive evaluation, including anterior and posterior coverage on the standing radiograph, is likely to have contributed to the high diagnostic performance of the AI system.

The screening performance of our AI system for HD was better overall than that of the residents. Clinicians who first see patients with hip pain are often not experts in hip diseases, and our results suggest that even general orthopaedic residents may show variable screening performance, including both missed diagnoses and overcalling of HD. One resident showed very high sensitivity but very low specificity, which may reflect a sensitivity‐oriented screening bias favoring over‐referral rather than uniformly poor diagnostic performance. Delayed diagnosis increases the likelihood of THA at a young age [38]. In contrast, early diagnosis allows timely surgical intervention with PAO, which can favorably modify the natural course of HD [12]. If implemented in routine clinical practice, our AI system could support primary care‐oriented radiographic screening of hips with possible dysplasia and facilitate timely referral to hip surgeons, including in settings where the initial assessor is not a hip specialist. This would increase the proportion of patients undergoing PAO at an appropriate stage and may reduce the number of young patients requiring THA.

The AI system showed excellent inter‐rater reliability with GT (ICCs: 0.91–0.95), except for AWI and AC, which are related to anterior acetabular coverage. These inter‐rater observer reliability values were comparable to those reported in previous studies [8, 14, 39]. Additionally, the AI system demonstrated perfect reproducibility for all parameters involving those reflecting anterior coverage, whereas surgeons—regardless of their level of hip expertise—showed considerable measurement variability, as noted previously (ICC = 0.70–0.98) [8, 14, 39]. This highlights the robustness of the AI system, which consistently yields identical results when applied repeatedly to the same images.

The current AI system achieved higher accuracy and diagnostic performance with considerably less training data than previous AI‐based approaches [40, 41]. Li et al. reported an AI system that evaluates HD status by automatically measuring the Sharp angle [40], but its classification performance was inferior to that of our system. Yang et al. described an AI system that can automatically measure parameters related to lateral acetabular coverage only [41]. Their AI system demonstrated excellent or good agreement with clinicians (ICCs: 0.83–0.93), whereas our system showed even better agreement across all parameters (ICCs: 0.91–0.95). Two factors may account for the superior performance of our system. First, previous studies assessed only parameters related to lateral acetabular coverage and did not incorporate indices of anterior and posterior coverage. Second, by using DeepLabCut as the detection model, our system could be trained effectively with relatively small datasets [17, 25, 26, 27]. We were thus able to obtain clinically valuable results from only 40 training cases, compared with 742 and 9,248 training cases in previous studies [40, 41].

This study has several limitations. First, the sample size was modest, the investigation was conducted at a single center, most participants were female, and all radiographs were acquired using a standardized protocol within one healthcare system. Accordingly, the generalizability of the present system remains uncertain, and external multicenter validation in more diverse populations and imaging settings is required before clinical deployment. Second, although the test cohort included asymptomatic volunteers and some borderline hips, the study was not specifically designed to validate performance in asymptomatic or borderline populations or in hips with acetabular rim crossover or acetabular retroversion. Further validation in these populations is required. Third, the GT was derived from the mean of four expert measurements. Although this approach was intended to reduce random measurement error, it does not eliminate residual human variability; therefore, both model training and performance evaluation may still have been influenced by expert‐dependent measurement variability. Fourth, AWI and PWI are radiograph‐based indices referenced to the femoral neck axis and may therefore be affected by femoral rotation, abduction‐adduction, and neck‐shaft morphology, unlike CT‐based three‐dimensional assessments that use a fixed coordinate system. Fifth, segmentation of the anterior acetabular wall did not always perform optimally. In some hips with markedly deficient anterior coverage, accurate segmentation was challenging owing to a blurred outline of the anterior acetabular wall caused by overlap with the ligamentum teres of the femoral head. Segmentation performance for the anterior wall may be improved by incorporating more training data from multiple centers, particularly from cases with deficient anterior coverage. Further multicenter studies are warranted to refine the AI system and evaluate its clinical impact.

In conclusion, the proposed AI‐based screening system can automatically identify hips likely to have HD without requiring large training datasets. Its clinical implementation may support primary care‐oriented radiographic screening and timely referral to hip surgeons, particularly for clinicians including non‐specialists, thereby potentially reducing the number of young patients who require THA.

Author Contributions

All authors have read and approved the final submitted manuscript. Contributions of the authors are based on [1] substantial contributions to research design, or the acquisition, analysis or interpretation of data [2], drafting the paper or revising it critically [3], and approval of the submitted and final versions. Kazuki Miyama: substantial contributions to research design, or the acquisition, analysis or interpretation of data, drafting the paper or revising it critically, approval of the submitted and final versions; Kenji Kitamura: substantial contributions to research design, or the acquisition, analysis or interpretation of data, drafting the paper or revising it critically, approval of the submitted and final versions; Masanori Fujii: substantial contributions to research design, or the acquisition, analysis or interpretation of data, drafting the paper or revising it critically, approval of the submitted and final versions; Goro Motomura: substantial contributions to research design, or the acquisition, analysis or interpretation of data, approval of the submitted and final versions; Satoshi Hamai: substantial contributions to research design, or the acquisition, analysis or interpretation of data, approval of the submitted and final versions; Shinya Kawahara: substantial contributions to research design, or the acquisition, analysis or interpretation of data, approval of the submitted and final versions; Ryosuke Yamaguchi: substantial contributions to research design, or the acquisition, analysis or interpretation of data, approval of the submitted and final versions; Takeshi Utsunomiya: substantial contributions to research design, or the acquisition, analysis or interpretation of data, approval of the submitted and final versions; Soichiro Yoshino: substantial contributions to research design, or the acquisition, analysis or interpretation of data, approval of the submitted and final versions; Makoto Endo: substantial contributions to research design, or the acquisition, analysis or interpretation of data, approval of the submitted and final versions; Yasuharu Nakashima: substantial contributions to research design, or the acquisition, analysis or interpretation of data, approval of the submitted and final versions.

Ethics Statement

Ethical approval for this study was obtained from the institutional review boards of all participating institutions.

Conflicts of Interest

The authors declare that they received no financial or material support that could be perceived as influencing the research, authorship, or publication of this article.

Supporting information

Supporting File

JOR-44-0-s001.docx (113.6KB, docx)

Acknowledgments

The authors disclose receipt of the following financial support for the research, authorship, and/or publication of this article: this work was supported by a Grant‐in‐Aid for Scientific Research from the Japan Society for the Promotion of Science (No. JP24K23057).

Data Availability Statement

The data that support the findings of this study are available on request from the corresponding author. The data are not publicly available due to privacy or ethical restrictions.

References

  • 1. Wilkin G. P., Ibrahim M. M., Smit K. M., and Beaulé P. E., “A Contemporary Definition of Hip Dysplasia and Structural Instability: Toward a Comprehensive Classification for Acetabular Dysplasia,” Journal of Arthroplasty 32, no. 9S (September, 2017): S20–S27. [DOI] [PubMed] [Google Scholar]
  • 2. Gala L., Clohisy J. C., and Beaulé P. E., “Hip Dysplasia in the Young Adult,” Journal of Bone and Joint Surgery 98, no. 1 (January, 2016): 63–73. [DOI] [PubMed] [Google Scholar]
  • 3. Troelsen A., “Assessment of Adult Hip Dysplasia and the Outcome of Surgical Treatment,” Danish Medical Journal 59, no. 6 (June, 2012): B4450. [PubMed] [Google Scholar]
  • 4. Kitamura K., Fujii M., Ikemura S., Hamai S., Motomura G., and Nakashima Y., “Factors Associated With Abnormal Joint Contact Pressure After Periacetabular Osteotomy: A Finite‐Element Analysis,” Journal of Arthroplasty 37, no. 10 (October, 2022): 2097–2105.e1. [DOI] [PubMed] [Google Scholar]
  • 5. Ganz R., Klaue K., Vinh T. S., and Mast J. W., “A New Periacetabular Osteotomy for the Treatment of Hip Dysplasias. Technique and Preliminary Results,” Clinical Orthopaedics and Related Research 232 (July, 1988): 26–36. [PubMed] [Google Scholar]
  • 6. Beltran L. S., Rosenberg Z. S., Mayo J. D., et al., “Imaging Evaluation of Developmental Hip Dysplasia in the Young Adult,” American Journal of Roentgenology 200, no. 5 (May, 2013): 1077–1088. [DOI] [PubMed] [Google Scholar]
  • 7. Clohisy J. C., Carlisle J. C., Trousdale R., et al., “Radiographic Evaluation of the Hip has Limited Reliability,” Clinical Orthopaedics & Related Research 467, no. 3 (March, 2009): 666–675. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Nelitz M., Guenther K. P., Gunkel S., and Puhl W., “Reliability of Radiological Measurements in the Assessment of Hip Dysplasia in Adults,” British Journal of Radiology 72, no. 856 (April, 1999): 331–334. [DOI] [PubMed] [Google Scholar]
  • 9. Schmitz M. R., Murtha A. S., and Clohisy J. C., “Developmental Dysplasia of the Hip in Adolescents and Young Adults,” Journal of the American Academy of Orthopaedic Surgeons 28, no. 3 (February, 2020): 91–101. [DOI] [PubMed] [Google Scholar]
  • 10. Tachibana T., Fujii M., Kitamura K., Nakamura T., and Nakashima Y., “Does Acetabular Coverage Vary Between the Supine and Standing Positions in Patients With Hip Dysplasia?,” Clinical Orthopaedics & Related Research 477, no. 11 (November, 2019): 2455–2466. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Kitamura K., Fujii M., Ikemura S., Hamai S., Motomura G., and Nakashima Y., “Does Patient‐Specific Functional Pelvic Tilt Affect Joint Contact Pressure in Hip Dysplasia? A Finite‐Element Analysis Study,” Clinical Orthopaedics & Related Research 479, no. 8 (August, 2021): 1712–1724. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Wyles C. C., Vargas J. S., Heidenreich M. J., et al., “Hitting the Target: Natural History of the Hip Based on Achieving an Acetabular Safe Zone Following Periacetabular Osteotomy,” Journal of Bone and Joint Surgery 102, no. 19 (October, 2020): 1734–1740. [DOI] [PubMed] [Google Scholar]
  • 13. Stetzelberger V. M., Leibold C. S., Steppacher S. D., Schwab J. M., Siebenrock K. A., and Tannast M., “The Acetabular Wall Index Is Associated With Long‐Term Conversion to THA After PAO,” Clinical Orthopaedics & Related Research 479, no. 5 (May, 2021): 1052–1065. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Siebenrock K. A., Kistler L., Schwab J. M., Büchler L., and Tannast M., “The Acetabular Wall Index for Assessing Anteroposterior Femoral Head Coverage in Symptomatic Patients,” Clinical Orthopaedics & Related Research 470, no. 12 (December, 2012): 3355–3360. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Tannast M., Hanke M. S., Zheng G., Steppacher S. D., and Siebenrock K. A., “What Are the Radiographic Reference Values for Acetabular Under‐ and Overcoverage?,” Clinical Orthopaedics & Related Research 473, no. 4 (April, 2015): 1234–1246. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Troelsen A., Rømer L., Kring S., Elmengaard B., and Søballe K., “Assessment of Hip Dysplasia and Osteoarthritis: Variability of Different Methods,” Acta Radiologica 51, no. 2 (March, 2010): 187–193. [DOI] [PubMed] [Google Scholar]
  • 17. Kurmis A. P. and Ianunzio J. R., “Artificial Intelligence in Orthopedic Surgery: Evolution, Current State and Future Directions,” Arthroplasty 4, no. 1 (March, 2022): 9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Nakamura S., Ninomiya S., and Nakamura T., “Primary Osteoarthritis of the Hip Joint in Japan,” Clinical Orthopaedics and Related Research 241 (April, 1989): 190–196. [PubMed] [Google Scholar]
  • 19. Nepple J. J., Fowler L. M., and Larson C. M., “Decision‐Making in the Borderline Hip,” Sports Medicine and Arthroscopy Review 29, no. 1 (March, 2021): 15–21. [DOI] [PubMed] [Google Scholar]
  • 20. Hara D., Nakashima Y., Hamai S., et al., “Kinematic Analysis of Healthy Hips During Weight‐Bearing Activities by 3D‐to‐2D Model‐To‐Image Registration Technique,” BioMed Research International 2014 (November, 2014): 457573. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Kitamura K., Fujii M., Utsunomiya T., et al., “Effect of Sagittal Pelvic Tilt on Joint Stress Distribution in Hip Dysplasia: A Finite Element Analysis,” Clinical Biomechanics 74 (April, 2020): 34–41. [DOI] [PubMed] [Google Scholar]
  • 22. Tönnis D., Congenital Dysplasia and Dislocation of the Hip in Children and Adults. (Springer, 1987), 24. [Google Scholar]
  • 23. Tannast M., Fritsch S., Zheng G., Siebenrock K. A., and Steppacher S. D., “Which Radiographic Hip Parameters Do not Have to be Corrected for Pelvic Rotation and Tilt?,” Clinical Orthopaedics & Related Research 473, no. 4 (April, 2015): 1255–1266. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Mathis A., Mamidanna P., Cury K. M., et al., “DeepLabCut: Markerless Pose Estimation of User‐Defined Body Parts With Deep Learning,” Nature Neuroscience 21, no. 9 (September, 2018): 1281–1289. [DOI] [PubMed] [Google Scholar]
  • 25. Miyama K., Bise R., Ikemura S., et al., “Deep Learning‐Based Automatic‐Bone‐Destruction‐Evaluation System Using Contextual Information From Other Joints,” Arthritis Research & Therapy 24, no. 1 (October, 2022): 227. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Williams S., Zhao Z., Hafeez A., et al., “The Discerning Eye of Computer Vision: Can It Measure Parkinson's Finger Tap Bradykinesia?,” Journal of the Neurological Sciences 416 (September, 2020): 117003. [DOI] [PubMed] [Google Scholar]
  • 27. Park S., Yoo H. J., Jang J. S., and Lee S. H., “Automated Non‐Contact Measurement of the Spine Curvature at the Sagittal Plane Using a Deep Neural Network,” Clinical Biomechanics 111 (January, 2024): 106146. [DOI] [PubMed] [Google Scholar]
  • 28. Ronneberger O., Fischer P., and Brox T., “U‐Net: Convolutional Networks for Biomedical Image Segmentation.” Medical Image Computing and Computer‐Assisted Intervention – MICCAI 2015. (Springer International Publishing, 2015), 234–241. [Google Scholar]
  • 29. Wada K. labelme: Image Polygonal Annotation With Python (Polygon, Rectangle, Circle, Line, Point and Image‐Level Flag Annotation) [Internet]. Github; [Cited 2022 Nov 24]. https://github.com/wkentaro/labelme.
  • 30. Kingma D. P. and Ba J., “Adam: A Method for Stochastic Optimization [Internet],” arXiv (2014), 10.48550/arXiv.1412.6980. [DOI] [Google Scholar]
  • 31. Ying X., “An Overview of Overfitting and Its Solutions,” Journal of Physics: Conference Series 1168, no. 2 (February, 2019): 022022. [Google Scholar]
  • 32. Koo T. K. and Li M. Y., “A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research,” Journal of Chiropractic Medicine 15, no. 2 (June, 2016): 155–163. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Willey M., Holland T., Thomas‐Aitken H., and Goetz J., “Diagnosis and Management of Borderline Hip Dysplasia and Acetabular Retroversion,” Journal of Hip Surgery 02, no. 4 (December, 2018): 156–166. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Wilson E. B., “Probable Inference, the Law of Succession, and Statistical Inference,” Journal of the American Statistical Association 22, no. 158 (June, 1927): 209–212. [Google Scholar]
  • 35. Rutter C. M., “Bootstrap Estimation of Diagnostic Accuracy With Patient‐Clustered Data,” Academic Radiology 7, no. 6 (June, 2000): 413–419. [DOI] [PubMed] [Google Scholar]
  • 36. McNEMAR Q., “Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages,” Psychometrika 12, no. 2 (June, 1947): 153–157. [DOI] [PubMed] [Google Scholar]
  • 37. Kitamura K., Fujii M., Iwamoto M., et al., “Is Anterior Rotation of the Acetabulum Necessary to Normalize Joint Contact Pressure in Periacetabular Osteotomy? A Finite‐Element Analysis Study,” Clinical Orthopaedics & Related Research 480, no. 1 (January, 2022): 67–78. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Nunley R. M., Prather H., Hunt D., Schoenecker P. L., and Clohisy J. C., “Clinical Presentation of Symptomatic Acetabular Dysplasia in Skeletally Mature Patients,” Journal of Bone and Joint Surgery 93, no. Suppl 2 (May, 2011): 17–21. [DOI] [PubMed] [Google Scholar]
  • 39. Wylie J. D., Ferrer M. G., McClincy M. P., et al., “What Is the Reliability and Accuracy of Intraoperative Fluoroscopy in Evaluating Anterior, Lateral, and Posterior Coverage During Periacetabular Osteotomy?,” Clinical Orthopaedics & Related Research 477, no. 5 (May, 2019): 1138–1144. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Li Q., Zhong L., Huang H., et al., “Auxiliary Diagnosis of Developmental Dysplasia of the Hip by Automated Detection of Sharp's Angle on Standardized Anteroposterior Pelvic Radiographs,” Medicine 98, no. 52 (December, 2019): e18500. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Yang W., Ye Q., Ming S., et al., “Feasibility of Automatic Measurements of Hip Joints Based on Pelvic Radiography and a Deep Learning Algorithm,” European Journal of Radiology 132 (November, 2020): 109303. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supporting File

JOR-44-0-s001.docx (113.6KB, docx)

Data Availability Statement

The data that support the findings of this study are available on request from the corresponding author. The data are not publicly available due to privacy or ethical restrictions.


Articles from Journal of Orthopaedic Research are provided here courtesy of Wiley

RESOURCES