Skip to main content
PLOS One logoLink to PLOS One
. 2025 Jul 24;20(7):e0328295. doi: 10.1371/journal.pone.0328295

Deep learning for pediatric chest x-ray diagnosis: Repurposing a commercial tool developed for adults

Prerana Agarwal 1,*, Alexander Rau 1, Helen Ngo 1, Ambika Seth 2, Fabian Bamberg 1, Elmar Kotter 1, Jakob Weiss 1
Editor: Shahriar Ahmed3
PMCID: PMC12289065  PMID: 40705715

Abstract

The number of commercially available artificial intelligence (AI) tools to support radiological workflows is constantly increasing, yet dedicated solutions for children are largely unavailable. Here, we repurposed an AI-tool developed for chest radiograph interpretation in adults (Lunit INSIGHT CXR) and investigated its diagnostic performance in a real-world pediatric clinical dataset. 958 consecutive frontal chest radiographs of children aged 2−14 years were included and analyzed with the commercially available AI-tool. The reference standard was determined in a dedicated reading session by a board-certified radiologist. The original reports validated by specialized pediatric radiologists, were considered as second readings. All discordant findings were reanalyzed in consensus. The diagnostic performance of the AI-tool was validated using standard measures of accuracy. For this, the continuous AI output (ranging from 0−100) was binarized using vendor recommended thresholds recommended for adults and optimized thresholds identified for children. Relevant findings were defined as consolidation, atelectasis, nodule, cardiomegaly, mediastinal widening due to mass, pleural effusion and pneumothorax. 200 radiographs [20.9%] demonstrated at least one relevant pathology. Using the adult threshold, the AI-tool showed a high performance for all relevant findings with an AUC 0.94 (95% CI: 0.92–0.95) and. In stratified analysis by age (2−7 vs. 7–14-years-old) a significantly higher performance (p < 0.001) was found for older children with an AUC of 0.96 (95% CI: 0.94–0.98) with a sensitivity and specificity of 87.5% and 82.3% respectively, which further increased using optimized thresholds for children. Repurposing existing AI-tools developed for adult application to pediatric patients could support clinical workflows until dedicated solutions become available.

Introduction

Chest radiography remains one of the most commonly used imaging tests in the pediatric population and serves as an important tool in the workup of various conditions involving the lungs, mediastinum and chest wall [1]. Over the past years, the number of commercially available artificial intelligence (AI) tools to support radiological workflows has substantially increased, especially in the field of chest radiography with applications encompassing a wide range of clinical scenarios such as nodule or pneumonia detection, tuberculosis screening, as well as triaging and streamlining workflow [2–6]. However, the focus of this rapidly evolving landscape has remained with the adult population and dedicated solutions for pediatric patients are limited. Among the currently FDA-cleared commercially available AI solutions, there are, to date, no computer-aided detection products specifically authorized for pediatric use [7]. This gap is attributable to several factors, including logistical challenges such as the limited availability of high-quality, open-access pediatric imaging datasets, and the scarcity of pediatric subspecialty radiologists required for accurate image annotation. Additional barriers include stricter regulatory requirements for pediatric research, smaller market potential compared to adult applications, and the inherent variability in clinical conditions and anatomical characteristics associated with the diverse age range within the pediatric population [8–10].

In this context, repurposing existing AI-tools developed for adult application to pediatric patients could enhance clinical workflows, aid in decision-making and support pediatricians in a resource constrained set-up until dedicated solutions become available. In recent years, efforts have been made to adapt adult chest AI algorithms for pediatric use in order to accelerate the development of pediatric imaging AI. Various approaches have been explored, including the exclusion of children under two years of age and the omission of certain findings such as cardiomegaly to improve performance, as well as optimizing the operating thresholds of AI tools [11–13]. Nonetheless, thorough and rigorous testing is mandatory to gain a better understanding of potential benefits and limitations.

Here, we investigated the diagnostic performance of a commercially available AI-tool developed for adult chest radiograph interpretation (Lunit INSIGHT CXR) in a real-world clinical dataset of children aged 2–14 years old. Our hypothesis was that when used for specific relevant pathologies, such as consolidation, pneumothorax or effusion, repurposed AI tools for chest radiograph analysis can reliably detect such pathologies and support clinical management.

Materials and methods

Patient population

In this single-center retrospective cohort study to independently externally validate a commercially available AI tool, we included frontal chest radiographs of 1000 consecutive children between 2–14 years old who received a clinically indicated chest radiograph at our tertiary care center between January 2021 and April 2022 (Fig 1). For the purpose of the study, the data was accessed between 01.02.2023 and 30.06.2023.

Fig 1. Brief summary highlighting the study methology. a) AI Tool Development and Pediatric Repurposing: The AI tool was originally trained and validated using a large dataset of adult chest radiographs.

Fig 1

For pediatric validation, the tool was retrospectively tested on 958 pediatric chest radiographs (CXR) from children aged 2–14 years. b) Diagnostic Performance Analysis: The AI tool’s diagnostic performance in children was assessed using vendor-recommended thresholds, stratified by age groups (2–6 and 7–14 years), and optimized pediatric-specific thresholds.

Inclusion criteria were chest radiographs acquired in a postero-anterior or antero-posterior projection in the routine clinical workup in children aged 2–14. Children under the age of 2 years were not included in our study since they often show substantially differing anatomical features and a different spectrum of pathologies compared to older children. Since the AI tool is already approved for use in children older than 14 years, we limited the age group for our study from 2–14 years. Exclusion criteria were cases with poor/corrupted image quality such as strong rotation, artifacts due to motion or studies not fully covering the chest.

This study was performed in line with the principles of the Declaration of Helsinki. Approval was granted by the Ethics Committee of University of Freiburg (22–1184-retro) and informed consent was waived.

Chest radiographs

All chest radiographs were acquired in clinical routine and were retrieved in DICOM format from the Picture Archiving and Communication System (PACS). The images were obtained using different radiography units (Siemens Mobilett XP, Philips Digital Diagnost and Mobile Diagnost and Samsung GM 60A and 85A). Only one chest radiograph per patient was included. After visually checking for impaired/corrupted image quality, the chest radiographs were sent to a dedicated image analysis post-processing platform NORA (www.nora-imaging.com) for further analysis.

Reference standard reading.

A board-certified radiologist specialized in chest imaging and with experience in pediatric imaging (P.A., 10 years of experience) analyzed the radiographs in a dedicated reading session, in which the presence/absence of predefined relevant pathologies was recorded for each chest radiograph and annotated in the image for later comparison and verification of the AI model outputs (Fig 2). Relevant findings were defined as: 1) consolidation, 2) atelectasis, 3) nodule, 4) cardiomegaly, 5) mediastinal widening due to mass, 6) pleural effusion and 7) pneumothorax. To generate the best possible reference reads for the clinical question, in addition to the chest radiograph, all additional information available (e.g., lab results, CT scans, clinical history) was taken into account. The original signed reports, which were validated by specialized pediatric radiologists, were considered as second readings. Finally, if any discordant findings were noted between the dedicated study reads and the original signed reports, the case was reanalyzed in consensus by P.A. and J.W. to determine the final reference standard read for this study.

Fig 2. Example of reference standard and AI output.

Fig 2

A) Reference standard with annotated finding by board-certified radiologist specializing in thoracic imaging. Blue marker indicates a consolidation in the right lower zone. The box in the upper right corner shows the annotation tool of the image-processing platform NORA. B) AI output with grayscale map showing consolidation (Csn) with an abnormality score of 68% in the right lower zone, considered as a true positive finding.

AI-Algorithm

The radiographs were evaluated with a commercially available fully automatic AI tool (Lunit INSIGHT CXR, Version 3.1.4.4), which was originally developed for chest radiograph (postero-anterior or antero-posterior projection) interpretation in adults. For training of the AI, 262,445 adult frontal chest radiographs (146,679 normal and 115,766 abnormal) were used from various vendors and with a wide range of acquisition parameters and across a wide range of geographic regions. Independent testing was performed on a total of 486 normal and 529 abnormal chest radiographs collected from five independent institutions (1 from each participant, 628 males and 387 females). Five board-certified radiologists with 7–14 years of experience participated in labeling the chest radiographs, providing the type and the exact location of the abnormalities. Further details on the original development and testing are provided elsewhere in detail [14]. The following ten findings are detected by the algorithm: consolidation, atelectasis, nodule, fibrosis, calcification, pleural effusion, pneumothorax, pneumoperitoneum, cardiomegaly and mediastinal widening. The output is provided as a heatmap or a grayscale map to indicate the location of the finding on the chest radiograph, along with with a probability or abnormality score between 0–100 for each finding, reflecting the certainty of the AI tool (Figs 2 and 3). The software classifies the lesions as positive using a cut-off of 15%, which is the optimized for the adult population.

Fig 3. Image examples representing strengths and limitations of the AI tool.

Fig 3

A) AI output in a 10-year-old patient with a history of Ewing sarcoma of the 1st right rib. The area was highlighted as pathologic with an abnormality score of 73% and classified as consolidation (Csn), fibrosis (Fib) and nodule (Ndl), which most closely resemble the findings the AI tool was developed for. B) Image of a 4-year-old child with a venolymphatic malformation of the chest wall. Similar to A an abnormality was correctly detected but erroneously classified as effusion (PEf) and consolidation (Csn) as it was beyond the application of the AI tool.

Thus far, the AI tool has European Clearance with a Medical Device Regulation Certificate (MDR) for individuals >=14-years-old, Korean Food and Drug Administration Clearance (KFDA) and Brazilian ANVISA clearance and has been widely tested in different clinical scenarios in its intended area of use [15–18]. For the current study with a repurposed application in children, Lunit provided technical support but was not involved in the study design, data collection or data analysis, or decision to publish.

Performance analysis of the AI tool

All chest radiographs meeting the inclusion criteria were sent to the AI tool for automatic analysis. After analysis, the manually annotated reference chest radiograph and the grayscale map overlaid radiographs generated by the AI tool were re-read by one of the authors (H.N.) to identify potential false positives by the AI tool (same findings but different locations in the image).

Statistical analysis

The assumption of normal data distribution of the data was tested with the Shapiro–Wilk test. Continuous variables are presented as mean±standard deviation (SD) or median and interquartile ranges (IQR) as appropriate. Categorical variables are given as frequencies and percentages. The diagnostic performance of the AI tool was investigated using the dedicated reference reads generated for this study as gold standard and reported as per the STARD statement [19]. The continuous output of the AI tool was used to calculate the area under the receiver operating characteristic curve (AUC) with 95% confidence intervals (95% CI) for the combined performance across all findings. In addition, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV) were calculated after binarizing the continuous AI output using a predefined, vendor recommended threshold cut-off of 15, which was the optimal threshold identified for adults. To account for the relatively small number of individual findings, consolidation, nodule, atelectasis, pleural effusion and pneumothorax were combined to pleuroparenchymal and cardiomegaly and mediastinal widening to mediastinal findings to allow for reasonable statistics. First, performance metrics were calculated for the entire dataset using the vendor recommended threshold recommend for adults. Furthermore, subanalyses stratified by were conducted in a similar fashion to investigate a potential age bias. As a final step, instead of using the predefined vendor recommended threshold, an overall optimal cut-off was calculated followed by separate cut-offs for pleuroparenchymal and mediastinal subgroups by maximizing the sum of sensitivity and specificity (R-Package cutpointr). AUCs were compared using the DeLong method [14]. P-value of less than 0.05 was considered to be statistically significant. All analyses were conducted with R (version 4.3.0; R Foundation for Statistical Computing, Vienna, Austria).

Results

Patient cohort

Out of the 1000 radiographs, 42 cases were excluded due to poor/corrupted image quality, resulting in a final study cohort of 958 patients (524 boys; 6.75 ± 3.6 years) with 200 radiographs [20.9%] demonstrating at least one relevant pathology (Table 1).

Table 1. Demographic information.

Parameter Total
(n = 958)
Boys
(n = 524)
Girls
(n = 434)
Age (years) * 6.75 ± 3.6 6.83 ± 3.6 6.65 ± 3.6
Projection (AP) 500/958 (52.2%) 274/524 (52.3%) 226/434 (52.1%)
Relevant overall pathology 200/958 (20.9%) 96/524 (18.3%) 104/434 (23.4%)
• Pleuroparenchymal pathology † 162/200 (81.0%) 83/200 (41.5%) 79/200 (39.5%)
• Mediastinal pathology ‡ 77/200 (38.5%) 29/200 (14.5%) 48/200 (24.0%)
• Both mediastinal and pleuroparenchymal pathology 39/200 (19.5%) 16/200 (8.0%) 23/200 (11.5%)
*

Data represents mean ± SD with range in brackets

†

Relevant pleuroparenchymal pathology includes consolidation, nodule, atelectasis, pleural effusion and pneumothorax

‡

Relevant mediastinal pathology includes cardiomegaly and mediastinal widening due to a mass.

AP = Anteroposterior

Algorithm performance

Performance analysis based on the thresholds recommended for adults:

A summary of the AI tool performance is presented in Table 2. The overall performance of the AI tool for identifying any relevant pathology was high with an AUC of 0.94 (95% CI: 0.92–0.95) and an accuracy of 83.4% (95% CI: 80.1–85.7%). The sensitivity, specificity, PPV and NPV were 87.5%, 82.3%, 56.6% and 96.1%, respectively.

Table 2. Results of the algorithm performance: For the calculation of the diagnostic performance of the AI tool, dedicated reference reads generated for this study were used as a gold standard.

Pathology Cases (n = 958) AUC (95%CI) Accuracy (95% CI) Sensitivity (%) Specificity (%) PPV (%) NPV (%)
Relevant pathology 200 (20.9%) 0.94 (0.92-0.95) 83.4 (80.1-85.7) 87.5 (175/200) 82.3 (624/758) 56.6 (175/309) 96.1 (624/649)
Pleuroparenchymal pathology 162 (16.9%) 0.94 (0.92-0.96) 85.3 (82.9-87.5) 87.7 (142/162) 84.8 (675/796) 54 (142/263) 97.1 (675/695)
Mediastinal pathology 77 (8%) 0.94 (0.92-0.96) 89 (86.9-90.9) 72.7 (56/77) 90.5 (797/881) 40 (56/140) 97.4 (797/818)
Individual pathologies * Cases True Positives False Negatives False Positives True Negatives
Consolidation 144 (15%) 84.7% (122/144) 15.3% (22/144) 14.6% (119/814) 85.4% (695/814)
Atelectasis 24 (2.6%) 41.7% (10/24) 58.3% (14/24) 1.7% (16/934) 98.3% (918/934)
Nodule 9 (1%) 88.9% (8/9) 11.1% (1/9) 6.4% (61/949) 93.6% (888/949)
Pleural effusion 27 (2.8%) 66.7% (18/27) 33.3% (9/27) 1.5% (14/931) 98.5% (917/931)
Pneumothorax 2 (0.2%) 100%(2/2) 0% (0/2) 0.1% (1/956) 99.9% (955/956)
Cardiomegaly 72 (7.5%) 73.6% (53/72) 26.4% (19/72) 8.4% (74/886) 91.6% (812/886)
Mediastinal widening 6 (0.6%) 100% (6/6) 0% (0/6) 7.5% (71/952) 92.5% (881/952)

AUC = Area under the receiver operating characteristic curve

PPV = Positive Predictive Value

NPV = Negative Predictive Value

*

For individual pathologies, the AI-tool performance is depicted in percentages and raw numbers as standard metric of accuracy could not be reliably calculated due to small sample sizes.

For the identification of relevant pleuroparenchymal pathologies (consolidation, nodule, atelectasis, pleural effusion and pneumothorax), the algorithm showed an AUC of 0.94 (95% CI: 0.92–0.96) and an accuracy of 85.3% (95% CI: 82.9–87.5%). The sensitivity and specificity were 87.7% and 84.8%, respectively. PPV and NPV were 54% and 97.1%, respectively.

For mediastinal pathologies (cardiomegaly and mediastinal widening due to a mass), the metrics were as follows: AUC 0.94 (95% CI: 0.92–0.96); accuracy 89% (95% CI: 86.6–90.9%); sensitivity, specificity, PPV and NPV: 72.7%, 90.4%, 40% and 97.4%, respectively.

Performance analysis stratified by age.

To investigate the impact of age on the algorithm performance, we performed a subgroup analysis between children aged 2–6 years vs. 7–14 years using the corrected performance metrics. This age cut-off was chosen as it allowed for a balanced distribution of patients, allowing for statistically meaningful comparison between early childhood and school-aged children. Furthermore, in a previous study evaluating the same AI tool, 7 years was identified as the median age for correct diagnosis [12].

In the younger age group (n = 509 [53.1%]; mean age 3.7 ± 1.4 years), the AUC for the presence of any relevant finding was 0.91 (95% CI: 0.88–0.94), with an accuracy of 78% (95% CI: 74.1–81.5%) and a sensitivity, specificity, PPV and NPV of 88.1%, 74%, 57% and 94.1%, respectively. In the older age group (n = 449 [46.9%]; mean age 10.2 ± 2 years), the performance metrics were as follows: AUC 0.96 (95% CI: 0.94–0.98), accuracy 89.7% (95% CI: 86.5–92.4%), sensitivity 86.2%, specificity 90.3%, PPV 56.8% and NPV 97.8%, respectively, which were significantly higher compared to the younger age groups (p < 0.001).

A detailed summary of all findings is presented in S1 and S2 Tables.

Optimal thresholds for pediatric patients.

The vendor-recommended threshold of 15 to binarize the continuous AI output (ranging from 0–100) is based on the optimal threshold identified for adults. Therefore, as a final step in our analysis, we calculated optimized cut-offs for children for the entire cohort. Overall, the optimal cut-off for the pediatric age group was found to be 44.5. This adjusted threshold yielded a similar AUC of 0.93 with an accuracy of 89.8% (95% CI: 87.7–91.6%), sensitivity of 80.0%, and specificity of 92.4%. Compared to the adult threshold, the pediatric-specific cut-off led to improvement in the specificity and the PPV, supporting a safer AI interpretation in children with a slight trade-off in the sensitivity at 80%. The sub-group specific thresholds for pleuroparenchymal and mediastinal subgroups are shown in the Fig 4 and for each pathology are shown in S1 Fig.

Fig 4. Definition of optimal thresholds for the performance of the AI-tool in children.

Fig 4

The dotted blue line represents the pre-defined vendor recommended threshold of 15, which is based on the optimal threshold identified for adults to dichotomize the continuous AI-output (0-100). The green diamonds show optimized cut-offs calculated for maximizing the sum of sensitivity and specificity.. The performance metrics based on adult threshold of 15 (blue) and optimized cutoffs for children (green) are shown in the column on the right side (sens = sensitivity, spec = specificity, PPV = positive predictive value, NPV = negative predictive value).

Discussion

In this study, we repurposed and externally validated the diagnostic performance of a commercially available AI-tool (Lunit INSIGHT CXR) which was originally developed for chest radiograph evaluation for adults in a real-world pediatric dataset. When benchmarked against dedicated reference reads generated for this study, the algorithm exhibited a high and clinically acceptable performance for relevant findings. Notably, subanalyses by age revealed a significantly higher performance for older children aged 7–14 years, compared to younger patients (2–6 years), presumably due to a substantially different anatomy in younger individuals. For example, the presence of a thymic shadow could pose a challenge, resulting in its erroneous identification as cardiomegaly or mediastinal enlargement. Additionally, perhaps the distinctly visible vascular patterns in the lungs of supine pediatric subjects were occasionally misinterpreted as pulmonary edema (the algorithm classifies pulmonary edema as consolidation).

Our results are of clinical importance, because most research on AI tools for pediatric chest radiograph interpretation over the past years focused on pneumonia detection and automatic segmentation of the lungs while comprehensive solutions evaluating more pathologies are still missing [20–22]. The concept of repurposing AI tools developed for the adult population in pediatric patients was already investigated by Morcos at., who explored the reliability of an algorithm from an open source library for chest radiograph datasets (TorchXRayVision) for pneumonia detection ([13,23]. The authors found a decent but substantially lower performance compared to our study (sensitivity 79.83%, specificity 67.66%, PPV 86.95%, NPV 55.41%). The authors highlighted that although models exclusively trained on pediatric images would likely perform better, AI research on pediatric chest radiograph interpretation could be expedited by leveraging adult-based algorithms.

A similar approach as in the presented study was previously reported by Shin et. al., who investigated the diagnostic performance of the same AI-tool in a pediatric cohort (0–18 years) [12]. Exclusion of children under the age of 2 years and cardiomegaly increased the accuracy to as high as 96.9%, comparable to the performance for adult patients. In keeping with the findings of our study, they observed that age of patients with incorrect diagnoses was significantly younger than those with correct diagnosis (median 1 year vs. 7 years, p < 0.001). Moreover, age emerged as a significant factor for incorrect diagnosis in the logistic regression test. While noting a better performance of the tool in this study, compared to our results, it is crucial to note a fundamental distinction in the study design. We conducted a dedicated reading session informed by clinical information and radiological expertise, followed by a separate assessment of the AI-tools performance to avoid any potential bias. In contrast, in the above mentioned study, the radiologist was presented with the AI-output and could then confirm the abnormalities in reference to the results of the AI-tool, which was considered as a reference read. Nonetheless, these results emphasize the need for a refined approach: while older children could benefit from an extended application of adult chest radiograph algorithms, simultaneous developmental efforts are needed for younger subjects. This is underscored by our analysis calculating optimized thresholds to binarize the AI-output, which, without any retraining of the AI-tool, allowed for a performance increase, especially in the patients aged 2–6 years. Although we refrained from pathology specific cut-offs due to a small sample size of individual pathologies, this additional step highlights the potential to further refine the diagnostic performance of the AI tool by calculating specific cut-offs for children in larger data sets. One such approach was used in another study where separate operating points were calculated for children based on lesion type, age and imaging method in a larger dataset to improve the diagnostic performance of the AI tool [11].

The spectrum of pathologies differs significantly in pediatric patients. In our study, we noted that in cases of rare pathologies, such as sequestration, Ewings-Sarcoma of the first rib, and a chest wall lymphangioma producing a pleural shadow; the algorithm demonstrated reasonable sensitivity in detection even though the classification was inaccurate, which is partly explained by overlapping radiological findings of various clinical entities (Fig 3). Furthermore, the performance of the AI tool cannot be extrapolated to certain respiratory conditions like tuberculosis, where radiological manifestation of TB are different among the pediatric and adult populations [24]. This further highlights the possible role of AI-algorithms in abnormality detection while leaving the role of interpretation to an expert for this patient population.

Our study has the following limitations. Firstly, though the sample was enrolled from clinical routine, the frequency of pathologies was relatively small in the study population. Secondly, although the newly defined optimized thresholds for children allowed for a performance increase, especially in younger individuals, further validation in larger and more diverse datasets is necessary to test for generalizability. Our study was conducted in a predominantly European pediatric population, reflecting the demographic composition of the study region. Future studies involving more ethnically diverse cohorts are warranted to assess the generalizability and fairness of AI applications in pediatric imaging, and to account for potential algorithmic biases related to ethnicity [25]. Finally, even though care was taken to generate high quality reference reads, minor variances in image interpretation such as missing/over diagnosing slight pulmonary edema or atelectasis could not be avoided. This could have been partially be off-set with a formal multi-reader consensus panel for ground truth reading to improve the replication of radiological analysis. Furthermore, during a focused review of cases to identify potential false positives, we identified a few cases such a discrete atelectasis or borderline cardiomegaly, which were flagged by the AI tool but not marked in the reference reads. While it is also important to note that AI may “overdiagnose” subtle findings that experienced radiologists would reasonably judge as clinically insignificant or not warranting formal reporting, this could suggest a potential role of AI tools in highlighting findings that could be overlooked in the clinical routine.

In conclusion, a repurposed AI tool developed for adult chest radiograph diagnosis but applied to a pediatric population showed high and clinically acceptable diagnostic performance for relevant findings in this independent validation study. Additional fine-tuning may help to further increase performance and support clinical decision-making in a particularly vulnerable patient population.

Supporting information

S1 Table. Results of the algorithm performance in children aged 2–6 years.

(DOCX)

pone.0328295.s001.docx (18.5KB, docx)
S2 Table. Results of the algorithm performance in children aged 7–14 years.

(DOCX)

pone.0328295.s002.docx (14.9KB, docx)
S1 Fig. Flow charts showing the performance for all the relevant pathologies (a) in all age groups, (b) 7–14 years, and (c) 2–6 years.

A) All relevant pathologies. B) Pleuroparenchymal pathologies. C) Mediastinal pathologies.

(TIF)

pone.0328295.s003.tif (390.3KB, tif)
S2 Fig. Definition of optimal thresholds for the performance of the AI-tool in children.

The dotted blue line represents the pre-defined vendor recommended threshold of 15, which is based on the optimal threshold identified for adults to dichotomize the continuous AI-output (0–100). The green diamonds show optimized cut-offs calculated for maximizing the sum of sensitivity and specificity. The performance metrics based on adult threshold of 15 (blue) and optimized cutoffs for children (green) are shown in the column on the right side (sens = sensitivity, spec = specificity, PPV = positive predictive value, NPV = negative predictive value)

(TIF)

pone.0328295.s004.tif (547.2KB, tif)

Abbreviations

AI

Artificial Intelligence

AUC

Area under the receiver operating characteristic curve

CI

Confidence Interval

CXR

Chest X-Ray

DL

Deep learning

NPV

Negative predictive value

PPV

Positive predictive value

SD

Standard deviation

Data Availability

Due to institutional data privacy regulations and patient confidentiality at the University Hospital Freiburg, the data cannot be shared publicly. However, the data may be made available upon reasonable request by contacting the institution at rdia.studienzentrum@uniklinik-freiburg.de.

Funding Statement

Our institute received a grant from Lunit for technical support of the study. The funders had no role in study design, data collection and analysis, decision to publish the manuscript. Some information regarding the AI tool (training data) was provided for manuscript preparation.

References

  • 1.Menashe SJ, Iyer RS, Parisi MT, Otto RK, Stanescu AL. Pediatric Chest Radiographs: Common and Less Common Errors. AJR Am J Roentgenol. 2016;207(4):903–11. doi: 10.2214/AJR.16.16449 [DOI] [PubMed] [Google Scholar]
  • 2.Rueckel J, Kunz WG, Hoppe BF, Patzig M, Notohamiprodjo M, Meinel FG, et al. Artificial Intelligence Algorithm Detecting Lung Infection in Supine Chest Radiographs of Critically Ill Patients With a Diagnostic Accuracy Similar to Board-Certified Radiologists. Crit Care Med. 2020;48(7):e574–83. doi: 10.1097/CCM.0000000000004397 [DOI] [PubMed] [Google Scholar]
  • 3.Hwang EJ, Park CM. Clinical Implementation of Deep Learning in Thoracic Radiology: Potential Applications and Challenges. Korean J Radiol. 2020;21(5):511–25. doi: 10.3348/kjr.2019.0821 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Lee JH, Sun HY, Park S, Kim H, Hwang EJ, Goo JM, et al. Performance of a Deep Learning Algorithm Compared with Radiologic Interpretation for Lung Cancer Detection on Chest Radiographs in a Health Screening Population. Radiology. 2020;297(3):687–96. doi: 10.1148/radiol.2020201240 [DOI] [PubMed] [Google Scholar]
  • 5.Guo R, Passi K, Jain CK. Tuberculosis Diagnostics and Localization in Chest X-Rays via Deep Learning Models. Front Artif Intell. 2020;3:583427. doi: 10.3389/frai.2020.583427 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Kao E-F, Liu G-C, Lee L-Y, Tsai H-Y, Jaw T-S. Computer-aided detection system for chest radiography: reducing report turnaround times of examinations with abnormalities. Acta Radiol. 2015;56(6):696–701. doi: 10.1177/0284185114538017 [DOI] [PubMed] [Google Scholar]
  • 7.Diagnostic Image Analysis Group. AI for radiology: an implementation guide. https://grand-challenge.org/aiforradiology/. 2020. [Google Scholar]
  • 8.Sammer MBK, Akbari YS, Barth RA, Blumer SL, Dillman JR, Farmakis SG, et al. Use of Artificial Intelligence in Radiology: Impact on Pediatric Patients, a White Paper From the ACR Pediatric AI Workgroup. J Am Coll Radiol. 2023;20(8):730–7. doi: 10.1016/j.jacr.2023.06.003 [DOI] [PubMed] [Google Scholar]
  • 9.Tierradentro-Garcia LO, Sotardi ST, Sammer MBK, Otero HJ. Commercially available artificial intelligence algorithms of interest to pediatric radiology: The growing gap between potential use and data training. Journal of the American College of Radiology. 2023;20(8):748–51. [DOI] [PubMed] [Google Scholar]
  • 10.Padash S, Mohebbian MR, Adams SJ, Henderson RDE, Babyn P. Pediatric chest radiograph interpretation: how far has artificial intelligence come? A systematic literature review. Pediatr Radiol. 2022;52(8):1568–80. doi: 10.1007/s00247-022-05368-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Shin HJ, Han K, Son N-H, Kim E-K, Kim MJ, Gatidis S, et al. Optimizing adult-oriented artificial intelligence for pediatric chest radiographs by adjusting operating points. Sci Rep. 2024;14(1):31329. doi: 10.1038/s41598-024-82775-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Shin HJ, Son N-H, Kim MJ, Kim E-K. Diagnostic performance of artificial intelligence approved for adults for the interpretation of pediatric chest radiographs. Sci Rep. 2022;12(1):10215. doi: 10.1038/s41598-022-14519-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Morcos G, Yi PH, Jeudy J. Applying Artificial Intelligence to Pediatric Chest Imaging: Reliability of Leveraging Adult-Based Artificial Intelligence Models. J Am Coll Radiol. 2023;20(8):742–7. [DOI] [PubMed] [Google Scholar]
  • 14.Hwang EJ, Park S, Jin KN, Kim JI, Choi SY, Lee JH. Development and validation of a deep learning–based automated detection algorithm for major thoracic diseases on chest radiographs. JAMA Network Open. 2019;2(3):e191095-e. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.van Beek EJR, Ahn JS, Kim MJ, Murchison JT. Validation study of machine-learning chest radiograph software in primary and emergency medicine. Clin Radiol. 2023;78(1):1–7. doi: 10.1016/j.crad.2022.08.129 [DOI] [PubMed] [Google Scholar]
  • 16.Hwang EJ, Kim H, Yoon SH, Goo JM, Park CM. Implementation of a Deep Learning-Based Computer-Aided Detection System for the Interpretation of Chest Radiographs in Patients Suspected for COVID-19. Korean J Radiol. 2020;21(10):1150–60. doi: 10.3348/kjr.2020.0536 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Kim JH, Kim JY, Kim GH, Kang D, Kim IJ, Seo J, et al. Clinical Validation of a Deep Learning Algorithm for Detection of Pneumonia on Chest Radiographs in Emergency Department Patients with Acute Febrile Respiratory Illness. J Clin Med. 2020;9(6):1981. doi: 10.3390/jcm9061981 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Hwang EJ, Lee JS, Lee JH, Lim WH, Kim JH, Choi KS, et al. Deep Learning for Detection of Pulmonary Metastasis on Chest Radiographs. Radiology. 2021;301(2):455–63. [DOI] [PubMed] [Google Scholar]
  • 19.Bossuyt PM, Reitsma JB, Bruns DE, Gatsonis CA, Glasziou PP, Irwig L, et al. STARD 2015: An Updated List of Essential Items for Reporting Diagnostic Accuracy Studies. Radiology. 2015;277(3):826–32. doi: 10.1148/radiol.2015151516 [DOI] [PubMed] [Google Scholar]
  • 20.Liang G, Zheng L. A transfer learning method with deep residual network for pediatric pneumonia diagnosis. Comput Methods Programs Biomed. 2020;187:104964. doi: 10.1016/j.cmpb.2019.06.023 [DOI] [PubMed] [Google Scholar]
  • 21.E L, Zhao B, Guo Y, Zheng C, Zhang M, Lin J, et al. Using deep-learning techniques for pulmonary-thoracic segmentations and improvement of pneumonia diagnosis in pediatric chest radiographs. Pediatr Pulmonol. 2019;54(10):1617–26. doi: 10.1002/ppul.24431 [DOI] [PubMed] [Google Scholar]
  • 22.Zucker EJ, Barnes ZA, Lungren MP, Shpanskaya Y, Seekins JM, Halabi SS, et al. Deep learning to automate Brasfield chest radiographic scoring for cystic fibrosis. J Cyst Fibros. 2020;19(1):131–8. doi: 10.1016/j.jcf.2019.04.016 [DOI] [PubMed] [Google Scholar]
  • 23.Cohen JPVJ, Hashir M, Bertrand H. TorchXRayVision: A library of chest X-ray datasets and models. 2020. https://github.com/mlmed/torchxrayvision [Google Scholar]
  • 24.Leung AN, Müller NL, Pineda PR, FitzGerald JM. Primary tuberculosis in childhood: radiographic manifestations. Radiology. 1992;182(1):87–91. doi: 10.1148/radiology.182.1.1727316 [DOI] [PubMed] [Google Scholar]
  • 25.Bachina P, Garin SP, Kulkarni P, Kanhere A, Sulam J, Parekh VS. Coarse race and ethnicity labels mask granular underdiagnosis disparities in deep learning models for chest radiograph diagnosis. Radiology. 2023;309(2):e231693. [DOI] [PMC free article] [PubMed] [Google Scholar]

Decision Letter 0

Shahriar Ahmed

PONE-D-25-04141Deep Learning for Pediatric Chest X-Ray Diagnosis: Repurposing a Commercial Tool Developed for AdultsPLOS ONE

Dear Dr. Agarwal,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

==============================

Thank you for submitting this interesting manuscript to our journal. The use of AI is the most talked about thing in healthcare industry right now and we highly appreciate your work in this field. There is also notable gaps in evidence base for its use case in childhood TB. We are very happy to inform you that this paper has been reviewed by relevant experts and I am also very happy to see the time and effort they have invested in this review. I believe that addressing their comments/feedback will significantly improve the scientific validity of the manuscript. We look forward to the revised submission. 

==============================

Please submit your revised manuscript by May 22 2025 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org . When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols . Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols .

We look forward to receiving your revised manuscript.

Kind regards,

Shahriar Ahmed, MBBS, MHE, MPhil

Academic Editor

PLOS ONE

Journal Requirements:

When submitting your revision, we need you to address these additional requirements.

1.Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2. Please note that PLOS ONE has specific guidelines on code sharing for submissions in which author-generated code underpins the findings in the manuscript. In these cases, we expect all author-generated code to be made available without restrictions upon publication of the work. Please review our guidelines at https://journals.plos.org/plosone/s/materials-and-software-sharing#loc-sharing-code and ensure that your code is shared in a way that follows best practice and facilitates reproducibility and reuse.

3. Please note that funding information should not appear in any section or other areas of your manuscript. We will only publish funding information present in the Funding Statement section of the online submission form. Please remove any funding-related text from the manuscript.

4. Thank you for stating the following financial disclosure:

“Our institute received a grant from Lunit for technical support of the study”

Please state what role the funders took in the study.  If the funders had no role, please state: "The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript."

If this statement is not correct you must amend it as needed.

Please include this amended Role of Funder statement in your cover letter; we will change the online submission form on your behalf.

5. We note that you have indicated that there are restrictions to data sharing for this study. PLOS only allows data to be available upon request if there are legal or ethical restrictions on sharing data publicly. For more information on unacceptable data access restrictions, please see http://journals.plos.org/plosone/s/data-availability#loc-unacceptable-data-access-restrictions.

Before we proceed with your manuscript, please address the following prompts:

a) If there are ethical or legal restrictions on sharing a de-identified data set, please explain them in detail (e.g., data contain potentially identifying or sensitive patient information, data are owned by a third-party organization, etc.) and who has imposed them (e.g., a Research Ethics Committee or Institutional Review Board, etc.). Please also provide contact information for a data access committee, ethics committee, or other institutional body to which data requests may be sent.

b) If there are no restrictions, please upload the minimal anonymized data set necessary to replicate your study findings to a stable, public repository and provide us with the relevant URLs, DOIs, or accession numbers. For a list of recommended repositories, please see

https://journals.plos.org/plosone/s/recommended-repositories. You also have the option of uploading the data as Supporting Information files, but we would recommend depositing data directly to a data repository if possible.

We will update your Data Availability statement on your behalf to reflect the information you provide.

6. Your ethics statement should only appear in the Methods section of your manuscript. If your ethics statement is written in any section besides the Methods, please move it to the Methods section and delete it from any other section. Please ensure that your ethics statement is included in your manuscript, as the ethics statement entered into the online submission form will not be published alongside your manuscript.

Additional Editor Comments:

Thank you for submitting this interesting manuscript to our journal. The use of AI is the most talked about thing in healthcare industry right now and we highly appreciate your work in this field. There is also notable gaps in evidence base for its use case in childhood TB. We are very happy to inform you that this paper has been reviewed by relevant experts and I am also very happy to see the time and effort they have invested in this review. I believe that addressing their comments/feedback will significantly improve the scientific validity of the manuscript. We look forward to the revised submission.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Yes

**********

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #2: Yes

**********

3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: No

Reviewer #2: No

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

**********

5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: This is important research that aims to advance the performance of AI CXR models among pediatric patients, a patient group that has not received attention in the digital X-ray space despite the accelerated development of this field for adults. Authors assessed how a CAD algorithm that was previously trained on adults performs among children and found that performance was clinically acceptable, though further cutoff tuning could help further improve.

I have a few comments:

Abstract:

• I recommend naming the commercial tool directly in the abstract so that researchers interested in this tool can identify your paper more quickly.

• Would mention out of the 958 how many had abnormal CXR (n=200) for a bit more context on the study population

• “The diagnostic performance of the AI-tool was validated using standard measures of accuracy using recommended and optimized thresholds to dichotomize the continuous AI output (0-100)” – clarify that the recommended thresholds are for adults.

• I am not a proponent of reporting on the accuracy of an algorithm. Accuracy can be highly misleading in imbalanced datasets as is the case with only 200 positives. For overall performance, I would report the AUC and then the sensitivity/specificity, specifying that this is against the adult-recommended threshold. For the age-stratified results, I would only report the AUCs in the abstract.

Introduction

• Again, would name the actual CXR device here. You don’t introduce what the tool is until the methods.

Methods

• The initials of the radiologists are still “XX”

• Inconsistent nomenclature for figure (“Fig.” vs “Figures”)

• “The following ten findings are detected by the algorithm: consolidation, atelectasis, nodule, fibrosis, calcification, pleural effusion, pneumothorax, pneumoperitoneum, cardiomegaly and mediastinal widening.” – could you provide the N for each of these that were used when training the model on the adult population. Would allow for a comparison with the pediatric cohort.

• You are missing the diagnostic performance metrics (AUC, sens, spec…) of the CAD in the adult population. I see it in figure 5, but this could be missed and should be in the main text to allow for comparison with what is achieved in pediatrics.

Results

• When discussing threshold-based metrics (accuracy, sens, spec, PPV, NPV) it is good to remind the reader that this is benchmarked against the adult threshold.

• The corrected analysis is interesting to highlight the imperfect reference standard (i.e., human radiologist). I think it highlights a limitation of the study that is not acknowledged in the discussion, that there was only one board-certified radiologist reviewing the X-rays compared to the 5 used in the adult study. This is evidenced by the fact that 6/7 corrected scores were actually when the AI was right but the labeller was wrong. I am not sure if re-calculating the diagnostic metrics for the corrected analysis makes sense as this is altering the definition of what the reference actually is. It may be sufficient to descriptively highlight these discrepant cases, and mention this lack of replication of radiological analysis as a limitation in the discussion.

• “we calculated optimized cut-offs for each pathology and their performance metrics” – name the two pathology groups because I initially thought it was for each individual pathology (which is underpowered). It would also be good to have an overall re-calculated thresholds for any abnormal vs normal for pediatric, as was done for adults.

Discussion

• As previously mentioned, acknowledgement of a lack of replication of radiological interpretation in the reference needs to be acknolweged

• Further limitations are that performance cannot be extrapolated to other respiratory conditions (e.g., tuberculosis CAD X-ray interpretation in adults vs in pediatrics cannot draw conclusions from this study). Could go along with the generalizability statement already made.

Tables

• Table 1: the pleuroparenchymal and mediastinal denominator should be 200. Would also be good to show the number who have both as a separate line .

• Inconsistent N for the relevant pathology, pleuro, mediastinal across table 1, 2a and 2b. I expect the number in 1 to match 2a since 2b is the correct. However relevant overall in 1 is N=200 while in 2a is N=198, etc.

• Figure 1: there is a lot of text, would be good to cut down. No definition of what (a) and (b) represent. Would be good to also show performance metrics for both adults and pediatrics so that it gives all the info in one place.

• Figure 2: doesn’t add much to the manuscript, can be deleted.

Reviewer #2: Dear Editor, authors,

I am grateful for the opportunity to review the manuscript “Deep Learning for Pediatric Chest X-Ray Diagnosis: Repurposing a Commercial Tool Developed for Adults”. The authors note that there are few dedicated AI tools for evaluating pediatric chest radiographs. They, therefore, set out to study how a tool developed for the analysis of chest radiographs of adults performs when used (“repurposed”) for pediatric chest radiographs.

Title:

Fine.

Keywords:

Please revisit the keywords. My suggestion: Pediatric; Chest Radiograph; Artificial Intelligence

Abstract:

Overall informative. Please consider the following suggestions:

Rare --> largely unavailable; Any discordant findings --> All discordant findings

The authors state: “The diagnostic performance of the AI-tool was validated using standard measures of accuracy using recommended and optimized thresholds to dichotomize the continuous AI output (0-100).” I think some readers might struggle with this and therefore ask the authors to elaborate on this.

The performance was high for relevant findings. Please elaborate what relevant means in this context.

The readers may find the categorization according to the 7-year cut-off somewhat arbitrary. The authors may want to touch on this point, too.

Competing Interests:

The authors have declared that no competing interests exist. With all due respect, this statement must be further justified given that the developer of the studied software tool supported the study and that one of the authors is affiliated with the company itself. I also notice that the authors report on page 17 that the developer financed two of the key authors through their institution. I warmly welcome full transparency. If possible, consider stating that although the developer funded the study and one of the authors is employed by the company, all authors had access to the data and the decision to publish was based on the findings, not by the request of the company.

Ethics Statement:

Please also explicitly state the need for written informed consent was waived by the ethics committee and the study adheres to the national laws and regulations, if applicable.

Abbreviations:

I note the authors use the terms chest radiograph and chest X-ray. Please stick to chest radiograph through and through.

Introduction:

Although the introduction is well-written and reasoned, I think a more thorough review of existing tools for pediatric chest radiographs is very much needed. I therefore suggest well-conducted literature search and addition of a table summarizing the key results.

Materials and Methods:

Patient population:

The authors should describe the study as a single-center retrospective cohort study.

Perhaps the authors would like to explain why they chose to include 1000 consecutive chest radiographs. What led them to think this is an appropriate size for the cohort, not for example 900 or 1100 chest radiographs?

I think it’s reasonable to omit the exclusion criteria 2 because the authors already stated they included radiographs of older children. Rather, please reason why the study was limited to children aged between 2-14 years.

Ethics statement: please make sure it’s concordant with the ethics statement as discussed earlier.

Reference standard reading:

Please consider the following suggestions:

Fig. 3 --> Figure 3.

Please explain why all the available information, including the lab results, CT scans, and clinical history were taken into consideration when laying out the reference reads. After all, this information was not available to the software. While I of course support the use of all available information when reading the radiographs, one might argue that a truer performance assessment could have been achieved without considering other information.

Chest radiographs:

Were the chest radiographs caputed in supine, standing or sitting position?

AI-Algorithm:

Rather than speaking of gender, please speak of sex. Men --> males, women --> females.

Was Lunit involved in the decision to publish the results?

Could you please also touch on the training set used in the validation study. I think it would be very important to discuss the homogeneity/heterogeneity of the dataset particularly in relation to sex, age and ethnicity of the patients.

Statistical Analysis:

I wish to applaud for the well-written chapter. Normally this section is overlooked. I find it easy to immerse into the Results section after reading this chapter. However, I ask the authors to elaborate on the age categorization.

Results:

All in all, this section is well written and nicely backed by the tables. However, as the tables should stand alone, I think the authors could think of further strengthening the tables by informing the reader where the numbers come from. For example, it should be clear what is being used as the gold standard and why the number of relevant overall pathologies differ between the tables.

Please also report the results for girls/boys separately in Table 1.

I’d also encourage the authors to report the ethnicities of the patients.

Discussion:

As I suggested earlier, I prompt a more thorough review and discussion of the existing tools for pediatric chest radiograph assessment.

I’d like the authors to also discuss whether the results would translate to all demographic subgroups of children. Particularly, does ethnicity have an effect?

Table 2b:

“a simplified manner to avoid bias due to low number of cases.” � please rewrite for clarity.

Figure 1:

Please also indicate the study setting: this is a single-center retrospective cohort study.

It said in the results 200 (20.9%) of the chest radiographs had at least one pathology. Figure 1 seems to suggest there were 198 chest radiographs with abnormal findings.

Of note, it seems the fonts are widely inconsistent between the Figures.

Figure 2:

Please name the Figure descriptively keeping in mind that the Figure and the description should be able to “stand alone”, i.e., be clear without reading the text.

If the name of the title refers to the CONSORT reporting system, the word should be fully capitalized.

Please correct the inconsistent spacings.

Figure 3:

There seems to be some kind of an annotation in the upper right corner of the A part. Was this intentional?

Figure 4:

OK, very informative.

Figure 5:

Not sure if the figure is needed. I would not oppose its inclusion, however, I feel it does not really add too much information. Should the authors want to keep it, I’d ask them to avoid abbreviations sens, PPV, NOV and spec.

**********

6. PLOS authors have the option to publish the peer review history of their article (what does this mean? ). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy .

Reviewer #1: No

Reviewer #2: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/ . PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org . Please note that Supporting Information files do not need this step.

PLoS One. 2025 Jul 24;20(7):e0328295. doi: 10.1371/journal.pone.0328295.r002

Author response to Decision Letter 1


25 May 2025

Point-by-Point Response to Reviewer Comments

PONE-D-25-04141

Deep Learning for Pediatric Chest Radiograph Diagnosis: Repurposing a Commercial Tool Developed for Adults

General Response: We thank the editor and reviewers for their thoughtful comments and have addressed all concerns in the revised version of the manuscript. We have also formatted the manuscript to suit the journal´s style and reorganized the layout to include the tables and figure texts in the manuscript. A point-by-point response to reviewers is provided below. The page and line numbers refer to the manuscript version with tracked changes.

Response to Reviewer 1

We thank the reviewer for the critical evaluation of our work and are grateful for their very helpful comments and suggestions that we have incorporated in our revised manuscript as detailed below:

Comment 1:

I recommend naming the commercial tool directly in the abstract so that researchers interested in this tool can identify your paper more quickly.

Response 1:

We have added the name of the tool to the abstract. (Page: 3, Line: 5)

Comment 2:

Would mention out of the 958 how many had abnormal CXR (n=200) for a bit more context on the study population.

Response 2:

We thank the reviewer for this suggestion. We have added the following sentence to the Abstract to provide this context (Page: 3, Line: 17):

“200 radiographs [20.9%] demonstrated at least one relevant pathology”.

Comment 3:

The diagnostic performance of the AI-tool was validated using standard measures of accuracy using recommended and optimized thresholds to dichotomize the continuous AI output (0-100)” – clarify that the recommended thresholds are for adults.

Response 3:

The sentence was changed as follows for clarity (Page: 3, Lines: 12-14 ):

“For this, the continuous AI output (ranging from 0-100) was binarized using vendor recommended thresholds recommended for adults and optimized thresholds identified for children.”

Comment 4:

I am not a proponent of reporting on the accuracy of an algorithm. Accuracy can be highly misleading in imbalanced datasets as is the case with only 200 positives. For overall performance, I would report the AUC and then the sensitivity/specificity, specifying that this is against the adult-recommended threshold. For the age-stratified results, I would only report the AUCs in the abstract.

Response 4:

Thank you for your valuable feedback.

We have changed the abstract as follows (Page: 3, Lines: 17-23)

“Using the adult threshold, the AI-tool showed a high performance for all relevant findings with an AUC 0.94 (95% CI: 0.92-0.95) and. In stratified analysis by age (2-7 vs. 7-14-years-old) a significantly higher performance (p<0.001) was found for older children with an AUC of 0.96 (95% CI: 0.94-0.98) with a sensitivity and specificity of 87.5% and 82.3% respectively, which further increased using optimized thresholds for children.”

Comment 5:

Again, would name the actual CXR device here (in introduction).

Response 5:

We have made the following changes to the introduction: Page: 5, Line: 2

“Here, we investigated the diagnostic performance of a commercially available AI-tool developed for adult chest radiograph interpretation (Lunit INSIGHT CXR) in a real-world clinical dataset of children aged 2-14 years old.”

Comment 6:

The initials of the radiologists are still “XX”

Response 6:

The initials were changed to “PA” (Page: 7, Line: 12)

Comment 7:

Inconsistent nomenclature for figure (“Fig.” vs “Figures”)

Response 7:

We changed “Fig.” and “Figures” to “Fig” to comply with the journal´s formatting guidelines.

Comment 8:

“The following ten findings are detected by the algorithm: consolidation, atelectasis, nodule, fibrosis, calcification, pleural effusion, pneumothorax, pneumoperitoneum, cardiomegaly and mediastinal widening.” – could you provide the N for each of these that were used when training the model on the adult population. Would allow for a comparison with the pediatric cohort.

Response 8:

We would like to thank the reviewer for this question. We have contacted Lunit for this question and have included their response here:

“While we are unable to disclose the exact number of annotated training cases for each individual finding due to proprietary limitations, I can confirm that Lunit INSIGHT CXR was trained on a dataset of approximately 280,000 chest X-rays with annotation and the algorithm was specifically trained to detect the following ten key findings: Consolidation, atelectasis, nodule, fibrosis, calcification, pleural effusion, pneumothorax, pneumoperitoneum, cardiomegaly, and mediastinal widening. All images used for training were confirmed by one of the following methods: Original radiology report + pathology confirmation or radiology report + CT confirmation, or independent image review by at least one expert radiologist. This approach ensured high-quality ground truth labeling for accurate model training. The dataset consisted of PA and AP images from adult patients, with data sourced from multiple countries and acquired using equipment from 24+ X-ray device manufacturers, providing significant diversity across pathologies and imaging conditions.”

Comment 9:

You are missing the diagnostic performance metrics (AUC, sens, spec…) of the CAD in the adult population. I see it in figure 5, but this could be missed and should be in the main text to allow for comparison with what is achieved in pediatrics.

Response 9:

We thank the reviewer for this comment. We would like to clarify that the diagnostic performance metrics shown in Figure 5 reflect the performance of the AI tool in the pediatric population, using the recommended threshold of 15, which is based on the optimal threshold identified for adults. To improve clarity, we have now explicitly stated this in the main text and the figure. We would like to highlight that performance metrics for the adult population were not the focus of this analysis.

Page 14, Lines: 12 -23

“The vendor-recommended threshold of 15 to binarize the continuous AI output (ranging from 0-100) is based on the optimal threshold identified for adults. Therefore, as a final step in our analysis, we calculated optimized cut-offs for children for the entire cohort. Overall, the optimal cut-off for the pediatric age group was found to be 44.5. This adjusted threshold yielded a similar AUC of 0.93 with an accuracy of 89.8% (95% CI: 87.7–91.6%), sensitivity of 80.0%, and specificity of 92.4%. Compared to the adult threshold, the pediatric-specific cut-off led to improvement in the specificity and the PPV, supporting a safer AI interpretation in children with a slight trade-off in the sensitivity at 80%. The sub-group specific thresholds for pleuroparenchymal and mediastinal subgroups are shown in the Fig 4 and for each pathology are shown in S1 Fig.

Page 14, Lines: 25-32

“Fig 5: Definition of optimal thresholds for the performance of the AI-tool in children. The dotted blue line represents the pre-defined vendor recommended threshold of 15, which is based on the optimal threshold identified for adults to dichotomize the continuous AI-output (0-100). The green diamonds show optimized cut-offs calculated for maximizing the sum of sensitivity and specificity for pleuroparenchymal and mediastinal subgroups in the entire cohort. The performance metrics based on adult threshold of 15 (blue) and optimized cutoffs for children (green) are shown in the column on the right side.”

Comment 10:

When discussing threshold-based metrics (accuracy, sens, spec, PPV, NPV) it is good to remind the reader that this is benchmarked against the adult threshold.

Response 10:

We thank the reviewer for this suggestion. Accordingly, we added the following sentence to Statistical analysis:

Page: 9, Lines 23-24

“In addition, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV) were calculated after binarizing the continuous AI output using a predefined, vendor recommended threshold cut-off of 15, which was the optimal threshold identified for adults.”

We also changed the subheading in the results section as follows:

Page: 11, Line 9

“Performance analysis based on the thresholds recommended for adults:”

Comment 11:

The corrected analysis is interesting to highlight the imperfect reference standard (i.e., human radiologist). I think it highlights a limitation of the study that is not acknowledged in the discussion, that there was only one board-certified radiologist reviewing the X-rays compared to the 5 used in the adult study. This is evidenced by the fact that 6/7 corrected scores were actually when the AI was right but the labeller was wrong. I am not sure if re-calculating the diagnostic metrics for the corrected analysis makes sense as this is altering the definition of what the reference actually is. It may be sufficient to descriptively highlight these discrepant cases, and mention this lack of replication of radiological analysis as a limitation in the discussion.

Response 11:

We would like to thank the reviewer for this important observation. While one board-certified radiologist performed the reference readings, we would like to clarify that the original signed reports by pediatric radiologists were considered as second readings. Discrepant cases were subsequently reviewed and resolved through a consensus reading to establish the final reference standard for this study. We have acknowledged the limitation of not using a formal multi-reader consensus panel, in the Discussion section (see below)

Regarding the corrected analysis, we agree with the reviewer that recalculating diagnostic metrics based on retrospective case by case comparison could change the definition of the reference standard. We have therefore chosen not to include recalculated metrics and have thus omitted the paragraph “corrected performance analysis”. In keeping with this, we have removed the term “crude” performance analysis from the previous sections and have deleted the figure 2b. We now report these discrepancies descriptively in the Discussion.

Page: 17, Lines: 12-20

“This could have been partially be off-set with a formal multi-reader consensus panel for ground truth reading to improve the replication of radiological analysis. Furthermore, during a focused review of cases to identify potential false positives, we identified a few cases such a discrete atelectasis or borderline cardiomegaly, which were flagged by the AI tool but not marked in the reference reads. While it is also important to note that AI may “overdiagnose” subtle findings that experienced radiologists would reasonably judge as clinically insignificant or not warranting formal reporting, this could suggest a potential role of AI tools in highlighting findings that could be overlooked in the clinical routine.”

Comment 12:

“we calculated optimized cut-offs for each pathology and their performance metrics” – name the two pathology groups because I initially thought it was for each individual pathology (which is underpowered).

Response 12:

We would like to thank the reviewer for bringing this to our notice. We have changed the sentences as follows: (Page: 14, Lines: 21-23)

“The sub-group specific thresholds for pleuroparenchymal and mediastinal subgroups are shown in the Fig 4 and for each pathology are shown in S 1 Fig.”

Comment 13:

It would also be good to have an overall re-calculated thresholds for any abnormal vs normal for pediatric, as was done for adults.

Response 13:

We would like to thank the reviewer for this comment. We have now calculated the overall optimal threshold for children (Page:14, Lines: 17-21).

“Overall, the optimal cut-off for the pediatric age group was found to be 44.5. This adjusted threshold yielded a similar AUC of 0.93 with an accuracy of 89.8% (95% CI: 87.7–91.6%), sensitivity of 80.0%, and specificity of 92.4%. Compared to the adult threshold, the pediatric-specific cut-off led to improvement in the specificity and the PPV, supporting a safer AI interpretation in children with a slight trade-off in the sensitivity at 80%.”

We have updated the figure 4 accordingly (Page: 26)

Comment 14:

As previously mentioned, acknowledgement of a lack of replication of radiological interpretation in the reference needs to be acknowledged.

Response 14:

We thank the reviewer for highlighting this point. The following changes were made to the discussion. (Page: 17, Lines: 9-14)

“Finally, even though care was taken to generate high quality reference reads, minor variances in image interpretation such as missing/overdiagnosing slight pulmonary edema or atelectasis could not be avoided. This could have been partially be off-set with a formal multi-reader consensus panel for ground truth reading to improve the replication of radiological analysis.”

Comment 15:

Further limitations are that performance cannot be extrapolated to other respiratory conditions (e.g., tuberculosis CAD X-ray interpretation in adults vs in pediatrics cannot draw conclusions from this study). Could go along with the generalizability statement already made.

Response 15:

Thank you for the valid comment highlighting another important difference in disease manifestation in the pediatric population. We have added the following sentence to discussion (Page:16, Lines: 30-32):

“Furthermore, the performance of the AI tool cannot be extrapolated to certain respiratory conditions like tuberculosis, where radiological manifestation of TB are different among the pediatric and adult populations”

Comment 16:

Table 1: the pleuroparenchymal and mediastinal denominator should be 200. Would also be good to show the number who have both as a separate line .

Response 16:

We thank the reviewer for this point. We have updated the table 1 accordingly (Page: 17)

Parameter Total

(n = 958) Boys

(n = 524) Girls

(n = 434)

Age (years) * 6.75 ± 3.6 6.83 ± 3.6 6.65 ± 3.6

Projection (AP) 500/958 (52.2 %) 274/524 (52.3%) 226/434 (52.1%)

Relevant overall pathology 200/958 (20.9 %) 96/524 (18.3%) 104 /434 (23.4%)

• Pleuroparenchymal pathology † 162/200 (81.0 %) 83/200 (41.5 %) 79/200 (39.5 %)

• Mediastinal pathology ‡ 77/200 (38.5 %) 29/200 (14.5 %) 48/200 (24.0 %)

• Both mediastinal and pleuroparenchymal pathology 39/200 (19.5 %) 16/200 (8.0 %) 23/200 (11.5 %)

Comment 17:

Inconsistent N for the relevant pathology, pleuro, mediastinal across table 1, 2a and 2b. I expect the number in 1 to match 2a since 2b is the correct. However relevant overall in 1 is N=200 while in 2a is N=198, etc.

Response 17:

We would like to thank the reviewer for this and would like to apologize for this mistake. The values have been corrected (please refer to response 16).

Comment 18:

Figure 1: there is a lot of text, would be good to cut down. No definition of what (a) and (b) represent. Would be good to also show performance metrics for both adults and pediatrics so that it gives all the info in one place.

Response 18:

We thank the reviewer for this comment. The figure has been adjusted accordingly. The figure description was also changed to reflect the definition of (a) and (b). Since we have only evaluated the diagnostic performance in the pediatric population, we have only depicted the study methology in figure 1 and not included results. (Page: 6, Lines: 9-15)

Fig 1: Brief summary highlighting the study methology. a) AI Tool Development and Pediatric Repurposing: The AI tool was originally trained and validated using a large dataset of adult chest radiographs. For pediatric validation, the tool was retrospectively tested on 958 pediatric chest radiographs (CXR) from children aged 2–14 years.

b) Diagnostic Performance Analysis: The AI tool’s diagnostic performance in children was assessed using vendor-recommended thresholds, stratified by age groups (2–6 and 7–14 years), and optimized pediatric-specific thresholds.

Comment 19:

Figure 2: doesn’t add much to the manuscript, can be deleted.

Response 19:

We thank the reviewer for this feedback and have removed Figure 2. The numbering of other figures has been changed accordingly.

Response to Reviewer 2

We thank the reviewer for the critical evaluation of our work and are grateful for the

Attachment

Submitted filename: CXR3_point-to-point_withresponse.docx

pone.0328295.s006.docx (530.2KB, docx)

Decision Letter 1

Shahriar Ahmed

Deep Learning for Pediatric Chest X-Ray Diagnosis: Repurposing a Commercial Tool Developed for Adults

PONE-D-25-04141R1

Dear Dr. Agarwal,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice will be generated when your article is formally accepted. Please note, if your institution has a publishing partnership with PLOS and your article meets the relevant criteria, all or part of your publication costs will be covered. Please make sure your user information is up-to-date by logging into Editorial Manager at Editorial Manager®  and clicking the ‘Update My Information' link at the top of the page. If you have any questions relating to publication charges, please contact our Author Billing department directly at authorbilling@plos.org.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Shahriar Ahmed, MBBS, MHE, MPhil

Academic Editor

PLOS ONE

Additional Editor Comments (optional):

We would like to congratulate the authors for successfully addressing all comments raised by the reviewers. Thank you for considering this journal for publishing your important work. We hope that you will consider us again for your future publications.

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #1: All comments have been addressed

Reviewer #2: All comments have been addressed

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Yes

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #2: Yes

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: No

Reviewer #2: No

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: (No Response)

Reviewer #2: Thank you for addressing the comments. I am happy with the revisions and support the publication of this interesting manuscript.

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean? ). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy .

Reviewer #1: No

Reviewer #2: No

**********

Acceptance letter

Shahriar Ahmed

PONE-D-25-04141R1

PLOS ONE

Dear Dr. Agarwal,

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now being handed over to our production team.

At this stage, our production department will prepare your paper for publication. This includes ensuring the following:

* All references, tables, and figures are properly cited

* All relevant supporting information is included in the manuscript submission,

* There are no issues that prevent the paper from being properly typeset

You will receive further instructions from the production team, including instructions on how to review your proof when it is ready. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few days to review your paper and let you know the next and final steps.

Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

If we can help with anything else, please email us at customercare@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Dr. Shahriar Ahmed

Academic Editor

PLOS ONE

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 Table. Results of the algorithm performance in children aged 2–6 years.

    (DOCX)

    pone.0328295.s001.docx (18.5KB, docx)
    S2 Table. Results of the algorithm performance in children aged 7–14 years.

    (DOCX)

    pone.0328295.s002.docx (14.9KB, docx)
    S1 Fig. Flow charts showing the performance for all the relevant pathologies (a) in all age groups, (b) 7–14 years, and (c) 2–6 years.

    A) All relevant pathologies. B) Pleuroparenchymal pathologies. C) Mediastinal pathologies.

    (TIF)

    pone.0328295.s003.tif (390.3KB, tif)
    S2 Fig. Definition of optimal thresholds for the performance of the AI-tool in children.

    The dotted blue line represents the pre-defined vendor recommended threshold of 15, which is based on the optimal threshold identified for adults to dichotomize the continuous AI-output (0–100). The green diamonds show optimized cut-offs calculated for maximizing the sum of sensitivity and specificity. The performance metrics based on adult threshold of 15 (blue) and optimized cutoffs for children (green) are shown in the column on the right side (sens = sensitivity, spec = specificity, PPV = positive predictive value, NPV = negative predictive value)

    (TIF)

    pone.0328295.s004.tif (547.2KB, tif)
    Attachment

    Submitted filename: CXR3_point-to-point_withresponse.docx

    pone.0328295.s006.docx (530.2KB, docx)

    Data Availability Statement

    Due to institutional data privacy regulations and patient confidentiality at the University Hospital Freiburg, the data cannot be shared publicly. However, the data may be made available upon reasonable request by contacting the institution at rdia.studienzentrum@uniklinik-freiburg.de.


    Articles from PLOS One are provided here courtesy of PLOS

    RESOURCES