Abstract
Background
Sybil, an open-source deep learning model that uses low-dose CT (LDCT) for lung cancer prediction, requires rigorous external testing to confirm generalizability. Additionally, its utility in identifying individuals with high risk who never smoked or have light smoking histories remains unanswered.
Purpose
To externally test Sybil for identifying individuals with high risk for lung cancer within an Asian health checkup cohort.
Materials and Methods
This retrospective study analyzed LDCT scans from a single medical checkup facility in a study sample of individuals aged 50–80 years, collected between January 2004 and December 2021, with at least one follow-up scan. The predictive performance of the model for lung cancer risk over a 6-year period was assessed using the time-dependent area under the receiver operating characteristic curve (AUC). These evaluations were conducted in the overall study sample and within subgroups of patients with heavy (at least 20 pack-years) and never- or light smoking histories (ie, ever smoking [median, 2 pack-years]; ineligible for lung cancer screening per 2021 U.S. Preventive Services Task Force recommendations). Additionally, performance was evaluated according to the visibility of lung cancers on baseline LDCT scans.
Results:
Among 18 057 individuals (median age, 56 years [IQR, 52–61 years]; 11 267 male), 92 lung cancers were diagnosed (0.5%) within 6 years. Of these, 2848 had heavy smoking histories and 9943 had never- or light smoking histories, with 24 (0.8%) and 41 (0.4%) lung cancers, respectively. Sybil achieved AUCs of 0.91 for 1-year risk and 0.74 for 6-year risk. In the heavy-smoking subgroup, 1-year AUC was 0.94 (for visible lung cancers) and 6-year AUC was 0.70 (for future lung cancers). For the never- or light-smoking subgroup, Sybil had an AUC of 0.89 for visible lung cancers and 0.56 for future lung cancers.
Conclusion:
Sybil demonstrated excellent discriminative performance for visible lung cancers and acceptable performance for future lung cancers in Asian individuals with heavy smoking history but demonstrated poor performance for future lung cancers in a never- or light-smoking subgroup.
© RSNA, 2025
Supplemental material is available for this article.
See also the editorial by Jacobson and Byrne in this issue.
Summary
Sybil, an open-source deep learning model, predicted future lung cancer risk in Asian individuals with heavy smoking histories of at least 20 pack-years using low-dose CT scans, supporting its role in risk-based screening strategies, but demonstrated poor performance in a never- or light-smoking subgroup (median, 2 pack-years).
Key Results
■ In this retrospective study of 18 057 Asian individuals undergoing low-dose CT (LDCT) scans for health checkups, an open-source deep learning model (Sybil) underwent external testing to identify high-risk individuals (areas under the receiver operating characteristic curves [AUCs], 0.91 for 1-year risk; 0.74 for 6-year risk).
■ On baseline LDCT scans, Sybil demonstrated excellent predictive performance for visible lung cancers (AUC, 0.94) and acceptable performance for future lung cancers (AUC, 0.70) in a subgroup with at least 20 pack-years of smoking (n = 2848).
■ In contrast, among a never- or light-smoking subgroup with a median of 2 pack-years (n = 9943), Sybil showed predictive performance only for visible cancers (AUC, 0.89), with poor discrimination for future cancers (AUC, 0.56).
Introduction
The National Lung Screening Trial and the Dutch-Belgian Lung Cancer Screening Trial have demonstrated the efficacy of lung cancer screening using low-dose CT (LDCT) in individuals with a history of heavy smoking (at least 30 pack-years for the National Lung Screening Trial; >15 cigarettes daily for over 25 years or 10 cigarettes daily for over 30 years for the Dutch-Belgian Lung Cancer Screening Trial), resulting in a reduction of lung cancer mortality by 20% and 24%, respectively (1,2). Currently, evidence supporting lung cancer screening is primarily limited to this group (3).
Recently, an open-source deep learning model named Sybil (4) was developed to predict the risk of lung cancer within 6 years using LDCT scans from individuals with heavy smoking histories. In internal and external test cohorts of these individuals, Sybil achieved high predictive performance (C-index, 0.75–0.81) (4). It also showed comparable performance (C-index, 0.80) in a Taiwanese cohort, although the lack of smoking history raises the possibility that many participants never smoked (4,5). This uncertainty limits the interpretation of subgroup performance of Sybil among Asian individuals with heavy smoking histories (ie, individuals with at least 20 pack-years of smoking, including those who currently smoke and those who had quit within the past 15 years), light smoking history (ie, individuals who ever smoked but do not meet the lung cancer screening criteria for heavy smoking), and those who never smoked (3). To assess its generalizability, external testing in a cohort with a well-defined smoking history is essential. If Sybil performs well in Asian individuals who never smoked or have light smoking histories, where lung cancer incidence is increasing (6–13), it could advance lung cancer risk stratification and support its use beyond traditional high-risk populations.
Opportunistic LDCT screening, which takes place outside of systematic screening programs, is a common practice in East Asia and is driven by the proactive health-seeking behavior of individuals (14–16). This means that LDCT scans are acquired regardless of smoking status to detect various noncommunicable diseases (13,17–21). This unique clinical context offers a valuable opportunity to assess the applicability and performance of Sybil in populations underrepresented in traditional screening frameworks.
Thus, the aim of this study was to externally test Sybil for identifying individuals at high risk for lung cancer within an Asian health checkup cohort. This included individuals both eligible and ineligible for lung cancer screening according to the 2021 U.S. Preventive Services Task Force (USPSTF) recommendations (3). The feasibility of using Sybil to predict lung cancer across subgroups defined by smoking history, including heavy and never- or light smoking history, was investigated.
Materials and Methods
This retrospective study was approved by the institutional review board of Seoul National University Hospital, and the requirement for informed consent was waived (no. H-2304-017-1419). The sample of this study has not been previously reported. The schematic illustration of this study is depicted in Figure 1A.
Figure 1:
(A) Analytical framework of the study. For the entire study sample (n = 18 057), heavy-smoking subgroup (n = 2848), and never- or light-smoking subgroup (n = 9943), the following evaluations were performed: (a) the ability of Sybil, an open-source deep learning model, to predict lung cancer within 6 years (green box); (b) the predictive performance of Sybil stratified by the visibility of lung cancers on baseline CT scans (blue box); (c) a comparison of sensitivity and specificity of Sybil with those of the Lung CT Screening Reporting and Data System (Lung-RADS) v2022 (gray box); and (d) Sybil's attention maps (red box). Additionally, for the heavy-smoking subgroup, the clinical utility of Sybil in lung cancer screening was investigated by assessing the added value to the U.S. Preventive Services Task Force (USPSTF) recommendation (yellow box). None of the individuals included in this study were previously involved in the development or external testing of Sybil. AUC = area under the receiver operating characteristic curve. (B) Flowchart of the study sample.
Study Sample
This study was conducted at the Seoul National University Hospital Healthcare System Gangnam Center, a single medical checkup facility in Seoul, Korea. The facility provides a comprehensive screening program for noncommunicable diseases, within which LDCT is selectively performed based on individual request, symptoms, risk factors, or physician recommendation (22). Accordingly, this study included individuals of various ages and smoking histories, ranging from those who never smoked to those with heavy smoking histories (Appendix S1).
Consecutive individuals aged 50–80 years who underwent LDCT from January 2004 to December 2021, with at least one follow-up scan acquired at a minimum interval of 1 year from the initial scan, were included. For individuals with multiple scans, the first scan was selected as the input for the Sybil model. Then, the following exclusion criteria were applied: (a) unavailable Digital Imaging and Communications in Medicine files and (b) CT scans with image dimension or positioning inconsistencies or missing required tags (Appendix S2). Baseline data, including age, sex, smoking status, and comorbidities, were collected. Individuals with heavy smoking histories were defined as those with at least 20 pack-years of smoking, including those who currently smoke and those who had quit within the past 15 years, in accordance with the 2021 USPSTF recommendations for lung cancer screening (3). All other individuals who ever smoked but did not meet these criteria were classified as having light smoking histories.
For individuals diagnosed with lung cancer, assessments included tumor size (longest diameter) on CT scans, nodule type (pure ground-glass nodule, part-solid nodule, or solid nodule), clinical stage, histologic diagnosis, epidermal growth factor receptor mutation status, and treatment information. Survival status of individuals diagnosed with lung cancer was obtained from the Ministry of the Interior and Safety, Korea. Survival time was censored on February 14, 2025.
Details on the LDCT scan acquisition are described in Appendix S3. The individuals in this study were not included in the development or prior test cohorts of Sybil (4).
Sybil
Sybil, an open-source deep learning model, was used to predict lung cancer risk within 6 years from a single LDCT scan without requiring additional clinical information (4). The model outputs six probability scores corresponding to each year of risk. Attention maps were generated to visualize regions prioritized by the model during risk estimation (Appendix S4).
The Sybil model is publicly available at https://github.com/reginabarzilaygroup/Sybil.git. To clarify, this study conducted a dedicated external testing of Sybil without fine-tuning or transfer learning.
Outcome
The study outcome was the incidence of lung cancer following LDCT scans of individuals. Individuals, diagnosed with lung cancer (C33 or C34 according to the International Classification of Diseases 10th Revision) by March 6, 2023 (23), were identified through a search of the lung cancer registry at Seoul National University Hospital and Seoul National University Hospital Healthcare System Gangnam Center, and the dates of lung cancer diagnoses were recorded (24,25).
Comparison of Sybil and the Lung CT Screening Reporting and Data System v2022
This analysis included individuals with available pack-year data and excluded individuals with future lung cancers (4). Only individuals with visible lung cancer (true-positive findings) or no cancer during follow-up (true-negative findings) were analyzed. Given that the health checkup CT scans were not originally assessed using the Lung CT Screening Reporting and Data System (Lung-RADS), an age- and sex-matched case-control study was conducted. True-positive and true-negative cases were selected at a 1:3 ratio within two strata: heavy-smoking and never- or light-smoking. Matching was performed within each stratum and then combined for overall analysis.
Five board-certified thoracic radiologists (6–18 years of experience) independently interpreted the CT scans using Lung-RADS v2022, blinded to lung cancer diagnosis but not to sex and age. Each scan was reviewed once by one radiologist; therefore, consensus was not required. Lung-RADS scores of 1 or 2 were considered negative, and scores of 3 or 4 were classified as positive (4,26).
Statistical Analysis
To evaluate the predictive performance of Sybil for lung cancers diagnosed within 6 years following LDCT scans, the time-dependent area under the receiver operating characteristic curve (AUC) with 95% CI up to 1, 2, 3, 4, 5, and 6 years was calculated (27). Additionally, the Uno C-index with 95% CIs for lung cancer incidence over a 6-year period was evaluated (4). These analyses were conducted in the entire study sample, as well as separately in subgroups of heavy smoking and never- or light smoking. Additionally, sex-specific analyses were conducted for individuals with never- or light-smoking histories.
To assess whether Sybil functions as a nodule detector or a true risk predictor for baseline-invisible future lung cancers, its performance was evaluated according to the visibility of lung cancers at baseline LDCT. Analyses were performed in the overall sample, individuals with pack-year data, a history of heavy smoking, and a history of never- or light-smoking. For visible cancers, 1-year AUCs were calculated, as long-term predictions are less meaningful when lesions are already present. For future cancers, 6-year AUCs were assessed to evaluate the potential of Sybil for long-term risk prediction.
For performance comparison between Sybil and Lung-RADS, threshold-based analyses were conducted by aligning either sensitivity or specificity between the two methods. First, the Sybil threshold that achieved the same (or the closest possible) sensitivity as Lung-RADS was identified, and the corresponding specificity values were compared. Conversely, the Sybil threshold with the same specificity as Lung-RADS was selected, and the associated sensitivities were compared. Point estimates and 95% CIs were calculated for sensitivity and specificity. Statistical significance was assessed using the McNemar test on the overall matched subset and within heavy-smoking and never- or light-smoking subsets.
A 6-year Sybil risk score cutoff yielding approximately 85% sensitivity was selected empirically, reflecting the baseline screening sensitivity reported in the National Lung Screening Trial (post hoc analysis using Lung-RADS) and Dutch-Belgian Lung Cancer Screening Trial (28,29). With this cutoff, the added value of refining screening eligibility as defined by the USPSTF was evaluated by incorporating Sybil-based risk stratification. The number needed to screen to detect one case of lung cancer was calculated as the total screened divided by true-positives. The trade-off between sensitivity and specificity was also assessed by quantifying the number of false-positives avoided relative to the number of true-positives missed.
All statistical analyses were performed using R, version 4.3.2 (packages: timeROC and survC1; R Foundation for Statistical Computing; https://www.R-project.org), and P < .05 was considered to indicate statistical significance. P values for multiple comparisons were adjusted with use of the Benjamini-Hochberg procedure, with the significance threshold remaining unchanged. Additional details are provided in Appendix S5.
Results
Baseline Characteristics of the Study Sample
Of 19 569 individuals aged 50–80 years who underwent LDCT scans with at least a 1-year interval between January 2004 and December 2021, one individual was excluded due to unavailable files and 1511 were excluded due to CT scans with image inconsistencies or missing required tags. After exclusions, a total of 18 057 individuals (median age, 56 years [IQR, 52–61 years]; 11 267 male, 6790 female) were included (Fig 1B, Table 1). Of these, 1.3% (228 of 18 057) were diagnosed with lung cancer over a median follow-up of 6.0 years (IQR, 3.0–10.3 years). Details for 228 lung cancers are described in Table S1. In the specified 6-year follow-up, 92 lung cancers were identified (0.5%). The heavy-smoking subgroup accounted for 15.8% (2848 of 18 057), with 24 lung cancer cases (0.8%) observed over the 6-year period. The never- or light-smoking subgroup comprised 55.1% (9943 of 18 057) of the study sample, with 41 lung cancer cases (0.4%) identified within 6 years: 0.4% among those who never smoked (20 of 5058) and 0.4% among those with light smoking histories (21 of 4885). The light-smoking group consisted of 1234 individuals who currently smoke and 3651 who formerly smoked, with a median of 2.0 pack-years of smoking (IQR, 1.0–10.0 pack-years). The remaining 5266 individuals (29.2%) could not be categorized due to missing pack-year information. Survival rates of individuals with lung cancers according to the clinical stage are described in Table S2 and Figure S1.
Table 1:
Baseline Characteristics of the Study Sample and Test Datasets of the Original Model Development Study
| Test Sets in the Original Model Development Study (4) | ||||||||
|---|---|---|---|---|---|---|---|---|
| This Study | NLST | MGH | CGMH | |||||
| Variable | Study Sample | Lung Cancer | Study Sample | Lung Cancer | Study Sample | Lung Cancer | Study Sample | Lung Cancer |
| Total | 18 057* | 228 (1.3)*, 92 (0.5)† | 6282 | 299 (4.8) | 8821 | 255 (2.9) | 12 280 | 126 (1.0) |
| Age (y) | ||||||||
| <50 | 0 (0) | 0 (0) | 0 (0) | 0 (0) | 9 (0.1) | 0 (0) | 4296 (35.0) | 24 (19.1) |
| 50–60 | 12 649 (70.1) | 126 (55.3) | 2318 (36.9) | 77 (25.8) | 2044 (23.2) | 63 (24.7) | 4258 (34.7) | 42 (33.3) |
| 60–70 | 4517 (25.0) | 89 (39.0) | 3212 (51.1) | 169 (56.5) | 4563 (51.7) | 139 (54.5) | 2878 (23.4) | 33 (26.2) |
| 70–80 | 891 (4.9) | 13 (5.7) | 752 (12.0) | 53 (17.7) | 2155 (24.4) | 52 (20.4) | 722 (5.9) | 19 (15.1) |
| >80 | 0 (0) | 0 (0) | 0 (0) | 0 (0) | 49 (0.6) | 1 (0.4) | 126 (1.0) | 8 (6.3) |
| Sex | ||||||||
| M | 11 267 (62.4) | 157 (68.9) | 3769 (60.0) | 190 (63.5) | 4662 (52.9) | 104 (40.8) | 7134 (58.1) | 59 (46.8) |
| F | 6790 (37.6) | 71 (31.1) | 2513 (40.0) | 109 (36.5) | 4159 (47.1) | 151 (59.2) | 5146 (41.9) | 67 (53.2) |
| Smoking status‡ | ||||||||
| Never | 5939 (41.2) | 66 (31.3) | 0 (0) | 0 (0) | NA | NA | NA | NA |
| Ever | 8486 (58.8) | 145 (68.7) | 6282 (100) | 299 (100) | NA | NA | NA | NA |
| Body mass index‡§ | 23.6 (21.9–25.5) | Not applicable | NA | NA | NA | NA | NA | NA |
| Presence of pulmonary disease history‡ | 1436 (14.3) | 28 (15.6) | NA | NA | NA | NA | NA | NA |
| Presence of cancer history‡ | 714 (7.5) | 18 (11.0) | NA | NA | NA | NA | NA | NA |
| Time to last follow-up (y) | ||||||||
| <1 | 17 (0.1) | 17 (7.5) | 101 (1.6) | 82 (27.4) | 3175 (36.0) | 152 (59.6) | 1549 (12.6) | 12 (9.5) |
| ≥1, <2 | 2345 (13.0) | 13 (5.7) | 81 (1.3) | 48 (16.1) | 2392 (27.1) | 57 (22.4) | 2799 (22.8) | 26 (20.6) |
| ≥2, <3 | 2081 (11.5) | 14 (6.1) | 112 (1.8) | 52 (17.4) | 1795 (20.3) | 29 (11.4) | 2538 (20.7) | 35 (27.8) |
| ≥3, <4 | 1855 (10.3) | 15 (6.6) | 294 (4.7) | 43 (14.4) | 986 (11.2) | 13 (5.1) | 1743 (14.2) | 24 (19.0) |
| ≥4, <5 | 1513 (8.4) | 16 (7.0) | 1586 (25.2) | 50 (16.7) | 473 (5.4) | 4 (1.6) | 1650 (13.4) | 18 (14.3) |
| ≥5, <6 | 1285 (7.1) | 17 (7.5) | 4108 (65.4) | 24 (8.0) | 0 (0) | 0 (0) | 2001 (16.3) | 11 (8.7) |
| ≥6 | 8961 (49.6) | 136 (59.6) | NI | NI | NI | NI | NI | NI |
Note.—Information from the National Lung Cancer Screening Trial (NLST), Massachusetts General Hospital (MGH), and Chang Gung Memorial Hospital (CGMH) was directly referenced from the original model development study. Unless otherwise specified, data are numbers of patients, with percentages in parentheses. NA = not available, NI = not included.
Individuals and lung cancers of the entire follow-up period (median, 6.0 years [IQR, 3.0–10.3 years]).
Lung cancers observed in 6 years.
Smoking status was available for 14 425 individuals, body mass index (calculated as weight in kilograms divided by height in meters squared) for 16 426 individuals, pulmonary disease history for 10 032 individuals, and cancer history for 9560 individuals in this study sample.
Body mass index is reported as the median, with IQR in parentheses.
Performance of Sybil to Predict Lung Cancer within 6 Years
The 1–6-year Sybil risk scores are summarized in Table S3. In the entire sample (n = 18 057), the AUCs were 0.91 (95% CI: 0.83, 1.00) for 1-year risk and 0.74 (95% CI: 0.68, 0.79) for 6-year risk (Table 2). In the heavy-smoking subgroup (n = 2848), AUCs were 0.94 (95% CI: 0.93, 0.95) for 1-year risk and 0.73 (95% CI: 0.63, 0.83) for 6-year risk. In the never- or light-smoking subgroup (n = 9943), Sybil achieved a 1-year AUC of 0.90 (95% CI: 0.78, 1.00) and a 6-year AUC of 0.71 (95% CI: 0.62, 0.79).
Table 2:
Performance of Sybil, an Open-Source Deep Learning Model, in Predicting Lung Cancer from Low-Dose Chest CT Scans
| Subgroup | 1-year Risk | 2-year Risk | 3-year Risk | 4-year Risk | 5-year Risk | 6-year Risk | Uno C-Index |
|---|---|---|---|---|---|---|---|
| Entire study sample (n = 18 057) | 0.91 (0.83, 1.00) | 0.87 (0.79, 0.95) | 0.86 (0.81, 0.92) | 0.79 (0.72, 0.85) | 0.76 (0.70, 0.82) | 0.74 (0.68, 0.79) | 0.73 (0.67, 0.79) |
| History of heavy smoking eligible for the USPSTF recommendation (n = 2848)* | 0.94 (0.93, 0.95) | 0.75 (0.50, 1.00) | 0.81 (0.69, 0.93) | 0.74 (0.62, 0.86) | 0.70 (0.59, 0.82) | 0.73 (0.63, 0.83) | 0.72 (0.62, 0.83) |
| History of never- or light-smoking ineligible for the USPSTF recommendation (n = 9943)* | 0.90 (0.78, 1.00) | 0.90 (0.79, 1.00) | 0.87 (0.77, 0.96) | 0.77 (0.65, 0.89) | 0.74 (0.63, 0.85) | 0.71 (0.62, 0.79) | 0.70 (0.61, 0.80) |
Note.—Unless otherwise specified, data are time-dependent areas under the receiver operating characteristic curves. Data in parentheses are 95% CIs. Sybil demonstrated good performance in external testing among Asian individuals undergoing low-dose CT for health checkups, including the entire study sample and subgroups with heavy smoking history or never- or light-smoking history. Smoking histories for 5266 of 18 057 individuals (29.2%) could not be categorized due to missing pack-year smoking information. The Sybil lung risk score is a deep learning–based risk score for predicting future lung risk within 6 years from a low-dose chest CT examination. USPSTF = U.S. Preventive Services Task force.
The USPSTF recommends lung cancer screening for individuals aged 50–80 years who have a 20-pack-year smoking history. Individuals with heavy smoking history were defined as those with at least 20 pack-years of smoking, including those who currently smoke and those who had quit within the past 15 years. All other individuals who ever smoked but did not meet these criteria were classified as having light smoking history (median, 2 pack-years in this study). Individuals with a never-smoking history were defined as those who had never smoked.
Representative Sybil attention maps are shown in Figures 2, 3, and S2–S4 and are explained in Appendix S6. The results of the separate assessment for the never- or light-smoking subgroup and sex-specific analysis are provided in Tables S4 and S5, respectively. Analyses on lung cancers are detailed in Appendix S7 and Tables S6 and S7.
Figure 2:
Examples of Sybil, an open-source deep learning model, correctly predicting lung cancers (future lung cancer). (A, B) Axial low-dose CT image of the lungs in a 55-year-old heavy-smoking man shows no abnormality, but the attention map of Sybil focused on the right upper lobe (arrow). The Sybil lung risk scores were 1.1% for 1-year risk, 2.4% for 2-year risk, 4.2% for 3-year risk, 5.6% for 4-year risk, 6.8% for 5-year risk, and 10.5% for 6-year risk. (C) Axial CT image after 5 years shows that the patient had a 7.3-cm mass in the right upper lobe, diagnosed as lung cancer (arrow; adenocarcinoma). The Sybil lung risk score is a deep learning–based risk score for predicting future lung cancer risk within 6 years from a low-dose chest CT scan.
Figure 3:
Examples of Sybil, an open-source deep learning model, correctly predicting lung cancers (visible lung cancer at baseline CT). (A) Axial low-dose CT image of the lungs in a 54-year-old heavy-smoking man shows a 1.5-cm subsolid nodule in the right lower lobe, with the (B) attention map focusing on this area (arrow). The Sybil lung risk scores were 13.1% for 1-year risk, 20.2% for 2-year risk, 20.7% for 3-year risk, 23.6% for 4-year risk, 24.8% for 5-year risk, and 31.2% for 6-year risk. (C) Axial CT image after 2 years shows that the subsolid nodule had grown to 2.0 cm and was subsequently diagnosed as adenocarcinoma (arrow). The Sybil lung risk score is a deep learning–based risk score for predicting future lung cancer risk within 6 years from a low-dose chest CT scan.
Predictive Performance According to Visibility of Lung Cancers at Baseline CT
Of the 92 lung cancers within a 6-year follow-up, 64.1% (59 of 92) were visible at baseline, including eight cases in the heavy-smoking subgroup and 30 cases in the never- or light-smoking subgroup. The remaining 35.9% (33 of 92) were future lung cancers, comprising 16 cases in the heavy-smoking subgroup and 11 cases in the never- or light-smoking subgroup. For visible lung cancers, Sybil demonstrated robust predictive performance. The AUCs were 0.91 (95% CI: 0.82, 1.00) in the entire study sample, 0.94 (95% CI: 0.93, 0.95) in the heavy-smoking subgroup, and 0.89 (95% CI: 0.76, 1.00) in the never- or light-smoking subgroup (Table 3).
Table 3:
Performance of Sybil, an Open-Source Deep Learning Model, for Predicting Lung Cancer from Low-Dose Chest CT Scans Stratified by Lung Cancer Visibility on Baseline CT Scans and Smoking Status
| Visibility of Lung Cancers at Baseline CT and Group | AUC* |
|---|---|
| Visible lung cancers | |
| Entire study sample (n = 17 969) | 0.91 (0.82, 1.00) |
| Individuals with available pack-year information (n = 12 728) | 0.89 (0.77, 1.00) |
| Heavy smoking (n = 2836) | 0.94 (0.93, 0.95) |
| Never- or light-smoking (n = 9892) | 0.89 (0.76, 1.00) |
| Future lung cancers | |
| Entire study sample (n = 17 917) | 0.67 (0.59, 0.75) |
| Individuals with available pack-year information (n = 12 668) | 0.65 (0.55, 0.74) |
| Heavy smoking (n = 2797) | 0.70 (0.58, 0.82) |
| Never- or light-smoking (n = 9871) | 0.56 (0.43, 0.70) |
Note.—Data in parentheses are 95% CIs. Sybil demonstrated robust predictive performance for visible lung cancers and good predictive performance for future lung cancers on baseline low-dose CT scans in the heavy-smoking subgroup. However, in the never- or light-smoking subgroup, its predictive performance was limited to visible cancers, with reduced discrimination for future cancers. Individuals with heavy smoking history were defined as those with at least 20 pack-years of smoking, including those who currently smoke and those who had quit within the past 15 years. All other individuals who ever smoked but did not meet these criteria were classified as having light smoking history (median, 2 pack-years in this study). Individuals with a never-smoking history were defined as those who had never smoked. For evaluation of visible lung cancer, future lung cancers on baseline CT scans were excluded. For evaluation of future lung cancer, visible lung cancers on baseline CT scans were excluded. The Sybil lung risk score is a deep learning–based risk score for predicting future lung risk within 6 years from a low-dose chest CT scan. AUC = time-dependent area under the receiver operating characteristic curve.
One-year AUC for visible lung cancers and 6-year AUC for future lung cancers on baseline CT scans.
However, its performance declined for future lung cancers, with AUCs of 0.67 (95% CI: 0.59, 0.75) in the overall sample, 0.70 (95% CI: 0.58, 0.82) in the heavy-smoking subgroup, and 0.56 (95% CI: 0.43, 0.70) in the never- or light-smoking subgroup (Table 3), indicating limited and statistically nonsignificant discriminative ability in the never- or light-smoking subgroup.
Comparison with Lung-RADS v2022
A subset of 250 individuals was selected (46 heavy-smoking and 204 never- or light-smoking). This subset comprised 63 cases of visible lung cancer and 187 noncancer controls. After matching, there was no evidence of differences between Sybil and Lung-RADS for either specificity (86.6% vs 95.2%; P = .14) or sensitivity (27.0% vs 44.4%; P = .19). Sybil had lower specificity than Lung-RADS when sensitivity was matched in a never- or light-smoking subset (83.0% vs 96.7%; P = .002) (Table 4).
Table 4:
Comparison of Specificity and Sensitivity of Sybil, an Open-Source Deep Learning Model, with Those of Lung-RADS v2022
| A: Matched Sensitivity | ||||
|---|---|---|---|---|
| Subgroup | Matched Sensitivity (%) | Specificity, Lung-RADS (%) | Specificity, Sybil Model (%) | P Value |
| Entire subset (n = 250) | 44.4 (31.9, 57.5) [28/63] | 95.2 (91.1, 97.8) [178/187] | 86.6 (80.9, 91.2) [162/187]* | .14 |
| Heavy smoking (n = 46) | 33.3 (9.9, 65.1) [4/12] | 88.2 (72.5, 96.7) [30/34] | 97.1 (84.7, 99.9) [33/34]† | .47 |
| Never- or light-smoking (n = 204) | 47.1 (32.9, 61.5) [24/51] | 96.7 (92.5, 98.9) [148/153] | 83.0 (76.1, 88.6) [127/153] | .002 |
| B: Matched Specificity | ||||
|---|---|---|---|---|
| Subgroup | Matched Sensitivity (%) | Specificity, Lung-RADS (%) | Specificity, Sybil Model (%) | P Value |
| Entire subset (n = 250) | 95.2 (91.1, 97.8) [178/187] | 44.4 (31.9, 57.5) [28/63] | 27.0 (16.6, 39.7) [17/63] | .19 |
| Heavy smoking (n = 46) | 88.2 (72.5, 96.7) [30/34] | 33.3 (9.9, 65.1) [4/12] | 50.0 (21.1, 78.9) [6/12] | >.99 |
| Never- or light-smoking (n = 204) | 96.7 (92.5, 98.9) [148/153] | 47.1 (32.9, 61.5) [24/51] | 25.5 (14.3, 39.6) [13/51] | .07 |
Note.—Data in parentheses are 95% CIs. Data in brackets indicate the numerator and denominator of true-positive findings over total positive findings for sensitivity and true-negative findings over total negative findings for specificity, as determined by the Lung CT Screening Reporting and Data System (Lung-RADS) or Sybil. Statistical comparisons were performed with the McNemar test. In a subset of 250 individuals—comprising 63 cases of visible lung cancer and 187 noncancer controls—no significant differences were observed between Sybil and Lung-RADS in terms of sensitivity or specificity. However, in the never- or light-smoking subgroup, Sybil showed lower specificity than Lung-RADS when sensitivity was matched. Individuals with heavy smoking history were defined as those with at least 20 pack-years of smoking, including those who currently smoke and those who had quit within the past 15 years. All other individuals who ever smoked but did not meet these criteria were classified as having a light smoking history (median, 2 pack-years in this study). Individuals with a neversmoking history were defined as those who had never smoked. An age- and sex-matched case-control sample was obtained at a 1:3 ratio (n = 250). Individuals with visible lung cancer were classified as true-positive cases (n = 63), while those without incident lung cancer were classified as true-negative cases (n = 187).
As an exact match in sensitivity with Lung-RADS was not available, the closest sensitivity (41.3%) was selected (Sybil cutoff, 1.2%).
As an exact match in sensitivity with Lung-RADS was not available, the closest sensitivity (25.0%) was selected (Sybil cutoff, 3.9%).
Added Value of Sybil to the USPSTF Recommendations
Among the heavy-smoking subgroup (n = 2848), 2840 individuals remained eligible for postbaseline screening under the 2021 USPSTF recommendations after the exclusion of eight baseline-visible lung cancer cases, leading to the detection of 16 additional cases. Applying a Sybil 6-year risk score cutoff of 3.3% to this group enabled further risk stratification. This strategy reduced the postbaseline screening cohort to 1779 individuals, detecting 13 of the 16 lung cancer cases. The three missed cancers developed 1.7, 3.5, and 4.9 years after the baseline. It reduced the number of false-positive findings by 1058. The trade-off ratio of 352.7:1 (1058/3) indicates that for each additional missed lung cancer case, 353 false-positive screening examinations could be avoided. The number needed to screen decreased from 177.5 (2840/16) under the USPSTF criteria alone to 136.8 (1779/13) with the combined strategy.
Discussion
Sybil is an open-source deep learning model that predicts lung cancer risk from LDCT, but its generalizability, particularly in individuals who never smoked or have light smoking histories, remains unclear. Thus, we conducted a retrospective study in an Asian health checkup sample that included individuals with smoking histories ranging from those who never smoked to those with heavy smoking histories of at least 20 pack-years. Sybil exhibited robust predictive performance for lung cancers, achieving an area under the receiver operating characteristic curve (AUC) of 0.91 for 1-year risk and good performance for 6-year risk, with an AUC of 0.74. In the heavy-smoking subgroup, the model demonstrated consistent performance, with an AUC of 0.94 for visible lung cancers and 0.70 for future lung cancers. Although Sybil achieved an AUC of 0.89 for visible lung cancers in the never- or light-smoking subgroup with a median of 2.0 pack-years, nonsignificant discriminative performance was demonstrated for future lung cancers (AUC, 0.56) in this group.
In the original model development study (4), Sybil showed higher specificity than Lung-RADS for visible lung cancers in the National Lung Screening Trial cohort (92% vs 86%) (4). However, the benefit of Sybil compared with Lung-RADS was not evident in our study, suggesting that its role in the assessment of baseline screening CT may be limited. Instead, the primary value of Sybil may lie beyond baseline evaluation, particularly in refining follow-up screening intervals or determining continued screening. Indeed, the three lung cancers missed in the heavy-smoking subgroup using 6-year risk score cutoff of Sybil of 3.3% developed 1.7, 3.5, and 4.9 years after the baseline LDCT examination. This suggests that integrating Sybil with USPSTF criteria may help tailor screening intervals; individuals classified as low risk by Sybil could be candidates for biennial or triennial screening. This approach may improve the cost-effectiveness of lung cancer screening. Several studies have proposed model-based strategies to adjust screening intervals. Models such as PLCO2012results and LCRAT+CT (30,31), which rely on clinical, demographic, and LDCT findings, have shown promise but did not incorporate full imaging data. More recently, a deep learning model was developed to predict 1-year lung cancer risk in individuals with lung nodules (32); however, it is not applicable to those without nodules. Sybil provides an alternative by using the entire baseline LDCT examination—regardless of nodule presence—to predict 6-year risk. Future studies should directly compare Sybil with clinical models and assess whether adding clinical variables could enhance its accuracy.
In individuals with heavy smoking histories, Sybil achieved an AUC of 0.70 for baseline-invisible future lung cancers. However, the attention maps of model showed limited ability to localize the eventual tumor sites in these cases, which raises concerns about the interpretability and clinical credibility of its predictions. This discrepancy between risk score–based performance and spatial localization underscores a limitation of attention-based explanations, particularly for cancers not yet visible at imaging.
In a never- or light-smoking subgroup, Sybil demonstrated comparable performance for visible lung cancers but limited predictive capability for future lung cancer risk. This limitation may arise from fundamental biologic differences between individuals who smoked heavily and those who never smoked or smoked lightly in terms of the anatomic or structural changes, such as emphysema, that are associated with lung cancer development (33–35). It is plausible that Sybil, trained on data from individuals with heavy smoking histories, did not sufficiently capture the distinct biologic characteristics and subtle imaging features relevant to lung carcinogenesis in those who never smoked or smoked lightly, in whom lung cancers more frequently manifest as subsolid nodules, often corresponding to epidermal growth factor receptor–mutated adenocarcinomas (6).
Several limitations should be noted. First, the retrospective study design might have introduced biases that could affect the validity of the findings. Second, we externally tested Sybil in an Asian health checkup sample undergoing opportunistic LDCT examinations, a setting that may be specific to East Asia and less common elsewhere. Additionally, our study sample consisted of health-conscious individuals who were financially able to undergo health checkups, which may have introduced selection bias. Third, the relatively small number of lung cancers might have limited the statistical power. Fourth, the comparison between Sybil and Lung-RADS was conducted using a matched case-control subset with an artificially elevated cancer prevalence. Although this does not affect sensitivity or specificity, which are prevalence-independent, it limits the generalizability to real-world screening populations. Finally, a key limitation is the absence of standardized follow-up protocols, as the study was based on a single-center real-world registry. Data on participant retention, loss to follow-up, or emigration were unavailable, possibly leading to an underestimation of true lung cancer incidence. However, as a large-scale cohort study from Korea reported lung cancer rates of 0.5% for individuals who never smoked and 0.6% for those who ever smoked (17), the extent of underestimation in our study may not be substantial.
In conclusion, Sybil, an open-source deep learning model, demonstrated potential for future lung cancer prediction among Asian individuals with heavy smoking histories of at least 20 pack-years and may support optimization of follow-up intervals. However, its utility for predicting future lung cancer in those who never smoked or smoked lightly was limited.
Acknowledgments
Acknowledgment
The authors acknowledge Soon Ho Yoon, MD, PhD; Eui Jin Hwang, MD, PhD; Woo Hyeon Lim, MD; Ji Young Lee, MD, PhD; and Hye Soo Cho, MD (all from Seoul National University Hospital) for their valuable assistance in interpreting low-dose chest CT scans using Lung-RADS v2022 in this study.
J.H.L. and K.J.C. contributed equally to this work.
S.H.C. and H.K. are co-senior authors.
Funding: This study was supported by the National Research Foundation of Korea (NRF) grant funded by the Korea government (Ministry of Science and ICT, South Korea) (no. RS-2023-00207978; NRF-2021R1C1C1009818). However, the funder had no role in the study design; in the collection, analysis, and interpretation of the data; in the writing of the report; or in the decision to submit the article for publication.
Data sharing: Data generated or analyzed during the study are available from the corresponding author by request.
Disclosures of conflicts of interest: J.H.L. Research grant from Corelinesoft and Kakao Brain; consulting fees from RadiSen. K.J.C. No relevant relationships. M.T.L. Grants to institution from the American Heart Association; AstraZeneca; Ionis; Johnson & Johnson Innovation; Kowa Pharmaceuticals America; National Academy of Medicine; National Heart, Lung, and Blood Institute; and Risk Management Foundation of the Harvard Medical Institutions. Y.C.C. No relevant relationships. S.L. No relevant relationships. J.M.G. Research grants from Corelinesoft and Taejoon Pharm; in an associate editor for Radiology. S.H.C. No relevant relationships. H.K. Grants to institution from Kakao Brain and RadiSen; consulting fees to author and institution from RadiSen; stock and stock options in Medical IP and Soombit.ai; medical director of Soombit.ai; member of the editorial board for Radiology Advances.
Abbreviations:
- AUC
- area under the receiver operating characteristic curve
- LDCT
- low-dose CT
- Lung-RADS
- Lung CT Screening Reporting and Data System
- USPSTF
- U.S. Preventive Services Task Force
References
- 1. Aberle DR , Adams AM , Berg CD , et al. ; National Lung Screening Trial Research Team . Reduced lung-cancer mortality with low-dose computed tomographic screening . N Engl J Med 2011. ; 365 ( 5 ): 395 – 409 . [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. de Koning HJ , van der Aalst CM , de Jong PA , et al . Reduced lung-cancer mortality with volume CT screening in a randomized trial . N Engl J Med 2020. ; 382 ( 6 ): 503 – 513 . [DOI] [PubMed] [Google Scholar]
- 3. Krist AH , Davidson KW , Mangione CM , et al. ; US Preventive Services Task Force . Screening for lung cancer: US Preventive Services Task Force recommendation statement . JAMA 2021. ; 325 ( 10 ): 962 – 970 . [DOI] [PubMed] [Google Scholar]
- 4. Mikhael PG , Wohlwend J , Yala A , et al . Sybil: a validated deep learning model to predict future lung cancer risk from a single low-dose chest computed tomography . J Clin Oncol 2023. ; 41 ( 12 ): 2191 – 2200 . [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Welch HG , Gao W , Wilder FG , Kim SY , Silvestri GA . Lung cancer screening in people who have never smoked: lessons from East Asia . BMJ 2025. ; 388 : e081674 . [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Lam DCL , Liam CK , Andarini S , et al . Lung cancer screening in Asia: an expert consensus report . J Thorac Oncol 2023. ; 18 ( 10 ): 1303 – 1322 . [DOI] [PubMed] [Google Scholar]
- 7. Pelosof L , Ahn C , Gao A , et al . Proportion of never-smoker non–small cell lung cancer patients at three diverse institutions . J Natl Cancer Inst 2017. ; 109 ( 7 ): djw295 . [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Cufari ME , Proli C , De Sousa P , et al . Increasing frequency of non-smoking lung cancer: presentation of patients with early disease to a tertiary institution in the UK . Eur J Cancer 2017. ; 84 : 55 – 59 . [DOI] [PubMed] [Google Scholar]
- 9. Sung H , Ferlay J , Siegel RL , et al . Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries . CA Cancer J Clin 2021. ; 71 ( 3 ): 209 – 249 . [DOI] [PubMed] [Google Scholar]
- 10. Khan S , Hatton N , Tough D , et al . Lung cancer in never smokers (LCINS): development of a UK national research strategy . BJC Rep 2023. ; 1 ( 1 ): 21 . [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. LoPiccolo J , Gusev A , Christiani DC , et al . Lung cancer in patients who have never smoked—an emerging disease . Nat Rev Clin Oncol 2024. ; 21 ( 2 ): 121 – 146 . [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Siegel DA , Fedewa SA , Henley SJ , et al . Proportion of never smokers among men and women with lung cancer in 7 US states . JAMA Oncol 2021. ; 7 ( 2 ): 302 – 304 . [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Chang GC , Chiu CH , Yu CJ , et al. ; TALENT Investigators . Low-dose CT screening among never-smokers with or without a family history of lung cancer in Taiwan: a prospective cohort study . Lancet Respir Med 2024. ; 12 ( 2 ): 141 – 152 . [DOI] [PubMed] [Google Scholar]
- 14. Kang HR , Cho JY , Lee SH , et al . Role of low-dose computerized tomography in lung cancer screening among never-smokers . J Thorac Oncol 2019. ; 14 ( 3 ): 436 – 444 . [DOI] [PubMed] [Google Scholar]
- 15. Zhang Y , Jheon S , Li H , et al . Results of low-dose computed tomography as a regular health examination among Chinese hospital employees . J Thorac Cardiovasc Surg 2020. ; 160 ( 3 ): 824 – 831.e4 . [DOI] [PubMed] [Google Scholar]
- 16. Kakinuma R , Muramatsu Y , Asamura H , et al . Low-dose CT lung cancer screening in never-smokers and smokers: results of an eight-year observational study . Transl Lung Cancer Res 2020. ; 9 ( 1 ): 10 – 22 . [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Kim YW , Kang HR , Kwon BS , et al . Low-dose chest computed tomographic screening and invasive diagnosis of pulmonary nodules for lung cancer in never-smokers . Eur Respir J 2020. ; 56 ( 5 ): 2000177 . [DOI] [PubMed] [Google Scholar]
- 18. Kondo R , Yoshida K , Kawakami S , et al . Efficacy of CT screening for lung cancer in never-smokers: analysis of Japanese cases detected using a low-dose CT screen . Lung Cancer 2011. ; 74 ( 3 ): 426 – 432 . [DOI] [PubMed] [Google Scholar]
- 19. Kim HY , Jung KW , Lim KY , et al . Lung cancer screening with low-dose CT in female never smokers: retrospective cohort study with long-term national data follow-up . Cancer Res Treat 2018. ; 50 ( 3 ): 748 – 756 . [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Wang L , Qi Y , Liu A , et al . Opportunistic screening with low-dose computed tomography and lung cancer mortality in China . JAMA Netw Open 2023. ; 6 ( 12 ): e2347176 . [Retraction in JAMA Netw Open. 2024 Sep 3;7(9):e2438532. doi: 10.1001/jamanetworkopen.2024.38532.] [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Tang W , Liu L , Huang Y , et al . Opportunistic lung cancer screening with low‐dose computed tomography in National Cancer Center of China: the first 14 years’ experience . Cancer Med 2024. ; 13 ( 3 ): e6914 . [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Lee C , Choe EK , Choi JM , et al . Health and Prevention Enhancement (H-PEACE): a retrospective, population-based cohort study conducted at the Seoul National University Hospital Gangnam Center, Korea . BMJ Open 2018. ; 8 ( 4 ): e019327 . [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. World Health Organization . ICD-10 Version:2019 Web site . https://icd.who.int/browse10/2019/en. Updated October, 2019. Accessed March 6, 2023.
- 24. Lee JH , Sun HY , Park S , et al . Performance of a deep learning algorithm compared with radiologic interpretation for lung cancer detection on chest radiographs in a health screening population . Radiology 2020. ; 297 ( 3 ): 687 – 696 . [DOI] [PubMed] [Google Scholar]
- 25. Lee JH , Lee D , Lu MT , et al . Deep learning to optimize candidate selection for lung cancer CT screening: advancing the 2021 USPSTF recommendations . Radiology 2022. ; 305 ( 1 ): 209 – 218 . [DOI] [PubMed] [Google Scholar]
- 26. Christensen J , Prosper AE , Wu CC , et al . ACR Lung-RADS v2022: assessment categories and management recommendations . J Am Coll Radiol 2024. ; 21 ( 3 ): 473 – 488 . [DOI] [PubMed] [Google Scholar]
- 27. Kamarudin AN , Cox T , Kolamunnage-Dona R . Time-dependent ROC curve analysis in medical research: current methods and applications . BMC Med Res Methodol 2017. ; 17 ( 1 ): 53 – 19 . [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Pinsky PF , Gierada DS , Black W , et al . Performance of Lung-RADS in the National Lung Screening Trial: a retrospective assessment . Ann Intern Med 2015. ; 162 ( 7 ): 485 – 491 . [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Horeweg N , Scholten ET , de Jong PA , et al . Detection of lung cancer through low-dose CT screening (NELSON): a prespecified analysis of screening test performance and interval cancers . Lancet Oncol 2014. ; 15 ( 12 ): 1342 – 1350 . [DOI] [PubMed] [Google Scholar]
- 30. Tammemägi MC , Ten Haaf K , Toumazis I , et al . Development and validation of a multivariable lung cancer risk prediction model that includes low-dose computed tomography screening results: a secondary analysis of data from the National Lung Screening Trial . JAMA Netw Open 2019. ; 2 ( 3 ): e190204 . [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Robbins HA , Berg CD , Cheung LC , Chaturvedi AK , Katki HA . Identification of candidates for longer lung cancer screening intervals following a negative low-dose computed tomography result . J Natl Cancer Inst 2019. ; 111 ( 9 ): 996 – 999 . [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Landy R , Wang VL , Baldwin DR , et al . Recalibration of a deep learning model for low-dose computed tomographic images to inform lung cancer screening intervals . JAMA Netw Open 2023. ; 6 ( 3 ): e233273 . [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Subramanian J , Govindan R . Lung cancer in never smokers: a review . J Clin Oncol 2007. ; 25 ( 5 ): 561 – 570 . [DOI] [PubMed] [Google Scholar]
- 34. Remy-Jardin M , Remy J , Gosselin B , Becette V , Edme JL . Lung parenchymal changes secondary to cigarette smoking: pathologic-CT correlations . Radiology 1993. ; 186 ( 3 ): 643 – 651 . [DOI] [PubMed] [Google Scholar]
- 35. Terzikhan N , Verhamme KMC , Hofman A , Stricker BH , Brusselle GG , Lahousse L . Prevalence and incidence of COPD in smokers and non-smokers: the Rotterdam Study . Eur J Epidemiol 2016. ; 31 ( 8 ): 785 – 792 . [DOI] [PMC free article] [PubMed] [Google Scholar]




