Abstract
Background
Appropriate risk stratification of indeterminate pulmonary nodules (IPNs) is necessary to direct diagnostic evaluation. Currently available models were developed in populations with lower cancer prevalence than that seen in thoracic surgery and pulmonology clinics and usually do not allow for missing data. We updated and expanded the Thoracic Research Evaluation and Treatment (TREAT) model into a more generalized, robust approach for lung cancer prediction in patients referred for specialty evaluation.
Research Question
Can clinic-level differences in nodule evaluation be incorporated to improve lung cancer prediction accuracy in patients seeking immediate specialty evaluation compared with currently available models?
Study Design and Methods
Clinical and radiographic data on patients with IPNs from six sites (N = 1,401) were collected retrospectively and divided into groups by clinical setting: pulmonary nodule clinic (n = 374; cancer prevalence, 42%), outpatient thoracic surgery clinic (n = 553; cancer prevalence, 73%), or inpatient surgical resection (n = 474; cancer prevalence, 90%). A new prediction model was developed using a missing data-driven pattern submodel approach. Discrimination and calibration were estimated with cross-validation and were compared with the original TREAT, Mayo Clinic, Herder, and Brock models. Reclassification was assessed with bias-corrected clinical net reclassification index and reclassification plots.
Results
Two-thirds of patients had missing data; nodule growth and fluorodeoxyglucose-PET scan avidity were missing most frequently. The TREAT version 2.0 mean area under the receiver operating characteristic curve across missingness patterns was 0.85 compared with that of the original TREAT (0.80), Herder (0.73), Mayo Clinic (0.72), and Brock (0.68) models with improved calibration. The bias-corrected clinical net reclassification index was 0.23.
Interpretation
The TREAT 2.0 model is more accurate and better calibrated for predicting lung cancer in high-risk IPNs than the Mayo, Herder, or Brock models. Nodule calculators such as TREAT 2.0 that account for varied lung cancer prevalence and that consider missing data may provide more accurate risk stratification for patients seeking evaluation at specialty nodule evaluation clinics.
Key Words: lung cancer, lung nodule, prediction model
Take-home Points.
Study Question: Can clinic-level differences in nodule evaluation be incorporated to improve lung cancer prediction accuracy in patients undergoing immediate specialty evaluation compared with currently available models?
Results: Using a novel pattern submodel approach, the Thoracic Research Evaluation and Treatment (TREAT) 2.0 model (area under the receiver operating characteristic curve [AUC], 0.85) was able to distinguish between cancerous and benign nodules better than the TREAT 1.0 model (AUC, 0.80), Mayo Clinic model (AUC, 0.72), Herder model (AUC, 0.73), or Brock model (AUC, 0.68) in 1,401 indeterminate pulmonary nodules in a clinical population with high cancer prevalence referred for specialty evaluation.
Interpretation: The TREAT 2.0 model predicted the probability of lung cancer in a high cancer prevalence clinical setting with high accuracy and good calibration, better than currently existing models, and may improve time to diagnosis for lung cancer if implemented in pulmonology and thoracic surgery clinics.
Lung cancer is the leading cause of cancer mortality in the United States.1, 2, 3 Timely identification and resection of early-stage lung cancer is critical to save lives.4,5 With 1.6 million incidental indeterminate pulmonary nodules (IPNs) discovered annually and the expansion of lung cancer screening, IPNs represent a significant diagnostic and evaluative burden on the medical system.6, 7, 8, 9, 10 The CHEST guidelines recommend use of a validated prediction model or clinical expertise to estimate cancer probability in IPNs.11 Accurate identification of moderate-risk to high-risk IPNs can minimize invasive procedures in benign disease and can reduce time to diagnosis and treatment for lung cancer.
Currently available models for IPN risk prediction, including the Mayo Clinic model, Herder model, Brock model, and Veterans Affairs Lung Cancer model, were developed in populations with lower cancer prevalence than that seen in pulmonary nodule or thoracic surgery clinics (Table 1).12, 13, 14, 15, 16, 17, 18 Additionally, many of these models were developed in single-site cohorts of only a few hundred patients and were validated with even fewer patients, limiting generalizability.12, 13, 14,17,19,20 With the increasing burden of IPNs and high rates of invasive testing on benign disease, the need for an accurate model well calibrated to the higher-risk population seen in surgical and pulmonology clinics has never been greater.6,21, 22, 23, 24, 25
Table 1.
Current Lung Nodule Risk Assessment Models
| Model | No. of Variables | No. of Patients in Development Dataset | Cancer Prevalence Development, % |
|---|---|---|---|
| Brock (parsimonious) | 4 | 1,871 | 5.5 |
| Brock (full) | 9 | 1,871 | 5.5 |
| Mayo | 6 | 419 | 23 |
| Herder | 7 | 106 | 57.5 |
| Veterans Affairs | 4 | 375 | 54 |
| TREAT 1.0 | 12 | 492 | 72 |
TREAT = Thoracic Research Evaluation and Treatment.
Furthermore, patients evaluated in specialty nodule clinics frequently have an extensive workup already complete and thus have more predictive information available than what is used in the more parsimonious models. However, patients often are missing one or more variables required for more complex risk calculators, such as the full Brock model, compromising the predictive accuracy of these models. To address these concerns, we propose using a novel missing data method for predictive modeling, the pattern submodel (PS) approach.26
We sought to expand our previously developed Thoracic Research Evaluation and Treatment (TREAT) version 1.0 lung cancer prediction model into a robust yet flexible model for use by pulmonologists and thoracic surgeons: the TREAT version 2.0 model. We compared the performance of TREAT 2.0 with that of the Mayo, Herder, and Brock models for the diagnosis of lung cancer.12,13,17
Study Design and Methods
Study Population
This multicenter retrospective cohort study included 1,401 patients with IPNs from six cohorts across four US states (Table 2). Cohorts 1 and 2 included patients who sought treatment at pulmonary nodule clinics (lung cancer prevalence, 42%). Cohorts 3 and 4 included patients referred to thoracic surgery clinics (lung cancer prevalence, 73%). Cohorts 5 and 6 included those who underwent an inpatient lung resection for known or suspected lung cancer (lung cancer prevalence, 90%). The overall lung cancer prevalence in the study was 70%.
Table 2.
Clinical Setting and Cohort Description
| Clinical Setting | No. | Cohort | Cohort Location | No. | Lung Cancer Prevalence, % |
|---|---|---|---|---|---|
| Pulmonary nodule clinic | 374 | 1 | Tennessee | 258 | 34 |
| 2 | Arizona | 116 | 60 | ||
| Thoracic surgery clinic | 553 | 3 | Tennessee | 492 | 72 |
| 4 | Massachusetts | 61 | 82 | ||
| Inpatient surgical resection | 474 | 5 | Tennessee | 216 | 93 |
| 6 | Virginia | 258 | 88 |
Cohort 1 (n = 258) included patients evaluated at a pulmonary nodule clinic at Vanderbilt University Medical Center in Nashville, Tennessee. Cohort 2 (n = 116) included patients evaluated at a pulmonary nodule clinic at Mayo Clinic in Phoenix, Arizona. Cohort 3 (n = 492) included patients evaluated at a thoracic surgery clinic at Vanderbilt University Medical Center in Nashville, Tennessee, from 2005 through 2010; this group comprised the original development cohort for the TREAT model. Cohort 4 (n = 61) included patients from Lahey Hospital and Medical Center in Burlington, Massachusetts, with a screening-detected Lung CT Screening Reporting and Data System class 4 nodule or mass who underwent subsequent evaluation in a surgical or multidisciplinary thoracic oncology clinic. Although these nodules were detected by screening, this represents a high cancer prevalence subset of screening-detected nodules, and as such, is not representative of a general screening population. Cohort 5 (n = 216) comprised patients who underwent lung resection for known or suspected lung cancer at the Tennessee Valley Healthcare System Veterans Affairs hospital in Nashville, Tennessee, from 2005 through 2015. Cohort 6 (n = 258) included patients who underwent lung resection for known or suspected lung cancer at the University of Virginia in Charlottesville, Virginia, as part of the Lung Cancer Biospecimen Resource Network collaboration. Besides cohort 4, all cohorts included patients with symptomatic, incidentally discovered, or screening-detected IPNs.
Because this model is intended for use among IPNs, nodules for which a diagnosis was immediately apparent in either clinic setting were excluded. Therefore, patients with evidence of clearly benign disease (fibrotic or calcified nodules) or evident metastatic disease (extrathoracic or lung primary) or who already had undergone neoadjuvant therapy were excluded. Exclusion criteria for surgical patients (cohorts 3-6) also included multiple or other suspicious nodules. Among patients undergoing a second resection for a new primary or recurrence during the study period, only the first resection was included. If pulmonary clinic patients demonstrated multiple nodules, risk was estimated for the most suspicious nodule based on clinicians’ reports. The institutional review board of each facility approved this study with waivers of individual patient consent.
Procedures and Definitions
The original TREAT model was developed using a prespecified set of demographic, clinical, and radiographic variables.20 The same variables (age, BMI, sex, pack-years of tobacco use, size of nodule, spiculation, growth over time, upper lobe location, prior cancer history, preoperative predicted FEV1, preoperative symptoms, positive fluorodeoxyglucose-PET [FDG-PET] scan findings) were included with the addition of clinical setting of evaluation. The clinical setting accounts for clinical environment cancer prevalence variation and acts as a surrogate variable for components of the clinical history that may trigger referral to a specialty clinic. Study data were collected using a common, secure Research Electronic Data Capture form hosted at Vanderbilt University Medical Center.27
Imaging data were abstracted from radiologist reports or directly from the most recent preoperative CT scans for size, location, growth over at least 60 days, spiculation, and FDG-PET scan avidity (< 2.5 or ≥ 2.5 standardized uptake value) by experienced medical reviewers.18,21,28,29 Nodules were considered PET scan avid based on radiologist suspicion in the absence of standardized uptake value data or when radiologist suspicion conflicted with a standardized uptake value of < 2.5. Records with insufficient data to determine lesion growth (fewer than 60 days between scans or no sequential scans) were treated as missing. Preoperative symptoms were defined as any documented evidence in the medical record of hemoptysis, shortness of breath, unplanned weight loss, fatigue, pain, or persistent pneumonia. Preoperative FEV1 was based on the most recent pulmonary function test before evaluation. The outcome of malignant or benign diagnosis was determined by pathologic tissue diagnosis or by 2 years of radiographic surveillance among patients not undergoing biopsy.
Statistical Analysis
Analyses were performed using R version 3.6.3 software and Rstudio version 1.1.456 (R Foundation for Statistical Computing) using the following packages: glmnet, randomForest, and gbm. Variables were summarized and examined using appropriate descriptive statistics. Several candidate predictive models were developed using traditional logistic regression, logistic regression with least absolute shrinkage and selection operator regularization, random forest, and gradient boosting.
The predictive accuracy of each model, as measured by the area under the receiver operating characteristic curve (AUC), was cross-validated with repeated 10-fold validation with the proportion of malignancy in each fold held constant. The features set consisted of patient age at time of nodule evaluation, tobacco use history, history of cancer, lesion size, spiculation, lesion location, sex, BMI, indication of lesion growth when available, and FDG-PET scan results. Restricted cubic splines were used for continuous features to allow for possible nonlinear association with the outcome. To improve calibration, we included a fixed covariate for clinical setting: 1 = pulmonary clinic, 2 = thoracic surgery clinic, and 3 = surgical resection.
To maximize predictive accuracy, missing covariate data were handled using the PS approach.26 The PS approach subsets the data into groups with identical missing data patterns and fits the prediction model within each subset. For example, all patients missing only FDG-PET scan data form one subset, whereas patients missing FDG-PET scan and FEV1 data comprise another. The PS method maximizes predictive accuracy and allows the missing data pattern to be predictively informative itself. The overall resulting prediction model is a set of multiple prediction models, one for each missing data pattern. By fitting each submodel to the specific missing data pattern, the PS approach incorporates the fact that missingness likely is not random. The predictive information contained in the presence or absence of a variable, which incorporates other unseen predictive clinical information that may have led a clinician to obtain imaging or testing such as FDG-PET imaging, pulmonary function testing, and so forth. We used the PS approach within several techniques to develop a set of candidate prediction models (e-Figs 1, 2). For each missing data pattern, the same type of prediction model was used, but the fit of the model varied by pattern. Predictive performance was assessed within each missing data pattern using repeated 10-fold cross validation. For missing data patterns not observed in the data set, the complete case subset was substituted with variables removed to match the estimated missing pattern. For missing data patterns with fewer than 50 observations or fewer than 10 observations of either benign or malignant nodules, the complete subset (n = 471) supplemented the existing observations.
The discrimination and calibration of the TREAT 2.0 prediction model was estimated and compared with the Mayo, Herder, Brock, and original TREAT 1.0 models within each missing data pattern subset. We measured the discrimination of the model (ability to distinguish cancer from benign disease) using the AUC. We summarized the large-scale calibration of the model, measuring how closely the model’s predicted probabilities match the observed outcome, using the Brier score. The mean and interquartile ranges from 50 repetitions of 10-fold cross validation of AUC and Brier scores are reported. Overall calibration was assessed with estimated logistic and nonparametric calibration curves.
Risk Group Reclassification
To estimate the clinical impact of TREAT 2.0, we evaluated the reclassification of the 1,401 patients in the nodule cohort using TREAT 2.0 compared with the Mayo model. The cohort was divided into those with malignant and with benign nodules. Reclassification was represented visually using reclassification plots and quantitatively by calculating the bias-corrected clinical net reclassification index. Nodules were classified into low-risk, intermediate-risk, or high-risk groups using cutoff values of 10% and 70% per British Thoracic Society guidelines (< 10% denotes low risk, 10%-70% denotes intermediate risk, and > 70% denotes high risk).
Results
Missing data frequency and summary data for each TREAT variable are presented in Table 3. e-Tables 1 and 2 contain the full descriptive statistics summarized by cohort and clinical setting. In total, 64 patterns of missing data were observed, although most of these patterns included only one or two patients (e-Table 3). A missingness pattern could involve patients missing one variable or patients missing multiple variables (eg, BMI and FDG-PET scan avidity). Four patterns accounted for 76% of patients and were used to illustrate the performance of the different prediction models in Figure 1. The largest pattern (n = 471) represents complete cases, that is, patients with no missing data.
Table 3.
TREAT Variable Characteristics and Missing Data
| Variable | Population (N = 1,401) | Missing Data |
|---|---|---|
| Age, y | 64 ± 11.7 | 0 |
| Male sex | 813 (58) | 0 |
| BMI, kg/m2 | 28 ± 6.1 | 129 (9) |
| Tobacco use history, pack-y | 41 ± 34 | 42 (3) |
| Symptomatic | 627 (45) | 149 (11) |
| Previous cancer | 438 (31) | 3 (0.2) |
| Predicted FEV1 | 77 ± 20 | 277 (20) |
| Nodule size, mm | 25 ± 17 | 18 (1) |
| Spiculation | 578 (41) | 60 (4) |
| Growth | 490 (35) | 635 (45) |
| Upper lobe | 743 (53) | 23 (2) |
| PET avid | 904 (65) | 311 (22) |
Data are presented as No. (%) or mean ± SD. TREAT = Thoracic Research Evaluation and Treatment.
Figure 1.
Graph comparing internally validated model discriminative ability. Distribution of repeated cross-validated AUC values within each model for each missing pattern (50 repeats of 10-fold cross validation). The points indicate the mean AUC for a certain model and missing pattern, with the size of the point corresponding to the size of the missing pattern in the observed data. For the four largest missing patterns, the interquartile range for the AUC distribution also is given as the vertical interval lines. The horizontal dashed lines are the weighted AUC means across the top four missing patterns, and the horizontal dotted lines are the weighted AUC means across all missing patterns (weighting according to pattern size, which is the number of observed patients in the missing pattern). AUC = area under the receiver operating characteristic curve; IQR = interquartile range; TREAT = Thoracic Research Evaluation and Treatment.
Model Performance
Using the PS approach, the model with the best overall performance was the logistic regression model with fully relaxed least absolute shrinkage and selection operator regularization with constrained interactions and potential nonlinear associations. The criteria for selecting the predictive model were based on the mean of the repeated cross-validated AUC values across the missing patterns, weighted according to sample size of the missing pattern (see e-Fig 1 for comparison of models tested). The selected model type allows any of the main effect, interaction, or nonlinear terms to drop from the model. This approach results in a model that keeps only information that contributes meaningfully to the risk prediction.
To compare model performance, the weighted average of the cross-validated AUCs across all missing patterns was used. The TREAT 2.0 model (mean AUC, 0.85) (Table 4) was able to distinguish between cancerous and benign nodules better than the TREAT 1.0 model (mean AUC, 0.80), Mayo Clinic model (mean AUC, 0.72), Herder model (mean AUC, 0.73), or Brock model (mean AUC, 0.68). Figure 1 illustrates the range of cross-validated AUCs for each model by missing pattern for the four most common missing patterns (see e-Fig 3 and e-Table 4 for each observed missing pattern).
Table 4.
Summary of Model Performance Across Missingness Patterns
| Model | Mean AUC | Mean Brier Score |
|---|---|---|
| TREAT 2.0 | 0.85 | 0.11 |
| TREAT 1.0 | 0.80 | 0.13 |
| Mayo | 0.72 | 0.20 |
| Herder | 0.73 | 0.28 |
| Brock | 0.68 | 0.33 |
The mean is the weighted average across all observed missing patterns, using weights proportional to the size of the missing pattern. The lower Brier score demonstrates improved discrimination. AUC = area under the receiver operating characteristic curve; TREAT = Thoracic Research Evaluation and Treatment.
The TREAT 2.0 model calibration was compared with that of the other models by missing pattern. The mean Brier score across missing patterns indicated that TREAT 2.0 has improved calibration (Table 4). Figure 2 illustrates the calibration curves for each model in the largest missing pattern, where all variables are observed, providing a more insightful picture of the overall calibration than the Brier score. The predicted probabilities in a well-calibrated model should follow closely the actual probabilities, illustrated by the diagonal line in the calibration plots. Figure 2 illustrates that the Mayo, Herder, Brock, and TREAT 1.0 models underestimate malignancy risk. These results only describe the complete case missing pattern; each of the observed missing patterns have their own set of curves (e-Figs 4, 5, e-Table 5). In most missing patterns, the TREAT 2.0 model was well calibrated, the Mayo and Herder models were generally not well calibrated, and the TREAT 1.0 model demonstrated mixed calibration performance. The calibration curve for TREAT 2.0 may be overly optimistic, given that it is assessed in the same sample used to fit the model. The cross-validated Brier scores provide additional support for the improved calibration of the TREAT 2.0 model.
Figure 2.
Graphs showing model calibration for the complete case missing pattern. The diagonal solid grey line shows the hypothetical ideal calibration, where actual probabilities match exactly the predicted probabilities. The solid curve is the calibration estimated via a logistic model as described in Harrell (2001).30 The dotted curve is the calibration estimated via nonparametric locally estimated scatterplot smoothing. Along the x-axis is a histogram of the predicted probabilities. Areas where the line is above the diagonal indicate that the model underestimates cancer risk, and areas where the line is below the diagonal indicate that the model overestimates cancer risk. TREAT = Thoracic Research Evaluation and Treatment.
Risk Group Reclassification
Nodule reclassification is summarized in Figure 3. Figure 3A demonstrates classification of benign nodules, as determined by the Mayo model vs the TREAT 2.0 model. The blue sections represent an appropriate downgrade in nodule risk category, where a benign nodule has moved from high-risk or intermediate-risk groups per the Mayo model to intermediate-risk or low-risk groups per the TREAT 2.0 model. The red sections of the graph represent benign nodules that have been upgraded incorrectly in risk stratification, nodules initially classified as low or intermediate risk by the Mayo model that were reclassified incorrectly as intermediate or high risk by the TREAT 2.0 model. The model correctly moved 32% (97/304) of benign nodules incorrectly classified as high or intermediate risk by the Mayo model into lower-risk groups.
Figure 3.
A, B, Risk group classification plots illustrated for benign nodules (A) and malignant nodules (B) for the TREAT 2.0 model compared with the Mayo model. A, Blue boxes demonstrate an appropriate downgrade in risk group for benign nodules by the TREAT 2.0 model compared with the Mayo model, whereas red boxes demonstrate an incorrect upgrade in risk group. B, Blue boxes represent an appropriate upgrade in risk group for malignant nodules by the TREAT 2.0 model compared with the Mayo model, whereas red boxes represent an incorrect downgrade in risk group. TREAT = Thoracic Research Evaluation and Treatment.
Figure 3B demonstrates the reclassification for malignant nodules. The blue sections represent an appropriate increase in risk group and the red sections represent an incorrect downgrade in risk group as assessed by the TREAT 2.0 model compared with the Mayo model. No malignant nodules initially classified as high risk were reclassified incorrectly as low risk using the TREAT 2.0 model. The TREAT 2.0 model correctly moved 79% (359/452) of incorrectly classified malignant nodules into higher-risk groups. The bias-corrected clinical net reclassification index was 0.23, denoting significant classification improvement.
Discussion
Clinicians face a deluge of IPNs from symptom-driven imaging, incidental findings, and, increasingly, lung cancer screening. Lung biopsy carries significant risk, with a 15% prevalence of pneumothorax requiring treatment and 1% to 2% risk of more severe complications such as hemorrhage, mechanical ventilation, or death.31 Patients proceeding to surgery without a definitive diagnosis are placed at even higher risk, with mortality rates of 1% to 3% at 30 days and up to 7% at 90 days after resection for a procedure with about a 20% probability of being nontherapeutic.6,32, 33, 34 However, the risk of morbidity and mortality often required to obtain a diagnosis must be balanced with the risk of delaying a cancer diagnosis with a surveillance strategy and potentially resulting in disease progression.
We previously developed a novel, validated clinical prediction model to diagnose primary lung cancer in a high-risk population of patients with IPNs undergoing surgical evaluation for lung biopsy or resection, with an AUC of 0.87.20 Our objective with TREAT 1.0 was to integrate the full spectrum of clinical and radiographic evidence available to surgeons at the juncture of evaluating a patient with a high-risk IPN without a definitive tissue diagnosis.
The present study expands on our original TREAT model by refitting the predictive model parameters across a range of clinical populations with high cancer prevalence and varying amounts of diagnostic information, represented in our model as missing data patterns. This approach provides greater flexibility in real-world clinical settings and leverages information contained in evaluation patterns represented by the patterns of available data themselves. For example, most screening-discovered nodules have multiple CT scans for determining growth. Far fewer incidental nodules, especially those with high-risk characteristics, had growth data, because the clinician may consider waiting to observe growth to be risk prohibitive. Thus, the pattern of missing data itself provides meaningful information about the nodule’s risk. Including the clinical evaluation setting and leveraging the information contained in missing data results in improved real-world applicability and generalizability across clinical practice settings while improving predictive accuracy.
Although many existing models demonstrated impressive discrimination with a high AUC in their respective development data sets, their performance declines in external validation data sets, particularly those representing clinical populations with a cancer prevalence that differs from the one in which they were developed, corroborated by a relatively poor performance in our high-prevalence dataset.18,35,36 The Brock model was developed in and initially intended for a screening population, and thus is calibrated to a low cancer prevalence setting and does not include tobacco exposure as a model predictor, which may limit generalizability to a nonscreening population. One of the strengths of the TREAT 2.0 model is that it was developed in a geographically diverse, multi-institutional dataset and was internally validated via a robust repeated cross-validation approach with retained accuracy and discrimination in high cancer prevalence referral settings. Most importantly, our results demonstrate improvement in risk group reclassification compared with the Mayo model, indicating that widespread use of TREAT 2.0 could translate into meaningful clinical benefits for patients in these high-prevalence populations.
Although this study fills a unique gap in the literature as the first accurate, well-calibrated model in patients undergoing specialist IPN evaluation, limitations exist that must be acknowledged. First, this study relied on retrospective data collection, although we standardized across sites using a common Research Electronic Data Capture data entry platform and, where possible, implemented direct review of clinical data by the same group of trained chart abstractors. Additionally, although the clinical settings represented a range of cancer prevalence, all cohorts demonstrated higher cancer prevalence than a typical screening or general pulmonary population. Specifically, cohort 4 represents a high-risk Lung CT Screening Reporting and Data System class 4 subgroup and is not meant to represent a general screening cohort. Our purpose was to design a model calibrated to the population with high cancer prevalence evaluated in specialty referral clinics. This model has not been studied in populations with lower cancer prevalence, such as a general lung cancer screening population, and as such, may be calibrated poorly to this population. The inclusion of other factors currently under investigation such as blood-based cancer and fungal biomarkers and imaging-based radiomic biomarkers may refine lung cancer diagnosis further.37, 38, 39, 40 Because race was not a predictive variable in the original TREAT model, it was not included in TREAT 2.0, and therefore was not collected in some cohorts. As such, it is unknown whether the performance of the model is comparable across sociodemographic subpopulations, including those defined by race. Finally, the TREAT 2.0 model, like nearly all prediction models, cannot easily be calculated manually. To allow for widespread use by clinicians, a web-based calculator is under development.
Interpretation
Multiple lung cancer risk prediction models currently are available for IPNs, but do not perform as well in dedicated lung nodule or thoracic surgery clinics because these models were developed and validated in populations with lower cancer prevalence. The TREAT 2.0 model predicts lung cancer in IPNs with higher accuracy and better discrimination (AUC, 0.85) than existing models in lung nodule and surgery clinics. It is designed to assess the risk of lung cancer in patients undergoing invasive evaluation, functions in the presence of missing variables by using a novel pattern submodel approach, and improves reclassification of both benign and malignant nodules. Importantly, because the model demonstrates improved risk classification, it may reduce unnecessary invasive procedures in benign disease and may improve the time to diagnosis for cancer if used in clinical practice.
Funding/Support
M. E. S. was supported by the Agency for Healthcare Research [Grant T32 HS026122]. A. W. M. was supported by the Office of Academic Affiliations, VA National Quality Scholars Program. C. M. G. was supported by the Surgical Oncology Training Grant [Grant T32CA106183-18]. E. L. G. previously was a recipient of the Department of VA Health Services Research and Development Service Career Development Award [Grant 10-024]. E. L. G. and S. A. D. are supported by the National Cancer Institute [Grant U01CA152662]. M. C. A. is supported by the National Institutes of Health [Grant R01CA251758]. F. M. is supported by the National Institutes of Health [Grant R01CA253923-01]. Research Electronic Data Capture is supported by the National Center for Advancing Translational Sciences, National Institutes of Health [Grant UL1 TR000445].
Financial/Nonfinancial Disclosures
The authors have reported to CHEST the following: J. M. I. discloses grants from Guardant Health and GRAIL, prior support for meeting attendance from Intuitive Surgical, planned or issued patents with AstraZeneca and Roche Genentech, and stock or stock options from LumaCyte, LLC. L. T. V. has received consulting fees from Ambu A/S. F. M. receives consulting fees from Medtronic, Johnson & Johnson, and Intuitive and additionally received research funding from Medtronic. The disclosures listed did not have any relation to the content of this manuscript. None declared (C. M. G., M. E. S., V. F. W., A. W. M., M. C. A., C. M., J. C., S. R., O. B. R., R. P., E. S. L., J. C. N., J. D. B., S. A. D., E. L. G.).
Acknowledgments
Author contributions: C. M. G. contributed to formal analysis, writing the original draft, reviewing and editing subsequent drafts, and visualization. M. E. S. contributed to conceptualization, methodology, data analysis, investigation, writing the original draft, and reviewing and editing subsequent drafts. V. F. W. contributed to conceptualization, methodology, investigation, formal analysis, validation, writing the original draft, reviewing and editing subsequent drafts, and visualization. A. W. M. contributed to conceptualization, data analysis, and review and editing subsequent drafts. M. C. A. contributed to reviewing and editing subsequent drafts. C. M. contributed to investigation, data analysis, and reviewing and editing subsequent drafts. J. C. contributed to investigation, data analysis, and reviewing and editing subsequent drafts. L. T. V. contributed to investigation, supervision, writing the original draft, and reviewing and editing subsequent drafts. S. R. contributed to investigation, supervision, writing the original draft, and reviewing and editing subsequent drafts. J. M. I. contributed to investigation, supervision, writing the original draft, and reviewing and editing subsequent drafts. O. B. R. contributed to investigation, supervision, writing the original draft, and reviewing and editing subsequent drafts. R. P. contributed to investigation, data analysis, and reviewing and editing subsequent drafts. E. S. L. contributed to writing the original draft and reviewing and editing subsequent drafts. J. C. N. contributed to writing the original draft and reviewing and editing subsequent drafts. F. M. contributed to reviewing and editing subsequent drafts. J. D. B. contributed to methodology, formal analysis, validation, reviewing and editing subsequent drafts, and supervision. S. A. D. contributed to conceptualization, methodology, and reviewing and editing subsequent drafts. E. L. G. contributed to conceptualization, methodology, writing the original draft, reviewing and editing subsequent drafts, and supervision and is the guarantor for the study.
Role of sponsors: The funding agencies had no role in the design and conduct of the study; collection, management, analysis, or interpretation of the data; preparation, review, or approval of the manuscript; or in the decision to submit the manuscript for publication.
Disclaimer: The views expressed in this article are those of the authors and do not necessarily represent the views of the VA.
Other contributions: The authors thank Melissa Potter, BA, for assistance with editing the manuscript.
Additional information: The e-Figures and e-Tables are available online under “Supplementary Data.”
Footnotes
Drs Godfrey and Shipe contributed equally to this manuscript.
Supplementary Data
References
- 1.Siegel R., Ward E., Brawley O., Jemal A. Cancer statistics, 2011: the impact of eliminating socioeconomic and racial disparities on premature cancer deaths. CA Cancer J Clin. 2011;61(4):212–236. doi: 10.3322/caac.20121. [DOI] [PubMed] [Google Scholar]
- 2.Jemal A., Siegel R., Xu J., Ward E. Cancer statistics, 2010. CA Cancer J Clin. 2010;60(5):277–300. doi: 10.3322/caac.20073. [DOI] [PubMed] [Google Scholar]
- 3.Singh S.D., Henley S.J., Ryerson A.B. Surveillance for Cancer Incidence and Mortality - United States, 2013. MMWR Surveill Summ. 2017;66(4):1–36. doi: 10.15585/mmwr.ss6604a1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Ost D.E., Jim Yeung S.C., Tanoue L.T., Gould M.K. Clinical and organizational factors in the initial evaluation of patients with lung cancer: Diagnosis and Management of Lung Cancer, 3rd ed: American College of Chest Physicians evidence-based clinical practice guidelines. Chest. 2013;143(5 suppl):e121S–e141S. doi: 10.1378/chest.12-2352. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Detterbeck F.C., Lewis S.Z., Diekemper R., Addrizzo-Harris D., Alberts W.M. Executive summary: Diagnosis and Management of Lung Cancer, 3rd ed: American College of Chest Physicians evidence-based clinical practice guidelines. Chest. 2013;143(5 suppl):7S–37S. doi: 10.1378/chest.12-2377. [DOI] [PubMed] [Google Scholar]
- 6.National Lung Screening Trial Research Team. Church T.R., Black W.C., et al. Results of initial low-dose computed tomographic screening for lung cancer. N Engl J Med. 2013;368(21):1980–1991. doi: 10.1056/NEJMoa1209120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Jamal A., King B.A., Neff L.J., Whitmill J., Babb S.D., Graffunder C.M. Current cigarette smoking among adults—United States, 2005-2015. MMWR Morb Mortal Wkly Rep. 2016;65(44):1205–1211. doi: 10.15585/mmwr.mm6544a2. [DOI] [PubMed] [Google Scholar]
- 8.Bach P.B., Kattan M.W., Thornquist M.D., et al. Variations in lung cancer risk among smokers. J Natl Cancer Inst. 2003;95(6):470–478. doi: 10.1093/jnci/95.6.470. [DOI] [PubMed] [Google Scholar]
- 9.Gould M.K., Tang T., Liu I.L., et al. Recent trends in the identification of incidental pulmonary nodules. Am J Respir Crit Care Med. 2015;192(10):1208–1214. doi: 10.1164/rccm.201505-0990OC. [DOI] [PubMed] [Google Scholar]
- 10.Dyer O. US task force recommends extending lung cancer screenings to over 50s. BMJ. 2021;372:n698. doi: 10.1136/bmj.n698. [DOI] [PubMed] [Google Scholar]
- 11.Gould M.K., Donington J., Lynch W.R., et al. Evaluation of individuals with pulmonary nodules: when is it lung cancer? Diagnosis and management of lung cancer, 3rd ed: American College of Chest Physicians evidence-based clinical practice guidelines. Chest. 2013;143(5 suppl):e93S–e120S. doi: 10.1378/chest.12-2351. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.McWilliams A., Tammemagi M.C., Mayo J.R., et al. Probability of cancer in pulmonary nodules detected on first screening CT. N Engl J Med. 2013;369(10):910–919. doi: 10.1056/NEJMoa1214726. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Swensen S.J., Silverstein M.D., Ilstrup D.M., Schleck C.D., Edell E.S. The probability of malignancy in solitary pulmonary nodules. Application to small radiologically indeterminate nodules. Arch Intern Med. 1997;157(8):849–855. [PubMed] [Google Scholar]
- 14.Schultz E.M., Sanders G.D., Trotter P.R., et al. Validation of two models to estimate the probability of malignancy in patients with solitary pulmonary nodules. Thorax. 2008;63(4):335–341. doi: 10.1136/thx.2007.084731. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Gurney J.W., Lyddon D.M., McKay J.A. Determining the likelihood of malignancy in solitary pulmonary nodules with Bayesian analysis. Part II. Application. Radiology. 1993;186(2):415–422. doi: 10.1148/radiology.186.2.8421744. [DOI] [PubMed] [Google Scholar]
- 16.Callister M.E., Baldwin D.R., Akram A.R., et al. British Thoracic Society guidelines for the investigation and management of pulmonary nodules. Thorax. 2015;70(suppl 2):ii1–ii54. doi: 10.1136/thoraxjnl-2015-207168. [DOI] [PubMed] [Google Scholar]
- 17.Herder G.J., van Tinteren H., Golding R.P., et al. Clinical prediction model to characterize pulmonary nodules: validation and added value of 18F-fluorodeoxyglucose positron emission tomography. Chest. 2005;128(4):2490–2496. doi: 10.1378/chest.128.4.2490. [DOI] [PubMed] [Google Scholar]
- 18.Isbell J.M., Deppen S., Putnam J.B., Jr., et al. Existing general population models inaccurately predict lung cancer risk in patients referred for surgical evaluation. Ann Thorac Surg. 2011;91(1):227–233. doi: 10.1016/j.athoracsur.2010.08.054. discussion 233. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Gould M.K., Ananth L., Barnett P.G. A clinical model to estimate the pretest probability of lung cancer in patients with solitary pulmonary nodules. Chest. 2007;131(2):383–388. doi: 10.1378/chest.06-1261. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Deppen S.A., Blume J.D., Aldrich M.C., et al. Predicting lung cancer prior to surgical resection in patients with lung nodules. J Thorac Oncol. 2014;9(10):1477–1484. doi: 10.1097/JTO.0000000000000287. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Deppen S.A., Blume J.D., Kensinger C.D., et al. Accuracy of FDG-PET to diagnose lung cancer in areas with infectious lung disease: a meta-analysis. JAMA. 2014;312(12):1227–1236. doi: 10.1001/jama.2014.11488. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Kozower B.D., Meyers B.F., Reed C.E., Jones D.R., Decker P.A., Putnam J.B., Jr. Does positron emission tomography prevent nontherapeutic pulmonary resections for clinical stage IA lung cancer? Ann Thorac Surg. 2008;85(4):1166–1169. doi: 10.1016/j.athoracsur.2008.01.018. discussion 1169-1170. [DOI] [PubMed] [Google Scholar]
- 23.Veronesi G., Bellomi M., Scanagatta P., et al. Difficulties encountered managing nodules detected during a computed tomography lung cancer screening program. J Thorac Cardiovasc Surg. 2008;136(3):611–617. doi: 10.1016/j.jtcvs.2008.02.082. [DOI] [PubMed] [Google Scholar]
- 24.Kuo E., Bharat A., Bontumasi N., et al. Impact of video-assisted thoracoscopic surgery on benign resections for solitary pulmonary nodules. Ann Thorac Surg. 2012;93(1):266–272. doi: 10.1016/j.athoracsur.2011.08.035. discussion 272-273. [DOI] [PubMed] [Google Scholar]
- 25.Starnes S.L., Reed M.F., Meyer C.A., et al. Can lung cancer screening by computed tomography be effective in areas with endemic histoplasmosis? J Thorac Cardiovasc Surg. 2011;141(3):688–693. doi: 10.1016/j.jtcvs.2010.08.045. [DOI] [PubMed] [Google Scholar]
- 26.Fletcher Mercaldo S., Blume J.D. Missing data and prediction: the pattern submodel. Biostatistics. 2020;21(2):236–252. doi: 10.1093/biostatistics/kxy040. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Harris P.A., Taylor R., Thielke R., Payne J., Gonzalez N., Conde J.G. Research electronic data capture (REDCap)—a metadata-driven methodology and workflow process for providing translational research informatics support. J Biomed Inform. 2009;42(2):377–381. doi: 10.1016/j.jbi.2008.08.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Grogan E.L., Weinstein J.J., Deppen S.A., et al. Thoracic operations for pulmonary nodules are frequently not futile in patients with benign disease. J Thorac Oncol. 2011;6(10):1720–1725. doi: 10.1097/JTO.0b013e318226b48a. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Grogan E.L., Deppen S.A., Ballman K.V., et al. Accuracy of fluorodeoxyglucose-positron emission tomography within the clinical practice of the American College of Surgeons Oncology Group Z4031 trial to diagnose clinical stage I non-small cell lung cancer. Ann Thorac Surg. 2014;97(4):1142–1148. doi: 10.1016/j.athoracsur.2013.12.043. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Harrell F.E. Regression Modeling Strategies. Springer; 2001. [Google Scholar]
- 31.Wiener R.S., Schwartz L.M., Woloshin S., Welch H.G. Population-based risk for complications after transthoracic needle lung biopsy of a pulmonary nodule: an analysis of discharge records. Ann Intern Med. 2011;155(3):137–144. doi: 10.1059/0003-4819-155-3-201108020-00003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Cheung M.C., Hamilton K., Sherman R., et al. Impact of teaching facility status and high-volume centers on outcomes for lung cancer resection: an examination of 13,469 surgical patients. Ann Surg Oncol. 2009;16(1):3–13. doi: 10.1245/s10434-008-0025-9. [DOI] [PubMed] [Google Scholar]
- 33.von Meyenfeldt E.M., Gooiker G.A., van Gijn W., et al. The relationship between volume or surgeon specialty and outcome in the surgical treatment of lung cancer: a systematic review and meta-analysis. J Thorac Oncol. 2012;7(7):1170–1178. doi: 10.1097/JTO.0b013e318257cc45. [DOI] [PubMed] [Google Scholar]
- 34.Powell H.A., Tata L.J., Baldwin D.R., Stanley R.A., Khakwani A., Hubbard R.B. Early mortality after surgical resection for lung cancer: an analysis of the English National Lung cancer audit. Thorax. 2013;68(9):826–834. doi: 10.1136/thoraxjnl-2012-203123. [DOI] [PubMed] [Google Scholar]
- 35.Susam S., Çinkooğlu A., Ceylan K.C., et al. Comparison of Brock University, Mayo Clinic and Herder models for pretest probability of cancer in solid pulmonary nodules. Clin Respir J. 2022;16(11):740–749. doi: 10.1111/crj.13546. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Vachani A., Zheng C., Amy Liu I.L., Huang B.Z., Osuji T.A., Gould M.K. The probability of lung cancer in patients with incidentally detected pulmonary nodules: clinical characteristics and accuracy of prediction models. Chest. 2022;161(2):562–571. doi: 10.1016/j.chest.2021.07.2168. [DOI] [PubMed] [Google Scholar]
- 37.Mathios D., Johansen J.S., Cristiano S., et al. Detection and characterization of lung cancer using cell-free DNA fragmentomes. Nat Commun. 2021;12(1):5060. doi: 10.1038/s41467-021-24994-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Kammer M.N., Lakhani D.A., Balar A.B., et al. Integrated biomarkers for the management of indeterminate pulmonary nodules. Am J Respir Crit Care Med. 2021;204(11):1306–1316. doi: 10.1164/rccm.202012-4438OC. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Kammer M.N., Massion P.P. Noninvasive biomarkers for lung cancer diagnosis, where do we stand? J Thorac Dis. 2020;12(6):3317–3330. doi: 10.21037/jtd-2019-ndt-10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Marmor H.N., Jackson L., Gawel S., et al. Improving malignancy risk prediction of indeterminate pulmonary nodules with imaging features and biomarkers. Clin Chim Acta. 2022;534:106–114. doi: 10.1016/j.cca.2022.07.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.



