Abstract
Background
Accurate identification of patients at high risk of pulmonary infection after thoracoscopic lung cancer resection is important for timely and targeted preventive measures. Methods for determining the risk of pulmonary infection after thoracoscopic lung cancer resection have not been well studied.
Methods
This study was a retrospective case–control research project. The information of 3219 hospitalised patients who underwent thoracoscopic lung cancer resection between January 2019 and December 2023 was obtained from the hospital electronic medical record system. 26 clinical characteristics were obtained from medical and nursing records. The variables were screened using the least absolute contraction and selection operator (LASSO) regression, and the risk prediction models for pulmonary infection after thoracoscopic lung cancer resection was constructed using the following 5 machine learning algorithms: logistic regression model (LR), artificial neural network (ANN), support vector machine (SVM), random forest (RF) and eXtreme gradient boosting (XGB). The model was evaluated using the following metrics: the area under the receiver operator characteristic curve (AUC), accuracy, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV) and F1 score (F1). Shapley additive explanation (SHAP) was used to interpret the machine learning models.
Results
There were 3219 enrolled patients, 2203 (70%) of whom were assigned to the training cohort and 966 (30%) of whom were assigned to the validation cohort. The AUC range of the five models was 0.883–0.951. The XGB model outperformed the others, with an AUC of 0.951 (95% confidence interval: 0.943–0.964), accuracy of 0.902 (95% confidence interval: 0.886–0.913), sensitivity of 0.927, specificity of 0.864, positive predictive value of 0.898, negative predictive value of 0.824, precision of 0.908, recall of 0.872 and F1 score of 0.815 in the validation group. The model’s prediction performance in the 45–65 age group was the best. The AUC of the logistic regression model was 0.948 (95% confidence interval: 0.931–0.957). We transformed the logistic regression model into a nomogram to help clinicians visualise the model and make them more likely to use it to identify the risk of pulmonary infection after thoracoscopic surgery in lung cancer patients.
Conclusions
The establishment of a risk prediction model based on machine learning can help clinical nursing staff identify high-risk patients for pulmonary infection after thoracoscopic lung cancer resection.
Clinical trial number
Not applicable.
Supplementary Information
The online version contains supplementary material available at https://doi.org/10.1186/s12911-026-03698-5.
Keywords: Lung cancer, Video-assisted thoracoscopic surgery, Pulmonary infection, Prediction model, Machine learning
Background
According to the International Agency for Research on Cancer of the World Health Organization [1], there were 2.5 million new cases of lung cancer and 1.8 million deaths worldwide in 2022. Lung cancer is the most common malignant tumour in the world, with the highest standardised mortality rate. The global economic cost of cancer from 2020 to 2050 was estimated and predicted by Chen [2]; namely, lung cancer will have the highest medical burden among all malignant tumours in the next 30 years, resulting in a global economic loss of 3.9 trillion US dollars. Early surgical intervention is considered to be the best treatment choice for lung cancer patients. With the development of medical technologies, thoracoscopy and other endoscopic techniques have become widely used in clinical practice. However, it has been found that intraoperative operation causes severe damage to patients’ cardiopulmonary function and that 33.2% of the patients had postoperative pulmonary complications [3]. Pulmonary infection is the primary complication in patients after thoracoscopic lung cancer resection. It not only leads to a prolonged hospital stay and increased medical costs but also leads to more serious complications if its treatment is not timely, which increases the re-hospitalisation rate and long-term mortality rate of patients [4]. It is a preventable postoperative complication [5]; therefore, early identification of high-risk factors for postoperative pulmonary infection and timely intervention are required to reduce the associated burden of downstream complications and costs, improve the utilisation of medical resources and enhance the overall prognosis of lung cancer patients after surgery.
The traditional scoring model relies on the scoring of parameters through logistic regression analysis, which has the advantage of strong model interpretation, but the technique fails to handle the interactions between complex variables and is prone to underfitting [6]. When dealing with complex data and nonlinear problems, other more complex models need to be considered [7]. Artificial intelligence has emerged as a valuable statistical analysis paradigm, and the medical field has made notable progress from its applications. Machine learning is an important branch of artificial intelligence that enables researchers to set specific rules and process the input information according to these rules to generate the output. Machines can process cumbersome clinical data, independently learn and make accurate predictions [8]. Machine learning algorithms have been used to process and mine medical data for different purposes, such as data collection, intelligent analysis and risk identification, to accomplish intelligent patient management [9]. They are also often used in patient morbidity risk assessment [10, 11] to assist clinical medical personnel in the diagnosis and detection of disease complications. Early detection and treatment can reduce the occurrence of disease complications, curtail the economic burden of patients and improve their prognosis.
Machine learning has been applied to the risk management of postoperative pulmonary infection in surgical patients, including the risk prediction of spinal cord injury–related postoperative pulmonary infection, patients with primary liver cancer and patients with kidney transplantation. The research results have shown that with the help of machine learning algorithms, nurses can take more accurate preventive measures to deal with the postoperative complications of their patients, thus reducing healthcare costs [12–14]. In summary, machine learning has become a common patient risk management aid in clinical care. Some studies have established risk prediction models for pulmonary infection in advanced cancer patients and lung cancer chemotherapy patients [15, 16], but it is not clear whether these models are applicable to postoperative patients who have undergone thoracoscopic lung cancer resection. To the best of our knowledge, this is the first attempt to build a predictive model of pulmonary infection in patients after thoracoscopic lung cancer resection based on machine learning algorithms.
This study developed and validated five machine learning–based models to predict pulmonary infections in patients following thoracoscopic lung cancer surgery. The potential benefits include increased healthcare professionals’ productivity, improved resource utilisation and better patient outcomes. The predictive power of our models was compared for patients of different ages. To address the difficulty to visualise its mechanism and thereby meet the clinical needs of some medical units that function without intelligent medical electronic systems, we transformed the prediction model into a nomogram to rapidly estimate the possibility of pulmonary infection in patients after thoracoscopic lung cancer resection by manual calculation.
Methods
Study setting and participants
A retrospective case–control study was conducted in a grade three hospital in Shanghai, China. This study was approved by Medical Ethics Committee of Shanghai Chest Hospital Affiliated to Shanghai Jiaotong University School of Medicine (IRB Approval No. KS23016). The study included patients who underwent thoracoscopic lung cancer resection between January 2019 and December 2023. Patients meeting the following criteria were included in this study: (1) patients ≥ 18 years of age; (2) patients who met the diagnostic criteria for primary lung cancer [17]; (3) patients with no distant lung cancer metastasis; (4) patients diagnosed for the first time in the clinic who had not undergone radiotherapy, chemotherapy or targeted drug therapy; and (5) Patients who had undergone thoracoscopic resection of lung cancer. The exclusion criteria were (1) patients with infectious disease before surgery; (2) patients treated with immunosuppressants or glucocorticoids 1 month before surgery; (3) patients with severe heart, liver or kidney failure; and (4) patients with a clear diagnosis of mental illness or lack of normal cognitive or communication skills. This study was a retrospective case–control study and was approved by the Institutional Medical Ethics Committee. Therefore, informed consent from patients was not required. All patients’ personal information was encrypted to prevent disclosure. The stratified sampling randomization and the computer random seed number was used to randomly divide the participants into the training cohort and the validation cohort at a ratio of 7:3.
Diagnosis of pulmonary infection
Postoperative pulmonary infection refers to the signs of pulmonary infection in patients with lung cancer during the period from the end of surgery to the time before discharge. In this study, according to the Diagnostic Criteria for Nosocomial Infection issued by the Ministry of Health of China and the Guidelines for the Diagnosis and Treatment of Chinese Adult Hospital-Acquired Pneumonia and Ventilators Associated Pneumonia [18], patients can be diagnosed with any three of the following five criteria: (1) cough, purulent sputum or aggravation of original respiratory symptoms; (2) body temperature ≥ 38 °C, with blood routine indicating that the white blood cell count was increased significantly; (3) physical examination indicating dry and wet rales or signs of lung consolidation; (4) chest imaging examination indicating inflammatory changes or new or progressive infiltrating shadow, patchy infiltrating shadow or interstitial changes; and (5) sputum culture testing positive for pathogenic bacteria.
Model input features
The following search terms were used to screen the candidate predictors: (“lung cancer” OR “lung malignancy” OR “lung tumour” OR “lung mass”) AND (thoracoscope OR “endoscopic surgery” OR “minimally invasive surgery” OR “Surgical operation”) AND: (pneumonia OR “pulmonary infection” OR “pulmonary inflammation” OR “nosocomial infection”) AND: (“risk factor” OR “influencing factor” OR “affecting factor” OR predictor). We searched five electronic databases, namely PubMed, Cochrane Library, Medline, EMBASE and Ovid. Then, we assembled the candidate predictors based on the frequency and importance of the predictors in the references as well as the operability and convenience of obtaining the predictors in clinical work.
Based on the literature review and clinical expertise, we developed a survey form to collect data on the predictors of pulmonary infection. The form contained 26 candidate pulmonary infection predictors after thoracoscopic lung cancer resection were selected: (1) Personal information: gender, age, body mass index (BMI), daily smoking, hypertension and diabetes; (2) Preoperative pulmonary function: forced expiratory volume in one second/forced vital capacity (FEV1/FVC), diffusing capacity of the lung for carbon monoxide percentage of predicted (DLCO%Pred), airway resistance percentage of predicted (Raw%Pred) and forced vital capacity (FVC); (3) Surgical information: surgical level, mode of operation, surgical site, intraoperative lymph node dissection, number of intraoperative indwelling drainage tubes, intraoperative pathological type, intraoperative hilar activity, intraoperative pleural adhesion degree, degree of tumour invasion and maximum tumour diameter; and (4) Perioperative nursing information: 24 h postoperative activities of daily living, 24 h postoperative critical Modified Early Warning score (MEWs), properties of thoracic fluid 24 h after surgery, thoracic fluid volume 24 h after surgery, perioperative oral nutritional supplements and postoperative urinary catheter. The details of the survey form can be found in eTable 1. Finally, the Personnel trained in standardised data collection were responsible for obtaining the relevant information from the hospital’s electronic medical record system and recording it in pre-designed forms. We excluded variables with a proportion of missing data > 5%. During data entry, our database prevents final validation of the case if the mandatory data is missing. K-nearest neighbor was used as the method to impute the remaining missing value. K similar samples are found and filled according to the value of neighbors.
Model development and validation
Considering that least absolute shrinkage and selection operator (LASSO) regression can select the most important features, and by selecting the features corresponding to the non-zero coefficients, the features with the greatest predictive ability for the target variables can be screened out to simplify the model and improve the generalisation ability of the model, we used LASSO for variable screening in this study. Five machine learning algorithms were used to construct prediction models for pulmonary infection after thoracoscopic lung cancer resection, namely logistic regression model (LR), artificial neural network (ANN), support vector machine (SVM), random forest (RF) and eXtreme gradient boosting (XGB). We use the grid search method to determine the final hyperparameter values of each model. The area under curve (AUC), accuracy, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV) and F1 score (F1) were calculated to assess the model’s performance. The model was calibrated by the Brier score and calibration curve.
Statistical analysis
Counts (proportions) and medians (interquartile ranges, i.e., IQRs) were used to represent the categorical variables. The chi-square test was used to compare the differences between the two groups, and the T-test or Mann–Whitney U test was used for testing differences between the continuous and categorical variables, respectively. Based on the output of the LR model, the nomogram was developed. Double-sided P < 0.05 was considered statistically significant. We used Shapley additive explanation (SHAP) method for interpreting machine learning models and providing insights into how individual variables influence predictions. The LR model was corrected by adjustment of the intercept and the regression cocfficients using the calibration intercept and calibration slope, the LR model was converted into a nomogram for the visualization in clinical applications. The analyses were implemented in R version 4.1.
Results
Patient characteristics
The recruitment process is shown in Fig. 1. A total of 3749 patients had thoracoscopic lung cancer resections, and 3219 (85.9%) met the inclusion criteria. Of these, 220 patients (6.8%) developed postoperative pulmonary infection, 1285 (39.9%) of whom were male. Their ages ranged from 22 to 89 years, with a median age of 59 years (IQR: 50–65 years).
Fig. 1.

Patient recruitment flowchart
Among these patients, 564 (17.5%) underwent cuneiform resection, 636 (19.8%) underwent segmentectomy, 1948 (60.5%) underwent lobectomy and 71 (2.2%) underwent total pneumonectomy. The training cohort comprised 2253 patients, with a median age of 59 years (IQR: 50–65 years). Of these patients, 907 (40.3%) were male, and 156 (6.9%) developed postoperative pulmonary infection. There were 996 patients in the validation cohort, with a median age of 59 years (IQR: 50–66 years). Of these patients, 378 (39.1%) were male, and 64 (6.6%) developed postoperative pulmonary infection. All model input features are shown in Table 1. The characteristics of the training and validation cohorts were not statistically different.
Table 1.
Features in the training cohort and the validation cohort
| Features | Training cohort (n = 2253) | Validation cohort (n = 966) |
P |
|---|---|---|---|
| Pulmonary infection | 0.613 | ||
| Yes | 158(7.01%) | 62(6.42%) | |
| No | 220(9.76%) | 904(93.58%) | |
| Gender (n, %) | 0.607 | ||
| Male | 1024(45.45%) | 427(44.20%) | |
| Female | 1229(54.55%) | 539(55.80%) | |
| Age (n, %) | 0.553 | ||
| < 45 years | 353(15.67%) | 147(15.22%) | |
| 45-65years | 1353(60.05%) | 575(59.52%) | |
| > 65 years | 547(24.28%) | 244(25.26%) | |
| BMI (n, %) | 0.475 | ||
| < 18.5 | 119(5.28%) | 41(4.24%) | |
| 18.5–24 | 1284(56.99%) | 562(58.18%) | |
| 24.1–28 | 712(31.60%) | 294(30.43%) | |
| > 28 | 138(6.13%) | 69(7.14%) | |
| Daily smoking (n, %) | 0.358 | ||
| No smoking | 2061(91.48%) | 892(92.34%) | |
| < 10 roots | 61(2.71%) | 24(2.48%) | |
| 10–20 roots | 101(4.48%) | 41(4.24%) | |
| > 20 roots | 30(1.33%) | 9(0.93%) | |
| Hypertension (n, %) | 0.638 | ||
| Yes | 150(6.66%) | 60(6.21%) | |
| No | 2103(93.34%) | 906(93.79%) | |
| Diabetes (n, %) | 0.893 | ||
| Yes | 45(2.00%) | 20(2.07%) | |
| No | 2208(98.00%) | 956(98.96%) | |
| FEV1/FVC (n, %) | 0.681 | ||
| FEV1/FVC≥80% | 1582(70.22%) | 681(70.50%) | |
| 70%≤FEV1/FVC < 80% | 235(10.43%) | 87(9.01%) | |
| 50%≤FEV1/FVC < 70% | 233(10.34%) | 104(10.77%) | |
| FEV1/FVC < 50% | 203(9.01%) | 94(9.73%) | |
| DLCO%Pred (n, %) | 0.757 | ||
| DLCO%Pred ≥ 80% | 1978(87.79%) | 842(87.16%) | |
| 60%≤ DLCO%Pred < 80% | 46(2.04%) | 24(2.48%) | |
| 40%≤ DLCO%Pred < 60% | 207(9.19%) | 91(9.42%) | |
| DLCO%Pred < 40% | 22(0.98%) | 9(0.93%) | |
| Raw%Pred (n, %) | 0.674 | ||
| Raw%Pred ≤ 120% | 810(35.95%) | 354(36.65%) | |
| 120%< Raw%Pred ≤ 140% | 585(25.97%) | 234(24.22%) | |
| 140%< Raw%Pred ≤ 160% | 777(34.49%) | 332(34.37%) | |
| Raw%Pred > 160% | 81(3.60%) | 46(4.76%) | |
| FVC (n, %) | 0.808 | ||
| FVC<3 L | 1184(52.55%) | 513(53.11%) | |
| 3 L < FVC<4 L | 836(37.11%) | 356(36.85%) | |
| FVC>4 L | 233(10.34%) | 97(10.04%) | |
| Surgical level (n, %) | 0.578 | ||
| Second-level surgery | 549(24.37%) | 226(23.40%) | |
| Third-level surgery | 1696(75.28%) | 737(76.29%) | |
| Fourth-level surgery | 8(0.36%) | 3(0.31%) | |
| Mode of operation (n, %) | 0.249 | ||
| Wedge resection | 405(17.98%) | 159(16.46%) | |
| Segmentectomy | 445(19.75%) | 191(19.77%) | |
| Lobectomy | 1357(60.23%) | 591(61.18%) | |
| Pneumonectomy | 46(2.04%) | 25(2.59%) | |
| Surgical site (n, %) | 0.618 | ||
| Left lung | 834(37.02%) | 366(37.89%) | |
| Right lung | 1396(61.96%) | 591(61.18%) | |
| Both lungs | 23(1.02%) | 9(0.93%) | |
| Intraoperative lymph node dissection (n, %) | 0.679 | ||
| Clean out | 1523(67.60%) | 658(68.12%) | |
| Partial cleaning | 433(19.22%) | 203(21.01%) | |
| No cleaning | 297(13.18%) | 105(10.87%) | |
| Number of intraoperative indwelling drainage tubes (n, %) | 0.178 | ||
| single | 1289(57.21%) | 585(60.56%) | |
| Two or more | 964(42.79%) | 381(39.44%) | |
| Intraoperative pathological type (n, %) | 0.488 | ||
| Lung squamous cell carcinoma | 1618(71.82%) | 686(71.01%) | |
| Lung adenocarcinoma | 556(24.68%) | 242(25.05%) | |
| Lung adenosquamous carcinoma | 79(3.51%) | 38(3.93%) | |
| Intraoperative hilar activity (n, %) | 0.116 | ||
| Normal motion | 556(24.68%) | 226(23.40%) | |
| Strong activity | 1697(75.32%) | 740(76.60%) | |
| Intraoperative pleural adhesion degree (n, %) | 0.219 | ||
| Non-adhesion | 2123(94.23%) | 899(93.06%) | |
| Minor adhesion | 46(2.04%) | 23(2.38%) | |
| Moderate adhesion | 35(1.55%) | 18(1.86%) | |
| Extensive adhesion | 49(2.17%) | 26(2.69%) | |
| Degree of tumor invasion (n, %) | 0.74 | ||
| Invasive carcinoma | 1164(51.66%) | 492(50.93%) | |
| Microinvasive carcinoma | 590(26.19%) | 277(28.67%) | |
| Carcinoma in situ | 499(22.15%) | 197(20.39%) | |
| Maximum tumor diameter (n, %) | 0.247 | ||
| < 1 cm | 874(38.79%) | 391(40.48%) | |
| 1–2 cm | 777(34.49%) | 335(34.68%) | |
| > 2 cm | 602(26.72%) | 240(24.84%) | |
| 24 h postoperative activities of daily living (n, %) | 0.407 | ||
| Mild dependence | 109(4.84%) | 52(5.38%) | |
| Moderate dependence | 2009(89.17%) | 861(89.13%) | |
| Heavily dependence | 135(5.99%) | 53(5.49%) | |
| 24 h postoperative critical MEWs (n, %) | 0.757 | ||
| < 4 points | 1536(68.18%) | 653(67.60%) | |
| 4–5 points | 712(31.60%) | 311(32.19%) | |
| > 5 points | 5(0.22%) | 2(0.21%) | |
| Properties of thoracic fluid 24 h after surgery (n, %) | 0.658 | ||
| serous | 60(2.66%) | 19(1.97%) | |
| Light bloody | 1721(76.39%) | 744(77.02%) | |
| Bloody | 472(20.95%) | 203(21.01%) | |
| Thoracic fluid volume 24 h after surgery (n, %) | 0.467 | ||
| < 100 ml | 300(13.32%) | 114(11.80%) | |
| 100-300 ml | 1874(83.18%) | 818(84.68%) | |
| 301-500 ml | 40(1.78%) | 20(2.07%) | |
| > 500 ml | 39(1.73%) | 14(1.45%) | |
| ONS (n, %) | 0.287 | ||
| Yes | 246(10.92%) | 118(12.22%) | |
| No | 2007(89.08%) | 848(87.78%) | |
| Postoperative urinary catheter (n, %) | 0.904 | ||
| Yes | 2087(92.63%) | 896(92.75%) | |
| No | 166(7.37%) | 70(7.25%) |
Feature screening
The choice of the parameter λ is crucial in the LASSO modeling and variable screening process, as it directly affects the degree to which the coefficients are compressed in the model. In this study, the method to determine the optimal value of the parameter λ is to divide the data into K groups of the same size, then select an extreme value of λ, use LASSO to fit the data. We choose the λ value corresponding to the maximum value within one standard error of the minimum mean square error as the optimal value in this study. We selected the λ value of 0.002126 under 10 cross-validations as the optimal value of the model for feature screening. After this step, the remaining 9 features were used for downstream logistic regression modelling. The details of LASSO feature selection are shown in Fig. 2; Table 2. The logistic regression results for the training cohort are shown in Table 3.
Fig. 2.

Lasso regression to generate the selected clinic features with iterative fitting using 10-fold cross-validation. A Variation of the hyperparameter λ and mean square error in Lasso regression. B The coefficient profiles of clinic features
Table 2.
LASSO regression results of important variables related to pulmonary infection after thoracoscopic lung cancer resection in the training cohort
| Variables | Coefficient | Lambda. min |
|---|---|---|
| IPT.LUSC | 1.041545 | 0.002126 |
| MTD | 0.298411 | |
| IPAD | 0.363495 | |
| Daily smoking | 0.856869 | |
| DLCO%Pred | 0.352636 | |
| Surgical level | 0.813694 | |
| TFV | 2.815829 | |
| ONS | 2.541024 | |
| PuC | 3.321835 |
IPT.LUSC: Intraoperative pathological type: Lung squamous cell carcinoma; MTD: Maximum tumor diameter; IPAD: Intraoperative pleural adhesion degree; DLCO%Pred: Diffusing capacity of the lung for carbon monoxide percentage of predicted; TFV: Thoracic fluid volume 24 h after surgery; ONS: Perioperative oral nutritional supplements; PUC: Postoperative urinary catheter
Table 3.
Multivariate logistic regression analysis results in the training cohort
| Variables | Adjusted OR | Upper | Lower | P |
|---|---|---|---|---|
| MTD | 2.715 | 1.856 | 3.825 | < 0.001 |
| Daily smoking | 3.209 | 1.976 | 3.962 | < 0.001 |
| DLCO%Pred | 1.513 | 1.308 | 2.933 | 0.003 |
| TFV | 2.642 | 1.850 | 3.761 | < 0.001 |
| ONS | 2.635 | 1.793 | 3.694 | < 0.001 |
| PIC | 1.926 | 1.535 | 2.961 | 0.002 |
MTD: Maximum tumor diameter; DLCO%Pred: Diffusing capacity of the lung for carbon monoxide percentage of predicted; TFV: Thoracic fluid volume 24 h after surgery; ONS: Perioperative oral nutritional supplements; PUC: Postoperative urinary catheter
Model performance
For the support vector machines, we tested five kernels for the SVM model: linear, polynomial, gaussian radial basis, sigmoid and laplacian. When the regularization parameter C value is 0.07, the kernel parameter γ value is 0.03, the error of the gaussian radial basis was minimal. For the artificial neural network, the learning rate was set to be 0.001, the number of training iterations was 200, and the lowest error rate was found with 3 layers and 64 neurons in the hidden layer. For the eXtreme gradient boosting, the final selected hyperparameter configuration was a maximum tree depth of 7, a learning rate of 0.2, a gamma value of 0.1, and a minimum child node weight of 2, which striked a good balance between model complexity and prediction performance. For the random forest, when the number of trees is 100 and the maximum depth of a tree is 14, the cross-validation error is minimal. The performance of the five models is summarised in Table 4. Figure 3(A and B) shows the AUC ranges of the five models in the training and validation cohorts. Based on the comparison of the AUC, accuracy, sensitivity, specificity, PPV, NPV, Recall and F1, XGB demonstrated the best performance among the five models. As shown in Fig. 3(C and D), our XGB model was reliable based on the Brier score for predicting pulmonary infection, which was 0.025 for the validation cohort.
Table 4.
Five model performances in training cohort and validation cohort
| AUC(95%CI) | Accuracy(95%CI) | Sensitivity | Specificity | PPV | NPV | Precision | Recall | F1 | Brier | |
|---|---|---|---|---|---|---|---|---|---|---|
| Training cohort | ||||||||||
| LR | 0.969(0.956,0.973) | 0.898(0.879,0.905) | 0.894 | 0.819 | 0.882 | 0.689 | 0.892 | 0.864 | 0.818 | 0.039 |
| ANN | 0.975(0.968,0.982) | 0.910(0.893,0.926) | 0.917 | 0.802 | 0.895 | 0.631 | 0.895 | 0.874 | 0.764 | 0.034 |
| SVM | 0.954(0.945,0.963) | 0.890(0.882,0.907) | 0.929 | 0.887 | 0.892 | 0.720 | 0.893 | 0.909 | 0.859 | 0.058 |
| RF | 0.982(0.974,0.988) | 0.905(0.898,0.920) | 0.933 | 0.859 | 0.878 | 0.629 | 0.903 | 0.876 | 0.887 | 0.030 |
| XGB | 0.986(0.977,0.991) | 0.905(0.898,0.921) | 0.946 | 0.899 | 0.908 | 0.765 | 0.918 | 0.906 | 0.907 | 0.028 |
| Validation cohort | ||||||||||
| LR | 0.948(0.931,0.957) | 0.891(0.873,0.905) | 0.891 | 0.852 | 0.832 | 0.715 | 0.868 | 0.861 | 0.780 | 0.042 |
| ANN | 0.949(0.939,0.955) | 0.896(0.879,0.908) | 0.905 | 0.793 | 0.845 | 0.668 | 0.835 | 0.869 | 0.742 | 0.030 |
| SVM | 0.883(0.873,0.891) | 0.872(0.852,0.889) | 0.893 | 0.809 | 0.818 | 0.782 | 0.885 | 0.843 | 0.770 | 0.040 |
| RF | 0.948(0.929,0.960) | 0.898(0.884,0.912) | 0.901 | 0.763 | 0.826 | 0.691 | 0.814 | 0.853 | 0.784 | 0.025 |
| XGB | 0.951(0.943,0.964) | 0.902(0.886,0.913) | 0.927 | 0.864 | 0.898 | 0.824 | 0.908 | 0.872 | 0.815 | 0.025 |
LR: Logistic regression; ANN: Artificial neural network; SVM: Support vector machine; RF: Random forest; XGB: eXtreme gradient boosting
Fig. 3.

ROC curves and calibration plots of five different machine learning models. A ROC curves of five different machine learning models in the training cohort. B ROC curves of five different machine learning models in the validation cohort. C The calibration plots of five different machine learning models in the training cohort. D The calibration plots of five different machine learning models in the validation cohort. LR: Logistic regression; ANN: Artificial neural network; SVM: Support vector machine; RF: Random forest; XGB: eXtreme gradient boosting
In our study, the XGB model demonstrated the highest AUC performance. We combine the methods of SHAP to assesses the relevance of every predictor variable. A high SHAP value indicates a positive influence on the model’s output, whereas a low value suggests the opposite effect. Figure 4A presents the SHAP summary plot, which ranks feature importance in the XGB model. Each point in the plot represents the Shapley value for a specific feature and instance. The y-axis position reflects the feature’s importance, while the x-axis corresponds to the Shapley value. Overlapping points are distributed along the y-axis, illustrating the density of Shapley values per feature. The horizontal position shows the degree to which the value was high or low relative to the predicted value: red dots indicate high risk, and blue dots indicate low risk. Additionally, Fig. 4B highlights the significance of each predictor variable in the XGB model’s predictions.
Fig. 4.

SHAP summary plots of key features in XGB model. A Each characteristic attribute in SHAP. B the SHAP-based feature importance ranking. TFV: Thoracic fluid volume 24 h after surgery; PUC: Postoperative urinary catheter; MTD: Maximum tumor diameter; ONS: Perioperative ONS supplements; DLCO%Pred: Diffusing capacity of the lung for carbon monoxide percentage of predicted; IPT.LUSC: Intraoperative pathological type: Lung squamous cell carcinoma; IPAD: Intraoperative pleural adhesion degree; IPT.LUASC: Intraoperative pathological type: Lung adenosquamous carcinoma
In addition, we compared five models’ performance for different age subgroups and found that they performed best for patients between 45 and 65 years of age and poorer for those younger than 45 years. Figure 5 illustrates the performance of these models stratified by age: less than 45 years, 45–65 years and more than 65 years; the AUC values of the XGB model for these three subgroups were 0.876, 0.946 and 0.903, respectively.
Fig. 5.

Area under the receiver operating characteristic curve (AUC) for machine learning models stratified by age. XGB: eXtreme gradient boosting; LR: logistic regression; ANN: artificial neural network; RF: random forest; SVM: support vector machine
Nomogram
Most risk prediction models based on machine learning algorithms are not amenable to visualisation and are not convenient for clinical medical staff to use directly. In this study, the performance of the LR model was good, with an AUC of 0.948 (95% confidence interval: 0.931–0.957), second only to the XGB model. The nomogram was constructed using the nine candidate factors screened by LASSO to improve the operability of the model in actual clinical settings (Fig. 6).
Fig. 6.

Nomogram for estimating pulmonary infection in patients with thoracoscopic lung cancer resection. IPT: Intraoperative pathological type; MTD: Maximum tumor diameter; IPAD: Intraoperative pleural adhesion degree; PDL: Preoperative diffusion of lung; TFV: Thoracic fluid volume 24 h after surgery; ONS: Perioperative ONS supplement; PIC: Postoperative indwelling catheters
Discussion
In this study, we systematically reviewed and screened the risk factors for pulmonary infection after thoracoscopic lung cancer resection. Through the hospital medical record system, we collected the relevant information of patients who underwent thoracoscopic lung cancer resection to ensure the integrity of the risk factors. Then, we used the nine best characteristic variables screened by LASSO regression to establish the prediction model of pulmonary infection in patients after thoracoscopic lung cancer resection.
In our cohort, 6.83% of patients experienced pulmonary infection after thoracoscopic lung cancer resection. Five machine learning–based risk models for predicting postoperative pulmonary infection risk were developed and internally validated. Quantifying the performance of the prediction models and comparing the advantages and disadvantages of different models are important aspects in gauging the model’s practical utility. In our study, the AUC of the models ranged from 0.823 to 0.946, suggesting that the models had good predictive ability. The sensitivity ranged from 0.891 to 0.927, and the positive predictive value ranged from 0.818 to 0.898, indicating a good sensitivity for patients at high risk for pulmonary infections after thoracoscopic lung cancer resection. The specificity of the five models was good (0.763–0.864), indicating that the models could distinguish pulmonary infection from other similar diseases and reduce the incidence of misdiagnosis. Given the adverse consequences of lung cancer patients developing pulmonary infections after surgery, the more important goal of these models is to capture as many potential cases of pulmonary infection as possible. These five models have the potential to be effective tools to assist in the clinical management of pulmonary infections in patients after thoracoscopic lung cancer resection, and their capabilities will eventually improve as clinical data accumulate and more input features are added.
It is worth noting that in the internal verification process, we found that the XGB model had the best prediction performance, with good discrimination and calibration characteristics. The main reason for this outcome is that XGB is an integrated learning algorithm with superior learning and generalisation ability. It can handle large datasets and high-dimensional features effectively and optimise model performance by using gradient lifting to provide accurate predictions with a small generalisation error. The regularisation terms can be used to control the complexity of the model and prevent overfitting. It also has unique advantages in dealing with high-dimensional variables and complex interactions and nonlinear relations between the variables. XGB has been used in many machine learning and data mining challenges [19, 20]. As shown in Fig. 4B, the top five predictors are our focus of observation in future clinical work. First, thoracic fluid volume 24 h after surgery is the first risk factor for lung infection. Massive blood loss promotes the release of pro-inflammatory factors, triggers systemic inflammatory response syndrome, and increases the probability of pulmonary inflammatory response after surgery [21]. In addition, catheterization increases the risk of urinary tract infections in patients, leading to retrograde infectious diseases. Catheterization can also lead to reduced mobility and a buildup of pleural fluid in the lungs, increasing the risk of rapid bacterial reproduction in the patient’s lungs [22]. Preoperative smoking usually damages the epithelial cilia of the respiratory tract, which can easily lead to sputum retention. Increased airway responsiveness in smoking patients and easy injury to the respiratory mucosa during intraoperative anesthesia make it easier for bacteria to invade the lungs, which is consistent with Yang’s findings [23]. Patients with larger tumor diameters have a greater degree of lung parenchyma destruction and greater damage to lung function. Expansion of surgical resection and prolonged operative time both increase the chances of pathogens multiplying in lung tissue [24]. Nutritional supplementation is also a factor that we should pay attention to when preventing postoperative lung infections. Studies have shown that inadequate protein supplementation in cancer patients is associated with an increased incidence of postoperative complications [25, 26]. This suggests that we should pay more attention to the above aspects in clinical work and formulate targeted preventive measures.
Unexpectedly, when we performed a subgroup analysis of the participating patients grouped by age, we found that these models performed best in predicting the risk of pulmonary infection in patients aged 45–65 years and showed low predictive performance in patients younger than 45 years of age. In recent years, more and more studies have found that there are differences in the proportion of gender, incidence characteristics, disease progression, prognostic influencing factors, pathological features, clinicopathological and gene mutation characteristics between the young (18–45 years) and old (> 45 years) lung cancer patients [27–30]. This is one of the important reasons for the differences in the performance of risk prediction models in different age groups in this study. Against the backdrop that the incidence rate of lung cancer patients is gradually getting younger, in the future, we should utilize big data and machine learning algorithms to establish separate models for different age groups, carry out individual prediction and hierarchical management of the risk of pulmonary infection in patients with thoracoscopic lung cancer resection, improve the efficiency of postoperative complication management in young patients, achieve more accurate prediction and precise intervention.
There are many obstacles to translating risk prediction models built on XGB models into clinical or other real-world settings. One of the main obstacles is the poor interpretability of machine learning model results; therefore, they are unlikely to be immediately accepted by clinical users [31]. A predictive model suitable for clinical use not only needs good predictive performance, but it also needs to have clinical universality and promotional value, along with being easy to use by medical personnel, which can help medical personnel accept and apply it in clinical practice more quickly. The LR model showed similar performance to the XGB model in this study; therefore, we translated it into a nomogram to contribute to extending its applicability and increasing its clinical significance. Nomograms offer several advantages, including visualisation, convenience and timeliness. In clinical practice, this has been more useful than other prediction systems for predicting the risk of pulmonary infection among medical personnel [32–34]. In the nomogram, multiple prediction indicators are integrated into a graph to show the relationship between the variables of the model in a scaled line segment on the same plane. A line segment’s length indicates how significant it is for the resulting event. As shown in the figure, ‘Point’ represents each variable’s score under different values, and ‘Total Points’ represents the sum of all the variables’ scores. The total score is converted into the probability of the outcome event based on a function conversion relationship. Thus, the probability of pulmonary infection after thoracoscopic lung cancer resection can be obtained. Patients at high risk for pulmonary infections can be quickly identified by caregivers using the tool.
Machine learning algorithms have obvious advantages and broad application prospects in the medical field. However, further research and exploration are needed to transform the research results and apply them in clinical situations. In subsequent studies, we plan to incorporate machine learning algorithms to optimize the model and integrate it into the hospital’s electronic medical record system. The prediction model is synchronously updated to provide automatic, real-time, efficient and intelligent early warning of pulmonary infection after thoracoscopic lung cancer resection, with the goal to assist clinicians in early assessment and judgment of the patient’s status.
Our study has several limitations. First, this was a retrospective study, and future prospective randomized controlled trials are needed to provide higher-quality evidence. Second, our data were obtained from a single large academic medical center, lacking external multicenter validation. In future studies, we aim to collect sufficient independent external datasets from multiple institutions across different time periods and geographic regions to comprehensively evaluate the model’s generalizability through temporal and geographical validation. Additionally, further research is needed to explore whether factors such as family history, biomarkers, and other potential variables influence the risk of pulmonary infection following thoracoscopic lung cancer resection. We intend to focus on addressing these issues in our future research.
Conclusions
In this study, we established a risk prediction model for pulmonary infection after thoracoscopic lung cancer resection based on five machine learning algorithms, of which XGBoost best predicted the occurrence of postoperative pulmonary infection in patients. We believe that the nomogram may be a useful tool for nursing staff to predict and manage pulmonary infections in patients after thoracoscopic lung cancer resection. Our model could eventually lead to more precise prevention of pulmonary infections in patients who have undergone thoracoscopic lung cancer resection.
Electronic supplementary material
Below is the link to the electronic supplementary material.
Acknowledgements
Not applicable.
Abbreviations
- LASSO
Least absolute shrinkage and selection operator
- LR
Logistic regression
- ANN
Artificial neural network
- SVM
Support vector machine
- RF
Random forest
- XGB
eXtreme gradient boosting
- BMI
Body mass index
- FEV1/FVC
Forced expiratory volume in one second/forced vital capacity
- DLCO%Pred
Diffusing capacity of the lung for carbon monoxide percentage of predicted
- Raw%Pred
Airway resistance percentage of predicted
- FVC
Forced vital capacity
- MTD
Maximum tumor diameter
- IPAD
Intraoperative pleural adhesion degree
- MEWs
Modified Early Warning score
- TFV
Thoracic fluid volume 24 h after surgery
- ONS
Perioperative oral nutritional supplements
- PUC
Postoperative urinary catheter
- IPT
Intraoperative pathological type
- IPT.LUSC
Intraoperative pathological type: Lung squamous cell carcinoma
- IPT.LUAD
Intraoperative pathological type: Lung adenocarcinoma
- IPT.LUASC
Intraoperative pathological type: Lung adenosquamous carcinoma
- AUC
Area under curve
- PPV
Positive predictive value
- NPV
Negative predictive value
- F1
F1 score
- SHAP
Shapley additive explanation
Author contributions
Conceptualization, JM, ZZ, BX, JF and XL; methodology, JM, ZZ, LY and HC; formal analysis, JM, ZZ and HC; data curation, JM, ZZ and XL; writing-original draft preparation, JM and XL; writing-review and editing, ZZ, BX, and XL; supervision, BX, JF and XL. All authors read and approved the final manuscript.
Funding
This study was supported by Shanghai Jiao Tong University School of Medicine: Nursing Development Program, Shanghai Jiao Tong University School of Medicine Nursing Research Project (Jyh2401), Shanghai 2024 “Science and Technology Innovation Action Plan” Science Popularization Special Project (24DZ2300700), the 2024 Shanghai Health System Young Talent Award Foundation’s First Jahwa-Nursing Special Technology Support Project, and the 2024 Shanghai Hospital Development Center Municipal Hospital Diagnosis and Treatment Technology Promotion and Optimization Management Project (SHDC22024210). The funding bodies only provided the financial means to allow the authors to carry out the study, and played no role in the design of the study and collection, analysis, and interpretation of data and in writing the manuscript.
Data availability
The datasets generated and/or analyzed during the current study are not publicly available due to ownership by the Nursing Department, Shanghai Chest Hospital affiliated to Shanghai Jiaotong University School of Medicine, Shanghai, China, but are available from the corresponding author on reasonable request.
Declarations
Ethics approval and consent to participate
The procedure mentioned in this retrospective study involving human participants was performed in accordance with the Declaration of Helsinki (as revised in 2013) and approved by the Medical Ethics Committee of Shanghai Chest Hospital Affiliated to Shanghai Jiaotong University School of Medicine (IRB Approval No. KS23016). The need for written informed consent from individual patients was waived by the Medical Ethics Committee of Shanghai Chest Hospital Affiliated to Shanghai Jiaotong University School of Medicine because all data were anonymised for research purposes.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Jiajia Ma and Zhengmin Zhang contributed equally to this work.
References
- 1.WHO. Global cancer burden growing, amidst mounting need for services. 2024. https://www.who.int/news/item. Accessed 17 Feb 2024. [PMC free article] [PubMed]
- 2.Chen S, Cao Z, Prettner K, Kuhn M, Yang J, Jiao L, et al. Estimates and Projections of the Global Economic Cost of 29 Cancers in 204 Countries and Territories From 2020 to 2050. JAMA Oncol. 2023;9(4):465–72. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Zheng Y, Mao M, Li F, Wang L, Zhang X, Zhang X, Wang H, Zhou H, Ji M, Wang Y, Liu L, Zhu Q, Reinhardt JD, Lu X. Effects of enhanced recovery after surgery plus pulmonary rehabilitation on complications after video-assisted lung cancer surgery: a multicentre randomised controlled trial. Thorax. 2023;78(6):574–86. [DOI] [PubMed] [Google Scholar]
- 4.Wang YQ, Liu X, Jia Y, Xie J. Impact of breathing exercises in subjects with lung cancer undergoing surgical resection: a systematic review and meta-analysis. J Clin Nurs. 2019;28(5–6):717–32. [DOI] [PubMed] [Google Scholar]
- 5.Semenkovich TR, Frederiksen C, Hudson JL, Subramanian M, Kollef MH, Patterson GA, Kreisel D, Meyers BF, Kozower BD, Puri V. Postoperative Pneumonia Prevention in Pulmonary Resections: A Feasibility Pilot Study. Ann Thorac Surg. 2019;107(1):262–70. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Cecen B, Topates G, Kara A, Akbulut SO, Havitcioglu H, Kozaci LD. Biocompatibility of silicon nitride produced via partial sintering & tape casting. Ceram Int. 2021;47(3):3938–45. [Google Scholar]
- 7.Deo RC. Machine learning in medicine. Circulation. 2015;132(20):1920–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Leeming J. How AI is helping the natural sciences. Nature. 2021;598(7880):S5–7. [Google Scholar]
- 9.Gao Z, Lou L, Wang M, Sun Z, Chen X, Zhang X, Pan Z, Hao H, Zhang Y, Quan S, Yin S, Lin C, Shen X. Application of machine learning in intelligent medical image diagnosis and construction of intelligent service process. Comput Intell Neurosci. 2022:9152605. [DOI] [PMC free article] [PubMed]
- 10.Schwalbe N, Wahl B. Artificial intelligence and the future of global health. Lancet. 2020;395(10236):1579–86. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Wang Y, zhu Y, Xue Q, Ji MH, Tong JH, Yang JJ, Zhou CM. Predicting chronic pain in post-operative breast cancer patients with multiple machine learning and deep learning models. J Clin Anesth. 2021;74:110423. [DOI] [PubMed] [Google Scholar]
- 12.Li MP, Liu WC, Wu JB, Luo K, Liu Y, Zhang Y, Xiao SN, Liu ZL, Huang SH, Liu JM. Machine learning for the prediction of postoperative nosocomial pulmonary infection in patients with spinal cord injury. Eur Spine J. 2023;32(11):3825–35. [DOI] [PubMed] [Google Scholar]
- 13.Lu C, Xing ZX, Xia XG, Long ZD, Chen B, Zhou P, Wang R. Development and validation of a postoperative pulmonary infection prediction model for patients with primary hepatic carcinoma. World J Gastrointest Oncol. 2023;15(7):1241–52. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Luo Y, Tang Z, Hu X, Lu S, Miao B, Hong S, Bai H, Sun C, Qiu J, Liang H, Na N. Machine learning for the prediction of severe pneumonia during posttransplant hospitalization in recipients of a deceased-donor kidney transplant. Ann Transl Med. 2020;8(4):82. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Dong Z, Liu G, Tu L, Su X, Yu Y. Establishment of a prediction model of postoperative infection complications in patients with gastric cancer and its impact on prognosis. J Gastrointest Oncol. 2023;14(3):1250–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Sun T, Liu J, Yuan H, Li X, Yan H. Construction of a risk prediction model for lung infection after chemotherapy in lung cancer patients based on the machine learning algorithm. Front Oncol. 2024;14:1403392. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Ettinger DS, Wood DE, Akerley W, Bazhenova LA, Borghaei H, Camidge DR, et al. Non-Small Cell Lung Cancer, Version 6.2015. J Natl Compr Canc Netw. 2015;13(5):515–24. [DOI] [PubMed] [Google Scholar]
- 18.Shi Y, Huang Y, Zhang TT, Cao B, Wang H, Zhuo C, et al. Chinese guidelines for the diagnosis and treatment of hospital-acquired pneumonia and ventilator-associated pneumonia in adults (2018 Edition). J Thorac Dis. 2019;11(6):2581–616. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Wang R, Zhang J, Shan B, He M, Xu J. XGBoost Machine Learning Algorithm for Prediction of Outcome in Aneurysmal Subarachnoid Hemorrhage. Neuropsychiatr Dis Treat. 2022;18:659–67. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Shin H. XGBoost Regression of the Most Significant Photoplethysmogram Features for Assessing Vascular Aging. IEEE J Biomed Health Inf. 2022;26(7):3354–61. [DOI] [PubMed] [Google Scholar]
- 21.Narala VR, Narala SR, Aiya Subramani P, Panati K, Kolliputi N. Role of mitochondria in inflammatory lung diseases. Front Pharmacol. 2024;15:1433961. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Nollen JM, Brunsveld-Reinders AH, Steyerberg EW, Peul W, van Furth WR. Improving postoperative care for neurosurgical patients by a standardised protocol for urinary catheter placement: a multicentre before-and-after implementation study. BMJ Open Qual. 2025;14(2):e003073. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Yang K, Wei H, Zhu W, Xu Y, Wang S, Fan F, Zhang K, Yuan Q, Wang H. Clinical characteristics and risk factors of late-stage lung adenocarcinoma patients with bacterial pulmonary infection and its relationship with cellular immune function. Front Immunol. 2025;16:1559211. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Chen S, Huang Q, Fu F, Wang Z, Zhang Y, Chen H. Segmentectomy for ground glass-dominant invasive lung cancer with tumour diameter of 2–3 cm: protocol for a single-arm, multicentre, phase III trial (ECTOP1012). BMJ Open. 2024;14(7):e087088. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Dingemans AM, van Walree N, Schramel F, Soud MY, Baltruškevičienė E, Lybaert W, Veldhorst M, van den Berg CA, Kaasa S. High Protein Oral Nutritional Supplements Enable the Majority of Cancer Patients to Meet Protein Intake Recommendations during Systemic Anti-Cancer Treatment: A Randomised Controlled Parallel-Group Study. Nutrients. 2023;15(24):5030. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Simonsen C, de Heer P, Bjerre ED, Suetta C, Hojman P, Pedersen BK, Svendsen LB, Christensen JF. Sarcopenia and Postoperative Complication Risk in Gastrointestinal Surgical Oncology: A Meta-analysis. Ann Surg. 2018;268(1):58–69. [DOI] [PubMed] [Google Scholar]
- 27.Li J, Pan B, Huang Q, Zhan C, Lin T, Qiu Y, Zhang H, Xie X, Lin X, Liu M, Wang L, Zhou C. A Nomogram for Predicting Cancer-Specific Survival in Young Patients With Advanced Lung Cancer Based on Competing Risk Model. Clin Respir J. 2024;18(8):e13800. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Dong Y, Zhou S, Li J, Zhang Y, Che G. Distant metastatic patterns in young and old non-small cell lung cancer patients: A dose–response analysis based on SEER population. Heliyon. 2024;10(17):e36657. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Lu X, Chen Y, Li Y, Tang M, Zheng X. Different clinicopathological features between young and older patients with pulmonary adenocarcinoma and ground-glass opacity. Sci Rep. 2024;14(1):15679. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Xia J, Li H, Zhang R, Wang J. Clinicopathological characteristics and prognosis of young patients aged ≤ 45 years old with non-small cell lung cancer. Open Med (Wars). 2023;18(1):20230684. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Chekroud AM, Bondar J, Delgadillo J, Doherty G, Wasil A, Fokkema M, et al. The promise of machine learning in predicting treatment outcomes in psychiatry. World Psychiatry. 2021;20(2):154–70. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Hu XY, Liu H, Zhao X, Sun X, Zhou J, Gao X, et al. Automated machine learning-based model predicts postoperative delirium using readily extractable perioperative collected electronic data. CNS Neurosci Ther. 2022;28(4):608–18. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Mao D, Fu LJ, Zhang WL. Construction and validation of an early prediction model of delirium in children after congenital heart surgery. Transl Pediatr. 2022;11(6):954–64. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Shi Y, Wang H, Zhang L, Zhang M, Shi XY, Pei HH, et al. Nomogram models for predicting delirium of patients in emergency intensive care unit: a retrospective cohort study. Int J Gen Med. 2022;15:4259–72. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The datasets generated and/or analyzed during the current study are not publicly available due to ownership by the Nursing Department, Shanghai Chest Hospital affiliated to Shanghai Jiaotong University School of Medicine, Shanghai, China, but are available from the corresponding author on reasonable request.
