Skip to main content
NPJ Digital Medicine logoLink to NPJ Digital Medicine
. 2025 Oct 17;8:617. doi: 10.1038/s41746-025-01989-1

Semi-automated surveillance of surgical site infections using machine learning and rule-based classification models

Américo Agostinho 1, Etienne Chalot 1, Daniel Teixeira 1, Davide Bosetti 1, Niccolò Buetti 1,2, Gaud Catho 1,2, Stephan Harbarth 1,2, Mohamed Abbas 1,2,3,
PMCID: PMC12534672  PMID: 41107441

Abstract

Surgical site infections (SSIs), among the most frequent healthcare-associated infections, require surveillance, but traditional methods are labour-intensive. We developed machine learning (ML) and rule-based models for the semi-automated detection of deep and organ/space SSIs using data from a prospective cohort of 3931 surgical patients. We assessed sensitivity and workload reduction (proportion of patients not requiring manual review) at a 0.5 decision threshold, and computed area under the receiver operating characteristic curve (AUROC) and area under the precision-recall curve (AUPRC). The best-performing ML models (Naïve Bayes and dense neural network) achieved sensitivity up to 0.90, AUROC up to 0.968, AUPRC up to 0.248, and workload reduction over 90%. The rule-based model showed higher sensitivity (1.000) but lower AUROC, AUPRC, and workload reduction. Our findings suggest that semi-automated approaches can support efficient and accurate SSI surveillance while reducing manual workload. Further validation in other settings is warranted.

Subject terms: Health care, Infectious diseases

Introduction

Surgical site infections (SSIs) are among the most common healthcare-associated infections (HAIs), contributing significantly to patient morbidity, prolonged hospital stays, and increased healthcare costs1. Surveillance of HAIs is recognised by the World Health Organization as a core component of effective infection prevention and control (IPC) programmes to help guide the implementation of preventive practices and measure their impact2. More specifically, participation in an SSI surveillance network has been shown to be associated with a decreased SSI risk over time3.

In most settings, SSI surveillance is performed by IPC professionals who perform a manual chart review of all patients undergoing a specific surgical procedure and determine the infection status. However, this approach is time-consuming and has low interrater reliability4. With the increasing adoption of electronic health records (eHRs), there is a growing interest in transitioning to an automated and semi-automated surveillance of SSI516.

Multiple studies have explored automated SSI surveillance using rule-based algorithms or artificial intelligence (AI), particularly machine learning (ML) models1621. Some approaches have demonstrated promising diagnostic accuracy2224, but few performed external validation25. Few have directly assessed the potential of these models to reduce workload20,21,26. Semi-automated approaches, that classify patients as being at low- or high-risk for SSI and where only high-risk cases undergo manual review, are particularly promising for real-world implementation, especially for regulatory surveillance, but evidence comparing machine learning and rule-based models in this context is limited26.

The primary aim of this study was to develop machine learning (ML) models for use in the semi-automated surveillance of deep and organ/space SSI and thus reduce the surveillance workload, whilst minimising undetected SSI cases.

Results

We included 3931 patients (male, 59%; median age, 62 years [interquartile range (IQR), 50-73]) who underwent SSI surveillance during the study period. Patient characteristics are shown in Table 1. The overall risk of deep and organ/space SSI was 4.5%. Incidence varied by surgical category: neurosurgery, 2.0% (35/1731); digestive, 9.3% (122/1318); and cardiac, 2.4% (20/847). Table 2 presents the association between patient characteristics (features) and SSI. Patients who developed SSI were generally older with more complex postoperative courses requiring longer hospital stay, a higher need for reoperations and scar revisions, and more likely to have a high contamination class (class 3–4) and less likely to have an implant. They had significantly more antibiotic days, readmissions, infectious disease consultations, C-reactive protein samples, and more frequent fever episodes.

Table 1.

Patient characteristics

Characteristic Colorectal surgery (N = 1318) Cardiac surgerya (N = 847) Laminectomy and spinal fusion (N = 1766) Total (N = 3931)
Age, years (mean, SD) 65.1 ( ± 15.6) 64.5 ( ± 13.0) 57.3 ( ± 16.1) 61.5 ± 15.8
Sex, female 635 (48.2) 198 (23.4) 780 (44.2) 1613 (41.0)
Weight, kg (mean, SD) 71.9 ( ± 16.4) 78.1 ( ± 16.2) 77.1 ( ± 16.4) 75.6 ± 16.5
Presence of implant 26 (2.0) 804 (94.9) 696 (39.40) 1526 (38.8)
ASA: 1-2 737 (55.9) 48 (5.7) 1360 (77.0) 2145 (54.6)
ASA: >=3 581 (44.1) 799 (94.3) 406 (23.0) 1786 (45.4)
Contamination class
 Clean/clean-contaminated 973 (73.8) 802 (94.7) 1746 (98.9) 3521 (89.6)
 Contaminated/dirty 345 (26.2) 45 (5.3) 20 (1.1) 410 (10.4)
Surgery duration > 75th percentile 770 (58.4) 346 (40.9) 728 (41.2) 1844 (46.9)
NNIS: 0 217 (16.5) 33 (3.9) 858 (48.6) 1108 (28.2)
NNIS: 1 601 (45.6) 456 (53.8) 669 (37.9) 1726 (43.9)
NNIS: 2 404 (30.7) 340 (40.1) 232 (13.1) 976 (24.8)
NNIS: 3 96 (7.3) 18 (2.1) 7 (0.4) 121 (3.1)
Follow-up: completed 1196 (90.7) 741 (87.5) 1682 (95.2) 3619 (92.1)
Follow-up: deceased patient 38 (2.9) 42 (5.0) 40 (2.3) 120 (3.1)
Antibiotic days (mean, SD) 3.0 (5.5) 4.4 (9.4) 2.0 (7.5) 2.9 (7.4)
No. of readmissions (mean, SD) 0.1 (0.3) 0.2 (0.5) 0.1 (0.4) 0.1 (0.4)
No. of ID consults (mean, SD) 0.4 (1.2) 0.9 (2.4) 0.2 (1.1) 0.4 (1.5)
No. of CRP samples (mean, SD) 5.7 (5.0) 8.5 (6.8) 1.4 (3.8) 4.3 (5.8)
No. of CRP samples > 50 mg/l (mean, SD) 3.5 (4.1) 5.3 (4.6) 0.7 (2.6) 2.6 (4.1)
Length of stay after surgery, days (mean, SD) 11.6 (11.0) 12.2 (11.7) 7.3 (8.7) 9.8 (10.4)
No. of days with fever (mean, SD) 0.8 (1.8) 1.1 (2.1) 0.4 (1.4) 0.7 (1.7)
No. of reoperations (mean, SD) 0.2 (0.9) 0.3 (1.0) 0.2 (0.6) 0.2 (0.8)
Scar revision (mean, SD) 0.0 (0.0) 0.0 (0.1) 0.0 (0.2) 0.0 (0.1)
No. of cultures (mean, SD) 1.3 (4.6) 1.9 (7.3) 0.6 (4.1) 1.1 (5.1)
No. of positive cultures (mean, SD) 0.7 (2.7) 0.5 (3.0) 0.2 (1.5) 0.4 (2.3)
No. of radiology tests (mean, SD) 0.7 (1.2) 0.8 (1.4) 0.4 (1.0) 0.5 (1.2)
No. of keywords (mean, SD) 1.5 (3.4) 4.4 (5.9) 0.9 (3.4) 1.8 (4.3)
Actual:expected duration ratio (mean, SD) 1.6 (0.7) 1.3 (0.6) 1.4 (0.8) 1.5 (0.7)
Follow-up: other 84 (6.4) 64 (7.6) 44 (2.5) 192 (4.9)
SSI-type: none 1196 (90.7) 827 (97.6) 1731 (98.0) 3754 (95.5)
SSI-type: deep 18 (1.4) 8 (0.9) 9 (0.5) 35 (0.9)
SSI-type: organ-space 104 (7.9) 12 (1.4) 26 (1.5) 142 (3.6)

aIncludes coronary artery bypass grafting (CABG) and non-CABG.

Values are numbers and percentages unless otherwise specified.

ASA American Society of Anesthesiologists, CRP C-reactive protein; ID infectious diseases, NNIS National Nosocomial Infection Surveillance, SSI surgical site infection.

Table 2.

Distribution of features (characteristics) according to SSI status

Variable No SSI (N = 3824) SSI (N = 107) p-value
Age 60.93 (15.73) 64.76 (14.14) 0.013
Antibiotic days 2.50 (6.86) 15.51 (13.90) <0.001
No. of readmissions 0.10 (0.38) 0.23 (0.47) <0.001
No. of ID consults 0.38 (1.45) 2.50 (2.75) <0.001
No. of CRP samples 4.09 (5.45) 13.69 (8.38) <0.001
No. of CRP samples > 50 mg/l 2.42 (3.81) 9.89 (6.93) <0.001
Length of stay after surgery 9.36 (9.70) 24.53 (20.37) <0.001
No. of days with fever 0.61 (1.60) 2.74 (3.40) <0.001
No. of reoperations 0.18 (0.64) 1.82 (2.30) <0.001
Scar revision 0.01 (0.10) 0.22 (0.52) <0.001
No. of cultures 0.80 (4.03) 11.63 (16.53) <0.001
No. of positive cultures 0.26 (1.59) 5.64 (9.03) <0.001
No. of radiology tests 0.51 (1.11) 1.93 (2.01) <0.001
No. of keywordsa 1.66 (3.95) 7.99 (9.40) <0.001
Actual:expected duration ratio 1.48 (0.71) 1.63 (0.60) 0.023
Presence of implant 0.39 (0.49) 0.21 (0.41) <0.001
Sex, female 0.41 (0.49) 0.46 (0.50) 0.315
Wound contamination class 3-4 0.10 (0.30) 0.32 (0.47) <0.001

aKeywords list: redness, discharge, infection, purulent, pus.

ASA American Society of Anesthesiologists, CRP C-reactive protein, ID infectious diseases, SSI surgical site infection.

The Naïve Bayes classifier and dense neural network (DNN) models achieved the best diagnostic performance (Table 3). The Naïve Bayes model had an NPV of 0.997 (95% CI, 0.996-0.998), an AUROC of 0.968 (95% CI, 0.964-0.971), an AUPRC of 0.248 (95% CI, 0.196-0.300), a FNR of 10.0% (95% CI, 7.42-12.6), and a workload reduction of 91.6%. The DNN model performed similarly, with an NPV, 0.996 (95% CI, 0.992–1.00), AUROC, 0.924 (95% CI, 0.891-0.958), AUPRC, 0.196 (95% CI, 0.152–0.241), FNR, 15.0% (95% CI, 6.67–23.6), and a workload reduction of 90.1% (see ROC curves of these models in Supplementary Figs. 3 and 4). There was no significant decrease in their performance between training and validation datasets, suggesting no substantial overfitting (Supplementary Table 3). Supplementary Table 4 shows model performance across different probability thresholds. The rule-based model yielded an NPV of 1.000 (95% CI, 0.998–1.00), AUROC, 0.859 (95% CI, 0.834–0.884), AUPRC, 0.085 (95% CI, 0.067–0.102), FNR, 0.00% (95% CI, 0.00-4.60), and a workload reduction of 70%.

Table 3.

Model performance in the validation dataset. Values in parentheses are 95%CI. Sensitivity (recall) and specificity are reported; note that these are equivalent to 1 − false negative rate (FNR) and 1 − false positive rate (FPR), respectively

Model Features Negative predictive value Workload reduction (%) Sensitivity Specificity AUROC MCC AUPRC F2 score
Naive Bayes All features 0.997 (0.996-0.998) 91.6 (90.6-92.6) 0.900 (0.874-0.926) 0.937 (0.929-0.946) 0.968 (0.964-0.971) 0.475 (0.426-0.525) 0.248 (0.196-0.300) 0.616 (0.587-0.644)
Dense neural network All features 0.996 (0.992-1.00) 90.1 (89.4-90.8) 0.850 (0.764-0.936) 0.920 (0.915-0.926) 0.924 (0.891-0.958) 0.406 (0.358-0.453) 0.196 (0.152-0.241) 0.548 (0.541-0.556)
Random forest All features 0.981 (0.975-0.986) 98.5 (98.2-98.7) 0.250 (0.218-0.282) 0.991 (0.989-0.993) 0.957 (0.933-0.981) 0.309 (0.242-0.377) 0.124 (0.077-0.170) 0.271 (0.184-0.357)
Support vector classifier All features 0.981 (0.974-0.987) 98.1 (98.1-98.1) 0.250 (0.250-0.250) 0.987 (0.987-0.987) 0.908 (0.884-0.933) 0.273 (0.273-0.273) 0.070 (0.02-0.119) 0.201(0.154-0.248)
Quadratic discriminant analysis All features 0.997 (0.993-1.00) 88.9 (88.1-89.8) 0.900 (0.842-0.958) 0.910 (0.903-0.917) 0.916 (0.898-0.935) 0.407 (0.387-0.426) 0.189 (0.178-0.200) 0.539 (0.499-0.579)
XGBoost classifier All features 0.984 (0.979-0.989) 97.2 (96.9-97.5) 0.400 (0.335-0.465) 0.982 (0.979-0.985) 0.951 (0.942-0.960) 0.364 (0.270-0.459) 0.161 (0.081-0.241) 0.392 (0.319-0.465)
Rule-based classification model Selection of 4 features 1.000 (0.998-1.00) 70 (68.3-71.7) 1.000 (0.954-1.00) 0.718 (0.701-0.735) 0.859 (0.834-0.884) 0.247 (0.222-0.272) 0.085 (0.067-0.102) 0.316 (0.267-0.366)

AUPRC area under the precision-recall curve, AUROC area under the receiver-operating characteristic curve, F2 F2 score (harmonic mean of precision and recall [i.e., sensitivity]), MCC Matthews correlation coefficient.

Sensitivity analyses

We performed one-way sensitivity analyses to assess the impact of removing contamination class and keyword frequency as these features are challenging to extract from eHRs. Omitting the contamination class feature had a minimal impact on performance for both the Naïve Bayes and DNN models (Supplementary Fig. 5). For Naive Bayes, sensitivity decreased from 0.900 to 0.850 and MCC dropped from 0.475 to 0.439, while NPV, AUROC, and workload reduction values remained relatively stable. Similarly, the DNN model showed minimal changes, with sensitivity remaining at 0.850, a slight decrease in AUROC from 0.924 to 0.915, and workload reduction remaining around 90.0%. Conversely, removing the keywords feature significantly impacted the Naïve Bayes model, reducing sensitivity from 0.900 to 0.800 and decreasing the MCC from 0.475 to 0.421 (Supplementary Fig. 5). However, the DNN model demonstrated resilience to this change, with workload reduction increasing slightly to 90.6% and MCC improving to 0.418, while maintaining a stable sensitivity at 0.850 and a marginal improvement in AUROC.

Feature importance

For the Naïve Bayes model (Fig. 1), SHAP values showed that the overall number of cultures and number of positive cultures were the most significant features, emphasising the importance of microbiological data. Other influential features included antibiotic days, reoperations, and the number of infectious disease consultations. Demographic features and contamination class had less impact. In contrast, the DNN model (Fig. 2) relied mostly on contamination class, sex and presence of an implant, suggesting a focus on contextual and patient-specific characteristics. It also assigned moderate weight to antibiotic days and the number of C-reactive protein samples, but relied less on culture data.

Fig. 1. SHAP values of the Naïve Bayes model (Bernoulli assumption).

Fig. 1

The barplot (left), shows the average magnitude of SHAP values, ranking features by their overall influence. The beeswarm plot (right), provides a detailed view of each feature’s SHAP value distribution, indicating how high or low values affect the model’s output and highlighting feature interactions.

Fig. 2. SHAP values of the dense neural network model.

Fig. 2

The barplot (left), shows the average magnitude of SHAP values, ranking features by their overall influence. The beeswarm plot (right), provides a detailed view of each feature’s SHAP value distribution, indicating how high or low values affect the model’s output and highlighting feature interactions.

In the sensitivity analysis with noise probes, all three random variables (Gaussian, uniform, and Bernoulli) received minimal importance across models, confirming that the algorithms did not assign predictive value to unrelated input features (Supplementary Fig. 6 and 7). This supports the robustness of the models and suggests a low risk of overfitting to noise.

Discussion

In this study, we evaluated the performance of ML models and a rule-based classification algorithm for SSI surveillance within a high-quality prospective cohort. Our findings showed that while both approaches substantially reduce manual workload, the rule-based algorithm demonstrated a higher sensitivity. Conversely, ML models outperformed the rule-based model in terms of AUROC, AUPRC, and workload reduction, but at the cost of a higher FNR. This trade-off between accuracy and operational efficiency will need careful consideration when selecting and implementing surveillance methodologies.

There is a growing interest in digitising HAI surveillance, particularly SSIs, as events are rare and the number of records needed to screen to identify one infection is large5,6,19. A systematic review of AI models used for HAI surveillance conducted in 202016 showed that SSI was the most frequently studied infection type11,13,18,2224,27,28. The primary motivation is to reduce the workload for IPC professionals, allowing them to allocate more time for preventive actions rather than retrospective case identification. However, there is significant heterogeneity in the algorithms used for SSI detection5, ranging from rule-based classification models20 to administrative data-driven17 and ML models16. A systematic review by Streefkerk et al.19 highlighted a substantial variability in the sensitivity and specificity of electronically-assisted surveillance systems for SSI detection, with sensitivity ranging from 0.02 to 1.0 and specificity from 0.59 to 1.0. The authors found that semi-automated approaches where algorithms flag high-risk cases for manual validation were the most commonly studied approaches. Fully automated detection, while desirable, remains rare and challenging as many models rely on structured data (e.g., ICD-10 codes, microbiological results), which may not capture all clinically relevant SSIs.

Our findings are consistent with other studies. A systematic review that evaluated a wide range of ML models showed that performance was generally good, with AUROC ranging from 0.719 to 0.963. For studies that reported sensitivity, values ranged from 61.4% to 93.5%, depending on the model type, although external validation was not reported. Since then, one US group has published external validation of a LASSO-penalized logistic regression model showing conserved performance (AUROC 0.905 for organ-space SSI)25. Although ML models outperformed traditional surveillance methods in many cases, the authors observed that rule-based models still played an important role in automated SSI detection due to their interpretability and ease of implementation, which was confirmed by our study. A more recent study using a large national dataset showed that random forest and DNN had high AUROC, but sensitivity was significantly increased when integrating clinical free text with data29. Interestingly, the authors found that some misclassifications were due to an imperfect “gold standard”, suggesting that models may identify otherwise missed SSIs.

In a similar approach to our work, Cho et al.26 developed and evaluated both ML and rule-based classification algorithms for SSI surveillance in colorectal surgery. In addition, they compared standalone ML models and a hybrid approach combining ML with a rule-based algorithm. Their findings mirrored some of our observations. While the best-performing ML model (neural networks with recursive feature elimination) demonstrated a higher AUROC (0.963 vs. 0.857) and greater workload reduction (78% vs. 72%) compared to the rule-based approach, sensitivity remained similar (0.951). The hybrid model further improved workload reduction (84%), while maintaining sensitivity at 0.940. These findings underscore an important consideration in automated SSI surveillance in that ML models may improve overall performance metrics, but they do not inherently replace rule-based algorithms, which often provide greater interpretability in clinical settings. Our study found that the rule-based classification exhibited a lower FNR, meaning that it was more reliable for ensuring no missed infections, whereas ML models offered a superior AUROC and workload reduction. The trade-off between sensitivity, efficiency and interpretability highlights the need for a balanced approach, rather than an exclusive reliance on ML models. In this context, ML models should not be viewed as a universal solution, but as one component of a broader strategy that integrates rule-based logic, expert validation, and data-driven decision-making to optimise automated surveillance systems.

Our study focused on the detection of SSI for surveillance purposes as opposed to the prediction of SSI for pre-surgical risk estimation. A recent systematic review30 found that detection models had a significantly higher sensitivity (pooled sensitivity, 0.89; 95% CI, 0.81–0.94) compared to predictive models (pooled sensitivity, 0.70; 95% CI, 0.61–0.78), although AUROCs were similar.

We used unstructured text data in the form of the occurrence of infection-related textual keywords when developing both our rule-based and ML models. Feature importance analysis showed that these textual keywords were critical predictors for the Naïve Bayes classifier, but not for DNN. This echoes findings from a systematic review30 showing that models using only structured data without the inclusion of textual data decreased their pooled sensitivity from 0.83 (95% CI, 0.78-0.87) to 0.56 (95% CI, 0.43-0.69). This highlights the need to pool results from different studies based on model type. It was unclear whether the studies included natural language processing, which we did not perform. A recent study suggested that there may only be a limited added value of natural language processing in terms of workload reduction, with a concomitant decrease in sensitivity31.

Despite the promising performance of models developed, substantial barriers remain for their broad implementation, notably the selection of appropriate denominator data. In a nationwide survey of the state of digitalisation of SSI surveillance in Germany, only 26% (216/821) of responding surgical departments reported the availability of digital denominator data32, with the main reason being a lack of IT support. Interestingly, data availability in this high-income country was not considered a barrier.

Our study has several strengths. First, we used a prospectively collected cohort with high-quality data (including post-discharge surveillance). Second, most input variables selected for the models were not overly complicated to obtain. In addition, we performed sensitivity analyses where we trained models without certain variables that may be more difficult to obtain; even without these variables, the models retained good performance.

The study has also some limitations. First, it was exclusively focused on deep and organ/space SSIs. This decision represents a strategic trade-off between exhaustive surveillance and a broader, less intense surveillance approach, thereby enabling more time for preventive actions. While this may limit the applicability of our findings to surveillance in our national network, we may be able to use these models internally for surveillance of other surgical categories. While our models demonstrated strong performance on the validation set, we acknowledge the need for external validation on temporally and geographically distinct datasets. Cross-validation and independent validation set performance were broadly consistent, and standard deviations across cross-validation folds were small, suggesting limited overfitting. In a sensitivity analysis, we also introduced random noise probes which were assigned negligible importance by all models, further supporting the robustness of our approach. We used SHAP values to assess model interpretability, and although SHAP’s additive formulation can be sensitive to feature collinearity and background selection, and its computation is more demanding for deep networks, it remains a widely adopted, state-of-the-art approach for structured-data models. Future work will further extend our interpretability analyses by integrating complementary explainable artificial intelligence (XAI) techniques to provide an even richer explanatory perspective. Finally, our models were trained on tabular data aggregated per surgical procedure, which may not capture fine-grained temporal dynamics. Recent studies have demonstrated the potential of transformer-based architectures to model temporal patterns in eHR data for outcome prediction33.

In conclusion, ML models and rule-based classification algorithms offer complementary advantages in SSI surveillance. ML models improve AUROC and workload reduction and can be used for sentinel surveillance purposes, but rule-based models are more sensitive and better suited for regulatory surveillance. By identifying high-risk procedures with high sensitivity while excluding a large proportion from manual review, semi-automated surveillance systems could free up time for IPC teams to focus on prevention. Future research should focus on external validation, hybrid approaches integrating rule-based and ML models, exploring sequence-based models such as transformers as time-resolved datasets become available, and assessing the feasibility of real-world implementation.

Methods

Design, setting, and study population

This prospective cohort study was performed between 1 October, 2016 and 30 September, 2022 at Geneva University Hospitals (Geneva, Switzerland) as part of a national quality improvement programme (Swissnoso, Swiss National Center for Infection Prevention). We included adult patients ( ≥18years old) undergoing elective or urgent surgery in the following surgical categories: cardiac; coronary artery bypass grafting; colorectal surgery; laminectomy; and spinal fusion. SSI surveillance was performed according to Swissnoso SSI surveillance system guidelines34, which follow the definitions of the United States Centers for Disease Control and Prevention35,36. Patients undergoing selected surgical procedures were prospectively included, with postoperative surveillance for 30 days, or 90 days if an implant was placed. Surveillance targeted deep incisional and organ/space SSIs, defined according to Swissnoso/CDC criteria. Deep SSIs are defined as infections involving the fascia or muscle layers and are diagnosed based on purulent drainage, wound dehiscence with clinical signs, or evidence of infection (e.g., abscess) on imaging or reoperation. Organ/space SSIs are defined as infections involving anatomical compartments (e.g., peritoneum, mediastinum, bone) and are diagnosed based on purulent drainage from a drain, a positive culture, or similar evidence of infection in the affected compartment. SSI status was determined by trained IPC professionals based on manual chart review and post-discharge telephone follow-up. Surveillance of surgical procedures involving implants was initially performed up to one year post-surgery but was shortened to 90 days after October 2021 in accordance with updated guidelines. Superficial infections were excluded as their detection introduces significant methodological and practical challenges and they are mostly detected by post-discharge surveillance37, which could potentially lead to the underdetection of SSI and misclassification if included in a (semi-)automated surveillance system. Post-discharge surveillance involved up to five telephone call attempts by IPC professionals who administered a standardised questionnaire. For study purposes, follow-up ended at 90 days for all patients undergoing surgery with implants. The clinical outcome was the occurrence of deep and organ/space SSI during the follow-up period.

The study was part of a quality improvement programme and did not require approval from an institutional review board or informed consent from participants. Data were de-identified prior to analysis.

Features

For each patient, we extracted an additional number of postoperative features from the eHR up to the end of the follow-up period including: number of days with antibiotic treatment; postoperative fever (temperature >38 °C); number of infectious disease consultations; number of bacteriological cultures (sterile fluids, biopsies, prosthetic materials, wound swabs, and aspirates from abscesses or joints); radiological examinations (diagnostic scans, targeted biopsies, and image-guided procedures); and frequency of occurrence of certain keywords in follow-up notes (e.g., “infection”, “redness”, “pus”). Supplementary Table 1 lists all input features, their encoding schemes, and preprocessing steps.

Outcome

The primary outcome was the classification performance and diagnostic accuracy of ML and rule-based models for detecting deep and organ/space SSIs in a semi-automated framework.

Statistical analysis

We split the dataset randomly into a training and a validation set using an 80/20 split. The validation set was used solely for model evaluation. The machine learning target variable was a binary outcome (y ∈ {0,1}), where y = 1 denoted the occurrence of a deep or organ/space SSI, and y = 0 indicated no SSI, as determined by IPC professionals following routine surveillance procedures. This final classification, performed after the end of the follow-up period, was used as the reference standard for model training and evaluation. All predictor features were derived from data available prior to this timepoint to avoid data leakage.

Each observation in the dataset corresponded to a unique surgical procedure, including associated perioperative and follow-up data. A schematic representation of the temporal alignment between surgery, hospital stay, and follow-up period for each procedure is shown in Supplementary Fig. 1. The train/test split was performed at the procedure level, in line with our aim to predict SSI occurrence following individual surgical interventions. All features and outcomes were specific to each procedure. The dataset was structured to minimise data sharing across procedures; only a single instance of overlapping follow-up between training and validation sets was identified. While the unique instance of overlap makes any practical data leakage negligible, dividing the data by procedure could, in theory, convey information whenever the same patient undergoes multiple operations. We acknowledge this methodological limitation and will therefore implement patient-level splits in forthcoming, larger cohorts to yield a more stringent assessment of model generalisability.

With the training set, we used five-fold cross-validation38, whilst ensuring similar proportions of SSI cases in each fold. We selected a combination of linear and non-linear models to balance interpretability and predictive performance39. Logistic regression and discriminant analysis were included for their transparency and widespread use in clinical research. Naïve Bayes was selected for its simplicity, speed, and robustness in high-dimensional data despite its strong independence assumption. Random forests and XGBoost were included for their capacity to capture non-linear relationships and complex interactions as ensemble models. The Dense Neural Network was included for its capacity to learn deeper representations from non-linear and complex relationships, due to its multiple fully connected layers with non-linear activation functions. Model implementation was performed using the Scikit-learn, TensorFlow, and XGBoost libraries in Python.

Hyperparameter tuning was performed using a grid search approach within the cross-validation framework, selecting the combination that maximised the area under the receiver operating characteristic curve (AUROC) and negative predictive value (NPV). All numerical predictors were then rescaled with minimum-maximum normalisation39. We favoured this operation over z-standardisation because many variables are naturally bounded counts or proportions, minimum-maximum scaling preserves the sparsity of binary indicators, and preliminary tests showed slightly faster neural-network convergence with identical AUROC. The final hyperparameters are displayed in Supplementary Table 2.

In the second stage, final models were trained on the entire training set using the optimised hyperparameters and minimum-maximum normalisation. Model performance was evaluated on the independent validation set using minimum-maximum normalisation parameters derived from the training set to prevent data leakage.

We assessed the impact of omitting variables that may be difficult to extract from eHRs. Specifically, we evaluated model performance when excluding contamination class and keyword frequency features.

To benchmark ML models, we developed a rule-based classification model for SSI detection based on established clinical indicators of SSI routinely documented in the eHR (Supplementary Fig. 2). Patients were classified as having an SSI if they met any of the following criteria: ≥ 5 days of postoperative antibiotics; ≥ 1 readmission; ≥ 2 positive cultures; or ≥1 infectious disease consultation. The predictive performance of the model was compared with ML models to quantify improvements in prediction accuracy and workload reduction.

Assessment of model performance

Model performance was assessed using essential indicators of diagnostic accuracy: AUROC, area under the precision-recall curve (AUPRC), NPV, false-negative rate (FNR), and workload reduction. The latter was defined as the proportion of patients classified as not having SSI and therefore not requiring manual chart review under a semi-automated surveillance framework. For the primary analysis, we applied a default classification threshold of 0.5 for ML models, at which sensitivity, NPV and workload reduction were calculated. To evaluate threshold-dependent performance, we assessed model metrics, including sensitivity, specificity, negative predictive value (NPV), workload reduction, and F2 score, across a range of classification thresholds (from 0.0 to 1.0 in increments of 0.1). Performance metrics were averaged across cross-validation folds for each ML model and reported with 95% confidence intervals (CIs). The best-performing linear and non-linear ML models were selected based on sensitivity, NPV and workload reduction.

Model interpretability was assessed using SHapley Additive exPlanations (SHAP)40, which quantifies the contribution of each characteristic (feature) to model predictions by considering all possible combinations in order to calculate their marginal impact on output. This ensures a fair attribution of prediction influence to each feature, providing both global feature importance insights and local interpretability for individual predictions. This approach is particularly beneficial for complex models, offering transparency in decision-making and feature relevance.

Model performance was assessed on a held-out validation set, and we compared cross-validation and validation performance to evaluate risk of overfitting. In a sensitivity analysis we introduced three random noise variables generated from Gaussian (N[0,1]), uniform (U[0,1]), and Bernoulli distributions (B[1,0.1]), into the training dataset. These probes were not correlated with the outcome and served as negative controls. We evaluated their relative importance across models using SHAP values to assess whether the models inappropriately prioritised irrelevant features.

Statistical analyses were performed using Python version 3.9. The following packages were used for ML: TensorFlow (version 2.8.4); XGBoost (version 1.5.0); and sklearn (version 1.1.1).

Supplementary information

Acknowledgements

This work was supported by a grant from the Geneva University Hospitals Research and Development Programme. The funder had no role in the design and analysis of the study. This project would not have been possible without the work of the IPC nurses and data managers that have been conducting SSI surveillance in our hospital for many years: Henrique Alves, Catherine Bandiera-Clerc, Nadia Colaizzi, Marlène Fraccaro Hamoneau, Monica Perez, Valérie Sauvan, and Isabelle Soulake. The authors would also like to thank Emmanuel Durand (IT Department, Geneva University Hospitals) for his support, as well as Dr. Jonathan Aryeh Sobel (PhD) and Prof. Douglas Teodoro for providing input. We would like to thank Ms. Rosemary Sudan for editing the manuscript.

Author contributions

A.A., E.C., D.T., D.B., N.B., G.C., S.H. and M.A. all contributed substantially to this work in accordance with the ICMJE criteria. Specifically, all authors were involved in: (1) the conception and design of the study, acquisition of data, or analysis and interpretation of data; (2) drafting the article or revising it critically for important intellectual content; and (3) final approval of the version to be submitted. CRediT author statement: A.A.: investigation, data curation, writing—original draft; E.C.: methodology, formal analysis, visualisation, writing—review and editing; D.T.: conceptualisation, validation, data curation; DB: investigation, formal analysis, writing—review and editing; N.B.: conceptualisation, methodology, supervision, writing—review and editing; G.C.: writing—review and editing S.H.: conceptualisation, supervision, writing—review and editing; M.A.: conceptualisation, supervision, project administration, writing—original draft, writing—review and editing.

Data availability

The code and a synthetic dataset can be found on GitHub (https://github.com/etiennechlt/ssi-detect). The datasets used and analysed during this study are not publicly available due to patient confidentiality and institutional regulations, but may be available from the corresponding author upon reasonable request.

Code availability

The code and a synthetic dataset can be found on GitHub (https://github.com/etiennechlt/ssi-detect). The datasets used and analysed during this study are not publicly available due to patient confidentiality and institutional regulations, but may be available from the corresponding author upon reasonable request.

Competing interests

The authors declare no competing interests.

Declaration of generative AI in scientific writing

During the preparation of this work the authors used ChatGPT (April 2025 version, OpenAI) in order to assist with writing, editing, and improving clarity. After using this tool/service, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Change history

11/21/2025

A Correction to this paper has been published: 10.1038/s41746-025-02153-5

Supplementary information

The online version contains supplementary material available at 10.1038/s41746-025-01989-1.

References

  • 1.Shambhu, S. et al. The burden of health care utilization, cost, and mortality associated with select surgical site infections. Jt Comm. J. Qual. Patient Saf.50, 857–866 (2024). [DOI] [PubMed] [Google Scholar]
  • 2.Storr, J. et al. Core components for effective infection prevention and control programmes: new WHO evidence-based recommendations. Antimicrob. Resist Infect. Control6, 6 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Abbas, M. et al. Impact of participation in a surgical site infection surveillance network: results from a large international cohort study. J. Hosp. Infect.102, 267–276 (2019). [DOI] [PubMed] [Google Scholar]
  • 4.Birgand, G. et al. Agreement among healthcare professionals in ten European countries in diagnosing case-vignettes of surgical-site infections. PLoS ONE8, e68618 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Verberk, J. D. M. et al. Automated surveillance systems for healthcare-associated infections: results from a European survey and experiences from real-life utilization. J. Hosp. Infect.122, 35–43 (2022). [DOI] [PubMed] [Google Scholar]
  • 6.van Mourik, M. S. M. et al. PRAISE: providing a roadmap for automated infection surveillance in Europe. Clin. Microbiol Infect.27, S3–S19 (2021). [DOI] [PubMed] [Google Scholar]
  • 7.Behnke, M. et al. Information technology aspects of large-scale implementation of automated surveillance of healthcare-associated infections. Clin. Microbiol. Infect.27, S29–S39 (2021). [DOI] [PubMed] [Google Scholar]
  • 8.van Rooden, S. M. et al. Governance aspects of large-scale implementation of automated surveillance of healthcare-associated infections. Clin. Microbiol. Infect.27, S20–S28 (2021). [DOI] [PubMed] [Google Scholar]
  • 9.Leclere, B. et al. Matching bacteriological and medico-administrative databases is efficient for a computer-enhanced surveillance of surgical site infections: retrospective analysis of 4,400 surgical procedures in a French university hospital. Infect. Control Hosp. Epidemiol.35, 1330–1335 (2014). [DOI] [PubMed] [Google Scholar]
  • 10.Rennert-May, E. et al. Validating administrative data to identify complex surgical site infections following cardiac implantable electronic device implantation: a comparison of traditional methods and machine learning. Antimicrob. Resist. Infect. Control11, 138 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Sanger, P. C. et al. A prognostic model of surgical site infection using daily clinical wound assessment. J. Am. Coll. Surg.223, 259–70.e2 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Tunthanathip, T. et al. Machine learning applications for the prediction of surgical site infection in neurological operations. Neurosurg. Focus47, E7 (2019). [DOI] [PubMed] [Google Scholar]
  • 13.Kuo, P. J. et al. Artificial neural network approach to predict surgical site infection after free-flap reconstruction in patients receiving surgery for head and neck cancer. Oncotarget9, 13768–13782 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Pindyck, T. et al. Validation of an electronic tool for flagging surgical site infections based on clinical practice patterns for triaging surveillance: operational successes and barriers. Am. J. Infect. Control46, 186–190 (2018). [DOI] [PubMed] [Google Scholar]
  • 15.Roth, J. A., Battegay, M., Juchler, F., Vogt, J. E. & Widmer, A. F. Introduction to machine learning in digital healthcare epidemiology. Infect. Control Hosp. Epidemiol.39, 1457–1462 (2018). [DOI] [PubMed] [Google Scholar]
  • 16.Scardoni, A., Balzarini, F., Signorelli, C., Cabitza, F. & Odone, A. Artificial intelligence-based tools to control healthcare associated infections: a systematic review of the literature. J. Infect. Public Health13, 1061–1077 (2020). [DOI] [PubMed] [Google Scholar]
  • 17.Bauer, J. M., Welling, S. E. & Bettinger, B. Can we automate spine fusion surgical site infection data capture?. Spine Deform11, 329–333 (2023). [DOI] [PubMed] [Google Scholar]
  • 18.Hu, Z. et al. Automated detection of postoperative surgical site infections using supervised methods with electronic health record data. Stud. Health Technol. Inf.216, 706–710 (2015). [Google Scholar]
  • 19.Streefkerk, H. R. A., Verkooijen, R. P., Bramer, W. M. & Verbrugh, H. A. Electronically assisted surveillance systems of healthcare-associated infections: a systematic review. Euro Surveill25, 1900321 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Sips, M. E., Bonten, M. J. M. & van Mourik, M. S. M. Semiautomated surveillance of deep surgical site infections after primary total hip or knee arthroplasty. Infect. Control Hosp. Epidemiol.38, 732–735 (2017). [DOI] [PubMed] [Google Scholar]
  • 21.Sips M. E., Bonten M. J. M., van Mourik M. S. M. Semi-automated surveillance of deep surgical site infections after cardiothoracic surgery. In Conference abstract OS0512. 27th European Conference on Clinical Microbiology and Infectious Diseases (ECCMID), (2017).
  • 22.Campillo-Gimenez, B., Garcelon, N., Jarno, P., Chapplain, J. M. & Cuggia, M. Full-text automated detection of surgical site infections secondary to neurosurgery in Rennes, France. Stud. Health Technol. Inf.192, 572–575 (2013). [Google Scholar]
  • 23.Sohn, S. et al. Detection of clinically important colorectal surgical site infection using Bayesian network. J. Surg. Res.209, 168–173 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Weller, G. B., Lovely, J., Larson, D. W., Earnshaw, B. A. & Huebner, M. Leveraging electronic health records for predictive modeling of post-surgical complications. Stat. Methods Med. Res.27, 3271–3285 (2018). [DOI] [PubMed] [Google Scholar]
  • 25.Zhu, Y. et al. Applying machine learning across sites: external validation of a surgical site infection detection algorithm. J. Am. Coll. Surg.232, 963–71.e1 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Cho, S. Y. et al. Development of machine learning models for the surveillance of colon surgical site infections. J. Hosp. Infect.146, 224–231 (2024). [DOI] [PubMed] [Google Scholar]
  • 27.Ke, C. et al. Prognostics of surgical site infections using dynamic health data. J. Biomed. Inf.65, 22–33 (2017). [Google Scholar]
  • 28.Soguero-Ruiz, C. et al. Data-driven temporal prediction of surgical site infection. AMIA Annu Symp. Proc.2015, 1164–1173 (2015). [PMC free article] [PubMed] [Google Scholar]
  • 29.Chakraborty, A. et al. Development and evaluation of machine learning models for the identification of surgical site infection in electronic health records. Surg. Infect. (Larchmt)26, 474–481 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Wu, G. et al. Performance of machine learning algorithms for surgical site infection case detection and prediction: a systematic review and meta-analysis. Ann. Med. Surg. (Lond.)84, 104956 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Verberk, J. D. M. et al. The augmented value of using clinical notes in semi-automated surveillance of deep surgical site infections after colorectal surgery. Antimicrob. Resist. Infect. Control12, 117 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Aghdassi, S. J. S. et al. Surgical site infection surveillance in German hospitals: a national survey to determine the status quo of digitalization. Antimicrob. Resist. Infect. Control12, 49 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Antikainen, E. et al. Transformers for cardiac patient mortality risk prediction from heterogeneous electronic health records. Sci. Rep.13, 3517 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Troillet, N., Aghayev, E., Eisenring, M. C. & Widmer, A. Swissnoso. First results of the swiss national surgical site infection surveillance program: who seeks shall find. Infect. Control Hosp. Epidemiol38, 697–704 (2017). [DOI] [PubMed] [Google Scholar]
  • 35.Horan, T. C., Gaynes, R. P., Martone, W. J. & Jarvis, W. R. Emori TG. CDC definitions of nosocomial surgical site infections, 1992: a modification of CDC definitions of surgical wound infections. Am. J. Infect. Control20, 271–274 (1992). [DOI] [PubMed] [Google Scholar]
  • 36.Kuster, S. P., Eisenring, M. C., Sax, H. & Troillet, N. Swissnoso. Structure, process, and outcome quality of surgical site infection surveillance in Switzerland. Infect. Control Hosp. Epidemiol.38, 1172–1181 (2017). [DOI] [PubMed] [Google Scholar]
  • 37.Ming, D. Y., Chen, L. F., Miller, B. A. & Anderson, D. J. The impact of depth of infection and postdischarge surveillance on rate of surgical-site infections in a network of community hospitals. Infect. Control Hosp. Epidemiol.33, 276–282 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Hastie T., Tibshirani R., Friedman J. The Elements of Statistical Learning 2nd edn (Springer, 2009).
  • 39.Ghori, K. M. et al. Performance analysis of different types of machine learning classifiers for non-technical loss detection. IEEE Access8, 16033–16048 (2020). [Google Scholar]
  • 40.Lundberg S. M., Lee S.-I. A unified approach to interpreting model predictions. In Proc. 31st International Conference on Neural Information Processing Systems 4768–4777 (ACM, 2017).

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

The code and a synthetic dataset can be found on GitHub (https://github.com/etiennechlt/ssi-detect). The datasets used and analysed during this study are not publicly available due to patient confidentiality and institutional regulations, but may be available from the corresponding author upon reasonable request.

The code and a synthetic dataset can be found on GitHub (https://github.com/etiennechlt/ssi-detect). The datasets used and analysed during this study are not publicly available due to patient confidentiality and institutional regulations, but may be available from the corresponding author upon reasonable request.


Articles from NPJ Digital Medicine are provided here courtesy of Nature Publishing Group

RESOURCES