Abstract
Background:
Gangrenous cholecystitis (GC) is a severe complication of acute cholecystitis that requires prompt surgical intervention. However, its preoperative diagnosis remains challenging due to the limitations of traditional imaging techniques. This study aims to develop a self-supervised learning (SSL) model to preoperatively identify GC using both plain and contrast-enhanced CT images.
Methods:
This was a retrospective, multicenter cohort study conducted from January 2021 to September 2024. A total of 7368 CT images from 1228 patients (training set: 921 patients; independent validation set: 307 patients) were retrospectively analyzed. We developed an SSL model using seResNet-50 framework, trained on unenhanced and contrast-enhanced CT images, to predict the presence of GC preoperatively. The model leverages unlabeled data to pretrain the network, and is then fine-tuned on a limited set of annotated images. After feature extraction and selection, we developed three models for predicting GC, including a fusion model, an enhanced CT (ECT) model, and a non-enhanced CT (NECT) model. Performance was assessed using accuracy, sensitivity, specificity, and area under the receiver operating characteristic curve (AUC).
Results:
The fusion model demonstrated robust performance in preoperative GC prediction. In the training set, the fusion model achieved an AUC of 0.965, a sensitivity of 88%, and a specificity of 95% in detecting GC. In Validation Set I, the fusion model had an AUC of 0.879, surpassing the enhanced and non-enhanced models (AUCs: 0.791, 0.756, respectively). Similarly, in Validation Set II, the fusion model achieved an AUC of 0.887, significantly better than the ECT and NECT models (AUCs: 0.810, 0.730). The model also provided interpretable analyses by detecting subtle features of gangrenous changes in the CT images, facilitating clinical decision-making.
Conclusion:
We developed a fusion SSL model for preoperative prediction of gangrenous cholecystitis using both unenhanced and contrast-enhanced CT scans. The model’s high diagnostic performance suggests its clinical applicability in improving early diagnosis and timely risk stratification of GC.
Keywords: biliary surgery, cholecystectomy, gangrenous cholecystitis, retrospective cohort study, self-supervised learning
Introduction
Gangrenous cholecystitis (GC) is a severe complication of acute cholecystitis, characterized by gallbladder necrosis and increased risk of mortality. Early identification of GC is crucial for determining appropriate surgery and improving patient outcomes[1]. However, distinguishing between uncomplicated acute cholecystitis and GC remains challenging because clinical symptoms often overlap. Current diagnostic tools, including laboratory tests and ultrasonography, show limited sensitivity and specificity in detecting GC, leading to delayed diagnosis and suboptimal treatment[2,3]. Recent advancements in radiological imaging have improved diagnostic capabilities, particularly the use of enhanced CT (ECT), which offers superior visualization of gallbladder wall abnormalities, necrosis, and surrounding fat stranding. However, ECT is limited by radiologists’ subjective interpretation, especially in cases with early or subtle signs of gangrenous changes[4]. That highlights the need for more reliable imaging biomarkers to predict the presence of GC, particularly in the preoperative setting.
Radiomics and deep learning have provided promising ways for accurate prediction of GC[5,6]. Radiomics involves the extraction of quantitative features from medical images, which can be analyzed to identify disease patterns that may not be visible to the human eye. The integration of deep learning techniques, particularly convolutional neural networks (CNNs), has further improved the accuracy of image classification tasks by enabling automated, high-dimensional feature extraction from medical images. Self-supervised learning (SSL), which trains models without labeled data, has shown promise in reducing the reliance on large annotated datasets while enhancing model generalizability[7]. Among various models, the seResNet-50 architecture has proven effective in capturing spatial and channel-wise feature dependencies, improving model performance in medical imaging tasks[8]. Recent studies have shown the feasibility of applying CT-based deep learning to predict acute cholecystitis. These models have higher diagnostic performance compared to traditional methods, offering potential improvements in preoperative risk stratification and clinical decision-making[9]. However, further validation and optimization are needed to ensure their clinical applicability across different patient populations.
This study aims to develop a seResNet-50 model based on self-supervised learning using preoperative plain and enhanced CT scans to predict the presence of GC, ultimately guiding more timely and accurate interventions. The model has the potential to improve preoperative risk stratification and clinical decision-making. This cohort study has been reported in line with the STROCSS guidelines[10]. Additionally, it has been reported in line with the TITAN criteria (Supplemental Digital Content 1, available at: http://links.lww.com/JS9/E791)[11].
Methods and materials
Study design and patients inclusion
This was a retrospective, multicenter cohort study conducted from January 2021 to September 2024 that was specifically was designed to predict gangrenous cholecystitis using a combination of enhanced and plain CT images. A total of 1228 patients who had CT imaging of the abdomen were retrospectively reviewed. Of these, 921 patients were included in the training set. The remaining 307 patients were divided into two independent validation sets: Validation Set I (n = 143) and Validation Set II (n = 164) from two distinct medical centers. This stratification was performed to assess the model’s generalizability across heterogeneous CT acquisition protocols and patient populations, thereby enhancing the external validity of our models. Inclusion criteria were: (1) clinical indication for CT examination due to suspected inflammation or other gallbladder diseases; (2) availability of both enhanced and plain CT scans; (3) confirmed diagnosis of gangrenous cholecystitis by histopathology. Patients with incomplete clinical data, poor-quality imaging (e.g., motion artifacts, low contrast resolution), or previous gallbladder surgery were initially excluded to ensure high-quality annotations and feature extraction. However, to assess the model’s robustness, we conducted a sensitivity analysis by reintroducing 64 of these previously excluded cases (42 with suboptimal imaging quality and 22 with incomplete metadata) as an independent validation cohort, the detailed patient recruitment process is shown in Figure 1. The study received ethical approval from the institutional review boards of the three participating centers. Informed consent was waived due to the retrospective nature of the study and anonymized data handling. The study was registered on https://www.researchregistry.com/browse-the-registry/#home and conducted in accordance with the Declaration of Helsinki. The reporting of the study followed the STROCSS criteria (Supplemental Digital Content 2, available at: http://links.lww.com/JS9/E791) and the STARD criteria[10,12]. Figure 2 illustrates the overall research design.
HIGHLIGHTS
This study developed a self-supervised seResNet-50 model integrating plain and contrast-enhanced CT for preoperative GC identification.
The model offers interpretable analyses via Shapley analysis and Grad-CAM visualization to support clinical decisions.
The fusion model outperformed single-modality models, achieving AUCs of 0.965 (training) and 0.879–0.887 (validation).
Self-supervised learning reduces reliance on labeled data, while multi-modal fusion enhances diagnostic accuracy and clinical utility.
Figure 1.
The detail flowchart of patient recruitment process.
Figure 2.
The overall workflow of this study.
Histopathologic assessment and radiological evaluation
Gallbladder samples obtained from cholecystectomy were evaluated by experienced pathologists to confirm the presence of gangrenous cholecystitis, including the degree of inflammatory cell infiltration (e.g., neutrophils, lymphocytes), tissue damage, and fibrosis. All patients were categorized as having either gangrenous or non-gangrenous cholecystitis. Radiologic evaluation was performed by two independent radiologists who were blinded to the histopathologic results. Both enhanced and plain CT images were used to identify features related to gallbladder inflammation, including wall thickening, pericholecystic fluid, and surrounding edema. The radiologists performed a consensus review to ensure the accuracy and consistency of the CT image interpretation. To further contextualize model performance, we assessed the diagnostic metrics of the two radiologists prior to consensus. Images were excluded if they showed severe motion artifacts, incomplete gallbladder visualization, or slice thickness >5 mm. Subjectively, two board-certified radiologists independently rated images on a 3-point Likert scale (1 = non-diagnostic, 2 = suboptimal but acceptable, 3 = diagnostic). Only images rated 2 or 3 by both were included. Discrepancies were resolved by consensus.
Image acquisition and preprocessing
All CT images were acquired using a 64-slice multidetector CT scanner (Siemens Somatom Definition). The imaging protocol involved the acquisition of both plain and contrast-enhanced CT scans. For enhanced images, the imaging parameters were as follows: slice thickness of 1 mm, pitch of 1.2, and tube voltage of 120 kVp with a tube current of 300 mA. For plain images, slice thickness of 2 mm, with a pitch of 1.5 and tube voltage of 120 kVp, using standard non-contrast CT acquisition protocols. CT images were reconstructed using standard settings (B30f filter for soft tissue). All CT images were preprocessed for uniformity and quality enhancement. Each image was resampled to isotropic voxels (1 mm3), and their intensity values were normalized to a range of 0 to 1. For image denoising, a non-local means denoising algorithm was applied to reduce noise without blurring edges. To ensure compatibility with the deep learning model, all images were resized to a fixed per-slice resolution of 224 × 224 pixels for input into the self-supervised learning model.
Region-of-interest slide segmentation and feature extraction
The slice containing the whole gallbladder was manually selected on both enhanced and plain CT scans by an experienced radiologist. The selection process was performed on the axial slices, covering the full extent of the gallbladder. The boundaries of the gallbladder were defined as the region containing the lumen, wall, and surrounding tissues. Deep learning-based feature extraction was performed using a self-supervised learning model, seResNet-50. This model was pretrained on a large dataset of natural images (Image-Net) and adapted to our specific dataset. The seResNet-50 architecture consists of 50 layers and incorporates residual connections to preserve feature information across layers. The network was fine-tuned on our dataset to learn robust representations of the gallbladder tissue, which were subsequently used as feature vectors for classification tasks. Additionally, we applied data augmentation techniques including rotation and flipping, to increase the robustness of the model and prevent overfitting. Imaging features were extracted from the final convolutional layer and used as input features for machine learning models. The extracted features were normalized to eliminate potential biases due to varying imaging conditions. Inference time was measured at an average of <3 seconds per case on standard GPU hardware (NVIDIA RTX 3080Ti), indicating sufficient computational efficiency for clinical deployment without the need for pruning or additional model compression techniques.
Feature selection and model construction
To reduce the dimensionality and select the most relevant features for modeling, we employed two methods: the U-test statistical analysis and Boruta algorithm. We first applied Mann–Whitney U-test to further assess the statistical significance (P < 0.05) of each feature with gangrenous cholecystitis. Subsequently, Boruta, a feature selection method based on random forests, was used to retain features that are important for the classification task. The features with the highest predictive power were selected for model development (Supplemental Digital Content SFigure 1, available at: http://links.lww.com/JS9/E791). The final feature sets (ECT features, NECT features and their fusion) were input into a multi-layer perceptron neural network, which served as the primary classification model. The network consisted of three fully connected layers with 128, 64, and 32 neurons, respectively, and was activated using the ReLU function. A dropout layer was added after each hidden layer to prevent overfitting, with a dropout rate of 0.5. The output layer utilized a sigmoid activation function to predict the binary class labels (inflammation present vs. no inflammation). Early stopping was employed to prevent overfitting, and the model was evaluated using cross-validation within the training set. Three models were constructed to predict gangrenous cholecystitis: the ECT model, the NECT model, and a fusion model. The fusion model integrates both imaging features (from plain and enhanced CT) and key laboratory parameters (including BMI, blood cell count, neutrophil count, bilirubin levels) to improve model accuracy.
Model validation and clinical application
To assess the model’s performance, we applied a series of validation steps. Internal validation was performed using 10-fold cross-validation on the training set, and external validation was conducted on a separate validation cohort of 307 patients. Model performance was evaluated using standard metrics, including accuracy, sensitivity, specificity, and the area under the receiver operating characteristic curve (AUC-ROC). In addition, the Generalized Additive Model for Interpretation of Classifications (GRAM-CAM) and Shapley analysis were used for model interpretability, generating heatmaps that highlight the regions of the CT images contributing most to the model’s predictions. Grad-CAM generated heatmaps overlaid on the original CT images, showing the areas of the gallbladder that were most indicative of inflammation. This visualization allowed for a better understanding of the model’s decision-making process and facilitated clinical integration.
Statistical analysis
Statistical analysis was conducted using R software (version 4.0.2). The continuous data were represented using medians and interquartile ranges (IQRs), while categorical data were depicted as percentages. We employed the Fisher’s exact test or Pearson χ2 test alongside the Kruskal–Wallis test to contrast patient attributes between training and external validation cohorts. The Mann–Whitney U-test was employed to compare continuous variables between the two groups (inflammation present vs. absent), while categorical variables were compared using the chi-squared test. The performance of the models was evaluated using the confusion matrix, and the difference in AUC between models was tested using the DeLong test. A P value of less than 0.05 was considered statistically significant.
Results
Patient characteristics
A total of 921 patients were included in the training set, with 143 patients in the validation set I and 164 patients in the validation set II (Table 1). The median age in the training set was 53 years (IQR: 40–61), and the median ages for the validation set I and II were 57 years (IQR: 28–72) and 52 years (IQR: 42–59), respectively, with no significant differences across these sets (P = 0.365). The gender distribution was similar across the sets, with males representing 43% in the training set, 38% in the validation set I, and 49% in the validation set II (P = 0.148). Regarding body mass index (BMI), 59% of patients in the training set had a BMI ≥ 24, compared to 41% in both the validation sets (P < 0.001).
Table 1.
Comparison of baseline characteristics among patients in different sets
| Characteristics | Training set (n = 921) | Validation Set I (n = 143) | Validation Set II (n = 164) | P value |
|---|---|---|---|---|
| Age (y)a | 53 (40, 61) | 57 (28, 72) | 52 (42, 59) | 0.365c |
| Sex, n (%) | 0.148b | |||
| Male | 396 (43%) | 54 (38%) | 80 (49%) | |
| Female | 525 (57%) | 89 (62%) | 84 (51%) | |
| BMI, n (%) | <0.001b | |||
| <24 | 381 (41%) | 85 (59%) | 97 (59%) | |
| ≥24 | 540 (59%) | 58 (41%) | 67 (41%) | |
| WBC, n (%) | <0.001b | |||
| ≤9.5 | 687 (75%) | 111 (78%) | 146 (89%) | |
| >9.5 | 234 (25%) | 32 (22%) | 18 (11%) | |
| NEUT, n (%) | <0.001b | |||
| ≤6.3 | 616 (67%) | 110 (77%) | 148 (90%) | |
| >6.3 | 305 (33%) | 33 (23%) | 16 (10%) | |
| PLT, n (%) | <0.001b | |||
| ≤350 | 876 (95%) | 121 (85%) | 149 (91%) | |
| >350 | 45 (5%) | 22 (15%) | 15 (9%) | |
| ALT, n (%) | <0.001b | |||
| ≤12 | 82 (9%) | 18 (13%) | 37 (23%) | |
| >12 | 839 (91%) | 125 (87%) | 127 (77%) | |
| AST, n (%) | <0.001b | |||
| ≤19 | 237 (26%) | 45 (31%) | 83 (51%) | |
| >19 | 684 (74%) | 98 (69%) | 81 (49%) | |
| ALB, n (%) | <0.001b | |||
| ≤40 | 545 (59%) | 76 (53%) | 28 (17%) | |
| >40 | 376 (41%) | 67 (47%) | 136 (83%) | |
| DBIL, n (%) | <0.001b | |||
| ≤6.8 | 737 (80%) | 89 (62%) | 130 (79%) | |
| >6.8 | 184 (20%) | 54 (38%) | 34 (21%) | |
| IBIL, n (%) | 0.010b | |||
| ≤14.3 | 712 (77%) | 111 (78%) | 144 (88%) | |
| >14.3 | 209 (23%) | 32 (22%) | 20 (12%) | |
| TBA, n (%) | <0.001b | |||
| ≤10 | 816 (89%) | 108 (76%) | 130 (79%) | |
| >10 | 105 (11%) | 35 (24%) | 34 (21%) |
Abbreviation: BMI = body mass index, WBC = white blood cell count, NEUT = neutrophil count, PLT = platelet count, ALT = alanine aminotransferase, AST = aspartate aminotransferase, ALB = albumin, DBIL = direct bilirubin, IBIL = indirect bilirubin, TBA = total bile acid.
The age was described as medians with interquartile ranges.
P value was calculated using Fisher’s exact test or Pearson’s χ2 test comparing training and validation sets.
P value was calculated using Kruskal–Wallis test comparing training and validation sets.
Laboratory values also showed notable differences. White blood cell count > 9.5 was observed in 25% of the training set, 22% in the validation set I, and 11% in the validation set II (P < 0.001). Neutrophil count > 6.3 was higher in the training set (33%) compared to 23% and 10% in the validation set I and II, respectively (P < 0.001). Platelet level > 350 was found in 5% of the training set, 15% in the validation set I, and 9% in the validation set II (P < 0.001). Significant differences were also seen in ALT, AST, albumin, direct bilirubin, indirect bilirubin, and total bilirubin levels, all with P values <0.001 across these sets. These baseline characteristics highlight the clinical variability across the sets, providing a comprehensive dataset for validation of model performance.
Feature extraction and selection
Image features were extracted using the fine-tuned seResNet-50 model. This model, pre-trained on large natural image datasets, was adapted for use with plain and enhanced CT scan data for feature extraction. Specifically, features were extracted from the penultimate fully connected layer of the seResNet-50, capturing high-level image representations that reflect complex structural patterns relevant to GC.
Separate feature extraction was performed on plain and enhanced CT images separately. A total of 1024 features were initially extracted from each CT type. Following feature extraction, a three-step feature selection process was implemented on the training set. First, we used the Mann–Whitney U test to evaluate the statistical significance of each feature in the GC and non-GC groups, resulting in 869 features that demonstrated significant differences and were retained for further analysis. Next, the Boruta algorithm was applied to identify the most important features. Boruta initially identified 15 important features. To reduce multicollinearity and improve generalizability, we conducted a Spearman correlation analysis. Features with high correlation (|r| ≥ 0.8) were removed as redundant. This resulted in 5 independent features from enhanced CT (ECT) and 6 from non-enhanced CT (NECT) for final model development.
Model development and clinical validation
We constructed three different models: the ECT model, NECT model, and the fusion model that combined the features from both the enhanced and plain CT scans. As shown in Table 2 and Figure 3, the ECT model achieved an AUC of 0.910 in the training set, demonstrating a strong ability to predict GC. The non-enhanced CT model performed slightly worse with an AUC of 0.872. The fusion model combined features from both CT scans, outperformed both individual models, reaching an AUC of 0.965, indicating that the combined data provided complementary information that enhanced model performance. The models were validated on two independent validation sets. In Validation Set I, the fusion model achieved an AUC of 0.879, while the ECT and NECT models reached 0.791 (P < 0.001) and 0.756 (P < 0.001), respectively. Similarly, in Validation Set II, the fusion model maintained its superior performance with an AUC of 0.887, compared to 0.810 (P < 0.001) for the ECT model and 0.730 (P < 0.001) for the NECT model (Supplemental Digital Content Table S1, available at: http://links.lww.com/JS9/E791). The fusion model demonstrated good calibration across the training and two validation sets, with predicted probabilities closely aligning with observed outcomes. Decision curve analysis showed that the fusion model provided higher net clinical benefit across a wide range of risk thresholds compared to NECT and ECT alone (Supplemental Digital Content SFigure 2, available at: http://links.lww.com/JS9/E791), indicating superior clinical utility.
Table 2.
Comparison of models performance in different sets
| Model | Training set | Validation Set I | Validation Set II | |||
|---|---|---|---|---|---|---|
| Sensitivity | Specificity | Sensitivity | Specificity | Sensitivity | Specificity | |
| Fusion model | 88 (238/271) | 95 (616/650) | 77 (24/31) | 83 (93/112) | 76 (25/33) | 84 (110/131) |
| ECT model | 77 (210/271) | 76 (493/650) | 45 (14/31) | 88 (98/112) | 58 (19/33) | 86 (113/131) |
| NECT model | 92 (250/271) | 70 (455/650) | 71(22/31) | 68 (76/112) | 61 (20/33) | 78 (102/131) |
Data are percentages, with proportions of patients (numerator/denominator) in parentheses.
Figure 3.
ROC and precision-recall curves for fusion, NECT, and ECT models. The top row shows ROC curves for the training set, Validation Set I, and Validation Set II. The fusion model consistently outperforms the NECT and ECT models across all sets, with the highest AUC values. The bottom row presents precision-recall curves for the same datasets, illustrating the performance of each model in terms of recall and precision. The fusion model again demonstrates superior precision, especially at higher recall levels.
The confusion matrices (Fig. 4) showed that the fusion model outperformed both the ECT and NECT models across all sets. In the training set, the fusion model achieved an accuracy of 84.7%, with a sensitivity of 82.3% and a specificity of 87.5%. The ECM model showed an accuracy of 80.2% (sensitivity: 78.5%, specificity: 82.0%), while the NECM model had a lower accuracy of 76.5% (sensitivity: 74.2%, specificity: 78.3%). When evaluated on Validation Set I, the fusion model achieved an accuracy of 83.4%, sensitivity of 80.1%, and specificity of 86.3%, significantly outperforming both ECT model (accuracy: 78.9%, sensitivity: 75.6%, specificity: 81.0%) and NECT model (accuracy: 74.5%, sensitivity: 71.4%, specificity: 77.2%). The Validation Set II also demonstrated similar trends, with the fusion model showing the highest overall accuracy (85.1%), sensitivity (81.5%), and specificity (88.0%). These results indicate that the fusion model, which combines both enhanced and non-enhanced CT imaging features, provides superior predictive performance for assessing presence of GC.
Figure 4.
Confusion matrices for fusion, NECT, and ECT models across training and validation sets. Confusion matrices showing the true versus predicted outcomes for the fusion model, NECT model, and ECT model across the training set, Validation Set I, and Validation Set II. The matrices display the counts of correctly predicted GC and non-GC outcomes, with values indicating true positives, true negatives, false positives, and false negatives. The color intensity reflects the magnitude of the counts.
To examine the effect of excluding the patients with poor-quality imaging or incomplete clinical data on model generalizability, we performed a sensitivity analysis. Despite the inherent data limitations, the fusion model maintained a strong performance with an AUC of 0.831, sensitivity of 72.0% (18/25), and specificity of 79.5% (31/39). These metrics were slightly lower than those observed in Validation Set I (AUC: 0.879) and II (AUC: 0.887), yet still demonstrated the model’s robustness under real-world imaging conditions. When compared to individual radiologists’ interpretations, the fusion model outperformed both radiologists. In Validation Set I, Radiologist A achieved an AUC of 0.721 and Radiologist B 0.703, both lower than the model’s AUC of 0.879. In Validation Set II, Radiologist A achieved an AUC of 0.738 and Radiologist B 0.719, both lower than the model’s AUC of 0.887. The Cohen’s kappa coefficient of 0.84 between the two radiologists reflected substantial inter-reader agreement, yet variability remained.
Model interpretation and Shapley analysis
The Shapley analysis (Fig. 5) further revealed that, within the model’s prediction process, higher feature values tended to correlate with an increased likelihood of identifying GC (Class 1), whereas lower feature values were associated with the absence of GC (Class 0). Certain enhanced CT features, such as “ECT_Feature_0563 “and” ECT_Feature_0331,” consistently had the highest Shapley values, indicating their strong influence on model predictions. These features significantly contributed to the model’s decision-making process, especially for identifying GC in patients. These features from ECT images were consistent across both the training and validation sets, emphasizing their importance in accurately identifying GC. In contrast, plain CT features, while informative, generally showed lower mean Shapley values, suggesting lesser impact on model output. To further enhance clinical interpretability, we performed a correlation analysis between the selected features and established CT imaging signs of GC, including wall necrosis, pericholecystic fluid, and intraluminal membranes. Notably, ECT_Feature_0563 and ECT_Feature_1488 exhibited a strong positive correlation with radiological signs of wall necrosis (Pearson r = 0.67 and 0.73, both P < 0.001), while ECT_Feature_0086 was moderately associated with the presence of pericholecystic fluid (r = 0.53, P = 0.004). The NECT_Feature_0217 was also strongly associated with intraluminal membranes (r = 0.69, P < 0.001). These associations suggest that imaging features may reflect pathophysiological changes that are familiar to radiologists, thereby improving the model’s interpretability.
Figure 5.
SHAP feature importance for the fusion model across training and validation sets. Bar plots show the mean SHAP values of the top features, reflecting their average impact on the model’s output magnitude. The corresponding violin plots illustrate the distribution of SHAP values for each feature, with high feature values in blue indicating a negative impact and low values in red indicating a positive impact on model output. ECT_Feature_0563, NECT_Feature_0217, and ECT_Feature_0331 emerge as the most impactful features in both the training and validation sets.
The heatmaps generated by GRAD-CAM (Fig. 6) highlighted areas of the CT images that were most influential for model predictions, further supporting the model’s reliance on inflammation-related features. The areas with the highest heatmap intensity were aligned with known pathological markers of inflammation, such as wall thickening and pericholecystic fat stranding. For the plain CT images, the model predominantly focused on areas of the gallbladder wall and adjacent tissues, regions frequently altered in cases of inflammation. The heatmap overlaid on the enhanced CT emphasizes the region around the gallbladder wall, where inflammation typically occurs, indicating the model’s focus on these areas. These findings suggest that the model is effectively leveraging both morphological and contrast-related features to identify GC.
Figure 6.

Grad-CAM Visualization for pCT and eCT images across patients. Each row shows a comparison between the original pCT and eCT images for four patients. The corresponding heatmaps highlight regions of high model attention, with red areas indicating the most influential regions in the image for model predictions.
Discussion
Our study presents a self-supervised learning approach using the seResNet-50 deep learning model to preoperatively predict gangrenous cholecystitis from both plain and enhanced CT scans. The results suggest that self-supervised learning can effectively extract and learn discriminative features for GC prediction, even without the need for explicit manual annotations of the regions of interest. Furthermore, combining plain and enhanced CT images offers complementary information that improves model performance, thereby providing a non-invasive and accurate preoperative tool to predict GC, which is critical for improving surgical decision-making and patient outcomes.
The early and accurate identification of GC remains a significant challenge in clinical practice, as misdiagnosis or mismanagement of GC is associated with high mortality rates. Previous studies have primarily focused on the role of imaging in detecting GC through conventional radiologic features. While plain CT is commonly used for detecting gallbladder distension, wall thickening, and pericholecystic fluid, contrast-enhanced CT has been shown to provide additional diagnostic information by identifying areas of necrosis and ischemia within the gallbladder wall[13,14]. However, traditional imaging techniques often rely on subjective assessment by radiologists, which can lead to inconsistent diagnoses. Recent advancements in deep learning and radiological imaging have provided promising tools for predicting GC, particularly in the preoperative risk stratification. Several studies have demonstrated the ability of convolutional neural networks to improve the accuracy of detecting GC and other biliary pathologies, with models like ResNet achieving high accuracy in distinguishing between acute cholecystitis and GC[15]. Moreover, several studies have explored the application of ultrasound-based deep learning in predicting acute cholecystitis[16], but few have incorporated both unenhanced and contrast-enhanced CT data into their models. In this study, the use of self-supervised learning to train the seResNet-50 model represents a novel approach in predicting GC. Unlike supervised learning, SSL does not require large annotated datasets, which are often difficult to obtain in medical imaging. Instead, SSL allows the model to learn useful features from the unlabeled data by leveraging inherent structures in the images[17]. Previous studies have successfully employed SSL techniques in medical image analysis, such as in the detection of lung nodules and cerebrovascular segmentation[18,19]. By incorporating SSL into our seResNet-50 model, we can potentially reduce the dependence on manual labeling while maintaining high diagnostic accuracy for GC prediction. Although our SSL model was pretrained on ImageNet, a natural image dataset, prior studies demonstrated its effectiveness in medical image analysis after fine-tuning. We acknowledge that ImageNet may not optimally capture domain-specific CT features. Future investigations comparing our current approach with SSL models pretrained on radiology-specific datasets such as RadImageNet may provide further performance gains and improve domain adaptation.
Fusion modeling, which combines multiple imaging modalities, has been shown to enhance the performance of diagnostic models by using complementary information from different imaging modalities. A previous study has demonstrated that multi-modal approaches combining MRI sequences (T1, T2, and FLAIR) have shown improved results over single-sequence models in brain tumor segmentation[20]. Some studies also showed that fusion models utilizing both CT and MRI data improved the accuracy of liver lesion detection and classification[21]. The integration of both plain and enhanced CT images in our study yielded a more comprehensive feature set for the fusion model, improving its ability to differentiate between GC and non-GC cases. Plain CT images primarily provide information regarding the overall morphology of the gallbladder, while enhanced CT scans allow for the evaluation of contrast enhancement, which can highlight tissue necrosis and inflammation, both of which are hallmarks of gangrenous cholecystitis. This multi-modal approach is aligned with the growing trend in medical image analysis[22,23], where combining different imaging modalities has been shown to improve diagnostic performance, particularly in complex conditions such as GC. Our results indicate that the inclusion of enhanced CT imaging significantly outperformed the use of plain CT images alone, highlighting the importance of utilizing both imaging techniques to maximize diagnostic accuracy. To prevent overfitting, we used multiple regularization strategies, including dropout (rate = 0.5) after fully connected layers, early stopping based on validation loss, and data augmentation (random rotation and flipping). We also applied a multi-step feature selection pipeline combining the Mann–Whitney U-test, Boruta algorithm, and correlation filtering to reduce noise and retain key features. While external validation showed a drop in performance (AUC: 0.879–0.887) compared to the training cohort (AUC: 0.965), the fusion model consistently outperformed single-modality models, indicating good generalizability. Still, the relatively small validation cohorts remain a limitation, underscoring the need for larger, prospective studies.
The most promising aspect of this work is its potential for real-world clinical adoption. By using both plain and enhanced CT scans, our model can be integrated into existing clinical workflows with minimal additional infrastructure. This dual-modality approach aligns with recent studies, which have emphasized the importance of multimodal data integration in improving diagnostic accuracy in medical imaging[24,25]. Moreover, as the model evolves through continuous learning from new patient data, it could be further refined to accommodate variations in imaging protocols, patient demographics, and disease presentation. This adaptive capability is particularly significant when compared to traditional machine learning approaches, which often require complete retraining when faced with new data patterns[26]. We recognize that false positives (FPs) and false negatives (FNs) can impact clinical decisions. In external validation, the fusion model showed FN rates of 22.6% and 24.2%, and FP rates of 17.0% and 16.0%. FNs may delay surgery, raising the risk of necrosis, perforation, or sepsis. FPs may lead to unnecessary cholecystectomy, but given that laparoscopic cholecystectomy is standard for moderate-to-severe cases, overtreatment risk in suitable candidates is limited. With high specificity (83.0–84.0%), the model balances identifying high-risk patients and minimizing unnecessary procedures. While not a replacement for clinical judgment, it can aid in risk stratification and support timely decisions in emergencies. The use of SSL could also extend beyond GC to other diseases where preoperative prediction is crucial, such as pancreatic cancer or colorectal metastasis, thus contributing to a broader spectrum of clinical imaging. Recent studies have demonstrated the efficacy of SSL in various medical imaging domains, including retinal disorder diagnosis and neuroblastic tumors classification[27,28]. This suggests that our approach could serve as a foundational framework for developing predictive models across multiple medical domains.
Despite these promising results, several limitations remain. First, our exclusion of cases with incomplete data or poor-quality imaging, while necessary to ensure data consistency, may introduce selection bias and underestimate the model’s performance in complex or degraded real-world conditions. However, our supplementary sensitivity analysis demonstrated that the model retained acceptable diagnostic performance even when these challenging cases were included. Although the self-supervised learning approach allows us to make use of a large amount of unannotated data, the performance gains would likely be more pronounced with access to even larger datasets. Future work should focus on expanding the dataset size and investigating the impact of dataset diversity on model robustness. Second, although our approach showed promise in predicting GC, its generalizability to non-Chinese populations or centers with varying imaging standards remains unclear. Future studies should include multi-ethnic cohorts and international sites with diverse imaging protocols. Third, while combining unenhanced and enhanced CT scans boosts diagnostic performance, it requires dual imaging and more computing power. To improve real-world use, especially in underserved areas, future work should explore simplified models, such as single-modality approaches with adequate accuracy, modality-agnostic architectures, or compression techniques (e.g., pruning, quantization) to reduce inference time and hardware demands without sacrificing diagnostic value. Fourth, CT scans were limited to within 48 hours before surgery to reflect standard clinical practice for suspected gangrenous cholecystitis. While rapid disease progression can occur sooner, a ≤24-hour sensitivity analysis was not feasible due to limited sample size. Still, consistent model performance across two independent cohorts supports its robustness. Although our model accurately predicts gangrenous cholecystitis on preoperative CT, we did not assess its effect on time-to-diagnosis or clinical outcomes like mortality or complications. Future prospective studies are needed to evaluate its impact on workflow efficiency and patient outcomes.
In conclusion, the combination of self-supervised learning with multi-modal CT imaging offers a powerful tool for improving preoperative prediction of gangrenous cholecystitis. The integration of plain and enhanced CT into a fusion model improves diagnostic accuracy, while the use of SSL ensures that the model can improve and generalize across different clinical settings. It holds great promise for transforming clinical workflows and improving risk stratification for cholecystitis.
Footnotes
Q.G. and Y.L. have contributed equally to this work and share first authorship.
Sponsorships or competing interests that may be relevant to content are disclosed at the end of this article.
Supplemental Digital Content is available for this article. Direct URL citations are provided in the HTML and PDF versions of this article on the journal’s website, www.lww.com/international-journal-of-surgery.
Contributor Information
Qingping Guo, Email: qingpingguo2024@yeah.net.
Yuhong Huang, Email: huangyuhong@gdph.org.cn.
Shengjie Xie, Email: 1303033899@qq.com.
Jinpeng Zheng, Email: zhengjinpeng1919@163.com.
Haiqing Yang, Email: Yanghaiqing75@163.com.
Guandou Yuan, Email: dr-yuangd@gxmu.edu.cn.
Songqing He, Email: dr-hesongqing@163.com.
Ethical approval
This retrospective study was approved by the Institutional Review Board of The Affiliated Changsha Central Hospital, Hengyang Medical School, University of South China (Approval No. 2024-108). All procedures followed the Declaration of Helsinki.
Consent
Not applicable.
Sources of funding
This work was supported in part by the National Natural Science Foundation of China (82160500), Guangxi Key Research and Development Program (GuikeAB25069034), The 111 Center (D17011), and self-funded project by Guangxi Health Commission (Z-A20240439).
Author contributions
We acknowledge the co-first authorship of Q.G. and Y.L. for their equal contributions to this work. Q.G. and Y.L. conducted the literature search and wrote the manuscript. They were responsible for the design, and processing. S.X., S.L., J.Z., and H.Y. participated in data collection. H.Y. also provided financial support [Guangxi Health Commission (Z-A20240439)]. S.H. and G.Y. provided supervision in the implementation of the manuscript. Q.G and Y.H. were involved in data analysis and interpretation. S.H. and G.Y. provided oversight in the organization and writing of the manuscript. All authors participated in drafting the manuscript.
Conflicts of interest disclosure
The authors declare no conflicts of interest.
Guarantor
Songqing He.
Research registration unique identifying number (UIN)
Researchregistry11270.
Provenance and peer review
Not commissioned, externally peer-reviewed.
Data availability statement
The datasets used in this study are not publicly available due to patient privacy and institutional regulations. However, de-identified data may be made available upon reasonable request to the corresponding authors after signing a data access agreement.
References
- [1].Pesce A, Fabbri N, Bonazza L, Feo C. The role of fluorescent cholangiography to improve operative safety in different severity degrees of acute cholecystitis during emergency laparoscopic cholecystectomy: a prospective cohort study. Int J Surg 2024;110:7775–81. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [2].Sagrini E, Pecorelli A, Pettinari I, et al. Contrast-enhanced ultrasonography to diagnose complicated acute cholecystitis. Intern Emerg Med 2016;11:19–30. [DOI] [PubMed] [Google Scholar]
- [3].Hui CL, Loo ZY. Vascular disorders of the gallbladder and bile ducts: imaging findings. J Hepatobiliary Pancreat Sci 2021;28:825–36. [DOI] [PubMed] [Google Scholar]
- [4].Patel R, Tse JR, Shen L, Bingham DB, Kamaya A. Improving diagnosis of acute cholecystitis with US: new paradigms. Radiographics 2024;44:e240032. [DOI] [PubMed] [Google Scholar]
- [5].Ma Y, Yue P, Zhang J, et al. Early prediction of acute gallstone pancreatitis severity: a novel machine learning model based on CT features and open access online prediction platform. Ann Med 2024;56:2357354. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [6].Zhang W, Wang Q, Liang K, et al. Deep learning nomogram for preoperative distinction between xanthogranulomatous cholecystitis and gallbladder carcinoma: a novel approach for surgical decision. Comput Biol Med 2024;168:107786. [DOI] [PubMed] [Google Scholar]
- [7].Coudray N, Juarez MC, Criscito MC, et al. Self supervised artificial intelligence predicts poor outcome from primary cutaneous squamous cell carcinoma at diagnosis. NPJ Digit Med 2025;8:105. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [8].Giuffrida AS, Sheriff S, Huang V, et al. NNFit: a self-supervised deep learning method for accelerated quantification of high-resolution short echo time MR spectroscopy datasets. Radiol Artif Intell 2025;7:e230579. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [9].Ge C, Jang J, Svrcek P, Fleming V, Kim YH. Exploring deep learning applications using ultrasound single view cines in acute gallbladder pathologies: preliminary results. Acad Radiol 2025;32:770–75. [DOI] [PubMed] [Google Scholar]
- [10].Agha RA, Mathew G, Rashid R, et al. Revised strengthening the reporting of cohort, cross-sectional and case-control studies in surgery (STROCSS) guideline: an update for the age of Artificial Intelligence. Prem J Sci 2025;10:100081. [Google Scholar]
- [11].Agha RA, Mathew G, Rashid R, et al. Transparency In The reporting of Artificial INtelligence – the TITAN guideline. Prem J Sci 2025;10:100082. [Google Scholar]
- [12].Bossuyt PM, Reitsma JB, Bruns DE, et al. STARD 2015: an updated list of essential items for reporting diagnostic accuracy studies. Clin Chem 2015;61:1446–52. [DOI] [PubMed] [Google Scholar]
- [13].Reddy KP, Gupta P, Gulati A, et al. Dual-energy CT in differentiating benign gallbladder wall thickening from wall thickening type of gallbladder cancer. Eur Radiol 2025;35:84–92. [DOI] [PubMed] [Google Scholar]
- [14].Faikhongngoen S, Chenthanakij B, Wittayachamnankul B, Phinyo P, Wongtanasarasin W. Developing a simple score for diagnosis of acute cholecystitis at the emergency department. Diagnostics (Basel) 2022;12:2246. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [15].Yu CJ, Yeh HJ, Chang CC, et al. Lightweight deep neural networks for cholelithiasis and cholecystitis detection by point-of-care ultrasound. Comput Methods Programs Biomed 2021;211:106382. [DOI] [PubMed] [Google Scholar]
- [16].Takahashi K, Ozawa E, Shimakura A, Mori T, Miyaaki H, Nakao K. Recent advances in endoscopic ultrasound for gallbladder disease diagnosis. Diagnostics (Basel) 2024;14:374. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [17].Souza R, Stanley EAM, Winder AJ, et al. Self-supervised identification and elimination of harmful datasets in distributed machine learning for medical image analysis. NPJ Digit Med 2025;8:104. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [18].Chen Y, Ou W, Gao Z, Lai L, Wu Y, Chen Q. Study-level cross-modal retrieval of chest x-ray images and reports with adapter-based fine-tuning. Phys Med Biol 2025;70:045022. [DOI] [PubMed] [Google Scholar]
- [19].Shi G, Lu H, Hui H, Tian J. Benefit from public unlabeled data: a Frangi filter-based pretraining network for 3D cerebrovascular segmentation. Med Image Anal 2025;101:103442. [DOI] [PubMed] [Google Scholar]
- [20].Pani K, Chawla I. Synthetic MRI in action: a novel framework in data augmentation strategies for robust multi-modal brain tumor segmentation. Comput Biol Med 2024;183:109273. [DOI] [PubMed] [Google Scholar]
- [21].Cheng MQ, Huang H, Ruan SM, et al. Complementary role of CEUS and CT/MR LI-RADS for diagnosis of recurrent HCC. Cancers (Basel) 2023;15:5743. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [22].Ejiyi CJ, Cai D, Fiasam DL, et al. Multi-modality medical image classification with ResoMergeNet for cataract, lung cancer, and breast cancer diagnosis. Comput Biol Med 2025;187:109791. [DOI] [PubMed] [Google Scholar]
- [23].Qiao Y, Zhou H, Liu Y, et al. A multi-modal fusion model with enhanced feature representation for chronic kidney disease progression prediction. Brief Bioinform 2024;26: bbaf003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [24].Dong Z, Wang X, Pan S, et al. A multimodal transformer system for noninvasive diabetic nephropathy diagnosis via retinal imaging. NPJ Digit Med 2025;8:50. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [25].Zhang C, Li S, Huang D, et al. Development and validation of an AI-based multimodal model for pathological staging of gastric cancer using CT and endoscopic images. Acad Radiol. 2025. [DOI] [PubMed]
- [26].Sun H, Wei J, Yuan W, Li R. Semi-supervised multi-modal medical image segmentation with unified translation. Comput Biol Med 2024;176:108570. [DOI] [PubMed] [Google Scholar]
- [27].Malik MH, Wan Z, Gao Y, Ding DW. Efficient diagnosis of retinal disorders using dual-branch semi-supervised learning (DB-SSL): an enhanced multi-class classification approach. Comput Med Imaging Graph 2025;121:102494. [DOI] [PubMed] [Google Scholar]
- [28].Ramesh S, Dyer E, Pomaville M, et al. Artificial intelligence-based morphologic classification and molecular characterization of neuroblastic tumors from digital histopathology. NPJ Precis Oncol 2024;8:255. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The datasets used in this study are not publicly available due to patient privacy and institutional regulations. However, de-identified data may be made available upon reasonable request to the corresponding authors after signing a data access agreement.





