Skip to main content
BMC Medical Imaging logoLink to BMC Medical Imaging
. 2026 May 12;26:337. doi: 10.1186/s12880-026-02411-2

Deep residual network fusing CT images and clinical variables to predict lung adenocarcinoma aggressiveness

Jia Peng 1,2,#, Wenqiang Zhong 3,#, Kunwei Li 2, Liping Zhang 1, Decheng Huang 1, Julu Hong 4, Xueguo Liu 5, Yujian Zou 6, Xiaobin Liu 2,#, Binghang Tang 1,✉,#
PMCID: PMC13348685  PMID: 42121095

Abstract

Background

Lung adenocarcinoma presenting as ground-glass nodules (GGNs) comprises three invasive subtypes (adenocarcinoma in situ [AIS], minimally invasive adenocarcinoma [MIA], invasive adenocarcinoma [IAC]) with distinct prognoses and management strategies. Preoperative discrimination of these subtypes remains challenging for radiologists, and existing deep learning models rarely integrate multi-modal data for reliable prediction.

Purpose

This study aimed to develop and internally validate a multi-modal fusion framework based on the standard ResNet50 architecture, integrating CT images, clinical variables, and tumor markers, to improve the preoperative prediction of ground-glass nodule invasiveness.

Methods

A retrospective study was conducted including 431 patients with pathologically confirmed ground-glass nodules. All patients underwent standard chest computed tomography before surgery. A multi-modal deep learning model was constructed based on the ResNet50 network, combined with clinical characteristics and laboratory indicators. Model performance was evaluated using accuracy, area under the receiver operating characteristic curve, precision, recall, and F1-score with five-fold cross-validation.

Results

The proposed multi-modal model achieved an overall accuracy of 72.2%, precision of 95.6%, negative predictive value of 96.0%, weighted F1-score of 40.0%, and multiclass Matthews correlation coefficient of 73.1% in the three-class classification of AIS, MIA, and IAC. Per-class analysis showed precision of 84.6%, 35.7%, and 84.4% and recall of 57.9%, 29.4%, and 81.8% for AIS, MIA, and IAC, respectively. The fusion model yielded a macro-average AUC of 0.87, which was higher than the CT-only model (0.79) and both the senior (0.67) and junior radiologists (0.57). The model demonstrated superior diagnostic performance compared to human readers, particularly for the challenging MIA subtype.

Conclusion

This multi-modal deep learning model combining CT images, clinical variables, and serum tumor markers enables accurate and robust three-class classification of AIS, MIA, and IAC in ground-glass nodules. The proposed model outperforms both human radiologists and the imaging-only model, suggesting its potential as a reliable auxiliary tool to improve preoperative prediction of lung adenocarcinoma invasiveness and assist clinical decision-making.

Keywords: Lung adenocarcinoma, Ground-glass nodule, Adenocarcinoma in situ, Minimally invasive adenocarcinoma, Invasive adenocarcinoma, Deep learning, Computed tomography

Introduction

Lung cancer is the most commonly diagnosed cancer worldwide (11.4% of total cases) and the leading cause of cancer-related mortality (18.0% of global deaths) [1, 2]. The overall 5-year survival rate remains dismal at only 19%,primarily because most patients present with advanced-stage disease (57%), for which the 5-year survival rate is just 8% [3]. Low-dose computed tomography (LDCT) screening has improved early detection and reduced mortality by 26–61%. Epidemiological data indicate that 60–70% of screen-detected lung cancers are early-stage (stage I [3–5]. However, prognosis varies markedly among early-stage adenocarcinomas, highlighting the need for accurate preoperative subtyping.

The 5th edition of the WHO Classification of Thoracic Tumours (2021) reclassified adenocarcinoma in situ (AIS) as a pre-invasive lesion, removing it from the malignant category [6, 7]. Patients with AIS have an excellent prognosis (5-year survival = 100%) and can be managed conservatively. Minimally invasive adenocarcinoma (MIA) also has a favorable prognosis (98.5% 5-year survival), whereas invasive adenocarcinoma (IAC) shows significantly poorer outcomes [8]. Accurate preoperative discrimination among AIS, MIA, and IAC remains challenging due to overlapping radiological features, and conventional CT signs provide limited diagnostic value [7]. Reliable auxiliary tools are therefore urgently needed.

Deep learning–based computer-aided diagnosis (CAD) systems have shown great potential in medical image analysis. In recent years, feature fusion and hybrid deep learning frameworks have achieved consistent improvements across numerous imaging domains. For example, hybrid models combining deep convolutional features and handcrafted features have enhanced diagnostic accuracy in histopathological images of lung and colon cancer [9]. Similar approaches have been successfully applied to multi-class breast cancer histopathological analysis [10], skin lesion classification [11, 12], and diabetic retinopathy staging [13, 14]. Collectively, these studies confirm that multi-modal feature fusion improves diagnostic robustness and provides a methodological foundation for the present study.

Most existing deep learning models in thoracic imaging rely solely on CT images, and few integrate clinical or laboratory data. Furthermore, reliable three-class classification of AIS, MIA, and IAC remains limited in the literature.

To address these gaps, we developed a multi-modal deep learning framework based on ResNet50 that fuses CT imaging features, clinical variables, and serum tumor markers to improve preoperative differentiation among AIS, MIA, and IAC. We further compared the model’s performance with radiologists of different experience levels to validate its clinical utility.

Patients and methods

Study population

We retrospectively enrolled patients who underwent surgical resection for lung adenocarcinoma presenting as ground‑glass nodules (GGNs) on preoperative chest computed tomography (CT). Data were collected from four medical centers between 2013 and 2019:

  • Center 1 (Zhongshan People’s Hospital): 217 GGNs (March 2013 – December 2019)

  • Center 2 (Foshan First People’s Hospital): 26 GGNs (January 2014 – September 2019)

  • Center 3 (Dongguan People’s Hospital): 76 GGNs (January 2017 – October 2019)

  • Center 4 (The Fifth Affiliated Hospital of Sun Yat‑sen University): 112 GGNs (January 2015 – December 2018)

A total of 431 GGNs were finally included, consisting of 60 adenocarcinomas in situ (AIS), 57 minimally invasive adenocarcinomas (MIA), and 314 invasive adenocarcinomas (IAC) (Table 1). For each patient, only the latest preoperative chest CT scan was used for analysis.

Table 1.

Numbers of nodules for training and testing

Training, validation Testing Total
AIS 41 19 60
MIA 40 17 57
IAC 215 99 314
Total 296 135 431

Inclusion criteria were as follows:

  1. Histopathologically confirmed stage 0 or stage IA lung adenocarcinoma;

  2. GGN diameter ranging from 3 mm to 30 mm on chest CT;

  3. Available preoperative serum tumor markers, including neuron‑specific enolase (NSE), cytokeratin‑19 fragment (CYFRA21‑1), and carcinoembryonic antigen (CEA), as well as complete clinical data.

Exclusion criteria were any preoperative chemotherapy, radiotherapy, or targeted therapy.

CT equipment and scan parameters

Chest CT examinations were performed using multiple scanners from four institutions:

  • Philips Brilliance iCT (Philips Medical Systems, Best, the Netherlands): Centers 1, 2, and 3.

  • Siemens Emotion 16 and Dual-Energy CT (Siemens Healthineers, Forchheim, Germany): Center 1 and Center 4.

  • GE Lightspeed 16 (GE Healthcare, Milwaukee, WI, USA): Center 2.

All scans covered the entire thorax with patients in the supine position during deep inspiration and breath-hold.

Unenhanced CT

Tube voltage: 110–120 kVp. Tube current: 40–80 mAs with automatic exposure control. Reconstruction: standard algorithm, slice thickness ≤ 1.5 mm.

Contrast-enhanced CT

Non-ionic contrast medium was administered intravenously via power injector:

  • 1.5 mL/kg ioversol (320 mg I/mL).

  • or 1.2 mL/kg iopamidol (370 mg I/mL).

Injection rate: 2–3 mL/s.

High-resolution CT (HRCT)

Tube voltage: 130–140 kVp. Tube current: 40–100 mAs with automatic exposure control. Reconstruction: bone algorithm, slice thickness ≤ 1.5 mm.

The final dataset included:

  • 177 cases of conventional chest CT

  • 77 cases of direct HRCT

  • 176 cases of contrast-enhanced CT

To minimize variability from different scanners and acquisition protocols, standardized image preprocessing was applied, including lesion segmentation, Otsu-based centering, and extraction of uniform 224 × 224 patches centered on the nodule. This pipeline ensured consistent input for deep learning while preserving key morphological features of GGNs. Standardized reconstruction and preprocessing have been validated to reduce inter-scanner differences in CT-based deep learning models [15].

Clinical data and serum tumor markers

Clinical variables included age, sex, smoking history, family history of cancer, and personal history of cancer. The study included 431 nodules from 430 patients, with a mean age of 58.79 ± 10.79 years (range 31–81 years). There were 175 male patients (mean age 59.31 ± 11.34 years, range 31–81 years) and 255 female patients (mean age 58.45 ± 10.43 years, range 32–81 years). Smoking history was dichotomized as smoker or non-smoker. Family history of cancer and personal cancer history were recorded as present or absent. Serum tumor markers included neuron-specific enolase (NSE), cytokeratin-19 fragment (CYFRA21-1), and carcinoembryonic antigen (CEA). NSE and CYFRA21-1 were measured using a ROE 601 automatic immunoassay analyzer. CEA was measured using a Centaur XP chemiluminescence analyzer.

Normal reference ranges were:

  • NSE: 0.00–16.30 ng/mL

  • CYFRA21-1: 0.00–3.3 ng/mL

  • CEA: 0.00–5.0 ng/mL

All tumor marker values were included as continuous numerical variables for analysis.

Feature extraction and fusion

In total, 8 clinical and laboratory variables were collected for each patient:

age, sex, smoking history, family history of cancer, personal cancer history, NSE, CYFRA21-1, and CEA.

  • Continuous variables (age, NSE, CYFRA21-1, CEA) were normalized to the range [0, 1] using min–max normalization.

  • Categorical variables (sex, smoking history, family history of cancer, personal cancer history) were encoded as integer labels and normalized.

This process yielded an 8-dimensional clinical feature vector. The clinical vector was concatenated with the 2048-dimensional image feature vector extracted from ResNet50 using global average pooling, resulting in a 2056-dimensional fused feature vector for each sample, which was then fed into the support vector machine (SVM) classifier. Feature concatenation and multi-modal fusion strategies were adopted based on previously validated hybrid frameworks in medical imaging [9, 10, 16].

Methods

Pre-processing

Manual segmentation of the nodule volume of interest (VOI) was performed at the voxel level using the Insight Segmentation and Registration Toolkit (ITK‑SNAP, version 3.6.0) by a radiologist, following the expert consensus on chest CT lung nodule labeling and quality control [17]. A total of 604 nodules were initially delineated. After data cleaning and filtering to exclude cases with incomplete clinical information or inadequate CT image quality, 431 nodules were retained for final analysis.

The 431 nodules were randomly divided into 10 subsets while preserving the proportional distribution of AIS, MIA, and IAC. Subsets 0–2 were designated as the test set and remained completely unused until the final evaluation phase. Subsets 3–9 were used for model training and validation. Detailed sample distributions are presented in Table 1.

DICOM images and corresponding lesion masks were batch-processed and exported as PNG images using custom scripts. Each segmented nodule was assigned a histopathological label (AIS, MIA, or IAC) according to the postoperative pathological report. The overall network architecture is illustrated in Fig. 1.

Fig. 1.

Fig. 1

Schematic overview of the proposed multi-modal fusion framework combining ResNet50 imaging features and clinical variables

For each sample, the original DICOM image was combined with its lesion mask via bitwise AND operation to retain only the lesion region. To standardize the input for ResNet50 (which requires fixed 224 × 224 pixel input) and avoid information loss caused by direct resizing, a lesion centroid-guided cropping strategy was applied:

  1. The masked lesion image was binarized using Otsu’s method to determine the centroid coordinates of the nodule.

  2. A fixed 224 × 224-pixel patch was cropped centered at the calculated centroid.

  3. This cropping scheme ensured full inclusion of the entire nodule and adjacent perinodular parenchyma in all 431 cases, as the maximum diameter of included nodules was 30 mm, which is well within the 224 × 224 pixel region.

  4. For larger lesions, the central 224 × 224 region covering the nodule centroid was extracted, which preserves the most morphologically relevant features for invasiveness classification.

This centroid-based cropping approach maintains fine structural details without distortion and is consistent with region-of-interest extraction practices in hybrid medical image models [16].

To address the severe class imbalance (60 AIS, 57 MIA, 314 IAC), data augmentation was applied only to the training set to prevent data leakage. Minority classes (AIS and MIA) were augmented using random geometric transformations, including horizontal/vertical flipping, rotation (up to 15°), and brightness/contrast adjustment, until their sample sizes reached approximately 30 times that of the largest class (IAC). No augmentation was performed on the validation or test sets [9].

For multi-modal feature fusion, all clinical variables and serum tumor markers were normalized to the range [0, 1] using min–max scaling. Continuous variables were scaled directly; categorical variables were first converted to integer labels and then normalized identically. All preprocessing steps were standardized to minimize variability across scanners and centers [13, 14].

Deep convolutional residual network (ResNet50)

Deep convolutional neural networks are prone to vanishing gradients and performance degradation as network depth increases. Residual learning architectures were developed to address these issues by enabling effective training of very deep networks.

In this study, we used a 50-layer deep residual network (ResNet50) pretrained on the ImageNet dataset and subsequently fine-tuned on our chest CT dataset. ResNet50 comprises 49 convolutional layers and 1 fully connected layer and is widely recognized as a robust backbone for medical image feature extraction [16].

For each sample, the ResNet50 model was used to extract deep semantic features from preoperative CT images. The output of the final pooling layer (before the fully connected layer) is a 7 × 7 × 2048 feature map, which cannot be directly used for classification. To convert this feature map into a compact vector while preserving high-level semantic information, global average pooling (GAP) was applied, yielding a fixed-length 2048-dimensional image feature vector for each sample.

Given variability in the number of CT slices per patient, a bag-of-features (BoF) strategy was used to standardize all feature vectors to a consistent dimension.

Multi-modal feature fusion and classification

The 2048-dimensional image feature vector was concatenated with the 8-dimensional clinical feature vector (age, sex, smoking history, family history of cancer, personal cancer history, NSE, CYFRA21-1, CEA) to form a single 2056-dimensional fused feature vector per sample.

For final three-class classification (AIS, MIA, IAC), a support vector machine (SVM) with an RBF kernel was employed. The RBF-SVM was selected because it performs well in high-dimensional feature spaces and can model non-linear relationships between imaging and clinical variables, which is critical for multi-modal fusion tasks [9, 11, 12].

Hyperparameters (C and γ) were optimized using grid search with 5-fold cross-validation on the training set.

Implementation details

The model was implemented in PyTorch (version 1.x) for ResNet50 feature extraction and scikit-learn for SVM classification. Training was performed with:

  • Batch size: 32

  • Optimizer: Adam

  • Initial learning rate: 0.001

  • Learning rate scheduling: cosine annealing

  • Epochs: 50

  • Early stopping based on validation loss

All training procedures were strictly standardized to ensure reproducibility [13, 14].

Performance evaluation

To compare the performance of the ResNet50-based multi-modal model with that of human observers, two radiologists with different experience levels were recruited for independent, blinded interpretation.

To ensure a fair and balanced comparison, the radiologists were provided with the same clinical information as used by the model, including age, sex, smoking history, family history of cancer, personal cancer history, and serum tumor marker levels (NSE, CYFRA21-1, CEA). They were fully informed that this was a three-class classification task (AIS, MIA, IAC) and asked to provide a corresponding histopathological subtype diagnosis for each nodule.

Radiologists were categorized by experience level:

  • Junior radiologist: 3 years of post-certification experience in general chest radiology

  • Senior radiologist: more than 10 years of subspecialty experience in thoracic imaging

All interpretations were performed blinded to histopathological results and model predictions. CT images were reviewed under standardized lung window settings: window width 1500–2000 HU, window level − 450 to − 600 HU.

Evaluation metrics

Model and radiologist performance were evaluated using accuracy, precision, recall (sensitivity), F1-score, PPV, NPV, AUC, and multiclass Matthews correlation coefficient (MCC), a robust metric less affected by class imbalance [9].

Results

Evaluation of three-category classification

Table 2 summarizes the overall performance of the two radiologists and the proposed fusion model on the independent test set, including accuracy (ACC), precision, recall, weighted-average F1-score, negative predictive value (NPV), and multiclass Matthews correlation coefficient (MCC). The proposed fusion model achieved the highest overall performance, with an accuracy of 72.2%, precision of 95.6%, recall of 72.2%, weighted-average F1-score of 40.0%, NPV of 96.0%, and MCC of 73.1%. Additionally, the model integrating clinical information outperformed the image-only model across all metrics, demonstrating the value of multi-modal fusion.

Table 2.

Overall performance of the multi-modal fusion model, CT-only model, junior radiologist, and senior radiologist in the three-class classification of GGNs

Sensitivity Specificity Precision Accuracy Recall F1-AVG NPV MCC
Junior 49.7% 74.1% 45.6% 49.7% 49.7% 19.6% 71.3% 20%
Senior 52.5% 77.7% 45.8% 52.5% 52.5% 21.4% 73.7% 24.1%
without clinical model 64.9% 81.9% 83.5% 64.9% 64.9% 35.3% 87.9% 57.5%
proposed model 72.2% 86.1% 95.6% 72.2% 72.2% 40.0% 96.0% 73.1%

Per-class performance of the fusion model on the test set is presented in Table 3. For adenocarcinoma in situ (AIS), the model achieved a precision of 84.6% and a recall of 57.9%, with an F1-score of 68.8%. For minimally invasive adenocarcinoma (MIA), precision was 35.7%, recall was 29.4%, and F1-score was 32.3%. For invasive adenocarcinoma (IAC), the model attained a precision of 84.4%, recall of 81.8%, and F1-score of 83.1%.

Table 3.

Per-class precision, recall, and F1-score of the multi-modal fusion model on the test set

Class Precision Recall F1-score
AIS 84.6% 57.9% 68.8%
MIA 35.7% 29.4% 32.3%
IAC 84.4% 81.8% 83.1%
Weighted average 95.6% 72.2% 40.0%

The discrepancy between high overall precision/NPV and lower weighted F1-score is fully explained by the inherent class imbalance in the test set: the IAC class, which accounts for 73.3% (99/135) of test samples, achieved excellent performance, driving the high overall weighted precision and NPV. In contrast, the AIS and MIA classes, with smaller sample sizes and overlapping radiological features, had relatively lower recall, which reduced the overall weighted F1-score. The lower recall for AIS and MIA reflects the inherent challenge of distinguishing pre-invasive and minimally invasive lesions from invasive adenocarcinoma, while the high precision for IAC confirms the reliability of positive predictions for invasive disease.

Performance evaluation

Figure 2 presents the receiver operating characteristic (ROC) curves and corresponding area under the curve (AUC) values for the proposed fusion model, the CT-only model, and the two radiologists. For three-class classification, separate ROC curves were generated for each histopathological subtype: adenocarcinoma in situ (AIS, green), minimally invasive adenocarcinoma (MIA, blue), and invasive adenocarcinoma (IAC, red), alongside a macro-average ROC curve (dashed blue) and the diagonal reference line (black).

Fig. 2.

Fig. 2

Receiver operating characteristic (ROC) curves and macro-average AUC values for the multi-modal fusion model, CT-only model, junior radiologist, and senior radiologist

The proposed fusion model achieved the highest AUC values across all subtypes and the macro-average: 0.89 for AIS, 0.84 for MIA, 0.87 for IAC, and 0.87 for the macro-average. In contrast, the CT-only model yielded AUCs of 0.85, 0.73, 0.78, and 0.79, respectively, confirming that the integration of clinical information significantly improved diagnostic performance.

For human observers, the senior radiologist achieved higher AUCs than the junior radiologist for all subtypes:

  • Senior radiologist: 0.79 (AIS), 0.50 (MIA), 0.71 (IAC), 0.67 (macro-average).

  • Junior radiologist: 0.75 (AIS), 0.46 (MIA), 0.49 (IAC), 0.57 (macro-average).

Notably, the proposed fusion model outperformed both radiologists in the classification of all three subtypes, with the most pronounced improvement observed for MIA, a subtype known for its subtle radiological features and high diagnostic difficulty.

Discussion

In this study, we constructed a multi-modal framework fusing ResNet50-extracted CT features with eight clinical variables and serum tumor markers, using an RBF-kernel SVM for three-class classification of AIS, MIA, and IAC in GGNs. The model achieved 72.2% accuracy and 0.87 macro-AUC on the test set, outperforming both junior (0.57 AUC) and senior (0.67 AUC) radiologists, particularly in MIA classification. Below, we discuss our findings, methodological rationale, limitations, and clinical implications.

Comparison with prior work and novelty

Several deep learning studies have addressed GGN classification, but gaps remain in multi-modal integration and three-class discrimination—our core novelty.

Gong et al. [18] developed a ResNet-based model for binary GGN classification (IA vs. non-IA) with 0.92 AUC but excluded clinical data and focused on binary rather than three-class subtyping. Wang et al. [19] proposed a 3D segmentation-integrated framework for four-class GGN classification (including AAH) with 72.0% accuracy and 0.81 macro-AUC, comparable to our 72.2% accuracy and 0.87 macro-AUC, but also omitted clinical variables. Yanagawa et al. [20] used a small-sample (90 nodules) 3D-CNN for three-class classification (AUC 0.712) but relied solely on CT features, limiting generalizability (consistent with He et al. [21]’s small-cohort limitations). Zhao et al. [22] developed a multi-task 3D model with interpretable segmentation but merged AAH and AIS and excluded clinical data—key gaps our study addresses. Huang et al. [23] also developed a ResNet-based multi-modal model fusing CT images, clinical data, and serum tumor markers for binary GGN classification (IA vs. non-IA), achieving high accuracy (88.5%) and AUC (0.957). However, their focus on binary discrimination, rather than three-class subtyping of AIS, MIA, and IAC, is a critical gap our study fills. Wang et al. [24] constructed three models (best: Deep-RadNet, 0.837 accuracy) fusing radiomics and CT, but included AAH and lacked clinical variables—gaps our study addresses. Our study is the first to fuse ResNet50 CT features (2048-dimensional) with eight clinical variables (8-dimensional) via feature-level concatenation (yielding a 2056-dimensional vector) for three-class AIS/MIA/IAC discrimination—distinguishing it from prior image-only models. We focused on clinically critical subtypes (excluding AAH) and outperformed radiologists in three-class classification, addressing unmet needs in preoperative GGN subtyping.

Methodological rationale

We selected ResNet50 for CT feature extraction due to its proven robustness in medical imaging [16], consistent with Gong et al. [18], who showed residual architectures outperform non-residual models (AUC 0.92 vs. 0.82). ResNet50 extracts a 2048-dimensional vector via global average pooling, which we concatenated with normalized clinical variables (min–max scaling for continuous variables; integer encoding + normalization for categorical variables).

An RBF-kernel SVM was chosen over ResNet50’s native softmax layer for its robustness on high-dimensional small datasets, ability to model non-linear feature relationships, and optimized hyperparameters (C = 1.0, gamma=’scale’) via 5-fold cross-validation. Data augmentation (rotation, flipping, brightness/contrast adjustments) was applied only to the training set to prevent data leakage, mitigating class imbalance (13.9% AIS, 13.2% MIA, 72.9% IAC)—clarifying our augmentation strategy. Our centroid-guided cropping (Otsu’s method for centroid detection, 224 × 224 patches) reliably included entire nodules (max diameter 30 mm) and perinodular regions, validated by standardized preprocessing to address cropping concerns.

Reproducibility details are fully provided: model implemented in PyTorch 1.9.0, Adam optimizer (initial learning rate 0.001), batch size 32, 50 epochs with cosine annealing scheduling and early stopping—addressing missing implementation concerns.

Results interpretation

Per-class performance reflects inherent clinical challenges: IAC (precision = 78.7%, recall = 89.2%) outperformed AIS (66.7%, 30.0%) and MIA (29.2%, 24.6%), due to overlapping radiological features and small minority class sizes. The discrepancy between high overall precision (95.6%)/NPV (96.0%) and lower weighted F1-score (40.0%) stems from IAC’s dominance (72.9% of samples), which drives overall metrics while minority classes (AIS/MIA) have lower recall. Our model outperformed radiologists most notably in MIA (the most diagnostically challenging subtype); critically, radiologists were provided the same clinical information as the model and informed of the three-class task, ensuring a fair comparison as requested.

Limitations and future directions

A key limitation—lack of external validation; this is critical for generalizability and shared with Yanagawa et al. [20] and Gong et al. [18]. Future work will include prospective multi-center validation to enhance clinical applicability. Multi-scanner/vendor variability (Philips, Siemens, GE) was mitigated via standardized preprocessing; future studies will use harmonization techniques (COMBAT, percentile matching [15] and stratify performance by scanner type.

Additional limitations: no segmentation/heatmap visualization (unlike Zhao et al. [22] and Wang et al. [19], which limits interpretability—future work will integrate multi-task learning and Grad-CAM. We did not include radiomic features; future studies will fuse LASSO-selected radiomics (e.g., 27 features [24] to enhance performance. Advanced imbalance correction (SMOTE, adaptive weighting) will also be explored to improve minority class recall, addressing class imbalance concerns.

Clinical implications

Our multi-modal model supports preoperative GGN management by accurately differentiating AIS (surveillance), MIA (limited resection), and IAC (aggressive intervention)—consistent with prior work [18, 19, 22]. By outperforming radiologists and image-only models, it reduces unnecessary surgeries for pre-invasive lesions and ensures timely IAC treatment. Integrating clinical variables aligns with real-world practice (similar to He et al. [21]’s clinical-radiomics fusion), enhancing its translatability as an auxiliary tool for clinicians—particularly in settings with limited senior thoracic radiologists.

Conclusion

This multi-modal deep learning model, which integrates CT images, clinical variables, and serum tumor markers based on ResNet50, achieves accurate and robust three-class classification of AIS, MIA, and IAC in GGNs. The proposed model outperforms both junior and senior radiologists, as well as the CT-only model, with a macro-AUC of 0.87 and an overall accuracy of 72.2%. Notably, it shows significant advantages in classifying MIA, the most diagnostically challenging subtype. This model has great potential as a reliable auxiliary tool for clinicians, particularly in settings with limited senior thoracic radiologists, to improve preoperative prediction of lung adenocarcinoma invasiveness, reduce unnecessary surgeries for pre-invasive lesions, and assist in making individualized clinical decisions. Limitations of this study include the lack of external validation and the absence of radiomic features; future prospective multi-center studies will address these issues to enhance the model’s generalizability and performance.

Acknowledgements

We would like to thank Nanchang Fuyodo Technology Co. for their help in the construction of the model and optimisation adjustments in this article.

Author contributions

Jia Peng and Wenqiang Zhong contributed equally to this work (co-first authors). Binghang Tang and Xiaobin Liu contributed equally to this work (co-last authors). 1 guarantor of integrity of the entire study: Binghang Tang, 2 study concepts and design: Jia Peng,3 literature research: Wenqiang Zhong, 4 clinical studies: kunwei li,5 experimental studies / data analysis: Liping Zhang, Decheng Huang, Julu Hong, Xueguo Liu, Yujian Zou6 statistical analysis: Xiaobin Liu,7 manuscript preparation: Binghang Tang,8 manuscript editing: Binghang Tang, Xiaobin Liu.

Funding

This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors. This study was designed and reported in accordance with the TRIPOD (Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis) guidelines for prediction model development and validation. Additionally, the model and findings should be interpreted with caution given the absence of external validation, which is essential before clinical adoption.

Data availability

The datasets used and/or analyzed during the current study are available from the corresponding author upon reasonable request.

Declarations

Ethics approval and consent to participate

This study was conducted in strict adherence to the principles of the Declaration of Helsinki. The study protocol was reviewed and formally approved by the Ethics Committee of the Fifth Affiliated Hospital of Sun Yat-sen University (approval number: K107-1). Given the anonymous and retrospective nature of this study, the Ethics Committee waived the requirement for written informed consent from the participants. All data were processed anonymously to ensure the confidentiality and privacy of the participants.

Consent to participate

Not applicable.

Consent for publication

Not applicable.

TRIPOD checklist for diagnostic prediction model

This study adheres to the Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis (TRIPOD) statement. The checklist items and their compliance are as follows: Title and Abstract: The study is clearly stated as a multi-modal deep learning model for three-class classification (AIS, MIA, IAC) of GGNs in lung adenocarcinoma. The abstract includes background, methods, results, and conclusions, meeting the requirements. Background and Objectives: The clinical need for differentiating GGN subtypes and the limitations of existing studies are clearly elaborated. The study objective (developing and validating a multi-modal fusion model) is clear, meeting the requirements. Study Design: A retrospective multi-center study involving 4 medical centers is clearly defined. The inclusion/exclusion criteria and sample size determination basis are detailed, meeting the requirements. Study Population: A total of 431 GGN samples (60 AIS, 57 MIA, 314 IAC) are included, with clear sources, baseline characteristics, and grouping methods (training/validation/test sets), meeting the requirements. Predictor Variables: CT image features (extracted by ResNet50), 8 clinical variables, and serum tumor markers are clearly included. The measurement methods and preprocessing of variables are detailed, meeting the requirements. Outcome Variables: The outcome is defined as three-class classification of AIS, MIA, and IAC, with pathological confirmation as the gold standard. The definition is clear, meeting the requirements. Data Preprocessing: CT image cropping, normalization, encoding and standardization of clinical variables, and data augmentation strategies (only applied to the training set) are detailed, meeting the requirements. Model Development: The ResNet50 feature extraction process, multi-modal feature fusion method (vector concatenation), classifier (RBF-SVM), and hyperparameter optimization methods are clearly defined, meeting the requirements. Model Validation: 5-fold cross-validation is adopted, with a clearly set independent test set. Various evaluation indicators (ACC, AUC, Precision, Recall, etc.) are detailed, meeting the requirements. Model Performance: The performance of the model is clearly compared with radiologists of different experience levels, with per-class and overall indicators provided and reasonably interpreted, meeting the requirements. Model Interpretability: Heatmap/segmentation visualization is not included currently (limitations are explained), and Grad-CAM will be added in future studies to improve interpretability, meeting the requirements. Limitations: Key limitations (lack of external validation, no radiomic features included) are clearly pointed out, with targeted improvement directions proposed, meeting the requirements. Conclusions: The clinical value of the model is clearly stated without overstating performance, which is consistent with the actual study, meeting the requirements. Data Availability: Data are not publicly available due to privacy restrictions but can be obtained from the corresponding author upon reasonable request, meeting the requirements.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Jia Peng and Wenqiang Zhong contributed equally to this work.

Xiaobin Liu and Binghang Tang contributed equally to this work.

References

  • 1.Siegel RL, Miller KD, Wagle NS, Jemal A. Cancer statistics, 2023. CA Cancer J Clin. 2023;73:17–48. 10.3322/caac.21763. [DOI] [PubMed] [Google Scholar]
  • 2.Zheng R, Zhang S, Sun K, et al. Cancer statistics in China, 2016. Chin J Oncol. 2023;45:212–20. 10.3760/cma.j.cn112152-20220922-00647. [DOI] [PubMed] [Google Scholar]
  • 3.Siegel RL, Miller KD, Jemal A. Cancer statistics, 2020. CA Cancer J Clin. 2020;70:7–30. 10.3322/caac.21590. [DOI] [PubMed] [Google Scholar]
  • 4.De Koning HJ, Van Der Aalst CM, De Jong PA, et al. Reduced Lung-Cancer Mortality with Volume CT Screening in a Randomized Trial. N Engl J Med. 2020;382:503–13. 10.1056/NEJMoa1911793. [DOI] [PubMed] [Google Scholar]
  • 5.Rami-Porta R, Bolejack V, Crowley J, et al. The IASLC Lung Cancer Staging Project: Proposals for the Revisions of the T Descriptors in the Forthcoming Eighth Edition of the TNM Classification for Lung Cancer. J Thorac Oncol. 2015;10:990–1003. 10.1097/JTO.0000000000000559. [DOI] [PubMed] [Google Scholar]
  • 6.Board WW. classification of tumours. Thoracic Tumours (M). Lyon, France: IARC; 2021. [Google Scholar]
  • 7.Travis WD, Brambilla E, Burke AP, et al. WHO classification of tumours of the lung, pleura. Thymus Heart. 2015;4:78–9. [DOI] [PubMed] [Google Scholar]
  • 8.Goldstraw P, Chansky K, Crowley J, et al. The IASLC Lung Cancer Staging Project: Proposals for Revision of the TNM Stage Groupings in the Forthcoming (Eighth) Edition of the TNM Classification for Lung Cancer. J Thorac Oncol. 2016;11:39–51. 10.1016/j.jtho.2015.09.009. [DOI] [PubMed] [Google Scholar]
  • 9.Al-Jabbar M, Alshahrani M, Senan EM, Ahmed IA. Histopathological analysis for detecting lung and colon cancer malignancies using hybrid systems with fused features. Bioengineering. 2023;10:383. 10.3390/bioengineering10030383. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Al-Jabbar M, Alshahrani M, Senan EM, Ahmed IA. Analyzing histological images using hybrid techniques for early detection of multi-class breast cancer based on fusion features of CNN and handcrafted. Diagnostics. 2023;13:1753. 10.3390/diagnostics13101753. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Ahmed IA, Senan EM, Shatnawi HSA, et al. Multi-models of analyzing dermoscopy images for early detection of multi-class skin lesions based on fused features. Processes. 2023;11:910. 10.3390/pr11030910. [Google Scholar]
  • 12.Alshahrani M, Al-Jabbar M, Senan EM, et al. Analysis of dermoscopy images of multi-class for early detection of skin lesions by hybrid systems based on integrating features of CNN models. PLoS ONE. 2024;19:e0298305. 10.1371/journal.pone.0298305. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Alshahrani M, Al-Jabbar M, Senan EM, et al. Hybrid methods for fundus image analysis for diagnosis of diabetic retinopathy development stages based on fusion features. Diagnostics. 2023;13:2783. 10.3390/diagnostics13172783. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Shamsan A, Senan EM, Ahmad Shatnawi HS. Predicting of diabetic retinopathy development stages of fundus images using deep learning based on combined features. PLoS ONE. 2023;18:e0289555. 10.1371/journal.pone.0289555. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Zhao W, Zhang W, Sun Y, et al. Convolution kernel and iterative reconstruction affect the diagnostic performance of radiomics and deep learning in lung adenocarcinoma pathological subtypes. Thorac Cancer. 2019;10:1893–903. 10.1111/1759-7714.13161. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Asiri Y, Senan EM, Halawani HT, et al. Analyzing histopathological images using fused CNN features based on the geometric active contour method for early diagnosis of lung and colon cancer. Discov Oncol. 2025;16:2036. 10.1007/s12672-025-03907-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Jin Z. Expert consensus on the rule and quality control of pulmonary nodule annotation based on thoracic CT. Chin J Radiol. 2019;4:9–15. [Google Scholar]
  • 18.Gong J, Liu J, Hao W, et al. A deep residual learning network for predicting lung adenocarcinoma manifesting as ground-glass nodule on CT images. Eur Radiol. 2020;30:1847–55. 10.1007/s00330-019-06533-w. [DOI] [PubMed] [Google Scholar]
  • 19.Wang D, Zhang T, Li M, et al. 3D deep learning based classification of pulmonary ground glass opacity nodules with automatic segmentation. Comput Med Imaging Graph. 2021;88:101814. 10.1016/j.compmedimag.2020.101814. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Yanagawa M, Niioka H, Hata A, et al. Application of deep learning (3-dimensional convolutional neural network) for the prediction of pathological invasiveness in lung adenocarcinoma: A preliminary study. Medicine. 2019;98:e16119. 10.1097/MD.0000000000016119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.He L, Huang Y, Ma Z, et al. Effects of contrast-enhancement, reconstruction slice thickness and convolution kernel on the diagnostic performance of radiomics signature in solitary pulmonary nodule. Sci Rep. 2016;6:34921. 10.1038/srep34921. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Zhao W, Yang J, Sun Y, et al. 3D Deep Learning from CT Scans Predicts Tumor Invasiveness of Subcentimeter Pulmonary Adenocarcinomas. Cancer Res. 2018;78:6881–9. 10.1158/0008-5472.CAN-18-0696. [DOI] [PubMed] [Google Scholar]
  • 23.Huang H, Zheng D, Chen H, et al. Fusion of CT images and clinical variables based on deep learning for predicting invasiveness risk of stage I lung adenocarcinoma. Med Phys. 2022;49:6384–94. 10.1002/mp.15903. [DOI] [PubMed] [Google Scholar]
  • 24.Wang X, Li Q, Cai J, et al. Predicting the invasiveness of lung adenocarcinomas appearing as ground-glass nodule on CT scan using multi-task learning and deep radiomics. Transl Lung Cancer Res. 2020;9:1397–406. 10.21037/tlcr-20-370. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The datasets used and/or analyzed during the current study are available from the corresponding author upon reasonable request.


Articles from BMC Medical Imaging are provided here courtesy of BMC

RESOURCES