Abstract
Background
Lung cancer ranks among the most lethal malignancies globally, and its traditional diagnosis suffers from strong subjectivity, high misdiagnosis rates and uneven medical resources. To overcome the poor feature alignment caused by simply concatenating CT images and clinical text, this paper proposes a lung cancer multimodal auxiliary diagnosis model based on entropy weight decision fusion.
Methods
This retrospective cohort study enrolled 5847 participants from 2020 to 2025, comprising 1823 lung cancer cases, 2253 normal controls and 1771 pulmonary nodule controls. All CT images and corresponding reports were analyzed, and three datasets were established via random sampling from the original dataset. The study incorporated Vision Transformer (ViT) and Bidirectional Encoder Representations from Transformers (BERT) as feature extractors for images and text, respectively, to extract high-dimensional semantic features from lung CT images and CT Imaging Report. Secondly, independent classifiers based on Multi-Layer Perceptron (MLP) were established to convert the embedding vectors of different modalities into predicted probability distributions (Logits). Finally, the entropy weight method was employed to adaptively fuse the decision results of images and text. The model performance evaluation indicators include area under the receiver operating characteristic curve (AUC), accuracy, precision, recall, and F1-score.
Results
The proposed method in this study can fully leverage the complementary information from CT images and imaging text multimodality. On the clinical dataset, it achieved an accuracy of 0.9375, a precision of 0.9324, a recall of 0.9322, and an F1-score of 0.9322, significantly improving diagnostic performance.
Conclusions
This study validates that, on a real-world lung cancer dataset, multimodal data decision fusion outperforms unimodal models and common fusion methods in terms of diagnostic accuracy, precision and recall. It provides a potential reference for the early auxiliary diagnosis of pulmonary nodules and lung cancer, and lays a foundation for subsequent clinical applications.
Graphical abstract

Keywords: Vision Transformer (ViT), Bidirectional Encoder Representations from Transformers (BERT), Deep learning, Entropy weight decision fusion, Lung cancer classification
Introduction
In 2022, there were approximately 2.5 million new lung cancer cases and 1.8 million lung cancer-related deaths worldwide. As the cancer with the highest incidence and mortality globally, it has ranked first in global cancer mortality for ten consecutive years. Early detection and intervention are crucial to reducing lung cancer mortality [1, 2]. The surge in lung cancer patients has also driven oncology into an artificial intelligence (AI)-driven era. AI-based technologies that simulate physicians’ diagnostic logic, mine potential information, and thereby reduce workload while improving efficiency have gradually become indispensable [3, 4]. Currently, most studies are limited to the use of single-modal data (such as computed tomography (CT) images, pathological biopsies, and medical text data) or their simple combination for diagnostic research, neglecting the multi-level manifestations of cancer and the correlations among these data. Meanwhile, existing studies fail to fully capture the complementary and differential relationships between multimodal data.
Deep learning has achieved breakthroughs in medical imaging and natural language processing [5]. Pre-trained models such as Vision Transformer (ViT) [6, 7] and Bidirectional Encoder Representations from Transformers (BERT) [8] have demonstrated powerful capabilities in image feature extraction and text semantic understanding. To address the limitations of single-modal data, researchers have proposed multimodal fusion methods to integrate CT images, pathological images, genetic data, and clinical data [9, 10]. Among these, entropy-weighted decision fusion [11] presents several key advantages in the fusion of medical images (e.g., CT images) and text (e.g., radiology reports). It quantifies diagnostic uncertainty caused by CT artifacts and subjective radiology reports via modal information entropy, and dynamically assigns weights according to sample multimodal quality. This approach perfectly matches the heterogeneous characteristics of medical data and effectively improves the reliability of the fused diagnostic results. These methods leverage pre-trained models to fully represent medical text and CT image data, and utilize effective fusion strategies to exploit the interconnections among multimodal data [12–14].
To tackle the effective fusion of multi-source data in lung cancer diagnosis, this study proposes a multimodal feature extraction framework based on the Transformer architecture. Specifically, the BERT model and ViT model are employed to extract deep features from imaging report texts and CT images, respectively, thereby fully mining disease-related information embedded in different modalities. To address the challenge of dynamic allocation of modal confidence in the multimodal fusion process, a decision-level fusion strategy based on entropy weight is proposed [15]. This strategy calculates the information entropy of the output results of each modality to quantify the uncertainty of different modalities in the diagnostic task for specific samples. It then adaptively assigns decision weights to each modality, realizing an intelligent fusion mechanism of "on-demand selection", which effectively solves the problem of inconsistent reliability of different samples across various modalities. To verify the effectiveness of the proposed model, comparative experiments and performance evaluations were conducted on a real-world clinical dataset of lung cancer. The experimental results demonstrate that the multimodal feature extraction framework combined with the entropy-weighted decision-level fusion strategy achieves an accuracy of 93.75% in lung cancer diagnosis tasks, which is significantly superior to single-modality models and common fusion methods. This validates its application potential in clinical auxiliary diagnosis scenarios.
Results
The experimental data comprise real-world imaging and examination report data from a tertiary hospital in Anhui Province, including not only data with obvious lung cancer lesion features but also data of pulmonary nodules and normal lung tissues. The datasets used in this study were all randomly sampled from 5,847 samples to form training, test, and validation sets with a ratio of 8:1:1. Within each dataset, the patient samples in the training, test, and validation sets were mutually independent and non-overlapping. Three datasets were obtained by repeating the random sampling process three times. Since each sample corresponds to the information of only one patient, the content and distribution of samples in the training, test, and validation sets differed across the three datasets. The training and validation sets were used for parameter training and constructing a model suitable for this study, while the test set was used for model evaluation.
To comprehensively compare the performance of the proposed method, we conducted comparative predictions using the BERT and ViT models. Meanwhile, we additionally set up a control group without the cross-attention mechanism for comparison, to verify whether introducing cross-attention can improve the diagnostic performance under multimodal feature fusion. Furthermore, to validate the effectiveness of the entropy-weighted fusion strategy in the experiments, we compared it with average fusion and gated fusion methods. The average fusion method obtains the final model prediction by computing the weighted average of the prediction results from text and image modalities. Gated fusion [16] assists the model in selecting and enhancing the most useful features during the fusion process while suppressing noise and irrelevant information, thereby effectively addressing challenges such as potential semantic conflicts, information redundancy, and modality missing among different modalities.
As shown in Table 1, among the six compared models, the proposed method achieves optimal accuracy and precision on three independent datasets in a lung disease prediction task with over 5,000 image-text samples, demonstrating outstanding classification performance.
Table 1.
Experimental comparison results of the proposed method with ViT, gated, averaging and BERT models
| Accuracy | Precision | Recall | F1-score | ||
|---|---|---|---|---|---|
| Dataset_1 | Proposed method | 0.9362 | 0.9309 | 0.9309 | 0.9308 |
| Without attention | 0.921 | 0.9165 | 0.9154 | 0.9157 | |
| ViT | 0.9105 | 0.9034 | 0.9032 | 0.903 | |
| BERT | 0.894 | 0.8872 | 0.8862 | 0.8866 | |
| Gated | 0.9271 | 0.9227 | 0.9253 | 0.9245 | |
| averaging | 0.9145 | 0.9137 | 0.9112 | 0.9158 | |
| Dataset_2 | Proposed method | 0.9362 | 0.9313 | 0.9306 | 0.9307 |
| Without attention | 0.9216 | 0.9166 | 0.9161 | 0.9159 | |
| ViT | 0.8979 | 0.8992 | 0.8978 | 0.8973 | |
| BERT | 0.9028 | 0.9034 | 0.9029 | 0.9026 | |
| Gated | 0.9214 | 0.9235 | 0.9246 | 0.9252 | |
| Averaging | 0.9179 | 0.9141 | 0.9115 | 0.9196 | |
| Dataset_3 | Proposed method | 0.9402 | 0.9351 | 0.9352 | 0.9351 |
| Without attention | 0.9222 | 0.9165 | 0.9165 | 0.9165 | |
| ViT | 0.9003 | 0.8948 | 0.8924 | 0.8917 | |
| BERT | 0.8912 | 0.8858 | 0.8832 | 0.8831 | |
| Gated | 0.9288 | 0.9223 | 0.9211 | 0.9246 | |
| Averaging | 0.9168 | 0.9104 | 0.9119 | 0.9166 |
Based on the prediction results, we calculated and plotted the ROC curves and corresponding AUC values. The results demonstrate that the proposed method achieves the highest AUC value, indicating that this model has important diagnostic value for the prediction of lung cancer and pulmonary nodules (Fig. 1).
Fig. 1.

Comparison of ROC curves and AUC metrics between the proposed method and without attention, ViT, gated, averaging, BERT models
To further investigate the prediction accuracy of the model for various diseases, we aggregated the test set data from three experiments, including data from 676 normal patients, 532 patients with pulmonary nodules, and 547 patients with lung cancer and conducted an in-depth analysis. The analysis results are presented in Table 2. It can be observed that both the proposed method, the Gated model, and the ViT model achieved 100% prediction recall on normal lung images. In addition, the entropy-weighted decision fusion method exhibited higher precision than both the gated fusion model and the averaging fusion model in the individual prediction of lung cancer, pulmonary nodules, and normal lung status.
Table 2.
Comparison of correct prediction counts, recall and specificity indicators between the proposed method and ViT, gated, averaging, and BERT models for each disease
| Lung cancer (547) | Pulmonary nodules (532) | Normal patients (676) | |||||
|---|---|---|---|---|---|---|---|
| Recall | Proposed method | 501 | 91.59% | 484 | 90.97% | 676 | 100% |
| ViT | 439 | 80.26% | 474 | 89.10% | 676 | 100% | |
| BERT | 459 | 83.91% | 472 | 88.72% | 664 | 98.22% | |
| Gated | 482 | 88.12% | 479 | 90.03% | 676 | 100% | |
| Averaging | 476 | 87.02% | 476 | 89.47% | 674 | 99.7% | |
| Specificity | Proposed method | 96.03% | 96.24% | 100% | |||
| ViT | 95.2% | 91.17% | 100% | ||||
| BERT | 94.04% | 91.82% | 98.24% | ||||
| Gated | 95.89% | 92.86% | 99.46% | ||||
| Averaging | 95.58% | 92.43% | 98.75% | ||||
To verify the statistical significance of performance differences among different models in the core three-class classification scenario, this study decomposes the three-class task into three binary classification subtasks: "lung cancer vs non-lung cancer", "pulmonary nodules vs non-pulmonary nodules", and "normal controls vs non-normal controls". The DeLong test is employed to conduct pairwise comparisons of the AUC values of each model. Herein, a P-value < 0.05 indicates that the difference in AUC between two models is statistically significant (i.e., a notable performance difference), while a P-value ≥ 0.05 indicates no significant difference. The results are presented in Table 3.
Table 3.
DeLong test results
| Proposed method | Averaging | Gated | ViT | BERT | |
|---|---|---|---|---|---|
| Lung cancer versus non-lung cancer | |||||
| Proposed method | 1 | 0.84 | 0.042 | 0 | 0 |
| ViT | 0.84 | 1 | 0.031 | 0 | 0 |
| BERT | 0.042 | 0.031 | 1 | 0.007 | 0.004 |
| Gated | 0 | 0 | 0.007 | 1 | 0.947 |
| Averaging | 0 | 0 | 0.004 | 0.947 | 1 |
| Pulmonary nodules versus non-pulmonary nodules | |||||
| Proposed method | 1 | 0.96 | 0.314 | 0 | 0.027 |
| ViT | 0.96 | 1 | 0.305 | 0 | 0.038 |
| BERT | 0.314 | 0.305 | 1 | 0 | 0.004 |
| Gated | 0 | 0 | 0 | 1 | 0.072 |
| Averaging | 0.027 | 0.038 | 0.004 | 0.072 | 1 |
| Normal controls versus non-normal controls | |||||
| Proposed method | 1 | 1 | 1 | 0 | 1 |
| ViT | 1 | 1 | 1 | 0 | 1 |
| BERT | 1 | 1 | 1 | 0 | 1 |
| Gated | 0 | 0 | 0 | 1 | 0 |
| Averaging | 1 | 1 | 1 | 0 | 1 |
In this study, t-distributed Stochastic Neighbor Embedding (t-SNE) was employed to project high-dimensional features into a low-dimensional space for classification visualization. The results verified that the proposed model can map raw high-dimensional data into a well-discriminated feature space for samples from normal controls, pulmonary nodules and lung cancer cases. As shown in Fig. 2, our method achieved 100% classification accuracy for normal lung images, with distinct feature boundaries clearly separating normal samples from pulmonary nodule and lung cancer samples. Nevertheless, lung cancer and pulmonary nodules exhibit considerable similarities in CT morphological manifestations and textual descriptions in clinical reports. Both lesions commonly present as round or quasi-round focal hyperdense shadows with shallow lobulation and short spiculation, meanwhile, overlapping descriptive terms such as focal spiculation and mild pleural traction are frequently used in imaging reports. These inherent similarities and feature overlaps between the two categories are also consistent with the partial mixing distribution of lung cancer and pulmonary nodule samples observed in the t-SNE visualization.
Fig. 2.

Visualization results of the model on image and text data
Discussion
This study aims to explore the application of deep learning models in the identification of lung cancer, pulmonary nodules and normal controls, and to analyze their effectiveness in diagnostic accuracy evaluation. The primary clinical goal is the early detection of lung cancer, that is, distinguishing malignant lesions from benign or healthy conditions. Accordingly, we define the core task as a three-class classification problem over samples of lung cancer, pulmonary nodules and normal controls. Meanwhile, considering clinical practicality, we also report the performance metrics under the binary classification scenario of lung cancer vs. non-lung cancer.
With the rapid development of artificial intelligence technology, research on the application of deep learning-based models such as BERT and ViT in the medical field has emerged one after another, initially demonstrating these technologies’ excellent ability to understand and process complex medical information. Chaudhari et al. [17] introduced the BERT model into the field of typo detection in imaging reports, and the pre-trained BERT model exhibited outstanding performance in identifying phonetic typos. For the description of "differentiated adenocarcinoma with vascular invasion" in pathological reports, researchers found that the BERT model can resolve the semantic correlations of key terms such as "differentiation" and "vascular invasion", and locate the corresponding abnormal vascular structures in CT images through attention weights [18]. BioBERT, a domain-adaptive pre-trained model fine-tuned using massive medical literature and clinical records, can accurately identify descriptions of complex biomarkers such as "EGFR mutation" and "PD-L1 expression", further improving the accuracy of text-image alignment [19]. Kumar et al. [20] explored a framework with Vision Transformer (ViT) as the core, where ViT layers are frozen as a feature extractor and a dedicated classifier head is added for lung cancer image classification. To establish robust generalization ability and balance local feature extraction with global context modeling, Singh [7] fused EfficientNetB0 with Vision Transformer (ViT) through a feature concatenation mechanism, utilizing local morphological clues and broader tissue-level information to achieve more discriminative lung cancer classification.
To better utilize multimodal data, the adoption of decision fusion approaches [21–23] enables each modal sub-model to undergo optimal training and parameter tuning for its respective modality, without the need to consider the data distribution of other modalities during training. Sathvik et al. [11] proposed an entropy-based weighted summation classification method to fuse the classification results of each unimodal modality for predicting the disease course of Alzheimer’s disease (AD). This reduces model coupling complexity, and when multimodal data are missing or of poor quality, the output of sub-models for other modalities can still be used to generate results, thereby reducing overall errors and improving generalization performance [24, 25].
As shown in Table 1, the global experimental metrics of accuracy, precision, recall and F1-score are 0.9375, 0.9324, 0.9322 and 0.9322, respectively, demonstrating that the model achieves satisfactory classification performance on lung cancer, pulmonary nodules and the normal control group, and presents more stable and superior performance compared with using the BERT or ViT model alone. In dataset sequence 3, the proposed method attains a high accuracy of 0.9402, which is 5% higher than that of the BERT model and 4% higher than that of the ViT model. The underlying reason is that the proposed method enables each modality to focus on its own key features via the self-attention mechanism, and then captures complementary information from the other modality through the cross-attention mechanism. Comparison with the other two fusion methods reveals that the entropy-weighted fusion strategy outperforms average fusion by 2.1% and gated fusion by 1%, demonstrating that the entropy-weighted fusion strategy can generate more reliable and stable decision results. The model constructed in this study achieves an F1-score of 0.9351 on the test set, reflecting high accuracy in disease diagnosis. As shown in Fig. 1, the AUC for diagnostic accuracy of the proposed method also achieves the optimal performance across the three datasets. Comparative experiments with and without the cross-attention mechanism show that model performance is improved by 2%. The cross-attention mechanism calculates the correlation weight between multimodal features, adaptively establishes the correspondence between lesion areas in CT images and semantic descriptions in clinical texts, and achieves precise cross-modal feature alignment. It enables the model to capture key lesion-related features and effectively improves the diagnostic accuracy of lung diseases.
To clarify the predictive performance of the model across different patient categories (normal patients, pulmonary nodules, and lung cancer), Table 2 presents the summarized experimental results for lung cancer, pulmonary nodules, and normal controls based on the three test sets. It can be observed that the proposed method achieves the optimal prediction performance for all three categories. Specifically, the recall rate of the proposed method for lung cancer prediction reaches 91.59%, representing an 8% improvement over the BERT model and an 11% improvement over the ViT model. In terms of specificity, the model achieves values of 96.03%, 96.24%, and 100% for lung cancer, pulmonary nodules, and normal patients, respectively. These results demonstrate that the model has favorable generalization ability.
Multimodal fusion strategies can effectively mitigate the influence of unimodal uncertainty on diagnostic classification performance. As shown in Table 2, all three fusion approaches achieve this effect in single-disease diagnostic tasks. Taking recall on the lung cancer dataset as an example: the entropy weight fusion, gated fusion, and average fusion strategies attain recall values of 91.59%, 88.12%, and 87.02%, respectively. All three schemes surpass the performance of unimodal ViT (80.26%) and BERT (83.91%) models. The fundamental explanation is that multimodal fusion allows the model to exploit complementary information between imaging and textual data, enhances comprehension of complex medical information, better adapts to feature discrepancies across different diseases, and maintains stable model performance.
To verify the statistical significance of AUC differences among various models, this study decomposed the three-class classification task into three independent binary classification scenarios using the DeLong test, with the results presented in Table 3. In the scenario of lung cancer vs non-lung cancer, the difference between the proposed method and the average fusion model (Averaging) is not statistically significant (P = 0.84 > 0.05). In contrast, the difference between the proposed method and the gated fusion model (Gated) is statistically significant (P = 0.042 < 0.05), and the differences between the proposed method and the unimodal models ViT and BERT are highly statistically significant (P = 0 < 0.001). These results indicate that the proposed method can significantly improve the ability to identify lung cancer and is superior to the gated fusion and unimodal models. In the scenario of pulmonary nodules vs non-pulmonary nodules: the differences between the proposed method and Averaging, Gated are not statistically significant (P = 0.96 and 0.314, both > 0.05), suggesting that its performance is consistent with existing fusion methods in the identification of benign lesions. The difference between the proposed method and ViT is highly statistically significant (P = 0 < 0.001), and the difference between the proposed method and BERT is statistically significant (P = 0.027 < 0.05), further confirming the value of multimodal fusion in enhancing the identification of benign lesions. All multimodal fusion models outperform unimodal counterparts significantly, verifying that multimodal information complementarity is critical for improving lung disease identification.
Despite the promising results, this study still has several limitations. First, the data were collected from a single center with a limited sample size, lacking external validation using real data from other hospitals. Future work will conduct multicenter trials and prospective clinical verification to further evaluate model performance. Second, practical clinical scenarios contain more complex and diverse case data rather than standardized encoded information. Further research will integrate additional data types such as genomics and clinical variables to improve predictive accuracy.
Conclusion
This paper proposes an entropy-weighted decision fusion-based multimodal auxiliary diagnosis model for lung cancer. ViT and BERT are used to extract deep features from CT images and clinical texts, respectively, and a MLP-based independent classifier outputs the prediction probability distribution. Self-attention is applied to capture critical features within each modality, and cross-attention is adopted to acquire cross-modal complementary information. Finally, the entropy-weighted fusion strategy adaptively assigns weights to bimodal decisions to obtain reliable and stable diagnostic results. Experimental results show that the accuracy of the proposed model in the classification task is improved from 0.85 (of traditional models) to 0.92, and the F1-score is raised from 0.77 to 0.91, representing increments of 7% and 14%, respectively. It achieves the best performance in predicting lung cancer, pulmonary nodules and the normal control group compared with unimodal models and other fusion models. These results verify that the proposed method can effectively integrate image and text information, and its classification performance outperforms unimodal models and common multimodal fusion methods. It provides a potential reference for the early auxiliary diagnosis of pulmonary nodules and lung cancer, and lays a foundation for subsequent clinical applications. In future research, we will further optimize the model structure and training strategy to improve the overall efficiency and performance of the model. Additionally, we will integrate the research on lung cancer diagnosis based on deep learning methods with experts’ domain knowledge, so as to better apply it to lung cancer diagnosis.
Materials and methods
Data sources
This study covers the period from 2020 to 2025, with data sourced from the medical records of hospitalized patients with lung cancer, pulmonary nodules, and normal patients at Hefei Hospital Affiliated to Anhui Medical University.
Inclusion criteria: patients were eligible if they met the Chinese Medical Doctor Association Clinical Guidelines for Lung Cancer (2025 Edition) and satisfied the following: (1) ‘Discharge Diagnosis’ in EMRs indicated lung nodules or cancer with available CT images; (2) primary cancer was histologically confirmed by biopsy; (3) age ≥ 18 years with comprehensive clinical profiles. Ethical approval was obtained.
Exclusion criteria: patients were excluded if they had: (1) incomplete or unstructured text reports; (2) other concurrent malignancies; (3) suboptimal CT image quality; (4) received prior therapy for the index lesion.
The total sample size was 6,000 patients. After excluding 153 individuals with missing or duplicate key variable data, the final analysis cohort consisted of 5847 patients, including 1,823 lung cancer patients, 2253 normal patients, and 1771 pulmonary nodule patients.
General information
This study is a retrospective data-based research. Since the data have been de-identified, informed consent from patients is not required. All participants involved in the study are clinical researchers from our hospital and researchers from cooperating universities. Patient data include clinical data, consisting of CT images and corresponding textual descriptions in examination reports.
Methods
The proposed method firstly manually filters clinically meaningful CT images and corresponding imaging report texts of lung cancer, pulmonary nodule and normal lung samples from the original dataset. Secondly, a bimodal feature extraction strategy is adopted, utilizing the Vision Transformer (ViT) and Bidirectional Encoder Representations from Transformers (BERT) models to extract deep features from CT images and imaging report texts, respectively. Then, the independent classifier module converts the features of each modality into predicted probability distributions (Logits) through a Multi-Layer Perceptron (MLP). Finally, the entropy weight-based decision fusion module calculates the entropy value of each modality, assigns weights accordingly, and fuses the results to obtain the final diagnostic outcome. To evaluate the prediction performance of the proposed method on the lung cancer dataset, three comparative models were constructed: (1) a model that performs prediction by inputting patients’ imaging report text data into BERT; (2) a model that performs prediction by inputting patients’ CT image data into ViT; (3) models that adopt average fusion and gated fusion to make predictions based on the patients’ imaging report text and CT image data processed by BERT and ViT. The above models were trained and evaluated separately, and their prediction performance was compared to more robustly assess the generalization ability of the proposed model.
Vision Transformer (ViT) and Bidirectional Encoder Representations from Transformers (BERT)
The Vision Transformer (ViT) transforms the lung cancer image classification task into an image patch sequence prediction task, thereby capturing the dependencies within the image. First, lung cancer CT images are preprocessed for size unification and pixel normalization. The processed images are then split into fixed-size non-overlapping patches, which are projected into low-dimensional embedding vectors through linear transformation. A special [CLS] token for classification is prepended to the embedding vector sequence, and positional encoding is integrated to preserve the spatial positional information of the patches. The resulting sequence is then fed into the Transformer encoder, which adaptively learns the global dependencies among different patches via the multi-head self-attention mechanism. Finally, the feature vector of the [CLS] token is extracted and fed into a fully connected layer for mapping to obtain classification results, enabling automatic classification of lung cancer and pulmonary nodule CT images.
Therefore, for an input medical image , it is first reshaped into a sequence of N flattened 2D patches. These patches are then mapped to a d-dimensional embedding space using a linear projection layer, with positional encoding superimposed. To aggregate global semantics, a learnable classification token is inserted at the beginning of the sequence. The encoder output consists of two parts:
| 1 |
where denotes the global visual representation, and denotes the sequence of local patch representations.
BERT (Bidirectional Encoder Representations from Transformers) captures the semantic associations and key medical information in examination report texts through a bidirectional attention mechanism. First, the raw report texts are tokenized and meaningless symbols are removed. Special tokens such as [CLS] (classification token) and [SEP] (sentence separation token) are added in accordance with the requirements of the BERT model. Meanwhile, tokens are mapped to their corresponding word embeddings, which are fused with positional and segment embeddings to form the final input sequence. After the sequence is input into the BERT encoder, the multi-layer bidirectional Transformer adopts multi-head self-attention to parallelly calculate dependencies among all tokens, so as to adaptively extract key lesion-related features from imaging reports. Finally, the global feature vector of the [CLS] token and local feature vectors of specific tokens are used for information extraction and classification of lung cancer and pulmonary nodules. The CT imaging report is tokenized and embedded, then input into the bidirectional Transformer encoder for feature extraction:
| 2 |
represents the global representation of the semantic meaning of the entire sentence, and represents the sequence of local representations at the word level.
Given the inherent semantic gap between medical images and reports, a cross-attention-based directional interaction module is proposed to explore multimodal complementarity. It consists of two symmetric pathways to enable mutual guidance and feature enhancement between the two modalities.
Text-guided visual enhancement. We take the refined visual features as the Query, and the textual features as the Key and Value. Through cross-attention, the visual features can borrow the explicit semantics from the text to eliminate their own ambiguity:
| 3 |
Vision-guided text enhancement. Symmetrically, we take the textual features as the Query to retrieve spatial cues from the visual features, aligning the abstract textual concepts with the concrete visual patterns:
| 4 |
The formula for the self-attention mechanism is as follows:
| 5 |
where , and denote the query, key and value matrices, respectively, and represents the dimension of the key vectors. The interacted feature vectors and are mapped to the category space via separate linear classifiers, respectively,
| 6 |
denotes the unnormalized logits.
Entropy-weighted decision fusion
The entropy-weighted decision fusion method adopts a bimodal model to extract key features from CT images and imaging reports and yield preliminary decisions. It quantifies the reliability of each modality via entropy for adaptive weight assignment, and achieves accurate diagnosis through weighted fusion. Epistemic uncertainty is further introduced to address the inconsistent modal reliability across different samples. The probability distribution is obtained using the Softmax function. Shannon entropy is then employed to quantify the uncertainty:
| 7 |
where is a numerical stability constant. A high entropy value indicates a flat prediction distribution (high uncertainty), while a low entropy value indicates a sharp prediction distribution (high confidence). We dynamically calculate the fusion weight according to the entropy value:
| 8 |
The final fused prediction probability is the weighted sum:
| 9 |
Model establishment and evaluation
This study applies a bimodal feature extraction strategy. Vision Transformer (ViT) and BERT are used to obtain the prediction logits from CT images and imaging reports, respectively. On this basis, the entropy-weighted decision fusion algorithm calculates the entropy of image and text modalities, allocates adaptive weights, and fuses multimodal outputs to obtain final diagnostic results. The experimental workflow consists of data input, preprocessing and dataset partitioning. The model training and evaluation process is illustrated in Fig. 3.
Fig. 3.

Model training and prediction process diagram
Experimental environment: the operating system was Ubuntu 20.04, the GPU was NVIDIA RTX 3090 (24 GB), and the software stack included the Python 3.8 experimental environment with PyTorch 1.8.2.
Model training parameters: the base learning rate was scaled by a factor of 0.1 for the BERT and ViT models, while the remaining components adopted the base learning rate. The model weights were initialized with bert_base_chinese and ViT_B_16_Weights IMAGENET1K_V1, respectively. The optimizer used was AdamW with a weight decay of 1e-3. The batch size was set to 16 and the number of training epochs was 50. An early stopping strategy was employed: training was terminated if the validation loss showed no decrease for 5 consecutive epochs.
The performance evaluation metrics included ROC curves and AUC metrics, accuracy, precision, recall, specificity, and the F1-score. As a core metric, accuracy assesses overall prediction performance by calculating the ratio of true positives and true negatives to the total number of samples. Precision quantifies the model’s ability to correctly identify true positives among all predicted positive cases, thus measuring its capacity to reduce false positive errors. Sensitivity—also referred to as recall or true positive rate—evaluates the model’s proficiency in detecting true positives from all actual positive samples. Specificity measures the ability of a classification model to correctly identify negative cases among all truly negative samples. In contrast, the F1-score acts as a balanced measure of model accuracy that accounts for both false positives and false negatives, and it is calculated as the harmonic mean of precision and recall. All the above metrics are reported as macro-averaged:
| 10 |
| 11 |
| 12 |
| 13 |
| 14 |
where TP: True Positives (correctly predicted positive instances); FP: False Positives (incorrectly predicted positive instances); TN: True Negatives (correctly predicted negative instances); FN: False Negatives (incorrectly predicted negative instances).
Acknowledgements
None.
Author contributions
HX Z: methodology, writing—original draft; YH T: data curation, formal analysis, conceptualization; PP L: supervision; XZ C and WJ F: writing—review and editing. All authors read and approved the final manuscript.
Funding
This study was supported by grants from the National Natural Science Foundation of China (Grant No. 62376085) and the Science and Technology Planning Project of Bengbu Medical College (Grant No. 2024byzd568sk).
Availability of data and materials
Not applicable.
Declarations
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing interests
The authors declare that they have no competing interests.
Footnotes
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Haixiang Zhang and Yuhong Tang have contributed equally to this work.
Contributor Information
Peipei Li, Email: peipeili@hfut.edu.cn.
Weijian Fan, Email: 335019360@qq.com.
References
- 1.Zhao M, Xue G, He B, et al. Integrated multiomics signatures to optimize the accurate diagnosis of lung cancer. Nat Commun. 2025;16(1):84–84. 10.1038/s41467-024-55594-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Yuan X, et al. Systematic review and meta-analysis of artificial intelligence for image-based lung cancer classification and prognostic evaluation. NPJ Precis Oncol. 2025;9(1):300. 10.1038/s41698-025-01095-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Priya A, Bharathi PS. SE-ResNeXt-50-CNN: a deep learning model for lung cancer classification. Appl Soft Comput. 2025;171:112696. 10.1016/j.asoc.2025.112696. [Google Scholar]
- 4.Gulsoy T, Kablan EB. FocalNeXt: a ConvNeXt augmented FocalNet architecture for lung cancer classification from CT-scan images. Expert Syst Appl. 2025. 10.1016/j.eswa.2024.125553. [Google Scholar]
- 5.Guney S. Enhanced lung cancer classification accuracy via hybrid sensor integration and optimized fuzzy logic-based electronic nose. Sensors (Basel). 2025;25(17):5271. 10.3390/s25175271. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Xu H, et al. Vision transformers for computational histopathology. IEEE Rev Biomed Eng. 2023;17:63–79. [DOI] [PubMed] [Google Scholar]
- 7.Singh O. FusionNet-ViT: hybrid deep learning model for lung cancer classification. SN Comput Sci. 2025. 10.1007/s42979-025-04484-2. [Google Scholar]
- 8.Devlin J, Chang MW, Lee K, et al. BERT: Pre-training of deep bidirectional transformers for language understanding. ArXiv, 2019, https://arxiv.org/abs/1810.04805. 10.18653/v1/n19-1423.
- 9.Simone AD, Sansone C. A multimodal deep learning based approach for Alzheimer's disease diagnosis. Lecture Notes in Computer Science, 2024:131–139. 10.1007/978-3-031-51026-7_12.
- 10.Huang Z, et al. A visual–language foundation model for pathology image analysis using medical twitter. Nat Med. 2023;29(9):2307–16. [DOI] [PubMed] [Google Scholar]
- 11.Prabhu SS, Berkebile JA, Rajagopalan N, et al. Multi-modal deep learning models for alzheimer's disease prediction using mri and ehr. In: 2022 IEEE 22nd international conference on bioinformatics and bioengineering (BIBE). IEEE, 2022: 168–173. 10.1109/BIBE55377.2022.00044.
- 12.Suter P, Dazert E, Kuipers J, et al. Multi-omics subtyping of hepatocellular carcinoma patients using a Bayesian network mixture model. PLoS Comput Biol. 2021. 10.1101/2021.12.16.473083. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Angelopoulos N, Chatzipli A, Nangalia J, et al. Bayesian networks elucidate complex genomic landscapes in cancer. Commun Biol. 2022. 10.1038/s42003-022-03243-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Herawan MI, Adriansjah R. Prostate specific antigen level and Gleason score in Indonesian prostate cancer patients. Majalah Kedokteran Bandung. 2024;56(3):209–13. 10.15395/mkb.v56.3571. [Google Scholar]
- 15.Liu L, Wan X, Li J, et al. An improved entropy-weighted topsis method for decision-level fusion evaluation system of multi-source data. Sensors. 2022;22(17):30. 10.3390/s22176391. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Wang C, Nie R, Cao J, et al. IGNFusion: an unsupervised information gate network for multimodal medical image fusion. IEEE J Sel Top Signal Process. 2022;16(4):854–68. [Google Scholar]
- 17.Chaudhari GR, Liu T, Chen TL, Joseph GB, Vella M, Lee YJ, et al. Application of a domain-specific BERT for detection of speech recognition errors in radiology reports. Radiol Artif Intell. 2022;4(4):e210185. 10.1148/ryai.210185. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Lei, H., et al. TCSA: A text-guided cross-view medical semantic alignment framework for adaptive multi-view visual representation learning. In: International symposium on bioinformatics research and applications. Singapore: Springer Nature Singapore, 2023:136–149. 10.1007/978-981-99-7074-2_11.
- 19.Jinhyuk L, Wonjin Y, Sungdong K, et al. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics (Oxford, England). 2020;36(4):1234–40. 10.1093/bioinformatics/btz682. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Kumar A, Mehta R, Reddy RB, et al. Vision transformer based effective model for early detection and classification of lung cancer. SN Comput Sci. 2024;5(7):839–839. 10.1007/s42979-024-03120-9. [Google Scholar]
- 21.Verma A, Yadav AK. FusionNet: dual input feature fusion network with ensemble based filter feature selection for enhanced brain tumor classification. Brain Res. 2025;1852:149507. 10.1016/j.brainres.2025.149507. [DOI] [PubMed] [Google Scholar]
- 22.Richard SS, Benjamin U, Jane S. Multimodal deep learning for biomedical data fusion: a review. Brief Bioinform. 2022. 10.1093/bib/bbab569. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Duan J, Xiong J, Li Y, et al. Deep learning based multimodal biomedical data fusion: an overview and comparative review. Inf Fusion. 2024. 10.1016/j.inffus.2024.102536. [Google Scholar]
- 24.Weihua X, Ke C, D. D W. A novel information fusion method using improved entropy measure in multi-source incomplete interval-valued datasets. Int J Approx Reason. 2024. 10.1016/j.ijar.2023.109081. [Google Scholar]
- 25.Wang X, Zhao Z, Pan D, et al. Deep cross entropy fusion for pulmonary nodule classification based on ultrasound Imagery. Front Oncol. 2025;15:1514779. 10.3389/fonc.2025.1514779. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
Not applicable.
