Abstract
Lymphoma histopathological diagnosis is complex due to rare subtypes, morphological overlaps, and poor tumor differentiation. In this paper, an AI-based system using deep transfer learning and simulated federated learning is developed to classify two lymphoma types i.e. Chronic Lymphocytic Leukemia (CLL) and Follicular Lymphoma (FL) from a dataset of 4500 histopathological images. Six models (VGG-16, VGG-19, MobileNetV2, ResNet50, DenseNet161, and Inception V3) were evaluated across four data thresholds (0.05 to 0.2). These models used fine-tuned convolutional layers to automatically extract high-level image features relevant to tissue morphology; the extracted features were processed internally through each model’s classifier, forming an end-to-end classification pipeline. DenseNet161 achieved the best classification performance across thresholds, while Inception V3 showed the highest accuracy (97.5%) and lowest RMSE (0.393) in the testing phase using deep learning. A simulated federated learning setup was also explored, where Inception V3 again outperformed other models, indicating its robustness in decentralized learning scenarios. The reported evaluation metrics loss, accuracy, precision, RMSE, F1 score, and recall, are derived from the testing phase, ensuring an accurate assessment of generalization performance. The findings highlight the efficacy of deep transfer learning in early and accurate lymphoma detection, with Inception V3 and DenseNet161 demonstrating strong performance across both learning paradigms. However, since federated learning was not fully deployed in a real-world distributed environment, its broader applicability remains a subject for future exploration.
Keywords: Lymphoma cancer, Histopathological diagnosis, Malignant lymphoma, Deep transfer learning, Leukemia, Simulated federated learning, DenseNet161
Subject terms: Diseases, Health care, Oncology, Risk factors, Signs and symptoms
Introduction
The human body has eleven organ systems in total; one of which is the Lymphatic system. A Lymphatic system comprises Lymph nodes, the Thoracic Duct Thymus gland, Spleen, and, Red Bone Marrow that collectively contributes to forming the germ-fighting network by housing White Blood Cells (WBC’s) or Lymphocytes required in protecting the human body from several infections1. It is responsible for transporting clean fluids with essential nutrients back to the body plus removing the excess fluids and wreckage from the tissues of the body. Lymphoma is the cancer of the lymph nodes that occurs when the lymphocytes multiply uncontrollably. It can spread to various tissues and organs throughout the body very swiftly. There are mainly two types of Lymphoma, namely, Non-Hodgkin Lymphoma and Hodgkin Lymphoma2.
Hodgkin lymphoma is defined as the cancer of the immune system that tends to move from one lymph node to the adjacent one. Although it is said to be a curable disease where the success rate reaches 80%, the toxicity released using the treatment remains a problem [oki, mottok]3. Non-Hodgkin Lymphoma is caused by the development of B and T lymphocytes in the lymph nodes or tissues and contributes 95% of the total Lymphoma cases4. The tumour growth may not affect every lymph node in this case. Mostly, the NHL patients may or may not have any precise symptoms that include fever, night sweating, fatigue, loss of weight, etc. As NHL can affect any body organ, there can be a wide range of symptoms plus other conditions. Hence, the diagnosis holds a lot of significance in the treatment5. As stated in GLOBOCAN 2020, the estimated number of Hodgkin lymphoma and non-hodgkin lymphoma cases in 2020 were recorded as 83,087 and 544,3526. In fact, they can be further categorized into two types: (a) Classical Hodgkin lymphoma - It is an aggressive form with 95% cases and four variants: Mixed Cellularity (MCHL), Modular Sclerosis (NSHL), Lymphocyte rich (LRCHL), and, Lymphocyte depleted (LDHL). Also, it is an extremely treatable disease; (b) Non-Classical Hodgkin Lymphoma - It is an indolent version that is managed with radiations and has only one variant known as Modular lymphocyte-predominant7.
However, histopathological diagnosis of lymphoma remains a highly challenging task. Subtle morphological differences between subtypes, overlapping features, and the lack of consistent slide preparation across laboratories often hinder reliable interpretation. Moreover, manual examination by pathologists is labor-intensive, time-consuming, and prone to subjectivity. The limited availability of annotated medical image data further complicates the training and evaluation of robust diagnostic models.
Artificial Intelligence techniques, particularly deep and transfer learning, have shown significant promise in automating medical diagnostics by effectively handling complex image interpretation tasks and aiding in histopathology analysis8. Neural network architecture is commonly used to accomplish deep learning and federated learning. Deep learning-based diagnostic systems have recently been applied to develop automated methods for interpreting histopathology images. Digital microscopy can uncover unique morphological traits and increase histopathology interpretation performance when paired with deep learning algorithms9. Visual object recognition is aided by deep learning in computing. Deep learning technologies have demonstrated superior performance in complex diagnostic tasks and can effectively assist pathologists in identifying diseases through medical image analysis10.
Although deep learning methods have shown promise in cancer detection, their performance can vary significantly across architectures, and their ability to generalize under limited training data remains questionable. Moreover, in clinical settings, patient data often cannot be centrally aggregated due to privacy regulations, prompting interest in federated learning frameworks. However, few studies have directly compared the efficacy of centralized versus federated learning in the context of histopathology, especially using multiple architectures. This paper addresses these gaps by evaluating and comparing six state-of-the-art deep transfer learning models such as VGG-16, VGG-19, ResNet50, DenseNet161, Inception V3, and MobileNetV2 across both centralized (deep learning) and decentralized (federated learning) training paradigms. By doing so, it aims to (1) assess which architectures are most suitable for this domain, (2) examine how performance scales with varying data availability via threshold sampling, and (3) explore the trade-offs in model effectiveness, privacy, and scalability. The proposed work employs an integration of deep transfer and federated learning techniques for the detection of Lymphoma The key contributions of this study are as follows;
Development of a robust AI-based diagnostic framework for classifying two lymphoma subtypes Chronic Lymphocytic Leukemia (CLL) and Follicular Lymphoma (FL) from histopathological images using deep transfer learning models fine-tuned on a curated dataset of 5,400 images.
Comparative evaluation of six advanced CNN architectures (VGG-16, VGG-19, MobileNetV2, ResNet50, DenseNet161, and Inception V3) across four data thresholds (0.05 to 0.2), enabling a systematic analysis of how input variability impacts classification performance.
Implementation of feature-based input enrichment, where contour- and intensity-level features such as area, perimeter, mean color, and aspect ratio were extracted and directly used as input to the deep learning pipelines, enhancing model sensitivity to morphological variations.
Simulation of a federated learning environment to assess how decentralized model training performs compared to centralized training, offering insights into the applicability of privacy-preserving AI in histopathology without direct patient data sharing.
Qualitative and visual error analysis using example predictions (TP, FP, FN) from real histopathological images to interpret classification decisions, identify model limitations, and support clinical transparency.
Related work
Diagnosing lymphoma is not always easy. Different subtypes look similar under a microscope, and doctors often need extra lab tests. But now, deep learning is a powerful type of artificial intelligence which is helping researchers build tools to make diagnosis faster and more accurate.
Kumar et al. (2025)11, who studied many existing deep learning methods used to detect lymphoma. They didn’t build a model themselves, but they carefully compared different approaches, datasets, and results. Their work showed the strengths and weaknesses of different models and gave helpful suggestions for future research. Then came Perry et al. (2023)12, who focused on detecting rare and aggressive types of lymphoma called double-hit and triple-hit lymphomas (DHL/THL). They built a model called the DHL-classifier using 57 biopsy slides. Their model achieved 100% sensitivity, 87% specificity, and an AUC of 0.95, with only two false positives. This was better than the usual FISH tests used in hospitals and showed that AI can help catch serious lymphoma cases early. Ferjaoui et al. (2024)13 took a different path by combining traditional image features (like texture and shape) with a deep learning model called BiLSTM. They tested it on a medical image dataset (LWBDWMRI) and compared it to simpler models like LSTM and VGG16. Their model achieved 96% accuracy, 97% F1-score, 98% recall, and had a fast execution time of 26.11 s. This made it useful for real-time decisions during treatment. Next, Tong et al. (2023)14 worked on predicting how patients with B-cell lymphoma would respond to CAR T-cell therapy. They analyzed 770 lesions from 39 patients using CT and PET scans before treatment. Their deep learning model, combined with a rule-based method, reached 82% lesion-level accuracy and AUC of 0.91. At the patient level, accuracy was 81%, much better than traditional scoring methods like IPI, which only had 54% accuracy. Aoki et al. (2025)15 took on a much larger problem: predicting who might develop chronic lymphocytic leukemia (CLL) years before symptoms appear. They used data from over 1 million patients, focusing on regular lab tests like white blood cell counts. Using a random forest model, they achieved an AUC of 0.92, showing that AI can help with early cancer warning from simple lab data. In the field of blood cancer diagnosis, Hagar et al. (2023)16 built two deep learning models to classify eight types of blood cancers, including CLL, FL, and MCL. They compared VGG16 and DenseNet-121. VGG16 performed better, reaching an impressive 98.2% accuracy, making it the preferred model for this multi-class task. Özgür & Saygılı (2024)17 designed a support system for diagnosing lymphoma using histopathological images. They extracted features using GLCM and reduced data size using PCA. Then they used both traditional ML models (like KNN, random forest) and deep learning (like VGG16, ResNet50, and DenseNet201). DenseNet201 gave the best result in binary classification (94% accuracy for CLL vs. FL). In the harder triple classification task, the highest accuracy was 82%, and the lowest (with KNN) was just 36%, showing that some models struggled with similar-looking subtypes. Back to deep learning models, Kumar et al. (2024)18 trained the popular VGG16 model on lymphoma subtypes using Kaggle histopathological images. With proper preprocessing and data augmentation, their model achieved 96.19% validation accuracy and 95.20% test accuracy. They also used a confusion matrix to find areas where the model needed improvement. Aly et al. (2024)19 made a stronger model by using DenseNet201 for extracting the features and Dense Neural Network (DNN) for classification. To improve results further, they applied Harris Hawks Optimization (HHO). Trained on 15,000 biopsy images, their model achieved 99.33% test accuracy, with excellent values for precision, recall, F1-score, and ROC-AUC—making it not only powerful but also explainable and ready for clinical use. Finally, Rajadurai et al. (2024)20 built both individual and ensemble models to classify CLL, FL, and MCL. They used pre-trained models like VGG16, VGG19, DenseNet201, InceptionV3, and Xception, then combined the best ones (InceptionV3 + Xception) into a two-level ensemble. This final model reached a peak accuracy of 99% over 300 training epochs, showing that carefully combined models can give the best results.
Table 1 provides an overview of recent state-of-the-art studies in lymphoma classification and related domains, highlighting key aspects such as dataset size, model type, performance metrics, strengths, and limitations for each work. This comparison helps position the current study within the broader research landscape and underscores its methodological contributions.
Table 1.
Summary of existing work on lymphoma classification using medical imaging and deep learning techniques.
| Author, year | Dataset size | Model highlights | Main rvaluation metrics | Distinct strengths | Limitations |
|---|---|---|---|---|---|
| Kumar et al11., 2025 | Multiple public datasets (not individually listed) | Survey of SOTA deep learning methods | - | Offers structured review, highlights gaps and future directions | No experimental results; insights based on secondary data |
| Perry et al12., 2023 | 57 biopsies (32 train, 25 val; rare DHL/THL cases) | Custom CNN (DHL-classifier) | AUC = 0.95, Sensitivity = 100%, Specificity = 87%, 2 false positives | AI can replace costly genetic screening; accurate on small, rare classes | Dataset size limited; no external validation |
| Ferjaoui et al13., 2024 | LWBDWMRI dataset (MRI-based) | Hybrid model (histogram + texture + BiLSTM) | Accuracy = 96%, F1-score = 97%, Recall = 98%, Time = 26.11s | Performs better than baseline models; multi-feature combination | Not tested on other datasets; lacks generalization study |
| Tong et al14., 2023 | 770 lesions from 39 patients (PET/CT scans) | Lesion-level DL + rule-based aggregation | Lesion Accuracy = 0.82, AUC = 0.91, Patient Accuracy = 0.81 | Outperforms IPI (accuracy = 0.54); enables early therapy decision | Small patient set; no multi-center or external validation |
| Aoki et al15., 2025 | > 1 million patients over 7 years | Random Forest Survival Model | AUC = 0.92 | Early prediction using only lab values; scalable across healthcare networks | No imaging data used; interpretability of model decisions is limited |
| Hagar et al16., 2023 | Microscopy images (number not specified) | CNNs: VGG16 and DenseNet-121 | VGG16 Accuracy = 98.2% | Multi-class capability; CNN comparison across architectures | Dataset size not disclosed; lacks external testing |
| Özgür & Saygılı17, 2024 | Histopathology images (CLL, FL, MCL) | ML (KNN, RF) + DL (VGG16, ResNet50, DenseNet201) | Binary Accuracy (DenseNet201) = 94%, Triple Accuracy = 82%, Lowest = 36% (KNN) | Shows strength of feature extraction + DL; good binary results | MCL classification weak; KNN performs poorly; no interpretability analysis |
| Kumar et al18., 2024 | Kaggle histopathology dataset | VGG16 with preprocessing and augmentation | Validation Accuracy = 96.19%, Test Accuracy = 95.20% | Good accuracy with simple architecture; misclassification tracked via confusion matrix | Limited to 3 lymphoma subtypes; no ensemble or external data |
| Aly et al19., 2024 | 15,000 biopsy images (CLL, FL, MCL) | DenseNet201 + Dense Neural Network + HHO | Test Accuracy = 99.33%, High Precision, Recall, F1, ROC-AUC | High interpretability; strong clinical potential | Does not report inference speed; lacks deployment-level analysis |
| Rajadurai et al20., 2024 | Multiclass histopathology dataset | Ensemble of InceptionV3 + Xception | Ensemble Accuracy = 99%, Individual Models > 90% | Highest accuracy among all; robust ensemble framework | Preprocessing and augmentation steps not described in detail |
Research methodology
In this section, the description regarding each phase such as collection of data, preprocessing of images, extraction of features, training the models, and the parameters used to evaluate the models in the identification and classification of lymphoma cancer has been mentioned.
The presented research work employs a novel methodology that utilizes different transfer learning models along with federated learning and deep learning models to detect lymphoma from an image and to classify its type. Initially, a set of pre-processing techniques are applied to enhance the input images. Further, exploratory data analysis is performed followed by extraction of significant features. Several data augmentation techniques are applied before categorizing training and testing datasets. Finally, an already trained transfer learning model is chosen to perform the classification procedure as illustrated in Fig. 1.
Fig. 1.
Proposed method for detecting and classifying lymphoma cancer.
Dataset description
The proposed work uses a dataset named “Malignant Lymphoma Classification”. There are images of lymphoma cancer available, as well as specimens generated by pathologists at various locations. In this work, we have worked on two dataset of lymphoma cancer i.e. Chronic Lymphocytic Leukemia (CLL) and Follicular Lymphoma (FL), as shown in Fig. 2. This dataset has the capability of identifying types of lymphoma from biopsies that are sectioned plus tinted with Hematoxylin/Eosin (H + E) as well as aids in extra compatibility and less challenging disease diagnosis21. Some of the dataset features are as follows:
Fig. 2.
Lymphoma samples taken from malignant lymphoma classification (a) CLL (b) FL.
4500 total images of size: 1388 × 1400;
The images are provided in.tiff format;
Images are in a standard RGB color space.
The table provides a breakdown of the dataset used to train and test a model for lymphoma subtype classification. It includes two classes: Follicular Lymphoma (FL) and Chronic Lymphocytic Leukemia (CLL). For CLL, a total of 2,500 histopathological images were available, of which 2,000 (80%) were used to train the model and the remaining 500 (20%) were reserved to test. Similarly, the FL class had 2,000 images in total, with 1,600 used for training and 400 for testing. Altogether, the dataset consisted of 4,500 images, with 3,600 used to train the model and 900 for evaluating its performance. This 80 − 20 split ensures that the model is trained on enough examples while retaining enough data to fairly assess its accuracy and generalizability.
Table 2 presents the class-wise distribution of images used in this paper, showing the total number of samples for each lymphoma subtype and their division into training and testing sets using an 80 to 20 split.
Table 2.
Class-wise distribution of histopathological images for CLL and FL.
| Class | Total images | Training set (80%) | Testing set (20%) |
|---|---|---|---|
| Chronic Lymphocytic Leukemia (CLL) | 2,500 | 2,000 | 500 |
| Follicular Lymphoma (FL) | 2,000 | 1,600 | 400 |
| Total | 4,500 | 3,600 | 900 |
Data pre-processing
This step is performed to clean and load the dataset images into the system. The main focus in this research work is based on two classes or two diseases of Lymphoma, namely, Chronic Lymphocytic Leukemia and Follicular Lymphoma. The number of images per disease is evaluated, 2500 for CLL and 2000 for FL as depicted in Fig. 3.
Fig. 3.
Sample dataset images.
Exploratory data analysis
This step is mainly executed to examine and study the input dataset mainly for their critical characteristics by employing several techniques of data visualization. The proposed work utilizes RGB Histograms for every image of both categories to represent the distribution of intensities graphically. Also, it enhances the image contrast as it spreads out the most frequent values of intensity. One image of each category with its corresponding histograms is shown in Fig. 4.
Fig. 4.
Input images (a) CLL and (b) FL with their corresponding histograms.
Extraction of features
The proposed system extracts the features to reduce and remove unnecessary information in the dataset and extract the most significant features required for the efficient performance of the system. Several contour-based and geometric features (Table 3) were computed from the histopathological images of Chronic Lymphocytic Leukemia (CLL) and Follicular Lymphoma (FL)22. These features included Epsilon, Area, Width, Height, Perimeter, Aspect Ratio, Extent, Mean Color Intensity, and Extreme Points, among others. These were designed to capture essential morphological traits of lymphoma cells and tissues, such as size, shape, and pixel intensity distribution. These combined features were then passed through the models’ fully connected layers for classification, enabling the models to learn from both raw image textures and structured morphological descriptors. This hybrid input strategy enriched the feature space and contributed to the models’ improved performance across various thresholds.
Table 3.
Values generated for contour features.
| Parameters | CLL | FL |
|---|---|---|
| Parameters | 0.0 | 0.0 |
| Area | 4.0 | 4.0 |
| Perimeter | 3 | 3 |
| Width | 1 | 1 |
| Height | 963,1039 | 1275,1039 |
| Extreme Rightmost Point | 3.0 | 3.0 |
| Aspect Ratio | 0.0 | 0.0 |
| Extent | 0.4 | 0.4 |
| Epsilon | 961,1039 | 1273,1039 |
| Extreme Topmost Point | 137.0 | 142.0 |
| Maximum Value | 963,1039 | 1275,1039 |
| Minimum Value Location | 961,1039 | 1273,1035 |
| Extreme Leftmost Point | 133.0 | 137.0 |
| Mean Color/Intensity | 0.0 | 0.0 |
| Equivalent Diameter | 129.0 | 128.0 |
| Minimum Value | 962,1039 | 1274,1039 |
| Maximum Value Location | 961,1039 | 1273,1039 |
Data augmentation
Data augmentation was applied to enhance the diversity of the training dataset and to reduce overfitting by generating transformed variants of existing images. Several augmentation techniques were employed, including horizontal flipping, zooming (± 15%), height and width shifts, rotation (up to 5 degrees), shearing, and contrast enhancement through normalization. These augmentations aimed to simulate real-world variations in histopathological imaging and improve the models’ generalization capabilities.
In addition to general augmentation, specific attention was given to addressing class imbalance between Chronic Lymphocytic Leukemia (CLL) and Follicular Lymphoma (FL). Since the dataset contains 2,500 CLL images and 2,000 FL images, a moderate imbalance existed (Fig. 5). To mitigate this:
Fig. 5.
Results of data augmentation (a) original images (b) augmented images.
Targeted data augmentation was applied more heavily to the minority class (FL) to increase its representation in the training set and reduce class imbalance. Each image in the training set underwent 4 to 5 augmentation transformations, resulting in an overall increase to approximately 13,500 training instances. This approach enhanced dataset diversity and improved the model’s ability to generalize.
During training, no oversampling or undersampling techniques were applied at the dataset level; instead, augmentation-based oversampling was used.
Additionally, the loss function was kept standard across experiments (e.g., categorical cross-entropy), and no class weighting was applied, given that the class imbalance was relatively mild and well-addressed through augmentation.
This approach ensured that both classes were adequately represented during training, without introducing bias or compromising the integrity of the dataset. The effects of this strategy were reflected in balanced precision and recall scores across both classes, as shown in the classification reports.
Model selection
In this section, the models that have been trained with the lymphoma dataset are briefly described and later evaluated with the various parameters to examine their performances. The selection of models is done on the basis of the nature as well as size of the dataset, computational resources, and time complexity. Besides this, it has been also found that these applied models have showed great performance in the wide range of dataset23–28.
VGG16 and VGG19 are known for their deep, sequential architecture using uniform 3 × 3 convolutional filters and max-pooling layers. These models were chosen as classical CNN baselines due to their widespread adoption in biomedical imaging. However, their lack of shortcut connections and high parameter count make them computationally intensive and prone to overfitting, particularly when training data is limited or contains high intra-class similarity as is the case with histopathological lymphoma images. While they achieved moderate performance in this study, their architecture restricts their ability to capture subtle textural variations between lymphoma subtypes. The architectural representation of both the networks is presented as:
VGG16
Covolutional layer
![]() |
![]() |
![]() |
![]() |
Fully connected layer
![]() |
VGG19
Covolutional layer
![]() |
![]() |
![]() |
![]() |
![]() |
Fully connected layer
![]() |
Here
covolutional layer with 3 × 3 kernel and 64 filters,
max pooling layer,
fully connected layer with 4096 neurons,
softmax activations function.
ResNet50 is a deep convolutional neural network that introduces residual or shortcut connections, allowing the network to skip certain layers by directly passing information forward. This architecture helps prevent issues like vanishing gradients, which commonly affect deeper networks during training. As a result, ResNet50 can effectively learn and retain both low-level and high-level features across its 50 layers. In the context of lymphoma classification, this ability is particularly important because histopathological images contain complex and subtle tissue patterns. ResNet50’s structure enables it to capture these fine-grained features more accurately, leading to improved performance metrics such as recall and F1-score. Its depth and design make it especially suited for medical image analysis where detail and precision are critical. Mathematically, it is represented by equations:
Basic residual block
![]() |
1 |
The residual function typically consists of
![]() |
2 |
![]() |
3 |
Here
()= residual function, BN= Batch Normalization. The architecture of ResNet50 is presented as:
Initial Convolution and max pooling
![]() |
Intermediate stages
![]() |
Final layers
![]() |
Inception V3, developed by Google, is a 48-layer deep convolutional neural network known for its architectural efficiency and strong performance in image classification. It incorporates advanced techniques such as inception modules, batch normalization, label smoothing, and the RMSProp optimizer, all aimed at reducing overfitting and improving convergence speed. Its hallmark design feature is the use of multi-scale convolutional filters in parallel, which allows the model to extract features at various resolutions simultaneously. This is particularly advantageous in lymphoma detection, where cell sizes and tissue textures vary widely across histopathological slides. Inception V3’s ability to efficiently process high-resolution medical images enables it to capture both local and global structures relevant to distinguishing between lymphoma subtypes. As observed in this study, its adaptability to data heterogeneity contributed to strong validation performance, especially in federated learning settings, making it a robust choice for accurate and scalable cancer detection. Its architecture can be represented as.
Initial layers
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
Inception modules
![]() |
Auxiliary classifier
where
![]() |
![]() |
Final layers
![]() |
MobileNet V2 is a lightweight architecture designed for deployment on mobile and embedded devices. It incorporates inverted residuals with linear bottlenecks and depthwise separable convolutions, significantly reducing model size and computation without heavily compromising accuracy. MobileNetV2 was selected for its efficiency and real-time applicability in low-resource environments. In this study, while it achieved respectable performance, its compactness may limit its ability to learn highly abstract or detailed features, which are essential for distinguishing between morphologically similar lymphoma subtypes such as CLL and FL. The architectural design of MobileNetV2 network is shown as.
Initial Convolution
![]() |
Sequence of inverted residual blocks
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
Pooling
![]() |
Fully connected layer
![]() |
DenseNet161 is a deep convolutional neural network comprising 161 layers and is well-regarded for its high accuracy in complex image classification tasks. Its core strength lies in the use of dense connectivity, where each layer receives feature maps from all preceding layers within a block. This design promotes feature reuse, efficient gradient flow, and reduced redundancy, enabling the network to learn richer and more diverse representations. In the context of lymphoma detection, where subtle morphological differences between subtypes like CLL and FL are critical, DenseNet161’s architecture allows it to capture fine-grained patterns in histopathological images with greater precision. By leveraging both high- and low-level features throughout the network, it enhances the ability of the model to distinguish between visually similar lymphoma types. This results in improved diagnostic accuracy, particularly in data-limited or noisy environments, making DenseNet161 one of the most effective models in this study.
Initial Convolution and pooling
![]() |
Dense block 1
![]() |
With 6 layers.
For each l layer in Dense Block
![]() |
.
Transition layer 1
![]() |
Dense block 2
![]() |
With 12 layers.
Transition layer 2
![]() |
Dense block 3
![]() |
With 36 layers.
Final layers
![]() |
Results
In this section, the performance of the applied deep learning and federated learning models is evaluated, and their results are computed for various disease categories, or threshold values. In addition, the models have been examined during the training as well as validation phases and are also analyzed graphically. Furthermore, the trained models are also assessed to determine their prediction accuracy.
Classification report
Table 4 represents the values of several evaluation parameters such as F1 score, precision, and, recall corresponding to various transfer learning models based on the two categories: CLL and FL. To assess model generalization under limited data, training was conducted using four threshold levels (0.05, 0.1, 0.15, 0.2), corresponding to 25%, 50%, 75%, and 100% of the dataset. These thresholds were chosen based on preliminary experimentation and help evaluate model robustness across varying data volumes.
Table 4.
Parameter evaluation for threshold value = 0.05.
| Models | Threshold Value | CLL | FL | ||||
|---|---|---|---|---|---|---|---|
|
|
|
|
|
|
||
| VGG-16 | 0.05 | 0.98 | 0.82 | 0.89 | 0.84 | 0.99 | 0.91 |
| MobileNet V2 | 0.99 | 0.62 | 0.76 | 0.72 | 0.99 | 0.84 | |
| ResNet-50 | 0.98 | 0.76 | 0.86 | 0.81 | 0.98 | 0.89 | |
| DenseNet-161 | 0.99 | 0.90 | 0.94 | 0.91 | 0.99 | 0.95 | |
| Inception V3 | 0.95 | 0.62 | 0.76 | 0.72 | 0.95 | 0.80 | |
| VGG-19 | 0.97 | 0.86 | 0.91 | 0.88 | 0.97 | 0.92 | |
| DenseNet-161 | 0.1 | 0.98 | 0.91 | 0.95 | 0.92 | 0.98 | 0.95 |
| VGG-19 | 0.96 | 0.89 | 0.92 | 0.90 | 0.96 | 0.93 | |
| VGG-16 | 0.98 | 0.85 | 0.91 | 0.87 | 0.98 | 0.92 | |
| ResNet-50 | 0.97 | 0.83 | 0.89 | 0.85 | 0.97 | 0.91 | |
| Inception V3 | 0.95 | 0.75 | 0.82 | 0.77 | 0.95 | 0.85 | |
| MobileNet V2 | 0.98 | 0.71 | 0.82 | 0.77 | 0.99 | 0.82 | |
|
0.15 | 0.98 | 0.87 | 0.92 | 0.88 | 0.98 | 0.93 |
|
0.98 | 0.76 | 0.86 | 0.80 | 0.98 | 0.88 | |
|
0.96 | 0.87 | 0.91 | 0.88 | 0.96 | 0.92 | |
|
0.97 | 0.92 | 0.94 | 0.95 | 0.97 | 0.95 | |
|
0.97 | 0.78 | 0.86 | 0.80 | 0.98 | 0.85 | |
|
0.95 | 0.91 | 0.93 | 0.91 | 0.96 | 0.93 | |
| VGG-16 | 0.20 | 0.97 | 0.88 | 0.93 | 0.89 | 0.98 | 0.93 |
| DenseNet-161 | 0.97 | 0.92 | 0.95 | 0.93 | 0.97 | 0.95 | |
|
0.95 | 0.92 | 0.93 | 0.92 | 0.95 | 0.94 | |
|
0.95 | 0.89 | 0.92 | 0.90 | 0.95 | 0.92 | |
|
0.97 | 0.76 | 0.86 | 0.80 | 0.98 | 0.88 | |
|
0.97 | 0.76 | 0.86 | 0.80 | 0.98 | 0.88 | |
For threshold value 0.05, DenseNet161 showed good performance for the CLL class of the Lymphoma dataset, achieving precision, recall, and F1 score values of 0.99, 0.90, and 0.94 respectively for 25% of the data. The model performed well on the FL class of lymphoma data, achieving precision, recall, and F1 scores of 0.91, 0.99, and 0.95, respectively. However, when comparing both classes, it has been observed that densenet161 performed better with the FL class. This implies that the model is better at recognizing FL cases compared to CLL cases in the dataset provided.
Similarly, in the cases where 50%, 75%, and 100% of the data was used, only DenseNet161 showed good performance when trained with images from both classes. As per the threshold value 0.1, DenseNet161 obtained the best precision value of 98%, followed by VGG16 and MobileNetV2. However, when it comes to recall as well as F1 score, only DenseNet161 performed the top with 91% recall and 95% F1 score specifically for the CLL class. But, when it comes to the FL class, DenseNet161 performed well in case of precision as well as F1 score only. However, for recall, MobileNet V2 outperformed all other models with a 99% score.
Based on the threshold value of 0.15, all of the models perform differently on the CLL (Chronic Lymphocytic Leukemia) and FL (Follicular Lymphoma) classes of the dataset, highlighting their unique strengths. VGG16 and MobileNetV2 had the best precision (98%) for CLL, meaning they were very good at reducing false positives. On the other hand, DenseNet161 performed well in recall and F1 score, which means it was effective at identifying CLL cases while maintaining a good balance. However, when it comes to the FL class, DenseNet161 showed better precision and F1 score. This demonstrates its capability to accurately identify FL cases. On the other hand, VGG16, MobileNetV2, and InceptionV3 showed impressive ability to accurately identify a large number of real FL cases by scoring the highest recall value.
At the end, for threshold value 0.20 the models showcased the same performance as shown in the FL class of 75% of the dataset but in case of CLL class, most of the models such as Vgg-16, MobileNetV2, DenseNet161, and Inception V3 did well by computing the best precision value of 0.97. But, on the other hand, DenseNet161 did a great execution of data and generated the highest recall and F1 score value of 0.92 and 0.95 respectively.
In a nutshell, from Tables 3, 4, 5 and 6, DenseNet-161 model has the highest average precision and recall rate for every threshold value when compared to all other models. Thus, it has the best F1 score as well. Hence, the classification report implies that this model classifies the disease accurately and efficiently.
Table 5.
Evaluation metrics for deep Learning.
| Models | Training | Validation | ||||
|---|---|---|---|---|---|---|
| Accuracy | Loss | RMSE | Accuracy | Loss | RMSE | |
| VGG-16 | 98.5 | 0.082 | 0.286 | 96.5 | 0.216 | 0.464 |
| MobileNet V2 | 94.5 | 0.199 | 0.446 | 92.5 | 0.255 | 0.504 |
| ResNet50 | 94.5 | 0.230 | 0.479 | 96.5 | 0.272 | 0.521 |
| DenseNet-161 | 98.2 | 0.081 | 0.284 | 95.6 | 0.183 | 0.427 |
| Inception V3 | 95.5 | 0.119 | 0.344 | 97.5 | 0.155 | 0.393 |
| VGG-19 | 96.5 | 0.123 | 0.350 | 94.5 | 0.213 | 0.461 |
Table 6.
Evaluation metrics for federated learning.
| Models |
|
|
||||
|---|---|---|---|---|---|---|
|
|
|
|
|
|
|
|
96.5 | 0.182 | 0.426 | 92 | 0.316 | 0.562 |
|
92.5 | 0.299 | 0.546 | 90.6 | 0.455 | 0.674 |
| ResNet50 | 96.5 | 0.130 | 0.360 | 95.6 | 0.172 | 0.414 |
|
95.6 | 0.113 | 0.336 | 91.0 | 0.298 | 0.545 |
| Inception V3 | 97.5 | 0.029 | 0.170 | 95.8 | 0.055 | 0.234 |
| VGG-19 | 94.5 | 0.223 | 0.472 | 92.4 | 0.313 | 0.559 |
[*bold value refers to the best values obtained by the model].
Performance evaluation of the applied models
Table 5 shows the best values of the models for the evaluation metrics namely; accuracy, loss, and, RMSE value for both training and validation. It can be observed that the DenseNet-161 model proved to be the best as it achieved an accuracy of 98.2% with a 0.081 loss value and RMSE value of 0.284. The validation phase observed the highest accuracy using the Inception V3 model (97.5%) and minimum loss value of 0.155. Also, the RMSE value of 0.393 is the least for Inception V3 model. Hence, Inception V3 model is the best for validation purpose.
Likewise, Table 6 shows the average best results in terms of their performance when evaluated using federated learning techniques. For training, Inception V3 model proved to be the best as it achieved 97.5% accuracy, a minimum loss value of 0.029, and, least RMSE value (0.170) when compared to other transfer learning models. Also, the Inception V3 model again proved to be the best for validation purposes as it observed the highest accuracy (95.6%) and a minimum loss value of 0.055. The least RMSE value of 0.234 is also achieved by Inception V3 model. Hence, the Inception V3 model is the most suitable when using federated learning techniques.
In Fig. 6, overfitting is demonstrated by the training and validation loss of VGG16, DenseNet161, and VGG19. It only happens when the training loss plot continues to drop with experience, and the validation loss plot falls to a point before increasing again. Because the plot of validation loss drops to the end of stability and has a small gap with the training loss, the training and validation loss curves displayed by ResNet 50, Inception V3, and MobileNetV2 are considerably closer to being described as perfect fit learning curves. For both illnesses, follicular lymphoma and chronic lymphocytic leukemia, the training accuracy curve is greater than the validation accuracy curve, indicating that it is easier to predict training dataset as compared to validation dataset.
Fig. 6.
Analysis of models based on their accuracy and loss values.
In Fig. 7, the confusion matrices illustrate the performance of Inception V3 and ResNet-50 in classifying two lymphoma subtypes: Chronic Lymphocytic Leukemia (CLL) and Follicular Lymphoma (FL). Inception V3 correctly classified 2,150 CLL cases but misclassified 350 as FL. For FL, it correctly identified 1,934 cases, with 66 misclassified as CLL. In contrast, ResNet-50 demonstrated slightly better accuracy for CLL, correctly classifying 2,225 cases and misclassifying 275 as FL. However, it showed a marginally lower performance on FL, with 1,883 correct classifications and 117 FL cases incorrectly predicted as CLL. Overall, Inception V3 was more effective at correctly identifying FL, while ResNet-50 performed better in accurately classifying CLL. These results highlight complementary strengths in the two models’ classification capabilities.
Fig. 7.
Confusion matrix of models.
In Fig. 8, the graphs labelled as “a” belongs to chronic lymphocytic leukemia (CLL) while as “b” belongs to follicular lymphoma (FL). These graphs are based on the ROC, precision, F1score, and recall parameters of the models such as VGG19, VGG16, ResNet50, MobileNetV2, Inception V3, and DenseNet161 for each individual type of diseases as mentioned earlier. It has been calculated that the ROC values obtained by VGG16 is 0.9880, MobileNet V2 is 0.9809, ResNet50 is 0.9820, DenseNet161 is 0.9920, and Inception V3 is 0.9820. In addition to this, it has been also observed that there is a constant relationship in between F1 score and the threshold values.
Fig. 8.
Examination of deep learning models for different parameters.
Besides this, the computational time to train each applied models under deep transfer learning and federated learning is also computed as shown in Table 7.
Table 7.
Computational time of models.
| Techniques | Computational Time (DL) | Computational Time (FL) |
|---|---|---|
| Inception V3 | 1 h 40 min | 2 h |
| MobileNet V2 | 1 h 46 min | 3 h 1 min |
| ResNet50 | 3 h | 1 h 40 min |
| DenseNet-161 | 1 h 50 min | 2 h 1 min |
| VGG-16 | 2 h | 2 h 5 min |
| VGG-19 | 2 h 2 min | 3 h 20 min |
The table provides a comparative overview of the computational time required for training several deep learning models using two different techniques: Deep Learning (DL) and Federated Learning (FL). In DL, which involves centralized training on a single, large dataset, the training times for the models are as follows:
To examine the time of the models under deep learning techniques, it has been observed that the maximum hours have been taken by ResNet50 (3 h) followed by VGG 19 and VGG16. Inception V3 is trained in the least time with 1 h 40 min and DenseNet161 also takes the time nearby to it with 1 h 50 min.
In the next scenario, the federated learning showed variation in the training times for the same models. Here VGG16 takes 2 h and 5 min which is little longer for the same model when applied under deep learning. Similarly, MobileNetV2 also trained the layers in much longer time of 3 h and 1 min as compared to its performance as deep learning model. VGG-19 also sees a notable increase in training time in FL, requiring 3 h and 20 min. On the contrary, ResNet50, surprisingly, trains faster in FL, taking only 1 h and 40 min.
To provide valuable qualitative insight into model performance and errors, Fig. 9 presents six histopathological image samples in which some correctly and others incorrectly classified are accompanied by prediction accuracies from various deep transfer and federated learning models. In Fig. 9(a), a Follicular Lymphoma (FL) case was correctly classified using VGG16, achieving 98% accuracy in deep learning and 99% in federated learning, representing a true positive (TP). Similarly, Fig. 9(b) shows another correctly classified FL image by MobileNet V2, with 100% accuracy under both paradigms, again a true positive. However, Fig. 9(c) presents a misclassified image of Chronic Lymphocytic Leukemia (CLL) predicted incorrectly by ResNet50 as FL, resulting in 85% accuracy represents a false negative (FN) for CLL, likely due to overlapping nuclear morphology. In Fig. 9(d), an FL sample was correctly predicted by DenseNet161 with 95% accuracy, a true positive, supported by its dense connectivity and strong feature reuse. Figure 9(e), which shows a CLL image, was misclassified as FL by Inception V3 in deep learning mode (95% accuracy) but corrected under federated learning (99% accuracy), suggesting a false positive (FP) in the deep learning setup. Finally, Fig. 9(f) depicts a correctly classified CLL image by VGG19, achieving 100% accuracy in deep learning and 98% in federated learning leads to another true positive. These examples highlight the importance of visual factors such as tissue contrast, cellular boundaries, and structural overlap in influencing both correct and incorrect classifications. Notably, false positives and false negatives occurred more frequently in CLL cases, underscoring the greater morphological ambiguity associated with this subtype compared to FL.
Fig. 9.
Prediction accuracy of lymphoma cancer images.
Discussion
The comparative performance analysis between deep transfer learning and federated learning models reveals nuanced insights into their applicability for lymphoma subtype classification. DenseNet161 emerged as the most robust model in the centralized deep learning setup, consistently delivering the highest F1-scores and precision-recall values across all data thresholds. It achieved a training accuracy of 98.2% and minimal RMSE, underscoring the effectiveness of dense feature reuse in capturing subtle morphological patterns within histopathological images. In contrast, Inception V3 exhibited superior performance in the validation phase, particularly under the federated learning paradigm, attaining 97.5% accuracy and the lowest RMSE (0.234). Its architectural ability to process multi-scale features appears advantageous in decentralized settings, where inter-institutional data heterogeneity and limited communication cycles often hinder convergence.
The findings also highlight computational trade-offs. Although MobileNetV2 is optimized for low-resource environments, it required over three hours to converge under federated learning and yielded lower accuracy than DenseNet161 or Inception V3. This suggests that MobileNetV2 may struggle with high-resolution medical image complexity when communication and synchronization overheads are involved. When benchmarked against state-of-the-art studies such as Aly et al. (2024) and Rajadurai et al. (2024), who employed ensemble or hybrid-optimized CNN architectures reporting accuracies around 99%, the present study achieves competitive results using individual models. Moreover, unlike prior works that evaluated performance using a single data threshold or architecture, this study presents a granular evaluation across multiple thresholds and models under both centralized and decentralized training regimes—adding significant methodological depth. Qualitative error analysis through confusion matrices and prediction visualizations further identified a higher rate of false positives and false negatives in CLL cases, consistent with clinical literature noting greater morphological heterogeneity in CLL than in FL. This finding reinforces the need for integrating multi-modal inputs (e.g., genomics or immunophenotyping) and incorporating explainable AI (XAI) techniques for interpretability and clinical trust. From a deployment perspective, centralized deep learning offers better accuracy and training efficiency but raises privacy concerns, particularly in healthcare, where data sharing is restricted by frameworks like HIPAA and GDPR28. Federated learning, by allowing local training and global model aggregation, addresses these concerns but introduces challenges like slower convergence, communication overhead, and model inconsistency across clients.
Conclusion
This section is divided into three parts to provide a comprehensive overview of the study. The first part presents a summary of the manuscript, outlining the key methods and findings. The second part highlights the limitations of the current work, pointing out areas where performance or applicability may be restricted. The third part discusses the future scope, suggesting possible improvements and directions for continued research in lymphoma classification using deep learning.
Summary
This paper proposed and evaluated an AI-based system for the classification of two lymphoma subtypes—Chronic Lymphocytic Leukemia (CLL) and Follicular Lymphoma (FL)—using deep transfer learning and federated learning frameworks on histopathological images. Six pre-trained convolutional neural networks (VGG16, VGG19, ResNet50, Inception V3, DenseNet161, and MobileNet V2) were fine-tuned and tested across multiple data thresholds. Among these, DenseNet161 consistently exhibited strong classification performance, while Inception V3 and MobileNet V2 achieved the highest accuracy rates of up to 100% in specific settings. Evaluation metrics such as accuracy, precision, recall, F1-score, loss, and RMSE were derived from the testing phase to ensure generalization assessment. The experimental results affirm the feasibility of using deep learning for early and accurate detection of lymphoma, highlighting the diagnostic potential of transfer learning and the added robustness from federated learning.
Limitations
Despite the promising results, the study has several limitations. The dataset used, although carefully curated and balanced between CLL and FL, may not capture the full diversity of histopathological variations found in real-world clinical settings. The system was evaluated only on two lymphoma subtypes; hence, the generalizability of the models to other types of hematological malignancies or solid tumors remains untested. Additionally, federated learning was explored primarily for performance benchmarking, and not fully deployed in a cross-institutional environment, limiting assessment of privacy-preserving benefits. Training deep models on high-resolution images also imposed significant computational requirements, which could pose challenges for deployment in resource-constrained settings. Optimization difficulties, particularly in model convergence and tuning across thresholds, were also encountered and warrant further exploration.
Future work
Several directions can be pursued to build upon this research. One promising avenue is the integration of multi-modal data, such as combining histopathological images with genomic profiles, clinical parameters, or immunohistochemical markers, to construct a more comprehensive and context-aware diagnostic framework. Future studies can also explore the deployment of lightweight, real-time models optimized for edge devices, enabling diagnostic support in remote or resource-constrained environments. In terms of data privacy and collaboration, implementing federated learning across multiple healthcare institutions is a critical next step toward practical, privacy-preserving AI deployment. Another key direction involves incorporating explainable AI (XAI) methods to interpret and visualize the decision-making process of deep learning models. Techniques such as Grad-CAM, SHAP, and LIME can help highlight regions of interest in histopathological images that contribute most to the model’s predictions. This interpretability is vital for improving clinical trust, model transparency, and regulatory acceptance, especially in high-stakes medical diagnostics. Finally, future work should expand the classification framework to include additional lymphoma subtypes, rare variants, or other cancer types, and evaluate the models on larger, more diverse datasets to enhance generalizability and clinical applicability.
Author contributions
AU and B.K. originally drafted the paper and performed the visualization and investigation, whereas P.J. provided supervision in the research, validated the findings, and contributed significantly to reviewing and editing the manuscript. Y.K. performed the experimental analysis of the research work and also analyzed the results for cancer detection. A.K. contributed to interpreting the results and preparing the result graphs and figures. All the authors equally contributed to the proofreading.
Funding
Open access funding provided by Symbiosis International (Deemed University).
Data availability
The datasets generated and/or analysed during the current study are available in the Kaggle repository and can be accessed via the following link: https://www.kaggle.com/datasets/andrewmvd/malignant-lymphoma-classification/data.
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Boes, K. M. & Durham, A. C. Bone marrow, blood cells, and the lymphoid/lymphatic system. Pathologic Basis Veterinary Disease, 9 (724), (2017).
- 2.Armitage, J. O., Gascoyne, R. D., Lunning, M. A. & Cavalli, F. Non-hodgkin lymphoma. Lancet390 (10091), 298–310 (2017). [DOI] [PubMed] [Google Scholar]
- 3.Flerlage, J. E. et al. Pediatric hodgkin lymphoma, version 3.2021. J. Natl. Compr. Canc. Netw.19 (6), 733–754 (2021). [DOI] [PubMed] [Google Scholar]
- 4.Francischetti, I. M. B., Alejo, J. C., Sivanandham, R., Davies-Hill, T., Fetsch, P., Pandrea, I., Jaffe, E. S., & Pittaluga, S. Neutrophil and eosinophil extracellular traps in Hodgkin lymphoma. HemaSphere, 5(9), e633, 1–10, (2021). [DOI] [PMC free article] [PubMed]
- 5.Kumar, Y., Gupta, S., Singla, R., & Hu, Y. A Systematic review of artificial intelligence techniques in cancer prediction and diagnosis. Archives of Computational Methods in Engineering, 29(4), 2043–2070, (2021). [DOI] [PMC free article] [PubMed]
- 6.Sung, H. et al. Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. Cancer J. Clin.71 (3), 209–249 (2021). [DOI] [PubMed] [Google Scholar]
- 7.Kattak, M. S. et al. Clinico-histological presentation of head and neck lesions in a tertiary care hospital. J. Rawalpindi Med. Coll.25 (1), 21–25 (2021). [Google Scholar]
- 8.Kumar, Y. & Singla, R. Federated learning systems for healthcare: perspective and recent progress. In M. Habib ur Rehman & M. M. Gaber (Eds.), Federated Learning Systems: Towards Next-Generation AI 141–156. Springer, 2021. [Google Scholar]
- 9.Komura, D. & Ishikawa, S. Machine learning methods for histopathological image analysis. Comput. Struct. Biotechnol. J.16, 34–42 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Bohr, A. & Memarzadeh, K. The rise of artificial intelligence in healthcare applications. In A. Bohr & K. Memarzadeh (Eds.), Artificial Intelligence in Healthcare 25–60. Academic Press, 2020.
- 11.HN, N. K. et al. Comprehensive study on lymphoma detection using deep learning. In 2025 3rd International Conference on Inventive Computing and Informatics (ICICI) (pp. 923–927). IEEE. (2025), June.
- 12.Perry, C., Greenberg, O., Haberman, S., Herskovitz, N., Gazy, I., Avinoam, A., … Avivi,I. (2023). Image-based deep learning detection of high-grade B-cell lymphomas directly from hematoxylin and eosin images. Cancers, 15(21), 5205. [DOI] [PMC free article] [PubMed]
- 13.Ferjaoui, R., Boujnah, S. & Khalifa, A. B. A novel handcrafted features and deep bilstm neural network for lymphoma recognition. In 2024 10th International Conference on Control, Decision and Information Technologies (CoDIT) (pp. 2266–2271). IEEE. (2024), July.
- 14.Tong, Y., Udupa, J. K., Chong, E., Winchell, N., Sun, C., Zou, Y., … Torigian, D.A. (2023). Prediction of lymphoma response to CAR T cells by deep learning-based image analysis. PLoS One, 18(7), e0282573. [DOI] [PMC free article] [PubMed]
- 15.Aoki, J., Khalid, O., Kaya, C. & Salama, M. E. Machine learning model predicts abnormal lymphocytosis associated with chronic lymphocytic leukemia. JCO Clin. Cancer Inf.9, e2400197 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Hagar, M., Elsheref, F. K. & Kamal, S. R. A new model for blood cancer classification based on deep learning techniques. International J. Adv. Comput. Sci. Applications, 14(6), 422–429, (2023).
- 17.Özgür, E. & Saygılı, A. A new approach for automatic classification of non-hodgkin lymphoma using deep learning and classical learning methods on histopathological images. Neural Comput. Appl.36 (32), 20537–20560 (2024). [Google Scholar]
- 18.Kumar, A., Nelson, L. & Arumugam, D. Blood cancer diagnosis using pretrained vgg16 transfer learning model with lymphoma dataset. In 2024 Asian Conference on Intelligent Technologies (ACOIT) (pp. 1–6). IEEE. (2024), September.
- 19.Aly, S. A., Bakhiet, A. & Balat, M. Diagnosis of malignant lymphoma cancer using hybrid optimized techniques based on dense neural networks. In 2024 International Conference on Computer and Applications (ICCA) (pp. 1–6). IEEE. (2024), December.
- 20.Rajadurai, S., Perumal, K., Ijaz, M. F. & Chowdhary, C. L. Precisionlymphonet: advancing malignant lymphoma diagnosis via ensemble transfer learning with Cnns. Diagnostics14 (5), 469 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Orlov, N. et al. Automatic classification of lymphoma images with transform-based global features. IEEE transactions on information technology in biomedicine: a publication of the IEEE engineering in medicine and biology society. 14. 1003–1013. (2010). 10.1109/TITB.2010.2050695 [DOI] [PMC free article] [PubMed]
- 22.Koul, A., Bawa, R. K. & Kumar, Y. An analysis of deep transfer learning-based approaches for prediction and prognosis of multiple respiratory diseases using pulmonary images. Arch. Comput. Methods Eng.31 (2), 1023–1049 (2024). [Google Scholar]
- 23.Alnowaiser, K., Saber, A., Hassan, E. & Awad, W. A. An optimized model based on adaptive convolutional neural network and grey Wolf algorithm for breast cancer diagnosis. PloS One, 19(8), e0304868, 1–16, (2024). [DOI] [PMC free article] [PubMed]
- 24.Saber, A., Elbedwehy, S., Awad, W. A. & Hassan, E. An optimized ensemble model based on meta-heuristic algorithms for effective detection and classification of breast tumors. Neural Comput. Appl.37 (6), 4881–4894 (2025). [Google Scholar]
- 25.Elbedwehy, S., Hassan, E., Saber, A. & Elmonier, R. Integrating neural networks with advanced optimization techniques for accurate kidney disease diagnosis. Sci. Rep.14 (1), 21740 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Khushi, H. M. T., Masood, T., Jaffar, A., Rashid, M. & Akram, S. Improved multiclass brain tumor detection via customized pretrained efficientnetb7 model. IEEE Access.11, 117210–117230 (2023). [Google Scholar]
- 27.Valizadeh, P., Jannatdoust, P., Pahlevan-Fallahy, M. T., Hassankhani, A., Amoukhteh,M., Bagherieh, S., … Gholamrezanezhad, A. (2025). Diagnostic accuracy of radiomics and artificial intelligence models in diagnosing lymph node metastasis in head and neck cancers: a systematic review and meta-analysis. Neuroradiology, 67(2), 449–467. [DOI] [PMC free article] [PubMed]
- 28.Ettaloui, N., Arezki, S. & Gadi, T. An overview of blockchain-based electronic health records and compliance with GDPR and HIPAA. Data Metadata. 2, 166–166 (2023). [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The datasets generated and/or analysed during the current study are available in the Kaggle repository and can be accessed via the following link: https://www.kaggle.com/datasets/andrewmvd/malignant-lymphoma-classification/data.


















































































