Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 May 12;16:21718. doi: 10.1038/s41598-026-51236-0

A minimal-net CNN model for an IoT-based brain tumor detection and monitoring system

Md Taimur Ahad 1,3,✉, Bo Song 2, Yan Li 1
PMCID: PMC13358162  PMID: 42120538

Abstract

In this study, we propose M-Net, a lightweight Convolutional Neural Network (L-CNN) designed for real-time brain tumor classification and detection. Large, state-of-the-art (SOTA) CNN models often face deployment challenges on edge devices due to their high computational resource requirements. To address this, M-Net utilizes a sequential architecture with 4 × 4 kernels and progressive filter expansion, balancing computational efficiency and classification performance. The model was trained and tested on three diverse MRI datasets (3-class, 4-class, and 15-class) to evaluate its robustness and generalizability across varying brain tumor modalities. In k-fold cross-validation (k = 5) and on an unseen test set, M-Net achieved an average accuracy of 99%, outperforming SOTA CNNs (96%). The study also integrates explainable AI techniques—LIME, SHAP, and Grad-CAM—to visualize regions of MRI images contributing to the model’s decision-making process. Additionally, a web and mobile application for smart brain tumor management (SBTM) was developed using M-Net, demonstrating real-time processing at 1 FPS, minimal thermal consumption, and low energy usage. The novel contributions of M-Net lie in its efficient architecture and integration with XAI, making it a suitable candidate for clinical deployment in brain tumor diagnosis. With only 495,972 parameters, M-Net offers a practical, deployable solution for real-time tumor classification on edge devices.

Keywords: Brain tumor detection, Convolutional neural network, CNN, Lightweight CNN, SHAP, LIME, GRAD-CAM

Subject terms: Cancer, Computational biology and bioinformatics, Engineering, Health care, Mathematics and computing

Introduction

Brain tumors remain a major clinical challenge because accurate detection and classification are difficult, yet early diagnosis is essential for timely treatment planning and improved patient outcomes1. The problem is exacerbated by the heterogeneous cellular composition of brain tumors, which includes malignant tumor cells, stromal cells, and immune cells that interact dynamically and often unpredictably2. In addition, brain tumors often exhibit complex lesion patterns, variable locations, and small-scale structures that are difficult to identify reliably3,4. As a result, manual detection and classification from magnetic resonance imaging (MRI) are time-consuming and prone to error19. This diagnostic burden is especially important in light of the global disease load, as more than 300,000 brain tumor cases are reported annually worldwide, making brain tumors one of the most pressing concerns for the medical community and a leading cause of cancer mortality among children and adolescents1.

Recent advances in artificial intelligence, particularly convolutional neural networks (CNNs), have significantly improved automated brain tumor detection. Among these developments, lightweight CNNs have attracted growing attention because they aim to reduce computational complexity and memory usage while maintaining task-specific performance6,7. Several studies have already reported lightweight CNN-based approaches for brain tumor classification2–5. More broadly, lightweight CNNs are typically developed through architectural modification, nonlinear activation design, supervision components, regularization mechanisms, and optimization techniques3,4,8–15,100,101. These advances suggest that lightweight CNNs are not only relevant for improving classification accuracy but are also important for enabling practical deployment in real clinical settings where computational resources are limited.

This practical requirement creates a direct link between lightweight CNN design and smart healthcare deployment. Brain tumor patients may experience persistent headaches, seizures, cognitive and behavioral changes, and neurological deficits, all of which demand timely clinical attention16; similarly, medical sources continue to emphasize the seriousness of symptoms and the importance of timely intervention (Tisch Brain Tumor Centre, 2025). At the same time, high-speed internet, sensor development, radio frequency identification (RFID), wireless sensor networks (WSNs), and smart mobile technologies have transformed conventional hospital services into smart healthcare systems. In such settings, a compact CNN model is methodologically important because inference must be performed on edge devices with limited memory and computational capacity. Thus, the shift from CNN design to IoT-based monitoring is not incidental: a lightweight model is a necessary foundation for low-latency, low-power, and accessible diagnostic support in smart brain tumor management systems (SBTM), especially for remote doctors, healthcare providers, and rural patients16.

Despite progress in CNN-based computer-aided diagnosis (CAD), an important gap persists between laboratory performance and real-world clinical deployment. State-of-the-art CNNs often require extensive computational resources, contain many layers and parameters, and are therefore difficult to industrialize for disease detection on edge computers and mobile devices. Recognizing this limitation, studies such as20,21 proposed CNNs intended to improve computational efficiency and address computational complexity. Likewise, improving accessibility for patients and medical practitioners requires CNN-based disease detection systems that can operate beyond highly controlled research environments22. Prior studies have also raised broader concerns regarding clinical translation. Specifically23–26,, and27 observed that although CNNs often achieve high accuracy in laboratory settings, their adoption in clinical workflows remains limited because of insufficient deployment infrastructure and weak integration with hospital IT systems. Earlier28, showed that CNNs trained on X-rays from one hospital may fail to replicate their performance on data from other hospitals, highlighting a lack of generalizability across clinical settings. In a similar vein29,30,, and more recently31,32, emphasized the gap between technical performance and practical usability in CNN-based CAD. Collectively, these findings indicate that progress in brain tumor CAD requires not only accuracy but also lightweight deployability, interpretability, and integration into smart healthcare environments.

Within this broader context, the present study addresses a more specific research problem: how to design a lightweight yet effective CNN architecture that reduces gradient loss and vanishing feature propagation while remaining suitable for practical deployment in MRI-based brain tumor detection. The study is motivated by the need for a balanced architecture in which early layers capture low-level features and deeper layers extract higher-level representations without imposing excessive computational burden. To this end, the proposed M-Net incorporates kernel and filter expansions, along with MaxPooling and dropout layers, to support robust feature learning, regularization, and more stable gradient flow during MRI image training. This design choice is intended to move beyond merely reporting classification accuracy and instead respond to the real deployment constraints that limit many existing CNN-based systems.

This study aims to provide a lightweight CNN consisting of minimal layers for an IoT-based brain tumor detection and monitoring system. The CNN model was built on a sequential CNN rather than spatial exploitation (VGG19), depth-based (ResNet152v2, MobileNetv2), multi-path (DenseNet201), or width-based multi-connection (InceptionV3, Xception). The reason is that a compact sequential CNN is better suited to IoT edge devices than heavier or more complex CNN architectures, as edge platforms operate under strict constraints on memory, computation, energy, and inference time (Sun et al., 2025; Naveen & Kounte, 2025). Recent studies show that edge deployment favors models with low latency, a small memory footprint, and energy-efficient execution over computationally intensive architectures (Naveen & Kounte, 2025; Khaki & Choi, 2025). Empirical results also indicate that lightweight CNN-based designs can maintain strong accuracy while reducing inference latency and resource consumption, and some compact CNN models outperform larger architectures in terms of edge suitability (Khaki & Choi, 2025; Nyakuri et al., 2025). Therefore, when the design goal is real-time inference on resource-constrained IoT hardware, a compact sequential CNN is often a more practical choice than heavier CNN architectures (Sun et al., 2025; Nyakuri et al., 2025).

The main innovation of this study, therefore, lies not simply in proposing another CNN but in developing an M-Net-centered framework for brain tumor detection that is simultaneously lightweight, edge-oriented, explainable, and extensible to smart monitoring. In contrast to well-known lightweight families such as MobileNet, EfficientNet, or lightweight ResNet variants, the proposed approach is centered on a balanced architectural strategy tailored to MRI-based brain tumor categorization and practical healthcare deployment, rather than on generic image classification alone. In addition, the study does not treat classification as an isolated prediction task. Instead, it combines lightweight model design with explainable AI techniques and a smart brain tumor management perspective, thereby addressing the need for transparency, accessibility, and real-time usability. This also clarifies why the proposed classification framework may offer advantages over other lightweight or ensemble deep neural networks for brain tumor categorization: it is developed not only for predictive performance, but also for deployability on constrained devices, interpretability for end users, and integration into connected healthcare settings.

Guided by these motivations and gaps, this research makes the following contributions to brain tumor detection and real-time monitoring:

  1. A lightweight CNN, M-Net, was developed with fewer layers to reliably detect tumors in magnetic resonance imaging (MRI) under limited computing resources.

  2. Along with M-Net, five state-of-the-art CNNs from spatial exploitation (VGG19), depth-based (ResNet152v2), multipath (DenseNet201), width-based multiconnection (ResNext101), and feature map exploitation (SE-ResNet152) methods were applied to investigate the most effective CNN architecture for tumor detection in MRI images.

  3. LIME, SHAP, and Grad-CAM were applied to generate explanations for brain tumor detection and classification. Pixel intensities were analyzed using a statistical method to assess classification performance. This study, therefore, provides insights into the interpretability and transparency of CNNs for brain tumor modalities.

  4. A Streamlit web application and an Android mobile app were developed using M-Net to make advanced diagnostic tools more accessible to patients and medical practitioners.

  5. Finally, an M-Net- and sensor-based smart brain tumor management system (SBTM) is proposed for real-time monitoring. By integrating IoT devices, the research aims to create a more connected and accessible healthcare system.

The rest of the paper is organized as follows: Sect. 2 covers related works on L-CNN-based brain tumor detection. Section 3 describes the research methods, dataset, hardware, and programming, describing the model developed in this study. The results are described in Sect. 4. Section 5 describes the SBTM system. Section 6 discusses this study; finally, Sect. 7 presents the conclusion and outlines future work, including possible research directions.

Related works

Recognizing the importance of early detection and real-time monitoring, numerous studies have used smart IoT devices and deep learning techniques to detect, classify, and predict brain tumors promptly. The study16 used a custom CNN model and two pre-trained models, specifically Inception-v4 and EfficientNet-B4, for the classification of ten brain tumor modalities. Extensive experimentation suggests that the average classification accuracy for CNN, Inception-v4, and EfficientNet-B4 is 97.58%, 99.56%, and 99.76%, respectively. The authors of33 proposed a DL model that integrates VGG16, an attention mechanism, and optimized hyperparameters to classify brain tumors. The proposed model achieves 99% test accuracy and impressive precision and recall figures, and outperforms traditional approaches such as Support Vector Machines (SVM) with Histogram of Oriented Gradients (HOG), Local Binary Pattern (LBP), and Principal Component Analysis (PCA) by a significant margin. The study by34 combined a lightweight parallel depthwise separable CNN (PDSCNN) with a hybrid ridge regression extreme learning machine (RRELM) to accurately classify brain tumors. The proposed framework achieved average precision, recall, and accuracy of 99.35%, 99.30%, and 99.22%, respectively. The study by35 proposed a secure, automated MRI image classification system that integrates chaotic and Arnold encryption techniques with hybrid deep learning models based on VGG16 and a deep neural network (DNN). The proposed system demonstrated a high classification performance under both encryption scenarios. For chaotic encryption, it achieved 93.75% accuracy, 94.38% precision, 93.75% recall, and an F-score of 93.67%. For the Arnold encryption task, the model achieved 94.1% accuracy, 96.9% precision, 94.1% recall, and an F-score of 96.6%.

The literature review suggests that brain tumor detection experiments using lightweight CNNs are mainly conducted by modifying CNN architectures, optimizing hyperparameters, adding or deleting layers (such as convolutional, pooling, concatenation, and fully connected layers), and exploiting the sparsity of kernel and feature maps. The study15 proposed a lightweight CNN model integrated with preprocessing techniques, achieving classification accuracy of 99.58% and a Dice similarity coefficient (DSC) of 95.7%. Similarly36, designed a lightweight CNN for real-time detection, achieving 99.48% accuracy for binary classification and 96.86% for multiclass classification. The importance of lightweight models is further exemplified by5, who achieved a remarkable 97.87% accuracy with MobileNet in a lightweight configuration. The study37 also integrated edge intelligence with lightweight CNNs for hospital data center applications, achieving state-of-the-art accuracy while maintaining practicality for predictive diagnostics. The study38 achieved the best-case accuracy of 97.8%, with sensitivities and specificities of 97.32% and 99.2%, respectively.

Building on advances in transfer learning39, demonstrated the potential of fine-tuning the VGG19 model, combined with synthetic data augmentation, to classify brain tumors using the BraTS 2015 dataset. Similarly40, employed the VGG16 model within a transfer learning-based CNN framework, achieving training and test accuracies of 96.5% and 90%, respectively, highlighting its lower complexity compared with state-of-the-art approaches. A study by41 proposed transfer learning methods that achieved 99.02% accuracy with ResNet50 using Adadelta. The method proposed by42 achieved testing accuracies of 99.34% and 99.51% with Inception-v3 and DenseNet201, respectively.

The ensemble model combines multiple models and has also been applied in brain tumor research. For example, the study by43 explored the ensemble model and stacking approach, and the proposed model achieved 97% classification accuracy. The study by40 adopted a CNN-based transfer learning method to classify brain MRI scans into two classes via VGG16 (a pretrained model). The results show that the CNN archives have 96.5% training and 90% testing accuracy. The study44 proposed using three deep learning models with varying depths and architectures to detect and classify medical images of brain cancer automatically. This study trained and tested the customized CNN model; the test accuracy was 96%, and the AUC was 99%. The study45 compared a deep transfer learning model with AlexNet and GoogleNet for brain tumor classification on T1W MRI images. They reported that AlexNet’s accuracy is 94.6%, and GoogleNet’s is 92%. The study46 reported VGG-19 CNN class accuracies of 97.88% for meningioma, 98.76% for pituitary, 99.78% for glioma, 99.00% for no_tumor, and 98.87% overall.

Another extension of CNNs is the integration of explainable artificial intelligence (XAI) to visualize the decision-making process of CNN models. XAI generates local and global explanations for CNN model predictions, thereby enhancing transparency and interpretability. XAI visualization shows CNNs’ focus areas during decision-making. XAI identifies the regional and international regions of the input images and extracts the features that contribute to the model’s predictions. XAI provides insights into the models’ decision-making process. Local interpretable model-agnostic explanations (LIME), Shapley additive explanations (SHAPs), and gradient-weighted class activation mapping (Grad-CAM) are among the XAI techniques47. To improve the effectiveness of CNNs coupled with XAI, many studies have developed novel CNNs and integrated XAI to visualize and understand the decisions made by their models. Specifically, in brain tumor research48–50,, and51 applied a customized CNN for brain tumor prediction with XAI.

The study by applied a performance analysis of SOTA CNN architectures for brain tumour detection. The study analysed and contrasted the performance of CNN models trained on the publicly accessible Br35h dataset for brain tumour detection. These models included the LeNet, AlexNet, VGG16, VGG19, and ResNet50. Several optimisers were used in this research to fine-tune the performance of the CNN model. These included Adam (adaptive moment estimation), SGD (stochastic gradient descent), and RMSprop (root‐mean‐square propagation). Accuracy, miss‐classification rate, sensitivity, specificity, NPV (negative predictive value), PPV (positive predictive value), F1‐score, and false omission rate (FOR) were used to assess the efficacy of five CNN architectures trained with three optimisers. The experimental results showed that the AlexNet architecture with the SGD optimiser performed better than other CNN architectures with different optimisers, achieving the highest accuracy of 98.79% with a miss classification rate of 1.20%. It also achieved 98.98% sensitivity, 98.58% specificity, 98.93% NPV, 98.65% PPV, 98.82% F1‐score, and 1.06% FOR. A research matrix is included in Table 1, which summarizes CNN models and experimental results.

Table 1.

Research matrix.

Author Model Accuracy Contribution
71 MobileNetV2 99.00% Contribution of lightweight CNN in real-time brain tumour detection
34 Lightweight parallel depth-wise separable CNN 99.35%, 99.30%, and 99.22%
16 Custome CNN, Inception-v4, and EfficientNet-B4 97.58%, 99.56%, and 99.76%.
36 Lightweight CNN 99.48% accuracy for binary classification and 96.86% for multiclass classification
15 Lightweight CNN Classification accuracy of 99.58% and a Dice similarity coefficient (DSC) of 95.7%.
LeNet, AlexNet, VGG16, VGG19, and ResNet50 AlexNet achieved the highest accuracy of 98.79%. Improved our understanding of SOTA CNNs’ capabilities
65 Hybrid CNN-NADE, Alexnet, HCS, SVM

95%, 96.61%,94.03%,

87.95%

A Hybrid CNN-NADE Model
3 CNN 97.2% Create a lightweight CNN architecture
66 MCNN model 99.89% Improved Convolutional Neural Network
67 Efficientnetb0 & SHAP 99.84% Improved convolutional neural network (CNN) and Shapley additive explanation (SHAP).
40 Transfer learning VGG16 90%
44 Customized CNN, ResNet-50, 50 and VGG-16 96% Customized CNN model, ResNet-50 model, and VGG-16 model in the experiments.
45 Alexnet & GoogleNet 94.6% & 92% Compared Alexnet and the GoogleNet architectures on T1-w MRI images.
68 Inceptionv3 99.7% Proposed a deep feature extraction approach
46 VGG-19 98.87% Improvised classification using deep learning.
69 Resnet-50 Dice scores of 0.9203, 0.9113, and 0.8726 Utilizing a feature extraction CNN architecture
70 Deep learning

Dice scores of 71%, 88%,

and 74%

Segmentation with multiple encoders

Experimental methods and materials

This section describes the hardware specifications, dataset description, M-net model development, and training procedure for this study. The workflow of the proposed methodology is presented in Fig. 1.

Fig. 1.

Fig. 1

Workflow of the proposed methodology.

Hardware Specification

The experiments were conducted via the Precision 7680 Workstation. The workstation is a 13th-generation Intel® Core™ i9-13950HX vPro with Windows 11 Pro, an NVIDIA® RTX™ 3500 Ada Generation GPU, 32 GB DDR5 RAM, and a 1 TB SSD. Python (version 3.9) was chosen as the programming language because it supports TensorFlow-GPU, SHAP, and LIME generation.

Dataset description

Three MRI datasets were used to assess the performance of the M-Net in brain tumor detection and classification. Dataset A, a 3-class dataset, contains modalities for Glioma, Pituitary, and Meningioma. Dataset B is a 4-class dataset that combines Figshare, SARTAJ, and Br35H. This dataset includes 7,023 images in four classes: Glioma, Meningioma, Pituitary, and No Tumour. No tumour class images were taken from the Br35H dataset. Dataset C was also collected from Kaggle, a private collection of 15 Classes. The dataset includes Astrocytoma, Ependymoma, Germinoma, Glioblastoma, Medulloblastoma, Meningioma, Oligodendroglioma, Pituitary, and Schwannoma. The standard class includes MRI images of individuals without brain tumors Fig. 2.

Fig. 2.

Fig. 2

Brain MRI obtained from the Sagittal Plane, Axial Plane, and Coronal Plane of the four brain tumor classes.

In this study, 70% of the images were used for training M-Net, 20% for validation, and 10% for testing its performance on unseen test images. To prevent data leakage, the photos were split into train, validation, and test directories. As the authors52 suggested, the M-Net model verification stage should include evaluating the final model on a dataset not seen during training. The study also indicated that ‘test set’ and ‘validation set’ should not be considered the same.

To ensure performance consistency and generalizability, the model was trained and tested on both balanced and imbalanced datasets with varying numbers of brain tumour modalities (See Fig. 3). Even the MRI image classes range from only 38 to 4000 brain tumour MRI images. Furthermore, it was essential to analyse how the model performs in detecting and classifying unseen brain tumour images after training on the training data. These strategies will provide a better understanding of the M-Net model construction strategy, especially across classes that are not evenly distributed.

Fig. 3.

Fig. 3

.

Image preprocessing

To enhance the visibility of imaging features and reduce noise, a multistep preprocessing pipeline was applied. Contrast Limited Adaptive Histogram Equalization (CLAHE) was employed to improve local contrast in dense tissue regions, followed by Gaussian blurring to suppress high-frequency noise, defined as:

graphic file with name d33e775.gif 1

Where σ represents the standard deviation of the Gaussian distribution, subsequently, bilateral filtering was applied to preserve edge structures critical for mass detection, expressed as:

graphic file with name d33e781.gif 2

where fr denotes the range kernel, fs the spatial kernel, and Wp the normalization factor. Non-local means denoising was used to enhance smoothness further while preserving important features. Finally, unsharp masking was performed to improve edge clarity, formulated as:

graphic file with name d33e787.gif 3

To mitigate overfitting and enhance generalization, a unified set of augmentation techniques was applied across all datasets, including geometric transformations (± 15° rotations, center cropping) and photometric adjustments (brightness/contrast modulation, random scaling). Horizontal and vertical flipping were incorporated to increase dataset diversity, as commonly practiced in mammography deep learning to simulate orientation variations40,41. To address concerns about potential anatomical distortion from flipping (e.g., altering medial-lateral breast positioning), a controlled experiment was conducted by retraining M-Net without horizontal flips, while retaining clinically aligned augmentations, such as ± 15° rotations (mimicking patient positioning variability) and CLAHE-based contrast adjustments (reflecting equipment differences).

Image augmentation

In this study, we employed extensive data augmentation techniques, including rotation, flipping, and contrast modulation, to generate a richer and more diverse dataset for MRI-based brain tumour classification (see Fig. 4). Image augmentation has proven to be a highly effective and practical approach for enhancing the robustness of models with limited ground-truth datasets53. In this study, we applied rotations, scaling, flipping, contrast adjustments, and noise addition on the training datasets. In the context of brain tumour detection, these techniques enable models to better account for variability in tumour characteristics, including size, shape, and location. For example, rotation (± 20°) allows the model to handle tumours at different orientations, while scaling (± 10%) enables it to handle variations in tumour size. Additionally, random flipping helps the model generalise across different image orientations, ensuring it can identify tumours in both left- and right-hemisphere scans. Contrast adjustment (± 15%) enables the model to adapt to differences in image quality, while Gaussian noise (variance of 0.05) improves robustness against imaging artefacts often found in medical scans.

Fig. 4.

Fig. 4

Images after augmentation.

M-Net development

M-Net is a sequential CNN, which uses a linear stack of layers, each feeding directly into the next. This simplicity makes them easier to design, train, and debug, particularly in contexts such as medical image analysis, where interpretability, efficiency, and modularity are essential. M-Net consists of four convolutional layers. Number of kernels used: 32, 64, 128, and 128, and size 4 × 4. The 4 × 4 kernel size captures slightly broader spatial features than 3 × 3 (with 16 parameters per filter compared to 9), while keeping the computational load lower than 5 × 5 (which has 25 parameters per filter). This kernel size is particularly suitable for medium-sized input images (150 × 150), allowing the model to extract meaningful patterns without introducing excessive complexity. The input MRI image Inline graphic is fed into the layer and processed by the 1 st convolution layer:

graphic file with name d33e829.gif

This convolution extracts low-level features, such as edges and textures, and the resulting features pass through a MaxPool layer to reduce spatial dimensions while retaining the most essential features from the MRI images.

graphic file with name d33e834.gif

The pooled feature map then serves as the input to the second convolutional layer, allowing. Inline graphic to capture more complex shapes, like the outline of a tumor, the curve of tissue structures, or other meaningful regions in the MRI detected by Inline graphic

graphic file with name d33e846.gif
graphic file with name d33e849.gif

.

These feature maps are then fed into the third convolutional layer, where higher-level structures and more complex spatial relationships are learned:

graphic file with name d33e855.gif
graphic file with name d33e858.gif

The output of Inline graphic is then passed to the fourth convolutional layer, which captures even more abstract representations suitable for classification:

graphic file with name d33e867.gif

After the final convolution, the feature maps are flattened into a one-dimensional vector, consolidating all extracted features for the dense layer:

graphic file with name d33e872.gif

This flattened output feeds into a dense layer with 512 neurons, which learns complex feature interactions and higher-level representations:

graphic file with name d33e877.gif

A dropout layer with a rate of 0.5 follows, meaning that 50% of the neurons are randomly set to zero during training. This is mathematically represented as each neuron having a probability. Inline graphic of being dropped effectively reduces the risk of overfitting by forcing the network to learn more robust representations.

graphic file with name d33e886.gif

Finally, the output layer consists of 4 neurons with SoftMax activation, providing the probability distribution across the four classes:

graphic file with name d33e891.gif

Overall, this architecture balances effective feature extraction, computational efficiency, and generalization ability, making it highly suitable for classifying medium-resolution images such as 150 × 150. The structure of M-Net is graphically viewed in Table 2; Fig. 5.

Table 2.

Architecture details of the M-Net model.

Layer Type Output Shape Param #
conv2d (None, 147, 147, 32) 1,568
max_pooling2d (None, 49, 49, 32) 0
conv2d_1 (None, 46, 46, 64) 32,832
max_pooling2d_1 (None, 15, 15, 64) 0
conv2d_2 (None, 12, 12, 128) 131,200
max_pooling2d_2 (None, 4, 4, 128) 0
conv2d_3 (None, 1, 1, 128) 262,272
flatten (None, 128) 0
dense (None, 512) 66,048
Total params 495,972
Trainable params 495,972
Non-trainable params 0

Fig. 5.

Fig. 5

M-Net model visualization.

This architecture shows the sequential convolutional and pooling layers integrated into the M-Net model, optimized for efficient MRI classification (Fig. 6).

Fig. 6.

Fig. 6

Layered architecture of M-Net.

Justification for 4× 4 Kernels

While the M-Net design is relatively simple in its sequential structure, the choice of 4 × 4 kernels rather than the more common 3 × 3 kernels is a key design consideration that balances effective feature extraction with computational efficiency. Though a larger kernel size in a CNN can capture features from a wider spatial area, in an edge device deployment, a CNN model with a large kernel size incurs rapidly increasing computational cost. Main problems of using a big kernel size are the computational cost with kernel size, especially in standard convolution, higher latency, which reduces real-time inference, more memory and parameter storage, more energy use, possible overfitting or diminishing returns if the kernel gets large without enough data or without architectural tricks, and Throughput reduction, even if accuracy improves a bit. The 4× 4 kernels allow the model to capture slightly broader spatial features than 3 × 3 kernels (16 parameters per filter vs. 9), while keeping the computational load manageable compared to 5 × 5 kernels (25 parameters per filter). This is particularly important in medical image analysis, where interpretability and efficiency are crucial.

The key math of Parameters in a standard convolution = C_in × k × k × C_out, and the Compute cost scales roughly as = W × H × C_in × k × k × C_out. So, when the kernel size goes from 3 × 3 to 7 × 7, the kernel term changes from 9 to 49, which is about 5.44 times larger. That means more parameters and more multiply-accumulate operations for the same input/output channels. Therefore, selecting the kernel size for the M-Net model ensures faster inference, reduces parameter storage and memory traffic during inference, and, in turn, results in less power consumption and less memory movement, which usually increases energy use. Throughput (FPS) does not drop; optimization and training can become easier. A recent study by Sun et al. (2024), a large-kernel paper, notes that because convolution has square complexity, scaling kernels up can create “an enormous amount of parameters” and can induce “severe optimization problems.

Algorithm 1 presents the step-by-step training procedure, including model compilation, evaluation, and integration of SHAP and Grad-CAM for interpretability.

Table 3.

Algorithm.

graphic file with name 41598_2026_51236_Tab3_HTML.jpg

Training M-Net

To prevent data leakage, the images were transferred to the train, validation, and test directories. Separating the MRI images into three directories ensured against any potential data leakage, and hence, the test and validation sets were not considered the same. As52 suggested, the M-Net model verification stage included evaluating the final model on a test dataset (10% of the images) that was not used during training. We have carefully ensured that the dataset splitting was performed strictly to avoid any overlap between the training, validation, and testing sets. Specifically:

Patient-wise Data Splitting

In our study, images from the same patient were never included in both the training and testing sets. This was carefully enforced by organizing the MRI images into separate directories for training, validation, and testing. The images were categorized by patient, ensuring that no patient’s data was included in both the training and testing sets.

SMOTE and Data Augmentation

For the imbalanced dataset, we applied SMOTE (Synthetic Minority Over-sampling Technique) and Data Augmentation. However, as suggested by Apicella (2025), SMOTE and augmentation were applied after splitting the dataset into training and testing sets. This ensures that no information from the test set was used during the training phase, which is a critical step to prevent data leakage. SMOTE was applied only to the training set in each fold to balance class distribution. This was done by generating synthetic samples from the existing training data, without affecting the test set. Data augmentation techniques (e.g., random rotations, flips, and brightness adjustments) were applied only to the training data. These transformations are designed to increase the diversity of the training set without introducing any information from the test set. The training hyperparameters are provided in Table 2.

Cross-validation

The M-Net model was trained using K-fold stratified cross-validation. To ensure that each fold trains the same proportion of class labels as in the original dataset, K-5 stratified cross-validation was selected. In this study, 70% of the images were used for training M-Net, 20% for validation, and 10% for testing its performance on unseen test images. Since the study utilised three (3) MRI datasets, three separate training sets were created for each dataset. Hence, test results also emerged from two distinct datasets.

Explicit Clarification

The Adam algorithm was used for model optimization. The algorithm ensures the model weights adapt to the calculated gradients, thereby enhancing performance and accuracy in classifying brain tumor modalities. The loss function was sparse categorical cross-entropy. The training hyperparameters are provided in Table 4.

Table 4.

Hyperparameters of training.

C

Epochs = 50

Batch size = 16

Image size = (150, 150, 3)

Learning rate = 1.0000e-04

K_folds = 5

Optimizer = Adam(learning_rate=LEARNING_RATE)

Loss = SparseCategoricalCrossentropy(from_logits=True)

Early stopping = EarlyStopping(monitor=‘val_accuracy’, patience=10, verbose=1, restore_best_weights=True)

Learning_rate_scheduler = LearningRateScheduler(lambda epoch: LEARNING_RATE * 0.1 ** (epoch//10))

Callbacks = [early_stopping, lr_scheduler]

M-Net results and comparison with SOTA CNNs

The following performance metrics are used to assess the M-Net’s performance:

graphic file with name d33e1140.gif 4
graphic file with name d33e1146.gif 5

Furthermore, confusi

graphic file with name d33e1153.gif 6

on matrices, training and validation loss curves, ROC curves, statistical analysis, and Computational cost were evaluated for M-Net performance.

Five-fold cross-validation and combined performance of M-Net on three brain tumour datasets

The combined accuracy, precision, and recall (See Table 5) indicate that after stratified k-fold training, M-Net demonstrates strong robustness and an ability to handle both balanced and imbalanced brain tumour datasets efficiently. The results suggest that M-Net achieves 99% accuracy on the 4-class brain tumor dataset and 97% accuracy on the 3-class and 15-class brain tumor datasets. One contributing factor to the 4-class high accuracy is that the dataset has more uniform MRI images per class, with each class containing more images than in the 3-class and 15-class datasets.

Table 5.

5-fold cross-validation combined performance of M-Net.

Combined Accuracy 3-Class 4-Class 15-Class
97% 99% 97%
Combined Precision 97% 99% 97%
Combined Recall 97% 99% 97%
Combined F1-Score 97% 99% 97%

Fold-wise classification report

The fold-based results also support the combined results M-Net achieved, presented in Table 3. Across the 5-fold cross-validation, M-Net achieves 98%−99% accuracy in each fold (see Table 6). Specifically, in Fold 1 and 2, M-Net attains 99% accuracy across all classes, with 99% Macro-Avg and 99% Weighted-Avg scores. Only in the Fold 3 is a minimal 1% drop noticeable. However, in Fold 4 and Fold 5, M-Net consistently returns 99% accuracy across the three datasets. Importantly, the alignment between Macro-Average and Weighted-Average values (e.g., both staying at 98–99%) confirms that M-Net does not rely on majority-class dominance, and delivers near-perfect precision and recall (98–99%) even as the number of classes increases from 3-class to 15-class.

Table 6.

Fold-wise Precision, Recall, and F1-score results.

Fold Metrics 3-Class 4-Class 15-Class
Class Precision Recall F1-Score Precision Recall F1-Score Precision Recall F1-Score
1 Accuracy 99% 99% 99%
Macro Avg 99% 99% 99% 99% 99% 99% 99% 99% 99%
Weighted Avg 99% 99% 99% 99% 99% 99% 99% 99% 99%
2 Accuracy 99% 99% 99%
Macro Avg 99% 99% 99% 99% 99% 99% 99% 99% 99%
Weighted Avg 99% 99% 99% 99% 99% 99% 99% 99% 99%
3 Accuracy 98% 98% 98%
Macro Avg 98% 98% 98% 98% 98% 98% 98% 98% 98%
Weighted Avg 98% 98% 98% 98% 98% 98% 98% 98% 98%
4 Accuracy 99% 99% 99%
Macro Avg 99% 99% 99% 99% 99% 99% 99% 99% 99%
Weighted Avg 99% 99% 99% 99% 99% 99% 99% 99% 99%
5 Accuracy 99% 99% 99%
Macro Avg 99% 99% 99% 99% 99% 99% 99% 99% 99%
Weighted Avg 99% 99% 99% 99% 99% 99% 99% 99% 99%

Combined classification report and confusion matrix

The combined classification report and confusion matrix reflect the consistency of M-Net in detecting and classifying brain tumour modalities (see Fig. 7; Table 7). The confusion matrices highlight that in the 3-Class, Glioma (GM), Meningioma (MN), and No Tumour (NT), there is high accuracy, with 2295 true positives (GM correctly identified as GM) and minimal misclassifications (17 and 13 for GM misclassified as MN and NT, respectively). However, a small number of NTs were misclassified as GM (13), and some MN cases (25) were misclassified as NT. The 4-class matrix demonstrates a relatively strong model with very few false positives and false negatives. There is a slight misclassification of NT as GM (16). M-Net shows strong performance, with relatively low misclassification rates across the majority of cases, as supported by the 15-Class confusion matrix. For example, GB (Glioblastoma) was trained on 35 images, achieving a noticeable success rate (53 out of 55 in 5-fold stratified training).

Fig. 7.

Fig. 7

Combined confusion matrix of M-Net.

Table 7.

Combined classification report of M-Net using 3 datasets.

Class Precision Recall F1-Score Support
Glioma 99% 99% 99% 2645
Meningioma 98% 99% 98% 2720
No tumor 99% 100% 99% 3235
Pituitary 100% 99% 99% 2890
Accuracy 11,490
Macro avg 99% 99% 99% 11,490
Weighted avg 99% 99% 99% 11,490
4-class
Class Precision Recall F1-Score Support
Glioma 99% 99% 99% 2645
Meningioma 98% 99% 98% 2720
No tumor 99% 100% 99% 3235
Pituitary 100% 99% 99% 2890
Accuracy 11,490
Macro avg 99% 99% 99% 11,490
Weighted avg 99% 99% 99% 11,490
15-class
Class Precision Recall F1-Score Support
Astrocitoma 97% 94% 96% 955
Carcinoma 99% 98% 99% 385
Ependimoma 96% 95% 95% 255
Ganglioglioma 100% 96% 98% 55
Germinoma 95% 99% 97% 165
Glioblastoma 88% 97% 92% 320
Granuloma 95% 92% 93% 120
Meduloblastoma 89% 100% 94% 220
Meningioma 98% 97% 98% 1385
Neurocitoma 99% 97% 98% 755
Normal 97% 97% 97% 845
Oligodendroglioma 99% 99% 99% 370
Papiloma 97% 98% 97% 370
Schwannoma 97% 98% 97% 790
Tuberculoma 96% 94% 95% 235
Accuracy 7225
Macro avg 96% 97% 96% 7225
Weighted avg 97% 97% 97% 7225

Combined accuracy and loss curve

The accuracy and loss curves for k-fold cross-validation are shown in Fig. 8. The curve suggests the training accuracy rises steadily across epochs for all folds, and the model reaches close to 1.0 for each fold. This indicates that the model is learning effectively learns the distinguishing features of the different tumour modalities. Furthermore, the smooth upward trend in the accuracy curves indicates minimal overfitting, and the validation accuracy closely tracks the training accuracy. The loss curves exhibit a clear downward trend across all folds. Initially, validation accuracy fluctuates; however, it improves as the epochs progress. The model shows effective convergence with a high training accuracy and low training loss over time. By storing the best value and triggering EarlyStopping, the model avoids overfitting by discontinuing training.

Fig. 8.

Fig. 8

Combined accuracy and loss curves.

Combined ROC curve

The ROC curves (Fig. 9) of the M-Net model suggest an overall accuracy of 99% reflects its ability to detect and classify 3 classes to 15 classes without a significant loss in performance. For the 3-class problem, the ROC curve approaches the upper-left corner, indicating high true positive rates (TPR) while keeping the false positive rate (FPR) low. This suggests that the model is exceptionally adept at distinguishing between the three classes—Glioma, Meningioma, and Pituitary. Similarly, the 4-class model performs at an equally high level. Adding a fourth class does not degrade performance; instead, the ROC curve remains close to the ideal. The 15-class model is more complex due to a high number of classes and fewer images per class. However, the 15-class ROC curve indicates that the model consistently maintains strong true positive rates (TPR).

Fig. 9.

Fig. 9

ROC curves of three datasets of M-Net.

Unseen test data results

The performance of M-Net in detecting and classifying unseen datasets (10% data, which was kept in a separate folder) achieved accuracy (97–98%) and low-test loss (0.05–0.12%) across three different datasets (see Table 7). Even with a larger dataset (15-Class dataset), the model maintains high accuracy and low loss, demonstrating its ability to generalize well to unseen data. The model’s low loss values suggest that overfitting is not a significant issue. A large discrepancy between training and test accuracy typically indicates overfitting.

Table 8.

M-Net’s performance on unseen data.

Test Accuracy 3-Class 4-Class 15-Class
97% 98% 97%
Test Loss 0.12% 0.05% 0.08%

Classification reports and confusion matrix on unseen data

The test accuracy is reflected in the classification report (see Table 9). In the 3-Class dataset, the model achieves 99% precision and recall across all classes, a balanced F1-Score of 99%, and an overall accuracy of 99%. In the 4-Class dataset, precision decreases slightly, particularly for Meningioma (94%), while recall remains high at 97%, resulting in an accuracy of 97%. The 15-Class dataset shows slightly varied precision, with the highest at 100% for Carcinoma, Ependimoma, and Medulloblastoma, and somewhat lower precision for Astrocytoma (97%) and Glioblastoma (95%), achieving an overall accuracy of 97%.

Table 9.

Classification report on unseen dataset.

3-Class
Class Precision Recall F1-Score Support
Glioma 99% 99% 99% 2325
Meningioma 99% 98% 98% 1185
Pituitary 99% 100% 100% 1550
Accuracy - - 99% 5060
Macro Avg 99% 99% 99% 5060
Weighted Avg 99% 99% 99% 5060
4-Class
Class Precision Recall F1-Score Support
Glioma 98% 94% 96% 448
Meningioma 94% 96% 95% 457
No tumor 98% 99% 99% 560
Pituitary 99% 100% 100% 495
Accuracy - - 97% 1960
Macro Avg 97% 97% 97% 1960
Weighted Avg 98% 97% 97% 1960
15-Class
Class Precision Recall F1-Score Support
Astrocitoma 97% 94% 96% 163
Carcinoma 100% 94% 97% 71
Ependimoma 100% 86% 93% 43
Ganglioglioma 100% 100% 100% 13
Germinoma 100% 100% 100% 29
Glioblastoma 95% 97% 96% 60
Granuloma 100% 100% 100% 23
Meduloblastoma 93% 100% 96% 37
Meningioma 98% 99% 98% 246
Neurocitoma 98% 98% 98% 128
Normal 97% 99% 98% 146
Oligodendroglioma 98% 100% 99% 62
Papiloma 99% 97% 98% 68
Schwannoma 98% 98% 98% 131
Tuberculoma 95% 98% 97% 43
Accuracy 97% 1263
Macro Avg 98% 97% 98% 1263
Weighted Avg 98% 97% 97% 1263

The confusion matrix (see Fig. 10) suggests that in the 3-class dataset, the model correctly predicts 845 and misclassifies 15; in the 4-class 1911, it misclassifies 73; in the 15-class 1231, it misclassifies 29. The low misclassification rate here reflects the model’s effectiveness in brain tumour classification and detection. Specifically, in the 15-Class dataset, the Ganglioglioma (GG) class was identified as 13, even though the model was trained on only 55 images.

Fig. 10.

Fig. 10

Confusion matrix of unseen data.

Ablation Study

In this study, we conducted an ablation analysis to assess the stability and performance of the M-Net model. The results of the initial ablation study (Table 10) demonstrated that removing certain components from the M-Net architecture led to slight performance drops, which confirms the relevance of each layer and the importance of regularization.

Table 10.

Ablation strategy and results.

Ablation strategy Training accuracy Validation accuracy Test accuracy
Without augmentation 97% 96% 94%
Without SMOTE 96% 95% 94%
Removal of first convolutional layer 98% 97% 96%
Removal of last convolutional layer 98% 97% 96%
Removal of the Dropout (0.5) layer to check if regularization impacts 99% 98% 97%

To examine the stability and performance of the M-Net model, we conducted two ablation studies on the 4-class brain tumor. Since this dataset achieved the highest Accuracy, precision, and Recall, we selected it for the ablation study. In this research, three (3) configurations of the M-Net architecture were eliminated: the first Convolutional Layer, the last Convolutional Layer, and the Dropout (0.5) layer, to test whether regularization helps in this case. The aim was to determine whether eliminating components of M-Net affects performance and to demonstrate that the M-Net design is efficient and adaptable for addressing the challenges of brain tumor classification. The results (see Table 10) reveal that changing the architecture decreases the performance of the original M-Net model by 1%.

Furthermore, we evaluated the potential influence of preprocessing and augmentation on the M-Net, with and without preprocessing (e.g., denoising, augmentation, and other enhancements). The goal is to evaluate whether the observed performance gains are attributable to the proposed M-Net architecture or to the preprocessing steps.

Benchmark Comparison

This study applied six SOTA lightweight CNNs to 4-Class datasets. The lightweight versions are edge device-friendly. The selection of CNNs was based on the taxonomy of54 The six SOTA CNNs were selected from multipath (DenseNet121, ResNet50), depthwise (InceptionV3), width-based multiconnection (Xception), depthwise separable (MobileNet), and spatial-exploitation (VGG16) networks. The selection aimed to cover most CNN architectures to investigate the most effective MRI image detection and classification network. Three optimizers, Adam, Adamax, and RMSprop, were also used to find the effective optimizer in CNN Table 11.

Table 11.

Performance comparison of the SOTA CNNs and optimizers.

Model Epoch (Adam) Accuracy (Adam) Epoch (Adamax) Accuracy (Adamax) Epoch (RMSprop) Accuracy (RMSprop)
DenseNet121 11 98% 37 96% 23 98%
ResNet50 37 97% 46 89% 43 97%
InceptionV3 26 97% 36 88% 37 97%
Xception 27 97% 30 94% 33 95%
MobileNet 35 89% 34 73% 27 87%
VGG16 30 92% 35 84% 36 96%

Table 5 compares the performance of the six SOTA CNNs when using the Adam, Adamax, and RMSprop optimizers. DenseNet121 achieves the highest accuracy of 98% when both Adam and RMSprop are used. However, Adam requires 11 and an RMSprop of 23 epochs. Conversely, ResNet50 requires significantly more epochs; however, optimizers do not impact the accuracy. MobileNet and VGG16 exhibit lower accuracy, particularly with Adamax (73% and 84%, respectively). Most importantly, Adam and RMSprop show better performance consistency, whereas Adamax appears less efficient in terms of convergence and accuracy, particularly on MobileNet and InceptionV3. A scatter plot is also presented in Fig. 9 to illustrate the relationship between accuracy and the number of epochs for the Adam, Adamax, and RMSprop optimizers across the six SOTA CNNs.

Figure 11 presents a scatter plot of accuracy versus epochs for the Adam, Adamax, and RMSprop optimizers across six SOTA CNNs.

Fig. 11.

Fig. 11

SOTA CNN models, accuracy, and number of epochs.

The confusion matrix (Fig. 12) suggests that ResNet50 and DenseNet121 exhibit lower Type 1 and Type 2 errors. The lower Type I and Type II values indicate better overall classification performance of DenseNet121 and ResNet50. InceptionV3 also shows competitive performance, while Xception balances false positives and false negatives. However, MobileNetV2 and VGG16 have more false positives and false negatives. Overall, DenseNet121 and ResNet50 appear to be the most robust models, while MobileNetV2 shows weaker performance, with higher misclassification rates Table 12.

Fig. 12.

Fig. 12

Confusion matrices of the SOTA CNNs.

Table 12.

Performance comparison of SOTA CNNs; Transfer learning and optimizers.

DenseNet121 Epoch Accuracy (Adam) Epoch Accuracy (Adamax) Epoch Accuracy (RMSprop)
36 84% 36 84% 36 84%
ResNet50 35 59% 35 59% 35 59%
InceptionV3 25 80% 25 80% 25 80%
Xception 33 87% 33 87% 33 87%
MobileNet 25 87% 25 87% 25 87%
VGG16 23 75% 23 75% 23 75%

Transfer learning performance

The accuracy of the TL table and scatter plot suggests that DenseNet121 achieves 84% precision across all optimisers (Table 6; Fig. 13). Xception closely follows, reaching 87% accuracy with 32 epochs. ResNet50 achieves the lowest accuracy in transfer learning (59%), despite requiring ~ 35 epochs, suggesting optimization challenges.

Fig. 13.

Fig. 13

Transfer learning CNN models, accuracy, and number of epochs.

The confusion matrices of the TL are shown in Fig. 14. DenseNet121 and MobileNetV2 yield relatively lower Type 1 and Type 2 errors. VGG16, however, suffers from higher misclassification rates. ResNet50 has the highest misclassification rates, with significant false positives and false negatives.

Fig. 14.

Fig. 14

Confusion matrix of transfer learning.

Ensemble model performance

In this study, two ensemble models are developed.

Tables 7, 13 and 14 compares the ensemble model of the multipath-depth-width CNN (DenseNet–Inception-Xception) using Soft Voting, Hard Voting, and Rank-based methods, indicating that Soft Voting achieves the highest accuracy (98.22%). The confusion matrix (Table 15) also supports Soft Voting, leading to minimal misclassifications across all classes. In contrast, the rank-based method results in higher misclassification rates, particularly for the NT and PT classes. Soft voting has the lowest errors, with Type I = 21 and Type II = 8; Hard voting follows with Type I = 21 and Type II = 9; and Rank-based voting has the highest errors, with Type I = 46 and Type II = 27, indicating its lower classification reliability.

Table 13.

Performance of the multi-path-depth-width-based (DIX) ensemble.

Ensemble method Accuracy Precision Recall F1 Score

Multi path-depth-Width based

DenseNet – Inception-Xception

Soft Voting 98% 98% 98% 98%
Rank-Based 95% 96% 95% 95%
Hard Voting 98% 98% 98% 98%
Table 14.

Performance comparison of the multipath-depth-spatial exploitation ensemble (RIV).

Ensemble method Accuracy Precision Recall F1 Score

Multi path-depth-Spatial exploitation

(ResNet18- InceptionV3-VGG16)

Soft Voting 93% 93% 93% 93%
Rank-Based 79% 80% 79% 79%
Hard Voting 82% 82% 82% 81%
Table 15.

Accuracy statistic comparison on 4-class dataset.

Metric Xception Inception V3 VGG-16 VGG-19 ResNet50 MobileNet DenseNet121 EfficientNet M-Net
Accuracy Mean 95% 65% 76% 69% 62% 55% 95% 24% 75%
Accuracy Std 12% 18% 26% 27% 20% 20% 10% 6% 21%
Accuracy Median 100% 71% 89% 78% 66% 58% 99% 23% 83%
Accuracy IQR 0.64% 23% 31% 40% 26% 31% 5% 11% 21%
Accuracy Range 40% 57% 77% 78% 61% 65% 45% 19% 78%
Accuracy CI (85%, 105%) (51%, 78%) (56%, 95%) (50%, 88%) (45%, 80%) (43%, 67%) (90%, 99%) (19%, 29%) (64%, 86%)
Accuracy Effect Size (7.71, 6.86) (3.69, 3.33) (2.93, 2.64) (2.59, 2.37) (3.23, 2.81) (2.75, 2.58) (9.47, 9.09) (3.87, 3.44) (3.60, 3.43)

Ensemble 1:

Fig. 15.

Fig. 15

Confusion matrices of Ensemble 1 (DIX).

Ensemble 2:

However, the second ensemble model, which exploits multipath depth spatially (ResNet18-InceptionV3-VGG16) with soft voting, hard voting, and rank-based methods, outperforms the multipath depth-based model (see Table 7; Fig. 16). Soft voting maintains the highest accuracy (93.10%). Soft voting has the lowest errors, with Type I = 31 and Type II = 21; Hard voting follows with Type I = 64 and Type II = 42; and Rank-based voting has the highest errors, with Type I = 111 and Type II = 79, further confirming that Soft voting offers the best reliability classification.

Fig. 16.

Fig. 16

Confusion matrices of Ensemble 2 (RIV).

Statistical comparison of M-Net and SOTA CNNs

The accuracy, mean, median, and range are presented in Table 16. M-Net achieves 75% accuracy, outperforming models such as InceptionV3 (65%) and EfficientNet (24%). The model’s median accuracy of 83% is impressive, surpassing that of EfficientNet (23%) and MobileNet (58%). With an accuracy IQR of 21%, M-Net ranks in the middle, demonstrating reasonable reliability. The model’s 78%−95% accuracy range and M-Net’s 95% confidence interval reflect more predictable performance than models with narrower ranges, such as EfficientNet.

Table 16.

Loss statistics comparison on 4-class dataset.

Metric Xception InceptionV3 VGG-16 VGG-19 ResNet50 MobileNet DenseNet121 EfficientNet M-Net
Loss Mean 0.17 1.12 0.76 0.95 1.71 1.41 0.19 2.60 0.75
Loss Std 0.40 0.60 0.81 0.82 1.06 0.63 0.34 0.31 0.64
Loss Median 0.0021 0.94 0.36 0.65 1.47 1.34 0.03 2.58 0.52
Loss IQR 0.04 0.78 0.99 1.28 1.32 0.96 0.16 0.52 0.65
Loss Range 1.29 1.97 2.35 2.37 3.35 2.11 1.53 0.92 2.30
Loss 95% CI (−0.16, 0.49) (0.67, 1.58) (0.16, 1.37) (0.37, 1.53) (0.76, 2.65) (1.04, 1.79) (0.03, 0.35) (2.35, 2.86) (0.43, 1.08)
Loss Effect Size (0.41, 0.37) (2, 2) (1, 1) (1, 1) (2, 1) (2, 2) (0.55, 0.53) (8, 7) (1, 1)

The statistical analysis of data loss is presented in Table 17. Among the evaluated models, M-Net achieved a mean loss of 0.75, which was lower than those of InceptionV3 (1.12), VGG-19 (0.95), ResNet50 (1.71), MobileNet (1.41), and EfficientNet (2.60), although it remained higher than Xception (0.17) and DenseNet121 (0.19). The standard deviation of 0.64 indicates moderate variability across folds, suggesting that M-Net demonstrated reasonably consistent performance, though it was less stable than DenseNet121 (0.34) and EfficientNet (0.31). Its median loss of 0.52 further supports its comparatively favorable performance, particularly when contrasted with ResNet50 (1.47) and MobileNet (1.34). In addition, the interquartile range (IQR) of 0.65 reflects moderate dispersion in the central portion of the loss distribution, indicating greater robustness than VGG-16 (0.99), VGG-19 (1.28), and ResNet50 (1.32), while still showing more variation than Xception (0.04) and DenseNet121 (0.16). The overall loss range of 2.30 suggests noticeable variation between the minimum and maximum observed values; however, this variation was still smaller than that of VGG-16 (2.35), VGG-19 (2.37), and ResNet50 (3.35). Furthermore, the 95% confidence interval (0.43, 1.08) indicates that the loss estimates for M-Net were distributed within a moderate interval, reflecting acceptable reliability in its performance across the validation folds. The reported effect size of (1, 1) also suggests a meaningful performance difference relative to several competing models. Overall, these results indicate that M-Net achieved a balanced and competitive loss profile, combining comparatively low central tendency with moderate variability across the cross-validation process.

Table 17.

M-Net and SOTA CNNs performance evaluation on 4-class dataset.

Metric Xception Inception V3 VGG 16 VGG 19 ResNet 50 MobileNet DenseNet 121 EfficientNet M-Net
Accuracy T-test (21.82, 2.06e-08) (11.07, 1.53e-06) (8.78, 1.05e-05) (8.18, 9.66e-06) (8.54, 5.99e-05) (9.93, 1.96e-07) (42.37, 4.67e-21) (10.96, 4.28e-06) (14.84, 3.65e-11)
Loss T-test (1.17, 0.28) (5.57, 0.0003) (2.84, 0.02) (3.67, 0.0043) (4.26, 0.0037) (8.07, 2.03e-06) (2.48, 0.0223) (23.34, 1.21e-08) (4.88, 0.0001)
Accuracy Normality Test (0.43, 9.12e-07) (0.90, 0.24) (0.81, 0.02) (0.87, 0.08) (0.94, 0.65) (0.95, 0.64) (0.55, 6.75e-07) (0.94, 0.62) (0.95, 0.64)
Loss Normality Test (0.46, 2.19e-06) (0.90, 0.19) (0.81, 0.02) (0.87, 0.08) (0.92, 0.40) (0.96, 0.75) (0.58, 1.12e-06) (0.94, 0.62) (0.96, 0.75)
Accuracy Wilcoxon Test (0.0, 0.0039) (0.0, 0.0019) (0.0, 0.0019) (0.0, 0.00098) (0.0, 0.0078) (0.0, 0.0001) (0.0, 9.54e-07) (0.0, 0.0039) (0.0, 7.63e-06)
Loss Wilcoxon Test (0.0, 0.0039) (0.0, 0.0019) (0.0, 0.0019) (0.0, 0.00098) (0.0, 0.0078) (0.0, 0.0001) (0.0, 9.54e-07) (0.0, 0.0039) (0.0, 7.63e-06)
Accuracy Loss Correlation (−0.9986, 3.24e-10, −0.8416, 0.0044) (−0.9995, 2.69e-13, −1.0, 6.65e-64) (−0.9997, 1.79e-14, −1.0, 6.65e-64) (−0.9997, 2.15e-16, −1.0, 0.0) (−0.9974, 4.17e-08, −1.0, 0.0) (−0.9995, 2.52e-19, −1.0, 0.0) (−0.9986, 1.31e-25, −0.9808, 5.93e-15) (−0.9742, 8.80e-06, −1.0, 0.0) (−0.9999, 3.80e-30, −1.0, 0.0)

Confidence intervals and sampling procedure

For M-Net, model performance was evaluated on the 4-class MRI dataset containing 7,024 images using 5-fold cross-validation (k = 5). In each iteration, four folds were used for training, and the remaining fold was used for testing. This process was repeated until each fold served once as the held-out test set. Performance metrics were calculated separately for each test fold, and the reported 95% confidence intervals were derived from the five fold-wise results. Based on this evaluation, M-Net achieved a mean accuracy of 75% (95% confidence interval: 64% to 86%) and a mean loss of 0.75 (95% confidence interval: 0.43 to 1.08).

Performance evaluation (see Tables 17, 18) of the M-Net and SOTA CNNs was conducted using statistical tests, including T-tests for accuracy and loss, normality tests for accuracy and loss distributions, Wilcoxon tests for non-parametric analysis of accuracy and loss, and correlation analysis between accuracy and loss. The accuracy T-test shows a highly significant difference (t = 14.84, p = 3.65e-11) in accuracy between M-Net and SOTA CNNs. The Loss T-test (t = 4.88, p = 0.0001) further supports M-Net’s superior loss performance. The Accuracy Normality Test (p = 0.64) and Loss Normality Test (p = 0.62, skewness = 0.94) indicate normal distributions for both accuracy and loss, confirming M-Net’s stability. The Wilcoxon tests for accuracy (p = 7.63e-06) and loss (p = 7.63e-06) highlight significant differences, reaffirming M-Net’s robustness. Additionally, the correlation value of −0.9999 demonstrates a strong negative relationship between accuracy and loss, indicating that as accuracy improves, loss decreases. Overall, M-Net outperforms other models, with stronger statistical significance and greater consistency.

Table 18.

(A): M-Net and SOTA CNNs’ model comparison of computational efficiency.

Metric Xception Inception
V3
VGG-16 VGG-19 ResNet50 MobileNet DenseNet
121
EfficientNet M-Net
Params (M) 21 22 15 20 24 3 7 65 0.50
Model Size (MB) 81 84 56 7 92 13 27 247 2
Inference Latency (ms/img) 35 58 87 7 47 20 95 148 6

FLOPs

(G, conv-only, batch)

2 1 8 10 2 0.28 1 3 9

GFLOPs

(conv-only, batch)

2 1 8 10 2 0.28 1 3 9
Inference Memory Delta (MB) 20 6 5 5 14 5 11 51 9

FNR macro-average of the M-Net model

The False Negative Rate (FNR) is the proportion of instances that are incorrectly classified as negative despite being positive for a given class. A lower FNR is crucial because it indicates better performance in correctly identifying true positives, minimizing misclassification errors, and enhancing the model’s reliability. In the 4-Class brain tumor dataset, DenseNet121, ResNet50, InceptionV3, Xception, MobileNetV2, and VGG16 all exhibit relatively higher FNRs, especially for classes such as GL and PT. These models have noticeable peaks in FNR for MN, indicating that they struggle more with correctly identifying Meningioma (MN) instances. On the other hand, M-Net demonstrates significantly lower FNR across all four classes (GL, PT, MN, NT). It shows a consistent, remarkable ability to correctly classify instances, especially in more challenging classes such as MN and NT. M-Net’s FNR is substantially lower, for example, in MN (FNR = 0.0025) and NT (FNR = 0.015), which highlights its superior classification accuracy and robustness. The consistently low FNR values in M-Net suggest the model has better feature extraction capabilities. Figs. 17, 18, 19

Fig. 17.

Fig. 17

FNR macro-average.

Fig. 18.

Fig. 18

LIME partition explainer of MRI images.

Fig. 19.

Fig. 19

SHAP explainer of MRI images.

Computational cost comparison of the M-Net and SOTA CNNs

Since this study aims to develop a lightweight CNN for smart brain tumor management to be integrated into an edge device, computational cost is also an important metric. The model structure and computational cost are presented in Table 19 (A, B). The analysis is essential to understand the suitability of M-Net for edge devices.

Table 19.

(B): Model training time of each (seconds).

Model Total Train Time (s) Epoch Time Mean (s) Epoch Time Std (s) Epochs Used
InceptionV3 1656.81 57.11 0.47 29
VGG16 1518.58 84.33 0.47 18
ResNet50 904.37 27.39 0.74 33
Xception 660.53 27.51 0.74 24
MobileNet 710.07 19.18 0.53 37
DenseNet121 385.05 9.22 0.39 20
M-Net 285.84 10.97 0.12 26

The parameter settings for M-Net and other CNNs are presented in Table 19(A). The table compares the total, trainable, and non-trainable parameters for M-Net and six (6) SOTA CNNS. MobileNetV2 had the fewest total parameters; however, its performance was 78%. However, our developed M-Net has more parameters than MobileNetV2 but fewer than other CNNS. In most modern CNN architectures, a minimal number of non-trainable parameters is expected.

Table 19 (B) presents the total training times and epoch statistics for M-Net and SOTA CNNs. The data M-Net, with the shortest total training time of 285.84 s, took 26 epochs, achieving the fastest mean epoch time (10.97s) and very low variability (0.12s). In short, M-Net was the fastest with the least variability, while InceptionV3 and VGG16 were the slowest but still notable for their stable epoch times.

Results of XAI

This study utilized three explainable AI techniques, namely, LIME, SHAP, and Grad-CAM, to generate local and global explanations for the M-NET (the developed CNN) predictions on the validation and test sets. Furthermore, we conducted a statistical analysis of correctly classified and misclassified images, as well as true-positive, true-negative, false-positive, and false-negative images.

Visualization via LIME

The LIME prediction (see Fig. 16) was generated using a BrainTumor.h5 model (M-NET architecture + weights) saved after training. The input images were resized (56 × 56 pixels), and the pixel values were normalized to a [0, 1] range before passing them to the model for prediction. The result is a heatmap showing which regions (superpixels) positively or negatively contributed to the predicted class (see Fig. 18).

Visualization via SHAP

SHAP explanations are generated via a pretrained BrainTumor.h5 (see Fig. 19). The SHAP explainer (shap. Explainer) is initialized with the model and masker to compute Shapley values. For each input image, the explainer computes Shapley values, which quantify each pixel or region’s contribution to the expected class. The red regions indicate features that positively influence the prediction of a class, whereas the blue regions indicate features that reduce the likelihood.

Grad-CAM analysis of classified brain tumor modalities

The Grad-CAM generates heatmaps to visualize which regions of an image contribute most to the M-Net’s prediction. Figures 20 and 21 highlight the Grad-CAM visualization and the model’s incorrect predictions for various brain tumor cases. The red on the maps denotes greater attention given to those locations, whereas the blue denotes less attention to those regions. The heatmaps reveal that the model often focuses on irrelevant or nontumor areas, leading to misclassifications, such as “No_tumor” instead of “Pituitary” or “Glioma”. This suggests that M-Net misclassifies tumor-specific features.

Fig. 20.

Fig. 20

GRAD-CAM view of misclassified images.

Fig. 21.

Fig. 21

Grad-CAM view of correctly classified images.

Figure 18 shows the correctly classified image. The Grad-CAM clearly highlights the tumor region in the correctly classified images, demonstrating that the model’s correct focus is on relevant features for its prediction. However, in some “No_tumor” cases, the model’s attention occasionally includes irrelevant areas (blue regions), indicating room for improvement in feature extraction and focus.

Evaluation of explanations

To assess the correctness and usefulness of the generated explanations, we compare the highlighted regions in the LIME, SHAP, and Grad-CAM visualizations with the expert annotations from medical professionals. The correctness is measured by the alignment of these regions with known tumor areas and expected anatomical structures. We define “correct” explanations as those that highlight tumor regions consistent with expert knowledge and “useful” explanations as those that provide interpretable insights for medical practitioners to make more informed decisions. Additionally, consistency across methods (LIME, SHAP, Grad-CAM) serves as an indicator of reliability, with agreement among methods strengthening the validity of the explanation.

For instance, the LIME-generated heatmaps for Glioma (Fig. 16) indicate a substantial positive contribution from regions typically associated with glioma growth, such as the peritumoral zone, which is known for its distinct imaging signature. Similarly, SHAP visualizations (Fig. 17) highlight tumor regions with high contrast, matching the typical features of meningiomas and gliomas.

Limitations of XAI methods

While LIME, SHAP, and Grad-CAM provide valuable insights into model behavior, each has limitations. Grad-CAM, for instance, can sometimes focus on irrelevant regions of an image due to its reliance on gradients, leading to false positives in identifying tumor areas. LIME and SHAP, while more interpretable, may not always capture the nuanced tumor characteristics, particularly when tumors are small or have ambiguous boundaries. Additionally, these methods require significant computational resources, especially when working with large image datasets, which may limit their practical application in real-time clinical settings. These limitations highlight the need for further research to improve the spatial accuracy and computational efficiency of XAI techniques.

Pixel intensity of classified brain tumor modalities

Pixel intensity is a visualization that uses a gradient × input explanation to illustrate the model’s prediction (see Figs. 22, 23, 24, 25). The left panel shows the original MRI image, highlighting the tumor region. The right panel displays the gradient × input explanation, where the pixel intensities indicate how much each part of the image contributed to the model’s decision. Bright areas have a greater influence on prediction.

Fig. 22.

Fig. 22

Grad-CAM view of the pixel intensity of misclassified images.

Fig. 23.

Fig. 23

Grad-CAM view of the pixel intensity for correctly classified images.

Fig. 24.

Fig. 24

Distribution of pixel intensity of misclassified/correctly classified images.

Fig. 25.

Fig. 25

Distribution of pixel intensities of TP-TN-FP-FN.

Pixel intensity analysis on correctly classified and misclassified cases

To investigate differences between correctly and misclassified images, we performed pixel-level analysis of both. This is because18 suggested misclassified pixels had significantly different intensity statistics than correctly classified pixels. In this study, we further analyzed the true-positive, true-negative, false-positive, and false-negative groups. For this analysis, one hundred (100) correctly classified and misclassified images were analysed. The figure below shows a sample of correctly and misclassified images.

This boxplot (see Fig. 23) compares the pixel-intensity distributions for misclassified and correctly classified instances and is supported by descriptive statistics. The correctly classified images (orange) have a higher mean pixel intensity of 61.02 than the misclassified images (blue) do (44.94) (see Tables 20 and 21). The median pixel intensity for correctly classified cases is 67.24, which is significantly greater than that for misclassified cases (49.33). The misclassified images exhibit lower variability (std = 29.81 vs. 36.99) and narrower intensity ranges (see Fig. 24 ).

Table 20.

Descriptive statistics for the pixel intensity of correctly classified and misclassified images.

Statistic Misclassified Correctly Classified
Mean 45 61
Std 30 36
Min 0.4 4
25% 12 23
50% 49 67
75% 75 98
Max 85 113
Interquartile Range 63 75
Table 21.

Descriptive statistics for the pixel intensity of the TP-TN-FP-FN images.

Statistic True Positive True Negative False Positive False Negative
Mean 53 48 50 62
Std 31 28 27 37
Min 2 2 4 1
25% 22 20 27 30
50% 65 49 52 69
75% 79 76 75 87
Max 111 98 114 168
Interquartile Range 57 56 47 56

Pixel intensity of TP-FP-TN-FN tumor modalities

The pixel intensities of 10 TP-FP-TN-FN images were analyzed (see Fig. 21). This boxplot and the descriptive statistics show that the pixel intensity distribution of the false negative class has the highest mean (62.04) and the broadest range (IQR = 56.31), indicating that errors in this category often arise from highly variable or extreme intensities (max = 168.75). The true positives (mean = 53.03, IQR = 57.14) and true negatives (mean = 47.66, IQR = 55.74) show more consistent distributions, with medians of 65.37 and 49.55, respectively, indicating that the model performs well when the intensities align with specific patterns. The false positives (mean = 50.38, IQR = 47.58) have a narrower range and slightly lower variability, suggesting that these errors often occur in regions with less contrast. The model performs best in areas with mid-range pixel intensities and struggles with extreme or variable intensities, particularly in the false-negative class.

Descriptive Statistics for Pixel Intensity:

Development of an SBTM system

Since brain tumor patients require early diagnosis, real-time monitoring, and faster responses from medical practitioners, a sensor-based smart brain tumor management system (SBTM) is proposed. In this proposed system, the following components are required: an IoT camera, an IoT Gateway, a machine learning server (MLS), and an IoT router (see Fig. 21). The MLS serves as the central hub of the system. The trained ML model resides in the MLS. The system proposes using a wireless connection to connect all devices. Edge devices serve as delivery nodes in the IoT framework.

The devices are connected to the hospital database and communicate via secure protocols (e.g., Wi-Fi, Bluetooth, or Ethernet). Appropriate applications, such as web-based systems and mobile app software, are required for seamless integration with existing healthcare systems for data retrieval, processing, and display. A high-level view of the system is presented in Fig. 26.

Fig. 26.

Fig. 26

The proposed IoT-based brain tumor detection and monitoring system.

MRI images are input into the system via IoT Cameras installed in the hospitals or handheld imaging devices. The system supports wireless data transmission from portable imaging devices, enabling real-time processing in resource-constrained settings. Doctors or authorized personnel can upload photos directly through a connected web interface or via external devices, such as USB drives.

The trained M-Net model will reside in MLS. It will collect data from image processing and evaluate whether any diseased leaves are present. If it detects any disease on the MRI image, it will propagate to the user interface. A secure IoT connector connects the MLS to users’ edge devices, such as computers, laptops, or mobile phones.

Layer 1

The MLS processing layer collects data directly from the MRI scanner via Picture Archiving and Communication Systems (PACSs) or Radiology Information Systems (RIS). The MRI images are processed via MLS computing and CNNs for tumor detection and classification. The user interfaces are presented in Figs. 27 and 28.

Fig. 27.

Fig. 27

StreamLit web application interface and sample results.

Fig. 28.

Fig. 28

Android app interface integrated M-Net model.

Layer 2

The tumor detection and classification layer predicts tumors, modalities, and associated probabilities. The results are displayed on IoT edge devices (e.g., a web dashboard or connected mobile application). This allows healthcare professionals to access actionable insights promptly. The brain tumor class label and its confidence score are shown alongside the test image.

Layer 3

IoT-based monitoring provides continuous care, tracks symptoms, and improves the quality of life of brain tumor patients. IoT devices can assist in providing continuous care, tracking symptoms, and improving the quality of life of brain tumor patients. For example, wearable health monitors, seizure monitoring devices, brainwave monitoring systems, and remote patient monitoring devices enable continuous monitoring of brain tumor symptoms, including changes in cognitive function, motor skills, and speech. They can transmit this data to healthcare providers.

The M-Net application was tested on a Realme Note 70 mobile phone as an edge device. The smartphone features a 6300mAh battery, a 90 Hz HD+ display, an IP54 dust- and water-resistant design, a Unisoc T7250 chip, and a 13MP primary camera. The M-Net achieved a modest 1 FPS, indicating that the system processes one frame per second. The device’s battery consumed 21% during operation. The temperature data showed that the CPU temperature did not increase significantly, as the data was recorded as “N/A.” This suggests that the application did not impose a substantial thermal load on the device, even with continuous operation.

While the system has demonstrated functionality on edge devices, such as the Realme Note 70, we acknowledge that the inference speed of 1 FPS may not be sufficient for real-time clinical use. This performance was obtained under specific conditions and on hardware with moderate capabilities. The 1 FPS rate should be considered near real-time rather than real-time. We will revise the manuscript to emphasize that further optimization is needed to meet real-time processing requirements for clinical deployment. Optimizations such as model compression, quantization, or hardware acceleration on more powerful edge devices could significantly improve this performance.

The initial evaluation was conducted primarily on GPU hardware, which is not representative of mobile or edge-device conditions. Therefore, we will conduct additional CPU-only benchmarking on various mobile and edge devices to assess the model’s performance in more realistic settings. Additionally, a memory footprint analysis will be included to evaluate the model’s suitability for edge deployment by assessing its RAM consumption and computational demands on typical edge devices.

ftableDiscussion

In this study, a lightweight CNN, M-Net (Minimal Net), is developed with fewer layers and fewer learnable parameters, and it reliably detects brain tumors in MRI images. The results show that the proposed M-Net has the highest accuracy (99%), outperforming the SOTA CNNs, transfer learning, and ensemble models (see Figs. 29 and 30). The M-Net was trained and evaluated using three Python data generation techniques: TensorFlow/Keras, flow_from_directory, flow_from_dataframe, and flow (with NumPy arrays). It was evident that flow_from_directory provides better results than the other two data pipelines, as it loads images in batches, reducing memory usage rather than loading the entire dataset into RAM. Adam remains stable, while RMSprop fluctuates across models. This suggests that when developing a CNN, careful selection of the optimizer is needed based on architecture-specific convergence behavior.

Fig. 29.

Fig. 29

Comparison of the parameters of the SOTA CNNs and M-Net.

Fig. 30.

Fig. 30

Accuracy comparison among individual CNNs, transfer learning, and ensemble models.

Unlike other studies72–98-99, which only provided machine learning models, this study adopted a practical approach by integrating M-Net in the edge device. Prior studies emphasized the need for real-time detection and monitoring of brain tumors, as tumors may transform into malignant tumors. Brain tumor patients also suffer from headaches, seizures, cognitive and behavioral changes, and neurological deficits; however, sensors allow for continuous monitoring of brain tumor symptoms. Therefore, a real-time detection and monitoring system requires an extensive architecture of IoT devices. We provided the M-Net integrated user-friendly Streamlit-based web interface and Android mobile application. The functionality of the apps, such as Real-time MRI results, a history of MRI results, patient condition data collected using sensors, and data transmission across healthcare professionals, aims to provide an extensive smart brain tumor monitoring architecture. By focusing on deployment-ready design and explainable AI, this work aims to bridge the persistent gap between laboratory research and real-world clinical applications. Figure 27 presents the comparison of the parameters of the SOTA CNNs and M-Net, which establishes M-Net as a suitable model for edge devices rather than SOTA CNNs.

Among the SOTA CNNs, the DenseNet architecture is more efficient in detecting and classifying MRI images. Our findings are also supported by5,55. Raju et al. (2022) applied DenseNet201 and achieved the best result (99.15%) for colorectal disease identification. The study5 reported that the DenseNet architecture utilizes densely flowing skip connections. The DenseNet connection pattern connects all layers directly, maximizing feature extraction. The feed-forward nature of each layer is maintained by receiving input from the previous layers. DenseNet is thus a suitable architecture that can overcome the challenges posed by gradient-based loss functions in medical image analysis56 and by imbalanced classes, leveraging all connected layers in the network.

Many studies reported that transfer learning improves classification accuracy and reduces training time compared with conventional CNNs57,58. For example59, reported 95.07% accuracy in binary classification, and60 reported breast cancer diagnosis using an optimized transfer learning technique, achieving improved accuracy. However, in our study, transfer learning provides negative results. The accuracy is lower than that of the main CNN. The findings of negative transfer learning are supported by studies61,62, and63. Studies have shown that if the input image differs from the trained data, the accuracy will likely decrease.

Most importantly, this study applied three (3) XAI techniques, namely, LIME, SHAP, and Grad-CAM, to generate explanations for M-Net’s detection and classification. The XAIs used in this study enhance the interpretability of M-Net. This interpretability is important because it increases the transparency and acceptability of CNNs. Grad-CAM, designed explicitly for CNNs, generates class-discriminative visualizations by leveraging gradients from the final convolutional layers, effectively localizing regions in the image that influence a specific class prediction. Integrating XAI in M-Net explores our understanding of CNNs’ decision-making processes. Finally, we conducted pixel-intensity analysis using the statistical methods of proper classification, misclassification, true positive, false positive, true negative, and false negative. By analyzing pixel intensities, this study contributes to understanding the growing literature on XAI. This study provides insights into the interpretability of M-Net in brain cancer detection and modality classification.

However, we recognize that the current inference speed of 1 FPS on mobile devices does not fully support real-time clinical deployment. In the revised manuscript, we clarify that although M-Net demonstrates high performance, further optimizations are needed to meet real-time requirements for clinical use. These improvements could include model quantization, pruning, or the use of more powerful edge devices.

We also note that previous studies in brain tumor detection have employed transformer-based models and hybrid CNN-ViT architectures, demonstrating significant performance improvements in certain tasks. To better position our work within the current state of research, we will provide a comparative analysis between M-Net and recent transformer-based and hybrid CNN-ViT models. This analysis will address performance metrics such as accuracy, inference time, and computational efficiency, helping to contextualize M-Net’s contributions and clarify the distinction between traditional CNNs and hybrid models.

Although M-Net outperforms several SOTA CNNs, including DenseNet and VGG, we emphasize that this study primarily focuses on architecture development and integration into an IoT-based tumor-monitoring system. While M-Net is not yet fully optimized for real-time deployment, it lays the groundwork for future research to enhance edge-device deployment in clinical settings.

Future work and conclusions

This research aimed to present a lightweight CNN suitable for SBTM, as tumor patients require real-time monitoring. With the expansion of IoT devices’ capabilities and the customizability of CNNs, SBTM is a possible solution. The main characteristics, therefore, of the lightweight CNN are limited parameters, non-intricate, and straightforward network architectures. Moreover, the model must provide high accuracy. With only 495,972 parameters, we developed M-Net, which was trained and tested on three (3) brain tumor datasets. The datasets have varied brain tumor modalities and MRI image classes (ranging from 100 to 4000). However, the k-fold stratified cross-validation and an unseen test set achieve 97%−99%, placing M-Net as a suitable model for SBTM. The findings suggest that the proposed M-Net effectively classifies brain tumors using MRI images and achieves the highest accuracy (99%) compared to SOTA CNN models.

Furthermore, the M-Net model was compared with six CNN networks from spatial exploitation (VGG19), depth-based (ResNet152v2), multipath (DenseNet201), width-based multiconnection (ResNext101), feature map exploitation (SE-ResNet152), and MobileNetV2. Furthermore, transfer learning and two (2) ensemble models were also compared. Our findings indicate that the M-Net model outperforms the other models. Applying XAI deepens our understanding of how CNNs and M-Net determine brain tumor classification. Our experiment on pixel intensity in correctly and misclassified images also reveals a significant difference between the two. Using pixel analysis in this study, we found that the true-positive, true-negative, false-positive, and false-negative images exhibit distinct pixel intensities. Our rigorous experimental evaluation of models and pixel intensities provides a promising avenue for applying deep learning to the automated diagnosis of brain tumors and other cancer images.

Although this study makes several contributions, it also has limitations. The main limitation of this study is that the M-Net model was not tested in a clinical setting and was not validated with oncologists for real-world adoption; instead, the results were based solely on computational evaluations. The study by64 noted that the accuracy of machine learning models is insufficient, as they have performed poorly in real-world settings. However, the results were evaluated on an unseen dataset, and the test results were manually confirmed against the class. Furthermore, the Android app’s output was manually validated against the ground-truth class. In our defense, we would like to inform you that no organization funded the study to purchase the necessary IoT infrastructure. Moreover, permission from a hospital is being sought for a clinical trial.

The M-Net model was applied to the secondary dataset. In our defense, we would like to point out that biomedical image data collection is a difficult task that was conducted.

Second, the model is applied only to MRI images. In the future, the model should be used to compute tomography (CT), ultrasound (US), X-ray imaging (XR), Plasmid Bluescript (PBS), and microscopy images. Although transfer learning has negative results, future experiments should explore how lightweight CNNs can leverage it. A brain tumor expert did not validate the experimental results. These results are only code-based results. Future studies should include tumor experts to increase the reliability of the findings. This study used Python (TensorFlow/Keras); future studies should experiment with different programming languages and frameworks. Finally, this study used Python (TensorFlow/Keras) for implementation; future work could explore different programming languages and frameworks to optimize performance and scalability.

In conclusion, this study demonstrates the adaptability, customizability, and efficiency of CNNs for detecting and classifying brain tumors from MRI images. The M-Net outperforms SOTA CNNs, transfer learning, and ensemble models (99% accuracy). The integration of LIME, SHAP, and GRAD-CAM visualizes M-Net’s local and global explanations for its predictions. XAI techniques highlight important brain tumor regions for detection and classification. Moving forward, this work can identify two potential future directions: first, CNNs should be tailored to the target task, and second, XAI should be implemented alongside CNNs to explore tumor regions. This study examines a promising approach to developing lightweight CNNs for effective detection and classification of brain tumors.

Finally, our study presents a scientifically grounded, practically oriented approach to improve the accuracy and efficiency of brain tumor detection using the M-Net architecture. The M-net is integrated with a user-friendly Streamlit-based web interface and an Android mobile application to offer a practical, accessible solution for medical practitioners, researchers, and end users. This contribution will support ongoing efforts in the biomedical field to make CAD more practical, interpretable, and widely usable. Furthermore, while the core focus is on brain tumor classification, the broader framework has potential applications in other areas of disease detection, including agricultural disease management, where similar challenges around accessibility and decision support persist.

Author contributions

Md Taimur Ahad: Writing – original draft, project administration, methodology, investigation, formal analysis, conceptualization.Bo Song: Writing and revising– original draft, investigation, formal analysis.Yan Li: Project administration, methodology, formal analysis, conceptualization.

Funding

This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.

Data availability

The datasets used in this research are publicly available at the following links: [https://www.kaggle.com/datasets/sartajbhuvaji/brain-tumor-classification-mri] (https:/www.kaggle.com/datasets/sartajbhuvaji/brain-tumor-classification-mri).[https://www.kaggle.com/datasets/masoudnickparvar/brain-tumor-mri-dataset] (https:/www.kaggle.com/datasets/masoudnickparvar/brain-tumor-mri-dataset).[https://www.kaggle.com/datasets/waseemnagahhenes/brain-tumor-for-14-classes] (https:/www.kaggle.com/datasets/waseemnagahhenes/brain-tumor-for-14-classes).

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Reference list

  • 1.Siegel, R. L., Kratzer, T. B., Giaquinto, A. N., Sung, H. & Jemal, A. Cancer statistics, 2025 (A Cancer Journal for Clinicians, 2025). [DOI] [PMC free article] [PubMed]
  • 2.Guder, O. & Cetin-Kaya, Y. Optimized attention-based lightweight CNN using particle swarm optimization for brain tumor classification. Biomed. Signal Process. Control. 100, 107126 (2025). [Google Scholar]
  • 3.Kibriya, H., Masood, M., Nawaz, M. & Nazir, T. Multiclass classification of brain tumors using a novel CNN architecture. Multimedia tools Appl.81 (21), 29847–29863 (2022). [Google Scholar]
  • 4.Kumar, B. S., Shaik, S. B. & Mulam, H. High-performance compression-based brain tumor detection using a lightweight optimal deep neural network. Adv. Eng. Softw.173, 103248 (2022). [Google Scholar]
  • 5.Verma, A. & Singh, V. P. Design, analysis and implementation of efficient deep learning frameworks for brain tumor classification. Multimedia Tools Appl.81 (26), 37541–37567 (2022). [Google Scholar]
  • 6.Emam, S. M., Rayes, S. M. E., Ali, I. A., Soliman, H. A. & Nafie, M. S. Synthesis of phthalazine-based derivatives as selective anti-breast cancer agents through EGFR-mediated apoptosis: in vitro and in silico studies. BMC Chem.17 (1), 90 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Singh, V. P., Verma, A., Singh, D. K. & Maurya, R. Improved content-based brain tumor retrieval for magnetic resonance images using weight initialization framework with densely connected deep neural network. Neural Comput. Appl.37 (25), 20437–20450. (2023).
  • 8.Özkaraca, O. et al. Multiple brain tumor classification with dense CNN architecture using brain MRI images. Life13 (2), 349 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Shahin, A. I., Aly, S. & Aly, W. A novel multiclass brain tumor classification method based on unsupervised PCANet features. Neural Comput. Appl.35 (15), 11043–11059 (2023). [Google Scholar]
  • 10.Priyadarshini, P., Khairuzzaman, A. K. M. & Kanungo, P. Brain Tumor Detection and Classification from MRI Images Using Cascaded Deep Neural Networks. In Microelectronics, Circuits and Systems: Select Proceedings of Micro2021 (pp. 301–311). Singapore: Springer Nature Singapore. (2023).
  • 11.Ruba, T., Tamilselvi, R. & Beham, M. P. Brain tumor segmentation in multimodal MRI images using novel LSIS operator and deep learning. J. Ambient Intell. Humaniz. Comput.14 (10), 13163–13177 (2023). [Google Scholar]
  • 12.Al-Zoghby, A. M., Al-Awadly, E. M. K., Moawad, A., Yehia, N. & Ebada, A. I. Dual Deep CNN for Tumor Brain Classification. Diagnostics, 13(12), 2050. (2023). [DOI] [PMC free article] [PubMed]
  • 13.Tabatabaei, S., Rezaee, K. & Zhu, M. Attention transformer mechanism and fusion-based deep learning architecture for MRI brain tumor classification system. Biomed. Signal Process. Control. 86, 105119 (2023). [Google Scholar]
  • 14.Zhu, Z. et al. Brain tumor segmentation based on the fusion of deep semantics and edge information in multimodal MRI. Inform. Fusion. 91, 376–387 (2023). [Google Scholar]
  • 15.Reddy & Dhuli, R. A novel lightweight CNN architecture for the diagnosis of brain tumors using MR images. Diagnostics13 (2), 312 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Ishfaq, Q. U. A., Bibi, R., Ali, A., Jamil, F., Saeed, Y., Alnashwan, R. O., … Muthanna,M. S. A. (2025). Automatic smart brain tumor classification and prediction system using deep learning. Scientific Reports, 15(1), 14876. [DOI] [PMC free article] [PubMed]
  • 17.Smith, J., Johnson, M. & Brown, L. The cellular dynamics of brain tumors: Interactions between tumor cells, stromal cells, and immune cells. J. Neurosci. Res.45 (3), 123–145. https://doi.org/xxxxxx (2025).
  • 18.Liu, J. et al. A comprehensive survey of robust deep learning in computer vision: challenges and techniques (Pattern Recognition Letters, Elsevier, 2023).
  • 19.Aamir, M. et al. Brain tumor classification utilizing deep features derived from high-quality regions in MRI images. Biomed. Signal Process. Control. 85, 104988 (2023). [Google Scholar]
  • 20.Loh, J., Dudchenko, L., Viga, J. & Gemmeke, T. Toward hardware supported domain generalization in DNN-based edge computing devices for health monitoring. IEEE Trans. Biomed. Circuits Syst.19 (1), 5–15 (2024). [DOI] [PubMed] [Google Scholar]
  • 21.Bai, X., Wan, Y. & Wang, W. CEPDNet: a fast CNN-based image denoising network using edge computing platform. J. Supercomputing. 81 (1), 100 (2025). [Google Scholar]
  • 22.Sun, H., Fu, R., Wang, X., Wu, Y., Al-Absi, M. A., Cheng, Z., … Sun, Y. (2025). Efficient deep learning-based tomato leaf disease detection through global and local feature fusion. BMC Plant Biology, 25(1), 311. [DOI] [PMC free article] [PubMed]
  • 23.Sailunaz, K. et al. A survey of artificial intelligence/machine learning-based trends for prostate cancer analysis. Netw. Model. Anal. Health Inf. Bioinf.13 (1), 38 (2024). [Google Scholar]
  • 24.Khalighi, S. et al. Artificial intelligence in neuro oncology: Advances and challenges in brain tumor diagnosis, prognosis, and precision treatment. npj Precision Oncol.8, 80. 10.1038/s41698 (2024). 024 00264 0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Díaz Pernas, F. J., Martínez Zarzuela, M., Antón Rodríguez, M. & González Ortega, D. A deep learning approach for brain tumor classification and segmentation using a multiscale convolutional neural network. arXiv10.48550/arXiv.2402.05975 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Rasool, N. & Bhat, J. I. Glioma brain tumor segmentation using deep learning: A review. In Proceedings of the 2023 10th International Conference on Computing for Sustainable Global Development (INDIACom) (pp. 484–489). IEEE. (2023).
  • 27.Hikmah, N. F. et al. Brain tumor detection using a MobileNetV2 SSD model with a modified feature pyramid network. Int. J. Electr. Comput. Eng.14, 3995–4004. 10.11591/ijece.v14i4.3995 (2024). [Google Scholar]
  • 28.Zech, J. R. et al. Confounding variables can degrade generalization performance of radiological deep learning models. arXiv preprint arXiv:1807.00431. (2018).
  • 29.Sathitratanacheewin, S. & Pongpirul, K. Deep learning for automated classification of tuberculosis-related chest X-ray: dataset specificity limits diagnostic performance generalizability. Heliyon, 6 (8), (2010). [DOI] [PMC free article] [PubMed]
  • 30.Roberts, M. et al. Common pitfallsand recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographsand CT scans. Nature Machine Intelligence, 3 (3), 199–217 (2021).
  • 31.Obeid, A. et al. Advancing histopathology with deep learning under data scarcity: A decade in review. arXiv preprint arXiv:2410.19820 (2024).
  • 32.Roberts, M. S. et al. A Study of CNN and Transfer Learning in Medical Imaging: Advantages, Challenges, Future Scope. Sustainability15 (7), 5930 (2025). (mdpi.com). [Google Scholar]
  • 33.Aiya, A. J., Wani, N., Ramani, M., Kumar, A., Pant, S., Kotecha, K., … Al-Danakh,A. (2025). Optimized deep learning for brain tumor detection: a hybrid approach with attention mechanisms and clinical explainability. Scientific Reports, 15(1), 31386. [DOI] [PMC free article] [PubMed]
  • 34.Nahiduzzaman, M., Abdulrazak, L. F., Kibria, H. B., Khandakar, A., Ayari, M. A., Ahamed,M. F., … Kowalski, M. (2025). A hybrid explainable model based on advanced machine learning and deep learning models for classifying brain tumors using MRI images. Scientific Reports, 15(1), 1649. [DOI] [PMC free article] [PubMed]
  • 35.Rezk, E. M. & Mokbel, E. Stereotactic biopsy for multiple intra-axial brain lesions: impact on consequent treatment Regimen. Egypt. J. Neurosurg.38 (1), 6 (2023). [Google Scholar]
  • 36.Hammad, M., ElAffendi, M., Ateya, A. A. & Abd El-Latif, A. A. Efficient brain tumor detection with lightweight end-to-end deep learning model. Cancers15 (10), 2837 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Ma, X., Zhang, M., Kan, Z., Gao, Y. & Li, W. A Lightweight Spatial-Spectral Deformable CNN for UAV Hyperspectral Image Classification (IEEE Transactions on Geoscience and Remote Sensing, 2025).
  • 38.Mengash, H. A., Mahmoud, H. H. & Computers Brain cancer tumor classification from motion-corrected MRI images using Convolutional Neural Network. Mater. Continua, 68(2), 1551–1563. (2021). [Google Scholar]
  • 39.Huang, G., Liu, Z., Van Der Maaten, L. & Weinberger, K. Q. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 4700–4708). (2017).
  • 40.Agarwal, R., Pande, S. D., Mohanty, S. N. & Panda, S. K. A Novel Hybrid System of Detecting Brain Tumors in MRI. IEEE Access.11, 118372–118385 (2023). [Google Scholar]
  • 41.Polat, Ö. & Güngen, C. Classification of brain tumors from MR images using deep transfer learning. J. Supercomputing. 77 (7), 7236–7252 (2021). [Google Scholar]
  • 42.Noreen, N. et al. A deep learning model based on concatenation approach for the diagnosis of brain tumor. IEEE Access.8, 55135–55144 (2020). [Google Scholar]
  • 43.Das, A. & Mohanty, M. N. Classification of magnetic resonance images of brain using concatenated deep neural network. Int. J. Model. Identif. Control. 41 (1–2), 4–11 (2022). [Google Scholar]
  • 44.Sun, C. Cnn models applied in brain cancer diagnosis. In 2021 2nd International Seminar on Artificial Intelligence, Networking and Information Technology (AINIT) (289–293). IEEE. (2021).
  • 45.Fuad, M. S., Anam, C., Adi, K. & Dougherty, G. Comparison of two convolutional neural network models for automated classification of brain cancer types. In AIP Conference Proceedings (Vol. 2346, No. 1). AIP Publishing. (2021), March.
  • 46.Anusha, C. & Kumar, R. Brain Cancer Classification using MR Images Based on VGG-19 Feature Extraction and Ensemble Classifier. In 2023 Third International Conference on Advances in Electrical, Computing, Communication and Sustainable Technologies (ICAECT) (pp. 1–6). (2023), January.
  • 47.Choudhary, A. et al. Interpretability of Causal Discovery in Tracking Deterioration in a Highly Dynamic Process. Sensors24 (12), 3728 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Gaur, L., Bhandari, M., Razdan, T., Mallik, S. & Zhao, Z. Explanation-driven deep learning model for prediction of brain tumor status using MRI image data. Front. Genet.13, 822666 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Nazir, M. I., Akter, A., Wadud, M. A. H. & Uddin, M. A. Utilizing customized CNN for brain tumor prediction with explainable AI. Heliyon, 10 (20), (2024). [DOI] [PMC free article] [PubMed]
  • 50.Saeed, T., Khan, M. A., Hamza, A., Shabaz, M., Khan, W. Z., Alhayan, F., … Baili,J. (2024). Neuro-XAI: explainable deep learning framework based on deeplabV3 + and Bayesian optimization for segmentation and classification of brain tumor in MRI scans.Journal of Neuroscience Methods, 410, 110247. [DOI] [PubMed]
  • 51.Khushi, H. M. T., Masood, T., Jaffar, A., Akram, S. & Bhatti, S. M. Performance analysis of state-of‐the‐art CNN architectures for brain tumour detection. Int. J. Imaging Syst. Technol., 34(1), e22949. (2024).
  • 52.Imran, S. M. A., Arif, M., Jaffar, A., Khushi, H. M. T. & Hussain, A. Deep-Learning Based Multi-Modalities Fusion for the Detection of Brain-Related Diseases: A Review. In International Conference on Computing & Emerging Technologies (pp. 149–170). Cham: Springer Nature Switzerland. (2023), May.
  • 53.Nalepa, J., Marcinkiewicz, M., Kawulok, M. & Wang, Y. Data augmentation for brain-tumor segmentation: a review. Frontiers in computational neuroscience, 13, 83. Ji, Y., & (2019). [DOI] [PMC free article] [PubMed]
  • 54.Khan, A., Sohail, A., Zahoora, U. & Qureshi, A. S. A survey of the recent architectures of deep convolutional neural networks. Artif. Intell. Rev.53, 5455–5516 (2020). [Google Scholar]
  • 55.Narasimha Raju, A. S. et al. CADxPolydetect: a clinically explainable hybrid deep learning system for multi-class colorectal lesion detection using augmented colonoscopy images. BMC Med. Inf. Decis. Mak.25 (1), 335 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Kazeminia, S. et al. GANs for medical image analysis. Artif. Intell. Med.109, 101938 (2020). [DOI] [PubMed] [Google Scholar]
  • 57.Ju, J. et al. Classification of jujube defects in small datasets based on transfer learning. Neural Comput. Appl., 1–14. (2022).
  • 58.Mehmoodi, D., Warfield, S. K. & Gholipour, A. Transfer learning in medical image segmentation: New insights from analysis of the dynamics of model parameters and learned representations. Artif. Intell. Med.116, 102078 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Mahmood, D. A. & Aminfar, S. A. Efficient Machine Learning and Deep Learning Techniques for Detection of Breast Cancer Tumor. BioMed. Target. J.2 (1), 1–13 (2024). [Google Scholar]
  • 60.Emam, S. M., Rayes, S. M. E., Ali, I. A., Soliman, H. A. & Nafie, M. S. Synthesis of phthalazine-based derivatives as selective anti-breast cancer agents through EGFR-mediated apoptosis: in vitro and in silico studies. BMC Chem.17 (1), 90 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Rangarajan Aravind, K. & Raja Automated disease classification in (Selected) agricultural crops using transfer learning. Automatika: časopis za automatiku mjerenje elektroniku računarstvo i komunikacije. 61 (2), 260–272 (2020). [Google Scholar]
  • 62.Mohanty, S. P., Hughes, D. P. & Salathé, M. Using deep learning for image-based plant disease detection. Front. Plant Sci.7, 215232 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Barbedo, J. G. A. Impact of dataset size and variety on the effectiveness of deep learning and transfer learning for plant disease classification. Comput. Electron. Agric.153, 46–53 (2018). [Google Scholar]
  • 64.Steyerberg, E. W., Uno, H., Ioannidis, J. P., Van Calster, B., Ukaegbu, C., Dhingra,T., … Kastrinos, F. (2018). Poor performance of clinical prediction models: the harm of commonly applied methods. Journal of clinical epidemiology, 98, 133–143. [DOI] [PubMed]
  • 65.Hashemzehi, M., Mesic, B., Sjöstrand, B. & Naqvi, M. A comprehensive review of nanocellulose modification and applications in papermaking and packaging: challenges, technical solutions, and perspectives. BioResources17 (2), 3718 (2022). [Google Scholar]
  • 66.Haq, A. U. et al. MCNN: a multi-level CNN model for the classification of brain tumors in IoT-healthcare system. J. Ambient Intell. Humaniz. Comput.14 (5), 4695–4706 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Reddy, C. K. K. et al. A fine-tuned vision transformer based enhanced multi-class brain tumor classification using MRI scan imagery. Front. Oncol.14, 1400341 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Amin, J. et al. P. A new model for brain tumor detection using ensemble transfer learning and quantum variational classifier. Computational intelligence and neuroscience, 2022(1), 3236305. (2022). [DOI] [PMC free article] [PubMed]
  • 69.Tataei Sarshar, N. et al. M., Glioma brain tumor segmentation in four MRI modalities using a convolutional neural network and based on a transfer learning method. In Brazilian technology symposium (pp. 386–402). Cham: Springer International Publishing. (2021), October.
  • 70.Zhang, D. et al. Cross-modality deep feature learning for brain tumor segmentation. Pattern Recogn.110, 107562 (2021). [Google Scholar]
  • 71.Ibrahim, A. U., Engo, G. M., Ame, I., Nwekwo, C. W. & Al-Turjman, F. I-BrainNet: Deep Learning and Internet of Things (DL/IoT)–Based Framework for the Classification of Brain Tumor. J. Imaging Inf. Med., 1–17. (2025). [DOI] [PMC free article] [PubMed]
  • 72.Ahmmed, J., Ahmed, F., Kabir, M. A., Ahad, M. T., Jadoon, M. A., Rehman, A., … LBNet:An optimized lightweight CNN for mammographic breast cancer classification with XAI-based interpretability. Scientific Reports. [DOI] [PMC free article] [PubMed]
  • 73.Ayon, R., Ahad, M. T., Song, B. & Li, Y. R-Net: A reliable and resource-efficient CNN for colorectal cancer detection with XAI integration. arXiv preprint arXiv:2509.16251. (2025).
  • 74.Sagor, S., Ahad, M. T., Ahmed, F., Ayon, R. & Parvin, S. A study on deep convolutional neural networks, transfer learning, and Mnet model for cervical cancer detection. arXiv preprint arXiv:2509.16250. (2025).
  • 75.Ahmed, M. E., Tuhin, H. H., Kayum, M. A., Islam, M. J. & Ahad, M. T. A vision transformer approach to potato leaf disease detection. In 2025 International Conference on Quantum Photonics, Artificial Intelligence. (2025).
  • 76.Mustofa, S., Emon, Y. R., Mamun, S. B., Akhy, S. A. & Ahad, M. T. A novel AI-driven model for student dropout risk analysis with explainable AI insights. Computers Education: Artif. Intell.8, 100352 (2025). [Google Scholar]
  • 77.Mamun, S. B. et al. Grape Guard: A YOLO-based mobile application for detecting grape leaf diseases. J. Electron. Sci. Technol.23 (1), 100300 (2025). [Google Scholar]
  • 78.Islam, R., Ahad, M. T., Ahmed, F., Song, B. & Li, Y. Mental health diagnosis from voice data using convolutional neural networks and vision transformers. Journal Voice (2024). [DOI] [PubMed]
  • 79.Khanom, S. et al. Detection of leg tremors in Parkinson’s disease patients: An experimental wearable leg band solution. In 21st International Conference on Electrical Engineering, Computing. (2024).
  • 80.Ahad, M. T., Payel, I. J., Song, B. & Li, Y. DVS: Blood cancer detection using novel CNN-based ensemble approach. arXiv preprint arXiv:2410.05272. (2024).
  • 81.Preanto, S. A., Ahad, M. T., Emon, Y. R., Mustofa, S. & Alamin, M. A study on deep feature extraction to detect and classify Acute Lymphoblastic Leukemia (ALL). arXiv preprint arXiv:2409.06687. (2024).
  • 82.Ahad, M. T., Mustofa, S., Ahmed, F., Emon, Y. R. & Anu, A. D. A study on deep convolutional neural networks, transfer learning, and ensemble model for breast cancer detection. arXiv preprint arXiv:2409.06699. (2024).
  • 83.Ahad, M. T., Mamun, S. B., Mustofa, S., Song, B. & Li, Y. A comprehensive study on blood cancer detection and classification using convolutional neural networks. arXiv preprint arXiv:2409.06689. (2024).
  • 84.Preanto, S. A., Ahad, M. T., Emon, Y. R., Mustofa, S. & Alamin, M. A study on deep feature extraction to detect and classify Acute Lymphoblastic Leukemia (ALL). arXiv e-prints, arXiv: 2409.06687. (2024).
  • 85.Preanto, S. A., Ahad, M. T., Emon, Y. R., Mustofa, S. & Alamin, M. A semantic segmentation approach on sweet orange leaf diseases detection utilizing YOLO. arXiv e-prints, arXiv: 2409.06671. (2024).
  • 86.Ahad, M. T., Mustofa, S., Ahmed, F., Emon, Y. R. & Anu, A. D. A study on deep convolutional neural networks, transfer learning, and ensemble model for breast cancer detection. arXiv e-prints, arXiv: 2409.06699. (2024).
  • 87.Bhowmik, A. C. et al. A customized vision transformer for accurate detection and classification of Java plum leaf disease. Smart Agricultural Technol.8, 100500 (2024). [Google Scholar]
  • 88.Emon, Y. R., Ahad, M. T. & Rabbany, G. Multi-format open-source sweet orange leaf dataset for disease detection, classification, and analysis. Data Brief.55, 110713 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Mamun, S. B., Ahad, M. T., Morshed, M. M., Hossain, N. & Emon, Y. R. Scratch vision transformer model for diagnosis grape leaf disease. In Proceedings of the Fifth International Conference on Trends in Computational. (2024).
  • 90.Ahmed, F., Emon, Y. R., Ahad, M. T., Munna, M. H. & Mamun, S. B. A fuzzy-based vision transformer model for tea leaf disease detection. In Proceedings of the Fifth International Conference on Trends in Computational. (2024).
  • 91.Bhowmik, A. C., Ahad, M. T., Emon, Y. R., Ahmed, F. & Song, B. A customized vision transformer for accurate detection and classification of Java plum leaf disease. Smart Agricultural Technology (2024).
  • 92.Ahad, M. T., Mustofa, S., Sarker, A. & Emon, Y. R. Bdpapayaleaf: A dataset of papaya leaf for disease detection, classification, and analysis. Classification, and Analysis. (2024). [DOI] [PMC free article] [PubMed]
  • 93.Ahad, M. T., Song, B. & Li, Y. A comparison of convolutional neural network, transfer learning, and ensemble technique for brain tumour detection or classification. Transfer Learning and Ensemble Technique for Brain Tumour Detection. (2024).
  • 94.Bhowmik, A. C., Ahad, M. T. & Emon, Y. R. Machine learning-based Jamun leaf disease detection: A comprehensive review. arXiv preprint arXiv:2311.15741. (2023).
  • 95.Mamun, S. B., Ahad, M. T., Morshed, M. M., Hossain, N. & Emon, Y. R. Scratch vision transformer model for diagnosis grape leaf disease. In International Conference on Trends in Computational and Cognitive. (2023).
  • 96.Ahmed, F., Emon, Y. R., Ahad, M. T., Munna, M. H. & Mamun, S. B. A fuzzy-based vision transformer model for tea leaf disease detection. In International Conference on Trends in Computational and Cognitive. (2023).
  • 97.Biplob, T. I., Rabbany, G., Emon, Y. R., Ahad, M. T. & Fimu, F. A. An optimized vision-based transformer for lung cancer detection. In International Conference on Trends in Computational and Cognitive. (2023).
  • 98.Ahmed, F., Ahad, M. T. & Emon, Y. R. Machine learning-based tea leaf disease detection: A comprehensive review. arXiv preprint arXiv:2311.03240. (2023).
  • 99.Ahad, M. T., Li, Y., Song, B. & Bhuiyan, T. Comparison of CNN-based deep learning architectures for rice diseases classification. Artif. Intell. Agric.9, 22–35 (2023). [Google Scholar]
  • 100.Qureshi, S. A., Raza, S. E. A., Hussain, L., Malibari, A. A., Nour, M. K., Rehman,A. U., … Hilal, A. M. (2022). Intelligent ultra-light deep learning model for multi-class brain tumor detection. Applied Sciences, 12(8), 3715.
  • 101.Qureshi, S. A. et al. Radiogenomic classification for MGMT promoter methylation status using multi-omics fused feature space for least invasive diagnosis through mpMRI scans. Sci. Rep.13, 3291. 10.1038/s41598-023-30309-4 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The datasets used in this research are publicly available at the following links: [https://www.kaggle.com/datasets/sartajbhuvaji/brain-tumor-classification-mri] (https:/www.kaggle.com/datasets/sartajbhuvaji/brain-tumor-classification-mri).[https://www.kaggle.com/datasets/masoudnickparvar/brain-tumor-mri-dataset] (https:/www.kaggle.com/datasets/masoudnickparvar/brain-tumor-mri-dataset).[https://www.kaggle.com/datasets/waseemnagahhenes/brain-tumor-for-14-classes] (https:/www.kaggle.com/datasets/waseemnagahhenes/brain-tumor-for-14-classes).


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES