Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2025 Nov 26;15:42060. doi: 10.1038/s41598-025-26144-4

High-performance parallel multi-scale attention network with explainable AI for intelligent diagnosis of leaf diseases in agricultural systems

R Sudhakar 1, K Nithya 2,✉, C R Dhivyaa 2, C Sharmila 2
PMCID: PMC12658209  PMID: 41298723

Abstract

Detecting leaf diseases is crucial for ensuring crop health and boosting agricultural productivity. An advanced deep learning-based framework is introduced for cassava and groundnut leaf disease detection, incorporating a suite of innovative techniques to enhance classification accuracy. Real-time leaf images are collected from various agricultural environments to capture a wide range of conditions. To improve image quality and segmentation precision, the Contextual Image Enhancement Wiener Filter (CIEWF) is employed for effective noise reduction. Data augmentation is performed using a Generative Adversarial Network (GAN), increasing dataset diversity and improving model generalization. A novel Region of Interest-based Multi-Dimensional Attention Network (ROI-MDAN) is developed to identify and segment critical disease-affected areas within the leaves. For robust feature extraction, the MSFNet-CAM model is proposed, leveraging parallel multi-scale features and incorporating Coordinate Attention to enhance feature fusion and improve classification performance. Furthermore, Gradient-weighted Class Activation Mapping (Grad-CAM) is used to interpret the model’s decision-making process by highlighting the influential regions contributing to disease classification. Experimental results validate the effectiveness of the proposed approach, setting a new benchmark for AI-assisted plant disease diagnosis.

Keywords: Attention network, Feature detection, Segmentation, Disease classification, Noise removal

Subject terms: Computational biology and bioinformatics, Engineering, Mathematics and computing, Plant sciences

Introduction

Plant diseases are a major challenge to global food production, affecting both crop yields and quality of the crops. These plant diseases1 can quickly propagate, causing significant financial losses for farmers and threatening global food security. Timely detection and accurate identification of plant diseases are essential to reducing their effects and preventing extensive harm. Conventionally, diagnosing plant diseases has depended on visual inspection by skilled plant pathologists. While this method has been effective, it is time-consuming, and impractical for monitoring large-scale agricultural fields2. In recent years, advancements in computer vision and Artificial Intelligence methods3 have revolutionized plant disease detection. These technologies can analyze plant images with high accuracy, reducing the need for manual inspections and enabling faster disease identification. Automated systems use AI algorithms to detect symptoms like spots, discoloration, or deformities on leaves, offering real-time insights to farmers. This not only saves time and effort but also helps in applying targeted treatments, reducing the overuse of pesticides and promoting sustainable farming practices.

The process of plant disease detection4–6 begins with preprocessing, which prepares the input images for further analysis. During preprocessing, images are resized to a consistent dimension, ensuring uniformity across the dataset. Noise reduction techniques7 are applied to eliminate irrelevant information, and image enhancement methods improve visibility. Preprocessing ensures the images are of high quality, allowing the model to detect significant patterns more efficiently. To further enhance the dataset, data augmentation is performed using a Generative Adversarial Network based model, which generates synthetic images by learning the distribution of the original data. Scientists8–10 have applied data augmentation strategies and computational models to expand datasets derived from captured plant disease images. This method helps create new samples with similar characteristics to the original data, improving the diversity and robustness of the dataset without the need to physically collect new data. It helps the model by providing a wider variety of images, enhancing its ability to generalize better across different plant disease conditions.

Segmentation11,12 is an essential process in image analysis that breaks down an image into distinct regions, making interpretation easier. In Leaf disease identification, segmentation helps to identify and find areas of interest, such as lesions, discoloration, or other symptoms of diseases on plant leaves. This method allows the model to concentrate on key areas of the image that hold critical information, enhancing both the precision and effectiveness of disease identification. Proper segmentation ensures that unnecessary data is excluded, which helps in reducing computation time and increasing model efficiency. Segmentation techniques such as U-Net13,14, Deep Convolutional Neural Networks15, and Mask R-CNN16 are widely used due to their ability to perform precise separation of disease-related areas. These techniques enable the model to focus on the regions of interest, such as the diseased parts of a leaf, while ignoring irrelevant background information. This improves the accuracy of disease detection by narrowing the analysis to the most critical sections of the image. Moreover, the model can better analyze the specific features indicative of plant diseases by isolating disease-related areas for more precise classification.

After segmentation, Convolutional Neural Networks (ConvNet) are extensively utilized for detecting plant diseases, as it effectively analyzes and recognizes local characteristics like lesions, color variations, and abnormalities in plant leaves. CNNs automatically learn spatial patterns from input images, making them well-suited for detecting early signs of disease without manual intervention. These neural networks are trained using labeled datasets, allowing the model to categorize leaf images as either healthy or affected by disease. Pretrained CNN models, such as VGG1617, ResNet-5018, and DenseNet-12119, have been extensively used in plant disease prediction due to their strong feature extraction capabilities. These models emphasize detecting important local features, like textures and edges, which are essential for identifying patterns linked to diseases. By analyzing these local features, CNNs can detect early signs of disease. These models are trained on large, diverse image datasets and can be fine-tuned to detect plant disease-specific features. By leveraging transfer learning, pretrained models help to decrease training time and enhance the accuracy of the detection system.

Transformers are increasingly being used for plant disease detection to find global features across the leaf image, providing a broader context beyond local patterns. Global feature extraction using Vision Transformers (ViT)20 and Swin Transformers21 has become a powerful approach for plant disease detection. ViT works by dividing images into patches and applying self-attention mechanisms across the entire image, which enables it to find long-range dependencies and identify global patterns. This is particularly useful for detecting symptoms that affect larger areas of the plant, providing a comprehensive understanding of disease spread. Swin Transformers, an improved version of ViT, incorporate a hierarchical structure with shifted windows, allowing them to capture features at multiple scales. The shift technique improves the model’s capability to concentrate on various regions of the image, improving its detection accuracy. By processing contextual information at various levels, Swin Transformers can identify complex disease patterns that span across regions. These global feature extraction techniques enable the model to perform better at detecting plant diseases across expansive areas.

Combining ConvNet for local feature detection and Transformers22 for global feature extraction creates a strong hybrid model for plant disease identification. CNNs are adept at identifying fine-grained, localized features such as lesions, discoloration, and irregularities on plant leaves, which are often indicative of disease. On the other hand, Transformers are designed to capture global contextual information, enabling the model to evaluate connections throughout the entire image. This ability to find both local details and global context enhances the model’s capability to detect diseases that manifest in diverse ways across plant images. The integration of CNNs and Transformers ensures the model can efficiently process complex patterns and it improves the model’s performance by combining the strengths of both architectures, making it more reliable and accurate in identifying plant diseases. The models VGG16, Variational Autoencoder, and Vision Transformer23 networks are utilized to find both local and global features of plant leaf images.

The final step in the plant disease detection pipeline is classification, where the model assigns the detected symptoms to specific disease categories. Advanced classification methods utilize softmax layers or fully connected layers in deep learning models to output probability distributions across predefined classes. By employing robust feature extraction techniques24 and integrating dimensionality reduction, the model attains great accuracy in detecting plant diseases. Furthermore, Explainable AI (XAI) enhances the reliability and trustworthiness of the classification process by providing clear, interpretable explanations for the model’s decisions. Methods such as LIME25, SHAP26, and Grad-CAM27 are essential for offering a deeper understanding of the model’s ability process. LIME explains the prediction by highlighting the most influential features, SHAP assigns importance scores to input features, and Grad-CAM generates heatmaps to visually identify the critical regions of plant images that contributed to the prediction.

The Explainable AI-based parallel Multi-Scale FeatureNet with Coordinate Attention Module (MSFNet-CAM) and ROI-MDAN is proposed to improve segmentation and feature representation for leaf disease classification. The ROI-based Multi-Dimensional Attention Network (ROI-MDAN) is employed for effective segmentation of leaf images. After segmentation, the Multi-Scale FeatureNet with Coordinate Attention Module (MSFNet-CAM) is used for optimal feature extraction and classification. The model leverages architectures like ResNet-50, DenseNet-121, and the Swin Transformer to capture both local and global features. The Coordinate Attention Module refines these features by enhancing spatial and channel dependencies. Finally, Grad-CAM is applied to provide explainability in the model’s decision-making process. By generating heatmaps that highlight the regions of the image contributing most to the disease classification.

The Major contributions of research work are,

  • Cassava plant images and Groundnut plant images were gathered in real-time from multiple agricultural locations, capturing a variety of growth conditions and environmental factors.

  • The Contextual Image Enhancement Wiener Filter (CIEWF) is adapted to improve the clarity of images by removing noise and it helps to make the dataset more suitable for accurate segmentation.

  • Generative Adversarial Network is used to expand the leaf image dataset by creating realistic synthetic sample leaf images, improving dataset diversity and quality for better model training.

  • ROI-based Multi-Dimensional Attention Network (ROI-MDAN) is introduced to detect significant region areas within the leaf images, optimizing the segmentation process by enhancing the focus on critical areas for accurate leaf disease identification.

  • The MSFNet-CAM model is proposed to capture parallel multi-scale feature extraction, incorporating both fine-grained and broad-scale features from leaf images. The Coordinate Attention module refines the fused features for efficient leaf disease classification.

  • Grad-CAM is utilized to analyze and explain the model’s decision process, emphasizing specific areas of leaf images that influence disease classification.

The paper is structured as follows: “Related works” section discusses previous research and related studies. “Explainable AI-driven MSFNet-CAM and ROI-MDAN for advanced plant leaf disease classification” section elaborates on the proposed methodology in detail. “Experimental results and analysis” section presents a comparative analysis of experimental results across various models. Lastly, “Conclusion” section summarizes the key findings and contributions of the study.

Related works

The domain of plant disease identification has evolved significantly over the years, with advancements in Artificial Intelligence and computer vision driving innovative solutions. Recently, several preprocessing methods have been employed to enhance plant leaf images for accurate disease identification. Earlier studies primarily focused on image processing methods, such as smoothing and sharpening filters, to improve image quality and highlight critical features. These techniques were further complemented by the application of specialized filters to effectively remove noise, ensuring that the images retained clarity and were suitable for accurate analysis and classification. The Optimization Assisted Cascaded Filtering Approach7 is introduced to address noise in plant leaf images. The method integrates a two-stage noise removal process to eliminate noise effectively and improve clarity. This approach outperforms traditional filters in terms of noise reduction and preprocessing efficiency.

Large-scale datasets are crucial for deep learning models to ensure robust performance and prevent overfitting. However, capturing real agricultural images for analysis is often limited and challenging. To address this, image augmentation serves as an effective alternative by generating synthetic images under various scenarios. Generative Adversarial Networks (GANs) are useful to create realistic synthetic images that enhance the diversity of training datasets. The SugarcaneGAN28 is proposed to generate the high-quality synthetic images for crop disease detection by addressing the limitations of existing GAN models. The Style-Generative Adversarial Network Adaptive Discriminator Augmentation29 is used for synthesizing rice leaf disease images to address the problem of limited and imbalanced datasets. The C3GAN30 model is utilized for data augmentation to generate synthetic maize disease images. This GAN-based approach enhances the dataset, addressing the issue of limited real-world data. The dual generative adversarial network31 approach is employed to generate high-quality rice leaf disease images.

Segmenting diseased leaf images is an essential process in automated disease identification and classification, forming the backbone of plant monitoring systems. Advanced image processing techniques combined with deep learning, such as U-Net and Mask R-CNN, have been utilized to enhance segmentation accuracy. The U-Net model32 is adapted for efficient segmentation of diseased areas in plant leaves. By integrating InceptionV3 through transfer learning, it leverages pre-trained features to enhance disease identification accuracy. The U-Net model33 is proposed that integrates a residual attention mechanism with improved segmentation techniques to enhance the detection of diseased areas. An improved U-Net model34 is proposed for plant diseased leaf image segmentation by incorporating residual blocks and residual paths to enhance feature transformation and network depth. The Mask R-CNN model35 is employed to accurately segment and classify plant leaf diseases by generating precise masks for diseased areas. It uses deep learning to extract spatial features, allowing for robust detection of disease symptoms, such as irregular shapes and varying sizes, in leaf images. The Mask R-CNN36 architecture is enhanced with hierarchical and threshold-based hierarchical masks to improve pest detection in crop leaves. By incorporating region-based features and a fault-tolerant mechanism, the model efficiently segments and classifies pest-invaded areas. The Mask R-CNN model37 is integrated with attention mechanisms38 for detecting defective pennywort leaves. These enhancements improved the model’s ability to extract and focus on crucial features in the images.

In plant disease identification, Convolutional Neural Networks have become essential due to their strong ability to analyze and capture fine details from images. By applying convolution operations, CNNs capture intricate patterns such as leaf textures, edges, and fine details critical for distinguishing healthy and diseased areas. These local features are extracted hierarchically by allowing the network to focus on small, localized significant region areas within the leaf images. To improve the effectiveness of CNNs in extracting localized features, researchers often employ knowledge transfer techniques, where pre-learned weights are adjusted and optimized for specific plant disease datasets. CNN architectures39, including VGG, Residual-Net, and DenseNet, demonstrate exceptional performance in capturing intricate local features such as surface details, contours, and structural patterns from tomato plant images. The VGG, ResNet, and DenseNet40 models utilize hierarchical feature learning to enhance classification accuracy and robustness. DenseNet201 model41 is utilized for detecting diseases in tomato, potato, and pepper plants, showcasing high accuracy in diagnosis. The research work42 highlights the effectiveness of a custom CNN for leaf disease detection, capable of extracting intricate features for precise classification. CNN-based model (LeafNet)43 is designed specifically for detecting seven common mango diseases in Bangladesh using region-specific leaf images for precise classification. The convolutional neural network-based models44 such as Densely connected network, Residual Network, VGG-16, and Inception model were fine-tuned for plant disease recognition using the PlantVillage dataset.

Transformer-based approaches have gained significant attention in plant disease detection for their capability to model long-range relationships and extract contextual features from images. Unlike conventional Convolutional Neural Networks (CNNs) that emphasize local feature extraction using convolutional filters, transformers utilize self-attention mechanisms to capture comprehensive relationships throughout the entire image. Vision Transformer (ViT)45 is applied to detect Java Plum leaf diseases by optimizing hyperparameters for improved performance. The MobileViT46 architecture is proposed with an inverted residual structure and CBAM for efficient plant disease detection in crops like wheat, coffee, and rice. Plant Leaf Detection Transformer47 is employed for efficient leaf disease detection in crops. It integrates CBAM with ResNet50 to achieve superior performance in detecting plant diseases using the PlantDoc dataset. A smartphone-based solution48 is developed using a Vision Transformer model with self-attention mechanisms for identifying healthy and diseased tomato plants. It extracts relevant features from leaf images and outperformed Inception V3 in testing accuracy. A lightweight and improved Vision Transformer network49 is adapted for identifying maize plant diseases. It incorporates a custom inception block and skip connections by enhancing feature extraction and performance in plant disease detection. A novel transformer block50 is introduced with soft split token embedding and inception architecture for improved plant disease prediction. It outperforms CNN and vision transformer-based models by demonstrating exceptional accuracy on datasets such as VillagePlant, Ibean, AI2018, and PlantDoc. The Vision Transformer-based approach for plant disease localization and classification51 incorporates parallel Multi-scale attention, collaborative attention, and cross-layer attention mechanisms to improve disease detection accuracy.

Recent research has investigated diverse deep learning models for identifying plant diseases, with Convolutional Neural Networks (CNNs) being extensively utilized for capturing essential features and performing categorization. However, Vision Transformers have become increasingly popular for their effectiveness in capturing long-range dependencies, achieving remarkable results in image classification. Hybrid CNN-Transformer models have also been proposed by integrating the advantages of both frameworks to enhance the performance. The proposed hybrid model52 integrates the advantages of both transformers and CNN to enhance the efficiency of plant disease identification. The integration of Vision Transformer and convolutional neural network blocks53 is employed for real-time automated plant disease classification. This model has been evaluated on multiple datasets by providing visual information for farmers to take necessary preventive measures in crops such as wheat and rice. The Leaf Disease Identification Network54 integrates both transformer and CNN architectures to accurately determine plant species, diagnose leaf diseases, and assess severity.

Recent advancements in plant disease detection have focused on integrating explainable AI (XAI) techniques to enhance model transparency and trustworthiness. Methods such as Grad-CAM, SHAP, and LIME have been applied to provide insights into deep learning model predictions, helping users to understand the basis for disease classification. The Grad-CAM55 explainable AI technique is employed to visualize important features in plant disease classification. It enhances model transparency by providing insights into prediction decisions for improved confidence in agricultural applications. LIME25 is integrated into the proposed system to provide transparency in model predictions. By visualizing the specific pixels influencing the model’s decision, XAI helps to identify both relevant and misleading features, providing transparent insights into the model’s decision process. LIME56 framework is employed to enhance the interpretability of deep learning models in plant disease classification.

Although many models have been developed for crop disease identification, accurately finding diseases remains a challenging task. Segmentation techniques like U-Net and Mask R-CNN are commonly used in plant disease detection. However, segmentation methods struggle with complex patterns like overlapping regions in diseased leaves. Minor symptoms, such as small lesions, are often hard to detect accurately. These challenges reduce effectiveness in real-world agricultural applications. Few hybrid models have been developed to extract both local and global features for plant disease categorization. However, effective hybrid models with optimization for accurate classification have not been fully explored. XAI methods such as Grad-CAM and LIME improve model transparency by highlighting the features that influence decision-making. However, the use of these methods in enhancing hybrid models for plant disease detection is still underutilized. Enhancing the integration of these methods can provide better interpretability and reliability in predictions. Further research is needed to leverage XAI for improving trust in hybrid models.

Explainable AI-driven MSFNet-CAM and ROI-MDAN for advanced plant leaf disease classification

Explainable AI-Driven MSFNet-CAM and ROI-MDAN is proposed for effective classification of plant leaf diseases. The Contextual Image Enhancement Wiener Filter (CIEWF) is an innovative image processing technique designed to enhance leaf segmentation by addressing both data cleaning and image enhancement. This filter improves image quality by removing noise such as dew drops, dust, and shadows, making the dataset more suitable for segmentation tasks. Additionally, Generative Adversarial Networks (GANs) are utilized to expand the leaf image dataset by creating realistic leaf images. The ROI-based Multi-Dimensional Attention Network (ROI-MDAN) is an advanced deep learning model designed for highly accurate segmentation of crop diseases. The model utilizes spatial and channel attention mechanisms to improve feature representation.

After feature extraction and attention mechanism application, the system uses a Region Proposal Network to detect significant region areas within the leaf images, enhancing the segmentation process. To further improve feature enhancement in leaf disease classification, the parallel Multi-Scale FeatureNet with Coordinate Attention Module (MSFNet-CAM) is proposed. This model integrates ResNet-50, DenseNet-121, and the Swin Transformer for parallel multi-scale feature extraction. ResNet-50 and DenseNet-121 are used to extract local features from leaf images, capturing fine-grained spatial details, while the Swin Transformer extracts global features to capture broader patterns and structures within the images. The local features from ResNet-50 and DenseNet-121 are fused with the global features from the Swin Transformer, creating a comprehensive feature representation that integrates both detailed and contextual information. To refine the fused features, the Coordinate Attention module is applied, enhancing spatial and channel-wise dependencies. This helps the model to focus on critical regions in the leaf images, which are crucial for identifying diseases. The refined feature map is processed through the Softmax layer for accurate classification of leaf diseases. Finally, Explainable AI with Grad-CAM is applied to interpret and visualize the decision process of the model, offering insights into the key regions of the images that contributed to the disease classification. The architectural layout of the proposed model and the flow diagram are illustrated in Fig. 1a, b.

Fig. 1.

Fig. 1

(a) Explainable AI-driven MSFNet-CAM and ROI-MDAN for plant leaf disease classification, (b) Flow diagram.

Advanced image enhancement technique for precise leaf analysis

The Contextual Image Enhancement Weiner Filter (CIEWF)57 is an innovative image processing technique designed to improve leaf segmentation by addressing both data cleaning and image enhancement. It is essential for eliminating unnecessary elements like tiny water droplets, particulate matter, dust, and ambient noise, which often affect the clarity of leaf images. By systematically eliminating these imperfections, CIEWF ensures that the dataset is clean and free from distortions, thus enhancing the accuracy of the subsequent segmentation process. Furthermore, the filter reduces the Mean Square Error (MSE) between the initial and enhanced images while retaining the key characteristics of the leaves. In the enhancement process, CIEWF effectively restores degraded leaf images by dynamically adjusting its parameters according to the distinct features of each image. The filter adapts its parameters in response to the specific attributes of each image, ensuring accurate enhancement. This adaptability makes it effective across different leaf textures and conditions. CIEWF’s adaptability is further demonstrated by its capacity to select the most suitable window substitute for filtering, allowing it to fine-tune its method to several regions of the crop leaf. This versatility ensures superior performance in both rough and smooth textured areas. Ultimately, CIEWF improves the quality of the leaf images, maintaining clarity and accuracy for effective segmentation.

The CIEWF method is applied to minimize image noise, with a focus on the pixel located at (× 1, × 2) as shown in Eq. (1).

graphic file with name d33e518.gif 1
graphic file with name d33e522.gif 2
graphic file with name d33e526.gif 3

where Inline graphic: Input image; Inline graphic: mean; Inline graphic: variance; Inline graphic: variance of noise; Inline graphic: set; Inline graphic: Surrounding region.

Generative adversarial networks for synthetic leaf image creation

Generative Adversarial Networks (GANs)31,58 are a powerful technique for augmenting leaf image datasets by utilizing two neural networks such as the synthetic image generator and fake image discriminator, as shown in Fig. 2. These networks collaborate to generate realistic synthetic leaf images, enhancing the diversity and quality of the dataset. The process includes three primary phases. In the first phase, the Generator takes a latent noise vector, denoted by h, and generates synthetic leaf images from this random input. The Discriminator evaluates these generated images to classify them as real or fake. The Generator strives to improve its output so that the Discriminator cannot distinguish between real and synthetic images. When the Discriminator successfully recognizes the generated images as fake, the Generator’s loss function applies a penalty.

Fig. 2.

Fig. 2

Generative adversarial networks.

The loss function for the Generator is represented as:

graphic file with name d33e579.gif 4

where Inline graphic : synthetic leaf image created by the Generator; Inline graphic: probability output of the Discriminator indicating whether the generated image is real; Inline graphic: number of synthetic images generated in a batch.

In the second phase, the Generator produces multiple synthetic leaf images by passing several noise vectors through it. The Discriminator evaluates these images and determines whether they are authentic or artificially generated. The Discriminator is trained to distinguish between real leaf images and artificially created ones produced by the Generator. When the Discriminator struggles to differentiate between real and generated images, it indicates that the Generator is becoming more effective.

The loss function for the Discriminator is formulated as:

graphic file with name d33e600.gif 5

where Inline graphic: real leaf images from the dataset; Inline graphic: synthetic image produced by the Generator; Inline graphic: Probability score produced by the Discriminator for a real image; Inline graphic: The probability value assigned by the Discriminator to a generated image.

During the third phase, the Discriminator is trained using authentic leaf images to enhance its ability to differentiate between real and generated images. As the Discriminator improves in identifying real images, it pushes the Generator to produce more lifelike images. Over time, this iterative process leads to increasingly accurate synthetic leaf images.

The final goal of the network training is represented by the following objective function

graphic file with name d33e625.gif 6

where Inline graphic: expectation value; Inline graphic: distribution of real images I; Inline graphic: distribution of latent noise vectors; Inline graphic: Discriminator’s output for a real image; Inline graphic: Discriminator’s output for a synthetic image.

ROI based multi-dimensional attention network (ROI-MDAN) for effective segmentation

The ROI-based Multi-Dimensional Attention Network59,60 (ROI-MDAN) is an advanced deep learning model designed for highly accurate leaf disease segmentation and the architecture of ROI-MDAN is illustrated in Fig. 3. It integrates several recent techniques to improve the segmentation process. The system uses ResNet50 as the backbone network, a powerful neural network with convolutional layers designed for hierarchical feature representation. Its multi-layered structure with residual connections facilitates the retention of crucial attributes, improving the model’s capability to identify intricate patterns in leaf images. ResNet50’s skip connections preserve crucial features during training, preventing the loss of important information. This design helps address challenges like vanishing gradients, ensuring efficient learning of complex patterns. This is crucial for effectively capturing the detailed information required for disease detection in leaf images. After the feature extraction by ResNet50, the model incorporates a Feature Pyramid Network (FPN). The FPN generates multi-level feature maps, allowing the model to find both large and small disease patterns effectively. This parallel multi-scale approach enhances the model’s robustness in detecting diseases at varying sizes and locations. It ensures the model can adapt to diverse leaf sizes and different disease manifestations.

Fig. 3.

Fig. 3

ROI based multi-dimensional attention network model.

To further refine the feature representations, the model employs position and channel attention mechanisms. The position attention block focuses on the spatial relationships between features across various positions in the image. It emphasizes the importance of each spatial location, enhancing the model’s understanding of how different parts of the leaf are related, even across varying distances. This attention to spatial context improves disease detection accuracy. Furthermore, the channel attention module refines feature representations by dynamically adjusting their importance, enabling the model to focus on the most significant attributes for disease segmentation. This enhances the model’s effectiveness in recognizing critical disease indicators. After extracting features and applying attention mechanisms, the system uses a Region Proposal Network to detect significant region areas within the leaf images. The RPN generates candidate regions likely to contain diseased areas, which are then filtered through Non-Maximum Suppression (NMS) to eliminate redundant and overlapping proposals, ensuring optimal selection. The remaining ROIs are processed through RoI Align, ensuring precise alignment and scaling, irrespective of their original size in the image. Finally, the model uses the aligned ROIs to generate segmentation masks that highlight the diseased portions of the leaf for accurate segmentation. The ROI-based Multi-Dimensional Attention Network (ROI-MDAN) combines ResNet50, FPN, position attention module and channel attention module, and RPN to increase the segmentation process. This powerful combination allows the model to extract robust features, capture multi-scale information in parallel, and focus on the most significant areas of the leaf image. The attention blocks refine the feature representation by prioritizing important spatial and channel-wise information. ROI-MDAN achieves superior accuracy in detecting and segmenting diseased regions in leaf images.

Parallel multi-scale FeatureNet with coordinate attention module (MSFNet-CAM) for optimal feature enhancement

Parallel Multi-Scale FeatureNet with Coordinate Attention Module (MSFNet-CAM) is proposed for optimal feature enhancement in leaf disease detection. It includes ResNet-50 and DenseNet-121 along with Swin Transformer for It includes ResNet-50 and DenseNet-121 along with Swin Transformer for multi-scale feature extraction in parallel multi-scale feature extraction. ResNet-50 and DenseNet-121 are utilized to extract local features from leaf images by capturing fine-grained spatial details, while the Swin Transformer captures global features to identify broader patterns and structures within the images.

The local features from ResNet-50 and DenseNet-121 are fused with the global features from the Swin Transformer to construct a detailed and robust feature representation by integrating both detailed and contextual information. To further refine the fused features, the Coordinate Attention block is introduced to the feature map, enhancing spatial-based and channel-wise dependencies. This module helps the model to focus on critical regions in the leaf images, which are crucial for identifying diseases. The optimized feature map is passed through the Softmax layer to identify and classify the leaf diseases accurately. Figure 4 illustrates the structure of MSFNet-CAM.

Fig. 4.

Fig. 4

Parallel multi-scale FeatureNet with coordinate attention module (MSFNet-CAM).

Skip connection-based deep residual network for diseased leaf feature analysis

The Deep Residual Network (ResNet-50)61 architecture is an advanced layered neural model that leverages skip connections to overcome the diminishing gradient challenge while improving feature learning and representation. The architecture is composed of 50 layers, grouped into five main blocks, where every module comprises multiple convolution layers and a residual connection. These skip connections enable the network to capture residual features rather than learning direct transformations, making the learning process more efficient. The output of a residual block is computed as the sum of the input and the residual function, represented by

graphic file with name d33e698.gif 7

where Inline graphic: Output of residual block; Inline graphic : Input to the residual block; Inline graphic: Residual function consisting of three convolutional layers, each followed by batch normalization and ReLU activation

graphic file with name d33e716.gif

Inline graphic : Weights of the convolutional layers such as W1,W2,W3; BN: Batch Normalization operation; Inline graphic: ReLU activation function.

Each residual block in ResNet-50 comprises three convolutional layers utilizing filter sizes of 1 × 1, followed by 3 × 3, and concluding with another 1 × 1. The initial 1 × 1 convolution reduces the input dimensions, the 3 × 3 convolution captures local spatial features, and the final stage of 1 × 1 convolution restores the original dimensionality. During feature extraction, the ResNet-50 model receives the input image for processing. The result generated by the designated convolutional layer which is responsible for capturing fine-grained local features and it is represented as:

graphic file with name d33e730.gif 8

Inline graphic: Feature map extracted from the final convolutional layer of Block 5; Inline graphic: input image; Inline graphic: Mapping corresponding to the output of the third block of the fifth stage.

Finally, the extracted local features Inline graphic are flattened into a 1D vector which is denoted by,

graphic file with name d33e753.gif 9

Densely connected neural network for enhanced leaf feature representation

Densely Connected Neural Network62 is an advanced deep learning model characterized by its tightly connected architecture, ensuring that every layer receives inputs from all previous layers, enhancing feature propagation and gradient flow. This design greatly enhances the reuse of extracted features and the smooth propagation of gradients, making it highly suitable for applications that demand detailed local feature analysis. DenseNet-121 consists of 121 layers organized into four dense modules, each succeeded by a transition layer that reduces quantity of extracted feature representations. The dense module allows the network to capture fine-grained local features while maintaining computational efficiency. In each dense block, every layer receives the output from all preceding layers. This leads to the concatenation of feature maps, rather than summing them.

The output of each dense block is computed as follows

graphic file with name d33e767.gif 10

where Inline graphic: output of the dense block; Inline graphic: input to the block; Inline graphic: output of the convolutional layers within the dense block; Inline graphic: weights of the convolutional layers.

Each dense block in DenseNet-121 comprises multiple layers, where each layer typically uses 1 × 1 convolution for feature compression and 3 × 3 convolutions for spatial feature extraction. The number of layers in each block varies but usually falls within the range of 6 to 12 layers, ensuring that each block captures increasingly abstract local features. The local feature extraction process begins with the input leaf image passing through the first layer of convolution, followed by a series of dense blocks. As the image propagates through the network, local features are progressively captured and refined by the concatenation of outputs from all preceding layers.

The result generated by the last layer of the third dense block, which is responsible for capturing detailed local features, is computed as:

graphic file with name d33e792.gif 11

where Inline graphic: output of the last layer of DenseBlock 4.

After DenseBlock 4, a global mean pooling layer reduces the spatial dimensions and outputs a 1D vector that represents the local features extracted from the image.

graphic file with name d33e803.gif 12
graphic file with name d33e807.gif 13

Swin transformer-based multi-head attention approach for global contextual feature extraction

The Swin Transformer model63 is a vision transformer model developed to efficiently capture global features. Initially, the input leaf image is segmented into distinct patches without overlap, then flattened and converted into feature representations through a linear embedding layer. These feature vectors are passed through hierarchical transformer blocks that apply self-attention within fixed-sized windows to process different regions of the image. To extract global features, the model shifts attention windows between consecutive blocks, enabling it to learn relationships across distant areas of the image. This shifting mechanism enables the system to find relationships across distant regions of the image, improving its ability to handle complex visual data. Its hierarchical structure aggregates the features progressively to build a comprehensive image representation. By efficiently capturing global dependencies, it balances computational cost and accuracy, ensuring robust performance.

After the initial transformer block, a patch merging layer is applied to reduce the spatial resolution by combining neighboring patches and stacking these patches depth-wise. This process not only reduces the spatial dimensions but also increases the depth of the feature map. As a result, the model is able to capture more abstract and high-level feature representations in the successive stages of processing. The hierarchical structure of the architecture enables a progressive reduction in spatial resolution while simultaneously increasing depth, which allows the model to find the detailed features and global features from the image. In the later stages, the swin Transformer generates a comprehensive, parallel multi-scale representation of the input leaf image by effectively capturing global context. The model’s capability to capture global relationships allows it to effectively extract essential features for leaf disease classification. Its efficient attention mechanism and hierarchical structure enable the processing of large, high-resolution leaf images while retaining important global details, ensuring accurate feature extraction for effective disease identification.

Figure 5 presents the Swin Transformer module, which consists of two key sub-modules that work together to extract global features effectively. The first sub-module incorporates layer normalization (LN) and a multi-layer perceptron (MLP), combined with a window-based Multi-Head Attention (WMHA) mechanism. The second sub-module follows a similar structure but replaces the WMHA with a shifted window multi-head attention (SWMHA) mechanism. This adaptive shifting approach allows the model to recognize connections between various areas within the image, facilitating comprehensive global feature extraction. Additionally, a Patch Merging module is integrated for downsampling, reducing spatial dimensions while increasing feature depth, which allows the model to efficiently extract hierarchical global representations.

Fig. 5.

Fig. 5

Architecture of swin transformer block with multi-head attention approach.

Window-based multi-head attention (WMHA) Attention mechanism22 is a crucial component in transformer-based models, which plays a significant role in improving performance in plant disease classification. In this model, images are divided into uniform-sized patches, which are then treated as sequences and handled together. The self-attention mechanism works by utilizing three learnable matrices: Queries (Inline graphic), Keys (Inline graphic,), and Values (Inline graphic),). These matrices are generated by multiplying the input sequence with corresponding weight matrices, producing Q, K, and V for the attention mechanism.

graphic file with name d33e853.gif 14

The attention score is calculated by,

graphic file with name d33e859.gif 15

where Inline graphic and Inline graphic : query and key matrices; Inline graphic: dimension of the key vectors.

The result is obtained by multiplying the attention score with the value matrix

graphic file with name d33e878.gif 16

The self-attention framework determines the dot product between the query and key elements, then applies softmax normalization to derive attention weights. The computed attention scores determine the weighted sum of all outputs in the sequence, with each score serving as a weight for its corresponding output.

In traditional self-attention, attention scores are calculated across the entire image. However, in this work, a window-based self-attention approach is used, which computes attention scores within localized windows. The image is initially segmented into n patches, with adjacent patches grouped together to form a window. This method is more computationally efficient, as calculating attention within a smaller window requires less computational power compared to computing it across the entire image. Moreover, a relative position bias is incorporated to improve the effectiveness of the self-attention mechanism. This bias takes into account the relative positions of patches within the image.

The attention mechanism by incorporating the relative position bias, is updated as follows

graphic file with name d33e888.gif 17

where Inline graphic represents the relative position bias, which enables system to acquire the spatial relationships between patches effectively. By employing these techniques, the transformer-based model can extract hierarchical global features that are vital for plant disease classification.

Shifted window-based multi-head attention (SWMHA) The Shifted Window-based Multi-Head Attention (SWMHA)22 mechanism enhances the standard Window Multi-Head Attention (WMHA) by addressing the issue of limited interaction between non-overlapping windows. In SWMHA, the input leaf image is divided into uniform-size patches that are grouped into windows, where self-attention is calculated locally within each window to reduce computational complexity. Unlike WMHA, SWMHA introduces a shift mechanism that cyclically shifts the windows by a fixed number of patches during subsequent attention layers. This shifting process creates overlaps between neighboring windows, allowing information to flow across window boundaries. By facilitating this interaction, SWMHA enables the system to better find global context-based information and spatial relationships within the image, addressing the limitations of localized attention in WMHA. In addition to the window-shifting mechanism, SWMHA incorporates a masking mechanism to manage attention calculations efficiently. When windows are shifted, the masking mechanism ensures that attention is computed only within patches belonging to the same window, preventing unnecessary computations for non-adjacent patches. This combination of shifting and masking mechanisms allows SWMHA to extract hierarchical global features effectively while keeping the computational workload manageable. These features make SWMHA well-suited for leaf disease classification, as it helps in understanding the patterns and spatial relationships in the image for accurate predictions.

SWMHA shifts the attention windows cyclically on the input feature map X, creating overlaps between neighboring windows in subsequent layers

graphic file with name d33e908.gif 18

The shifted windows allow the model to capture information across window boundaries, facilitating global feature extraction.

After each attention block, a patch merging operation is applied to reduce the spatial resolution by combining neighboring patches

graphic file with name d33e916.gif 19

This process reduces the spatial resolution while enhancing the depth of the feature map, enabling the model to capture more intricate and high-level features.

After passing through several transformer blocks and patch merging layers, the resulting feature map Inline graphic is flattened for feature concatenation and final classification

graphic file with name d33e928.gif 20

The Fig. 6 demonstrates the step-by-step process of applying WMHA and SWMHA mechanisms for feature extraction. In Fig. 6a, the input image is segmented into constant size non-overlapping windows, where multiple attention heads are employed within each window independently. Figure 6b introduces a structured shift in window positions, enabling better interaction between adjacent regions and improving feature continuity. The Fig. 6c illustrates the cyclic shift operation by enhancing long-range dependencies and maintaining spatial consistency. Finally, Fig. 6d presents the output with applied masks, preserving spatial alignment and ensuring effective attention computation across the entire image.

Fig. 6.

Fig. 6

Process of WMHA and SWMHA approaches.

Hyperparameter optimization algorithm

Bayesian optimization64 is a sequential approach that fine-tunes the optimal set of hyperparameters by leveraging prior information to estimate the posterior distribution of unknown functions, thereby minimizing redundant computations. This method can identify the optimal value of intricate objective functions with fewer assessments. In the research work, hyperparameters such as batch size, learning rate, and epoch size were optimized using Bayesian optimization to enhance model performance and training efficiency.

To optimize the hyperparameters, the research work followed these steps using HyperOpt library in Python:

  • The objective function is defined to minimize the loss function based on the current hyperparameter settings.

  • Hyperparameters such as batch size, learning rate, and epoch size are specified, along with their search spaces.

  • The TPE algorithm is selected for its ability to efficiently tune hyperparameters by utilizing probability distributions. It iteratively samples and evaluates various hyperparameter combinations to optimize the performance of the system.

  • The algorithm aims to reduce the loss function and converge to the optimal hyperparameters efficiently.

Coordinate attention module for spatial feature enhancement

The Coordinate Attention Module (CAM)65 is introduced to improve the system’s ability to focus on critical spatial locations in the image and it is given in Fig. 7. This module begins by calculating the mean values of the feature maps along the x and y axes using global pooling methods. These mean values are then concatenated to capture the contextual information from both axes. The concatenated features are passed through a feedforward network to generate attention weights for each spatial position. These weights highlight the importance of each feature, enabling the model to prioritize key areas. The selective attention mechanism ensures that the most important regions of the image are given priority during further processing.

Fig. 7.

Fig. 7

Coordinate attention module.

The attention-based weights produced by CAM are used to adjust the original feature maps by enhancing the significant regions while minimizing the less important ones. This rescaling improves the model’s understanding of spatial relationships within the image, which is vital for leaf disease classification. By focusing on the spatially important regions and suppressing irrelevant ones, the Coordinate Attention Module improves decision-making within the model. Moreover, this attention mechanism is computationally efficient, requiring fewer parameters than traditional attention methods. This allows for faster training and inference without affecting the performance. Its low computational cost and high effectiveness make it a valuable addition to the proposed architecture.

Classification

The resulting feature map is processed through a Softmax layer to calculate the likelihood for each class label. The model is optimized with the categorical cross-entropy loss function and the Adam optimizer to facilitate effective training. This enables the model to accurately classify leaf diseases by predicting the correct class labels for each disease type.

The likelihood for each class label is calculated in the following manner

graphic file with name d33e1006.gif 21

where Inline graphic: probability of class Inline graphic the input Inline graphic; Inline graphic: raw output score for class Inline graphic; Inline graphic: total number of classes.

Grad-CAM-based visualization technique for explainable leaf disease classification

Gradient-weighted Class Activation Mapping (Grad-CAM)27,66 is integrated to enhance the interpretability of the proposed MSFNet-CAM and ROI-MDAN model. Figure 8 shows the working of Grad-CAM model which highlights the most relevant regions in the input leaf images that significantly contribute to classification, ensuring transparency and reliability in disease identification. The Grad-CAM process begins with the forward pass of the MSFNet-CAM model, where an input leaf image is processed to generate prediction scores for each disease class. During this stage, local features are extracted from ResNet-50 and DenseNet-121, capturing fine-grained spatial details, while global features are extracted using the Swin Transformer, capturing broader patterns and structures. These multi-scale features are fused and further refined using the Coordinate Attention Module, which enhances spatial and channel-wise dependencies to highlight the most critical region areas of the image. Once the classification is obtained, the backpropagation procedure begins to calculate the gradients of the predicted class scores in relation to the feature maps from the final convolutional layer. These gradients indicate the importance of different spatial locations in determining the final classification. Gradients with higher values highlight regions that had a significant influence on the model’s decision, allowing Grad-CAM to generate a class-specific heatmap.

Fig. 8.

Fig. 8

Visualization of model decision-making using grad-CAM to highlight disease-relevant regions in plant leaf images.

To effectively highlight disease-affected regions, the computed gradients undergo spatial averaging, followed by a weighted aggregation of feature maps. The Rectified Linear Unit function (ReLU) is employed to retain only the positive influences, ensuring that only features contributing to disease classification are preserved. The generated heatmap is then resized to match the dimensions of the original leaf image, allowing for accurate overlay visualization.

The Grad-CAM steps are given as follows,

  • The class score is denoted by Inline graphic and gradient is computed by,
    graphic file with name d33e1069.gif 22
  • Global average pooling is measured by,
    graphic file with name d33e1076.gif 23
    where Inline graphic: Weight representing the significance of feature map h from the final convolutional layer; Inline graphic: Total number of pixels in the feature map; Inline graphic: Spatial locations in the feature map; Inline graphic: Model output for class Inline graphic; Inline graphic: Activation of the feature map at spatial location; Inline graphic: Gradient of the class score in relation to the activation.
  • Weighted Feature Map Combination is represented by,
    graphic file with name d33e1113.gif 24
    where Inline graphic: Grad-CAM heatmap representing the important regions influencing the model’s decision; Inline graphic: Rectified Linear Unit, used to retain only positive influences; Inline graphic: Importance weight for feature map Inline graphic computed from gradients; Inline graphic: Activation of the Inline graphic th feature map.

Experimental results and analysis

The proposed framework for leaf disease detection is examined through its experimental configuration and performance assessment. Various datasets are used in the experiments to assess the system’s effectiveness across different conditions. The framework is tested to validate its robustness and generalizability in detecting leaf diseases.

Model evaluation setup

The benchmark environment for evaluating the proposed Explainable AI-based parallel Multi-Scale FeatureNet with Coordinate Attention Module (MSFNet-CAM) and ROI-MDAN for crop disease identification includes training the neural network with convolutional layers using the Tensor-Based Computational Framework (PyTorch). The system operates on an 11th-generation Intel processor with 8 GB of RAM, ensuring efficient computation. Experiments are conducted using Python with version 3.7 on a Windows operating system. This setup ensures a reliable and efficient platform for testing the proposed framework’s real-world applicability. The hyperparameters of the CNN models, including ResNet-50, DenseNet-121, and the Swin Transformer, are fine-tuned using the Hyperopt library to optimize their performance for effective leaf disease detection. The best hyperparameters obtained after applying Hyperopt include a batch size of 32, number of epochs 30, Adam optimizer, ReLU activation function, learning rate of 0.001, early stopping, and Categorical Cross-Entropy loss function. Hyperopt was used to automatically determine the optimal learning rate and weight decay, allowing independent fine-tuning of these parameters. Early stopping was applied to prevent overfitting and reduce unnecessary computation. These optimizations collectively ensured stable convergence, improved classification accuracy, and robust feature learning across all disease classes.

Dataset collection and augmentation

The model’s performance is evaluated using two datasets: Benchmark Classification dataset and a hand-picked diverse dataset. The data is organized into 80% training data and 20% testing data, ensuring effective training and accurate evaluation. The hand-picked diverse dataset contains a limited number of images, which could impact model performance. To mitigate this limitation, a GAN is adopted to create realistic synthetic images. These generated images help expand the dataset by providing more varied examples for training, enhancing the model’s ability to recognize diverse leaf conditions.

Benchmark classification datasets

The Cassava Disease Classification Dataset67 is adopted to ensure the accurate identification of cassava diseases. This dataset comprises 9430 meticulously annotated images, organized into five classes: bacterial blight label, brown streak label, mosaic label, green mite label, and Healthy class label. The variety within this dataset plays a significant role in enhancing the system’s performance in recognizing cassava diseases. Additionally, the Groundnut Leaf Dataset68, a standard dataset, is employed to evaluate the model’s performance. It includes 10,361 pre-processed images, classified into six categories: Healthy, Early Leaf Spot, Late Leaf Spot, Nutritional Deficiency, Rust, and Early Rust. This dataset is an essential resource for advancing groundnut leaf disease classification. To further boost the diversity and robustness of the training data, 2,000 synthetic images are generated for each dataset using GAN. Sample cassava and groundnut leaf images are depicted in Figs. 9 and 10.

Fig. 9.

Fig. 9

Sample cassava plant leaf disease images from the benchmark Cassava Disease Classification Dataset.

Fig. 10.

Fig. 10

Sample groundnut plant leaf disease images from the benchmark Cassava Disease Classification Dataset.

Hand-picked diverse datasets

The cassava and groundnut leaf images were collected to ensure a diverse and accurate representation of different health conditions and growth stages. All the images were gathered from various locations in the southern region of India, ensuring geographical diversity. The dataset consists of a total of 1000 images, with 500 images representing cassava plants and the other 500 images representing groundnut plants. These images are organized into training sets of data and testing sets of data for optimal model training and evaluation. The cassava dataset is categorized into five groups: one class label for healthy leaves and four class labels for other leaf diseases, while the groundnut dataset contains six categories: one class label for healthy leaves and five class labels for various diseases. These several types are designed to highlight key variations in leaf health, enabling the model to more accurately identify specific conditions. Sample cassava and groundnut leaf images are shown in Figs. 11 and 12.

Fig. 11.

Fig. 11

Hand-picked diverse cassava plant leaf disease images.

Fig. 12.

Fig. 12

Hand-picked diverse groundnut plant leaf disease images.

To enhance the dataset for cassava and groundnut leaf classification, GAN was employed to generate synthetic images. The initial dataset consisted of 1000 images, with 500 images for each plant type such as cassava and groundnut. Through GAN augmentation, an additional 3000 images were generated for each plant type, bringing the total number of images to 7000. This increase in dataset size provided a more diverse set of images, capturing variations in leaf conditions, including healthy leaves and various disease types. The augmentation process allowed the model to better learn the distinguishing features of different leaf diseases. As a result, the model’s generalization and accuracy in disease detection were significantly improved. Sample generated leaf disease images of cassava plant and groundnut plant are depicted in Fig. 13.

Fig. 13.

Fig. 13

Synthetic leaf disease images of cassava and groundnut plants generated using a Generative Adversarial Network.

Performance assessment metrics

The effectiveness of the plant disease identification model is evaluated using various essential measures, including accuracy, precision, recall, F1-score, Intersection over Union (IoU), and Dice coefficient.

graphic file with name d33e1227.gif 25
graphic file with name d33e1231.gif 26
graphic file with name d33e1235.gif 27
graphic file with name d33e1239.gif 28
graphic file with name d33e1243.gif 29
graphic file with name d33e1247.gif 30

Results discussion

The effectiveness of the research framework is evaluated through the results of segmentation models, results of classification models, accuracy and loss graphs analysis, and a detailed comparison with existing methods to highlight its effectiveness.

Comparative analysis of segmentation models

Table 1 presents a comparison of different segmentation frameworks analyzed using standard datasets, specifically the Cassava Disease Classification and Groundnut Leaf datasets. Performance of the system is computed in terms of Intersection over Union (IoU) and Dice coefficients. The Efficient Neural Network (ENet) model shows solid performance across both datasets, achieving 77.62 IoU and 87.40 Dice for the Cassava dataset, and 78.54 IoU and 87.98 Dice for the Groundnut dataset. U-Net outperforms other models with 83.12 IoU and 90.78 Dice for Cassava, and 80.11 IoU and 88.96 Dice for Groundnut, indicating its robustness in segmenting leaf diseases. Mask R-CNN and DeconvNet also deliver competitive results with high Dice and IoU scores, especially in the Cassava dataset. The ROI-MDAN model stands out as the top performer, achieving the highest IoU (90.38) and Dice (94.95) for Cassava, and also delivering strong results on the Groundnut dataset.

Table 1.

Comparison of segmentation frameworks on standard datasets.

Segmentation approaches Cassava crop leaves Groundnut crop leaves
IoU Dice IoU Dice
Efficient neural network (ENet) 77.62 87.40 78.54 87.98
U Net 83.12 90.78 80.11 88.96
DeconvNet 80.12 88.96 78.55 87.99
Mask R-CNN 81.74 89.95 79.14 88.36
ResUNet 78.65 88.05 78.98 88.26
ROI-MDAN 90.38 94.95 82.00 90.11

ROI-MDAN shows the best performance on the hand-picked diverse dataset, achieving the highest IoU and Dice scores of 84.74 and 91.74 for Cassava, and 85.23 and 92.02 for Groundnut, respectively, demonstrating its ability to segment diseased regions accurately. Mask R-CNN and U-Net also perform well, with Mask R-CNN slightly outperforming U-Net on the Groundnut dataset, achieving IoU scores of 81.25 and 80.68 for Cassava, and 82.01 and 81.98 for Groundnut, respectively. Efficient Neural Network (ENet) and DeconvNet provide moderate segmentation results, with DeconvNet showing a slight advantage over ENet, scoring 78.76 and 88.12 for Cassava, and 76.87 and 86.92 for Groundnut. ResUNet performs similarly to DeconvNet, particularly on the Groundnut dataset, with IoU scores of 77.29 for Cassava and 79.11 for Groundnut. These results highlight the effectiveness of ROI-MDAN in handling a variety of leaf disease segmentation tasks, with superior performance across both Cassava and Groundnut datasets, as demonstrated in Table 2.

Table 2.

Comparison of segmentation frameworks on hand-picked diverse dataset.

Segmentation approaches Cassava crop leaves Groundnut crop leaves
IoU Dice IoU Dice
Efficient neural network (ENet) 75.25 85.88 78.25 87.80
U Net 80.68 89.31 81.98 90.10
DeconvNet 78.76 88.12 76.87 86.92
Mask R-CNN 81.25 89.66 82.01 90.12
ResUNet 77.29 87.19 79.11 88.34
ROI-MDAN 84.74 91.74 85.23 92.02

Figure 14 illustrates the segmented diseased areas in Cassava and Groundnut plant leaf images. It highlights the effectiveness of the segmentation model in accurately identifying affected regions. The visual representation helps in analyzing disease patterns for improved diagnosis and classification.

Fig. 14.

Fig. 14

Segmented diseased regions of cassava and groundnut plant leaf images.

Comparative analysis of classification models

The effectiveness of the proposed ROI-MDAN-based MSFNet-CAM framework is validated through a comparative analysis of different CNN architectures for Cassava and Groundnut leaf disease identification. Traditional pre-developed CNN models show competitive results based on accuracy and F1 scores for standard datasets (Cassava Disease Classification and Groundnut Leaf datasets). The ResNet-50 achieving the highest accuracy among the other models for the cassava and groundnut datasets. However, the proposed model significantly outperforms all baseline models, achieving 97.01% accuracy with an F1 score of 98.44% for cassava disease classification and 98.11% accuracy with an F1 score of 99.03% for groundnut leaf classification. This improvement highlights the impact of integrating ROI-based Multi-Dimensional Attention with MSFNet and Coordinate Attention in extracting multi-scale features and refining classification performance. The detailed results are presented in Table 3.

Table 3.

Comparative evaluation of CNN architectures for classification on standard datasets.

Classification approaches Cassava crop leaves Groundnut crop leaves
Accuracy F1 score Accuracy F1 score
VGG16 90.50 94.88 91.41 95.43
ResNet-50 94.22 96.94 95.04 97.41
DenseNet-121 93.69 96.69 94.31 97.02
Xception 91.56 95.48 92.38 95.97
EfficientNet-B0 92.89 96.21 94.08 96.89
Inception v3 88.38 93.66 91.41 95.43
Proposed ROI-MDAN based MSFNet-CAM model 97.01 98.44 98.11 99.03

The performance comparison of various CNN approaces on a hand-picked diverse dataset for cassava and groundnut leaf disease classification highlights the superiority of the proposed ROI-MDAN-based MSFNet-CAM model. Traditional pre-learned CNN approaches show competitive results based on accuracy and F1 scores, their performance varies across datasets(Cassava and Groundnut Datasets). The ResNet-50 yields 92.51% accuracy for cassava and 93.65% for groundnut. However, the proposed model significantly outperforms all baseline models, achieving 95.28% accuracy with an F1 score of 97.52% for cassava disease classification and 96.94% accuracy with an F1 score of 98.43% for groundnut leaf classification. This substantial improvement demonstrates the effectiveness of integrating ROI-based Multi-Dimensional Attention with MSFNet and Coordinate Attention for extracting critical features and enhancing classification accuracy. The detailed results are presented in Table 4.

Table 4.

Performance comparison of CNN models for classification on hand-picked diverse dataset.

Classification approaches Cassava crop leaves Groundnut crop leaves
Accuracy F1 score Accuracy F1 score
VGG16 91.08 95.20 88.80 93.99
ResNet-50 92.51 96.00 93.65 96.68
DenseNet-121 92.2.8 95.71 92.88 96.26
Xception 89.65 94.39 91.08 95.28
EfficientNet-B0 90.94 95.12 91.08 95.28
Inception v3 86.80 92.73 88.22 93.67
Proposed ROI-MDAN based MSFNet-CAM model 95.28 97.52 96.94 98.43

Figures 15 and 16 present the classification output images, displaying both predicted and actual class labels. The images are categorized into five classes for Cassava and six classes for Groundnut, allowing for a detailed evaluation of the model’s accuracy. By analyzing the predicted outputs against the ground truth labels, these values clearly demonstrate the framework’s ability to differentiate various disease categories, emphasizing its classification effectiveness.

Fig. 15.

Fig. 15

Classification results of groundnut plant leaf diseases for accurate identification of various leaf disease types.

Fig. 16.

Fig. 16

Classification results of cassava plant leaf diseases for accurate identification of various leaf disease types.

Figures 17a and 18a present the accuracy and loss graphs of the ROI-MDAN-based MSFNet-CAM model on benchmark datasets. The accuracy curve illustrates the model’s performance over training epochs and it explains its ability to enhance classification accuracy progressively. The loss curve represents the reduction in model error, with a steady decline indicating effective learning. Figures 17b and 18b display the accuracy graph values and loss graph values for the model applied to a hand-picked diverse dataset. The accuracy curve demonstrates the model’s capacity to adapt to varied data distributions, maintaining consistent classification performance across different leaf conditions. A smooth increase in accuracy with minimal fluctuations suggests strong generalization. The loss curve reflects the gradual minimization of classification errors across diverse data samples. Figure 19 provides an overview of the classification accuracy results for Cassava and Groundnut leaf diseases.

Fig. 17.

Fig. 17

Accuracy and loss graph of ROI-MDAN based MSFNet-CAM model on benchmark datasets.

Fig. 18.

Fig. 18

Accuracy and loss graph of ROI-MDAN based MSFNet-CAM model on hand-picked diverse dataset.

Fig. 19.

Fig. 19

Accuracy of cassava leaf disease images and groundnut leaf disease images for evaluating the performance of the proposed classification model.

Ablation study

Table 5 presents an ablation study of the proposed model, evaluating its accuracy across different combinations of ROI-MDAN with various deep learning architectures on both standard and a hand-picked diverse dataset. The results show that the integration of ROI-MDAN with MSFNet and MSFNet-CAM yields the highest accuracy, with MSFNet-CAM achieving 97.01% accuracy for cassava and 98.11% accuracy for groundnut in the benchmark datasets. The model’s performance is further demonstrated in the hand-picked diverse dataset, where MSFNet-CAM achieves 95.28% accuracy for cassava and 96.94% accuracy for groundnut, highlighting its effectiveness in capturing parallel multi-scale features and enhancing classification performance.

Table 5.

Ablation analysis of the proposed framework based on accuracy.

Approaches Benchmark datasets Hand-picked diverse dataset
Cassava crop leaves Groundnut crop leaves Cassava crop leaves Groundnut crop leaves
ROI-MDAN + ResNet-50 93.19 95.14 92.94 93.98
ROI-MDAN + DenseNet-121 93.54 95.02 91.97 92.98
ROI-MDAN + EfficientNet-B0 92.05 94.84 91.72 91.95
ROI-MDAN + VGG16 91.24 91.22 91.27 90.67
ROI-MDAN + Xception 92.38 92.98 90.39 90.44
ROI-MDAN + Inception v3 90.11 92.37 88.56 88.35
ROI-MDAN + MSFNet 95.57 97.15 94.42 95.63
ROI-MDAN + MSFNet-CAM 97.01 98.11 95.28 96.94

Grad-CAM visualization results

Grad-CAM is employed in the proposed model to visualize and interpret the decision-making process of the ROI-based Multi-Dimensional Attention Network (ROI-MDAN) and MSFNet-CAM for leaf disease classification. By analyzing the gradients from the final convolutional layer, Grad-CAM generates heatmaps that highlight the regions of the input leaf images most influential in disease detection. The heatmaps are superimposed on the original images, with warmer colors (reds and yellows) indicating areas of higher importance and cooler colors (blues) representing less significant regions. In this case, Grad-CAM helps to identify the critical regions within the leaf images that the model focuses on during classification. The results reveal that the model effectively prioritizes the disease-affected areas, enhancing its ability to detect and classify various leaf diseases. This visualization provides valuable insights into the model’s internal workings, ensuring transparency and improving the trustworthiness of the disease classification process. The Visualization results using Grad-CAM is shown in Fig. 20.

Fig. 20.

Fig. 20

Visualization results of plant leaf disease classification using Grad-CAM, highlighting the regions of the leaves that most influence the model’s predictions.

SOTA analysis

Various plant disease detection models have been evaluated across different datasets, with accuracy ranging from 82.54% to 96.51% for models like RS-UNet, Inception V3, MC-UNet, and Transformer-embedded ResNet. The proposed Graph-based Dual Optimization-DeepFeatureNet with ROI-CAN model outperforms these approaches, achieving 97.01% accuracy for cassava plant, 98.11% accuracy for groundnut plant on standard datasets, 95.28% accuracy for Cassava plant, and 96.94% for Groundnut on Hand-picked diverse dataset, demonstrating its superior performance in plant disease classification. Table 6 presents a comparative evaluation of different frameworks.

Table 6.

Comparative performance analysis of state-of-the-art models.

Relevant research Techniques Dataset leaf names Accuracy value (%)
Fu et al.69 U-shaped network model Plant village 92.3
Moupojou et al.70 Inception pre-learned V3 model Collected field type plants 83
Yubao et al.71 MC based U-Network model Plant village dataset 88
Hasan et al.72 SVM technique with Gaussian mixture model Betel leaf images 84
Abinaya et al.73 CAAR approach with U-network model Plant village images + coffee leaf images 95.3
Malek et al.74 Inception pre developed V3 model with knowledge transfer methods Betel leaf images 95
Wang et al.75 MFBP-UNet Pear leaf disease image dataset 90.89
Zhou et al.76 Restructured Deep Residual Dense Network Tomato test dataset 95
Yang et al.77 DeepLabv3 model with dense network Open source image dataset of plant leaf 93
Paul et al.78 VGG-19, VGG-16, and custom model Tomato leaf disease dataset 95
Elfatimi et al.79 Mobile pre developed network model Bean leaf image disease dataset 92
Sambasivam et al.80 CNN model Cassava disease classification 93
Chhetri et al.81 MobileNet V2 and VGG16 (Hybrid kernel) Cassava leaves dataset 90.11
Zhong et al.82 Transformer-embedded ResNet Cassava leaf disease dataset 91.12
Sasmal et al.83 InceptionV3 model Groundnut leaf disease dataset 96.51
Proposed model ROI-MDAN based MSFNet-CAM model

Cassava Benchmark dataset,

Groundnut Benchmark dataset,

Cassava Hand-picked diverse dataset,

Groundnut Hand-picked diverse dataset

97.01

98.11

95.28

96.94

Limitations and future work

While the proposed ROI-MDAN based MSFNet-CAM framework has shown high accuracy and strong performance in detecting and classifying cassava and groundnut leaf diseases, there are still some limitations that should be considered. The use of multi-scale feature extraction and attention mechanisms, although effective in improving detection and classification, increases the computational complexity of the model and may require significant computational resources, which could make it difficult to deploy on devices with limited resources. For future work, the framework can be further optimized to reduce computational requirements for deployment on resource-constrained devices, and it can be integrated with real-time monitoring systems to support early disease detection and timely intervention for farmers.

Conclusion

The proposed ROI-MDAN based MSFNet-CAM model has demonstrated remarkable performance in detecting and classifying cassava and groundnut leaf diseases. By leveraging a combination of advanced techniques, including ROI-based Multi-Dimensional Attention Network (ROI-MDAN) for region detection and the MSFNet-CAM model for parallel multi-scale feature extraction, the model efficiently addresses the challenges of disease detection in agricultural environments. Evaluation on multiple datasets, such as the Cassava Benchmark, Groundnut Benchmark, and hand-picked diverse datasets, resulted in impressive accuracy scores of 97.01%, 98.11%, 95.28%, and 96.94%, showcasing its robustness across varying conditions and diverse data. Furthermore, the integration of Grad-CAM for model interpretability enables better visualization of the decision-making process, highlighting critical parts of the leaf images contributes to leaf disease identification. The proposed framework demonstrates high accuracy and strong robustness when tested on multiple datasets, effectively identifying complex disease patterns and even small lesions that are difficult to detect, while the integration of Grad-CAM provides clear visualization of the areas influencing the model’s decisions. These findings showcase the strong potential of the proposed system in refining the accuracy and robustness of plant disease diagnosis. By offering a powerful tool for early diagnosis, this approach contributes to enhancing agricultural practices and promoting sustainable crop management, ensuring better protection for crops like cassava and groundnut.

Author contributions

R. wrote the main manuscript, K. wrote the main manuscript, C.R. wrote the main manuscript and C. wrote the main manuscript.

Data availability

The datasets are publicly available in a standard repository. Cassava Disease Classification Dataset is available at https://www.kaggle.com/c/cassava-disease/data and Groundnut Leaf Dataset is available at https://data.mendeley.com/datasets/22p2vcbxfk/3.

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Change history

12/18/2025

The original online version of this Article was revised: In the original version of this Article R. Sudhakar was incorrectly affiliated with ‘School of Computer Science and Engineering, Vellore Institute of Technology, Vellore, Tamilnadu, India.’ Their correct affiliation is ‘Department of Computer Science and Engineering, Nandha College of Technology, Erode, Tamilnadu, India.’ Furthermore, K. Nithya, C. R. Dhivyaa and C. Sharmila were incorrectly affiliated with ‘Department of Computer Science and Engineering, Nandha College of Technology, Erode, Tamilnadu, India.’ Their correct affiliation is ‘School of Computer Science and Engineering, Vellore Institute of Technology, Vellore, Tamilnadu, India.’ The original Article has been corrected.

References

  • 1.Upadhyay, N. & Gupta, N. Detecting fungi-affected multi-crop disease on heterogeneous region dataset using modified ResNeXt approach. Environ. Monit. Assess.196, 610. 10.1007/s10661-024-12790-0 (2024). [DOI] [PubMed] [Google Scholar]
  • 2.Upadhyay, N. & Bhargava, A. Artificial intelligence in agriculture: Applications, approaches, and adversities across pre-harvesting, harvesting, and post-harvesting phases. Iran. J. Comput. Sci.8, 749–772. 10.1007/s42044-025-00264-6 (2025). [Google Scholar]
  • 3.Bhargava, A. et al. Plant leaf disease detection, classification, and diagnosis using computer vision and artificial intelligence: A review. IEEE Access12, 37443–37469. 10.1109/ACCESS.2024.3373001 (2024). [Google Scholar]
  • 4.Kotwal, J., Kashyap, R. & Pathan, S. Agricultural plant diseases identification: From traditional approach to deep learning. Mater. Today Proc.80(1), 344–356. 10.1016/j.matpr.2023.02.370 (2023). [Google Scholar]
  • 5.Kotwal, J. G., Kashyap, R. & Shafi, P. M. Artificial driving based EfficientNet for Automatic plant leaf disease classification. Multimed. Tools Appl.83, 38209–38240. 10.1007/s11042-023-16882-w (2024). [Google Scholar]
  • 6.Kotwal, J., Kashyap, R., Shafi, P. M. & Kimbahune, V. Enhanced leaf disease detection: UNet for segmentation and optimized EfficientNet for disease classification. Softw. Impacts22, 100701. 10.1016/j.simpa.2024.100701 (2024). [Google Scholar]
  • 7.Bhujade, V. G., Sambhe, V. & Banerjee, B. Digital image noise removal towards soybean and otton plant disease using image processing filters. Expert Syst. Appl.246, 123031. 10.1016/j.eswa.2023.123031 (2024). [Google Scholar]
  • 8.Antwi, K., Bennin, K. E., Asiedu, D. K. P. & Tekinerdogan, B. On the application of image augmentation for plant disease detection: A systematic literature review. Smart Agric. Technol.9, 100590. 10.1016/j.atech.2024.100590 (2024). [Google Scholar]
  • 9.Min, B., Kim, T., Shin, D. & Shin, D. Data augmentation method for plant leaf disease recognition. Appl. Sci.13(3), 1465. 10.3390/app13031465 (2023). [Google Scholar]
  • 10.Li, M. et al. FWDGAN-based data ugmentation for tomato leaf disease identification. Comput. Electron. Agric.194, 106779. 10.1016/j.compag.2022.106779 (2022). [Google Scholar]
  • 11.Polly, R. & Devi, E. A. Semantic segmentation for plant leaf disease classification and damage detection: A deep learning approach. Smart Agric. Technol.9, 100526. 10.1016/j.atech.2024.100526 (2024). [Google Scholar]
  • 12.Upadhyay, N. & Gupta, N. SegLearner: A segmentation based approach for predicting disease severity in infected leaves. Multimed. Tools Appl.10.1007/s11042-025-20838-7 (2025). [Google Scholar]
  • 13.Wang, H. et al. MFBP-UNet: A network for pear leaf disease segmentation in natural agricultural environments. Plants.12(18), 3209. 10.3390/plants12183209 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Muhammad, S. et al. Deep learning-based segmentation and classification of leaf images for detection of tomato plant disease. Front. Plant Sci.10.3389/fpls.2022.1031748 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Trivedi, P. et al. Plant leaf disease detection and classification using segmentation encoder techniques. Open Agric. J.18, e18743315321139. 10.2174/0118743315321139240627092707 (2024). [Google Scholar]
  • 16.Bondre, S. V., Chinnaiah, K. & Bondre, V. D. Application of M-RCNN for prompt segmentation between infected tomato leaves and healthy tomato leaves. J. Phytopathol.10.1111/jph.13363 (2024). [Google Scholar]
  • 17.Maria, T. et al. Corn leaf disease: Insightful diagnosis using VGG16 empowered by explainable AI. Front. Plant Sci.10.3389/fpls.2024.1402835 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Askr, H., El-dosuky, M., Darwish, A. & Hassanien, A. E. Explainable ResNet50 learning model based on copula entropy for cotton plant disease prediction. Appl. Soft Comput.164, 112009. 10.1016/j.asoc.2024.112009 (2024). [Google Scholar]
  • 19.Pandey, A. & Jain, K. A robust deep attention dense convolutional neural network for plant leaf disease identification and classification from smart phone captured real world images. Ecol. Inform.70, 101725. 10.1016/j.ecoinf.2022.101725 (2022). [Google Scholar]
  • 20.Chang, B., Wang, Y., Zhao, X., Li, G. & Yuan, P. A general-purpose edge-feature guidance module to enhance vision transformers for plant disease identification. Expert Syst. Appl.237(C), 121638. 10.1016/j.eswa.2023.121638 (2024). [Google Scholar]
  • 21.Guo, Y., Lan, Y. & Chen, X. CST: Convolutional Swin Transformer for detecting the degree and types of plant diseases. Comput. Electron. Agric.202, 107407. 10.1016/j.compag.2022.107407 (2022). [Google Scholar]
  • 22.Si, H. et al. A dual-branch model integrating CNN and swin transformer for efficient apple leaf disease classification. Agriculture14(1), 142. 10.3390/agriculture14010142 (2024). [Google Scholar]
  • 23.Nagachandrika, B., Prasath, R. & PraveenJoe, I. R. An automatic classification framework for identifying type of plant leaf diseases using multi-scale feature fusion-based adaptive deep network. Biomed. Signal Process. Control95(A), 106316. 10.1016/j.bspc.2024.106316 (2024). [Google Scholar]
  • 24.Upadhyay, N., Sharma, D. K. & Bhargava, A. 3SW-Net: A feature fusion network for semantic weed detection in precision agriculture. Food Anal. Methods18, 2241–2257. 10.1007/s12161-025-02852-5 (2025). [Google Scholar]
  • 25.Oad, A. et al. Plant leaf disease detection using ensemble learning and explainable AI. IEEE Access12, 156038–156049. 10.1109/ACCESS.2024.3484574 (2024). [Google Scholar]
  • 26.Ahmed, I., Ahmad, M., Ghazouani, H., Barhoumi, W. & Jeon, G. Intelligent computing for crop monitoring in CIoT: Leveraging AI and big data technologies. Expert Syst.10.1111/exsy.13786 (2025). [Google Scholar]
  • 27.Karim, M. J. et al. Enhancing agriculture through real-time grape leaf disease classification via an edge device with a lightweight CNN architecture and Grad-CAM. Sci Rep14, 16022. 10.1038/s41598-024-66989-9 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Li, X. et al. SugarcaneGAN: A novel dataset generating approach for sugarcane leaf diseases based on lightweight hybrid CNN-Transformer network. Comput. Electron. Agric.219, 108762. 10.1016/j.compag.2024.108762 (2024). [Google Scholar]
  • 29.Haruna, Y., Qin, S. & Mbyamm Kiki, M. J. An improved approach to detection of rice leaf disease with GAN-based data augmentation pipeline. Appl. Sci.13(3), 1346. 10.3390/app13031346 (2023). [Google Scholar]
  • 30.Thakur, P. S., Chaturvedi, S., Khanna, P., Sheorey, T. & Ojha, A. Real-time plant disease identification: Fusion of vision transformer and conditional convolutional network with C3GAN-based data augmentation. IEEE Trans. AgriFood Electron.2(2), 576–586. 10.1109/TAFE.2024.3447792 (2024). [Google Scholar]
  • 31.Zhang, Z., Gao, Q., Liu, L. & He, Y. A high-quality rice leaf disease image data augmentation method based on a dual GAN. IEEE Access11, 21176–21191. 10.1109/ACCESS.2023.3251098 (2023). [Google Scholar]
  • 32.AlArfaj, A. A. et al. Multi-step preprocessing with UNet segmentation and transfer learning model for pepper bell leaf disease detection. IEEE Access11, 132254–132267. 10.1109/ACCESS.2023.3334428 (2023). [Google Scholar]
  • 33.Jianlong, W., Junhao, J., Yake, Z., Haotian, W. & Shisong, Z. RAAWC-UNet: An apple leaf and disease segmentation method based on residual attention and atrous spatial pyramid pooling improved UNet with weight compression loss. Front. Plant Sci.10.3389/fpls.2024.1305358 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Zhang, S. & Zhang, C. Modified U-Net for plant diseased leaf image segmentation. Comput. Electron. Agric.204, 107511. 10.1016/j.compag.2022.107511 (2023). [Google Scholar]
  • 35.Selvam, N. & Joy, J. K. Plant leaf disease detection with multivariable feature selection using deep learning AEN and mask R-CNN in PLANT-DOC data. Biotech. Res. Asia21(4), 1649 (2024). [Google Scholar]
  • 36.Wu, X. et al. Image segmentation for pest detection of crop leaves by improvement of regional convolutional neural network. Sci. Rep.14, 24160. 10.1038/s41598-024-75391-4 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Chowdhury, M. et al. Defective Pennywort leaf detection using machine vision and mask R-CNN model. Agronomy14(10), 2313. 10.3390/agronomy14102313 (2024). [Google Scholar]
  • 38.Kotwal, J. G., Kashyap, R. & Shafi, P. M. Yolov5-based convolutional feature attention neural network for plant disease classification. Int. J. Intell. Syst. Technol. Appl. (IJISTA)10.1504/IJISTA.2024.140949 (2024). [Google Scholar]
  • 39.Islam, S. U. et al. Enhanced deep learning architecture for rapid and accurate tomato plant disease diagnosis. AgriEngineering.6(1), 375–395. 10.3390/agriengineering6010023 (2024). [Google Scholar]
  • 40.Sarkar, C., Gupta, D., Gupta, U. & Hazarika, B. B. Leaf disease detection using machine learning and deep learning: Review and challenges. Appl. Soft Comput.145, 110534. 10.1016/j.asoc.2023.110534 (2023). [Google Scholar]
  • 41.Bakr, M., Abdel-Gaber, S., Nasr, M. & Hazman, M. DenseNet based model for plant diseases diagnosis. Eur. J. Electr. Eng. Comput. Sci.6(5), 1–9. 10.24018/ejece.2022.6.5.458 (2022). [Google Scholar]
  • 42.Alam, T. S., Jowthi, C. B. & Pathak, A. Comparing pre-trained models for efficient leaf disease detection: A study on custom CNN. J. Electr. Syst. Inf. Technol.11, 12. 10.1186/s43067-024-00137-1 (2024). [Google Scholar]
  • 43.Rizvee, R. A. et al. LeafNet: A proficient convolutional neural network for detecting seven prominent mango leaf diseases. J. Agric. Food Res.14, 100787. 10.1016/j.jafr.2023.100787 (2023). [Google Scholar]
  • 44.Eunice, J., Popescu, D. E., Chowdary, M. K. & Hemanth, J. Deep learning-based leaf disease detection in crops using images for agricultural applications. Agronomy.12(10), 2395. 10.3390/agronomy12102395 (2022). [Google Scholar]
  • 45.Bhowmik, A. C. et al. A customised vision transformer for accurate detection and classification of Java Plum leaf disease. Smart Agric. Technol.8, 100500. 10.1016/j.atech.2024.100500 (2024). [Google Scholar]
  • 46.Guoqiang, L., Yuchao, W., Qing, Z., Peiyan, Y. & Baofang, C. PMVT: A lightweight vision transformer for plant disease identification on mobile devices. Front. Plant Sci.10.3389/fpls.2023.1256773 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Li, W., Zhu, L. & Liu, J. PL-DINO: An improved transformer-based method for plant leaf disease detection. Agriculture14, 691. 10.3390/agriculture14050691 (2024). [Google Scholar]
  • 48.Barman, U. et al. ViT-SmartAgri: Vision transformer and smartphone-based plant disease detection for smart agriculture. Agronomy14, 327. 10.3390/agronomy14020327 (2024). [Google Scholar]
  • 49.Pushkar, G., Punam, B., Sudeep, M., Ashraful, H. M. & Kumar, D. C. TrIncNet: A lightweight vision transformer network for identification of plant diseases. Front. Plant Sci.10.3389/fpls.2023.1221557 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Yu, S., Xie, L. & Huang, Q. Inception convolutional vision transformers for plant disease identification. Internet Things21, 100650. 10.1016/j.iot.2022.100650 (2023). [Google Scholar]
  • 51.Hemalatha, S. & Jayachandran, J. J. B. A multitask learning-based vision transformer for plant disease localization and classification. Int. J. Comput. Intell. Syst.17, 188. 10.1007/s44196-024-00597-3 (2024). [Google Scholar]
  • 52.Thakur, P. S., Chaturvedi, S., Khanna, P., Sheorey, T. & Ojha, A. Vision transformer meets convolutional neural network for plant disease classification. Ecol. Inform.77, 102245. 10.1016/j.ecoinf.2023.102245 (2023). [Google Scholar]
  • 53.Borhani, Y., Khoramdel, J. & Najafi, E. A deep learning based approach for automated plant disease classification using vision transformer. Sci. Rep.12, 11554. 10.1038/s41598-022-15163-0 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Yang, B. et al. A novel plant type, leaf disease and severity identification framework using CNN and transformer with multi-label method. Sci. Rep.14, 11664. 10.1038/s41598-024-62452-x (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Natarajan, S., Chakrabarti, P. & Margala, M. Robust diagnosis and meta visualizations of plant diseases through deep neural architecture with explainable AI. Sci. Rep.14, 13695. 10.1038/s41598-024-64601-8 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Nigar, N., Muhammad Faisal, H., Umer, M., Oki, O. & Manappattukunnel Lukose, J. Improving plant disease classification with deep-learning-based prediction model using explainable artificial intelligence. IEEE Access12, 100005–100014. 10.1109/ACCESS.2024.3428553 (2024). [Google Scholar]
  • 57.Allaoua Chelloug, S., Alkanhel, R., Muthanna, M. S. A., Aziz, A. & Muthanna, A. MULTINET: A multi-agent DRL and EfficientNet assisted framework for 3D plant leaf disease identification and severity quantification. IEEE Access11, 86770–86789. 10.1109/ACCESS.2023.3303868 (2023). [Google Scholar]
  • 58.Haruna, Y., Qin, S. & Mbyamm Kiki, M. J. An improved approach to detection of rice leaf disease with GAN-based data augmentation pipeline. Appl. Sci.13, 1346. 10.3390/app13031346 (2023). [Google Scholar]
  • 59.Natesan, B. et al. Channel-spatial segmentation network for classifying leaf diseases. Agriculture12(11), 1886. 10.3390/agriculture12111886 (2022). [Google Scholar]
  • 60.Wang, D. & He, D. Fusion of Mask RCNN and attention mechanism for instance segmentation of apples under complex background. Comput. Electron. Agric.196, 106864. 10.1016/j.compag.2022.106864 (2022). [Google Scholar]
  • 61.Vallabhajosyula, S., Sistla, V. & Kolli, V. K. K. A novel hierarchical framework for plant leaf disease detection using residual vision transformer. Heliyon10(9), e29912. 10.1016/j.heliyon.2024.e29912 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Banothu, S. et al. Plant disease identification and pesticides recommendation using Dense Net. Cogent Eng.10.1080/23311916.2024.2353080 (2024). [Google Scholar]
  • 63.Kalpana, P. et al. Plant disease recognition using residual convolutional enlightened Swin transformer networks. Sci. Rep.14, 8660. 10.1038/s41598-024-56393-8 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar] [Retracted]
  • 64.Şahin, E., Özdemir, D. & Temurtaş, H. Multi-objective optimization of ViT architecture for efficient brain tumor classification. Biomed. Signal Process. Control91, 105938. 10.1016/j.bspc.2023.105938 (2024). [Google Scholar]
  • 65.Karthik, R. et al. GrapeLeafNet: A dual-track feature fusion network with inception-ResNet and shuffle-transformer for accurate grape leaf disease identification. IEEE Access12, 19612–19624. 10.1109/ACCESS.2024.3361044 (2024). [Google Scholar]
  • 66.Assaduzzaman, M. et al. XSE-TomatoNet: An explainable AI based tomato leaf disease classification method using EfficientNetB0 with squeeze-and-excitation blocks and multi-scale feature fusion. MethodsX14, 103159. 10.1016/j.mex.2025.103159 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.https://www.kaggle.com/c/cassava-disease/data.
  • 68.Manvikar, A. & Reddy, P. Dataset of groundnut plant leaf images for classification and detection. Mendeley Data10.17632/22p2vcbxfk.3 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Fu, J., Zhao, Y. & Wu, G. Potato leaf disease segmentation method based on improved UNet. Appl. Sci.13(20), 11179. 10.3390/app132011179 (2023). [Google Scholar]
  • 70.Moupojou, E. et al. FieldPlant: A dataset of field plant images for plant disease detection and classification with deep learning. IEEE Access11, 35398–35410. 10.1109/ACCESS.2023.3263042 (2023). [Google Scholar]
  • 71.Deng, Y. et al. An effective image-based tomato leaf disease segmentation method using MC-UNet. Plant Phenomics.15(5), 0049. 10.34133/plantphenomics.0049 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Hasan, M. Z., Zeba, N., Malek, M. A. & Reya, S. S. A leaf disease classification model in betel vine using machine learning techniques. In 2021 2nd International Conference on Robotics, Electrical and Signal Processing Techniques (ICREST), DHAKA, Bangladesh, 362–366. 10.1109/ICREST51555.2021.9331142 (2023).
  • 73.Abinaya, S., Kumar, K. U. & Alphonse, A. S. Cascading autoencoder with attention residual U-Net for multi-class plant leaf disease segmentation and classification. IEEE Access11, 98153–98170. 10.1109/ACCESS.2023.3312718 (2023). [Google Scholar]
  • 74.Malek, A., Basak, S. & Reya, S. S. An approach to identify diseases in betel leaf using deep learning techniques. In 2022 4th International Conference on Sustainable Technologies for Industry 4.0 (STI). Location of Conference, Bangladesh Date of Conference, 1–5 (2022).
  • 75.Wang, H. et al. MFBP-UNet: A network for pear leaf disease segmentation in natural agricultural environments. Plants12(18), 3209. 10.3390/plants12183209 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Zhou, C., Zhou, S., Xing, J. & Song, J. Tomato leaf disease identification by restructured deep residual dense network. IEEE Access9, 28822–28831. 10.1109/ACCESS.2021.3058947 (2021). [Google Scholar]
  • 77.Yang, T., Zhou, S., Xu, A., Ye, J. & Yin, J. An approach for plant leaf image segmentation based on YOLOV8 and the improved DEEPLABV3+. Plants12(19), 3438. 10.3390/plants12193438;2023 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Paul, S. G. et al. A real-time application-based convolutional neural network approach for tomato leaf disease classification. Array19, 100313. 10.1016/j.array.2023.100313 (2023). [Google Scholar]
  • 79.Elfatimi, E., Eryigit, R. & Elfatimi, L. Beans leaf diseases classification using MobileNet models. IEEE Access10, 9471–9482. 10.1109/ACCESS.2022.3142817 (2022). [Google Scholar]
  • 80.Sambasivam, G. & Opiyo, G. D. A predictive machine learning application in agriculture: Cassava disease detection and classification with imbalanced dataset using convolutional neural networks. Egypt. Inform. J.22(1), 27–34. 10.1016/j.eij.2020.02.007 (2021). [Google Scholar]
  • 81.Chhetri, T. R., Hohenegger, A., Fensel, A., Kasali, M. A. & Adekunle, A. A. Towards improving prediction accuracy and user-level explainability using deep learning and knowledge graphs: A study on cassava disease. Expert Syst. Appl.233, 120955. 10.1016/j.eswa.2023.120955 (2023). [Google Scholar]
  • 82.Zhong, Y., Huang, B. & Tang, C. Classification of Cassava leaf disease based on a non-balanced dataset using transformer-embedded ResNet. Agriculture12(9), 1360. 10.3390/agriculture12091360 (2022). [Google Scholar]
  • 83.Sasmal, B. et al. A novel groundnut leaf dataset for detection and classification of groundnut leaf diseases. Data Brief55, 110763. 10.1016/j.dib.2024.110763 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Ahmad, A., Gamal, A. E. & Saraswat, D. Toward generalization of deep learning-based plant disease identification under controlled and field conditions. IEEE Access11, 9042–9057. 10.1109/ACCESS.2023.3240100 (2023). [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The datasets are publicly available in a standard repository. Cassava Disease Classification Dataset is available at https://www.kaggle.com/c/cassava-disease/data and Groundnut Leaf Dataset is available at https://data.mendeley.com/datasets/22p2vcbxfk/3.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES