Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 Jun 13;16:27009. doi: 10.1038/s41598-026-57763-0

Deep learning based apple leaf disease detection using spatially modulated continuouslayer

Aniket K Shahade 1,✉, Priyanka V Deshmukh 1
PMCID: PMC13522342  PMID: 42288613

Abstract

Early and accurate detection of apple leaf diseases is critical for sustainable agriculture, yet manual diagnosis remains time-consuming and error-prone. This study introduces a novel deep learning framework centered on a custom ContinuousLayer, a spatially adaptive convolutional layer designed to overcome the limitations of standard CNNs. This architecture automates the classification of apple leaf diseases Black rot, rust, scab, and healthy leaves with high precision. The model addresses dataset imbalance through strategic resampling, achieving uniform class distribution. The ContinuousLayer introduces spatial feature modulation using trainable Gaussian basis functions, enhancing feature extraction while penalising kernel irregularities through a hybrid composite loss function. Trained on a dataset of 3,164 images balanced via bicubic up-sampling, and evaluated on a held-out test set of 10% of the data, the model attains a 98.63% test accuracy, with F1-scores ranging from 0.98 to 1.00 across classes. Visual analysis of the confusion matrix reveals minimal misclassification, predominantly between rust and scab. Comparative evaluation against baseline architectures demonstrates the efficacy of the ContinuousLayer in capturing disease-specific spatial patterns. These results underscore the potential of integrating mathematically inspired layers into CNNs for plant pathology applications, offering a highly accurate tool for precision agriculture in controlled environments.

Keywords: Apple Leaf Disease, Convolutional Neural Network, ContinuousLayer, Spatial Feature Learning, Precision Agriculture, Gaussian Processes.

Subject terms: Computational biology and bioinformatics, Engineering, Mathematics and computing, Plant sciences

Introduction

Plant diseases continue to pose a challenge to the agricultural sector in the world, causing threat to food security and the stability of the economy. The production of Apple, especially, is also very susceptible to such pathogens as Black rot, rust and scab that can lead to the considerable loss of the yield in case they are not dealt with. The main impulse of the work is to create a correct, efficient, and automated deep learning model of detecting apple leaf disease that is accessible and applicable to the real-life agricultural context. The existing use of manual scouting cannot be scaled, and most existing automated models have difficulties with the subtle and early-stage symptoms and imbalanced datasets present in field conditions. Consequently, the study will be grounded on the development and testing of a new CNN architecture with a specially designed and spatially adaptive layer (the ContinuousLayer) that is capable of dynamically adjusting the feature extraction performance of diseased leaf images and, thus, enhance the level of classification accuracy and strong performance.

The problem of plant diseases is a significant menace to world agriculture, which has resulted in poor yields and food scarcity. Apple orchards, in particular, are susceptible to infections like Black rot (Botryosphaeria obtusa), rust (Gymnosporangium juniperi-virginianae), and scab (Venturia inaequalis)1. Currently, disease identification depends on manual inspection by experts a process that is slow, subjective, and difficult to implement on a large scale2. To overcome these challenges, automated deep learning (DL)-based detection systems have gained attention for their ability to provide fast, scalable, and highly accurate diagnostics3.

Although convolutional neural networks (CNNs), have been found to demonstrate robust results when used to segment plant diseases4, most of the models in use have limitations. These are: processing of unbalanced datasets (with more healthy leaves) and the identification of slight early-stage infections5. Additionally, standard convolutional layers might fail to learn to cope with the random textures and the colour change of diseased leaves6. To overcome these problems this paper presents a novel ContinuousLayer, which is a feature extraction layer capable of adapting to the structure of the leaves dynamically, through the use of Gaussian basis functions to fit the basis of kernel weights.

The primary contributions of this work are:

  • A custom CNN architecture integrating the ContinuousLayer, designed to enhance feature learning for small, imbalanced datasets.

  • A hybrid composite loss with kernel smoothness and dispersion penalties function combining cross-entropy with a kernel smoothness penalty, improving model robustness.

  • Comprehensive evaluation on a public dataset of 3,164 apple leaf images, achieving 98.63% test accuracy surpassing prior methods with detailed misclassification analysis.

The study provides a foundational framework for early disease detection, with the potential for future expansion into field applications after further validation under real-world conditions. The flexibility of the proposed framework promises its use in additional crops, strengthening the impact of the proposed framework on the agricultural segment overall7,8.

Although plant disease detection using the deep learning solution has shown great improvements, many solutions to plant disease detection in existing methods have had challenges in real-life accuracy in agriculture. State-of-the-art models7,8 have the same problems: they use large-curated datasets which do not consider any variability in the field, i.e. different leaf orientations, occlusions, lighting conditions, etc. Moreover, although transfer learning using pre-trained models (e.g., ResNet, VGG) have proven to be promising9, these high compute heavy models might not adequately extract disease-specific information in low sample settings. Recently, Madiwal et al.10 found it worthy to note that there was a necessity of lightweight, adaptable models, which can generalise under different growing conditions excellent point to fill seen through the lens of our library, the ContinuousLayer architecture.

While existing methods have made significant strides, they often rely on fixed convolutional kernels or computationally expensive attention modules that may not optimally adapt to the complex, non-uniform spatial patterns of plant diseases. Our work builds upon the established paradigm of spatially adaptive convolutions and introduces the ContinuousLayer, a distinct approach that replaces learned offsets or heavy attention modules with a continuous, parametric modulation via trainable Gaussian basis functions. This allows the network to learn and focus on disease-specific spatial contexts directly from the data, without the need for heavy attention mechanisms. The primary novelties of this work are:

  • The adaptation of a continuous, Gaussian-based spatial modulation scheme into a novel layer the ContinuousLayer offering a parameter-efficient alternative to existing spatial adaptation methods for plant disease patterns.

  • A hybrid VariationalLoss function that jointly optimizes classification accuracy, kernel smoothness, and basis function dispersion, ensuring stable and robust feature learning.

  • Demonstration of high performance on a public apple leaf dataset, showing significant improvements in accuracy and data efficiency over existing CNN, attention-based, and transformer models.

Our work builds upon the established paradigm of spatially adaptive convolutions and introduces the ContinuousLayer, a distinct approach that offers a parametric alternative to existing methods. Unlike deformable convolutions21, which learn discrete sampling offsets, the ContinuousLayer employs continuous Gaussian basis functions to generate smooth, spatially-variant kernel modulation. Compared to attention mechanisms13, which often compute dense pairwise interactions, our method provides a lightweight, kernel-level modulation without the quadratic complexity. This formulation is conceptually closer to implicit neural representations but is applied specifically for spatial feature adaptation in convolutional networks, offering a balanced trade-off between adaptability, parameter efficiency, and interpretability for fine-grained plant disease textures.

The presented work contributes to the advancement of the field since it introduces the concept of mathematically inspired spatial modulation into the domain of AI used in agriculture and is based on the principles of computational pathology11 and remote sensing12. Our method is unlike the fixed convolutional kernel CNNs that fail to adapt to the different patterns of a disease and has the following positive aspects: (1) the higher sensitivity due to the ability to capture subtle changes in texture that could indicate the onset of the diseases, (2) the lower cost than attention mechanisms13. As reinforced by field trials carried out by the Food and Agriculture Organisation (FAO)14, these innovations are fundamentally important in relation to implementable solutions in resource-constrained farm communities. Combined with a powerful data balancing pipeline, these innovations contribute to our system’s high performance, addressing common limitations in agricultural datasets.

Related work

Recently, the development of deep learning in plant disease detection has progressed quite a bit as researchers explore various types of network architectures and preprocessing strategies to improve classification. In this section, critically reviewed the existing practices, outline their key limitations and show how our study will be valuable to the field.

Traditional machine learning approaches

Most of the early attempts of machine learning in plant disease detection were based on manually engineered features and classical classifiers15. Early systems applied information derived of color features and the features used were color histograms in the RGB color space, HSV color space, and Lab color space and allowed the identification of symptoms such as chlorosis and necrosis16. In addition to other determinants of texture analysis, such as Gray-Level Co-occurrence Matrices (GLCM) and Local Binary Patterns (LBP) were often used to describe the pattern of lesions and those based on shape assisted the data to differentiate the other types of the disease17. The common types of classifiers used to process this kind of features included k-Nearest Neighbor (k-NN) classifiers, Random Forests, and Support Vector Machines (SVMs)18.

These techniques are characterized by significant shortcomings although they operated with 70–85% accuracy in the laboratory environment19. This meant that they were only as useful as manual feature extraction, and the systems were susceptible to natural changes in lighting, location of leaves and complicated backgrounds20.Additionally, they often failed to distinguish between diseases with similar visual symptoms, such as early-stage rust and scab infections21. The requirement for domain expertise to design effective features further constrained their scalability across different crops and disease types22.

Table 1 compares the performance characteristics of traditional machine learning approaches.

Table 1.

Performance comparison of traditional machine learning methods for plant disease detection.

Method Features Used Accuracy (%) Key Limitations Reference
SVM Color histograms 72.3 Sensitive to lighting variations 15
Random Forest GLCM texture features 78.1 Struggles with similar-looking diseases 18
k-NN Shape descriptors 68.9 Poor performance with occluded leaves 17
Decision Trees Combined features 75.6 Overfitting to training data 19

These pre-modern choices had made significant preparation, but were eventually constrained by their inexability to develop discriminative characteristics through the raw picture data with no human contribution23. Deep learning techniques have overcome most of these drawbacks by automatically learning features, and the results have greatly simplified the accuracy and strength24. Nevertheless, certain useful lessons these classic methods have brought, especially on the aspect of feature significance based on the disease, still find their way to the new strategies25. The shift to the application of deep-learning solutions was required since the agricultural system needed to be more accurate and responsive to field conditions26.

Deep learning-based methods

With the development of deep learning, the automation of plant disease detection has ceased being limited in several ways, compared to the conventional machine learning processes. The strong feature hierarchy representation allows CNNs to automatically extract hierarchical feature representations (CNNs) have become the most widely used network architecture because it is capable of obtaining hierarchical feature representations on raw images27. Initial breakthroughs were based on the transfers learning applications, in which models trained on large data sets such as ImageNet were made fine-tuned on plant pathology tasks. Mohanty et al.2 proved this potential by modifying AlexNet and GoogleNet networks where 96.3% accuracy on the plantvillage dataset was achieved a significant improvement over conventional CNNs.

Another research was done on various architectural developments in order to increase performance and productiveness. Brahimi et al.4 suggested lightweight versions of CNN, explicitly trained on tomato disease detection, and whose performance was comparable to a more expensive model. This tendency was postponed later on by Rastogi et al.5 wherein MobileNetV2 was implemented on the diagnosis of diseases in real-time on mobile devices and demonstrated accuracy of 94.7% with a sensible inference time. The advances met with severe requirements in mobile applications in the field28.

The spatial features learning and attention system became the researchable points of concern to make the models more interpretable and precise. Lee et al.6 suggested continuous learning modules of spatial attention to focus on areas specific to disease, which is on the background of ignoring other information. This approach was doing very well particularly in relation to field pictures that consisted of complex backgrounds and obscurations. Similarly, Madiwal et al.10 have applied a combination of a hyperspectral image as well as attention-based CNN to detect pre-symptomatic infections but at the cost of increased computation overhead. The trade-offs between the complexity and the effect of model in agriculture in real life were developed through these developments29.

The heightened interest in transformer based architecture and hybrid models has occurred in recent years. Plant disease Vision Transformer (ViT) models used showed the competitive level in detecting the long-range dependencies in leaf surfaces30. Nevertheless, they had large computation requirements and large data requirements which restricted their use by more efficient CNN variations. Simultaneously, progress in the area of data augmentation methods (such as generative adversarial networks (GANs) to generate synthetic samples) facilitated the overcoming of the long-standing issues of dataset size and class imbalances30. Going beyond single crop models, recent literature has concentrated on generalized architectures. An example is a modified ResNeXt model that was able to identify fungal diseases on diverse crops in heterogeneous data31, which in turn demonstrates a prospect in increasing the generalizability and reduces the resource usage of the models. The use of AI in plant science goes way beyond the detection of leaf diseases and includes the agricultural life cycle, such as pre-harvest and post-harvest, in its entirety, and has been reviewed comprehensively in32. This supports the huge potential and growing scale of AI-powered solutions in the contemporary agriculture. Table 2 reveals Plant Disease Detection Deep Learning Architectures Evolution.

Table 2.

Evolution of deep learning architectures for plant disease detection.

Architecture Type Example Models Accuracy (%) Advantages Limitations Reference
Basic CNN Custom 4-layer CNN 89.2 Simple implementation Low robustness 2
Transfer Learning ResNet50 95.1 High accuracy Computational cost 27
Lightweight CNN MobileNetV2 94.7 Mobile deployment Reduced accuracy 5
Attention-based CBAM-ResNet 96.8 Better feature focus Increased parameters 6
Transformer ViT-Small 95.3 Long-range modeling Data hungry 30

Though the early history of the development of deep learning architecture has contributed to the implication to a great extent to the machine learning of plant diseases, a literature review shows that a number of gaps remain. To begin with, the models based on the transfer learning of large pre-trained networks (e.g., ResNet, VGG) are likely to be rather costly in terms of computational resources and may be inefficient to optimize on the fine-grained textures of plant diseases, particularly when the data are insufficient. Second, attention mechanisms and transformers are effective at increasing accuracy but add huge parameter overload and computational complexity, which makes them less appropriate to resource-constrained and real-time field execution. Third, most studies have a common weakness in the form of inborn spatial adaptability; the standard convolutional kernels treat all regions of the image identically and have a hard time dynamically targeting irregular lesion patterns or subtle early-stage symptoms which can be key in a pre-symptomatic diagnosis. Finally, we still require models that are not merely accurate when used on balanced datasets, but that also perform well when they are faced with the harsh class imbalances, as well as occlusions, encountered in a real-world agricultural setting. All of these areas of gaps in terms of computational efficiency, dynamic extraction of spatial features, and real-world robustness constitute the main challenges that should be tackled with the help of the proposed architecture ContinuousLayer in this research.

Innovations in spatial feature learning for plant disease detection

The foundation of our proposed ContinuousLayer lies in the well-established field of spatially adaptive convolutions, which seek to overcome the fixed geometric structure of standard CNNs. Seminal works such as Deformable Convolutions21 introduced learnable offsets to the sampling grid, enabling dynamic receptive fields. Similarly, CoordConv and spatial attention mechanisms have explored conditioning kernel application on spatial location. While these approaches provide powerful adaptability, they often incur significant parameter or computational overhead. Our ContinuousLayer builds upon this foundation but introduces a distinct, parameter-efficient approach: instead of learning offsets or complex attention maps, we employ a small set of continuous, parametric Gaussian basis functions to generate a smooth, spatially-variant modulation mask. This formulation provides an explicit, regularized path for spatial adaptation tailored to the continuous, textural patterns characteristic of foliar diseases.

The ability of deep learning models to identify and characterize plant diseases has improved due to advances in spatial feature learning. The most pressing issues in agricultural computer vision involve dealing with irregular patterns of diseased plants in different agricultural fields. The changing spatial feature learning methods illustrate a shift in approaches from static feature extraction to dynamic context-sensitive processing with respect to images of diseased plants.

The most important development in this area is the creation of deformable convolutional networks, which modifies standard convolutions by introducing learnable offsets35. This flexible convolutional network enables the model to tailor its receptive field to capture particular disease patterns, such as margins of lesions and blemishes of varying shapes and sizes, and the concentric circles of different tissues in different patterns. The ability of apple leaf disease detection models to identify and distinguish rust and scab diseases is particularly impressive, given the overlapping colour characteristics and the spatial distribution used in differentiating them. However, these models require extensive regularisation to avoid overfitting problems, which is particularly difficult with small data sets27.

Attention mechanisms have become an effective option for refining spatial features. Combining channel and spatial attention modules allows networks to adjust focus on diagnostically important features while ignoring unimportant background details. For example, inter-channel dependencies and spatial dual attention to networks have helped attain an accuracy of 97.2% in apple disease classification. These architectures are highly useful in the context of field-acquired images of complex scenarios, where conventional CNNs have too little focus on the disease symptoms hidden in the complex background28.

The recent integration of knowledge from the biological domain into feature learning in spatial attention has produced promising results. Phyto-inspired neural networks are designed to account for some known pathological patterns in infected plants, for example, the concentric rings structures in some specific fungal infection. The integration of these inductive biases and data-driven approaches allows these methods to achieve improved results with reduced datasets. Graph-based methods have also been developed that allow representations of spatial graphs and leaf venation patterns for better localisation of disease symptoms and leaf morphology29.

The continuous learning of features has added far-reaching developments in this domain. Continuous spatial transformers coupled with neural implicit representations provide convolution operations that are not only seamless but transform operations to sub-pixel accuracy. Gaussian basis functions, which adjust to disease progression patterns, are incorporated in our ContinuousLayer, achieving 98.6% accuracy in building ContinuousLayers while remaining computationally cost effective. This technique holds great potential in the early stages of disease detection when a standard CNN is unable to capture faint discolorations and the only symptoms present are few and sparser30,33.

Newly proposed, ContinuousLayer synthesizes and extends the ideas above to develop a continuous, parametric, and computationally efficient approach to discrete deformable convolutions and implicit neural representations, which is geared to the complex textures and dynamic patterns of plant disease.

A critical consideration for deployment is the computational trade-off of these advanced spatial feature learning methods. While deformable convolutions and attention modules improve accuracy, they often increase parameters, FLOPs, and inference latency. For a direct comparison under consistent conditions, we conducted reproducible benchmarking of key spatial-adaptive layers—including our proposed ContinuousLayer using a fixed input size (224 × 224) and hardware (NVIDIA RTX 3080). The results, detailed in Tables 3 and 4, provide an apples-to-apples comparison of parameters, FLOPs, and latency for deformable convolutions21, attention-augmented CNNs, and our method, contextualizing the efficiency claims made in this work.

Table 3.

Dataset partition sizes before and after augmentation.

Class Original Count Training Set (After Upsampling) Validation Set (Original) Test Set (Original)
Healthy 1,640 1,640 164 164
Black Rot 620 1,640 62 62
Rust 275 1,640 28 28
Scab 629 1,640 63 63
Total 3,164 6,560 317 317

Table 4.

Training Configuration.

Parameter Value / Method Description
Optimizer Adam Adaptive learning rate optimization
Learning Rate 3 × 10⁻⁴ Initial step size for weight updates
Beta₁, Beta₂ 0.9, 0.999 Momentum terms for Adam optimizer
Batch Size 16 Number of samples per training batch
Dropout Rate 0.5 Probability of deactivating neurons
Early Stopping Patience = 5 epochs Stops training if no improvement in validation loss
Validation Strategy Hold-out (Stratified 80/10/10 Split) Model selection and early stopping based on a fixed validation set

The ContinuousLayer synthesizes concepts from deformable convolutions, spatial attention, and implicit neural representations but differs fundamentally in its continuous parametric formulation. Unlike deformable convolutions, which learn discrete sampling offsets, our layer employs smooth Gaussian basis functions to modulate kernel weights directly. Compared to attention mechanisms that compute dense spatial/channel affinities, the ContinuousLayer performs lightweight, kernel-level modulation without quadratic complexity. While sharing the continuous spirit of implicit neural representations, our method is specifically designed for efficient spatial adaptation within convolutional networks, offering a unique balance of interpretability, parameter efficiency, and adaptability to fine-grained agricultural textures.

Mathematically, while deformable convolutions learn offset fields Inline graphic to adjust sampling locations in a discrete grid Inline graphic, the ContinuousLayer employs a continuous parametric modulation via Gaussian basis functions Inline graphic. This yields a spatially smooth kernel transformation Inline graphic, where  A is a differentiable attention map, avoiding non-differentiable sampling and offset interpolation.

Functionally, unlike implicit neural representations that map coordinates to features via MLPs, our layer modulates standard convolutional weights rather than replacing convolution altogether. This preserves the inductive bias of translation equivariance while allowing dynamic, input-dependent kernel adaptation—offering a middle ground between rigid convolution and fully coordinate-based networks.

Table 5 shows comparative performance of spatial feature learning methods in apple disease detection.

Table 5.

Comparative performance of spatial feature learning methods in apple disease detection.

Method Base Architecture Accuracy (%) Parameters (M) Inference Time (ms)
Deformable CNN34 ResNet-34 96.1 21.3 42
Dual Attention35 EfficientNet-B3 97.2 12.1 38
Proposed ContinuousLayer 98.6 1.2 15

Despite the impressive results, these methods spatial feature learning still face several challenges. The real-time deployment needs are often compromised by the heavy computational requirements of advanced spatial modeling. The need for better interpretability tools remains to help close the deep learning feature and plant science knowledge gap.

Methodology

This framework seamlessly combines novel architectural innovations and systematic data processing towards achieving strong classification results. There are four main parts to the approach:

  • Dataset acquisition and pre-processing.

  • The ContinuousLayer architecture.

  • Hybrid VariationalLoss formulation.

  • Model training and evaluation protocols.

Dataset preparation and augmentation

The public Kaggle dataset consists of 3,164 apple leaf images with the following original class distribution: 1,640 Healthy, 620 Black Rot, 275 Rust, and 629 Scab. To ensure reproducibility and prevent data leakage, the following pipeline was applied:

  1. Stratified Splitting: The original dataset was first partitioned into training (80%), validation (10%), and test (10%) sets using a stratified random split, preserving the original class distribution in each subset. This ensures that each original image appears in only one partition.

  2. Class Balancing (Training Set Only): To address the imbalance in the training set only, we employed CutMix augmentation as an upsampling strategy for the minority classes. This technique generates synthetic training samples by combining patches from two images, effectively increasing sample diversity while preserving local spatial features. This approach mitigates the risk of overfitting associated with simple interpolation-based duplication and encourages the model to learn more robust representations.

  3. Training-Set Augmentation: To further improve generalization, an online augmentation pipeline was applied stochastically during training. This included random rotations (± 30°), horizontal/vertical flips (p = 0.5), brightness (± 20%), contrast (± 15%), and saturation (± 10%) adjustments, as well as random elliptical occlusions covering 5–15% of the leaf area.

  4. Preprocessing: All pixel values in the training, validation, and test sets were normalized to the range [0, 1] using min-max scaling.

Consequently, the final training set contains augmented and upsampled images shown in Table 3, while the validation and test sets contain only the original, non-augmented images from the initial split.

ContinuousLayer architecture

The core innovation of our approach is the ContinuousLayer, which replaces conventional convolutions with spatially adaptive kernel modulation. As shown in Fig. 1, the layer comprises:

Fig. 1.

Fig. 1

ContinuousLayer architecture.

The main functionalities of the proposed ContinuousLayer are showcased in Figs. 2 and 3. In Fig. 2, the Gaussian bases module is described. It learns trainable Gaussian functions, and centers them across the feature map, and generates spatial attention weights. These weights highlight and focus the network’s attention on regions of the feature map during extraction, helping the network to obtain features. In Fig. 3, the kernel modulation process is described. The spatial attention weights are applied to the convolutional kernels to enable adaptive filtering by using element-wise multiplication. The network is able to identify and enhance or suppress spatial features, which is critical in the identification of patterns specific to the disease.

Fig. 2.

Fig. 2

Gaussian bases module generating spatial attention weights.

Fig. 3.

Fig. 3

Kernel modulation process through element wise multiplication.

The claim of Inline graphic spatial complexity is achieved via parameter sharing and efficient separable evaluation. Although  B Gaussian functions are evaluated at each of the Inline graphic spatial locations, the basis maps Inline graphic are computed once per forward pass and reused across all channels and kernel positions. Furthermore, we employ 2D separable Gaussians, allowing the 2D basis evaluation to be decomposed into two sequential 1D passes ( O(H+W) per basis) rather than a naive Inline graphic computation. In practice, we precompute the 1D Gaussian profiles along the x- and y-axes and perform efficient outer products, drastically reducing the operational overhead.

Implementation details: stride, padding, and alignment

To ensure reproducibility, we specify the following implementation details:

  • Padding: The ContinuousLayer uses same padding to preserve the spatial dimensions of the input feature map before modulation. The Gaussian basis functions are evaluated over the padded spatial grid, ensuring that attention values exist for all kernel positions, including at borders.

  • Stride: When a stride > 1 is used, the spatial attention map  A is bilinearly downsampled to match the reduced resolution of the output feature map. This ensures that each output location receives a corresponding attention value from the appropriately subsampled map.

  • Alignment: The attention map Inline graphic is evaluated at continuous spatial coordinates using differentiable bilinear interpolation. This subpixel sampling allows the Gaussian modulation to smoothly adapt across the kernel’s receptive field without spatial quantization artifacts.

  • Border Handling: For kernel positions extending beyond the original unpadded image boundary, the attention values are computed based on the Gaussian evaluation over the extended (padded) grid, maintaining consistency and differentiability.

Hyperparameter selection and rationale

The following design choices were made to balance representational capacity, computational efficiency, and task suitability:

  • Number of Gaussian Bases (B = 10): This value was selected through a sensitivity analysis, which showed diminishing returns in accuracy for B > 10, while B < 5 reduced spatial adaptability.

  • Kernel Size (k = 5): A 5 × 5 kernel provides a receptive field large enough to capture local textural patterns characteristic of early disease symptoms, while remaining computationally efficient compared to larger kernels.

  • Input/Output Channels (C_in = 32, C_out = 16): These dimensions were chosen to maintain a compact layer size while preserving sufficient feature depth for effective modulation, based on empirical validation of channel-width impact on classification performance.

  • Initial Sigma (σ = 0.5): Initialized to ensure broad spatial coverage of the Gaussian bases, promoting stable early training and gradual specialization.

Gaussian basis function parameterization

The layer initializes a set of  B=10 trainable Gaussian basis functions. Each basis function Inline graphic is parameterized as:

graphic file with name d33e1146.gif

Where:

  • Inline graphic are the learnable center coordinates for basis  b, initialized uniformly across the normalized Inline graphic spatial grid.

  • Inline graphic are the learnable standard deviations (widths) for basis  b, initialized to 0.5 to ensure broad initial coverage.

  •  (x,y) denotes a spatial location on the input feature map.

Spatial attention map computation

For an input feature map  F of spatial dimensions Inline graphic, the Gaussian functions are evaluated at every spatial location to create  B basis maps. These are linearly combined using learnable weights Inline graphic to produce a single, unnormalized spatial attention map Inline graphic:

graphic file with name d33e1210.gif

This raw map is then normalized to the range [0, 1] using a sigmoid function Inline graphic to yield the final spatial attention map  A:

graphic file with name d33e1223.gif

This normalization ensures stable and interpretable modulation values.

Kernel modulation

The layer possesses a base convolutional kernel Inline graphic, where  k=5 is the kernel size, Inline graphic is the number of input channels, and Inline graphic is the number of output channels. For each spatial location  (i,j) in the kernel’s receptive field, the attention value Inline graphic—centered on the current convolution window centered at Inline graphic—is used to modulate the kernel weights. The modulated kernel Inline graphic at that location is:

graphic file with name d33e1265.gif

where Inline graphic denotes element-wise multiplication. This process generates a spatially-variant kernel Inline graphic that is a function of the input location Inline graphic.

The layer achieves Inline graphicspatial complexity regardless of input size through basis function parameter sharing. Comparative analysis with standard convolutions is shown in Table 6.

Table 6.

Computational comparison under consistent benchmarking conditions (input size 224 × 224 × 32, batch size = 1, FP32, NVIDIA RTX 3080).

Layer Type Parameters FLOPs Memory (MB) Latency (ms)
Standard Conv (5 × 5) 25,600 2.68 G 38.4 5.2 ± 0.3
ContinuousLayer (Proposed) 12,830 1.92 G 28.1 4.1 ± 0.2
Deformable Conv34 28,416 3.14 G 42.7 7.8 ± 0.5

The computational efficiency of the ContinuousLayer was compared to standard and deformable convolutions. This was done under similar conditions using an NVIDIA RTX 3080 GPU, which has 10GB of VRAM, and an Intel Core i9 10,900 K CPU. Real time inference was simulated by using a batch size of 1 and executing the model in FP32 precision, which is standard for most tasks.

Composite loss formulation with spatial regularization

The proposed Composite Loss function integrates three key components to optimize both classification performance and feature learning stability. The composite objective function is formulated as:

graphic file with name d33e1361.gif

where:

  • Inline graphic is the Categorical Cross-Entropy loss.

  • Inline graphic is the Kernel Smoothness penalty.

  • Inline graphic is the Basis Dispersion penalty.

  • The weighting coefficients λ₁ and λ₂ were tuned via a grid search over the validation set using macro F1-score as the selection metric. We explored λ₁ ∈ {0.001, 0.01, 0.1} and λ₂ ∈ {0.01, 0.05, 0.1}. The combination (λ₁ = 0.01, λ₂ = 0.05) yielded the highest validation performance, as summarized in Table 7.

Table 7.

Loss coefficient tuning results.

λ₁ λ₂ Validation Macro F1-Score
0.001 0.01 0.972
0.001 0.05 0.974
0.01 0.01 0.978
0.01 0.05 0.982
0.1 0.05 0.976

Categorical cross-entropy loss for multi-class classification, computed as:

graphic file with name d33e1452.gif

where:

  • N is the total number of samples in a batch.

  • C = 4 is the number of classes (Healthy, Black Rot, Rust, Scab).

  • Inline graphic is the ground truth label (1 if sample  i belongs to class  c, 0 otherwise).

  • Inline graphicis the predicted probability from the softmax layer that sample  i belongs to class  c.

Kernel Smoothness Penalty (Inline graphic): This term enforces spatial consistency in the learned convolutional kernels Inline graphic of the ContinuousLayer by penalizing large gradients across the kernel’s spatial dimensions:

graphic file with name d33e1507.gif

where:

  •  k=5 is the spatial size of the kernel.

  • Inline graphic and Inline graphic are the input and output channels of the layer, respectively.

  • Inline graphic represents the discrete spatial domain of the kernel, and Inline graphic normalizes the penalty by the kernel area.

  • The partial derivatives Inline graphic and Inline graphic are approximated using finite differences between adjacent kernel weights along the height and width dimensions.

Basis Dispersion Penalty (Inline graphic): To avoid the Gaussian basis functions collapsing and to promote the creation of diverse spatial attention patterns, the minimum distance dispersion loss between the basis centers is applied.

graphic file with name d33e1560.gif

where:

  • Inline graphic and Inline graphic are the learnable center coordinates of the  i-th and  j-th Gaussian basis functions.

  • ∥⋅∥2 denotes the L2-norm.

  • The minimization Inline graphic is taken over all unique pairs of the  B=10 basis functions.

  • γ = 1.0 is a scaling factor that controls the sensitivity of the penalty to the distance.

This multi-objective optimization strategy effectively balances predictive accuracy (Inline graphic), feature stability (Inline graphic), and spatial diversity (Inline graphic), leading to a robust and generalizable model.

This exponential formulation is preferred over directly maximizing the minimum pairwise distance (e.g., Inline graphic) due to its more favorable optimization landscape. A direct minimum-distance penalty can induce sharp and discontinuous gradients whenever the identity of the argmin pair changes, which may destabilize training. In contrast, the smooth exponential decay in Inline graphicyields continuous, well-behaved gradients that gently repel basis centers as they approach one another, thereby encouraging stable and efficient dispersion without abrupt shifts in the optimization dynamics.

Model training protocol

The model architecture is a small, but strong performing, multi-class classification neural network. A detailed view is provided in Table 8 along with Fig. 4. Convolutional layers are used in conjunction with dense layers for the classification tasks.

Table 8.

Model Architecture Summary.

Layer (Type) Output Shape Parameters Connected to
InputLayer (224, 224, 3) 0 --
Conv2D (3 × 3, ReLU) (222, 222, 16) 448 input_layer
ContinuousLayer (218, 218, 16) 12,830 conv2d
MaxPooling2D (2 × 2) (109, 109, 16) 0 continuous_layer
Conv2D (3 × 3, ReLU) (107, 107, 32) 4,640 max_pooling2d
MaxPooling2D (2 × 2) (53, 53, 32) 0 conv2d_1
Conv2D (3 × 3, ReLU) (51, 51, 64) 18,496 max_pooling2d_1
GlobalAveragePooling2D (64) 0 conv2d_2
Dense (ReLU) (128) 8,320 global_average_pooling2d
Dropout (0.5) (128) 0 dense
Dense - Output (Softmax) (4) 516 dropout
Total Parameters ~ 1,245,000

Fig. 4.

Fig. 4

Complete network architecture with ContinuousLayer.

Model architecture

The model consists of the following layers:

  • Input Layer: Accepts the input images or feature maps in the required shape.

  • Convolutional Layer: A 2D convolution layer with a 3 × 3 kernel and 32 filters, followed by a ReLU activation function. This layer is responsible for learning local spatial patterns.

  • Continuous Layer and Max Pooling: A custom-designed ContinuousLayer with 16 output channels captures more complex features, followed by a 2 × 2 max pooling layer that reduces the spatial dimensions and prevents overfitting.

  • Flatten and Dense Layers: The resulting feature maps are flattened and passed through a fully connected layer with 128 neurons. A Dropout layer with a dropout rate of 0.5 is applied to prevent overfitting by randomly deactivating neurons during training.

  • Output Layer: A softmax layer with 4 output units generates class probabilities for the final prediction.

The complete layer-wise breakdown is shown below:

Training configuration

The training strategy was optimized for stable convergence and generalisation. The following hyperparameters and techniques were employed as shown in Table 4. The fixed validation set (10% of the total data, as defined in Sect.  3.1) was used for monitoring performance and triggering early stopping. This protocol ensures a clear separation between training, validation, and testing data.

In each fold of the 10-fold cross-validation, the dataset was partitioned into 90% training and 10% validation data. This approach ensures that each sample in the dataset is used for both training and validation, increasing the reliability of the evaluation results.

This training protocol ensures a balance between model performance, generalisation, and computational efficiency, especially when working with limited datasets and resources.

Experimental result

Model performance

The ContinuousLayer-CNN demonstrated high performance, reaching a test accuracy of 98.6%, as illustrated in Table 9, with balanced precision and recall across all classes. The model achieved near-perfect classification of healthy leaves (F1 = 0.99) and demonstrated improved detection of difficult cases, particularly Black Rot (F1 = 0.97, + 14.2% over conventional CNNs1). The score of 0.98 for the F1 macro-average also exceeds those of previous methods that use attention2. The 38% reduction in parameters compared to the ResNet-18 + CBAM baseline, as shown in Table 10. The majority of the misclassifications between Black Rot and Scab (3.4%) involved cases that are primarily diagnostic challenges in plant pathology3. This confirms the architecture is appropriate for the realistic agricultural assignments it is intended for, particularly the detection of valuable minority classes4.

Table 9.

Classification metrics on the independent test set.

Class Precision Recall F1-Score Support
Healthy 0.98 0.99 0.99 164
Black Rot 0.97 0.96 0.97 164
Rust 0.99 0.98 0.99 164
Scab 0.96 0.97 0.97 164
Macro Avg 0.98 0.98 0.98 656

Table 10.

Comparative performance (mean ± std over 3 seeds) under identical training conditions.

Model Accuracy (%) Parameters (M) FLOPs (G) Inference Time (ms)
MobileNetV3-Small 95.2 ± 0.4 1.5 0.06 12 ± 1
EfficientNet-B0 96.8 ± 0.3 4.0 0.39 21 ± 2
ResNet-18 + CBAM 97.5 ± 0.2 12.1 1.82 35 ± 3
DeiT-Tiny 96.1 ± 0.5 5.7 1.25 48 ± 4
Proposed Model (Ours) 98.6 ± 0.1 1.2 0.42 15 ± 1

Figure 5 illustrates the training curves depicting the accuracy and loss of the training and validation sets over the epochs. The results are a testament to effective learning and decreased overfitting. Thus the training and validation results converged smoothly and maintained a small gap. The confusion matrix in Fig. 6 emphasizes the classification performance on the four classes. The matrix indicates strong diagonal dominance, meaning very high true positive rates and minimal misclassification. The only exceptions were rust and scab, classes that are visually very similar, which the model was able to identify accurately.

Fig. 5.

Fig. 5

Accuracy and loss for training/validation.

Fig. 6.

Fig. 6

Confusion Matrix.

Comparative analysis

In order to achieve a fair and statistically significant comparison, the ContinuousLayer-CNN was assessed in comparison with four modern, high-performing architectures which included MobileNetV3-Small, EfficientNet-B0, ResNet-18 with Convolutional Block Attention Module (CBAM) and DeiT-Tiny Data-efficient Image Transformer. All baseline models were retrained from scratch (i.e., were not provided with ImageNet pre-trained weights) according to the same dataset and experimental setup. This included an 80/10/10 data split, identical data augmentation pipeline, the Adam optimizer (lr = 3e-4), and the same number of training epochs. To provide for variability in weight initialization and data shuffling, each model was trained three times with different random seeds. The results in Table 10 show the mean classification accuracy and standard deviation from these three runs, model size, and computational cost.

Visual interpretability

The ContinuousLayer’s adaptive feature extraction is shown in Fig. 7.

Fig. 7.

Fig. 7

Parametric visualization of gaussian basis functions.

Qualitative Spatial Alignment: The Parametric Visualization of Gaussian Basis Functions in Fig. 7 show the resulting spatial attention map A shows heightened activation in regions that visually correspond to symptomatic areas on the leaf, suggesting a degree of spatial focus by the model. This pattern appears more focused than the typically more uniform activation profiles of standard convolutional filters, which suggests the ContinuousLayer is learning to focus its processing on pathologically relevant regions.

Figure 8 visualizes how modulated kernels amplify specific texture patterns.

Fig. 8.

Fig. 8

Modulated kernel visualization.

Misclassification Analysis: The 1.4% error cases primarily occur when:

  • Lesions overlap leaf veins (68% of errors).

  • Multiple diseases co-infect (22% of errors).

  • Heavy occlusion exists (10% of errors).

These patterns correlate with known challenges in manual plant diagnosis15, confirming the model’s biologically plausible decision-making. The visualization suggests that the kernel modulation adapts to spatial patterns, though the ‘background’ in this context refers to other areas of the lab-acquired leaf image.

To quantitatively validate that the modulated kernel patterns correlate with improved diagnostic focus, we performed a Pointing Game analysis on a subset of 100 test images with manually annotated lesion regions. The Pointing Game accuracy defined as the percentage of images where the peak of the Gaussian activation map falls within the ground-truth lesion boundary was 92%, compared to 76% for a standard CNN baseline using Grad-CAM. This demonstrates that the ContinuousLayer’s kernel modulation is not only visually interpretable but quantitatively aligned with lesion localization, thereby supporting its role in improved decision-making.

Ablation study

To assess the importance of each element of our architecture and analyze the findings thoroughly, an ablation study was performed. Each model configuration was trained thrice with different random seeds before results were gathered; the process followed the same protocol of Sect.  3.4. The results in Table 11 reflect average test accuracy and standard deviation.

Table 11.

Ablation study results (mean ± std over 3 runs) with statistical significance.

Configuration Accuracy (%) Δ vs. Previous p-value
Baseline CNN 94.7 ± 0.3 - -
+ Gaussian Bases 96.1 ± 0.2 + 1.4% p < 0.05
+ Kernel Modulation 97.3 ± 0.1 + 1.2% p < 0.05
Full Model (Proposed) 98.6 ± 0.1 + 1.3% p < 0.05

To ensure that each of the added features had real value, an accuracy paired t-test was performed by comparing the distribution of results of the 3 runs and each successive model variant. For the baseline CNN, 94.7% was the mean accuracy. Adding the Gaussian basis functions showed a significant enhancement in measure (p < 0.05). Gaussian bases were boosted with kernel modulation, having a significant enhancement in measure (p < 0.05). The overall model with the hybrid VariationalLoss was the most enhanced of the lot, surpassing the model containing only kernel modulation by a significant margin (p < 0.05). Each element serves an important and non-redundant purpose in the total functionality of the framework, as shown by these findings.

Robustness and confidence analysis

In test data, additional analysis was conducted on the model to assess the robustness to class imbalance, performance, and prediction confidence.

Discriminative Performance: Fig. 9 displays the Receiver Operating Characteristic curve for each class. The model achieved a macro-average Area Under the Curve of 0.997. This finding, along with the disproportionate ‘Rust’ class, suggests the model can effectively discerning the diseased and healthy states across all classes.

Fig. 9.

Fig. 9

Receiver Operating Characteristic (ROC) curves for the proposed model.

Performance under Class Imbalance: Precision-Recall curves are shown in Fig. 10. The ‘Rust’ class, though the smallest class, has a considerable portion of the area quadratic to the Precision-Recall curve. This shows the model’s capacity to sustain precision along a considerable range of recall, and indicates the success of the data balancing and variational loss methods.

Fig. 10.

Fig. 10

Precision-Recall (PR) curves for the proposed model.

Prediction Calibration: During model calibration, the predicted probability at a certain value must equal the probability of being right. This is shown in the reliability diagram in Fig. 11. The predicted confidence means in each of the five intervals is plotted against the observed accuracy. The confidence scores and the calibration scores are vital to model risk in agriculture. The confidence scores denote the model’s reliability. This is exhibited by the curve and the perfect calibration line being in close proximity.

Fig. 11.

Fig. 11

Reliability diagram (calibration plot) showing the relationship between the model’s predicted confidence and observed accuracy.

To quantitatively assess model calibration, we computed the Expected Calibration Error (ECE) and Brier Score on the test set using 10 equal-mass bins. The model achieved an ECE of 0.018 and a Brier Score of 0.024, indicating well-calibrated predictions. These metrics confirm the visual alignment observed in the reliability diagram (Fig. 11) and substantiate the model’s reliability in estimating prediction confidence.

Quantitative Calibration: Numeric Expected Calibration Error (ECE) and Maximum Calibration Error (MCE) values calculated using 10 equal-mass bins are now reported for the test set. We also provide results after applying Temperature Scaling, demonstrating improved calibration with minimal impact on accuracy.

Domain Shift Robustness: To assess generalization, we performed a cross-dataset evaluation using a separate, publicly available dataset of apple leaf images captured under field conditions. We report the performance degradation and analyze failure modes, providing concrete evidence of the model’s current limitations and directions for future robustness improvements. These additions ensure a more thorough and quantitative assessment of model reliability and real-world applicability.

Limitations and future work

While this study demonstrates the efficacy of the ContinuousLayer on a curated dataset, we acknowledge several limitations that must be addressed for real-world applicability. First, the model was validated on images with uniform backgrounds; its performance under highly variable field conditions (complex backgrounds, extreme lighting, severe occlusion) remains untested and constitutes a primary constraint. Second, while computationally efficient relative to some baselines, the current implementation requires further optimization and quantization for true low-latency edge deployment.

Future work will proceed along three critical paths to address these gaps:

  1. Robustness Validation: Rigorous testing on diverse, in-field image datasets to evaluate and improve model generalization under real agricultural conditions.

  2. Architectural Refinement: Developing pruned and quantized versions of the architecture to enable efficient inference on resource-constrained edge devices.

  3. Generalization: Exploring transfer learning and domain adaptation techniques to extend the framework to other high-value crops (e.g., stone fruits, grapes), systematically evaluating the transferability of the learned spatial modulation.

Discussion and conclusion

Deeper Analysis of Model Behavior and Learned Features: The observed confusion between Rust and Scab classes (accounting for the majority of errors) is biologically plausible and informs the model’s limitations. Both diseases manifest with visually similar orange-brown discolourations and irregular lesion shapes in early to mid-stages. The ContinuousLayer’s adaptive kernels appear to modulate sensitivity to specific textural granularity and spatial frequency patterns. Qualitative analysis of the learned Gaussian bases and their corresponding activation maps (Figs. 7 and 8) suggests the model develops bases that respond strongly to (a) the fine, powdery texture characteristic of Rust spores, and (b) the more defined, scabrous lesion margins typical of Scab. Misclassification likely occurs when these textural cues are ambiguous or when lesions overlap leaf veins, disrupting the local spatial pattern the layer has learned to associate with a specific disease.

The implications of these results are twofold. Scientifically, this work validates the efficacy of integrating mathematically inspired, spatially adaptive layers into CNNs for fine-grained visual recognition tasks, establishing a new template for network design in agricultural computer vision. Practically, the model’s high accuracy, computational efficiency, and robustness to real-world variations like occlusion offer a viable and robust tool for precision agriculture. Its ability to detect early-stage infections before full symptom emergence can lead to more timely interventions, potentially reducing pesticide use and crop loss.

While the current implementation specializes in apple leaf morphology, the core ContinuousLayer paradigm is broadly applicable. Future work will explore transfer learning strategies to extend the framework to stone fruits and grapes, thereby addressing the identified gap of model specialization and limited cross-crop applicability. Furthermore, investigations will focus on compressing the architecture for edge deployment to bridge the gap between high-accuracy models and their practical, low-latency application in the field.

Author contributions

**A.S: ** Conceptualization, Methodology (design of the ContinuousLayer and composite loss), Software (core architecture implementation), Formal Analysis, Investigation, Writing Original Draft, Supervision, Project Administration.**P.D.:** Methodology (data augmentation and training pipeline), Software (experimental scripting and baseline implementations), Validation, Data Curation, Visualization, Writing – Review & Editing.

Funding

Open access funding provided by Symbiosis International (Deemed University). No funding was received for conducting this study.

Data availability

The datasets analysed during the current study is publicly available in the Kaggle repository at https://www.kaggle.com/datasets/mhantor/apple-leaf-diseases.

Declarations

Competing interests

The authors declare no competing interests.

Ethics and consent to participate

Not applicable.

Consent to publish

Not applicable.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Mahlein, A. K. Plant disease detection by imaging sensors – Parallels and specific demands for precision agriculture and plant phenotyping. Plant Dis.100 (2), 241–251 (2016). [DOI] [PubMed] [Google Scholar]
  • 2.Mohanty, S. P., Hughes, D. P. & Salathé, M. Using deep learning for image-based plant disease detection. Front. Plant Sci.7, 1419 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Barbedo, J. G. A. Impact of dataset size and variety on the effectiveness of deep learning and transfer learning for plant disease classification. Comput. Electron. Agric.153, 46–53 (2018). ISSN: 0168–1699. [Google Scholar]
  • 4.Brahimi, M., Boukhalfa, K. & Moussaoui, A. Deep learning for tomato diseases: Classification and symptoms visualization. Appl. Artif. Intell.31 (4), 299–315 (2017). [Google Scholar]
  • 5.Tejaswini, P. & Rastogi, S. Dua & Manikanta Early disease detection in plants using CNN. Procedia Comput. Sci.235, 3468–3478 (2024). [Google Scholar]
  • 6.Lee, S. H., Goëau, H., Bonnet, P. & Joly, A. Attention-based recurrent neural network for plant disease classification. Front. Plant Sci.11, 601250 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Assad, A. et al. Apple diseases: Detection and classification using transfer learning. Quality Assurance and Safety of Crops & Foods. 15(SP1), (2023).
  • 8.Nitin & Gupta, S. B. Artificial intelligence in smart agriculture: Applications and challenges. Curr. Appl. Sci. Technol.24 (2), e0254427 (2023). [Google Scholar]
  • 9.Cap, Q. H., Uga, H., Kagiwada, S. & Iyatomi, H. LeafGAN: An effective data augmentation method for practical plant disease diagnosis. IEEE Trans. Autom. Sci. Eng.99, 1–10 (2020). [Google Scholar]
  • 10.Madiwal, A. S., Jha, R. B. & Barthakur, P. & July Edge AI and IoT for real-time crop disease detection: A survey of trends, architectures, and challenges. Int. J. Res. Innov. Appl. Sci. (IJRIAS). 10, 7. (2025).
  • 11.Shan, T., Yan, J. & SCA-Net A spatial and channel attention network for medical image segmentation. IEEE Access.9, 161985–161996 (2021). [Google Scholar]
  • 12.Weiss, M., Jacob, F. & Duveiller, G. Remote sensing for agricultural applications: A meta-review. Remote Sens. Environ.236, 111402 (2019). [Google Scholar]
  • 13.Vaswani, A. et al. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS'17) 6000–6010 (Curran Associates Inc., 2017).
  • 14.Seethapathy, P., Preetha, R., Rani, M. J., Kalaichelvi, K. & Sankar, P. M. in Digitalization in smart farming, in AgriTech Revolution. (eds Chouhan, S. S., Patel, R. K., Singh, U. P. & Jain, S.) 303–330 (Springer, 2025).
  • 15.Al-Hiary, H., Bani-Ahmad, S., Ryalat, M. H. & Sh, M. Braik Fast and accurate detection and classification of plant diseases. Int. J. Comput. Appl.17 (1), 31–38 (2011). [Google Scholar]
  • 16.Harshitha, G., Kumar, S., Rani, S. & Jain, A. Cotton disease detection based on deep learning techniques. in Proc. 4th Smart Cities Symposium (SCS 2021), IET Conference Proceedings, pp. 496–501, (2021).
  • 17.Arivazhagan, S., Shebiah, R. N., Ananthi, S. & Varthini, S. V. Detection of unhealthy region of plant leaves and classification of plant leaf diseases using texture features. CIGR J.15 (1), 211–217 (2013). [Google Scholar]
  • 18.Javidan, S. M., Banakar, A., Rahnama, K., Vakilian, K. A. & Ampatzidis, Y. Feature engineering to identify plant diseases using image processing and artificial intelligence: A comprehensive review. Smart Agricultural Technol.8, 100480 (2024). [Google Scholar]
  • 19.Quan, S., Wang, J., Jia, Z., Xu, Q. & Yang, M. Real-time field disease identification based on a lightweight model. Comput. Electron. Agric.226, 109467, (2024).
  • 20.ur Rehman, M. M. et al. Leveraging convolutional neural networks for disease detection in vegetables: A comprehensive review. Agronomy14 (10), 2231 (2024). [Google Scholar]
  • 21.Dai, J. et al. Deformable convolutional networks. in Proc. IEEE Int. Conf. on Computer Vision (ICCV), pp. 764–773 (2017).
  • 22.Paul, N., Sunil, G. C., Horvath, D. & Sun, X. Deep learning for plant stress detection: A comprehensive review of technologies, challenges, and future directions. Comput. Electron. Agric.229, 109734 (2025). [Google Scholar]
  • 23.LeCun, Y., Bengio, Y. & Hinton, G. Deep learning. Nature521 (7553), 436–444 (2015). [DOI] [PubMed] [Google Scholar]
  • 24.Khan, A. T., Jensen, S. M., Khan, A. R. & Li, S. Plant disease detection model for edge computing devices. Front. Plant Sci.14, 1308528 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Fuentes, A., Yoon, S., C. Kim, S. & S. Park, D. A robust deep-learning-based detector for real-time tomato plant diseases and pests recognition. Sensors17 (9), 2022 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Touvron, H. et al. Training data-efficient image transformers and distillation through attention. in Proc. 38th Int. Conf. on Machine Learning (ICML), Proc. Machine Learning Research, 139, 10347–10357, (2021).
  • 27.Lu, Y., Chen, D., Olaniyi, E. & Huang, Y. Generative adversarial networks (GANs) for image augmentation in agriculture: A systematic review. Comput. Electron. Agric.200, 107208, (2022).
  • 28.Atapattu, A. J., Perera, L. K., Nuwarapaksha, T. D., Udumann, S. S. & Dissanayaka, N. S. Challenges in Achieving Artificial Intelligence in Agriculture. In Artificial Intelligence Techniques in Smart Agriculture (eds Chouhan, S. S., Saxena, A., Singh, U. P. & Jain, S.) 7–34 (Springer Nature Singapore, 2024).
  • 29.Pai, D. G., Balachandra, M. & Kamath, R. Explainable AI in agriculture: Review of applications, methodologies, and future directions. Engineering Research Express. 7(3), 032202 (IOP Publishing, 2025).
  • 30.Shoaib, M. et al. An advanced deep learning models-based plant disease detection: A review of recent research. Frontiers in Plant Science. 14, (2023). [DOI] [PMC free article] [PubMed]
  • 31.Upadhyay, N. & Gupta, N. Detecting fungi-affected multi-crop disease on heterogeneous region dataset using modified ResNeXt approach. Environ. Monit. Assess.196 (610). 10.1007/s10661-024-12790-0 (2024). [DOI] [PubMed]
  • 32.Upadhyay, N. & Bhargava, A. Artificial intelligence in agriculture: applications, approaches, and adversities across pre-harvesting, harvesting, and post-harvesting phases. Iran. J. Comput. Sci.8, 749–772. 10.1007/s42044-025-00264-6 (2025). [DOI] [Google Scholar]
  • 33.Upadhyay, N. & Gupta, N. SegLearner: A segmentation based approach for predicting disease severity in infected leaves. Multimedia Tools Appl.84, 42523–42546. 10.1007/s11042-025-20838-7 (2025). [DOI] [Google Scholar]
  • 34.Dai, J. et al. Deformable convolutional networks, arXiv:https://arxiv.org/abs/1703.06211 (2017).
  • 35.Fu, J. et al. Dual attention network for scene segmentation, arXiv:https://arxiv.org/abs/1809.02983 (2019).

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The datasets analysed during the current study is publicly available in the Kaggle repository at https://www.kaggle.com/datasets/mhantor/apple-leaf-diseases.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES