Abstract
Rice leaf diseases pose a major challenge to crop health and agricultural productivity, particularly when timely and accurate diagnosis is required under natural field conditions. The development of automated disease recognition systems depends heavily on the availability of large, well-annotated image datasets. However, many existing rice leaf disease datasets are limited in terms of environmental variability, disease representation, and real-field imaging conditions. To address this gap, this paper presents BanglaRiceLeaf, an original rice leaf image dataset collected and curated by the authors from the experimental fields of the Bangladesh Rice Research Institute (BRRI), Gazipur, Bangladesh, between July 2023 and July 2024. The dataset contains 4152 images belonging to five classes: Bacterial Leaf Blight, Bacterial Leaf Streak, Sheath Blight, Leaf Blast, and Healthy Leaf. The images were acquired from two rice varieties, BR11 and BRRI dhan34, under natural field conditions across varying illumination environments in order to reflect practical disease recognition scenarios. All images were manually annotated by trained annotators under expert supervision. The dataset is systematically organized and publicly released to support reproducible research in rice disease classification. In addition to dataset presentation, benchmark experiments using Xception, NASNetMobile, and InceptionV3 are provided to demonstrate its applicability for deep learning-based disease recognition. BanglaRiceLeaf is expected to serve as a useful resource for plant disease analysis, comparative model evaluation, and future research in precision agriculture and agricultural computer vision.
Keywords: Rice leaf disease, Disease classification, Rice health monitoring, Precision agriculture, Computer vision, Deep learning, Bangladesh agriculture, Crop management
Specifications Table
| Subject | Computer Sciences |
| Specific subject area | Image Processing, Disease Detection, Crop Management, Plant Pathology, Precision Agriculture |
| Type of data | The data consist of image files. Images are stored in JPG format. The image resolution is 640 × 480 pixels. |
| Data collection | The images were collected over a one-year period from July 2023 to July 2024 under varying natural lighting and weather conditions. Data collection was conducted in the experimental fields of the Bangladesh Rice Research Institute (BRRI), Gazipur. Images were captured using the rear cameras of seven smartphones, including five iPhone 11 Pro Max and two iPhone 11 devices. The dataset is organized into five classes representing four rice leaf diseases, Bacterial Leaf Blight (BLB), Bacterial Leaf Streak (BLS), Sheath Blight, and Leaf Blast, along with Healthy Leaf samples. In total, the dataset contains 4152 images. |
| Data source location | Data for this study were collected from the rice fields of the Bangladesh Rice Research Institute, Gazipur, Dhaka, Bangladesh (Latitude: 23°59′29.6″ N, Longitude: 90°24′29.2″ E). |
| Data accessibility | Repository Name: BanglaRiceLeaf: A Benchmark Dataset for Automated Rice Leaf Disease Detection and Health Classification in Bangladesh. Data Identification Number: https://doi.org/10.7910/DVN/XAOBYW Direct URL to Data: https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/XAOBYW Access Instructions: This dataset is publicly available on the Harvard Dataverse repository and can be accessed for academic, research, and instructional purposes. |
| Related research article | None. |
1. Value of the Data
-
•
BanglaRiceLeaf provides a curated image dataset of major rice leaf diseases together with healthy samples, supporting research in plant pathology and rice crop health assessment.
-
•
The dataset offers field-collected images acquired under natural environmental conditions, making it suitable for developing and evaluating image-based disease recognition methods under real field-captured settings with clearly visible symptoms.
-
•
It can serve as a benchmark resource for comparing image-based disease classification approaches and for studying robustness under realistic variations in leaf appearance and illumination.
-
•
The dataset may support future work on early disease detection, crop monitoring, and digital decision-support tools for rice disease management.
-
•
It is also useful for academic and agricultural training in visual disease recognition and image-based analysis.
-
•
By linking agricultural data collection with computational analysis, the dataset encourages interdisciplinary research across agriculture, plant science, and computer vision.
2. Background
Being the most widely consumed staple food worldwide, rice ranks third in global agricultural production by volume, with annual consumption exceeding 510 million metric tons of milled rice [1]. According to the United Nations Food and Agriculture Organization, over 50% of the global population depends on rice as a dietary staple [2]. However, rice cultivation faces significant threats from various biotic stresses, with leaf diseases being the most destructive. Annually, rice suffers yield losses of 10–30% due to plant diseases, posing a major threat to global food security. Projections indicate that by 2050, food production must increase by nearly 70% to sustain the expanding global population [3].
Therefore, accurate detection of rice leaf diseases is essential to minimize losses and ensure stable production. To enable such detection, researchers in [4] introduced a dataset containing 19,000 images across seven classes, but only 2753 are original while the remainder are augmented, limiting its ability to reflect real-world variability. Another dataset proposed in [5] includes multiple disease classes but excludes BLS and relies heavily on augmentation, with nearly 75% of images being synthetic. In contrast, the Dhan-Shomadhan [6] dataset contains just 1106 images across five classes and omits key diseases such as BLB and BLS, further constraining its applicability. In [7], the researchers introduced a dataset comprising three classes, BLB, Brown Spot, and Leaf Smut, with 40 images per class. The limited number of samples increases the risk of overfitting and reduces the ability of models to generalize to unseen data. In [8], the dataset includes two classes, Brown Spot and Healthy, with images captured under varying lighting conditions, some showing dew on the leaves. All images were carefully annotated and verified by domain experts; however, the dataset is restricted in scope due to having only two categories. According to [9], a collection of 10,766 samples was compiled, of which only 2508 are original images while the remainder were produced through augmentation. While this increases the total number of samples, it does not substantially improve the diversity of the dataset.
3. Data Description
The BanglaRiceLeaf dataset is an original rice leaf image dataset, collected and curated by the authors, comprising four economically important rice leaf diseases together with healthy samples to support research on automated rice disease identification, crop health monitoring, agricultural computer vision, and other related applications.
BanglaRiceLeaf contains 4152 images organized into five classes, each representing a distinct rice leaf condition: Bacterial Leaf Blight (BLB), Bacterial Leaf Streak (BLS), Sheath Blight, Leaf Blast, and Healthy Leaf. The dataset includes images sourced from two widely cultivated Bangladeshi rice varieties: BR11 and BRRI dhan34 to capture potential varietal differences. Despite potential varietal differences, the characteristic symptoms of rice diseases are primarily determined by the pathogen involved [10]. Leaf Blast images were collected from BRRI dhan34, whereas BLB, BLS, and Sheath Blight images were collected from BR11 plants. Healthy leaf images were captured from both BR11 and BRRI dhan34 plants. Images were captured under field conditions and subsequently organized into five disease categories for dataset construction. Table 1 presents the class categories, their defining visual characteristics, and the number of images in each category. This class composition provides a suitable basis for training and evaluating image-based rice disease classification models.
Table 1.
BanglaRiceLeaf dataset overview.
| Class | Description | Number of Images | Sample Image |
|---|---|---|---|
| BLB | Water-soaked lesions with yellow halos appear on leaves, eventually merging into irregular streaks, caused by Xanthomonas oryzae pv. oryzae. This disease accelerates leaf drying, disrupts photosynthesis, and can result in significant yield loss under severe infection [10]. | 1093 | ![]() |
| BLS | Long, narrow streaks develop between leaf veins, initially water-soaked and later turning gray or whitish, as a result of Xanthomonas oryzae infection. The disease reduces chlorophyll content, weakens plant vigor, and diminishes grain production efficiency [10]. | 1055 | ![]() |
| Sheath Blight | Irregular lesions form on leaf sheaths, blades, and occasionally culms due to Rhizoctonia solani. As the infection progresses, basal rotting and lodging occur, lowering photosynthetic efficiency and reducing both grain quality and overall yield [10]. | 418 | ![]() |
| Leaf Blast | Distinct lesions with gray centers and dark margins appear on leaves, stems, and panicles when infected by Magnaporthe oryzae. The disease spreads rapidly under favorable conditions, causing extensive tissue damage, spikelet sterility, and severe reductions in harvestable grain [10]. | 1086 | ![]() |
| Healthy Leaf | Leaves display a uniform, vibrant green color with smooth surfaces, free from visible lesions, scratches, or damage. The absence of disease symptoms or stress indicators reflects normal growth and sufficient nutrient availability [4]. | 500 | ![]() |
All images in the dataset were resized and maintained at a fixed resolution of 640 × 480 pixels, ensuring consistency across all classes and simplifying preprocessing and model input preparation. The dataset includes images acquired under varied illumination conditions, which Improves its value for developing models under field-collected conditions.
The BanglaRiceLeaf dataset is designed as a reusable benchmark resource for comparative evaluation in plant disease detection and may be used for studies involving domain adaptation, transfer learning, and explainable artificial intelligence. The structured organization of raw images and class annotations supports reproducible experimentation and broader reuse in agricultural image analysis research.
Fig. 1, illustrates the folder structure of the BanglaRiceLeaf repository. The dataset is organized into five class-specific subfolders under the main dataset directory. Images in each class are named systematically using the class abbreviation followed by a four-digit numerical identifier with leading zeros, for example, BLB_0001, BLS_0001, SB_0001, LB_0001, and Healthy_0001. This structured naming scheme supports consistent organization, efficient sorting, and convenient reuse of the dataset. The complete dataset is publicly accessible via the Harvard Dataverse repository [11].
Fig. 1.
The structure of the dataset folder in the data repository.
4. Experimental Design, Materials and Methods
4.1. Dataset acquisition
The image collection spanned a one-year period, from July 2023 to July 2024, capturing temporal variations across different stages of the growing season from a real rice field as shown in Fig. 2. The images in the dataset were captured under natural field conditions to reflect realistic variability in rice cultivation environments.
Fig. 2.
Geographical location of the study area and the rice field used for dataset acquisition.
Photographs were taken at a typical distance of approximately 30–60 cm from the leaf surface, with the camera held at an angle roughly perpendicular to the leaf plane to maximize leaf visibility and minimize occlusion from other plant parts. However, minor variations in viewing angle were unavoidable due to field constraints.
Images were acquired throughout different times of the day to capture natural variations in illumination and shadowing, and no artificial lighting or controlled backgrounds were used. This approach ensures that the dataset represents authentic field conditions, supporting the development of models under controlled field variability and enabling other researchers to reuse the images for benchmarking and automated plant disease recognition tasks.
The images were captured using the rear cameras of seven smartphones and are provided in high quality JPG format. All devices were operated using default automatic camera settings, including autofocus, automatic exposure, and automatic white balance, without any manual adjustments. This approach was intentionally adopted to reflect realistic field acquisition conditions, ensuring that the dataset captures practical, real-world variability rather than controlled laboratory settings. As summarized in Table 2, five devices were iPhone 11 Pro Max models, while the remaining two were iPhone 11 models.
Table 2.
Technical specifications of imaging devices.
| Specifications | Device 1 | Device 2 |
|---|---|---|
| Manufacturer | Apple | Apple |
| Model | iPhone 11 Pro Max | iPhone 11 |
| Resolution | 4032 × 3024 | 4032 × 3024 |
| Color space | sRGB | sRGB |
| Focal length | 26 mm | 24 mm |
| Aperture (f-number) | f/1.8 | f/1.8 |
| Exposure time | 1/10 s to 1/8000s | 1/10 s to 1/8000s |
During field data collection, only rice leaves exhibiting clearly visible and distinguishable symptoms of the target diseases were included in the dataset. Leaves that were severely occluded, partially visible, blurred, or physically damaged due to non-disease factors such as mechanical injury or insect feeding were excluded to maintain data quality and labelling reliability. To capture natural variability under real field conditions, multiple images of the same affected leaf were intentionally taken from different angles, orientations, distances, and lighting conditions. In addition, each image was assigned a single disease label corresponding to the predominant disease symptom visible on the leaf. Images exhibiting severe overlap of multiple diseases or ambiguous symptoms were excluded during dataset curation to ensure reliable annotations. This approach enhances model robustness by reflecting real-world deployment scenarios. However, exact or near-identical duplicate images were carefully reviewed and removed during post-collection screening to minimize redundancy in the final dataset. Table 3 summarizes the image collection conditions and quality control criteria used in the dataset.
Table 3.
Image collection conditions and quality control criteria.
| Parameter | Details |
|---|---|
| Data Collection Period | July 2023 – July 2024 |
| Field Environment | Real rice field (natural conditions) |
| Target | Rice leaves with clearly visible disease symptoms |
| Excluded Leaves | Severely occluded, partially visible, blurred, or damaged by non-disease factors (mechanical injury, insects) |
| Camera Type | Rear cameras of seven smartphones |
| Image Resolution | 640 × 480 pixels |
| Image Format | High quality JPG |
| Camera Settings | Default automatic settings (autofocus, auto exposure, auto white balance) |
| Camera-to-Leaf Distance | Approximately 30–60 cm |
| Camera Angle | Roughly perpendicular to leaf plane |
| Number of Images per Leaf | Multiple images from different angles, orientations, distances, and lighting conditions |
| Lighting Conditions | Natural illumination; no artificial lighting or controlled background |
| Time of Day | Various, to capture natural illumination and shadow variations |
| Post-Collection Processing | Near-identical duplicates removed to reduce redundancy |
4.2. Data annotation
After collection, a total of 4216 rice leaf images were organized into five disease-condition categories: Bacterial Leaf Blight (BLB), Bacterial Leaf Streak (BLS), Sheath Blight, Leaf Blast, and Healthy Leaf. The annotation process was performed independently by three trained annotators under the supervision of a domain expert, the Head of the Department of Plant Pathology at BRRI.
To assess annotation reliability, inter-annotator agreement was measured on the complete set of annotated images using Fleiss' kappa, yielding a score of κ = 0.9863, indicating almost perfect agreement among the annotators.
Beyond kappa analysis, the distribution of annotation agreement across disease classes was further examined. As illustrated in Fig. 3, the vast majority of images achieved unanimous agreement among the three annotators. Of the 4216 annotated images, 4152 (98.48%) received identical labels from all annotators. Among the remaining samples, 58 images (1.38%) exhibited partial agreement, while only 6 images (0.14%) showed complete disagreement among annotators.
Fig. 3.
Class-wise distribution of annotation agreement patterns among the annotators.
Class-wise analysis revealed consistently high annotation consistency across all disease categories. Disagreements were observed only in the Bacterial Leaf Blight, Bacterial Leaf Streak, and Leaf Blast classes, whereas all Healthy Leaf and Sheath Blight samples achieved unanimous agreement. These findings demonstrate the high reliability of the annotation process and the clear visual distinction of the disease categories.
All 64 non-unanimous cases were subsequently reviewed by the domain expert as part of the quality-control procedure. To ensure maximum label reliability and eliminate potentially ambiguous samples, these images were excluded from the final released dataset. Consequently, the released dataset consists of 4152 images with unanimous annotator agreement, providing a highly reliable benchmark for rice disease classification.
4.3. Benchmark model development and evaluation
To demonstrate the usability of the BanglaRiceLeaf dataset for image-based rice leaf disease classification, three deep learning models, Xception, NASNetMobile, and InceptionV3, were trained and evaluated as benchmark models. These architectures were imported from the TensorFlow Keras Applications library, leveraging their pre-trained weights for feature extraction. To adapt them for the target classification task, a Global Average Pooling 2D layer was added to reduce spatial dimensions and help prevent overfitting. A fully connected Dense layer with five output neurons and a Softmax activation function was then appended to perform multi-class classification. The detailed model configuration and training parameters used for all architectures are presented in Table 4.
Table 4.
Model training and configuration details.
| Parameter | Configuration |
|---|---|
| Input Image Size | 224 × 224 × 3 |
| Batch Size | 16 |
| Epochs | 10 |
| Optimizer | Adam |
| Learning Rate | 0.001 |
| Loss Function | Categorical Crossentropy |
| Number of Classes | 5 |
To ensure robustness and generalizability, the dataset was divided into training, validation, and testing subsets using an 80:10:10 ratio. Prior to splitting, the data were randomly shuffled, and class-wise distribution was preserved across all subsets to ensure that each split remained representative of the overall distribution and to minimize selection bias. The benchmark workflow is illustrated in Fig. 4. During model development, the training subset was used for learning, the validation subset was used to monitor performance during training, and the independent test subset was used for final evaluation on previously unseen data. The same data partitioning and evaluation protocol were applied consistently across all three architectures to ensure fair comparison.
Fig. 4.
Workflow of model training.
All experiments were conducted on a high-performance computing environment. The system configuration is summarized in Table 5.
Table 5.
Hardware and system configuration.
| Component | Specification |
|---|---|
| Processor | 13th Generation Intel Core i7 |
| RAM | 64 GB |
| Storage | 256 GB SSD and 2 TB HDD |
| GPU | NVIDIA RTX 4080 with 16 GB VRAM |
Model performance was assessed using four standard multi-class classification metrics: accuracy, precision, recall, and F1-score. In addition, training and validation accuracy-loss curves were recorded for each model and are presented in Fig. 5, Fig. 6, Fig. 7.
Fig. 5.
Accuracy and loss curve of Xception model.
Fig. 6.
Accuracy and loss curve of NASNetMobile model.
Fig. 7.
Accuracy and loss curve of InceptionV3 model.
For the benchmark experiment, Xception achieved 100.00% validation accuracy, precision, recall, and F1-score, and 99.76% for all four metrics on the test set. NASNetMobile achieved 99.26% validation accuracy, 99.29% precision, 99.26% recall, and 99.27% F1-score; on the test set, it achieved 97.60% accuracy, 97.69% precision, 97.60% recall, and 97.58% F1-score. InceptionV3 achieved 99.26% validation accuracy, 99.28% precision, 99.26% recall, and 99.26% F1-score, while its test accuracy, precision, recall, and F1-score were each 99.52%.
The findings suggest that while all three models achieved high performance, differences in architecture can still affect predictive outcomes. Careful selection and optimization of model structures may further refine classification accuracy and reliability in detecting rice leaf diseases.
These benchmark results are provided to illustrate the applicability of the dataset for deep learning-based rice leaf disease classification and to offer a reference point for future comparative studies using the BanglaRiceLeaf dataset.
Limitations
The BanglaRiceLeaf dataset was collected using a limited set of smartphone cameras and exclusively from experimental plots in Gazipur, Bangladesh. While this controlled setup ensured consistency in image acquisition, it may restrict the generalizability of models trained on the dataset when applied to images captured with different devices or under varying geographic and environmental conditions. Variations in camera sensors, soil backgrounds, and regional cultivation practices, which are not represented in the dataset, could significantly influence image characteristics, potentially affecting model performance in real-world deployments. Moreover, practical field conditions may involve early-stage disease symptoms, mixed infections, nutrient deficiencies, insect interference, and other confounding factors that are not fully represented in the current dataset. Therefore, the dataset primarily supports evaluation under conditions with clearly visible disease symptoms, and performance under more complex and unconstrained field scenarios remains an open challenge.
Ethics Statement
This research does not involve human subjects, animal experiments, or any data collected from social media platforms.
Credit Author Statement
Abu Bakar Siddique Mahi: Methodology, Validation, Investigation, Data Curation, Writing-Original Draft, Visualization. Apurba Datta: Methodology, Validation, Investigation, Data Curation, Writing Original Draft, Visualization. Safi Ullah Chowdhury: Methodology, Validation, Investigation, Data Curation, Writing-Original Draft, Visualization. Tasnim Jahin Mowla: Data Curation, Writing-Original Draft, Visualization. Tanjina Helaly: Validation, Supervision, Writing-Review and Editing, Project Administration. Tahmid Hossain Ansari: Validation and Review, Nasima Begum: Validation and Review.
Declaration of Competing Interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Data availability
References
- 1.M. Solh, “Foreword,” Food and Agriculture Organization of the United Nations (FAO). [PubMed]
- 2.Rice consumption by country 2026. 2026. https://worldpopulationreview.com/country-rankings/rice-consumption-by-country
- 3.Hossain M.M., Sultana F., Mostafa M., Ferdus H., Rahman M., Rana J.A., Islam S.S., et al. Plant disease dynamics in a changing climate: impacts, molecular mechanisms, and climate-informed strategies for sustainable management. Disc. Agricult. 2024;2(1):132. [Google Scholar]
- 4.Hasan A., Layes T.A., Afridi A.S., Rifat S.H., Nur F.N., Moon N.N. A comprehensive dataset of Rice leaf images for disease detection using machine learning. Data Brief. 2025;62 doi: 10.1016/j.dib.2025.111977. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.M. Hasan, S. Khatun, M.A. Raihan and A.H. Uddin, “Rice leaf bacterial and fungal disease dataset,” 2023. [Online]. Available: 10.17632/hx6f852hw4.2. [DOI]
- 6.M.F. Hossain, S. Abujar, S.R.H. Noori and S.A. Hossain, “Dhan-Shomadhan: a dataset of rice leaf disease classification for Bangladeshi local rice,” 2021. [Online]. Available: 10.17632/znsxdctwtt.1. [DOI]
- 7.J. Shah, H. Prajapati and V. Dabhi, “Rice leaf diseases,” 2017. [Online]. Available: 10.24432/C5R013. [DOI]
- 8.C. Pal, I. Mukherjee, S. Chatterji, S. Pratihar, P. Mitra and P.P. Chakrabarti, “Indian rice disease dataset (IRDD),” 2023. [Online]. Available: 10.21227/4rmf-gd63. [DOI]
- 9.R. Labib, S. Mim and M.U. Mojumdar, “Rice leaf and crop disease detection dataset,” 2024. [Online]. Available: 10.17632/g7tcwvshff.1. [DOI]
- 10.Simhadri C.G., Kondaveeti H.K., Vatsavayi V.K., Mitra A., Ananthachari P. Deep learning for rice leaf disease detection: a systematic literature review on emerging trends, methodologies and techniques. Inform. Process. Agricult. 2025;12(2):151–168. [Google Scholar]
- 11.A.B. Siddique, A. Datta, S.U. Chowdhury, T.J. Mowla, T.H. Ansari, T. Helaly and N. Begum, “BanglaRiceLeaf: a benchmark dataset for automated Rice leaf disease detection and health classification in Bangladesh,” 2025. [Online]. Available: 10.7910/DVN/XAOBYW. [DOI]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.












