Abstract
Manual diagnosis of retinal diseases from fundus images is slow, subjective, and highly dependent on specialist availability, an ongoing challenge in low-resource regions that leads to delayed treatment and preventable vision loss. This study aims to develop a high-accuracy, automated deep learning framework for early detection of multiple retinal diseases using fundus images, with emphasis on methodological transparency, comparative model evaluation, and applicability in real clinical environments. A dataset of 3,848 labeled fundus images from the University of Gondar Referral Hospital was preprocessed using histogram equalization, noise reduction, normalization, and targeted augmentation to address class imbalance. Five pre-trained CNN architectures, VGG19, ResNet50, InceptionV3, MobileNetV2, and DenseNet201, were fine-tuned using transfer learning. Models were evaluated using accuracy, precision, recall, F1-score, AUC-ROC, and confusion matrices. DenseNet201 achieved the best performance with 92.78% accuracy, demonstrating balanced and superior class-wise detection: F1-scores of 0.98 (Normal), 0.92 (Glaucoma), 0.92 (DR), and 0.86 (AMD). The confusion matrix showed minimal misclassification, with high recall for the underrepresented AMD class (0.91). DenseNet201 consistently outperformed all competing models. The study validates DenseNet201 as a robust, high-accuracy model for early multi-disease retinal screening on a novel, region-specific dataset. The proposed framework offers a scalable, automated diagnostic tool capable of reducing clinical workload and enhancing screening coverage in underserved settings.
Keywords: Deep learning, Fundus images, Retinal disease, DenseNet201, Automated screening, Convolutional neural networks
Subject terms: Computational biology and bioinformatics, Diseases, Health care, Mathematics and computing, Medical research
Introduction
Retinal diseases like Diabetic Retinopathy (DR), Age-related Macular Degeneration (AMD), and Glaucoma are leading causes of irreversible visual impairment globally1–3. Early diagnosis is critical for preventing progression and mitigating healthcare costs4,5. Diagnosis typically relies on funduscopy (ophthalmoscopy), a foundational, non-invasive technique valued for its simplicity6–8. However, the manual interpretation of fundus images is time-consuming, subjective, and highly dependent on specialist availability9,10. These challenges are most acute in low-resource settings, leading to delayed diagnosis and preventable vision loss11.
The last five years have seen deep learning (DL) transform retinal image analysis. Convolutional Neural Networks (CNNs) automatically learn hierarchical features, significantly boosting diagnostic accuracy, efficiency, and accessibility12,13. Numerous studies confirm high performance in multi-disease detection using CNN-based and hybrid architectures14–16, with specialized models employing advanced optimization (e.g., AG-DCNN with RSO17 and explainable AI18 reaching expert-level classification19–21.
Despite these breakthroughs, critical gaps persist. A major limitation is the poor generalizability of models trained on high-quality public datasets when applied to the diverse populations or lower-quality images prevalent in rural clinics11,14,22. Research consistently highlights issues with dataset bias, class imbalance, and model interpretability12,13,23. Furthermore, most successful models focus narrowly on single-disease detection (DR only) or binary classification24,25, leaving multi-class frameworks underexplored. This creates a significant barrier to deploying DL-based screening where it is most needed: a validated, high-accuracy, multi-class model trained on region-specific clinical data, especially from African settings, remains critically lacking.
To address these gaps, this study develops and evaluates a deep learning framework for early retinal disease detection using a novel, region-specific dataset of 3,848 fundus images collected at the University of Gondar Referral Hospital (UOGRH). The study systematically compares five pre-trained CNN architectures (VGG19, ResNet50, InceptionV3, MobileNetV2, and DenseNet201) after robust preprocessing (including histogram equalization and data augmentation) to identify the best-performing, scalable architecture. The findings aim to deliver a practical, locally validated AI tool to support early detection, triage, and improved patient outcomes in resource-limited regions.
Statements of the problem
Despite advances in eye care, retinal diseases such as diabetic retinopathy (DR), age-related macular degeneration (AMD), and glaucoma often remain undetected until significant vision loss occurs. Diagnosis relies heavily on the expertise of individual practitioners, introducing variability that can delay detection and treatment. Current eye health systems predominantly depend on fundus imaging, optical evaluations, and diverse diagnostic methods, which may result in misdiagnoses or missed opportunities for early intervention. Moreover, the costs associated with imaging and eye care screening can further hinder timely access to care. These challenges are particularly pronounced in resource-limited regions, where trained eye health professionals and efficient diagnostic systems are scarce, exacerbating delays in patient care and contributing to preventable vision impairment11,26.
The shortage of eye health professionals and clinics has further compounded these challenges, as limited accessibility to specialized services forces patients to navigate logistical barriers such as long travel distances and constrained clinic availability, delaying or preventing timely treatment11. Many current diagnostic approaches also rely on outdated imaging systems and manual evaluations, which are limited in precision and may fail to detect subtle or early signs of retinal disease. These constraints highlight the need for automated diagnostic systems capable of efficiently analyzing patient data and prioritizing care, ultimately improving diagnostic accuracy, optimizing resource allocation, and enhancing patient outcomes14.
Early identification of retinal disorders is essential to preserving vision before irreversible damage occurs. Conventional diagnostic methods often lack the sensitivity required to detect these conditions at an early stage. Convolutional neural networks (CNNs) have the potential to enhance early detection by identifying subtle patterns in retinal images that may be overlooked by human observers15. However, the generalizability of deep learning models across diverse populations and imaging modalities remains a major challenge. Many models are trained on relatively homogeneous datasets that do not capture the variability seen in real-world clinical environments, raising concerns about their broader applicability11. Furthermore, acquiring large-scale, high-quality datasets with accurately labeled retinal fundus images is difficult due to privacy concerns and reliance on expert annotations. Researchers often employ data augmentation or generate synthetic images to enhance model performance14, yet CNN models trained on specific datasets frequently fail to generalize to unseen data due to differences in imaging equipment, patient demographics, or disease pathophysiology16. Developing CNN models capable of robust generalization across diverse datasets and clinical settings remains a critical hurdle.
This study investigates the application of CNNs for early retinal disease detection, aiming to develop precise and reliable computerized screening tools. The research focuses on evaluating whether CNNs can accurately classify multiple retinal diseases from fundus images, assessing the impact of preprocessing techniques on model performance, comparing different CNN architectures in terms of diagnostic accuracy, and exploring the feasibility of integrating these models into clinical workflows to support ophthalmologists in early diagnosis. By addressing these questions, this work seeks to advance automated systems that improve the speed, accuracy, and accessibility of retinal disease screening.
Contribution of study
This study delivers a pivotal, validated deep learning framework designed for the automated, early classification of multiple retinal diseases (Diabetic Retinopathy, Glaucoma, and Age-related Macular Degeneration) from color fundus images. It introduces a rigorous comparative performance evaluation of five leading Convolutional Neural Network (CNN) architectures VGG19, InceptionV3, ResNet50, DenseNet201, and MobileNetV2 to address the critical research gap concerning limited datasets and the urgent need for early detection. Uniquely utilizing a localized dataset from the University of Gondar Referral Hospital Eye Clinic, the research confirms the superior efficacy of efficient image preprocessing, including class weighting for managing imbalance, combined with transfer learning. This methodological rigor ultimately validates the proposed system’s strength for timely diagnosis and intervention.
The findings establish a new benchmark in automated diagnostics, definitively demonstrating that DenseNet201 is the most powerful architectural tool for early retinal screening, achieving an outstanding overall accuracy of 92.78% after fine-tuning with pre-trained ImageNet weights. Beyond scientific contribution, this research yields a cost-effective and scalable diagnostic solution that significantly advances public health. The model offers the potential to democratize diagnostic technology and support clinical decision-making, particularly in low-resource and rural settings with scarce ophthalmologist access. It directly addresses healthcare accessibility by providing accurate, automated capabilities that can mitigate diagnostic backlogs and ultimately prevent irreversible vision loss.
Finally, the study provides a clear, actionable roadmap for future research, advocating for strategies that will enhance the model’s clinical adoption and generalizability. Key recommendations include expanding the dataset diversity, integrating Explainable AI (XAI) techniques for building clinical trust, exploring ensemble methods, and focusing on the real-time deployment of efficient models like MobileNetV2 on portable screening devices. This research is not merely an addition to the literature; it is a transformative step toward influencing clinical guidelines and policy changes in retinal disease screening worldwide.
Literature review
Funduscopy, or ophthalmoscopy, is the non-invasive, fundamental technique for examining the retina and intraocular structures6–8. It is widely used due to its simplicity and affordability6, serving as the essential method for diagnosing major ocular diseases, including diabetic retinopathy, glaucoma, and AMD6–8. Figure 1 provides a detailed visualization of retinal vessels and pathological signs (e.g., microaneurysms) necessary for timely intervention8.
Fig. 1.

Diabetic retinopathy8.
The integration of Artificial Intelligence (AI) has significantly enhanced funduscopy’s utility, enabling automated analysis (segmentation, classification, prediction) of fundus images with higher accuracy than traditional assessment6. This automation reduces clinician workload and offers a scalable solution for improving patient outcomes, especially where ophthalmologist access is limited6.
However, funduscopy can miss subtle or early-stage changes27. Therefore, for comprehensive evaluation, it is often complemented by advanced modalities like Optical Coherence Tomography (OCT) and fluorescein angiography (FA)27.
As we see in Fig. 2 bel Glaucoma is an eye condition where the optic nerve is damaged, often from raised intraocular pressure. Funduscopy helps detect changes in the optic disc and nerve fiber layer7.
Fig. 2.

Glaucoma7.
Funduscopy is a valuable tool for detecting several retinal conditions. In age-related macular degeneration (AMD), it can reveal characteristic drusen deposits and other changes in the macula. For hypertensive retinopathy, funduscopy helps identify retinal vessel narrowing and hemorrhages caused by high blood pressure. In cases of retinal detachment, it can show the separation of the retina from the underlying tissue along with any associated retinal tears7 see Fig. 3.
Fig. 3.

Rental Detachment7.
Early diagnosis is paramount in ophthalmology, demonstrably improving outcomes and reducing healthcare costs4,5. This need is met by Convolutional Neural Networks (CNNs), which excel in medical imaging by automatically learning features for accurate disease detection12,13. While CNNs significantly enhance diagnostic accuracy and efficiency, their full adoption is hindered by challenges: limited annotated data, poor generalizability, and a critical lack of explainability12,13,23. Overcoming these requires more robust, universally applicable, and clinically interpretable AI solutions.
Recent advances in deep learning have significantly improved automated retinal disease detection and classification using fundus images. Mohamed Elkholy26 proposed a CNN-based framework for early detection of diabetic retinopathy (DR), media haze (MH), and optic disc cupping (ODC), demonstrating that a 12-layer CNN achieved the highest validation accuracy (89.81% for Dataset A). Stewart Muchuchuti28 reviewed deep learning techniques for AMD, DR, glaucoma, and multiple retinal conditions, highlighting CNNs’ hierarchical feature extraction capabilities and the emerging role of vision transformers (ViTs) while discussing challenges like limited labeled data and interpretability.
Toan Duc Nguyen29 developed a deep convolutional ensemble (DCE) model using five InceptionV3 networks, achieving higher accuracy than seven board-certified ophthalmologists in classifying normal, DR, glaucoma, and AMD fundus images. Veena Mayya30 reviewed deep learning approaches, noting sensitivities above 90% and specificities over 95% across public datasets (MESSIDOR, AREDS, APTOS, IDRiD) and emphasizing transfer learning, data augmentation, and ensemble methods to address data scarcity. Anvesha Mongia31 and Ahsan Bin Tufail32 demonstrated the application of deep learning on ultra-wide-field fundus images (UFI), with ResNet152 achieving AUCs above 96%, while heatmaps validated disease localization by ophthalmologists. Santhosh Kumar and Shiva33 and Veena Mayya30 underscored the importance of ensemble architectures, explainable AI, and standardized preprocessing for reliable retinal image analysis.
Balla Goutam22 analyzed preprocessing strategies for automated central ocular disease diagnosis, showing that RoI segmentation and ensemble preprocessing improved F1 and Kappa scores. Ling-Ping Cen20 and Prashant U. Pandey19 proposed hierarchical CNN frameworks for multi-disease classification across 39 ocular conditions, achieving F1-scores above 0.92, high sensitivity, specificity, and AUC, and demonstrating generalization across multi-hospital and public datasets. Mohamed Akil and Rostom Kachouri21 extended AI applications from single-disease detection to comprehensive multi-disease retinal screening using hierarchical CNNs and data augmentation.
Sara Ejaz34 developed a multi-class detection framework, showing that a 12-layer CNN outperformed deeper models on the RFMiD datasets. Ademola Ilesanmi35 systematically reviewed CNN-based retinal segmentation and classification, categorizing models into optic disc/cup, arteries/vein, and retinal vessel segmentation tasks. Mehadi Hasan Hasib36 applied CNNs to classify CNV, DME, drusen, and normal retina images, demonstrating strong generalization across tele-reading and multi-hospital datasets.
Recent advances in deep learning have facilitated automatic glaucoma screening from color fundus images. Notably, Elangovan and Nath (2022)37 proposed En-ConvNet, an ensemble-based method that leverages 13 different pre-trained CNN architectures, explores multiple configuration strategies, and aggregates model outputs using probability averaging followed by an SVM classifier. On publicly available datasets (DRISHTI-GS1, ORIGA, RIM-ONE2, ACRIMA, LAG), they report classification accuracies up to 99.6%, demonstrating superior performance compared to many single-CNN or segmentation-based approaches. By avoiding explicit optic disc/cup segmentation and instead directly classifying raw fundus images, En-ConvNet offers a simpler yet effective pipeline, supporting the potential of ensemble learning + transfer learning as a robust paradigm for large-scale glaucoma screening.
COVID-19Net, proposed by Elangovan et al. (2023)38, introduces a hybrid ensemble model combining a custom ConvNet-24 with several fine-tuned pre-trained CNN architectures. The authors evaluated 126 ensemble configurations and selected the best-performing models for fusion. This approach achieved about 98.5% accuracy on chest X-ray datasets, outperforming many single-CNN methods. COVID-19Net highlights the effectiveness of ensemble learning and transfer learning for robust and reliable COVID-19 detection.
Elangovan et al.39, introduced a compact yet powerful ensemble of CNN models for meat quality assessment, demonstrating that combining lightweight architectures can significantly improve classification accuracy while reducing computational cost. Their method shows the effectiveness of compact ensemble CNNs for reliable and efficient food-quality evaluation.
Hybrid CNN-Transformer models have emerged as a powerful paradigm for medical image analysis, aiming to capitalize on the complementary strengths of both architectures. In the specific domain of DR grading, Khan et al.40 conducted a comprehensive benchmarking comparison of various CNNs (VGG, ResNet, DenseNet, EfficientNet), a standard Vision Transformer, and their proposed hybrids. Their work demonstrated that while a standalone ResNet-50 achieved a strong performance (97.88%), their hybrid EfficientNet-B0 + ViT model surpassed the state-of-the-art accuracy, reaching 98.78%. This highlights the efficacy of using a CNN as a local feature extractor for the transformer’s global attention mechanism, a design insight that informs our architectural choices.
Recent research pioneers hybrid models for retinal diagnostics. Kumar et al. (2023) engineered an advanced deep learning system, marrying Convolutional Neural Networks (CNNs) and transfer learning with specialized feature extraction41. This potent combination was demonstrated to significantly boost diagnostic accuracy, sensitivity, and specificity by adeptly identifying subtle retinal anomalies, thereby validating the critical role of integrated deep learning and image analysis in transforming early detection capabilities.
Kamber et al.42 pioneered a task-centric taxonomy for Explainable AI (XAI) in medical imaging, challenging traditional algorithmic-centric methods by mandating that explanation strategies align directly with specific clinical objectives (classification, prognosis). This paradigm shift is essential for providing verifiable rationales that validate model decisions against the specific clinical task, thereby directly confronting the “black box” problem and critically fortifying clinical trust necessary for AI adoption in healthcare.
Imran et al.43 introduced a Hybrid Convolutional and Recurrent Neural Network (HCRNN) for fundus image-based cataract classification. This pioneering architecture strategically integrates CNNs for robust feature extraction with LSTM units to effectively model the inter-spatial dependencies within the feature maps, overcoming the limitations of purely convolutional methods. The HCRNN achieved superior performance metrics, validating the efficacy of fusing spatial and sequential deep learning for enhanced diagnostic precision in ocular analysis.
Bilal et al.44 introduced DeepSVDNet, an efficacious deep learning framework designed for the detection and multi-stage classification of vision-threatening Diabetic Retinopathy (DR) from retinal fundus images. This novel architecture is specifically optimized to discern subtle pathological features indicative of severe stages (PDR, severe NPDR). DeepSVDNet achieved state-of-the-art performance, validating its utility as a highly precise and reliable automated tool to substantially enhance clinical mass screening efforts.
Ikram and Imran45 introduced the ResViT FusionNet Model, a pioneering Explainable AI (XAI) approach for automated, granular grading of Diabetic Retinopathy (DR). This model achieves superior predictive accuracy by synergistically fusing a ResNet backbone with the global contextual power of a Vision Transformer (ViT). Crucially, the integrated XAI component provides unprecedented transparency, offering clinician’s actionable insights by visually substantiating the diagnostic decision, thereby significantly advancing the trust and clinical efficacy of AI in complex DR screening.
Bilal et al.46 developed a robust hybrid methodology for the precise detection and classification of vision-threatening Diabetic Retinopathy (DR), utilizing an Improved Support Vector Machine (SVM) integrated with CNN-SVD. Their approach employs a CNN for powerful feature extraction, followed by Singular Value Decomposition (SVD) to generate discriminative, compact feature vectors. This synergistic strategy yielded superior performance in differentiating DR stages, validating an effective combination of deep learning and kernel-based classification precision.
Ikram and Imran47. Delivered a comprehensive systematic review on the current status and future trajectory of AI in fundus image-based Diabetic Retinopathy (DR) detection and grading. The work meticulously charts the field’s evolution, critically evaluating trends from CNNs to Vision Transformers. Crucially, it delineates translational challenges and underscores the imperative for robust multimodal learning to enhance the clinical efficacy and global adoption of DR screening AI.
Abbas et al.48 developed an efficient Transfer Learning (TL) methodology for robust detection and grading of cataract using fundus images. By fine-tuning large, pre-trained Convolutional Neural Networks on smaller, domain-specific ocular datasets, their approach effectively mitigates data scarcity issues and the risk of overfitting. This TL strategy demonstrated highly competitive and pragmatic performance for multi-class cataract grading, validating its efficacy in rapidly deploying high-performing models for essential ocular diagnostics.
Khan et al.49 introduced a novel fusion methodology for precise classification of Diabetic Eye Diseases (DED), utilizing Genetic Grey Wolf Optimization (GGWO) to enhance Kernel Extreme Learning Machines (KELM). GGWO served as a powerful meta-heuristic to judiciously select optimal kernel parameters, maximizing the KELM’s discriminative capabilities while maintaining speed. The resultant GGWO-KELM fusion demonstrated superior and robust classification performance, validating a highly effective strategy combining advanced optimization with efficient kernel-based learning for DED diagnostics.
Bilal et al.50 proposed a highly optimized method for breast cancer diagnosis using a Support Vector Machine (SVM) classifier tuned by an Improved Quantum-Inspired Grey Wolf Optimization (IQGWO) algorithm. By leveraging the enhanced exploration capabilities of IQGWO, the study carefully selected the optimal SVM kernel parameters. This innovative IQGWO-SVM combination showed superior classification accuracy and robustness across benchmark datasets, confirming it as an effective and high-precision diagnostic tool that combines SVM’s boundary capabilities with advanced metaheuristic optimization.
Hassan et al51. introduced a multimodal deep learning approach for the robust detection and classification of Alzheimer’s Disease (AD), strategically integrating diverse data sources (e.g., MRI and clinical scores). This methodology leverages deep neural networks optimized for effective feature fusion, capturing complex, interrelated AD pathological markers with enhanced sensitivity and specificity. The model achieved superior classification performance across AD stages (including MCI), validating a highly promising pathway toward more comprehensive and accurate automated diagnosis of this neurodegenerative disorder.
Ikram and Imran52 introduced FastDRNet, a specialized deep learning model for automated Diabetic Retinopathy (DR) detection using Optical Coherence Tomography (OCT) imaging. Architecturally optimized for rapid and accurate processing of complex OCT scans, FastDRNet focuses on efficiently extracting features relevant to sublayer DR changes (macular edema). The model demonstrated both high-speed inference and superior diagnostic accuracy, validating its potential as a time-critical and robust tool for screening and monitoring DR pathology using the volumetric precision of OCT data.
Collectively, Deep learning has revolutionized ophthalmology, moving toward comprehensive, multi-disease retinal screening with high-precision grading. Recent work overwhelmingly validates the pivotal role of CNNs for early detection of DR, AMD, and Glaucoma, achieving near-perfect accuracy via transfer learning and ensemble methods (En-ConvNet). The field’s frontier is defined by hybrid models (CNN-ViT fusions) which establish state-of-the-art accuracy, surpassing standalone CNNs. Crucially, the integration of Explainable AI (XAI) provides necessary transparency and clinical validation, moving AI beyond the “black box” problem. The focus is now on practical scalability, utilizing lighter architectures and innovative fusion techniques to deploy highly robust, generalized, and efficient diagnostic tools for detecting subtle, early-stage retinal anomalies in mass screening.
As stated in Table 1, while convolutional neural networks have demonstrated considerable promise in automating the detection of retinal diseases, prevailing approaches remain constrained by their reliance on homogeneous datasets, limited architectural comparisons, and insufficient attention to clinical deployment in diverse settings53–55. This study directly addresses these critical gaps by introducing a robust deep learning framework distinguished by its foundation on a novel, localized dataset from a tertiary care hospital in Ethiopia, thereby enhancing demographic representation and real-world applicability.
Table 1.
Summary of related works.
| Authors | Objective | Methods/Techniques Used | Dataset(s) Used | Key Results | Limitations/Future Work |
|---|---|---|---|---|---|
| Rupali Chavan24 | DR detection using deep CNNs | CNN ensemble, image preprocessing, TTA | Kaggle DR dataset | High AUC; state-of-the-art on DR detection | Limited to DR only |
| Ahmed Aizaldeen et al.25 | Detect DR and DME using deep learning | CNN trained on retinal images | EyePACS, Messidor-2 | AUC: 0.991; better than ophthalmologists | Need for large annotated datasets |
| Sara Ejaz et al.34 | Early detection of multiple retinal diseases | Deep CNN with data augmentation | RFMiD and RFMiD 2.0 | accuracy of 89.81% for validation, 88.72% | Moderate accuracy; requires improvement; limited dataset |
| Ademola E. Ilesanmi et al.35 | Analyze CNN methods from 2015 onwards | CNN | APTOS, EyePACS, RIM-ONE, ORIGA | ResNet50 outperformed others; AUC ~ 0.96 | Use a limited dataset |
| Prashant U Pandey et al.19 | Classify 39 ocular diseases via hierarchical CNN | 2-level hierarchy, CNNs, multi-source images | 249,620 images (internal & external) | F1-score: 0.923; AUC: 0.9984; expert-level | More data needed for clinical use |
| Stewart Muchuchuti28 | Review on DL for retinal diseases | CNNs, ViTs, segmentation/classification | Literature Review | Highlights ViTs’ potential, challenges like explainability | Need better clinical validation |
| Toan Duc Nguyen29 | Classify DR, AMD, glaucoma, normal | Ensemble of 5 InceptionV3 models | 43,055 images (12 datasets) | Outperformed 7 ophthalmologists; high agreement | Limited test set size |
| Veena Mayya et al.30 | Review of DL in retinal disease detection | CNNs, ViTs, TL, augmentation | Public datasets (MESSIDOR, APTOS, etc.) | Sensitivity > 90%, specificity > 95% | Data imbalance, interpretability |
| Anvesha Mongia et al.31 | Disease detection from UFI images | CNNs (ResNet152, ViT, ConVNext), preprocessing | 4697 UFI images | AUC: 96.47% (ResNet152); heatmaps validated | Need models for specific diseases |
| Santhosh Kumar & Shiva33 | Review on DL for segmentation & classification | CNNs, ViTs, ensembles | MESSIDOR, AREDS, IDRiD, etc. | Comprehensive analysis of techniques | Emphasis on explainability and ensemble CNNs |
| Balla Goutam et al.22 | Evaluate preprocessing on CNN performance | 9 preprocessing techniques + CNNs | ODIR-5 K | RoI + green channel + MSR best combo; F1 ↑ 3%, Kappa ↑ 30% | High compute cost for MSR, MIRNET |
| Ahsan Bin Tufail et al.32 | Classify COD: myopia, DR, AMD, glaucoma | CNNs (ResNet152, ViT, etc.) + preprocessing | 4697 images | ResNet152 AUC: 96.47%; expert-aligned heatmaps | Needs validation on larger sets |
| Ling-Ping Cen et al.20 | Classify 39 fundus conditions | Hierarchical CNN, 2-stage classifier | 249,620 images | F1: 0.923, AUC: 0.9984; strong generalization | More clinical validation required |
| Mohamed Akil & Rostom Kachouri21 | DL platform for 39 fundus conditions | CNN hierarchy + augmentation | 249,620 images, 275,543 labels | Expert-level performance; real-world compatible | Elastic deformations, low-data setting in future |
| Mehadi Hasan Hasib & Chandrika Chowdhury36 | Classify 4 classes (CNV, DME, DRUSEN, normal) | CNNs (ResNet152 best), ANN comparison | 84,495 images | AUC: 96.47%; effective in 7 hospitals | Improve generalization, lightweight models |
We systematically benchmarked five state-of-the-art architectures, VGG19, ResNet50, InceptionV3, MobileNetV2, and DenseNet201 to authoritatively identify the optimal network for this specific clinical task. Our approach is further unique in its bespoke preprocessing pipeline, which integrates advanced histogram equalization and anatomically-preserving augmentation to counteract class imbalance and accentuate subtle pathological features. The result is a highly optimized DenseNet201 model that achieves a benchmark accuracy of 92.78% in multi-disease classification. This work thus delivers a validated, clinically viable tool that not only establishes a new performance standard but is also explicitly engineered for scalable deployment in resource-limited environments, offering a transformative solution to mitigate global diagnostic disparities and preventable vision loss.
Research design and methodology
Research design
The research design used in this study is experimental. This type of design is particularly well-suited for evaluating the performance of deep learning models under controlled conditions. The research focuses on designing, implementing, and assessing Convolutional Neural Networks (CNNs) for the early classification of retinal diseases using fundus images. The entire study is guided by four research questions that cover the process from data preparation to model evaluation and potential clinical integration. The experimental design ensures the scientific rigor of model evaluation and the applied relevance of its potential healthcare impact.
Proposed model architecture
The proposed model architecture for this study is predicated on Convolutional Neural Networks (CNNs), a strategic selection justified by their intrinsic suitability for medical image analysis. While advanced deep learning paradigms exist, CNNs remain the cornerstone for fundus image interpretation due to their unparalleled capacity for learning spatial hierarchies of features through convolutional and pooling operations, which is critical for identifying localized pathological patterns such as microaneurysms, exudates, and optic disc cupping.
This research conducts a systematic comparative analysis of five pre-trained CNN architectures: VGG19, ResNet50, InceptionV3, DenseNet201, and MobileNetV2, selected for their proven efficacy in visual tasks. A key technique employed was transfer learning; we leveraged the convolutional bases of these models, pre-trained on ImageNet weights, to extract robust feature representations. The original classification layers were subsequently replaced and fine-tuned for our specific multi-disease task, with the DenseNet201 model achieving the best performance.
The conceptual system architecture is illustrated in Fig. 4. To enhance domain-specific adaptation, deeper convolutional layers were unfrozen in later training stages, allowing the models to refine their features for retinal characteristics.
Fig. 4.
The proposed architecture.
The final model architecture selected was a fine-tuned DenseNet201 leveraging transfer learning with ImageNet pre-training. The network processes input fundus images that have been resized to 225 × 225 pixels. The architecture comprises the frozen (and later fine-tuned) DenseNet201 convolutional base for feature extraction, followed by a custom classification head. This head consists of a Flatten layer, succeeded by a dense layer and a Dropout layer with a rate of 0.5 for regularization, and culminates in a final dense output layer with a Softmax activation function for the 4-class classification task (DR, AMD, Glaucoma, and Normal). The model was trained using the Adam optimizer and the Categorical Cross-Entropy loss function, with an initial learning rate of 0.0001 and a batch size of 32. Training employed early stopping to monitor validation loss for up to 10 epochs as a further regularization measure.
Data collection and preprocessing
Data source
The data utilized in this study, comprising retinal fundus images, was primarily sourced from the University of Gondar Referral Hospital (UOGRH) Eye Clinic in Ethiopia. The final dataset consisted of 3,848 labeled images, partitioned into four distinct clinical categories: Diabetic Retinopathy (DR) (31% or 1,192 images), Normal (29% or 1,116 images), Glaucoma (26% or 999 images), and Age-related Macular Degeneration (AMD) (14% or 541 images). This collection was split into a training (80%), validation (10%), and testing (10%) set. Before model ingestion, all images were subjected to a rigorous preprocessing pipeline: they were uniformly resized to 225 × 225 pixels, followed by normalization, histogram equalization for contrast enhancement, and noise reduction. Furthermore, to both address the inherent class imbalance and enhance model generalization, the training set was augmented using techniques that included rotation, shifts (horizontal/vertical), zooming, and random flipping. Finally, class weighting was employed during the supervised learning phase to account for the imbalanced distribution. These dataset classes is visually represented in Fig. 5.
Fig. 5.
Dataset classes (a) diabetic retinopathy (DR), (b) glaucoma, (c) age-related macular degeneration (AMD), (d) normal. Source (University of Gondar Referral Hospital Eye Clinic funduscopy images).
Dataset preprocessing
Before training the deep learning model, the retinal fundus images underwent a standardized preprocessing pipeline designed to enhance image quality and standardize inputs for the Convolutional Neural Network (CNN). This structured approach aimed to ensure the model was trained on clean, diverse, and representative images, promoting generalization in retinal disease classification. The preprocessing steps included: resizing all images to a uniform resolution of 225 × 225 pixels to meet the input requirements of the deep learning models. Histogram equalization was applied to enhance contrast and accentuate essential retinal features such as veins and lesions. Noise reduction filters were utilized to eliminate irrelevant visual data, thereby decreasing noise during the feature learning process. Data normalization scaled pixel values to a standard range, typically, which helps in accelerating model convergence and improving performance. To address class imbalance and increase the training sample size, various data augmentation techniques were introduced, including horizontal flipping, rotation (up to 5 degrees), zoom (zoom range = 0.05), brightness changes (brightness range=[0.95, 1.05]), width and height shifts (width shift range = 0.05 and height shift range = 0.05), and shear transformations (shear range = 0.05). Images were also rescaled to 1./255, and new pixels created during transformations were filled using ‘nearest’ mode. Notably, horizontal and vertical flipping were explicitly turned off to preserve the anatomical integrity of the retinal images. Finally, label encoding was performed to convert categorical disease classes into numerical classes suitable for neural network processing. After these preprocessing steps, the datasets were split into three sets: 80% for training, 10% for validation, and 10% for testing.
Data training, splitting and validation
Training and validation were meticulously structured to ensure robust model performance and generalizability. The fundus images were critically partitioned into a rigorous 80% training, 10% validation, and 10% testing split. To mitigate overfitting and ensure reliable performance metrics, the model selection process incorporated a 5-fold Cross-Validation (5-fold CV) regimen on the training and validation data. The high-performing DenseNet201 model was optimized using the Adam optimizer and the Categorical Cross-Entropy loss function. Training hyperparameters were tightly controlled: a stable learning rate of 0.0001, a batch size of 32, and a maximum of 50 epochs. Rigorous regularization was enforced via early stopping (10-epoch patience on validation loss) and a dropout layer (rate 0.5). The entire process was implemented using Keras with a TensorFlow backend, leveraging GPU acceleration to ensure effective, high-fidelity learning on unseen data.
Performance metrics and statistical reliability
The final DenseNet201 model’s effectiveness was rigorously confirmed on the held-out test set by a comprehensive suite of metrics. The architecture achieved an overall test Accuracy of 92.78%. More critically, its multi-class performance demonstrated high statistical reliability, with stable macro-averaged metrics of Precision {0.92}, Recall {0.93}, and F1-score {0.92}. This close agreement among macro-averaged values confirms robust and balanced classification across all four categories (DR, AMD, Glaucoma, and Normal), minimizing bias towards majority classes. The model’s discriminative capacity was further evaluated using the Area under the Receiver Operating Characteristic Curve (ROC-AUC, with the stability of the reported macro-averages validating the high generalizability and reliability of the final results.
Tools and environment
The study was implemented in Python 3.8, using TensorFlow and Keras to develop and train CNN models. Image processing was performed with OpenCV, PIL, and NumPy, while Pandas managed metadata and labels. Scikit-learn facilitated evaluation and comparison with classical algorithms, and Matplotlib visualized data and model performance. Experiments ran on a Windows 11 system with an Intel Core i7 CPU, 16 GB RAM, and an NVIDIA RTX 3060 GPU. Anaconda 2024 managed the environment, and Jupyter Notebook provided an interactive interface for model development and refinement.
Experimental setup
Hardware and software setup
The experiments were conducted on a system with an Intel Core i7 processor, 16GB RAM, and an NVIDIA RTX 3060 GPU. Python 3.8 with TensorFlow and Keras was used for model development, supported by NumPy, Pandas, OpenCV, Scikit-learn, and Matplotlib for data handling, preprocessing, evaluation, and visualization. Anaconda (2024) managed the environment, and Jupyter Notebook facilitated iterative testing.
Dataset setup and preprocessing
Dataset classes
The study utilized a curated and preprocessed dataset of retinal fundus images sourced primarily from the University of Gondar Referral Hospital Eye Clinic. This dataset comprised approximately 4,000 labeled images (specifically 3,848 images), which were stratified into four primary classes: Diabetic Retinopathy (DR), Age-related Macular Degeneration (AMD), Glaucoma, and Normal (healthy retina). Each image in the dataset was associated with a unique class label, serving as the ground truth for supervised learning.
Dataset distribution
The dataset, obtained from the University of Gondar Referral Hospital Eye Clinic, comprised 3,848 retinal fundus images categorized into Diabetic Retinopathy (31%), Normal (29%), Glaucoma (26%), and Age-related Macular Degeneration (14%). Class imbalance, particularly the underrepresentation of AMD, was mitigated using class weighting (1.83) and data augmentation techniques, including flipping, rotation, zooming, and brightness adjustment.
The dataset was split into training, validation, and testing sets in an 80:10:10 ratio to ensure balanced model development and evaluation. Figure 6 illustrates the distribution of these images across the different categories in the training dataset.
Fig. 6.
Dataset distribution.
Dataset characteristics of train/validation/test data set
The dataset of 3,848 retinal fundus images was divided into training, validation, and testing sets using an 80:10:10 ratio to ensure robust model development and unbiased evaluation.The dataset consists of four classes: Age-related Macular Degeneration (AMD), Diabetic Retinopathy (DR), Glaucoma, and Normal (Normal). Specifically, DR comprised 963/120/121 images, Normal 889/111/112, Glaucoma 805/100/102, and AMD 420/≈50/≈50 across training, validation, and testing, respectively. This systematic partitioning ensured sufficient data for training while maintaining independent sets for performance assessment. The number of fundus images assigned to the training, testing, and validation sets for each disease category is shown in Fig. 7.
Fig. 7.
Database distribution per class.
Class weights
As shown in the dataset distribution, class imbalance was most evident in AMD. To mitigate this, class weights were used together with augmentation, with weights assigned inversely proportional to class frequencies. As a result, AMD received the highest weight (1.83), while DR, Glaucoma, and Normal were weighted 0.80, 0.96, and 0.87, respectively (Fig. 8), helping to reduce imbalance effects during training.
Fig. 8.
Class weights.
The weighting highlights the severe underrepresentation of AMD, with only 420 training samples compared to DR (963), Glaucoma (805), and Normal (889). Without correction, the model would likely bias toward the majority classes, increasing false negatives for AMD. By assigning a weight of 1.83 over twice that of DR the model emphasizes AMD features during training, thereby mitigating class imbalance and improving detection performance.
Data augmentation
To address issues of class imbalance and enhance the generalization capabilities of the model, data augmentation techniques were extensively applied to the training dataset. This process involved synthetically increasing the variation and number of training samples, particularly for underrepresented classes, which helped to reduce potential overfitting and allowed the model to learn the variability in fundus images more effectively, preventing bias towards majority classes. The data augmentation was implemented using the ImageDataGenerator class from Keras. Specific parameters included rescaling all pixel values to a range of (rescale = 1./255), a rotation_range of 5 degrees to account for angular differences, width_shift_range = 0.05 and height_shift_range = 0.05 for minor translations, and shear_range = 0.05 and zoom_range = 0.05 for slight geometric distortions. Additionally, brightness_range = [0.95, 1.05] was used to simulate lighting changes. Notably, both horizontal flip and vertical flip were explicitly turned off to preserve the anatomical integrity of the retinal images, and fill_mode=’nearest’ was employed for new pixels created during transformations. These techniques, alongside class weighting, were critical in mitigating the impact of the lower representation of the Age-related Macular Degeneration (AMD) class, thereby improving CNN learning and minimizing overfitting. A sample-generated image is shown in Fig. 9 blow.
Fig. 9.
Sample augmented images (Augmented dataset).
Model training and evaluation
This study trained, validated, and tested multiple CNN architectures for retinal disease detection using the preprocessed and augmented dataset, with performance evaluated through standard metrics. The selected models, chosen for their proven success in computer vision tasks, are described in detail to highlight their design characteristics and intended purposes.
Training Performance (accuracy/loss plots per model):
Among the tested CNN architectures, DenseNet201 achieved the highest training and validation performance, with its dense connectivity enabling efficient feature reuse and stable convergence.
DenseNet201
As represented Fig. 10, among the tested CNN architectures, DenseNet201 achieved the highest training and validation performance, with its dense connectivity enabling efficient feature reuse and stable convergence.
Fig. 10.
DenseNet201 Training and validation accuracy.
DenseNet201 exhibited steadily increasing training and validation accuracy with minimal gap between the curves, indicating effective learning and strong generalization. The absence of oscillations or divergence suggests the model avoided overfitting during training.
The training and validation loss curves for DenseNet201 steadily decreased in tandem, with no sudden jumps, indicating stable learning and proper fitting of both distributions (See Fig. 11). Overall, DenseNet201 achieved high accuracy, smooth convergence, and minimal overfitting, making it well-suited for tasks requiring deep feature extraction and generalization. Future work may explore hyperparameter tuning or learning rate schedules to further enhance performance.
Fig. 11.
DenseNet201 Training and validation loss.
InceptionV3
Consistent with prior studies, InceptionV356, included for its inception modules and computational efficiency, achieved acceptable classification results. However, its performance was slightly less stable and consistent compared to DenseNet201, as reflected in training and evaluation curves, which provide insight into learning behavior, generalization, and future optimization directions.
As illustrated in Fig. 12, the training accuracy of InceptionV3 showed steady growth, while validation accuracy followed a similar trend with greater variability, indicating weaker generalization to unseen data. This may reflect the model’s complexity and the need for more careful hyperparameter tuning, such as learning rate and batch size, to achieve consistent performance.
Fig. 12.
InceptionV3 Training and validation accuracy.
As depicted in Fig. 13, the training loss of InceptionV3 steadily decreased, while validation loss showed minor fluctuations and occasional increases, indicating some overfitting. Despite this, the model achieved strong accuracy due to its inception modules, which capture multi-scale, hierarchical features. However, InceptionV3 is sensitive to hyperparameters and training conditions, leading to slightly unstable validation performance compared to the more consistent DenseNet201.
Fig. 13.
InceptionV3 Training and validation loss.
MobileNetV2
MobileNetV2, designed for computational efficiency and mobile applications57, showed a modest overall accuracy but demonstrated a stable and predictable training pattern compared to deeper CNN architectures.
As presented in Fig. 14, MobileNetV2 exhibited a stable, linear increase in training accuracy, reflecting rapid convergence despite its simpler architecture and fewer parameters. Validation accuracy closely followed the training curve with minor fluctuations, indicating good generalization and minimal overfitting.
Fig. 14.
MobileNetV2 training and validation accuracy.
The loss curves for MobileNetV2 further confirmed its stable learning, with training and validation loss decreasing steadily and only minor fluctuations observed. The absence of spikes in validation loss indicates effective regularization, demonstrating rapid learning while maintaining strong generalization (see Fig. 15).
Fig. 15.
MobileNetV2 training and validation loss.
In summary, MobileNetV2, while not achieving the highest accuracy, demonstrated stable training, minimal overfitting, and reliable convergence. Its efficiency and compact size make it well-suited for real-time applications where speed and model footprint are prioritized, despite a modest trade-off in accuracy.
ResNet50
ResNet50, a 50-layer deep CNN utilizing residual learning and identity shortcuts, mitigates the vanishing gradient problem and enables effective training of deeper networks.
As demonstrated in Fig. 16, the ResNet50 model showed slow improvement in training accuracy, rising from ~ 23% to just over 31% across 20 epochs, suggesting underfitting potentially due to insufficient fine-tuning or excessive frozen layers. Validation accuracy was erratic, oscillating between 19% and 58%, indicating generalization issues likely related to class imbalance, data augmentation inconsistencies, or suboptimal learning rate settings, with occasional validation values exceeding training accuracy.
Fig. 16.
ResNet50 training and validation accuracy.
As Fig. 17 indicates, for ResNet50, training loss decreased steadily from ~ 1.59 to 1.37 and validation loss from ~ 1.39 to 1.34, indicating effective learning and stable generalization with minimal overfitting. The consistent gap between the curves and plateauing of validation loss toward the end suggest incomplete convergence, highlighting opportunities for further hyperparameter tuning.
Fig. 17.
ResNet50 training and validation loss.
VGG19
VGG19, a straightforward 19-layer CNN, was evaluated alongside other models. Its stacked 3 × 3 convolutional filters and max-pooling layers enable deeper feature extraction while preserving fine spatial resolution.
As observed in Fig. 18, VGG19 training accuracy steadily increased from ~ 28% to ~ 68% over 20 epochs, while validation accuracy rose rapidly from ~ 37% to above 80% by epoch 7 and remained stable, indicating strong generalization. The large gap between training and validation accuracy may suggest underfitting or early regularization effects.
Fig. 18.
VGG19 training and validation accuracy.
As demonstrated in Fig. 19, VGG19 training and validation loss steadily decreased from ~ 1.44 to 1.04 and ~ 1.35 to 0.99, respectively, indicating productive learning with minimal overfitting. While its simpler, regular architecture led to slower convergence and slight susceptibility to overfitting, VGG19 demonstrated solid performance and serves as a reliable benchmark for comparing classical CNNs with newer architectures.
Fig. 19.
VGG19 training and validation loss.
Comparative analysis of CNN models and best model selection
Validation accuracy for the selected CNN architectures DenseNet201, MobileNetV2, InceptionV3, VGG19, and ResNet50 is presented in Fig. 20, providing a clear comparative view of each model’s performance on the validation dataset.
Fig. 20.
Accuracy comparison of CNN models.
DenseNet201 achieved the highest validation accuracy (0.89), benefiting from dense connectivity that enhances feature propagation and generalization. MobileNetV2 and InceptionV3 followed closely with accuracies of 0.86 and 0.85, respectively, with MobileNetV2 offering an efficient trade-off between accuracy and computational cost, and InceptionV3 leveraging multi-scale processing. VGG19 reached 0.83, providing a solid but less advanced benchmark, while ResNet50 performed the poorest at 0.58, likely due to suboptimal hyperparameters or limited dataset generalization. Overall, the results highlight both the importance of architecture and its suitability to the dataset in achieving optimal CNN performance.
As illustrated in Table 2, while DenseNet201 already demonstrated the highest validation accuracy among the tested models, its performance was further enhanced through fine-tuning. By freezing earlier layers and training on a new dataset for 10 epochs, validation accuracy improved from 0.89 to 92.78%, reflecting the benefits of dense connectivity, feature propagation, and mitigated vanishing gradients in achieving superior generalization.
Table 2.
Summary of comparative analysis of CNN models.
| Model | Final Val Accuracy | Overfitting | Convergence | Generalization |
|---|---|---|---|---|
| DenseNet201 | High | No | Smooth | Excellent |
| InceptionV3 | High | Mild | Good | Very Good |
| MobileNetV2 | Moderate | Slight | Slower | Fair |
| ResNet50 | High | Yes | Fast | Good |
| VGG19 | Low | Yes | Fast | Poor |
Comparative analysis and baseline benchmarking
To establish a fair performance benchmark for the multi-class retinal classification task, a systematic comparative analysis was conducted across five leading pre-trained CNN architectures: VGG19, ResNet50, InceptionV3, MobileNetV2, and DenseNet201. All models were implemented using the same transfer learning approach, leveraging ImageNet weights and fine-tuning a custom classification head. The performance comparison, using validation accuracy, definitively established the baseline for the study. The DenseNet201 architecture achieved the highest validation accuracy of 0.89, demonstrating superior feature reuse and generalization. In contrast, VGG19 provided a classical benchmark with an accuracy of 0.83, while ResNet50 performed the poorest at 0.58. The final optimized DenseNet201 model ultimately achieved an overall test accuracy of 92.78%. Note that architectures such as DenseNet121 and the EfficientNet family were not included in this direct comparative evaluation. The comparative performance benchmark discussed in Table 3.
Table 3.
Comparative performance benchmark (Validation accuracy).
| CNN architecture | Validation accuracy | Description |
|---|---|---|
| DenseNet201 | 0.89 | Highest performance; selected as the final model. |
| MobileNetV2 | 0.86 | Strong efficiency-accuracy trade-off. |
| InceptionV3 | 0.85 | Utilizes multi-scale feature processing. |
| VGG19 | 0.83 | Solid, established classical benchmark. |
| ResNet50 | 0.58 | Lowest performance among tested models. |
Significant values are in bold.
Evaluation metrics analysis for best model
DenseNet201 performance in classifying retinal fundus images was evaluated using a normalized confusion matrix and a classification report, providing detailed and summary assessments across the four diagnostic categories: AMD, DR, glaucoma, and normal.
Figure 21 presents the normalized confusion matrix for DenseNet201, illustrating classification performance across four classes: AMD, DR, glaucoma, and normal. Each cell shows the percentage of predictions for a given true label (rows) against predicted labels (columns), providing a detailed assessment of multi-class performance.
Fig. 21.
Confusion matrix for DenseNet201.
The DenseNet201 model demonstrated strong performance in classifying retinal diseases, achieving the highest accuracy in identifying normal fundus images at 97.3%, effectively distinguishing healthy retinal anatomy from pathological cases. Among disease categories, glaucoma achieved the second-highest accuracy of 92.2%, with misclassifications mainly into AMD and DR, but no cases were misclassified as normal, reflecting the model’s reliability in detecting structural optic nerve damage. AMD and DR were classified with accuracies of 90.6% and 90.1%, respectively, with misclassifications occurring primarily due to overlapping visual features, such as drusen, hemorrhages, exudates, and vascular changes, between these disease classes. Overall, the model showed robust ability to identify both healthy and diseased retinal images, with most errors arising from subtle similarities in pathological features across disease categories.
To transition from a black-box model to a transparent, clinically actionable diagnostic system, Gradient-weighted Class Activation Mapping (Grad-CAM) was systematically implemented. This critical interpretability step generates high-resolution localization maps that precisely delineate the regions within the input fundus image most influential to the DenseNet201’s classification output. The process begins by forward propagating the input image through the trained network to obtain the class score. Subsequently, the gradient of this predicted score is computed with respect to the final convolutional feature maps. These gradients are then spatially averaged to derive neuron importance weights, which quantify each feature map’s contribution to the final decision. Finally, the weighted feature maps are summed and passed through a ReLU activation to construct the final, high-fidelity Grad-CAM heatmap. This heatmap is then overlaid onto the original fundus image (as exemplified in Fig. 22) to visually validate that the model is fixating on authentic pathological indicators such as microaneurysms, hemorrhages, or optic disc cupping rather than on non-pathological image artifacts. This robust, systematic visualization is indispensable for establishing the model’s trustworthiness and accelerating its clinical adoption.
Fig. 22.
Confusion matrix for DenseNet201.
To complement the global insights provided by Grad-CAM and offer local, model-agnostic explanations, the Local Interpretable Model-agnostic Explanations (LIME) technique was also utilized. LIME specifically addresses interpretability by explaining the prediction of the complex DenseNet201 model for a single, specific fundus image. The methodology begins by segmenting the input image into superpixels, which represent small, perceptually homogeneous regions. The core of LIME involves generating numerous perturbations of the original image by randomly turning off (masking) a subset of these superpixels. Each perturbed image is then passed through the deep learning model to obtain corresponding predictions. Finally, a weighted linear model is fitted locally around the instance being explained, approximating the CNN’s complex behavior in that specific local region of the feature space. The resulting output highlights the superpixels most contributing to the predicted class, providing a critical visual map that is independent of the CNN’s internal architecture. This dual-explanation approach using both Grad-CAM and LIME provides robust, verifiable evidence for the model’s decision-making process, significantly enhancing clinical trust and utility.
Classification report analysis by accuracy, precision, recall, and F1-score
Table 4 presents the DenseNet201 model’s classification results on a held-out test set of 388 retinal fundus images. The evaluation utilized standard metrics, precision, recall, F1-score, and support for each class (AMD, DR, Glaucoma, and Normal), offering a detailed assessment of the model’s discriminative ability and its generalizability to unseen data.
Table 4.
Classification report.
| Precision | Recall | F1-Score | Support | |
|---|---|---|---|---|
| Amd | 0.83 | 0.91 | 0.86 | 53 |
| Dr | 0.93 | 0.90 | 0.92 | 121 |
| Glaucoma | 0.92 | 0.92 | 0.92 | 102 |
| Normal | 0.98 | 0.97 | 0.98 | 112 |
| Accuracy | 0.93 | 388 | ||
| Macro Avg | 0.92 | 0.93 | 0.92 | 388 |
| Weighted Avg | 0.93 | 0.93 | 0.93 | 388 |
Accuracy: 92.78%.
Significant values are in bold.
AMD (Age-related macular degeneration), DR (diabetic Retinopathy), Glaucoma, and Normal (Healthy)
The model demonstrated strong and balanced performance across all retinal categories. For AMD, it achieved an accuracy of 0.83, a recall of 0.91, and an F1 score of 0.86, effectively identifying most AMD cases, with occasional misclassification of non-AMD images likely due to retinal features such as drusen or pigmentary changes resembling early signs of other pathologies. In the DR category (support = 121, the largest test set), the model attained a precision of 0.93, a recall of 0.90, and an F1 score of 0.92, reflecting both the clear separability of DR from other conditions and successful detection of nearly all DR cases, capturing key features such as microaneurysms, hemorrhages, and exudates. For Glaucoma, the model achieved a precision of 0.92, a recall of 0.92, and an F1 score of 0.92, indicating balanced classification with minimal false positives and false negatives, despite the subtle nature of glaucomatous changes, such as optic disc cupping and nerve fiber layer thinning. The Normal class achieved very high performance, with a precision of 0.98, a recall of 0.97, and an F1 score of 0.98, demonstrating the model’s strong ability to accurately classify healthy retinal images, which is critical for avoiding unnecessary clinical follow-up or patient anxiety. Overall, the model achieved an accuracy of 92.78%, correctly classifying 360 out of 388 test cases. Macro-averaged metrics were stable at 0.92 for precision, 0.93 for recall, and 0.92 for F1, while weighted averages were similarly high at 0.93, indicating consistent performance across classes. The close agreement between macro and weighted averages suggests that the model maintained balanced performance across both majority and minority classes, without disproportionately favoring more prevalent categories such as DR. Collectively, these results demonstrate the model’s strong capability to identify critical retinal pathologies, supporting its potential utility in clinical screening workflows.
Interpretation and implications
The DenseNet201 model demonstrated robust performance in multi-class retinal disease classification, achieving an overall accuracy of 92.78%, correctly classifying 360 out of 388 test cases. Macro-averaged metrics were stable at 0.92 for precision, 0.93 for recall, and 0.92 for F1 score, while weighted averages were similarly high at 0.93, indicating consistent performance across all classes and minimal bias toward more prevalent categories such as DR.
Class-specific performance further highlighted the model’s effectiveness. For AMD, it achieved an accuracy of 0.83, a recall of 0.91, and an F1 score of 0.86, with occasional misclassification of non-AMD images likely due to retinal features such as drusen or pigmentary changes resembling early signs of other pathologies. DR (support = 121) attained a precision of 0.93, a recall of 0.90, and an F1 score of 0.92, reflecting successful detection of nearly all cases and accurate identification of microaneurysms, hemorrhages, and exudates. Glaucoma achieved a precision of 0.92, a recall of 0.92, and an F1 score of 0.92, indicating balanced classification despite the subtle nature of glaucomatous changes such as optic disc cupping and nerve fiber layer thinning. The Normal class exhibited the highest performance, with a precision of 0.98, a recall of 0.97, and an F1 score of 0.98, demonstrating the model’s ability to accurately classify healthy retinal images, crucial for avoiding unnecessary clinical follow-up or patient anxiety.
Analysis of the confusion matrix revealed a strong diagonal structure, with the vast majority of samples correctly classified across all categories. High recall rates suggest reliable detection of pathological cases, while high precision for Normal and DR adds confidence in minimizing false positives, which is critical in real-world screening and triage workflows. Some misclassification, particularly between DR and Glaucoma, indicates room for improvement, potentially achievable through additional preprocessing, class-feature augmentation, incorporation of clinical metadata, or advanced methods such as attention mechanisms.
Overall, the DenseNet201 model demonstrates strong capability in distinguishing between pathological and healthy retinal images, validating its potential for automated retinal disease screening. Its ability to identify patients with immediate care needs can support more efficient triage and clinical decision-making, particularly in settings with limited access to expert ophthalmologists, establishing DenseNet201 as a promising tool for computer-aided diagnosis in ophthalmology.
Interpretation and novelty of findings
This study delivers a pivotal advancement in automated ophthalmology by transcending conventional model benchmarking to address the critical translational gaps of demographic bias and clinical practicality. Our framework’s principal innovation lies in its foundation on a novel, localized dataset from an Ethiopian population, directly countering the Western-centric data hegemony that limits global AI applicability. We authoritatively establish DenseNet201 as the optimal architecture through a definitive comparative analysis, providing an empirical cornerstone for future clinical implementations. Beyond accuracy, our contribution is distinguished by a deployment-conscious design integrating anatomically-grounded preprocessing and computational efficiency to forge a robust, equitable diagnostic tool expressly engineered for real-world impact in resource-constrained settings, thereby bridging a critical chasm between algorithmic potential and scalable, life-changing healthcare delivery.
Comparison of related works with our study
Table 5 abov summarizes a comparative study of retinal disease detection from 2021 to 2025, emphasizing how the current research addresses specific limitations found in previous works by Chavan, Ejaz, and Nguyen. While earlier models often struggled with data imbalance, high computational costs, or limited geographic scope, the 2025 study achieves a 92.78% accuracy using a fine-tuned DenseNet201 model. Key contributions include the use of a novel Ethiopian dataset, anatomically-aware augmentation, and efficient preprocessing designed specifically for low-resource clinical settings.
Table 5.
Comparison of related works with our study.
| S.No | Authors | Study Focus/Title | Model Used | Country | Result (Accuracy/Metric) | Year | Gap Identified | Our Study Contribution |
|---|---|---|---|---|---|---|---|---|
| 1 | Rupali Chavan24 | DR detection using deep CNNs | CNN ensemble + preprocessing | Not specified | High AUC | 2024 | Limited to DR only | Multi-disease classification (DR, AMD, Glaucoma, Normal) |
| 2 | Sara Ejaz et al34. | Early detection of multiple retinal diseases | Deep CNN + augmentation | Not specified | 89.81% validation accuracy | 2024 | Moderate accuracy; limited dataset | Higher accuracy (92.78%) on localized Ethiopian dataset; comprehensive preprocessing pipeline |
| 3 | Toan Duc Nguyen et al29. | Classify DR, AMD, glaucoma, normal | Ensemble of 5 InceptionV3 | Not specified | Outperformed 7 ophthalmologists | 2023 | Limited test set size | Used larger local dataset (3,848 images); systematic benchmarking of 5 architectures |
| 4 | Prashant U Pandey et al19. | Classify 39 ocular diseases via hierarchical CNN | Hierarchical CNN | Not specified | F1-score: 0.923, AUC: 0.9984 | 2023 | More data needed for clinical use | Focused on early-stage detection of common retinal diseases; validated on clinical data from Ethiopian hospital |
| 5 | Veena Mayya et al30. | Review of DL in retinal disease detection | CNNs, ViTs, TL, augmentation | Not specified | Sensitivity > 90%, specificity > 95% | 2022 | Data imbalance, interpretability issues | Implemented class weighting and anatomically-preserving augmentation to address imbalance; high recall for minority class (AMD: 0.91) |
| 6 | Anvesha Mongia et al31. | Disease detection from UFI images | ResNet152, ViT, ConvNeXt | Not specified | AUC: 96.47% | 2022 | Need models for specific diseases | Provided disease-specific performance metrics; DenseNet201 achieved balanced F1-scores across all classes |
| 7 | Balla Goutam et al22. | Evaluate preprocessing on CNN performance | 9 preprocessing + CNNs | Not specified | F1 ↑ 3%, Kappa ↑ 30% | 2022 | High compute cost for some methods | Used efficient preprocessing (histogram equalization, noise reduction) without heavy computation; suitable for resource-limited settings |
| 8 | Ling-Ping Cen et al20. | Classify 39 fundus conditions | Hierarchical CNN | Not specified | F1: 0.923, AUC: 0.9984 | 2021 | More clinical validation required | Clinical validation on Ethiopian population; model tailored for early detection in underrepresented regions |
| 9 | Our Study | Early retinal disease detection from fundus images using deep neural networks | DenseNet201 (fine-tuned) | Ethiopia | Accuracy: 92.78%, F1-scores: 0.86–0.98 | 2025 | — | Novel localized dataset; systematic CNN benchmarking; anatomically-aware augmentation; high performance on minority classes; designed for deployment in low-resource settings |
Limitations of the study
While this study demonstrates the effectiveness of deep learning models for early retinal disease detection, several limitations should be acknowledged: The model was trained and tested on a single-institution dataset from one hospital, which may limit generalizability across populations and imaging devices. Class imbalance persisted despite augmentation, particularly for AMD. The dataset size (~ 4,000 images) is modest, and no external validation was performed. The work did not explore ensemble methods, which could boost robustness, nor integrate explainable AI for clinical transparency. Additionally, hyperparameters were not systematically optimized, and the model only classifies four retinal conditions, omitting other important diseases. These constraints highlight the need for more diverse data, external validation, and interpretable, ensemble-based approaches in future work.
Results and discussion
The comparative analysis of deep learning architectures on the localized UOGRH dataset validates the fine-tuned DenseNet201 as the optimal model for early, multi-class retinal disease detection. This model achieved the highest benchmark accuracy of 92.78% on the held-out test set, demonstrating robust performance across the four critical classes (Normal, Glaucoma, Diabetic Retinopathy (DR), and Age-related Macular Degeneration (AMD)).
Critically, the model exhibited balanced diagnostic capability, evidenced by high macro-averaged metrics (0.92 precision, 0.93 recall, 0.92 F1-score). While the Normal class showed near-perfect reliability (F1-score: 0.98), the strong and balanced F1-scores of 0.92 for both DR and Glaucoma confirm the model’s clinical utility. The most challenging class, AMD, achieved a notable F1-score of 0.86 and a high recall of 0.91, which is significant for a minority class, demonstrating the model’s effectiveness in minimizing false negatives, a key requirement for early screening. Analysis of the confusion matrix indicated minor residual confusion between DR and Glaucoma, likely stemming from shared or overlapping visual features of vascular and disc changes in early-stage disease.
The findings establish a significant advancement compared to existing literature, primarily by addressing the prevailing issues of data generalization and single-disease focus. Many state-of-the-art models are validated on large, public, high-resource datasets, which often fail to translate effectively to diverse populations and low-quality images found in rural or low-resource clinical settings. This study leverages a novel, localized dataset of 3,848 images from the University of Gondar Referral Hospital (UOGRH) in Ethiopia, thereby providing a regionally validated, high-accuracy multi-class framework for the most prevalent conditions, a substantial contribution for ophthalmic care in the African context. The achieved 92.78% accuracy compares favorably against similar multi-class frameworks reported in the literature, which typically report validation accuracies in the 89–91% range.
The superior performance of DenseNet201 over competing architectures (VGG19, ResNet50, InceptionV3, and MobileNetV2) is directly attributable to its core architectural innovation: the dense connectivity pattern. DenseNet’s design promotes feature reuse by connecting each layer to every subsequent layer in a feed-forward manner. This mechanism is highly effective in image recognition tasks involving subtle, localized features, such as those defining early-stage microaneurysms or minute optic disc cupping. By ensuring that fine-grained information from earlier convolutional layers is passed directly to deeper layers, feature degradation is mitigated, and the network can construct highly discriminative representations. Furthermore, this structure facilitates excellent gradient flow, reducing the vanishing gradient problem and enabling deeper, more effective training than standard residual networks. This inherent feature-extraction power was amplified by the studies robust, anatomically aware preprocessing pipeline, which included histogram equalization and noise reduction, thereby enhancing the contrast of key retinal features. This allowed DenseNet201 to maximize feature learning from challenging, real-world images.
Clinical implications
The deployment of this high-accuracy (92.78%) DenseNet201 model signifies a paradigm shift for proactive blindness prevention. In resource-constrained environments, where the ophthalmologist-to-patient ratio is critically low, this validated, automated screening tool offers immediate, reliable, and triage-level diagnostic support. The early and precise multi-class detection of Glaucoma, Diabetic Retinopathy (DR), and Age-related Macular Degeneration (AMD) is crucial for interrupting disease progression and preventing irreversible vision loss through timely therapeutic intervention (laser photocoagulation or pressure-lowering agents). The model’s 98% F1-score reliability in correctly identifying the Normal class rapidly clears healthy individuals, drastically reducing the bottleneck in mass screening programs. By enabling accessible, cost-effective, and non-invasive identification of high-risk patients, this system directly addresses global healthcare disparities, offering a translational solution to reduce the growing burden of preventable blindness, particularly within underserved populations.
Conclusion
This study successfully developed a robust and translationally relevant deep learning framework for the early, multi-class diagnosis of retinal diseases using fundus images. We systematically compared five state-of-the-art CNNs, demonstrating that the fine-tuned DenseNet201 architecture achieved the highest performance, attaining a benchmark multi-class accuracy of 92.78% on our localized UOGRH test dataset. This technical contribution validates the superior efficacy of DenseNet’s design specifically its dense connectivity and feature reuse for extracting subtle, critical pathological indicators from real-world, contextually challenging images.
The clinical relevance of this system is immediate and significant, particularly in addressing global healthcare disparities. By providing a highly accurate, automated, and non-invasive diagnostic tool, the system delivers reliable triage-level support, drastically improving the ophthalmologist-to-patient ratio in resource-constrained settings. This capability allows for the early interception of major causes of preventable blindness Diabetic Retinopathy, Glaucoma, and Age-related Macular Degeneration through timely referral and therapeutic intervention.
To advance this technology toward widespread clinical adoption, future research will strategically focus on three key directions: (1) establishing global generalization and robustness via rigorous multi-site external validation, (2) improving clinical acceptance and trust by integrating Explainable AI (XAI) techniques, and (3) significantly boosting diagnostic precision, particularly for challenging early-stage conditions, through multimodal fusion of fundus images with structural data from Optical Coherence Tomography (OCT) in a subsequent prospective clinical trial.
Limitations and future work
Limitations
The study’s primary constraint is the dataset’s size and regional domain specificity. While novel, the 3,848-image UOGRH dataset limits comprehensive assessment of generalization across vastly different global populations or rare, less-represented pathologies. Consequently, the high-accuracy DenseNet201 model, optimized for prevalent African-context conditions, requires rigorous external validation to confirm its efficacy in distinct clinical and ethnic environments. The current scope is restricted to four major disease classes; expanding the diagnostic spectrum to include a wider range of ocular diseases remains an imminent challenge.
Future work
Future research must prioritize avenues that translate this foundational success into broader clinical utility.
Generalization and Interpretability: Future work will focus on multi-site deployment and evaluation to establish global generalization capabilities. Crucially, integrating Explainable Artificial Intelligence (XAI), such as Grad-CAM, will be essential to provide clinicians with visual evidence of the model’s decision-making process, thereby enhancing clinical trust and facilitating crucial error analysis.
Multimodal Imaging Integration: A critical next step is the adoption of multimodal data fusion. Future deep learning models will be engineered to integrate features from fundus images and high-resolution Optical Coherence Tomography (OCT) scans. This fusion is expected to yield superior diagnostic accuracy, particularly for subtle conditions like early Age-related Macular Degeneration (AMD) and Glaucoma, which benefit significantly from structural OCT data (e.g., nerve fiber layer thickness).
Prospective Clinical Validation: The ultimate translational step is a prospective, randomized clinical trial to validate the AI system in a real-time, resource-constrained clinical workflow. This will provide necessary comparative data against human expert diagnoses, paving the way for regulatory approval and routine integration into ophthalmic screening protocols.
Abbreviations
- AG-DCNNL
Active gradient deep convolutional neural network
- AI
Artificial intelligence
- AMD
Age-related macular degeneration
- CNN
Convolutional neural network
- DCE
Deep convolutional ensemble
- DLP
Deep learning platform
- DR
Diabetic retinopathy
- HR
Hypertensive retinopathy
- OCT
Optical coherence tomography
- RD
Retinal detachment
- ReLU
Rectified linear unit
- RSO
Red spider optimization
- UOGRH
University of Gondar referral hospital
- XAI
Explainable AI
Author contributions
All authors contributed equally to the conception, design, implementation, and analysis of the study. Specifically, Abenet Alazar Hailu and Esubalew Asmare Desta contributed to the conceptualization, model development, and manuscript drafting. Atsedemaryam Mulugeta, Fikadu Berie Adugna, and Melsew Belachew were involved in the performance evaluation, experimental analysis, and data interpretation.Ayodeji Olalekan Salau and Lamesgin Addis Almaw provided methodological guidance, reviewed the manuscript critically, and contributed to the refinement of the final paper. All authors reviewed and approved the final version of the manuscript.
Data availability
The datasets generated and/or analysed during the current study are publicly available in the Kaggle repository, https://www.kaggle.com/datasets/esubalewasmare/retinal-disease-fundus-images-datasets.To ensure full reproducibility, the complete source code, model configurations, training logs, and pre-trained weights have been made publicly available on Google Drive at: https://drive.google.com/file/d/1EyxpomtXip2Q04Dq7w-2pYgw0_QNgWL5/view? usp=sharing.
Declarations
Competing interests
The authors declare no competing interests.
Consent to participate
The waiver of informed consent was approved by the Institutional Review Board (IRB) of the University of Gondar College of Medicine and Health Sciences.
Ethical considerations and data privacy
The use of retinal fundus images from the University of Gondar Referral Hospital (UOGRH) Eye Clinic received ethical clearance from the University of Gondar Institutional Review Board (IRB) for retrospective data use. To ensure patient confidentiality, the dataset was fully de-identified before analysis. This process removed all personally identifiable information (PII) and any auxiliary patient clinical data, ensuring the study relied solely on anonymized fundus images and their corresponding diagnostic labels.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Institute, N. E. Diabetic retinopathy. US Department of Health, Education, and Welfare, Public Health Service ….
- 2.Simon, S. D., Centers for disease control and prevention (CDC) centers for disease control and prevention (CDC). In Encyclopedia of Big Data 158–161 (Springer, 2022). [Google Scholar]
- 3.Organization, W. H. World report on vision, (2019).
- 4.David, S. E. S. & French, P. Rachael Powell promoting early detection and screening for disease., (Springer, 2018). 10.1007/978-0-387-93826-4_18.
- 5.Aron Halfin, M. Depression: The benefits of early and appropriate treatment. The Am. J. Managed Care, 13, 4, (2007). [PubMed]
- 6.Driban, M. et al. Artificial intelligence in chorioretinal pathology through fundoscopy: A comprehensive review. Int. J. Retina Vitreous.10(1), 36 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Sanghavi, J. & K., M. Ocular disease detection systems based on fundus images: A survey. Multimed. Tools Appl.83, 21471–21496. 10.1007/s11042-023-16366-x (2023). [Google Scholar]
- 8.Xu, Y. W. Y. et al. The diagnostic accuracy of an intelligent and automated fundus disease image assessment system with lesion quantitative function (SmartEye) in diabetic patients. BMC. Ophthalmol.10.1186/s12886-019-1196-9 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Kumar, N. P. S. K. S. Retinal disease prediction through blood vessel segmentation and classification using ensemble-based deep learning approaches. Neural Comput. Appl.35, 2495–12511. 10.1007/s00521-023-08402-6 (2023). [Google Scholar]
- 10.Abràmoff, M. D., Lavin, P. T., Birch, M., Shah, N. & Folk, J. C. Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices. npj Digit. Med.1(1), 39. (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Sorrentino, F. S. et al. Novel approaches for early detection of retinal diseases using artificial intelligence. J. Pers. Med.14(7), 690 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Salehi, S. K. A. W. et al. A study of CNN and transfer learning in medical imaging: Advantages, challenges, future scope. Sustainability10.3390/su15075930 (2023). [Google Scholar]
- 13.Daniella Convolutional Neural Network: operation, advantages and applications in AI. https://en.innovatiana.com/post/convolutional-neural-network (accessed.
- 14.Ejaz, S. et al. A deep learning framework for the early detection of multi-retinal diseases. PLoS One.19(7), e0307317 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Thanki, R. A deep neural network and machine learning approach for retinal fundus image classification. Healthc. Anal.10.1016/j.health.2023.100140 (2023). [Google Scholar]
- 16.Jeevani, M. Y. M., Vaishnavi, K., Bhavitha, P. T. K. & Harini, N. Prediction of retinal diseases using convolutional neural networks. Int. J. Res. Appl. Sci. Eng. Technol. (IJRASET). 10 (7). 10.22214/ijraset (2022).
- 17.Subramaniam, A. N. K. Enhancing retinal fundus image classification through active gradient deep convolutional neural network and red spider optimization. Neural Comput. Appl.10.1007/s00521-024-09989-0 (2024). [Google Scholar]
- 18.Fatema Tuj, M. B. M., Faria, J., Debnath, P., Fahim, A. I. & Shah, F. M. Explainable Convolutional Neural Networks for Retinal Fundus Classification and Cutting-Edge Segmentation Models for Retinal Blood Vessels from Fundus Images. Electr. Eng. Syst. Sci.1210.48550/arXiv.2405.07338 (2024).
- 19.Prashant, B. G. B. et al. Ensemble of deep convolutional neural networks is more accurate and reliable than board-certified ophthalmologists at detecting multiple diseases in retinal fundus photographs. Open. Access. Clin. Sci. 417–423. 10.1136/bjo-2022-322183 (2023). [DOI] [PMC free article] [PubMed]
- 20.Ling-Ping, J. J. et al. Yuqiang Huang, Tsz Kin Ng, Haoyu Chen, Weiqi Chen, Chi Pui Pang & Mingzhi Zhang, Automatic detection of 39 fundus diseases and conditions in retinal photographs using deep neural networks. Nat. Commun.10.1038/s41467-021-2513 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Mohamed Akil, Y. E. & Kachouri, R. Detection of retinal abnormalities in fundus image using CNN deep learning networks, (2020).
- 22.Balla Goutam, M. F. H., Geem, Z. O. N. G. W. O. O. & Bokde, N. E. E. R. A. J. D. H. A. N. R. A. A comprehensive review of deep learning strategies in retinal disease diagnosis using fundus images. IEEE Access10.1109/ACCESS.2022.3178372 (2022). [Google Scholar]
- 23.Kourounis, A. A. E. G., Thomson, B., Hunter, J., Ugail, H. & Wilson, C. Computer image analysis with artificial intelligence: A practical introduction to convolutional neural networks for medical professionals. Postgrad. Med. J.99(1178), 1287–1294. 10.1093/postmj/qgad095 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Chavan, D. P. R. Automatic multi‑disease classification on retinal images using multilevel glowworm swarm convolutional neural network. Journal of Engineering and Applied Science71, 26. 10.1186/s44147-023-00335-0 (2024). [Google Scholar]
- 25.Ahmed Aizaldeen, A. A., Abdullah, Hanaa, M. & Al Abboodi Review of Eye Diseases Detection and Classification Using Deep Learning Techniques, BIO Web of Conferences 9, (2024). 10.1051/bioconf/20249700012
- 26.Mohamed Elkholy, M. A. M. Deep learning-based classification of eye diseases using convolutional neural network for OCT images. Front. Comput. Sci.10.3389/fcomp.2023.1252295 (2024). [Google Scholar]
- 27.Steyn, E. C. Sensitivity and Specificity of Funduscopy in Determining the Need for Brain Computed Tomography in Patients Presenting With New-Onset Headache (2020)., University of the Witwatersrand, Johannesburg (South Africa).
- 28.Muchuchuti, S. V. S. Retinal disease detection using deep learning techniques: A comprehensive review. J. Imaging.10.3390/jimaging9040084 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Nguyen, K. J. T. D. et al. Retinal disease early detection using deep learning on ultra-wide-field fundus images. medRxiv10.1101/2023.03.09.23287058 (2023).38168353 [Google Scholar]
- 30.Veena, H. N., Muruganandham, A. & Kumaran, T. S. A novel optic disc and optic cup segmentation technique to diagnose glaucoma using deep learning convolutional neural network over retinal fundus images. J. King Saud Univ.10.1016/j.jksuci.2021.02.003 (2022). [Google Scholar]
- 31.Mongia, R. J. A. & Hussain, S. A. I. Eye disease detection and classification from retinal images using convolutional neural networks. Int. J. Software & Hardware Res. Eng. (IJSHRE)10(10), 106–118 (2022). [Google Scholar]
- 32.Ahsan Bin, I. U. et al. Rahim Khan, Kalimullah, Md. Sadek Ali, Diagnosis of diabetic retinopathy through retinal fundus images and 3D convolutional neural networks with limited number of samples, Wireless Communications and Mobile Computing, 2021, (2021). 10.1155/2021/6013448
- 33.SANTHOSH KUMAR M, S. M. & Retinal image processing using neural network with deep learning, (2022).
- 34.Sara, R. B., Ejaz, Z., AshrafI, M. M. & Alnfiai Mona Mohammed Alnahari, Reemiah Muneer Alotaibi, A deep learning framework for the early detection of multi-retinal diseases. OPEN. ACCESS.10.1371/journal.pone.0307317 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Ilesanmi, T. I. A. E. & Gbotoso, G. A. A systematic review of retinal fundus image segmentation and classification methods using convolutional neural networks. Healthc. Anal.10.1016/j.health.2023.100261 (2023). [Google Scholar]
- 36.Mehadi Hasan, T. S., Hasib, C. & Chowdhury Efficient Image Processing and Machine Learning Approach for Predicting Retinal Diseases, (2020).
- 37.Elangovan, P. & Nath, M. K. En-ConvNet: A novel approach for glaucoma detection from color fundus images using an ensemble of deep convolutional neural networks. Int. J. Imaging Syst. Technol.32(6), 2034–2048. 10.1002/ima.22761 (2022). [Google Scholar]
- 38.Elangovan, P., Vijayalakshmi, D. & Nath, M. K. COVID-19Net: An effective and robust approach for COVID-19 detection using an ensemble of ConvNet-24 and customized pre-trained models. Circuits Syst. Signal Process.43, 2385–2408. 10.1007/s00034-023-02564-3 (2023). [Google Scholar]
- 39.Elangovan, P. & Ramanathan, N. A novel approach for meat quality assessment using an ensemble of compact convolutional neural networks. Multimed. Tools Appl.82, 21563–21582. 10.1007/s11042-022-12852-y (2023). [Google Scholar]
- 40.Khan, M. A. et al. Diabetic retinopathy detection and analysis using convolutional neural networks and vision transformers. Comput. Mater. Contin.75(1), 1883–1903. 10.32604/cmc.2023.036041 (2023). [Google Scholar]
- 41.Kumar, A. et al. An Advanced Deep Learning Approach Combining Image Analysis for Precise Retinal Disease Detection, in International Conference on Artificial Intelligence for Innovations in Healthcare Industries (ICAIIHI), Raipur, India. (2023). 10.1109/ICAIIHI57871.2023.10489139
- 42.Kamber, A. N., Alkaabi, H., Al-Rekabi, M. & Jasim, A. K. Explainable AI for medical imaging: A taxonomy based on clinical task requirements. PERFECT: Journal of Smart Algorithms2(2), 72–77. 10.62671/perfect.v2i2.115 (2025). [Google Scholar]
- 43.Imran, A. et al. Fundus image-based cataract classification using a hybrid convolutional and recurrent neural network. Vis. Comput.37(8), 2407–2417 (2021). [Google Scholar]
- 44.Bilal, A. et al. DeepSVDNet: A deep learning-based approach for detecting and classifying vision-threatening diabetic retinopathy in retinal fundus images. Comput. Syst. Sci. Eng.48(2), 511–528 (2024). [Google Scholar]
- 45.Ikram, A. & Imran, A. ResViT FusionNet Model: An explainable AI-driven approach for automated grading of diabetic retinopathy in retinal images. Comput. Biol. Med.186, 109656 (2025). [DOI] [PubMed] [Google Scholar]
- 46.Bilal, A. et al. Improved support vector machine based on CNN-SVD for vision-threatening diabetic retinopathy detection and classification. PLoS One19(1), e0295951 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar] [Retracted]
- 47.Ikram, A. & Imran, A. A systematic review on fundus image-based diabetic retinopathy detection and grading: Current status and future directions. IEEE Access10.1109/access.2024.3427394 (2024). [Google Scholar]
- 48.Abbas, A. et al. A transfer learning based detection and grading of cataract using fundus images, in Proc. 25th Int. Multitopic Conf. (INMIC). 1–6. (2023).
- 49.Khan, A. Q. et al. A novel fusion of genetic grey wolf optimization and kernel extreme learning machines for precise diabetic eye disease classification. PLoS ONE, 19, 5, (2024).e0303094. [DOI] [PMC free article] [PubMed]
- 50.Bilal, A. et al. Breast cancer diagnosis using support vector machine optimized by improved quantum-inspired grey wolf optimization, Sci. Rep., 14, no. 1, Art. no. 10714, (2024). [DOI] [PMC free article] [PubMed]
- 51.Hassan, A., Imran, A., Yasin, A. U., Waqas, M. A. & Fazal, R. A multimodal approach for Alzheimer’s disease detection and classification using deep learning. J. Comput. Biomed. Inform.6(2), 441–450 (2024). [Google Scholar]
- 52.Ikram, A. & Imran, A. Automated diabetic retinopathy detection with FastDRNet on OCT imaging, In: Proc. Int. Conf. Emerging Technol. Electron., Comput., Commun. (ICETECC). 1–6. (2025).
- 53.Rafay, A., Asghar, Z., Manzoor, H. & Hussain, W. EyeCNN: Exploring the potential of convolutional neural networks for identification of multiple eye diseases through retinal imagery. Int. Ophthalmol.43(10), 3569–3586 (2023). [DOI] [PubMed] [Google Scholar]
- 54.Faria, F. T. J., Moin, M. B., Debnath, P., Fahim, A. I. & Shah, F. M. Explainable convolutional neural networks for retinal fundus classification and cutting-edge segmentation models for retinal blood vessels from fundus images, Preprint at https://arXiv.org/ abs/2405.07338, (2024).
- 55.Jain, S. & Salau, A. O. Detection of glaucoma using two dimensional tensor empirical wavelet transform. SN Appl. Sci.1(11), 1417. 10.1007/s42452-019-1467-3 (2019). [Google Scholar]
- 56.Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J. & Wojna, Z. Rethinking the inception architecture for computer vision, In:Proceedings of the IEEE conference on computer vision and pattern recognition, 2818–2826. (2016).
- 57.Sandler, M., Howard, A., Zhu, M., Zhmoginov, A. & Chen, L. C. Mobilenetv2: Inverted residuals and linear bottlenecks, In: Proceedings of the IEEE conference on computer vision and pattern recognition, 4510–4520. (2018).
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The datasets generated and/or analysed during the current study are publicly available in the Kaggle repository, https://www.kaggle.com/datasets/esubalewasmare/retinal-disease-fundus-images-datasets.To ensure full reproducibility, the complete source code, model configurations, training logs, and pre-trained weights have been made publicly available on Google Drive at: https://drive.google.com/file/d/1EyxpomtXip2Q04Dq7w-2pYgw0_QNgWL5/view? usp=sharing.



















