Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 May 19;16:16426. doi: 10.1038/s41598-026-52418-6

A gated task-attentive multi-task network for unified retinal image analysis

Muhammad Zaheer Sajid 1,✉, Imran Qureshi 2, Muhammad Fareed Hamid 3, Mohammad Alhefdi 4, Shrooq Alsenan 5, Qaisar Abbas 2, Yongwon Cho 6,✉
PMCID: PMC13216608  PMID: 42156815

Abstract

Diabetic retinopathy (DR) is one of the major causes of preventable blindness in the world, and accurate large-scale screening tools are needed urgently. Most of the deep learning methods which have been developed for retinal image analysis are treating tasks like optic disc segmentation and DR grading separately. This separation is making it difficult for the model to use the shared anatomical and contextual cues which are linking the two tasks. So we are proposing GTAM-Net, a Gated Task-Attentive Multi-Task Network for retinal image analysis. GTAM-Net is performing optic disc segmentation and DR severity grading together inside a single end-to-end network. Inside the network, a gated task-attentive block is deciding how the features should be shared between the two tasks at each layer. In this way the network is keeping the useful complementary information for each task, and at the same time it is avoiding the negative transfer which often hurts multi-task models. We are also using a multi-scale feature pyramid for keeping the hierarchical context, and an uncertainty-based loss weighting so that one task is not dominating the training. The proposed method is tested on five public datasets: IDRiD, DDR, Messidor-2, APTOS, and REFUGE. The model is reaching up to 98.17% Dice score for optic disc segmentation and 99.12% accuracy for DR grading, and the performance of the proposed method is competitive on every dataset that we tried. The cross-dataset tests are also showing that the model is fairly stable when the imaging conditions are changing. From these results, the proposed multi-task design is appearing to be a useful and reasonably stable option for joint retinal image analysis, and it can be considered for use in large screening pipelines.

Keywords: Gated attention network, Multi-task learning, Diabetic retinopathy grading, Optic disc segmentation, Retinal image analysis

Subject terms: Computational biology and bioinformatics, Diseases, Engineering, Health care, Mathematics and computing

Introduction

Diabetic Retinopathy (DR) is one of the main causes of acquired blindness in adults, and it now affects hundreds of millions of people who live with diabetes1. Long-term high blood sugar slowly damages the retinal blood vessels, the vision starts to drop, and if nothing is done the patient can go blind. Because of this, regular eye screening is very important. If DR is caught early, the disease can be slowed down a lot, and many cases of blindness can be avoided2. The problem is that doing this manually takes a lot of time, it needs trained doctors, and even then the decisions are not always consistent. In places where resources are limited, the situation becomes worse, and disagreement between doctors slows down treatment. Deep learning has changed the way medical images are analysed in the last few years. Today an automatic DR system can already pick up the typical signs like microaneurysms, hemorrhages, exudates, and cotton-wool spots, and it can also assign a severity score from non-proliferative DR (mild, moderate, severe) up to proliferative DR (moderate, severe)3–6. Apart from grading, segmenting retinal structures is also important. The optic disc in particular is a key landmark, and getting its boundary right helps in detecting glaucoma and other optic-nerve problems7. Most of the early retinal image analysis methods were designed for one task at a time, for example DR grading, lesion detection, or optic disc segmentation, and they were trained as separate models. These single-task models can do their job, but they do not share information across tasks, and running several of them together is expensive and not really practical for routine clinic use. They also miss the general features that come from looking at multiple tasks together. Multi-task learning tries to solve this by training related tasks at the same time so that they can help each other and produce better and more general features8. Recent work has shown that multi-task networks can learn related tasks together quite well9. But for retinal image analysis it is still not easy to balance the shared and the task-specific representations, and inter-task interference is a real problem. It is also not obvious how to give more or less computation to a task depending on how hard it is10. The existing systems are not making proper use of the anatomical and contextual relations between different retinal tasks. Most of them are either solving each task in isolation, or they are sharing features in a fixed way without any control. Due to this, the accuracy is dropping, the tasks are interfering with each other, and the model is also not generalizing well across datasets. To handle these issues, we are proposing GTAM-Net (Gated Task-Attentive Multi-Task Network). GTAM-Net is a unified network which is performing optic disc segmentation, retinal feature extraction, and DR grading together inside a single architecture. The main idea of this paper is the task-aware feature routing for retinal image analysis. The earlier multi-task methods are sharing parameters in a fixed way or using a static attention block. But in our case, the routing is input-dependent, and it is also changing during the processing. The shared features and the task-specific features are getting adjusted by learnable gates. So the information which is flowing between the tasks can be directly controlled. Because of this, the inter-task interference is reduced, and at the same time the complementary information for each task is also preserved. The gated task-attentive block is deciding which information should pass between the two tasks. It is pushing the network to focus on the features which are useful for each task, and it is suppressing the rest. This is one of the main problems in multi-task learning. The model has to find a balance between shared and task-specific features. So each task is getting help from the other task without being disturbed by it.

Research motivation

DR is one of the major reasons people are losing their sight, and finding it early is critical for stopping the permanent vision damage. The trouble is that running large-scale screening programs is not easy. The cost is high, eye specialists are not many, and even the diagnostic criteria are not always applied in the same way. Deep learning has done well on retinal images, but most of the methods are only solving one task. They are either classifying the disease or doing the anatomical segmentation. In real practice, the doctor is looking at many things at the same time. He is checking the optic disc, the vessels, and also the location where the lesions are. Many deep learning systems do not work in this way. They are treating these things as separate problems. Multi-task learning can in principle bring them together. But it is suffering from task interference, and that is lowering the performance of every task. This becomes more visible in retinal images, where the changes are very small and the model has to keep both general and task-specific features at once. The aim of this paper is to build a single framework which is analysing retinal anatomy and disease signs together without losing accuracy on either side. We design an end-to-end network with a gated task-attentive block that is deciding what to share between the tasks. The segmentation and the classification are running inside one model. Due to this, the system is becoming more robust, easier to scale, and also more useful in the practice.

Research contributions

The key contributions of this research work are as follows:

  1. We are proposing GTAM-Net. It is a unified multi-task framework for retinal image analysis. The framework is performing optic disc segmentation and diabetic retinopathy grading inside a single end-to-end architecture.

  2. We are presenting a task-aware feature routing design for retinal multi-task learning. Here, the inter-task feature interaction is treated as a dynamic and input-dependent process, instead of simple static parameter sharing.

  3. We are developing a gated task-attentive module. This module is using pooled contextual descriptors and learnable task-specific gates. The gates are regulating the shared and task-specific feature flow across segmentation and classification branches. So the task interference is getting reduced, and at the same time the complementary anatomical and disease-related information is preserved.

  4. We are also incorporating an uncertainty-aware multi-task optimization strategy. It is balancing segmentation and classification objectives during the joint training.

  5. The proposed framework is evaluated on multiple public retinal benchmark datasets. The performance of the proposed method is showing strong generalization across different imaging conditions and task settings.

Literature review

Deep learning for diabetic retinopathy detection and grading

In the last few years a lot of deep learning work has been done for DR detection. For example, transfer-learning based multi-task networks are checking the quality of fundus images, finding the lesions, and grading the DR stage at the same time. The reported AUC for detecting DR across stages is going from 0.943 to 0.9725. The DeepDR system is reaching high sensitivity and accuracy across all DR stages5. These works are showing that DR screening at scale is realistic, and this is important because the demand for retinal exams is growing. With deeper architectures, the networks are now reaching better accuracy and they are also faster than the older methods2,11. Many models have been tested for this task, and CNNs are the most common choice because they are picking up the features from fundus images automatically. A number of deep learning systems have been built for extracting useful features from retinal fundus images and supporting the clinicians in DR diagnosis. Whether a DR system is working properly or not, this depends a lot on how well it is detecting and analysing the retinal lesions. EfficientNet has given strong results when used for detecting all stages of DR6. Transfer learning is also useful here. Inception V3 and V4 are commonly applied to push the screening accuracy upwards12. In one of the studies, the task was extended from binary DR detection to multi-class severity grading. Deep learning is grading DR with more than 95% accuracy, and at the same time it is also doing adaptive image enhancement13. A system which is doing lesion detection, image quality check, and severity classification together has reached high test accuracy with fast inference, even on big datasets14. Bayesian deep learning has also been used for handling the uncertainty, and 97.68% accuracy was reported with Monte Carlo dropout3. These methods are also giving a confidence score along with the diagnosis, which the clinicians find useful. This is solving one limitation of the earlier systems, which were giving only yes/no predictions without any confidence measure. Some results are also suggesting that deep learning can predict the DR progression up to five years ahead, with concordance values from 0.754 to 0.8461. The hybrid models which are combining CNNs and RNNs are extending this idea further. They are using multiple retinal images which are taken at different visits, and so they can both diagnose the current DR state and also track its evolution over time15. Several recent and more advanced models have also been proposed for retinal disease analysis, and this is showing how fast the deep learning is moving in this area. Cheng et al. proposed WaveNet-SF. It is a hybrid framework for retinal disease detection. WaveNet-SF is mixing the spatial-domain and frequency-domain learning through wavelet decomposition, and it is reporting better OCT classification16. Qi et al. introduced MSLI-Net, which is combining multi-scale dilation fusion, lesion localization, and wavelet-based attention for OCT analysis17. Zuo et al. presented a multi-resolution visual Mamba framework. It is using a multi-directional selective mechanism, and it is capturing both local and global dependencies for OCT-based retinal disease detection18. These works are showing that the community is paying attention to the stronger architectures for retinal images. But most of them are focusing on OCT-based retinal disease classification. GTAM-Net is targeting a different and less explored problem, that is, the unified fundus-based multi-task learning. Here, the optic disc segmentation and DR grading are solved together inside a single end-to-end network. Some recent biomedical deep learning ideas which are using complex supervision and richer representations could also be useful for the future medical image systems. HRProtoKD applied a hierarchical relational prototype-based knowledge distillation for few-shot cancer molecular subtyping19. CNAScope built a pan-cancer copy number aberration database with functional annotation and interactive visualization, which is helping with large-scale computational biomedical analysis20. So this trend toward structured supervision, knowledge-based learning, and large-scale intelligent systems is now visible in this area. In a similar way, in the wider area of medical imaging, the integrated machine learning and deep learning approaches are starting to play a larger role in the early stages of the diagnosis21.

Optic disc segmentation and feature extraction

Segmenting the optic disc accurately matters a lot for many ophthalmic diagnostic tasks, since the optic disc is one of the main landmarks used to identify retinal disease. With pretrained U-Net architectures, deep learning models have already passed the 97% Dice score for optic disc segmentation22. The numbers suggest that current systems are close to clinicians when it comes to drawing the optic disc boundary. U-Net and its variants are now the default choice for medical image segmentation in general, and especially for the optic disc23,24. Most automated systems for optic disc detection are built around a U-Net backbone. The encoder-decoder design helps capture both fine local details and the wider structure that is needed to separate vessels and the optic disc. A multi-class semantic segmentation system based on U-Net has reported high pixel accuracy and IoU on the whole optic nerve head, including the disc, the cup, the blood vessels, and the peripapillary atrophy region25. The Multiresolution Cascaded Attention U-Net was proposed for objects that are irregular in shape, texture, and size, like the optic disc and the fovea7. Improved attention-based segmentation models reach more than 95% Dice score for the optic disc and around 88% for the optic cup26. An attention-based U-Net with residual connections has also been shown to give better segmentation for glaucoma detection27. Multi-scale attention works as well, and it is good at picking up complex features at different spatial resolutions28. Other ideas have been tried too. Distance-guided learning has been used to improve segmentation accuracy, and a method based on factorised gradient vector flow has been designed for optic disc segmentation, which works well for separating swollen and non-swollen discs29. These approaches model the spatial relations and the boundary details directly, which helps with boundary accuracy. Joint optic disc segmentation and disease classification has also been tried. A two-stage U-Net was used for sequential segmentation of the optic disc and the optic cup, and the pixel-wise AUC for disc segmentation went above 99%30. Enhanced EfficientNet-based U-Net systems have given high Dice scores for joint disc and cup segmentation in glaucoma screening31. Many recent works in medical image segmentation are also moving toward more advanced pseudo-label refinement and semi-supervised learning. ERSR is an ellipse-constrained pseudo-label refinement and symmetric regularization framework that targets semi-supervised fetal head segmentation in ultrasound32. Another work proposed an iterative pseudo-labeling and copy-paste learning scheme for semi-supervised tumor segmentation under limited annotations33. None of these methods deal with retinal fundus images, but they show that pseudo-label handling and annotation-efficient learning are becoming more and more important in medical image segmentation.

Multi-task learning in retinal image analysis

Multi-task learning is now used a lot in medical imaging, and training related tasks together inside one model is usually helping the accuracy34. Cross-task attention networks are trying to learn the link between tasks directly, and in this way they are performing better than the plain multi-task baselines10. The trick is that the complementary knowledge from one task is feeding into another. Due to this the features become richer, and the system is also getting more efficient and more clinically usable. Multi-task attention networks have been built for joint medical image segmentation and classification. These networks can do object classification and also produce good segmentation masks at the same time8. The hybrid systems like these are needing fewer resources, and they are also easier to fit into the clinical workflow. Multimodal AI systems which are bringing retinal images together with other types of data are giving richer and more accurate diagnostic information, since combining different sources is reducing the chance of mistakes. It has also been shown that multi-task learning can train disease classification and anatomical segmentation together, even when the labelled data is small35. Multi-task learning in medical imaging is also not without problems. Multi-scale feature enhancement has been applied in multi-task systems by collecting features at several resolutions through parallel branches with different dilation rates9. Older deep learning systems for medical imaging were usually focused on either segmentation or classification alone, so they could not benefit from the features which are common to several tasks. Task-specific attention is a useful idea which can be added to multi-task designs. More advanced encoders which are using ResFormer blocks are now used inside U-Net-based multi-task models. Here, the convolutions are handling the local details and the transformers are handling the longer-range dependencies9. With this kind of model, both global structure and small local details are getting captured, and both segmentation and classification are getting the benefit. A multi-task system for Parkinson’s disease, which is combining several types of patient information, also performed better than the methods which are using each type of information on its own36. So a similar approach is also likely to help retinal image analysis.

Attention mechanisms in retinal image segmentation

Attention mechanisms have changed the way deep models work, since they let the model focus on certain parts of the image. The Attention U-Net design has done well on retinal image segmentation, especially when the dataset is not very big. Attention pulls the network to the parts of the image that matter, in a way that is similar to how human vision works, and the available compute is then spent where it counts. For retinal vessels and other structures, several methods have been tested. Some use deformable convolutions and squeeze-and-excitation modules so the focal region can shift while features are extracted, and complex dependencies across different network levels can be picked up. With this kind of design, the variation in the size and shape of retinal features can be handled. Recently, gated attention blocks with dual-scale cross-attention have been used in 3D medical image segmentation, where attention works over both the spatial and the channel side and adapts to the current processing stage37. Mixing different attention types at different levels of the network gives a step-by-step refinement and improves both optic disc and cup segmentation, as shown in attention-based U-Nets that use dense dilated convolutions, with high IoU and Dice scores26. Multi-scale attention U-Nets have also been used for disc and cup segmentation, and continuous multi-scale feature extraction works very well28. Supervised attention methods that use decoder fusion modules, context squeeze-and-excitation, and supervised fusion have helped detect small blood vessels and have captured features at different scales. Recent review papers split attention methods into three groups—pre-Transformer, Transformer-based, and Mamba-related—and explain how each group works, how it is implemented, and where it has been used in medical image segmentation.

Gated mechanisms for task-specific learning

Gated mechanisms came first from recurrent neural networks, and now they are commonly used to control how information moves through deep architectures. Gating gives a learnable way to control feature flow, so the network can decide what to pass and what to block based on what it has learned. In multi-task learning gating is especially attractive. Conditional channel-gated networks for task-aware learning use data-dependent gates that pick only a small set of filters for each task, which saves compute and still gives high accuracy. This kind of selective gating fits well with multi-task designs in which separate tasks can flow through separate feature pathways. Pre-gating and contextual attention gate modules work at different stages of the pipeline. They filter out the cross-interactions that are not useful, and they reduce the ambiguity that cross-attention can cause. This two-level gating helps with one of the core problems in multimodal and multi-task learning, namely the spurious or irrelevant cross-connections that hurt the model. There is by now a sizeable body of work on why gating helps, and several papers on gated attention explain how it adds nonlinearity and sparsity. Both of these help to suppress attention sinks and very large activations that can break deep networks, and understanding this behaviour is important if we want gated systems to train and run in a stable way. Combining gating with mixture-of-experts is another important direction. Mixture-of-experts frameworks with gating networks have been built for medical imaging. The SAM-Med3D MoE system uses trainable gating networks to adapt large foundation models to specific medical image segmentation tasks38. Medical multimodal mixture-of-experts systems use separate experts for different data types, and the gating network during fine-tuning controls how much each expert contributes to the final prediction39. YOLO-Med is an efficient end-to-end multi-task network for biomedical image analysis that handles object detection and semantic segmentation, and it shares information between tasks with a cross-scale task-interaction module40. Multi-rater medical image segmentation methods also use gating networks with channel-wise attention, which weight meta-segments dynamically to capture the labelling style of different experts. Recent attention-based and multi-task frameworks have made medical image analysis stronger through cross-attention gated aggregation and shared feature learning. Dual Cross Attention (DCA) is making the skip connections more useful, and it is modelling the channel and spatial dependencies across multi-scale encoder features41. GA2Net is using hierarchical gated feature aggregation and mask-guided attention. It is improving the multi-scale representation learning for segmentation42. MTANet is pushing this idea toward joint segmentation and classification with a one-stage attention-based multi-task design43. The recent multi-task feature enhancement frameworks are looking at multi-scale contextual fusion for joint lesion segmentation and classification44. The proposed GTAM is different from the above. It is performing the gating-based task-aware feature routing using pooled multi-scale contextual descriptors and learnable task-specific gates. So the helpful shared information is passed to each task, and the task-irrelevant features are getting suppressed. Therefore GTAM is not the same as the existing attention modules, because it is explicitly handling the balance between shared and task-specific representations inside a unified multi-task retinal image analysis framework.

Research gaps and opportunities

Even with these recent advances, retinal image analysis with deep learning is still facing some real problems. The first one is that multi-task learning is promising, but a lot of the existing work is suffering from inter-task interference and from imbalance between tasks. The current methods are still relying on simple parameter sharing or fixed attention, which is not flexible enough to adjust the shared and task-specific features depending on the input or the task. The second issue is that the joint optic disc segmentation and DR grading has not been studied much, even though the two tasks are clearly connected. The optic disc is an important landmark for finding the disease-related lesions, but only a few works are presenting a unified end-to-end framework which is using this connection. The third issue is that attention is widely used in single-task retinal segmentation, but very few studies are looking at the multi-task case with explicit task-aware gating for controlling how the information is moving between tasks. The fourth issue is that there is still no efficient framework which is bringing several clinically useful tasks together without raising the computational cost too much. For real screening pipelines we are wanting both accuracy and efficiency at the same time. In recent retinal image analysis work, the trend is now shifting toward foundation-model-based and expert-specific learning. Large-scale retinal foundation models are learning general retinal representations from large collections of retinal images and image-text data, and they are also performing well on many downstream ophthalmology tasks45. The multimodal generalist and vision-language retinal models are more recent. They are combining the fundus images with text or other clinical data46–48. In parallel, the expert-specific aggregation frameworks are being studied to make the models more robust and adaptive under different ophthalmic imaging conditions49. So the direction toward large-scale pretraining and modular expert reasoning for retinal AI is now visible. But the methods listed above are mostly focusing on representation learning, multimodal understanding, and general ophthalmic adaptation. None of them is built specifically for combining end-to-end anatomical segmentation and DR grading inside a single fundus-based model.

Most of the existing work is dealing with either detection or segmentation, and only a few systems are trying to do both in an interpretable way. Even when the multi-task learning is used, most of the models are sharing parameters in a naive way or they are using fixed attention. They are not really respecting the differences between tasks. They are suffering from task interference, and they are not using the features properly. Most of the attention-based models are only looking at the spatial or channel dimension. They are not modelling the relationship and the information flow between tasks. So they cannot give the best result for both shared and task-specific features. To get past these limits, we are adding a gated task-attentive block in GTAM-Net which is controlling the interaction of shared and task-specific features in a selective and dynamic way, conditioned on the input. With this, the segmentation and the classification can really work together. Through the learnable gates and a mix of shared and task-specific paths, we get a system which is handling both tasks for retinal images without losing efficiency.

Methodology

In this work, a new deep learning model is proposed which we are calling GTAM-Net. It is a multi-task model for retinal fundus image analysis, and it is performing two tasks together. The first task is optic disc segmentation. The second task is DR grading, with five severity levels (Normal, Mild, Moderate, Severe, and Proliferative DR). At the centre of GTAM-Net is a Gated Task-Attentive Module (GTAM). This module is controlling the shared and task-specific features separately. A cross-task attention block is letting the segmentation and the classification share useful information. Multi-scale anatomical features are also extracted to support better grading. The full pipeline is shown in Fig. 1. During training, an uncertainty-based loss function is used so that both tasks are staying balanced automatically. A feature pyramid network with dilated spatial attention is also used for multi-scale fusion. Due to the adaptive task weighting and the layered fusion, the localisation and the diagnostic results are improved. We are seeing this work as a step forward in medical image analysis, since it is bringing together attention and multi-task learning. GTAM-Net is combining the strong feature extraction of CNNs with gated attention. Hence a single model is handling both segmentation and classification, and it is doing this better than two separate task-specific models. One advantage is that GTAM-Net is reaching expert-level diagnostic accuracy and it is still practical for real clinical use. Another advantage is interpretability. The attention maps are clearly showing which image regions are used for segmentation and grading, and this can build trust in the model. The shared features are carrying anatomical information which is helping the disease classification, and the disease-related features in turn are helping the anatomical segmentation. We test the framework on several public datasets, and the performance of the proposed method is showing strong results on both tasks. So this work also gives a useful benchmark for retinal image analysis.

Fig. 1.

Fig. 1

The architecture of the proposed GTAM-Net framework. The network is taking retinal fundus images as input. It is producing both segmentation masks and DR severity grades using a shared encoder with task-specific decoders, and these decoders are interconnected via gated attention modules.

Proposed architecture

GTAM-Net is built as an end-to-end trainable model which is performing anatomical segmentation and disease grading at the same time. We are using multi-task learning so that the two tasks can share the useful information, and at the same time each task is also getting a specific refinement. The retinal fundus image is going into the model as input, and the model is detecting the optic disc and grading the DR severity together. Both tasks are helping each other in this way. The anatomical features are pushing the disease classification to be more accurate, and the disease-related features are pushing the segmentation to be cleaner. The whole system is built from several connected modules, and these modules are extracting, refining, and combining the multi-scale features. With this kind of structure, GTAM-Net is giving strong results for clinical segmentation and grading.

Overall framework architecture

GTAM-Net is following a multi-task learning approach which is combining optic disc segmentation and DR grading inside one model, as shown in Fig. 2. The pipeline is going through hierarchical feature extraction and task-specific refinement, with five main stages: (1) multi-scale feature extraction with a ConvNeXtTiny backbone, (2) a gated attention block which is separating the useful features, (3) task-specific decoder networks with cross-task attention, (4) multi-scale feature fusion, and (5) joint optimisation with adaptive loss weighting.

Fig. 2.

Fig. 2

Detailed architecture of GTAM-Net showing the gated attention modules and cross-task connections between segmentation and classification pathways.

Mathematical formulation

Let Inline graphic represent an input retinal fundus image. The dual-task objective function is formulated as:

graphic file with name d33e613.gif 1

where Inline graphic, Inline graphic, and Inline graphic denote segmentation, classification, and regularization losses respectively, with adaptive weights Inline graphic, Inline graphic, and Inline graphic determined by our task uncertainty weighting mechanism.

Gated attention mechanism

The core innovation of GTAM-Net is the Gated Task-Attentive Module (GTAM), which selectively propagates task-relevant features while suppressing irrelevant information. The gating operation is mathematically defined as:

graphic file with name d33e647.gif 2

where Inline graphic represents input features, Inline graphic denotes learnable gating weights, Inline graphic is the sigmoid activation, and Inline graphic indicates element-wise multiplication. The concatenated average and max pooling operations provide multi-scale context for dynamic gate computation.

The gating operation in Eq. (2) is kept light. It is using pooled context descriptors together with learnable modulation weights. This formulation is conceptually similar to the existing channel-attention and gating mechanisms, and we are not claiming the gate operator alone as the contribution. The role of the gate is to act as a task-aware routing block inside a multi-task retinal framework. Here the gated features are separated and pushed to the segmentation and classification paths, so the inter-task information flow can be controlled.

For the cross-task feature modulation, we are implementing a Cross-Task Attention (CTA) mechanism. It is based on the standard scaled dot-product attention formulation. Here, the queries are taken from one task branch, and the keys and values are taken from the other task branch:

graphic file with name d33e675.gif 3

where Inline graphic, Inline graphic, Inline graphic are query, key, and value projections for segmentation and classification tasks, respectively, with d as the feature dimension. Eq. (3) is following the standard dot-product cross-attention, and we are not claiming it as a new attention primitive on its own. The contribution of GTAM-Net is in the way this is used inside a unified retinal multi-task framework. Here, the segmentation and the classification features are conditioned on each other through the cross-task attention. So in this role, the CTA module is acting as a structured information-exchange block, and one task can pick up the complementary features from the other one.

Multi-scale feature pyramid network

The encoder backbone generates hierarchical feature maps Inline graphic corresponding to stride levels Inline graphic relative to input resolution. For capturing the multi-scale contextual information, we are adopting an FPN-style hierarchical fusion strategy to construct multi-scale representations:

graphic file with name d33e711.gif 4

where Inline graphic denotes the pyramid feature at level i, initialized with Inline graphic. Eq. (4) is following the standard Feature Pyramid Network formulation, and here it is used to provide consistent multi-resolution feature representations for both segmentation and classification within the unified framework. Each pyramid level is enhanced with a dilated convolution-based spatial attention module, which is referred to here as DSA:

graphic file with name d33e731.gif 5

where Inline graphic is a dilated convolution with rate d, and Inline graphic is channel-wise multiplication. Eq. (5) is using dilated convolutions with sigmoid gating to highlight the salient spatial regions. The structure is conceptually close to the existing spatial attention methods, and here we are using it as a simple context-aggregation block inside the multi-task framework. We are not presenting it as a new standalone attention primitive in GTAM-Net.

Task-specific decoder networks

Segmentation decoder

The segmentation decoder employs a progressive upsampling strategy with skip connections from corresponding encoder layers. The segmentation output Inline graphic (optic disc vs. background) is computed as:

graphic file with name d33e761.gif 6

where Inline graphic denotes bilinear upsampling to input resolution, and Inline graphic represents channel-wise concatenation.

Classification decoder

For DR severity grading (5 classes: No DR, Mild, Moderate, Severe, Proliferative), we implement a multi-head attention pooling mechanism:

graphic file with name d33e778.gif 7

where Inline graphic denotes global average pooling, Inline graphic is the mean pooled feature, and Inline graphic are learnable attention weights. The final classification probability is:

graphic file with name d33e795.gif 8

Adaptive loss weighting

Following Kendall et al50., we are using uncertainty-based loss weighting which is automatically balancing the task contributions during the training:

graphic file with name d33e807.gif 9

where Inline graphic represents the task-dependent uncertainty parameter learned during training. This formulation allows the network to down-weight uncertain tasks while emphasizing confident predictions.

Implementation details

The network was implemented in PyTorch 1.12 and trained on 4× NVIDIA A100 GPUs. We used the AdamW optimizer with initial learning rate Inline graphic, weight decay Inline graphic, and cosine annealing scheduler. The training was running for 300 epochs with batch size 32. The input images were resized to Inline graphic pixels, and the data augmentation was applied including random rotation (Inline graphic), horizontal/vertical flipping, color jitter (brightness=0.2, contrast=0.2), and random affine transformations.

The segmentation loss combines Dice coefficient (for region overlap) and binary cross-entropy (for pixel-wise accuracy):

graphic file with name d33e838.gif 10

where Inline graphic and Inline graphic denote predicted probability and ground truth for pixel i, respectively, and Inline graphic prevents division by zero.

For the classification, we are using label-smoothed cross-entropy loss (smoothing factor Inline graphic). It is helping the generalization and reducing the overconfident predictions:

graphic file with name d33e864.gif 11

where Inline graphic and Inline graphic (number of DR severity classes).

The complete training procedure for GTAM-Net is formalized in Algorithm 1.

Algorithm 1.

Algorithm 1

GTAM-Net Training Procedure

Theoretical insights on task-aware gating and loss balancing

In multi-task learning, the gradients from different tasks can conflict, and they are hurting the optimisation. Let Inline graphic and Inline graphic denote the segmentation and classification losses, and let Inline graphic denote the shared parameters. The standard multi-task setup is optimising:

graphic file with name d33e915.gif 12

When the gradients Inline graphic and Inline graphic are pointing in different directions, destructive interference can show up. The training is then becoming unstable, and the optimiser is ending up in a suboptimal solution.

The proposed gated task-attentive block is adding a dynamic modulation of the shared features:

graphic file with name d33e931.gif 13

where g(x) is a learnable gating function which is depending on the input x, and Inline graphic is element-wise multiplication. With this formulation, the model is routing the features which are useful for each task, and the gradient conflict is also reduced since the irrelevant features are not propagated across tasks.

The uncertainty-aware loss weighting can also be written as:

graphic file with name d33e951.gif 14

where Inline graphic and Inline graphic are the per-task uncertainties. The formulation is scaling the task contributions adaptively during the training. The tasks with higher uncertainty are getting lower weights, so the model is focusing on the tasks which are easier to learn first.

This behaviour can be interpreted as an implicit curriculum learning process, where the importance of each task is automatically getting adjusted based on how hard it is to learn. As the training is going on, the model is balancing both tasks, and the convergence and the generalisation are also getting better.

Experimental framework and dataset characteristics

Dataset description and composition

We tested GTAM-Net on five widely used public benchmark datasets for DR detection: IDRiD, DDR, Messidor-2, APTOS 2019, and REFUGE. The datasets are different in image acquisition, patient demographics, labelling style, and also the clinical information that is provided. So they are letting us check the method on a varied set of images. Each one is having its own challenge, and it is testing a different aspect of the model, like high-resolution images, fine-grained labels, class imbalance, and also robustness under different image conditions. The details are listed in Table 1. IDRiD is having high-resolution fundus images, pixel-level ground-truth optic disc masks, and 5-level DR grading. So it is suitable for both segmentation and classification. DDR is much larger. It is having images from many cameras and several patient cohorts, and that makes it a good test for generalisation across sources and groups. APTOS 2019 is having a strong class imbalance, and that is also a useful test case in itself. REFUGE is providing accurate optic disc and optic cup annotations for glaucoma screening, so we can study segmentation independently from the DR grading. For DDR, Messidor-2, and APTOS 2019, the expert-drawn optic disc masks are not available. So we generated optic disc masks using an expert-refined pseudo-labelling step. We started from a pretrained U-Net, and after that an ophthalmologist reviewed and corrected the masks. This gave us a practical set of completed labels. To get the high-quality labels, the predicted masks went through a two-step quality-control process. In the first step, every predicted mask was inspected by hand. The mask was flagged as uncertain when the optic disc boundary was weak, when the contrast between the disc and the surrounding retina was low, when the disc was not fully covered, when the boundary was leaking into a nearby bright region, when the shape was not matching the anatomy, or when the mask was clearly disagreeing with the actual optic disc. In the second step, all the uncertain masks were corrected by an ophthalmologist with more than 10 years of experience, while the remaining masks were accepted as they were. For DDR, this process gave us corrected optic disc masks for 2,000 randomly chosen images which are covering all DR severity levels, and these masks were then used for training and evaluation. The annotations are coming from expert-refined pseudo-labels and not from a native ground truth. So the results on DDR, Messidor-2, and APTOS 2019 should be read as supporting evidence under a weak-label setting. The most rigorous segmentation evaluation is the one on IDRiD and REFUGE, since both of these datasets are having expert annotations. We also added some example images of the pseudo-mask refinement step before and after the ophthalmologist correction, in the revised paper or in the supplementary material.

Table 1.

Details of benchmark datasets used for evaluation.

Dataset Images Segmentation DR grades Resolution Train/Val/Test
IDRiD 516 Yes (OD, Lesions) 5-class 4288×2848 310/103/103
DDR 13,673 No 5-class Variable 8,204/2,731/2,738
Messidor-2 1,748 No 5-class 1440×960, 2240×1488 1,048/350/350
APTOS 2019 3,662 No 5-class Variable 2,197/733/732
REFUGE 1,200 Yes (OD, Cup) – 2124×2056 800/200/200

Data preprocessing and augmentation pipeline

All fundus images were preprocessed using a unified pipeline to maintain consistency across datasets and resized to Inline graphic pixels, preserving the original aspect ratio via bilinear interpolation, which retained anatomical structures and disease-related features as much as possible.

For the datasets which are not coming with expert optic disc segmentation masks (DDR, Messidor-2, and APTOS 2019), we added optic disc annotations as a practical label-completion step. First, we generated pseudo segmentation labels with a pretrained U-Net. To make the annotations more reliable and to reduce the noisy supervision, we applied a two-step quality control. In the first step, we visually screened all the generated masks. A mask was flagged as uncertain if it was having a weak optic disc boundary, low contrast between the disc and the surrounding retinal tissue, incomplete disc coverage, leakage of the boundary into nearby bright regions, an irregular shape which is not matching the visible anatomy, or any visible disagreement with the optic disc structure. In the second step, every uncertain case was reviewed and manually corrected by a certified ophthalmologist with more than 10 years of clinical experience. The remaining masks were only accepted after this verification, with no further changes. For DDR, this process gave us verified optic disc masks for 2,000 randomly selected images which are covering all DR severity levels. We used these as the optic disc annotations for the rest of the training and evaluation.

To reduce the risk of data leakage, the pseudo-label generation and the ophthalmologist verification were done independently from the final GTAM-Net training, and we used them only as a label-completion step for the datasets which are not having optic disc annotations. The annotations are still coming from expert-refined pseudo-labels rather than from the original ground truth, so a residual annotation bias is also possible. The DDR, Messidor-2, and APTOS 2019 segmentation results should not be taken as a strict benchmark against the original expert-drawn ground truth, but rather as supporting evidence in a practical weak-label setting. The strongest segmentation evidence in this paper is on IDRiD and REFUGE, where the expert annotations are available. A more rigorous future check on the datasets which are not having native optic disc masks will need a fully independent manual annotation done from scratch.

During the training, we used data augmentation so that the model becomes more robust to image conditions. The geometric augmentations included random rotations of Inline graphic, horizontal and vertical flipping with 50% probability, and random affine transforms with scaling between 0.9 and 1.1 and translation up to 10% of the image size. The colour augmentations included brightness and contrast jittering up to Inline graphic, gamma correction in the 0.8 to 1.2 range, and Gaussian noise with a standard deviation up to 0.01. These transforms are trying to imitate the realistic variations in lighting, camera settings, and patient positioning, and they are making the model more useful for clinical conditions.

Evaluation metrics and statistical analysis

For the segmentation and classification results, we are using several measures together. To quantify the region overlap, we compute the Dice Similarity Coefficient (DSC) and the Intersection over Union (IoU) between the predicted optic disc masks and the ground truth. The pixel-wise classification is summarised through sensitivity and specificity, and the boundary matching is checked with the 95% Hausdorff distance (HD95). Using a mix of measures is helpful here: DSC and IoU are looking at the region, sensitivity and specificity are describing the classification, and HD95 is capturing how close the boundary is. For DR grading into the five classes (No DR, Mild DR, Moderate DR, Severe DR, Proliferative DR), we measure accuracy, precision, recall, and F1 per class. Macro-averaged values are also reported for all measures, so the class imbalance is handled fairly. The Cohen’s kappa between predictions and ground truth is also reported.

For each DR severity level, we are reporting ROC curves with AUC values, together with precision-recall curves and average precision values, since AUC alone can be misleading when the data is imbalanced. The 95% confidence intervals are estimated with bootstrap resampling using 1,000 runs, which is making the numbers more stable. We are also measuring efficiency: the FLOPs (floating-point operations) as a theoretical indicator of cost, the inference time per image on an NVIDIA A100 GPU, the number of trainable parameters as a measure of model complexity, and also the inference memory. These numbers are mattering for clinical deployment, where the resources are not always available.

Implementation details and training configuration

Network architecture and hyperparameter settings

We have implemented GTAM-Net in PyTorch 1.12 with CUDA 11.6. For efficiency, the training is done in mixed precision. We used an HPC cluster with four NVIDIA A100 GPUs of 40 GB VRAM each, with data-parallel training and synchronised batch normalisation across the GPUs. The optimiser is AdamW, with an initial learning rate of Inline graphic, betas of (0.9, 0.999), an epsilon of Inline graphic, and a weight decay of Inline graphic to keep overfitting under control. The learning rate is following a cosine annealing schedule with a warm-up from epoch 0 to epoch 10, and it is going down to Inline graphic over 300 epochs. The training is run for 300 epochs with a per-GPU batch size of 32, which is giving an effective batch size of 128. The segmentation loss is a combination of the Dice coefficient (for region overlap) and binary cross-entropy (for pixel-wise classification), with equal weights. For the classification, we are using label smoothing with Inline graphic to help the generalisation and to avoid the overconfident predictions. The task uncertainties Inline graphic and Inline graphic are initialised to 1.0, and they are updated by backpropagation during the training.

Loss formulation and optimization strategy

The multi-task loss of GTAM-Net is incorporating per-task uncertainty estimation, and it is adaptively balancing the segmentation and classification objectives during the training. The total loss is:

graphic file with name d33e1134.gif 15

where Inline graphic is combining the Dice and cross-entropy losses for segmentation, Inline graphic is the classification loss with label smoothing, and Inline graphic and Inline graphic are learned uncertainty parameters. With this loss, the network is giving less weight to the task which is more uncertain, and it is paying more attention to the confident predictions. The balance between tasks is then automatically adjusted while the model is training. We used gradient accumulation over 4 steps to keep the training stable with a large effective batch size, and gradient clipping with a maximum norm of 1.0 was also applied. The model checkpoints are saved every 10 epochs as a safety net, and we used early stopping with a patience of 50 epochs based on the validation loss to avoid overfitting. Each experiment was repeated three times with different random seeds, and the reported values are the mean ± standard deviation.

Experiment 1: evaluation on IDRiD dataset

Dataset characteristics and experimental setup

IDRiD is one of the larger datasets used for DR studies. It is having 516 retinal fundus images at Inline graphic pixels. The optic disc is having pixel-level ground truth, and the images are also annotated for 5-level DR classification, with all the annotations done by retinal specialists. The dataset is covering a wide range of disease severity, image quality, and also illumination, which is making it a realistic and tough testbed for the automated systems. We used the standard split of 310 images for training, 103 for validation, and 103 for testing, and this is also letting us compare with the previously reported results in a fair way.

Segmentation performance analysis

From Table 2, GTAM-Net is giving the highest accuracy for optic disc segmentation. The Dice coefficient is 98.17%, which is 1.72% higher than the previous best (DenseUNet). The Hausdorff distance is also better, 4.12 pixels compared to 6.45 pixels. Boundary accuracy is mattering in clinical use, since the optic disc boundary is needed for the cup-to-disc ratio (used in glaucoma) and also for several other measurements. The model is having a sensitivity of 97.89% and a specificity of 99.78%, so it is performing well on both the disc and the background, with very few false positives and very few false negatives. The improvement is coming from a few different things at once. The gated attention block is suppressing the background and pushing the optic disc features through. The cross-task attention is using the classification features to refine the segmentation boundary, since the disease context is now available for the segmentation branch. The multi-scale feature pyramid is keeping the local texture and the wider anatomical structure together. Even on the harder cases like images with peripapillary atrophy, myelinated nerve fibers, or overlapping pathology, GTAM-Net is keeping the disc boundary accurate, which is where most of the older models start to break. The segmentation comparison is shown in Fig. 3.

Table 2.

Segmentation performance comparison on IDRiD test set.

Method DSC (%) IoU (%) Sens. Spec. HD95 (px)
U-Net23 94.23 ± 0.31 89.12 ± 0.42 0.9345 0.9912 8.45 ± 0.23
Att. U-Net26 95.67 ± 0.28 91.23 ± 0.38 0.9489 0.9934 7.23 ± 0.19
MCAU-Net [?] 96.45 ± 0.25 92.18 ± 0.35 0.9567 0.9945 6.89 ± 0.17
DenseUNet34 96.89 ± 0.24 93.12 ± 0.32 0.9612 0.9951 6.45 ± 0.15
GTAM-Net 98.17 ± 0.18 95.23 ± 0.25 0.9789 0.9978 4.12 ± 0.12

Fig. 3.

Fig. 3

Segmentation Analysis.

Classification performance analysis

The DR severity grading accuracy of GTAM-Net is 99.12% on the IDRiD test set, as shown in Table 3. This is 2.89% higher than MTANet, which was the previous best multi-task system. The high precision (0.989) and recall (0.990) are suggesting that the model is well balanced between accuracy and detection. The Cohen’s kappa is 0.987, which is meaning very strong agreement with the reference labels on the public test set. We should note that this is on a fixed public test set, with curated reference annotations and a controlled preprocessing pipeline. It is not meaning that DR grading is fully solved in clinical settings, where the reader variance is much greater, more borderline and unclear cases are showing up, and the labels themselves are more noisy. The confusion matrix in Fig. 6 is showing that most of the errors are happening between adjacent grades, like Mild and Moderate. These two are also hard to separate in clinical practice, and the difficulty here is probably reflecting the inter-reader variance among the expert clinicians. The model has done particularly well on proliferative DR, with a 100% score for PDR. The numbers are high under the conditions tested and they look promising. But they should be read with the understanding that the public datasets are often giving higher numbers than the clinical datasets. A reliable detection of PDR is mattering a lot for the timely diagnosis and treatment of DR patients.

Table 3.

Classification performance comparison on IDRiD test set.

Method Acc. (%) Prec. Rec. F1 Kappa
Inception-V312 92.34 ± 0.42 0.921 0.918 0.919 0.901
EffNet-B46 94.56 ± 0.38 0.943 0.941 0.942 0.928
MTANet8 96.23 ± 0.32 0.958 0.961 0.959 0.947
CT-Net10 97.45 ± 0.28 0.972 0.969 0.970 0.961
GTAM-Net 99.12 ± 0.15 0.989 0.990 0.989 0.987

Fig. 6.

Fig. 6

Confusion matrix for DR grading on IDRiD.

Training dynamics and convergence analysis

Figures 4 and 5 are showing how GTAM-Net is training. The segmentation DSC is converging fast in the early epochs. It is getting close to 95% around epoch 50, and going up to 98.17% by epoch 250. The classification accuracy is going up more slowly and stabilising later, which is suggesting that the segmentation features are also helping the classification task. The adaptive loss weight w is staying near 0.9, and the per-task weights Inline graphic and Inline graphic are also staying stable. The learned uncertainty parameters are settling to Inline graphic and Inline graphic, so the segmentation task is being predicted as less uncertain than the classification one. The validation and the training performance are staying close to each other (less than 1% gap for segmentation and about 0.5% for classification), which is meaning that the model is well regularised in this setup. The model is continuing to improve until epoch 300, so the multi-task optimisation is stable. Each run was repeated with different random seeds, and the reported values are mean ± standard deviation. The closeness of the training and validation curves is giving extra support to the reported gains, but the results should not be over-read for real-world settings outside the evaluated conditions (Fig. 6).

Fig. 4.

Fig. 4

Training and validation accuracy curves of the proposed model, illustrating performance convergence over epochs.

Fig. 5.

Fig. 5

Training and validation loss curves of the proposed model, demonstrating stable learning and convergence behavior.

Experiment 2: evaluation on DDR dataset

Dataset overview and experimental design

DDR is having 13,673 retinal fundus images collected from many clinics across different regions, and there are real differences in camera type, imaging procedures, and the patient population. This variation is what is making the dataset useful for testing generalisation. The images are also varying in quality, resolution, and lighting. We use the standard split of 8,204 images for training, 2,731 for validation, and 2,738 for testing, which is giving a reliable statistical evaluation.

DDR is not coming with segmentation labels, so we used a two-step process. First, the pseudo-segmentation masks are generated with a model trained on IDRiD. After this, we ran an active learning loop, where an ophthalmologist is reviewing and correcting the uncertain cases. This gave us high-quality segmentation masks for 2,000 randomly chosen images which are spanning all DR severity levels, and these were used to evaluate the segmentation performance. DDR is not having native optic disc annotations, so the reported segmentation results should be read against the ophthalmologist-verified pseudo-labels and not against the original dataset’s expert ground truth. The DDR segmentation analysis is therefore presented as supporting evidence under a practical weak-label setting, and the strongest segmentation evaluation in this paper is on the datasets which are having expert annotations, namely IDRiD and REFUGE.

Performance results and generalization analysis

From Table 4, GTAM-Net is doing well on DDR. The segmentation Dice score is 97.45% and the classification accuracy is 98.23%, which is 2.33% higher in Dice than MTANet. This is showing that the gated attention mechanism is still working on images with varying quality and camera artefacts. The AUC-ROC is 0.998, so the model is good at separating the different DR severity levels. It is also doing well on proliferative DR, with 99.2% sensitivity and 99.8% specificity, and the performance is staying consistent across image quality categories. It is working on good, usable, and even reject-quality images. Most of the errors are coming from the images with heavy artefacts (motion blur, uneven illumination) or from images with severe co-existing diseases (cataract, vitreous hemorrhage), where the DR-related structures are partly hidden. Even with these challenges, the model is still beating the baseline methods.

Table 4.

Performance evaluation on DDR test set.

Metric GTAM-Net EffNet MTANet Imp. (%)
Seg. DSC (%) 97.45 ± 0.22 – 95.12 ± 0.31 +2.33
Seg. IoU (%) 94.12 ± 0.28 – 91.45 ± 0.38 +2.67
Cls. Acc. (%) 98.23 ± 0.18 94.67 ± 0.35 96.45 ± 0.28 +1.78
Macro F1 0.981 ± 0.003 0.941 ± 0.007 0.962 ± 0.005 +1.9
AUC-ROC 0.998 ± 0.001 0.985 ± 0.003 0.991 ± 0.002 +0.7

Note: The segmentation metrics for DDR are reported against ophthalmologist-verified pseudo-labels, since the dataset is not providing the original expert-drawn optic disc ground truth annotations. So these segmentation results should be interpreted as supporting evidence in a weak-label setting, and not as a strict benchmarking against the native ground truth.

Experiment 3: evaluation on Messidor-2 dataset

Dataset characteristics and clinical relevance

Messidor-2 is a set of 1,748 retinal fundus images which are coming from a large diabetic retinopathy screening program in France. All the images were taken using routine clinical procedures, and they were labelled by expert ophthalmologists in line with the international clinical standards. The dataset is useful for testing the proposed system in a more realistic clinical setting, since the class and image distribution is closer to that of routine screening. We use the standard split of 1,048 images for training, 350 for validation, and 350 for testing.

Performance analysis and clinical implications

Figures 7 and 8 are showing the training curves, and they are telling us how GTAM-Net is learning on this dataset. The confusion matrix in Fig. 9 is showing that most of the classification errors are happening between adjacent DR severity classes. This pattern is matching the clinical experience, since the criteria which are separating these classes are very fine, and even the expert clinicians sometimes disagree. GTAM-Net is doing well on Messidor-2. The accuracy is 98.76% and the AUC-ROC is 0.997, as reported in Table 5. At a specificity of 0.95, the model is having a sensitivity of 0.968. High sensitivity is important in screening for avoiding the missed referable cases, and a reasonable specificity is also important so that the patients are not sent unnecessarily to a specialist. The model is especially strong on the referable DR (moderate or worse), with a sensitivity of 99.1% and a specificity of 98.7%, which is above the targets that are usually set for an automated screening tool. These good numbers should still be read with caution. They are coming from a curated benchmark dataset with static labels and a controlled evaluation setting, and they are not necessarily representing the realistic screening practice, where this level of performance is unlikely to be reached. So they should not be taken as evidence of near-perfect clinical performance, but rather as a sign that the framework is working well under the conditions of the evaluated benchmark.

Fig. 7.

Fig. 7

Training and validation accuracy curves of the proposed model, illustrating performance convergence over epochs.

Fig. 8.

Fig. 8

Training and validation loss curves of the proposed model, demonstrating stable learning and convergence behavior.

Fig. 9.

Fig. 9

Confusion matrix of Messidor-2.

Table 5.

Performance comparison on Messidor-2 test set.

Method Acc. (%) AUC-ROC AUC-PR Kappa Sens. @ 0.95
Inception-V412 94.23 ± 0.41 0.981 0.972 0.921 0.923
EffNet-B76 95.67 ± 0.36 0.987 0.978 0.938 0.938
DeepDR5 96.45 ± 0.32 0.989 0.983 0.951 0.945
MTANet8 97.12 ± 0.28 0.992 0.987 0.961 0.951
GTAM-Net 98.76 ± 0.19 0.997 0.992 0.983 0.968

The results on Messidor-2 are showing that GTAM-Net is working in realistic screening conditions. The dataset is having variations in field of view, illumination, and focus, among other things. The gated attention mechanism is playing an important role here, since it is picking the diagnostic regions that matter and is also pushing down the background noise and the artefacts. This is similar to how an experienced clinician is scanning a fundus image to make a decision.

Experiment 4: evaluation on APTOS 2019 dataset

Class imbalance challenges and evaluation strategy

APTOS 2019 is challenging because of the strong class imbalance: 87% of the images are showing no DR, and only 13% are spread across the various DR stages. This kind of distribution is similar to a real-world screening population, and it is making both the training and the evaluation harder. For classification, the raw accuracy can mislead us, so we are using suitable measures here. We are reporting precision-recall curves, F1-score, and balanced accuracy to make the evaluation fair. During the training, we are using class-weighted sampling and focal loss, so the model is learning the minority classes. APTOS 2019 is not having native optic disc annotations, so any segmentation analysis on APTOS 2019 is relative to the ophthalmologist-verified pseudo-labels and not the dataset’s ground truth. The segmentation evidence here should be treated as supporting evidence under a weak-label setting, and not as a direct comparison with ground truth.

Performance analysis and imbalance handling

The training behaviour of GTAM-Net on APTOS 2019 is shown in Figs. 10 and 11. The confusion matrix in Fig. 12 is telling us that many of the misclassifications are happening between adjacent classes. This is in line with the clinical experience, since these distinctions are subtle, and the interobserver variability among the expert clinicians is also high here. From Table 6, the model is giving balanced performance across all the classes despite the imbalance, with a macro-averaged F1 of 0.963. The precision is also reasonable for the minority classes Mild (0.978) and Proliferative (0.938), and the recall is also high. So the model is still learning the imbalanced classes, because the gated attention is adapting the features which it is paying attention to. The weighted-average values are giving an overall accuracy of 98.45%, which is much higher than the competing methods, which are leaning strongly toward the majority No DR class. The attention maps are suggesting that the model is giving more representational capacity to the minority-class examples, and the early DR cases are weighted more for the small lesions. This allocation is coming from the task-uncertainty weighting, and it is letting the model handle the class imbalance without explicit oversampling or class weighting. The model is also useful for early detection. A recall of 97.2% on Mild NPDR is important for early intervention and also for avoiding the major vision loss.

Fig. 10.

Fig. 10

Training and validation accuracy curves of the proposed model, illustrating performance convergence over epochs.

Fig. 11.

Fig. 11

Training and validation loss curves of the proposed model, demonstrating stable learning and convergence behavior.

Fig. 12.

Fig. 12

Confusion matrix of APTOS 2019.

Table 6.

Detailed classification performance on APTOS 2019 test set.

Class Prec. Rec. F1 AUC-ROC
No DR 0.996 0.995 0.995 0.999
Mild 0.978 0.972 0.975 0.994
Moderate 0.961 0.968 0.964 0.989
Severe 0.943 0.956 0.949 0.981
Prolif. 0.938 0.924 0.931 0.976
Macro 0.963 0.963 0.963 0.988
Wtd. 0.989 0.989 0.989 0.998

Experiment 5: evaluation on REFUGE dataset

Optic disc and cup segmentation performance

REFUGE is built for glaucoma evaluation, and it is giving accurate segmentation boundaries for both the optic disc and the optic cup. It is having 1,200 images with manually drawn segmentation masks for the optic disc and the optic cup. The dataset is not aimed at DR grading. But it is good for testing the segmentation accuracy on a different clinical task and also for checking the model generalisation. We use the standard split of 800 images for training, 200 for validation, and 200 for testing.

Segmentation generalization and clinical utility

From Table 7, GTAM-Net is reaching a Dice score of 98.05% for disc segmentation and 93.67% for cup segmentation on REFUGE. The cup-to-disc ratio (CDR) error is 0.048, which is acceptable for glaucoma screening, and 95% of the estimates are falling within Inline graphic of the expert measurements. So the model is generalising well across different segmentation tasks beyond DR, and the gated attention is also useful for medical image segmentation in general.

Table 7.

Optic disc and cup segmentation on REFUGE test set.

Method Disc DSC Cup DSC CDR Error Time (ms)
U-Net23 95.12 ± 0.35 88.45 ± 0.52 0.085 ± 0.012 42.3
Att. U-Net26 96.34 ± 0.31 90.12 ± 0.48 0.072 ± 0.010 45.6
MCAU-Net7 97.12 ± 0.28 91.45 ± 0.45 0.065 ± 0.009 48.2
GTAM-Net 98.05 ± 0.22 93.67 ± 0.38 0.048 ± 0.007 46.8

The optimised model is having an inference time of 46.8 ms, which is fine for real-time clinical applications. The improvements on disc and cup segmentation are suggesting that the multi-scale feature pyramid and the attention together are capturing the retinal structures, and they are recovering both the fine boundaries and the overall shape. Beyond the tasks tested here, the same design could also be used for other retinal layers like the blood vessels, the fovea, and the retinal lesions.

Cross-dataset generalization study

Experimental design and evaluation protocol

For checking how well the model is generalising, we ran cross-dataset experiments. Each model is trained on one dataset and tested on another without any fine-tuning. This setup is similar to a real clinic, where the test data can differ from the training data in terms of patient population, imaging device, or annotation protocol. We tried three settings: (1) train on IDRiD and test on DDR and REFUGE, (2) train on DDR and test on Messidor-2 and APTOS, and (3) train on a mixture of all the datasets and then test on each of them.

Generalization analysis and insights

From Table 8, the model which is trained on the mixed data is having the best overall performance. The drop on each dataset is small, so the diverse training data is really helping. The drop is also smaller for segmentation than for classification, and this is suggesting that the anatomical features are transferring between datasets more easily. The disease patterns are varying more, probably because of differences in cameras and imaging conditions. GTAM-Net is generalising well across datasets: when tested on unseen data, the drop is only around 1–4%, while the baseline methods are usually dropping 5–10%. The model which is trained on the combined data is the most stable, and this is again supporting the use of diverse training data. The gated attention is also helping the generalisation, since the model is learning to focus on the general and transferable features, and it is ignoring the artefacts which are specific to the training set. The attention analysis is also showing that the model is attending to the anatomically general features like vessel structure and disc shape, rather than to the domain-specific image cues, when it is tested on new data. This kind of robustness and adaptability is useful for the clinical deployment, where the data is coming from different cameras and protocols.

Table 8.

Cross-dataset generalization performance.

Train Inline graphic Test Seg. DSC Cls. Acc. Drop (%)
IDRiD Inline graphic DDR 96.78 ± 0.25 94.23 ± 0.32 1.39/3.89
IDRiD Inline graphic REFUGE 97.12 ± 0.23 – 0.93/–
DDR Inline graphic Messidor-2 – 95.45 ± 0.29 –/3.31
DDR Inline graphic APTOS – 94.67 ± 0.31 –/3.56
Mixed Inline graphic IDRiD 98.05 ± 0.20 98.12 ± 0.18 0.12/1.00
Mixed Inline graphic DDR 97.45 ± 0.22 97.23 ± 0.20 0.00/1.00

Computational efficiency and deployment analysis

Complexity analysis and comparison

We also did a detailed evaluation of the computational efficiency of GTAM-Net for clinical use, and Table 9 is reporting the numbers. The computational complexity of the network is 42.3 GFLOPs for a Inline graphic input image, which is only about 11% higher than running two separate models for classification (16.7 GFLOPs) and segmentation (21.4 GFLOPs). The reason is that the features are shared in the early layers, and the gated attention is only activating the task-specific paths when they are needed. The inference time is 46.8 ms per image, so on a single NVIDIA A100 GPU we can process about 21 images per second, which is fine for high-throughput screening. The inference memory is 1.38 GB, so the model is fitting on lower-end hardware and can be deployed in either large screening centres or point-of-care setups.

Table 9.

Computational cost comparison across methods.

Method Params (M) FLOPs (G) Time (ms) Memory (MB)
U-Net+EffNet 58.4 38.1 45.2 1,245
MTANet8 62.3 40.5 48.7 1,387
CTAN10 65.8 43.2 52.3 1,523
YOLO-Med40 59.7 41.8 47.6 1,412
GTAM-Net 61.2 42.3 46.8 1,378

Efficiency-accuracy trade-off analysis

From a Pareto-frontier analysis, GTAM-Net is giving a better accuracy-efficiency trade-off than the existing methods, and it is sitting in the favourable region of high accuracy at moderate cost. The ablation studies are showing that the gated attention block is contributing the largest part of the efficiency gain, and it is removing about 30% of the redundant computation compared to always-on attention. The cross-task attention is adding only a small overhead (around 5% of the total computation). But it is giving a clear accuracy gain through the shared information between tasks. The adaptive computation through task-uncertainty weighting is also helping the training, since more capacity is going to the harder examples and less to the easy ones. This is similar to how the human experts are working, and it is leading to faster convergence and a stronger final result. These efficiency advantages are making GTAM-Net suitable for resource-constrained settings, which are common in global health programs. The multi-scale feature pyramid is also helping and is adding about 0.8% to the segmentation Dice score, since it is capturing both the small local details and the wider retinal structure.

Ablation studies

Experimental design and component analysis

We ran a set of ablation studies for measuring the contribution of each architectural component in GTAM-Net. All the experiments were on IDRiD with the same training settings, so the comparison is fair. We started from a baseline shared-encoder model, and the main components were added one by one for seeing how each design choice is helping. Each experiment was repeated three times, and the results are reported as mean ± standard deviation for statistical reliability.

Component contributions and insights

From Table 10, every component of the framework is adding something to the final performance. The task-specific decoders are improving both segmentation and classification compared to the baseline shared-encoder model, so separating the two tasks after the common feature extraction is mattering. The gated attention block is giving the largest single improvement, with a 1.45% gain in DSC and a 2.22% gain in classification accuracy. So the gated, task-aware feature selection is very effective at adapting the shared representation to each task. The activation maps of the gates are also showing a clear difference between the two branches. The segmentation gate is emphasising the edge and texture features close to the optic disc boundary, while the classification gate is focusing on the lesion-related and broader anatomical features. Hence the gating mechanism is really learning to pick the task-relevant features, instead of giving the same weight to both tasks.

Table 10.

Systematic ablation study results.

Configuration DSC (%) Acc. (%) Inline graphic DSC Inline graphic Acc.
Baseline 94.23 ± 0.31 93.45 ± 0.35 – –
+ Task Decoders 95.67 ± 0.28 95.23 ± 0.31 +1.44 +1.78
+ Gated Att. 97.12 ± 0.25 97.45 ± 0.28 +1.45 +2.22
+ Cross-Task Att. 97.89 ± 0.22 98.34 ± 0.24 +0.77 +0.89
+ Uncertainty 98.17 ± 0.18 99.12 ± 0.18 +0.28 +0.78
Full GTAM-Net 98.17 ± 0.18 99.12 ± 0.18 +3.94 +5.67

The cross-task attention block is letting the useful information move between the segmentation and classification branches. The classification features are helping the segmentation refine the optic disc boundary in the difficult diseased regions, and the segmentation features are bringing the anatomical context which is helping the DR grading. This kind of mutual information transfer is similar to clinical practice, where the localisation and the diagnosis are supporting each other. The uncertainty-aware loss weighting is also improving the model, since it is balancing the two task contributions automatically during the training. It is not letting either task dominate the optimisation, and it is also stabilising the convergence. From Table 10, adding the uncertainty-aware weighting on top of the previous configuration is further improving both segmentation and classification. So the adaptive task balancing is really working for retinal multi-task learning. The learned uncertainty parameters for the segmentation and the classification branches are also evolving smoothly and stabilising during the training, and this is suggesting that the proposed loss weighting is numerically stable.

We also did a few more ablation studies on the key design choices. We tested different numbers of attention heads (2, 4, 8, and 16), and 8 heads is giving the best performance, with a good trade-off between representational power and cost. Among different gating functions (sigmoid, tanh, and softmax), the sigmoid is giving the best balance between selective feature control and stable gradient flow. The multi-scale feature pyramid is adding about 0.8% in segmentation DSC, which is showing that it is useful for picking up both the fine local features and the larger structures of the retina.

Qualitative results and interpretability analysis

Segmentation visualization and error analysis

From the qualitative examples in Fig. 13, GTAM-Net is able to delineate the optic disc boundary in many different imaging scenarios. This is including the harder cases like peripapillary atrophy, tilted discs, and overlapping pathology. Even when the boundary is partly hidden behind hemorrhages or exudates, the model is still keeping a good delineation, because the multi-scale network is using both the local boundary features and the wider anatomical structure. From the error analysis, most of the segmentation errors are falling into two groups. The first group is having heavily diseased eyes, like proliferative DR cases where more than 50% of the optic disc is covered by hemorrhage. The second group is having low-quality images with strong blur or non-uniform lighting. On the harder images, the model is conservative and is tending to slightly under-segment the disc, rather than producing a spurious or unstable boundary. This behaviour is fine for clinical screening, since it is generally safer to have a false positive than to miss a disease in this kind of setting.

Fig. 13.

Fig. 13

Segmentation visualization and error analysis.

Attention visualization and model interpretability

The attention maps are giving us some insight into how GTAM-Net is reasoning, and they are making the model easier to interpret for the clinical users. For segmentation, the attention is concentrated on the optic disc boundary, with weaker activations near the related anatomy like the origin of the vessels and the cup region. For classification, the attention is spreading over the pathological lesions, and it is concentrating on the optic disc region. This is similar to how an expert clinician is reading the image, by looking at the lesions and at the surrounding anatomical context. The visualisation of the gating activations is also showing that the importance of the selected features is changing with the input image. The gates are producing a balanced distribution of activations across features in clear, anatomically simple images, while in more complex images they are focusing on the robust features and are avoiding being misled by the artefacts or strong pathology.

Failure case analysis and limitations

Looking at the failure cases is giving us some insight into the current limits of the method and into possible directions for future work. The most common failures are falling into three groups:

  1. severe pathology which is completely hiding the anatomical boundary,

  2. rare anatomical variations like megalopapilla and tilted disc syndrome, and

  3. images with very heavy artefacts.

The most frequent classification errors are happening between two adjacent DR severity classes, like Mild and Moderate, which are differing by very subtle clinical or subjective criteria, and where even the expert clinicians are sometimes disagreeing.

The failure cases above are suggesting a few possible improvements. Training with more examples of these rare presentations should help. Adding extra clinical information like patient history or visual acuity to the model may also help. Ensemble techniques and the fusion of several imaging modalities are other useful directions. The model is already giving state-of-the-art performance, but it should still be used as a decision-support tool, and not as a full replacement for clinical judgement, especially in the difficult or borderline cases. Even with strong results on the benchmark tasks, GTAM-Net should not be taken as a sign that the problems of retinal image analysis are solved. The reported numbers are dataset-dependent, and they were obtained under tightly controlled conditions on standard public benchmarks. In the real clinical use, the cases will be less clear, the image diversity will be larger, the labels will sometimes be uncertain, and the population shifts can also pull the performance down. So the model should be used as a decision-support tool, and not as a replacement for an expert clinician, especially not in marginal or difficult cases.

Statistical significance and clinical validation

Statistical significance testing

We also did rigorous statistical testing for making sure that the GTAM-Net improvements are real and not from chance. We ran paired t-tests for comparing GTAM-Net with the second-best method on each dataset, and all the p-values are less than 0.001. So the gains are significant at the 99.9% confidence level. We also ran bootstrap resampling with 1,000 runs for computing 95% confidence intervals for the accuracy improvements. The intervals are [2.61%, 3.17%] for IDRiD, [1.52%, 2.04%] for DDR, [1.48%, 1.96%] for Messidor-2, and [1.92%, 2.48%] for APTOS. The DeLong test was used for comparing the ROC curves, and the differences are significant at Inline graphic between GTAM-Net and all the baselines on every dataset. The McNemar test was used for the paired classification results, and it is confirming that the error rate is significantly reduced (Inline graphic). Hence the gains in GTAM-Net are coming from a real improvement in the model, and not from random variation.

Limitations, clinical implications, and future directions

Current limitations and challenges

Even though GTAM-Net is giving state-of-the-art performance, there are still some limitations that we should mention. The model is needing segmentation annotations for training at its best, and that is making it harder to use on datasets which are only having classification labels. The performance is also dropping when the images are having strong artefacts or rare anatomical variations which are not common in the training data. The current model is only processing single images, and it is not using the temporal information from the follow-up visits, even though this would be useful for tracking the disease progression. The computational cost is reasonable for cloud-based use. But it is still a bit too high for very low-power devices in the resource-limited settings. The model is also not using clinical information like HbA1c, the duration of diabetes, or blood pressure, which the doctors are normally using during the diagnosis and treatment planning. These are not fundamental weaknesses of the framework, but they are showing clear directions for future improvement.

Clinical implications and deployment considerations

GTAM-Net is doing well in automated DR screening, with sensitivity above 99% and specificity above 98% for referable DR across several datasets. So the model could be useful in screening programs, it can take some load off the specialists, and it can also improve access to eye care in regions where the resources are scarce. The high consistency across datasets is also encouraging for monitoring the disease progression, and the attention maps are making the predictions more transparent, which can build trust in the model and help in integrating it into clinical workflows. The multi-task design is also letting us do a complete retinal assessment in a single image pass, which is speeding up the screening. Some deployment concerns are still there. The system will need regulatory approval, integration with the existing healthcare IT infrastructure, and also proper training of the medical staff, so that they can read the model predictions correctly.

Future research directions

This work is pointing to several useful directions for future research. The system can be extended to more retinal structures, like blood vessels, the fovea, and the individual lesions, which would give a more complete retinal assessment. The temporal information from follow-up scans of the same patient can also be added for predicting the disease progression over time. Semi-supervised and self-supervised learning methods can be explored for reducing the need of manual annotations. The performance on harder cases may also improve when different data types are combined, like OCT images, angiography, and clinical records. The same architecture can be adapted to other medical imaging fields like dermatology, radiology, and pathology, since the core ideas of gated attention and multi-task learning are likely to remain useful in those areas too. Finally, prospective clinical evaluation in real-world screening programs across diverse populations is important for confirming the practical clinical value of the system, and also for finding the deployment challenges which are not visible in retrospective datasets.

Conclusion

In this work, we proposed GTAM-Net, a unified gated task-attentive multi-task framework for retinal image analysis, which is doing optic disc segmentation and DR severity grading together inside a single architecture. By combining the shared feature learning with task-specific refinement, the framework is letting the anatomical and disease-related representations interact, and at the same time the gated feature control is reducing the inter-task interference. The experimental results on five public benchmark datasets are showing strong performance on both segmentation and classification under different imaging conditions and disease distributions. The framework is also reasonably robust under class imbalance, and it is giving an efficient unified design for retinal image analysis. So this work is presenting a useful task-aware multi-task framework for joint retinal segmentation and grading. The benchmark numbers we are reporting should be read in the context of the public datasets and the controlled experimental protocol used here, and not as evidence of universal clinical performance. Further validation on larger multi-centre cohorts, on more diverse imaging conditions, and also on additional retinal modalities is still needed before the real-world deployment. In future work, we are also planning to look at stronger cross-dataset generalisation, better task-balancing strategies, and lighter designs for practical screening applications.

Author contributions

Conceptualization, M.Z.S.; methodology, M.Z.S, I.Q, and M.F.H.; software, M.Z.S.; validation, M.Z.S., I.Q., M.A., S.A. and Q.A.; formal analysis, M.Z.S. and M.F.H.; investigation, M.Z.S.; resources, I.Q., M.A., S.A. and Q.A.; data curation, M.Z.S. and M.F.H.; writing—original draft preparation, M.Z.S., I.Q., and M.F.H.; writing—review and editing, M.F.H., I.Q., M.A., S.A., Y.C and Q.A.; visualization, M.Z.S. and Y.C.; supervision, I.Q., M.Z.S., and Q.A; project administration, M.Z.S. and Y.C.; funding acquisition, Y.C.

Funding

This work was supported by the National Research Foundation of Korea(NRF) grant funded by the Korea government(MSIT) (No. RS-2023-00218176) and the Soonchunhyang University Research Fund. Authors like to thank Princess Nourah bint Abdulrahman University Researchers Supporting Project number (PNURSP2026R506), Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia. The authors extend their appreciation to the deanship of research and graduate studies at King Khalid University for funding this work through a large research project under grant number RGP2/603/45.

Data availability

All retinal imaging datasets used in this study are publicly available and can be accessed from their respective official sources. The IDRiD is available at https://www.kaggle.com/datasets/mariaherrerot/idrid-dataset, the DDR dataset is accessible at https://www.kaggle.com/datasets/mariaherrerot/ddrdataset, the Messidor-2 dataset can be downloaded from https://www.adcis.net/en/third-party/messidor2/, the APTOS 2019 dataset is available at https://www.kaggle.com/competitions/aptos2019-blindness-detection/data, and the REFUGE dataset can be accessed via https://refuge.grand-challenge.org/. These datasets were originally collected with appropriate ethical approvals and informed consent and are released for research purposes.

Declarations

Competing interests

The authors declare no competing interests.

Informed consent 

This study exclusively utilized publicly available, anonymized datasets. No new human participants were recruited, and no additional data were collected directly from human subjects. All datasets used in this study were originally collected with informed consent and made publicly available for research purposes.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Muhammad Zaheer Sajid, Email: msajid4@gmu.edu.

Yongwon Cho, Email: dragon1won@sch.ac.kr.

References

  • 1.Dai, L. et al. A deep learning system for predicting time to progression of diabetic retinopathy. Nat. Med.30, 584–594. 10.1038/s41591-023-02702-z (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Senapati, A., Tripathy, H. K., Sharma, V. & Gandomi, A. H. Artificial intelligence for diabetic retinopathy detection: A systematic review. Inform. Med. Unlocked45, 101445 (2024). [Google Scholar]
  • 3.Akram, M. et al. Uncertainty-aware diabetic retinopathy detection using deep learning enhanced by Bayesian approaches. Sci. Rep.15, 1342. 10.1038/s41598-024-84478-x (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Gulshan, V. et al. Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA316(22), 2402–2410. 10.1001/jama.2016.17216 (2016). [DOI] [PubMed] [Google Scholar]
  • 5.Dai, L. et al. A deep learning system for detecting diabetic retinopathy across the disease spectrum. Nat. Commun.12, 3242. 10.1038/s41467-021-23458-5 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Arora, L. et al. Ensemble deep learning and EfficientNet for accurate diagnosis of diabetic retinopathy. Sci. Rep.14(1), 30554. 10.1038/s41598-024-81132-4 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Shalini, R. & Gopi, V. P. Multiresolution cascaded attention U-Net for localization and segmentation of optic disc and fovea in fundus images. Sci. Rep.14, 23107. 10.1038/s41598-024-73493-7 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Lin, Y. et al. Mtanet: Multi-task attention network for automatic medical image segmentation and classification. IEEE Trans. Med. Imaging43(2), 692–703 (2024). [DOI] [PubMed] [Google Scholar]
  • 9.Bui, P.-N., Vu, M.-H., Nguyen, N.-T., & Tran-Ngoc, A. Multi-scale feature enhancement in multi-task learning for medical image analysis, arXiv preprint arXiv:2412.00351, 2024. [Online]. Available: https://arxiv.org/abs/2412.00351 [DOI] [PubMed]
  • 10.Kim, S., Li, Y., Kang, H., Lee, Y., & Park, S. H. Cross-task attention network: Improving multi-task learning for medical imaging applications, arXiv preprint arXiv:2309.03837, 2023. [Online]. Available: https://arxiv.org/abs/2309.03837
  • 11.Muthusamy, D. & Palani, P. Deep learning model using classification for diabetic retinopathy detection: An overview. Artif. Intell. Rev.57, 185. 10.1007/s10462-024-10806-2 (2024). [Google Scholar]
  • 12.Yang, J. et al. Optimizing diabetic retinopathy detection with Inception-v4 and dynamic version of snow leopard optimization algorithm. Biomed. Signal Process. Control96, 106501. 10.1016/j.bspc.2024.106501 (2024). [Google Scholar]
  • 13.Abbasi, R. et al. Diabetic retinopathy detection using adaptive deep convolutional neural networks on fundus images. Sci. Rep.15(1), 24647. 10.1038/s41598-025-09394-0 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Wang, V. Y. et al. A deep learning-based adrppa algorithm for the prediction of diabetic retinopathy progression. Sci. Rep.14(1), 31772. 10.1038/s41598-024-82884-9 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Sushith, M. et al. A hybrid deep learning framework for early detection of diabetic retinopathy using retinal fundus images. Sci. Rep.15, 15166. 10.1038/s41598-025-99309-w (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Cheng, J. et al. Wavenet-sf: A hybrid network for retinal disease detection based on wavelet transform in spatial-frequency domain. Neural Netw.194, 108189 (2026). [DOI] [PubMed] [Google Scholar]
  • 17.Qi, Z. et al. Msli-net: Retinal disease detection network based on multi-segment localization and multi-scale interaction. Front. Cell Dev. Biol.13, 1608325 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Zuo, Q. et al. Multi-resolution visual mamba with multi-directional selective mechanism for retinal disease detection. Front. Cell Dev. Biol.12, 1484880 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.F. Guo, et al., Hrprotokd: A hierarchical and relational prototype based knowledge distillation framework for few-shot cancer molecular subtyping, IEEE Journal of Biomedical and Health Informatics, (2025), published online September 18, 2025. [DOI] [PubMed]
  • 20.Feng, X. et al. Cnascope: Pan-cancer copy number aberration database with functional annotation and interactive visualization. Nucleic Acids Res.54(D1), D1364–D1375 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Umirzakova, S., Muksimova, S., Iskhakova, N. & Anorova, S. Y. Im Cho. Enhancing early alzheimer’s disease detection: Integrative approaches using machine learning and deep learning in neuroimaging, Journal of Artificial Intelligence Research and Applications1(2), 113–126 (2024). [Google Scholar]
  • 22.Navaneethan, R. et al. Enhancing diabetic retinopathy detection through a novel mga-csg based approach. Expert Systems with Applications234, 121844 (2024).[Online]. Available: 10.1016/j.eswa.2024.123418
  • 23.Ronneberger, O., Fischer, P., & Brox, T. U-net: Convolutional networks for biomedical image segmentation, in Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015, ser. Lecture Notes in Computer Science. 9351, 234–241. (Cham: Springer, 2015) [Online]. Available: 10.1007/978-3-319-24574-4_28
  • 24.Alhendi, N., & Hamad, H. Optic disc segmentation in fundus images using u-net, in 2025 International Conference on New Trends in Computing Sciences (ICTCS). 14–19 (Amman, Jordan: IEEE, 2025) [Online]. Available: 10.1109/ICTCS65341.2025.10989330
  • 25.Kumar, V. V. N. S., Reddy, G. H. & Giriprasad, M. N. A novel glaucoma detection model using unet++-based segmentation and resnet with gru-based optimized deep learning. Biomed. Signal Process. Control.86, 105069 (2023). [Online]. Available: https://api.semanticscholar.org/CorpusID:259569512
  • 26.Kumar, G. B. & Kumar, S. Enhanced segmentation of optic disc and cup using attention-based u-net with dense dilated series convolutions. Neural Comput. Appl.37, 6831–6847. 10.1007/s00521-025-10989-x (2025). [Google Scholar]
  • 27.Chen, Y., Bai, Y. & Zhang, Y. Optic disc and cup segmentation for glaucoma detection using attention u-net incorporating residual mechanism. PeerJ Comput. Sci.10, e1941. 10.7717/peerj-cs.1941 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Meas, C., Guo, W., & Miah, M.H. Multi-scale attention u-net for optic disc and optic cup segmentation in retinal fundus images, in 2024 2nd International Conference on Advancement in Computation & Computer Technologies (InCACCT). 760–765. (Gharuan, India: IEEE, 2024) [Online]. Available: 10.1109/InCACCT61598.2024.10551123
  • 29.Naing, S. L. & Aimmanee, P. Automated optic disk segmentation for optic disk edema classification using factorized gradient vector flow. Sci. Rep.14, 371. 10.1038/s41598-023-50908-5 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Yu, H., & Ying, W. Two-stage u-net for optic disc/cup segmentation, in 2022 IEEE 2nd International Conference on Data Science and Computer Application (ICDSCA). 275–278 (Dalian, China: IEEE, 2022) [Online]. Available: 10.1109/ICDSCA56264.2022.9987816
  • 31.Wang, J., Li, X. & Cheng, Y. Towards an extended efficientnet-based u-net framework for joint optic disc and cup segmentation in the fundus image. Biomed. Signal Process. Control85, 104906. 10.1016/j.bspc.2023.104906 (2023). [Google Scholar]
  • 32.Zhou, L., et al. Ersr: An ellipse-constrained pseudo-label refinement and symmetric regularization framework for semi-supervised fetal head segmentation in ultrasound images, IEEE Journal of Biomedical and Health Informatics, (2025). published online August 25, 2025. [DOI] [PubMed]
  • 33.Jin, Q. et al. Iterative pseudo-labeling based adaptive copy-paste supervision for semi-supervised tumor segmentation. Knowl. Based Syst.324, 113785 (2025). [Google Scholar]
  • 34.Zhou, Y. et al. A foundation model for generalizable disease detection from retinal images. Nature622, 156–163. 10.1038/s41586-023-06555-x (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Haque, A., Hotan, N., Mondol, T.C., & Imran, A.Z. Generalized multi-task learning from substantially unlabeled multi-source medical image data, arXiv preprint arXiv:2110.13185, 2021. [Online]. Available: https://arxiv.org/abs/2110.13185
  • 36.Vlachostergiou, A., Tagaris, A., Stafylopatis, A., & Kollias, S. Multi-task learning for predicting parkinson’s disease based on medical imaging information, in 2018 25th IEEE International Conference on Image Processing (ICIP). 2052–2056 (Athens, Greece: IEEE, 2018) [Online]. Available: 10.1109/ICIP.2018.8451398
  • 37.Jiang, C., Wang, Y., Yuan, Q., Qu, P. & Li, H. A 3d medical image segmentation network based on gated attention blocks and dual-scale cross-attention mechanism. Sci. Rep.15(1), 6159. 10.1038/s41598-025-90339-y (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Wang, G., et al. Sam-med3d-moe: Towards a non-forgetting segment anything model via mixture of experts for 3d medical image segmentation, 2024. [Online]. Available: https://arxiv.org/abs/2407.04938
  • 39.Jiang, Y., & Shen, Y. Moe: A foundation model for medical multimodal image segmentation with mixture of experts, arXiv preprint arXiv:2405.09446, 2024. [Online]. Available: https://arxiv.org/abs/2405.09446
  • 40.Zhou, Z., Gao, S., Wang, Y. & Zhang, L. Yolo-med: Multi-task interaction network for biomedical images, arXiv preprint arXiv:2403.00245, 2024. [Online]. Available: https://arxiv.org/abs/2403.00245
  • 41.Ates, G. C., Mohan, P. & Celik, E. Dual cross-attention for medical image segmentation. Eng. Appl. Artif. Intell.126, 107139 (2023). [Google Scholar]
  • 42.Fiaz, M. et al. Guided-attention and gated-aggregation network for medical image segmentation. Pattern Recognit.156, 110812 (2024). [Google Scholar]
  • 43.Ling, Y. et al. Mtanet: Multi-task attention network for automatic medical image segmentation and classification. IEEE Trans. Med. Imaging43(2), 674–685 (2024). [DOI] [PubMed] [Google Scholar]
  • 44.Bui, P.N., Nguyen, H.V., Tran, T.N., & Nguyen, T.T. Multi-scale feature enhancement in multi-task learning for medical image analysis, Artificial Intelligence in Medicine, 2025, bibliographic details such as volume, pages, and DOI were not fully verified. [DOI] [PubMed]
  • 45.Wu, R., et al. Mm-retinal: Knowledge-enhanced foundational pretraining with fundus image-text expertise, (2024) arXiv preprint arXiv:2405.11793,
  • 46.Shi, D., et al. Eyefound: A multimodal generalist foundation model for ophthalmic imaging,(2024) arXiv preprint arXiv:2405.11338
  • 47.Shi, D.,et al. A multimodal visual-language foundation model for computational ophthalmology, npj Digital Medicine, 8, 381 (2025). [DOI] [PMC free article] [PubMed]
  • 48.Wang, M. et al. Enhancing diagnostic accuracy in rare and common fundus diseases with a knowledge-rich vision-language model. Nat. Commun.16, 5528 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Wang, X. et al. Automatic detection of 30 fundus diseases using ultra-widefield fluorescein angiography with deep experts aggregation. Ophthalmol. Ther.13, 1125–1144 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Kendall, A., Gal, Y., & Cipolla, R. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 7482–7491 (2018).

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

All retinal imaging datasets used in this study are publicly available and can be accessed from their respective official sources. The IDRiD is available at https://www.kaggle.com/datasets/mariaherrerot/idrid-dataset, the DDR dataset is accessible at https://www.kaggle.com/datasets/mariaherrerot/ddrdataset, the Messidor-2 dataset can be downloaded from https://www.adcis.net/en/third-party/messidor2/, the APTOS 2019 dataset is available at https://www.kaggle.com/competitions/aptos2019-blindness-detection/data, and the REFUGE dataset can be accessed via https://refuge.grand-challenge.org/. These datasets were originally collected with appropriate ethical approvals and informed consent and are released for research purposes.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES