Abstract
Diabetic macular edema (DME) and age-related macular degeneration (AMD) have emerged as leading causes of vision impairment worldwide; optical coherence tomography (OCT) has proven to be a crucial diagnostic tool for these diseases, for its rapid, non-invasive high-resolution imaging of retinal structures. Further accurate assessment of lesions in retinal OCT images plays a pivotal role in the early diagnosis of DME and AMD. However, current diagnosing of AMD and DME is restricted to utilizations of OCT two-dimensional images, for crucial three-dimensional (3D) lesion information inherent in OCT 3D images cannot be effectively extracted and utilized without appropriate methods. Here, we proposed an innovative deep-learning network characterized by fusing multi-scale feature extraction-aggregation and channel-spatial joint attention for high-accuracy 3D lesion segmentations of DME and AMD. Extensive experiment results demonstrated that our proposed method has commendable 3D segmentation performances and robust generalization capabilities, probably helping to understand DME and AMD diseases better and providing great convenience for clinical diagnosis and treatment.
1. Introduction
Diabetic macular edema (DME), a leading cause of vision loss in the working-age population of most developed countries [1–3], is a condition characterized by thickening and swelling of the macula caused by microvascular leakage [4,5]. Age-related macular degeneration (AMD) is characterized by various pathological changes, including the deposition of metabolites between the retinal pigment epithelium (RPE) and Bruch’s membrane (BM) [6], resulting in a macular lesion with symptoms such as blurred or loss of central vision [7]. Intraretinal fluid (IRF) is considered one of the biomarkers for DME [8], while drusen are widely accepted as contributing factors to the pathogenesis of AMD [9]. To ensure effective management of retinal disease and prevent its progression to blindness, it is crucial to closely monitor IRF and drusen, which are recognized as critical clinical indicators and significant risk factors for the development of DME and AMD [10].
Optical coherence tomography (OCT) is a non-invasive three-dimensional (3D) imaging technology [11–13]and plays an important role in the diagnosis and monitoring of retinal diseases relying on its ability to provide high-resolution cross-sectional images for the retinal structures. With B-scans of OCT contained in OCT 3D scanning volume, it can accurately visualize the lesions of AMD and DME in the retina. Precise segmentation of corresponding lesion OCT images further holds substantial clinical significance, but manual segmentation by ophthalmologists suffers from inefficiency and subjectivity. Consequently, automatic lesion segmentation of retinal OCT images is crucial and progressively sought after. However, existing segmentations for multiple two-dimensional (2D) B-scans are inadequate in diagnosing definitive the development of retinal lesions. Previous automatic diagnosis techniques did not efficiently harness the 3D imaging characteristics of OCT to provide comprehensive 3D lesion information.
Over the past few years, deep learning has been successfully applied to the task of medical image processing [14–18], including the fluid region segmentation problem [19–21], Unet [22], Seg-Net [23], Deeplabv3+ [24], and Unet++ [25] are widely recognized and notable deep learning architectures developed for image segmentation and used in medical applications. Inspired by these networks, numerous efficient architectures have been proposed and show superior performance in retinal lesion 2D segmentation [26,27]. Particularly, A. Roy et al. proposed ReLayNet [19], which utilized a contracting path of convolution blocks to learn a hierarchy of contextual features, followed by an expansive path of convolutional blocks for semantic segmentation. X. Yu et al. [27] introduced a novel loss-balanced joint-loss function to handle imbalance in the roles of different branches. G. Xing et al. [28] adopted a novel curvature regularization term to incorporate shape-prior information for OCT fluid segmentation. Suchetha et al. [29] presented a deep learning-based predictive algorithm, which applied Region Convolutional Neural Network (R-CNN) and faster R-CNN to improve the accuracy of AMD lesion segmentation. M. Wang et al. [10] proposed a novel multi-scale transformer global attention network for drusen segmentation in retinal OCT images. W. Liu et al. [30] proposed an attention-based method for DME segmentation, and they integrated multi-scale inputs, multi-side-outputs, and attention methods into Unet++ for better performance. In a study by X. Liu et al., [31] an attention-based Unet model was developed for intraretinal cystoid fluid (IRC) segmentation in DME images with an attention structure that automatically considers suspected regions. R. Rasti et al. [32] proposed the RetiFluidNet model, a multi-attention and self-adaptive deep convolutional network based on an end-to-end training scheme for retinal OCT fluid segmentation, which has demonstrated high performance in fluid segmentation. Moradi et al. [33] introduced the ensemble mechanism to improve non-advanced AMD classification. M. Li et al. [34] employed multi-scale and cross-channel feature extraction and channel attention to segment wet AMD lesions.
Although these newly developed methods have achieved remarkable improvements in robustness and accuracy in 2D segmentations, the majority of them primarily consider the properties of OCT 2D images, therefore such approaches are limited in their ability to account for the complex and various retinal layers and lesions. Above all, as no 3D context is taken into account, 2D results suffer from a lack of 3D spatial consistency which leads to edge discontinuity in the long-axis direction, and false positive predictions at a slice containing other lesions similar to target lesions. Browsing through all 2D images within OCT 3D volume takes time, while reviewing only a few may pose the risk of missing lesions. Given the rapid advancement of technology and the increasing emphasis on precision, 3D results that can show more comprehensive and intuitive information associated with the lesions, are more favored by specialists. Recently, several 3D segmentation methods have been proposed for OCT image segmentation. For instance, Y. He et al. [35] introduced a Unet based on LSTM, which utilizes longitudinal information for OCT layer segmentation. S. Mukherjee et al. [36] proposed a 3D deep convolutional regression network capable of preserving semantic information through the utilization of spatial context, continuity, and anatomical relationships, thereby enabling retinal layer segmentation. Additionally, S. Mukherjee et al. [37] employed a 3D Unet to extract high-dimensional continuous features from OCT volume data, achieving promising performance in retinal layer segmentation tasks for AMD patients. Based on these advancements, accurate and automated 3D segmentation for IRF and drusen is critical and increasingly demanding.
To overcome this limitation, we proposed a novel multicategory, multi-scale, and attention-based network structure that can synergistically explore DME and AMD lesions from OCT 3D volume images and obtain accurate 3D morphology of IRF and drusen. Specifically, a channel and spatial joint attention (CSJA) module is proposed to adaptively identify and select the most discriminative semantic and spatial information to segment lesions from surrounding speckle noise. Further, we proposed an efficient multi-scale feature extraction (MsFE) module to exploit multi-scale features by employing multiple parallel filters with different rates. Meanwhile, a novel multi-scale feature aggregation (MsFA) module is designed to guide the model to fuse adjacent features in the encoding and decoding stages and improve 3D segmentation performance consequently.
The overall workflow of this paper is as follows: Section 2 provides detailed information about the dataset utilized in this study, as well as an exposition of the methodologies proposed and the training strategies employed. Section 3 outlines the experimental procedure, encompassing the data augmentation methods, the configuration environment in which the experiments were conducted, and the evaluation metrics utilized. The results of our proposed model are presented in section 4. In Section 5, we compare the proposed method with existing methods through comparative experiments. We also conduct ablation experiments to evaluate the effectiveness of our model. Finally, we discuss the potential clinical applications of our method. Section 6 is the conclusion related to the findings in this study.
2. Methods
2.1. Datasets
This study was approved by the Institutional Review Board (IRB) of Sichuan Provincial People’s Hospital (IRB-2022-258). The datasets comprising OCT volumes of patients with DME and AMD were obtained using the Swept Source OCT system (BM-400 K BMizar, TowardPi Medical Technology, Beijing, China). This system utilizes a vertically sweeping laser with a wavelength of 1060 nm and a scanning speed of 400,000 A-scans per second. Additionally, TowardPi’s OCT angiographic scanning mode can generate angiographic images covering a 6 mm × 6 mm area centered on the macula. Within this 6 mm × 6 mm scan range, there are 512 A-scans and 512 B-scans.
Considering that the limited training data annotated with a quality golden standard is also a significant factor restricting the development of 3D segmentation of DME and AMD, we constructed a new 3D dataset containing 120 cases (61,440 slices) of DME and AMD volumes and corresponding manual pixel-level annotation from specialized ophthalmologists. Firstly, three ophthalmologists manually annotated the IRF and drusen regions frame by frame in the dataset using ITK-SNAP software. After the annotations were completed, a senior ophthalmologist reviewed and rated them, and the final annotations were selected based on the highest ratings. All participants underwent a comprehensive ocular examination including refraction and best-corrected visual acuity, non-contact intraocular pressure (IOP), ocular axis, slit lamp, wide-angle fundus imaging, and OCT. This study was approved by the Ethics Committee of Sichuan Provincial People’s Hospital and adhered to the Declaration of Helsinki. All participants provided written informed consent after the nature and possible consequences of the study were explained. More detailed information can be seen in Table 1. In the study, we used a fixed split on both the DME and AMD datasets, with 45 samples for training and the remaining 15 samples for validation. And since data that simultaneously exhibits both IRF and drusen is rare, there is only one set in the dataset used as the validation set.
Table 1. Information of the employed datasets.
| Dataset | DME | AMD | DME&AMD |
|---|---|---|---|
| Scanner device | BM-400 K Bmizar | BM-400 K Bmizar | BM-400 K Bmizar |
| Scan range | 6 mm × 6 mm | 6 mm × 6 mm | 6 mm × 6 mm |
| Total | 60 volumes × 512 slice | 60 volumes × 512 slice | 1 volumes × 512 slice |
| Size | 512 × 1044 | 512 × 1044 | 512 × 1044 |
| Resize | 512 × 512 | 512 × 512 | 512 × 512 |
| Dimension | 3D | 3D | 3D |
| Training / Test | 90/10 | 90/10 | - |
As shown in Fig. 1, (a1) represents the raw 3D data of the case related to DME, along with the corresponding 3D visualization of the labels. (a2) demonstrates the B-scan images that correspond to the blue slice of the raw data mentioned in (a1). (a3) exhibits the label images corresponding to the green dotted line observed in the 3D visualization of the labels. Additionally, (b) and (c) display cases of AMD, as well as a case where both DME and AMD lesions are present.
Fig. 1.
3D visualization and B-scans of examples in datasets. The first row shows the 3D OCT original image and the 3D lesion annotation, while the second row displays the blue cross-section from the OCT image and the corresponding 2D annotation with green dashed lines.
2.2. Overview architecture
To perform proper and robust 3D segmentation of DME and AMD, we proposed a novel network architecture, which adopts an encoder-decoder U-shaped network as the basic framework and integrates three core modules, as illustrated in Fig. 2, our main contributions are the synergistic effect among the three fused modules: CSJA, MsFE, and MsFA. The CSJA module is adopted in skip-connection to effectively extract multi-scale context information and enhance the proposed network’s ability to segment diverse retinal lesions. The MsFE module is proposed to adaptively identify and choose useful semantic information and spatial information for retinal multi-sort segmentation. In addition, the MsFA module is used to learn more semantic representations to refine the DME and AMD segmentation map.
Fig. 2.
The overall architecture of the proposed network is as follows: (a) shows the complete architecture of our model, which is capable of automatically performing 3D segmentation of OCT retinal lesions, and (b) illustrates the CSJA module, which leverages the attention mechanism to capture more accurate target information.
To dynamically determine the pertinent semantic and spatial information, we introduced the CSJA module (refer to Fig. 2(b)), aimed at facilitating the effective fusion of adjacent hierarchical features to more efficiently capture underlying semantic information. Initially, the features extracted from the layers, possessing identical dimensions, undergo a channel concatenation operation. Subsequently, to enhance the utilization of the most pertinent feature channels, channel attention is introduced to mitigate interference from irrelevant noise. Global max-pooling and global average pooling operations are employed to acquire the global information of each channel, and their outcomes are integrated. The resulting is processed in the spatial channel to obtain . Finally, we add the reweighted low-level features to the high-level features to yield the final result M.
| (1) |
| (2) |
| (3) |
where refers to the sigmoid activation used for deriving the attention map at the present scale, denotes the Multi-Layer Perceptron (MLP) operator. Moreover, represents the global max-pooling operation, stands for the global average pooling operation, and signifies the convolution with the 7 × 7 kernel. represents the operation of concatenation, and denote extract the average and maximum value in the channel, respectively. indicates the multiplication. Additionally, F denotes the initial input feature.
At the bottom of our network, the MsFE module is devised to extract concealed background information and consolidate features across multiple scales. As shown in Fig. 3, the MSFE module contains multi-scale feature extraction (see Fig. 3. (a)) and adaptive feature selection (see Fig. 3. (b)). The multi-scale feature extraction component generates a more comprehensive feature map to effectively capture scale variations in images, enabling the model to better handle feature at different scales. First, the feature map generated by is equally divided into four parallel feature groups, each of which is subjected to a dilated convolution operation (dilated rate = 1, 6, 12 and 18, respectively) to derive a more abundant feature map with various receptive fields. Thus, features at different scales in DME and AMD lesions can be adequately perceived by this module. The operations for extracting parallel features are as follows:
| (4) |
| (5) |
where represents the dilated convolution operation with dilated rate of , denotes the element-wise addition. Finally, these multi-scale feature information F are merged by the element-level addition and delivered to the adaptive feature selection component to further obtain global contextual features and suppress the interference of irrelevant local features adaptively. After a layer normalized, the input feature (where H, W and C indicate the height, width and channel of F respectively) first produces query (Q), key (K), and value (V) projections, enriched with context. It is achieved by applying 1 × 1 convolutions to aggregate pixel-wise cross-channel context followed by 3 × 3 depth-wise convolution to encode channel-wise spatial features, yielding , , . where is the 1 × 1 point-wise convolution and is the 3 × 3 depth-wise convolution. Next, we reshape query and key projections such that their dot-product interaction generates a transposed-attention map A of size C × C, instead of the huge regular attention map of size HW × HW. In addition, a residual connection is developed to aggregate the multi-scale feature maps. Overall, the adaptive feature selection process is defined as:
| (6) |
| (7) |
Fig. 3.
Multi-scale Feature extraction Module. The MsFA module is designed to extract feature at multiple scales, enabling it to preserve more detailed global information. First, the input feature map is divided into four parallel branches. Then, dilated convolution operations with different dilation rates (1, 6, 12, and 18) are applied to these feature maps to extract a richer feature map with various receptive fields.
Here, is a learnable scaling parameter to control the magnitude of the dot product of K and Q before applying the softmax function.
Several previous studies have shown that aggregating multi-scale features can effectively improve network performance [38,39]. Here we propose the MsFA module (see Fig. 4) to obtain more semantic details and refine the segmentation results. First, the input feature channel generated by the decoding layers is compressed to 8 using convolution.
| (8) |
where is a convolution with 1 × 1 kernel. The process involves concatenating features from neighboring scales in a top-down manner. Subsequently, an attention mechanism is employed to refine the features at various scales, resulting in a more discriminative representation of features.
| (9) |
| (10) |
| (11) |
| (12) |
where denotes the element-wise multiplication; represents the up-sampling operation; denotes the sigmoid activation to obtain the attention map of the current scale; and has a convolution with the 3 × 3 kernel, which is used to generate features with rich, detailed information and pass it to the next attention block until the final prediction of the category. Finally, the output of the MsFA module is obtained using the convolution of the two output channels followed by the sigmoid.
| (13) |
Fig. 4.
Multi-scale Feature Aggregation Module. a convolution is applied to reduce the channel dimension of the input features from the decoding layers to 8. Next, features from adjacent scales are concatenated in a top-down fashion. This process enhances the feature representation, making it more discriminative.
Through the utilization of the MsFA module to steer the gradual aggregation of features across neighboring scales, our proposed model exhibits improved efficacy in lesion segmentation.
2.3. Loss function
The absence of distinct labels for the DME and AMD regions presents a significant challenge, as the background noise and blurred boundaries make up a substantial portion of the retinal image. This lack of labeling complicates the accurate identification and delineation of these specific regions within the images, thereby posing a significant obstacle to our analysis. The Dice loss function [40] is often used to mitigate the negative effects of the imbalance between foreground and background in a sample. As the Dice loss tends to apply to only one kind of segmentation task and may not exhibit optimal performance in multi-task segmentation. To tackle this problem, we choose the Tversky loss [41] as the loss function, which introduces two coefficients and to better balance false negative and false positive. When both and are equal to 0.5, the Tversky loss also becomes the Dice loss. The Tversky loss can be expressed by:
| (14) |
where the is the probability of voxel i be a lesion and is the probability of voxel i be a non-lesion. Also, is 1 for a lesion voxel and 0 for a non-lesion voxel and vice versa for the . To optimize the model, we further introduced binary cross entropy (BCE) to optimize the model, which is defined:
| (15) |
where y is the annotated ground truth, x represents the predicted output by our model. Finally, we define the joint objective function of our network for DME and AMD segmentation as:
| (16) |
where is a balance parameter that is set as 0.5 for all experiments.
3. Experiment setup
3.1. Data augmentation strategy
Our data was augmented with random rotations, shifts, rolls and resizes, resulting in seven augmented images for each image in the training datasets.
3.2. Implementation details
The proposed model was implemented in the TensorFlow v2.5.0 with NVIDIA RTX A5000 (24 G) GPUs, and Python 3.9. In all comparative experiments, we set the maximum number of epochs, batch size and initial learning rate as 50, 4 and 0.0001, respectively. The learning rate was multiplied by 0.8 every 5 epochs, and Adam serves as the optimizer. In addition, since all the B-scans contained a lot of background area without information, we cropped the original image to 512 × 512 to improve the training efficiency.
3.3. Performance measures
All comparison networks discussed in this paper were trained with consistent training process and utilized the same dataset to ensure accuracy and impartiality in the evaluation. To comprehensively evaluate the model's performance, we employed commonly used metrics to quantitatively evaluate the segmentation performance of various methods in this study. These metrics include the Dice similarity coefficient (DSC), balanced accuracy (ACC), precision, recall, specificity and intersection over union (IOU). These metrics yield values within the range of 0 to 1, with higher values indicating more accurate segmentation. Their calculation formula is as follows:
| (17) |
| (18) |
| (19) |
| (20) |
| (21) |
| (22) |
where TP, FP, TN and FN denote true positive, false positive, true negative and false negative variables, respectively.
4. Result
In this section, we first present the 3D segmentation results for IRF and drusen using the proposed model, as shown in Fig. 5 and Fig. 6. Additionally, we display the segmentation results of our proposed network on new datasets, as illustrated in Fig. 7. Detailed files of 3D segmentation results and corresponding original 3D OCT images of Fig. 5–8 can be seen and downloaded from https://tianchi.aliyun.com/dataset/169157.
Fig. 5.
3D visualization of IRF segmentation results and some segmentation results in B-scans from freestanding individuals with DME in validation sets.
Fig. 6.
3D visualization of drusen segmentation results and some segmentation results in B-scans from freestanding individuals with AMD in validation sets.
Fig. 7.
3D segmentation results are obtained from a new dataset. The DME and AMD segmentations are shown in orange and yellow, respectively. (d), (e), and (f) display the 3D visualization of segmentation results associated with DME, AMD, and both conditions, respectively.
Fig. 8.
3D visualization of IRF segmentation results from different methods over the DME datasets. 3D images in the second and third rows correspond to the zoomed-in patches of the olive-green and aqua regions of the first row, respectively. The blue dashed circles in the second and third rows indicate that the comparison models may exhibit over-segmentation or under-segmentation.
Figure 5 and Fig. 6 show the 3D segmentation results of our network for patients with DME and AMD, respectively. Images in the first to third rows correspond to the OCT 3D images consisting of massive B-scans, the 3D segmentation results by our model and the corresponding manually annotated 3D images. As evidenced in Fig. 5 and Fig. 6, the visually intuitive results demonstrate that the 3D segmentation results of fluid lesions achieved by our proposed network closely align with the original labeled images, thereby furnishing more precise references for ophthalmologists to identify fluid lesions accurately. What's more, compared with conventional 2D segmentation methods, 3D segmentation results can provide a more intuitive and comprehensive representation of retinal disease progression. We can measure the central macular thickness (CMT) based on the volumetric visualization of fluid, which is the most widely employed among optical coherence tomography parameters to evaluate the efficacy of treatments against macular edema [42].
Our proposed model can also be utilized for segmenting images in which both DME and AMD lesions are present. Figure 7 illustrates the 3D segmentation results in OCT retinal images using our proposed model, obtained from randomly selected datasets (TowardPi, 512 × 1024 × 512), which are distinct from the training and validation sets. The IRF and drusen segmentation results are depicted in orange and yellow in the B-scan, respectively. As depicted in Fig. 7(f), our proposed network consistently delivers commendable 3D segmentation performance in the 3D images with two types of lesions. These findings substantiate the efficacy and resilience of our network for fluid segmentation in OCT retinal images, thus underscoring its potential applicability in clinical practice by furnishing precise lesion segmentation to facilitate ophthalmologists in diagnosis and alleviate their workload.
5. Discussion
In this section, a detailed comparison was conducted with a variety of classical medical image segmentation methods and the most advanced retinal fluid segmentation methods to offer a more thorough evaluation of the proposed approach’s performance. In order to demonstrate the efficacy of the proposed network, the average and standard values of DSC, precision, accuracy, recall, specificity, and IOU were utilized for a quantitative assessment of the DME and AMD 3D segmentation performance on our dataset. We tested all 15 samples in the validation set and performed statistical analysis on the obtained results. The standard deviation in the table represents the variability of the segmentation accuracy across the entire validation set.
5.1. Quantitative evaluation
Table 2 presents a quantitative metric results comparison among six metrics, our proposed model outperformed the comparative methods, with the average DSC, precision, specificity, and IOU values of 89.76%, 92.14%, 99.62%, and 78.57% respectively on DME dataset. Compared with those existing methods, the average DSC of our model outperforms Unet [22], ResUnet [43], AttentionUNet [44], Unet++ [25], CE-Net [45], RetifluidNet [32], 3D Unet [46] by 5.31%, 2.52%, 2.5%, 6.87%, 3.63%, and 3.42%, 1.69% respectively. In addition, the average precision, accuracy, specificity, and IOU of the proposed model are 3.64%, 0.04%, 0.1%, and 0.43% higher than RetifluidNet, which is state-of-art approaches for fluid segmentation on OCT images. The transformer-based model, Swin-Transformer [47], achieved an average DSC score of 85.22%, which is slightly lower than the performance of the model we proposed. The results consistently demonstrate that our method yields superior performance metrics.
Table 2. Quantitative results of different methods on the DME dataset.
| Network | DSC | Precision | Accuracy | Specificity | IOU |
|---|---|---|---|---|---|
| Unet [22] | 84.45 ± 6.39 | 75.83 ± 9.36 | 98.84 ± 0.56 | 98.92 ± 0.55 | 73.60 ± 9.14 |
| Unet++ [25] | 87.24 ± 5.52 | 84.74 ± 8.28 | 99.14 ± 0.42 | 99.45 ± 0.3 | 77.77 ± 8.17 |
| AttentionUNet [44] | 87.26 ± 5.21 | 82.47 ± 7.09 | 99.10 ± 0.45 | 99.26 ± 0.43 | 77.77 ± 7.85 |
| ResUnet [43] | 82.89 ± 6.75 | 82.87 ± 8.34 | 98.76 ± 0.77 | 99.41 ± 0.33 | 71.34 ± 9.77 |
| CE-Net [45] | 86.13 ± 5.99 | 80.54 ± 9.17 | 99.01 ± 0.50 | 99.23 ± 0.43 | 76.09 ± 8.62 |
| RetiFluidNet [32] | 86.34 ± 6.24 | 84.17 ± 7.94 | 99.04 ± 0.59 | 99.41 ± 0.36 | 76.47 ± 9.07 |
| 3D Unet [46] | 88.07 ± 5.73 | 87.69 ± 5.31 | 98.87 ± 0.63 | 99.48 ± 0.25 | 77.55 ± 8.37 |
| SwinTransformer [47] | 85.22 ± 6.34 | 83.71 ± 7.54 | 98.93 ± 0.71 | 99.27 ± 0.39 | 74.85 ± 8.41 |
| Ours | 89.76 ± 5.13 | 92.14 ± 3.41 | 99.07 ± 0.52 | 99.62 ± 0.17 | 78.57 ± 7.94 |
We also conducted similar evaluations on the AMD 3D dataset, as Table 3 shows. The methods compared in this study include classical CNN-based medical image segmentation models, such as ResUnet [43], AttentionUNet [44], Unet++ [25], RetifluidNet [32], and 3D Unet [46] as well as transformer-based approaches, including Swin-Unet [48] and Swin-Transformer [47]. The average DSC, precision, accuracy, recall, and IoU of our method are 82.59%, 86.91%, 99.27%, 86.51%, and 70.36%, respectively, which convincingly demonstrates that our model is effective in drusen segmentation. In addition, the average DSC, precision, accuracy, recall, and IOU of the proposed model are 3.96%, 11%, 0.13%, 1.04% and 3.28% higher than RetifluidNet. The transformer-based models, Swin-Unet and Swin-Transformer, achieved average DSC scores of 65.76% and 71.82%, respectively, which are slightly lower than the performance of the model we proposed. On the whole, the proposed method demonstrates enhanced reliability in DME and AMD 3D segmentation across various metrics, indicating its significant superiority over other methods in terms of 3D segmentation performance.
Table 3. Quantitative results of different methods on the AMD dataset.
| Network | DSC | Precision | Accuracy | Recall | IOU |
|---|---|---|---|---|---|
| Unet++ [25] | 56.41 ± 23.44 | 82.69 ± 22.20 | 98.57 ± 0.71 | 44.66 ± 21.04 | 42.63 ± 20.35 |
| AttentionUNet [44] | 58.99 ± 23.25 | 73.10 ± 27.66 | 98.56 ± 0.69 | 50.23 ± 20.26 | 45.18 ± 20.22 |
| ResUnet [43] | 46.25 ± 26.21 | 54.70 ± 29.54 | 98.04 ± 0.79 | 40.76 ± 24.20 | 33.61 ± 20.61 |
| Swin-Unet [48] | 65.76 ± 18.33 | 83.68 ± 17.33 | 98.64 ± 0.77 | 55.72 ± 18.53 | 51.34 ± 17.15 |
| RetiFluidNet [32] | 78.63 ± 15.44 | 75.91 ± 20.59 | 99.14 ± 0.39 | 85.47 ± 8.28 | 67.08 ± 18.08 |
| 3D Unet [46] | 77.58 ± 16.34 | 80.94 ± 21.63 | 99.16 ± 0.45 | 84.13 ± 9.82 | 68.91 ± 16.04 |
| SwinTransformer [47] | 71.82 ± 16.93 | 79.31 ± 18.57 | 98.87 ± 0.62 | 77.19 ± 13.20 | 65.45 ± 21.09 |
| Ours | 82.59 ± 11.35 | 86.91 ± 10.21 | 99.27 ± 0.31 | 86.51 ± 8.13 | 70.36 ± 15.43 |
5.2. Qualitative analysis
The qualitative experimental results of our method and other competitors on typical cases are illustrated in Fig. 8 and Fig. 9. We visualize results and compare qualitative results with zoomed-in patches. As shown in Fig. 8 and Fig. 9, most methods can generate the lesion of DME and AMD with high accuracy. However, in the case of fluid lesions, competing models generate fragments that do not belong to lesions of DME and AMD, particularly in the areas indicated by blue circles. Furthermore, they often fail to capture the finer details of the lesions, particularly in the regions highlighted by blue circles. Compared with other methods, our proposed method has achieved well-pleasing 3D results, as shown in the zoomed-in patches. As for the Unet++ method, it obtains more approving results than Unet but still has some fragments. AttentionUNet and RetifluidNet have achieved good results, but there is still a clear difference from the ground truth. Moreover, even speckle noise and blurry edges negatively affect the results, the proposed model achieves better segmentation performance than 3D UNet, AttentionUNet, Unet++, CE-Net, ResUnet, and Unet. Based on quantitative and qualitative analysis, we can see that the proposed model achieves a significant and consistent performance benefit, due to the synergistic effect among CSJA, MsFE, and MsFA modules by seamless integration. These results demonstrate the effectiveness and robustness of our proposed method when performing 3D segmentation of DME and AMD.
Fig. 9.
3D visualization of drusen segmentation results from different methods over the AMD datasets. 3D images in the second and third rows correspond to the zoomed-in patches of the olive-green and aqua regions of the first row, respectively. The blue dashed circles in the second and third rows indicate that the comparison models may exhibit over-segmentation or under-segmentation.
To further evaluate the performance of our model, we have conducted additional experiments using the publicly available RETOUCH dataset [49]. This dataset we used comprises 24 volumetric scans acquired with the Cirrus HD-OCT (Zeiss Meditec), each with dimensions of 512 × 1024 × 128. For preprocessing, all B-scan images across the various volumes were resized to 256 × 256 pixels to standardize the field of view. Subsequently, the dataset was then randomly split into training and validation sets in a 20:4 ratio to ensure a balanced approach for model training and evaluation, facilitating effective performance assessment. As shown in Table 4, our method achieved DSC, precision, accuracy, recall, and IOU scores of 86.47%, 83.20%, 98.65%, 88.26%, and 75.13%, respectively. These results demonstrate superior performance compared to AttentionUNet, 3D Unet, and Swin-Transformer, with our approach leading across all evaluation metrics. Figure 10 illustrates several B-scan segmentation results, where our method provides the segmentation outcomes most closely aligned with the ground truth. These findings highlight the significant potential and effectiveness of our approach in segmentation tasks.
Table 4. Quantitative results of different methods on the RETOUCH dataset.
| Network | DSC | Precision | Accuracy | Recall | IOU |
|---|---|---|---|---|---|
| AttentionUNet [44] | 78.71 ± 12.36 | 75.82 ± 16.04 | 97.25 ± 1.13 | 79.63 ± 10.39 | 70.82 ± 13.67 |
| Swin-Transformer [47] | 81.92 ± 10.04 | 76.18 ± 15.61 | 97.83 ± 1.06 | 84.19 ± 8.16 | 72.06 ± 12.23 |
| 3D Unet [46] | 83.25 ± 9.67 | 81.13 ± 13.08 | 98.24 ± 0.91 | 85.27 ± 7.45 | 73.58 ± 10.78 |
| Ours | 86.47 ± 7.83 | 83.20 ± 9.16 | 98.65 ± 0.82 | 88.26 ± 5.73 | 75.13 ± 9.25 |
Fig. 10.
Comparison of different method IRF segmentation results on the RETOUCH dataset
In certain challenging scenarios, the performance of our model may be limited to some extent. For instance, OCT scans can suffer from jitter during the scanning process, which disrupts the three-dimensional continuity of the data. This disruption can significantly affect the model’s segmentation accuracy, especially when comparing frames before and after the jitter occurs. Additionally, in OCT images obtained from low-quality scans with low contrast, the continuity between different lesion regions becomes more blurred, which negatively impacts the model’s ability to accurately extract spatial information of distinct lesions. Future work will focus on developing specialized image registration and enhancement algorithms to address these issues.
5.3. Model parameters and inference times
In this section, we compare the parameter counts and inference times of different models to evaluate their computational efficiency. By analyzing both the number of parameters and the time required for predicting a single image, we aim to provide a comprehensive understanding of the trade-offs between model complexity and inference speed. As shown in Table 5, Unet, with its relatively simple structure, contains the fewest parameters. AttentionUNet achieves the fastest inference speed, with an average prediction time of 0.0299 seconds per image. However, it comes with a higher computational cost. The model we proposed requires a series of attention mechanism operations, which results in slightly higher parameter counts and inference times. In future work, optimizing for lightweight and more efficient segmentation models, with reduced computational overhead and faster inference times, will be a primary focus of our research.
Table 5. The comparison of model parameters and average inference time per image.
5.4. Ablation experiment
Here we conducted a detailed ablation study on the proposed model’s architecture. Firstly, we took a U-shaped network consisting of an encoder with residual blocks [50] and a feature decoder as a baseline. Subsequently, we individually incorporated modules CSJA, MsFE, and MsFA into the baseline to obtain Model 1, Model 2, and Model 3. As shown in Table 6, all three models achieve commendable improvement over the base in terms of all four metrics. To further analyze the effectiveness of the proposed modules, Model 4 combined CSJA with MsFE, Model 5 combined CSJA with MsFA, and Model 6 combined MsFE with MsFA were compared. Finally, Model 7 was integrated with the three modules of CSJA, MsFE, and MsFA into the baseline. As shown in Table 6, compared with the baseline, DSC in Model 7 has improved by 5.31%, far greater than the performance of Model 1 (↑ 0.91%), Model 2 (↑ 1.44%), Model 3 (↑ 1.71%), Model 4 (↑ 2.33%), Model 5 (↑ 3.08%), and Model 6 (↑ 1.89%). The results of the ablation experiments conclusively demonstrated that the proposed modules significantly improve the 3D segmentation performance in OCT images.
Table 6. Ablation study results on the DME dataset.
| Network | Module | Metrics | |||||
|---|---|---|---|---|---|---|---|
|
| |||||||
| CSJA | MsFE | MsFA | DSC | Precision | Accuracy | IOU | |
| Baseline | 84.45 ± 6.39 | 75.83 ± 9.36 | 98.84 ± 0.56 | 73.59 ± 9.14 | |||
| Model 1 | √ | 85.38 ± 6.39 | 80.48 ± 9.34 | 98.97 ± 0.54 | 74.97 ± 8.84 | ||
| Model 2 | √ | 85.89 ± 5.77 | 87.84 ± 6.09 | 99.0 ± 0.69 | 75.7 ± 8.46 | ||
| Model 3 | √ | 86.16 ± 5.93 | 91.72 ± 4.93 | 99.07 ± 0.52 | 76.13 ± 8.53 | ||
|
| |||||||
| Model 4 | √ | √ | 86.78 ± 5.50 | 81.65 ± 7.60 | 99.07 ± 0.45 | 77.02 ± 7.91 | |
| Model 5 | √ | √ | 87.53 ± 4.12 | 85.00 ± 4.98 | 98.67 ± 0.55 | 78.06 ± 6.37 | |
| Model 6 | √ | √ | 86.34 ± 6.24 | 84.17 ± 7.94 | 99.04 ± 0.59 | 76.46 ± 9.07 | |
|
| |||||||
| Model 7 | √ | √ | √ | 89.76 ± 5.13 | 92.14 ± 3.41 | 99.11 ± 0.53 | 78.56 ± 7.92 |
5.5. Clinical trial results
Figure 11 illustrates the progress of the DME and AMD patients over three months, during which time the patient received two intravitreal injections. These pictures were arranged by time from left to right, representing different stages of treatment. As shown in Fig. 11(a1) and 11(g1), the patient had significant abnormalities in central macular thickness (CMT) before initial treatment. At this point, the patient's CMT was measured to be 515 µm, and patients with CMT measured as ≥ 300 µm may undergo intravitreal anti-VEGF or steroid injections [51,52]. The second row of Fig. 10 shows the B-scan images corresponding to the same scanning positions. Figure 11(g1)-(i1) show the heatmap of CMT in the three stages. The green parts represent the thickness at the normal value and critical value, respectively. However, the red part indicates that the part is thicker than the normal range, and the redder the color, the thicker the thickness. One month after receiving the first treatment, the patient's CMT was reduced to 488 µm (see Fig. 11(h1)). A further diagnosis two months later showed the patient's CMT to have dropped significantly to 428 µm (see Fig. 11(i1)). Meanwhile, as shown in Fig. 11(a1)-(c1), the patient's DME lesions were significantly reduced after receiving the treatment. The above results further demonstrate the accuracy of our 3D segmentation results. The same follow-up findings are seen in patients with AMD (see Fig. 11(a2)-(i2)), The patient had significant CMT excess before treatment and this symptom improved after treatment, which is consistent with the segmentation results of our method.
Fig. 11.
Condition monitoring of a patient with DME (a1-i1) and AMD (a2-i2) at different stages of treatment. (a), (b) and (c) are the 3D visualization results of the DME lesions of this patient before, during, and after treatment, respectively. (d), (e) and (f) are the B-scan images corresponding to the same section. (g), (h) and (i) are the corresponding macular thickness heatmaps of (a), (b) and (c).
The above experiment demonstrates that the alterations in fundus conditions of patients with DME or AMD can be accurately and visually represented through the 3D segmentation results. Moreover, the 3D segmentation results align with the diagnostic findings obtained through existing clinical diagnostic methods. This approach can additionally offer comprehensive information regarding DME and AMD, thereby facilitating convenient diagnosis and monitoring of therapeutic efficacy. Furthermore, it enhances the acceptance and recognition of diagnostic outcomes among clinical patients. The 3D segmentation results of DME and AMD lesions also assist ophthalmologists in effectively communicating with patients by providing them with more visualized information.
6. Conclusion
In this paper, we proposed a novel multi-sort, attention-based, multi-scale network for high-accuracy 3D segmentation of IRF and drusen in OCT images. Our proposed network effectively leverages multi-attention mechanisms to select valuable information from multi-scale features. It fuses features at various levels to attain enhanced semantic representations for 3D segmentation. Three core modules are proposed, namely, CSJA, MsFE, and MsFA. The CSJA module is adopted in skip-connection to effectively extract multi-scale context information and enhance the proposed network’s ability to segment diverse retinal lesions. The MsFE module is proposed to adaptively identify and choose useful semantic information and spatial information for retinal multi-sort segmentation. In addition, the MsFA module is used to learn more semantic representations to refine the DME and AMD segmentation map. Compared with other state-of-the-art (SOTA) methods, the proposed method achieves promising performance in lesion segmentation on datasets acquired by the swept source OCT scanner. The proposed model achieves more accurate 3D prediction of fluid regions, and its DSC outperforms the SOTA methods by 5.31%, 2.52%, 2.5%, 6.87%, 3.63%, and 3.42% for DME, respectively. Likewise, our model demonstrates exceptional proficiency in the segmentation of AMD cases. Results demonstrate that our proposed method has commendable 3D segmentation performances and robust generalization capabilities, probably helping to understand DME and AMD diseases better and providing great convenience for clinical diagnosis and treatment.
In our future work, we plan to investigate domain adaptation algorithms to enhance both the performance and generalization of the model. This is motivated by the variability in image distributions captured by different OCT scanning devices. A model trained on a single dataset may struggle to maintain consistent high performance when applied to out-of-domain test data.
Funding
National Natural Science Foundation of China 10.13039/501100001809 ( 61905036); China Postdoctoral Science Foundation 10.13039/501100002858 ( 2021T140090, 2019M663465); Fundamental Research Funds for the Central Universities 10.13039/501100012226 ( ZYGX2021J012); Medico-Engineering Cooperation Funds from University of Electronic Science and Technology of China ( ZYGX2021YGCX019).
Disclosures
The authors declare no conflicts of interest.
Data availability
Data underlying the results presented in this paper are not publicly available at this time but may be obtained from the authors upon reasonable request.
References
- 1.Abraham J. R., Wykoff C. C., Arepalli S., et al. , “Aqueous Cytokine Expression and Higher Order OCT Biomarkers: Assessment of the Anatomic-Biologic Bridge in the IMAGINE DME Study,” Am. J. Ophthalmol. 222, 328–339 (2021). 10.1016/j.ajo.2020.08.047 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Khalifa M., Albadawy M., “Artificial intelligence for diabetes: Enhancing prevention, diagnosis, and effective management,” Computer Methods and Programs in Biomedicine Update 5, 100141 (2024). 10.1016/j.cmpbup.2024.100141 [DOI] [Google Scholar]
- 3.Vitiello L., Salerno G., Coppola A., et al. , “Switching to an Intravitreal Dexamethasone Implant after Intravitreal Anti-VEGF Therapy for Diabetic Macular Edema: A Review,” Life (Basel, Switz.) 14(6), 725 (2024). 10.3390/life14060725 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Musat O., Cernat C., Labib M., et al. , “Diabetic macular edema,” Rom. J. Opthalmol. 59, 133–136 (2015). [PMC free article] [PubMed] [Google Scholar]
- 5.Wong T. Y., Sabanayagam C., “The War on Diabetic Retinopathy: Where Are We Now?” Asia-Pac. J. Ophthalmol. 8(6), 448–456 (2019). 10.1097/APO.0000000000000267 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Brandl C., Günther F., Zimmermann M. E., et al. , “Incidence, progression and risk factors of age-related macular degeneration in 35–95-year-old individuals from three jointly designed German cohort studies,” BMJ Open Ophthalmology 7(1), e000912 (2022). 10.1136/bmjophth-2021-000912 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Merl-Pham J., Gruhn F., Hauck S. M., “Proteomic profiling of cigarette smoke induced changes in retinal pigment epithelium cells,” in Retinal Degenerative Diseases: Mechanisms and Experimental Therapy , (Springer, 2016), pp. 785–791. [DOI] [PubMed] [Google Scholar]
- 8.Kalur A., Iyer A. I., Muste J. C., et al. , “Impact of retinal fluid in patients with diabetic macular edema treated with anti-VEGF in routine clinical practice,” Can. J. Ophthalmol. 58(4), 271–277 (2023). 10.1016/j.jcjo.2022.03.003 [DOI] [PubMed] [Google Scholar]
- 9.Flores-Bellver M., Mighty J., Aparicio-Domingo S., et al. , “Extracellular vesicles released by human retinal pigment epithelium mediate increased polarised secretion of drusen proteins in response to AMD stressors,” J. Extracell. Vesicles 10(13), e12165 (2021). 10.1002/jev2.12165 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Wang M., Zhu W., Shi F., et al. , “MsTGANet: Automatic Drusen Segmentation From Retinal OCT Images,” IEEE Trans. Med. Imaging 41(2), 394–406 (2022). 10.1109/TMI.2021.3112716 [DOI] [PubMed] [Google Scholar]
- 11.Ni G., Wu R., Zheng F., et al. , “Toward ground-truth optical coherence tomography via three-dimensional unsupervised deep learning processing and data,” IEEE Trans. Med. Imaging 43(6), 2395–2407 (2024). 10.1109/TMI.2024.3363416 [DOI] [PubMed] [Google Scholar]
- 12.Ni G., Chen Y., Wu R., et al. , “Sm-Net OCT: a deep-learning-based speckle-modulating optical coherence tomography,” Opt. Express 29(16), 25511–25523 (2021). 10.1364/OE.431475 [DOI] [PubMed] [Google Scholar]
- 13.Ni G., Zhang J., Liu L., et al. , “Detection and compensation of dispersion mismatch for frequency-domain optical coherence tomography based on A-scan’s spectrogram,” Opt. Express 28(13), 19229–19241 (2020). 10.1364/OE.393870 [DOI] [PubMed] [Google Scholar]
- 14.Tan X., Chen X., Meng Q., et al. , “OCT2Former: A retinal OCT-angiography vessel segmentation transformer,” Comput. Methods Programs Biomed. 233, 107454 (2023). 10.1016/j.cmpb.2023.107454 [DOI] [PubMed] [Google Scholar]
- 15.Hu K., Jiang S., Zhang Y., et al. , “Joint-Seg: Treat Foveal Avascular Zone and Retinal Vessel Segmentation in OCTA Images as a Joint Task,” IEEE Trans. Instrum. Meas. 71, 1–13 (2022). 10.1109/TIM.2022.3193188 [DOI] [Google Scholar]
- 16.Wang X., Tang F., Chen H., et al. , “Deep semi-supervised multiple instance learning with self-correction for DME classification from OCT images,” Med. Image Anal. 83, 102673 (2023). 10.1016/j.media.2022.102673 [DOI] [PubMed] [Google Scholar]
- 17.Shen D., Wu G., Suk H.-I., “Deep learning in medical image analysis,” Annu. Rev. Biomed. Eng. 19(1), 221–248 (2017). 10.1146/annurev-bioeng-071516-044442 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Litjens G., Kooi T., Bejnordi B. E., et al. , “A survey on deep learning in medical image analysis,” Med. Image Anal. 42, 60–88 (2017). 10.1016/j.media.2017.07.005 [DOI] [PubMed] [Google Scholar]
- 19.Roy A. G., Conjeti S., Karri S. P. K., et al. , “ReLayNet: retinal layer and fluid segmentation of macular optical coherence tomography using fully convolutional networks,” Biomed. Opt. Express 8(8), 3627–3642 (2017). 10.1364/BOE.8.003627 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Sappa L. B., Okuwobi I. P., Li M., et al. , “RetFluidNet: Retinal fluid segmentation for SD-OCT images using convolutional neural network,” J. Digital Imaging 34(3), 691–704 (2021). 10.1007/s10278-021-00459-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Teja R. V., Manne S. R., Goud A., et al. , “Classification and quantification of retinal cysts in OCT B-scans: Efficacy of machine learning methods,” in 41st Ann. Int. Conf. IEEE Eng. Med. Bio. Soc. (EMBC), (IEEE, 2019), pp. 48–51. [DOI] [PubMed] [Google Scholar]
- 22.Ronneberger O., Fischer P., Brox T., “U-net: Convolutional networks for biomedical image segmentation,” in Proc. MICCAI , (Springer, 2015), pp. 234–241. [Google Scholar]
- 23.Badrinarayanan V., Kendall A., Cipolla R. J. I. T. O. P. A., “Segnet: A deep convolutional encoder-decoder architecture for image segmentation,” IEEE Trans. Pattern Anal. Mach. Intell. 39(12), 2481–2495 (2017). 10.1109/TPAMI.2016.2644615 [DOI] [PubMed] [Google Scholar]
- 24.Chen L.-C., Zhu Y., Papandreou G., et al. , “Encoder-decoder with atrous separable convolution for semantic image segmentation,” in Proc. Euro. conf. computer vision (ECCV), (2018), pp. 801–818. [Google Scholar]
- 25.Zhou Z., Siddiquee M. M. R., Tajbakhsh N., et al. , “Unet++: Redesigning skip connections to exploit multiscale features in image segmentation,” IEEE Trans. Med. Imaging 39(6), 1856–1867 (2020). 10.1109/TMI.2019.2959609 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Perdomo O., Otálora S., González F. A., et al. , “Oct-net: A convolutional network for automatic classification of normal and diabetic macular edema using sd-oct volumes,” in IEEE 15th int. sym. biom. imag. (ISBI) , (IEEE, 2018), pp. 1423–1426. [Google Scholar]
- 27.Yu X., Li M., Ge C., et al. , “Loss-balanced parallel decoding network for retinal fluid segmentation in OCT,” Comput. Biol. Med. 165, 107319 (2023). 10.1016/j.compbiomed.2023.107319 [DOI] [PubMed] [Google Scholar]
- 28.Xing G., Chen L., Wang H., et al. , “Multi-Scale Pathological Fluid Segmentation in OCT With a Novel Curvature Loss in Convolutional Neural Network,” IEEE Transactions on Medical Imaging 41(6), 1547–1559 (2022). 10.1109/TMI.2022.3142048 [DOI] [PubMed] [Google Scholar]
- 29.Suchetha M., Ganesh N. S., Raman R., et al. , “Region of interest-based predictive algorithm for subretinal hemorrhage detection using faster R-CNN,” Soft Computing 25(24), 15255–15268 (2021). 10.1007/s00500-021-06098-1 [DOI] [PMC free article] [PubMed] [Google Scholar] [Retracted]
- 30.Liu W., Sun Y., Ji Q., “Mdan-unet: multi-scale and dual attention enhanced nested u-net architecture for segmentation of optical coherence tomography images,” Algorithms 13(3), 60 (2020). 10.3390/a13030060 [DOI] [Google Scholar]
- 31.Liu X., Wang S., Zhang Y., et al. , “Automatic fluid segmentation in retinal optical coherence tomography images using attention based deep learning,” Neurocomputing 452, 576–591 (2021). 10.1016/j.neucom.2020.07.143 [DOI] [Google Scholar]
- 32.Rasti R., Biglari A., Rezapourian M., et al. , “RetiFluidNet: A Self-Adaptive and Multi-Attention Deep Convolutional Network for Retinal OCT Fluid Segmentation,” IEEE Trans. Med. Imaging 42(5), 1413–1423 (2023). 10.1109/TMI.2022.3228285 [DOI] [PubMed] [Google Scholar]
- 33.Moradi M., Chen Y., Du X., et al. , “Deep ensemble learning for automated non-advanced AMD classification using optimized retinal layer segmentation and SD-OCT scans,” Comput. Biol. Med. 154, 106512 (2023). 10.1016/j.compbiomed.2022.106512 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Li M., Shen Y., Wu R., et al. , “High-accuracy 3D segmentation of wet age-related macular degeneration via multi-scale and cross-channel feature extraction and channel attention,” Biomed. Opt. Express 15(2), 1115–1131 (2024). 10.1364/BOE.513619 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.He Y., Carass A., Liu Y., et al. , “Longitudinal deep network for consistent OCT layer segmentation,” Biomed. Opt. Express 14(5), 1874–1893 (2023). 10.1364/BOE.487518 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Mukherjee S., De Silva T., Grisso P., et al. , “Retinal layer segmentation in optical coherence tomography (OCT) using a 3D deep-convolutional regression network for patients with age-related macular degeneration,” Biomed. Opt. Express 13(6), 3195–3210 (2022). 10.1364/BOE.450193 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Mukherjee S., De Silva T., Jayakar G., et al. , “Retinal layer segmentation for age-related macular degeneration patients with 3D-UNet,” in Medical Imaging 2022: Computer-Aided Diagnosis , (SPIE, 2022), 916–921. [Google Scholar]
- 38.Yan Q., Wang B., Zhang W., et al. , “Attention-guided deep neural network with multi-scale feature fusion for liver vessel segmentation,” IEEE J. Biomed. Health Inform. 25(7), 2629–2642 (2021). 10.1109/JBHI.2020.3042069 [DOI] [PubMed] [Google Scholar]
- 39.Dong C., Xu S., Dai D., et al. , “A novel multi-attention, multi-scale 3D deep network for coronary artery segmentation,” Med. Image Anal. 85, 102745 (2023). 10.1016/j.media.2023.102745 [DOI] [PubMed] [Google Scholar]
- 40.Milletari F., Navab N., Ahmadi S. A., “V-Net: Fully convolutional neural networks for volumetric medical image segmentation,” in Proce. 4th Int. Conf. 3D Vision, 565–571 (2016). [Google Scholar]
- 41.Salehi S. S. M., Erdogmus D., Gholipour A., “Tversky loss function for image segmentation using 3D fully convolutional deep networks,” in Int. Workshop on Machine Learning in Medical Imaging , (Springer, 2017), pp. 379–387. [Google Scholar]
- 42.Maggio E., Sartore M., Attanasio M., et al. , “Anti–Vascular Endothelial Growth Factor Treatment for Diabetic Macular Edema in a Real-World Clinical Setting,” Am. J. Ophthalmol. 195, 209–222 (2018). 10.1016/j.ajo.2018.08.004 [DOI] [PubMed] [Google Scholar]
- 43.Diakogiannis F. I., Waldner F., Caccetta P., et al. , “ResUNet-a: A deep learning framework for semantic segmentation of remotely sensed data,” ISPRS J. Photo. Remote Sensing 162, 94–114 (2020). 10.1016/j.isprsjprs.2020.01.013 [DOI] [Google Scholar]
- 44.Oktay O., Schlemper J., Folgoc L. L., et al. , “Attention u-net: Learning where to look for the pancreas,” arXiv (2018). 10.48550/arXiv.1804.03999 [DOI]
- 45.Gu Z., Cheng J., Fu H., et al. , “Ce-net: Context encoder network for 2d medical image segmentation,” IEEE Trans. Med. Imaging 38(10), 2281–2292 (2019). 10.1109/TMI.2019.2903562 [DOI] [PubMed] [Google Scholar]
- 46.Çiçek Ö., Abdulkadir A., Lienkamp S. S., et al. , “3D U-Net: learning dense volumetric segmentation from sparse annotation,” in Medical Image Computing and Computer-Assisted Intervention–MICCAI 2016: 19th International Conference, Athens, Greece, October 17-21, 2016, Proceedings, Part II 19, (Springer, 2016), 424–432. [Google Scholar]
- 47.Liu Z., Lin Y., Cao Y., et al. , “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021), 10012–10022. [Google Scholar]
- 48.Cao H., Wang Y., Chen J., et al. , “Swin-unet: Unet-like pure transformer for medical image segmentation,” in Proc. Euro. conf. computer vision (ECCV), (Springer, 2022), pp. 205–218. [Google Scholar]
- 49.Bogunović H., Venhuizen F., Klimscha S., et al. , “RETOUCH: The retinal OCT fluid detection and segmentation benchmark and challenge,” IEEE transactions on medical imaging 38(8), 1858–1874 (2019). 10.1109/TMI.2019.2901398 [DOI] [PubMed] [Google Scholar]
- 50.He K., Zhang X., Ren S., et al. , “Deep residual learning for image recognition,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), (2016), pp. 770–778. [Google Scholar]
- 51.Murakami T., Nishijima K., Akagi T., et al. , “Optical coherence tomographic reflectivity of photoreceptors beneath cystoid spaces in diabetic macular edema,” Invest. Ophthalmol. Vis. Sci. 53(3), 1506–1511 (2012). 10.1167/iovs.11-9231 [DOI] [PubMed] [Google Scholar]
- 52.Massin P., Bandello F., Garweg J. G., et al. , “Safety and efficacy of ranibizumab in diabetic macular edema (RESOLVE Study): a 12-month, randomized, controlled, double-masked, multicenter phase II study,” Diabetes Care 33(11), 2399–2405 (2010). 10.2337/dc10-0493 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
Data underlying the results presented in this paper are not publicly available at this time but may be obtained from the authors upon reasonable request.











