Skip to main content
Translational Vision Science & Technology logoLink to Translational Vision Science & Technology
. 2026 Jun 25;15(6):32. doi: 10.1167/tvst.15.6.32

Automated Identification and Segmentation of Diabetic Macular Edema Subtypes Using Deep Learning

Ming Yan 1,*, Xianggui Zhang 1,*, Ruilong Li 1,2, Qin Ding 1, Ya Ye 1, Zhen Huang 1, Cong Chen 1, Wenjing Zhang 1, Lulu Tang 3, Yanping Song 1,2,
PMCID: PMC13359094  PMID: 42345636

Abstract

Purpose

The purpose of this study was to develop a deep learning algorithm capable of accurately classifying diabetic macular edema (DME) subtypes and segmenting the lesions in patients with diabetic retinopathy (DR) using structural optical coherence tomography (OCT) images.

Methods

We retrospectively collected 3120 DME OCT B-scan images from 823 eyes of patients with DME, acquired from 4 different spectral-domain and swept-source OCT devices (Topcon 3D-2000, Topcon Triton, BK400K UWF SS-OCT, and VG200 SS-OCT) to enhance device diversity and evaluate cross-device generalizability. An annotation team, consisting of two mid-career ophthalmologists and one senior retinal specialist, performed meticulous multi-label annotations, including DME subtype categories, detection bounding boxes, and pixel-level segmentation masks, to build the DME-Seg dataset. Based on this dataset, we fine-tuned the YOLO11x-Seg model for the detection and segmentation tasks.

Results

The fine-tuned model achieved promising performance on the DME-Seg dataset. For lesion detection, it attained an average mAP50(B) of 0.82, mAP50-95(B) of 0.56, and Dice coefficients of 0.82 ± 0.20 (95% confidence interval [CI] = 0.81–0.83). For segmentation, it achieved an mAP50(M) of 0.84, mAP50-95(M) of 0.54, and Dice coefficients of 0.79 ± 0.18 (95% CI = 0.78–0.80).

Conclusions

The constructed DME-Seg dataset and the validated model demonstrate promising performance in the automated detection and segmentation of DME subtypes, with encouraging cross-device generalization capability. This resource provides a foundation for advancing artificial intelligence (AI)-assisted diagnosis and personalized treatment planning for DME.

Translational Relevance

This automated quantification tool bridges the gap between AI research and clinical utility by assisting ophthalmologists in the precise diagnosis and treatment of DME.

Keywords: diabetic macular edema (DME), optical coherence tomography (OCT) imaging, instance segmentation, deep learning

Introduction

Diabetic macular edema (DME), one of the most prevalent and vision-threatening complications of diabetic retinopathy (DR), represents the leading cause of irreversible visual loss among the working-age population worldwide.1 Epidemiological studies indicate that the global prevalence of DR among patients with diabetes is 34.6%, whereas the prevalence of center-involving DME is 6.81%.2 In 2020, approximately 18.83 million adults were affected by clinically significant macular edema, a figure projected to rise to 28.61 million by 2045.3 Thus, early and precise screening, coupled with personalized interventions, are essential for optimizing visual outcomes. Due to its noninvasive nature, high resolution, and quantitative capabilities, optical coherence tomography (OCT) has become the gold standard imaging modality for DME diagnosis and monitoring.4,5 Based on distinct morphological features observed on OCT, DME is conventionally classified into three primary phenotypic subtypes: cystoid macular edema (CME), diffuse retinal thickening (DRT), and serous retinal detachment (SRD).68 These subtypes demonstrate substantial heterogeneity in underlying pathophysiological mechanisms, responses to pharmacologic therapies, and clinical prognoses.911 Therefore, accurate recognition and localization of DME subtypes are crucial for developing individualized treatment strategies, evaluating prognosis, and predicting therapeutic efficacy, as illustrated in Figure 1.

Figure 1.

Figure 1.

Schematic diagram of AI-assisted ophthalmic diagnosis.

Traditional DME diagnosis relies on ophthalmologists manually interpreting OCT images, a process that is time-consuming, subjective, and exhibits poor inter-observer consistency.12 Recently, artificial intelligence (AI) technologies have emerged as a promising solution for OCT image analysis, utilizing extensive labeled retinal data to train deep learning-based algorithms. These approaches enable: (1) lesion detection and quantification, where AI models automatically identify DME-related biomarkers like intraretinal fluid (IRF) and subretinal fluid (SRF), surpassing manual assessment efficiency13,14; (2) disease diagnosis and stratification, which integrates multi-lesion features across algorithms to achieve high sensitivity and specificity for DME detection15,16 and accurate subtype classification17; (3) functional prediction, with convolutional neural networks (CNNs) estimating visual acuity from OCT images to guide personalized treatment18; and (4) multimodal image fusion and AI-assisted prediction of treatment responses, further enhancing clinical utility.19

Despite these advances, existing studies still face several limitations: (1) constrained datasets, as high-quality, fine-grained annotations remain scarce due to limited labeling expertise and high costs15; (2) poor generalization with small samples, as algorithms trained on single-source datasets exhibit limited robustness to image quality variations (e.g. hemorrhage occlusion) and inter-device differences20; and (3) inadequate subtype-specific segmentation, with most models focusing on aggregate fluid volumes (e.g. total IRF/SRF) rather than precise delineation of CME, DRT, and SRD subtypes.1,21 Notably, DRT is particularly prone to algorithmic oversight owing to its diffuse, spongiform thickening, indistinct boundaries, and grayscale similarity to healthy retina.22 These challenges substantially restrict AI model applicability.

Among DME subtypes, DRT poses a particular challenge for automated segmentation. Unlike cystoid spaces or subretinal fluid, DRT presents as spongiform, low-contrast retinal thickening with indistinct boundaries that are difficult to distinguish from healthy retinal tissue, even for experienced graders. Nevertheless, precise delineation of DRT is clinically important, as its extent correlates with visual function and may influence treatment decisions.9,11 These challenges motivate the exploration of deep learning methods capable of capturing the subtle textural and morphological features associated with DRT.

To address these challenges, this study makes two primary contributions: (1) dataset construction: we developed DME-Seg, a high-quality, multi-device OCT dataset specifically designed for DME subtype analysis. Retrospectively collected from 4 mainstream OCT devices—Topcon 3D-2000, Topcon Triton swept-source DRI-OCT, BK400K UWF SS-OCT, and VG200 SS-OCT—at a tertiary ophthalmology center between February 2020 and May 2025, the dataset encompasses 3120 images from 823 eyes, capturing diverse imaging principles, resolutions, and disease stages across the DME spectrum. An annotation team comprising two junior ophthalmologists and one senior retinal specialist conducted double-blinded, multi-label annotations using X-AnyLabeling, providing each image with subtype class labels (CME, DRT, and SRD), bounding boxes, and pixel-level segmentation masks following standardized protocols. (2) Model development: leveraging the newly constructed DME-Seg dataset, we fine-tuned a YOLO11x-Seg model for instance segmentation, optimized for clinical-grade performance in automated DME subtype detection and segmentation.

This work provides cross-device evidence for precise DME diagnosis and demonstrates a promising AI model that may accelerate subtype identification, enable quantitative lesion assessment, and inform personalized treatment planning. However, prospective, multi-center clinical evaluation is necessary before clinical integration.

Methods

DME Data Curation

Data Collection

In this study, we collected 3120 OCT images of DME from 823 eyes, acquired from four OCT devices at the department of ophthalmology ([removed for peer review]), from February 2020 to May 2025. The devices included Topcon 3D-2000 (spectral-domain OCT), Topcon Triton (swept-source OCT), BK400K (ultra-widefield swept-source OCT), and VG200 (ultra-widefield swept-source OCT), with specifications including axial resolution (Axial Res.), transverse resolution (Trans. Res.), scan resolution (Scan Res.), and scan area detailed in Table 1.

Table 1.

OCT Devices and Image Acquisition Specifications

Source Device Images Axial Res. Trans. Res. Scan Res. and Area
Topcon 3D-2000 353 5 10 512 × 650 pixels, 6 × 6 mm
Topcon Triton 184 6.3∼8 ∼15 512 × 256 pixels, 6 × 6 mm
BK400K 1473 3.8 10 1536 × 1280 pixels, 24 × 20 mm
VG200 1110 3.8 10 1024 × 1024 pixels, 16 × 16 mm

Each image in this study refers to a single B-scan extracted from a fovea-centered radial scan protocol. Because different OCT devices were used, the number of B-scans per scan session varied by device: Topcon 3D-2000 and Topcon Triton acquired 12 B-scans per session, whereas BK400K and VG200 acquired 18 B-scans per session. All included B-scans passed through the foveal center. We did not analyze all B-scans from each scan session; instead, only B-scans exhibiting characteristic DME features (intraretinal fluid, subretinal fluid, or diffuse retinal thickening) were selected and confirmed by a senior ophthalmologist. The distribution of selected B-scans per eye was: median = 4, interquartile range (IQR) = 2 to 5, and range = 1 to 9. For eyes with multiple clinical visits, only the baseline visit (i.e. the first available OCT examination) was included to maintain data independence and avoid temporal correlation.

To ensure data reliability, we implemented rigorous quality control criteria: (1) signal strength: only scans with a manufacturer-provided signal strength index (SSI) ≥7 were included; scans degraded by significant media opacities (dense cataract, vitreous hemorrhage, or corneal opacity) that reduced signal quality below this threshold were excluded. (2) Artifacts and positioning: images with significant motion artifacts, blink artifacts, or those not centered on the fovea were excluded. (3) Shadow masking: scans in which vitreous hemorrhage or other media opacities obscured >30% of the central macular retinal structures were further excluded, ensuring that all analyzable B-scans clearly displayed the principal pathological features. Representative examples of excluded images (e.g. vitreous hemorrhage shadowing) are provided in Supplementary Figure S1.

All imaging procedures followed standardized protocols. The study was conducted in accordance with the principles of the Declaration of Helsinki and was approved by the Institutional Ethics Committee ([removed for peer review]). All patient images were anonymized prior to analysis.

Data Annotation

During the image annotation phase, each image was independently annotated by two mid-career ophthalmologists (each with >5 years of clinical experience). The annotation tasks involved identifying the three clinical subtypes of DME: DRT, CME, and SRD. For each subtype, the ophthalmologists recorded its presence or absence (“yes” or “no”), allowing each image to receive one or more subtype labels. A total of 3120 images were confirmed to contain target instances (i.e. lesion regions).

Furthermore, each annotator independently delineated the lesion regions to generate precise segmentation masks. All annotations were subsequently reviewed by a senior retinal specialist (>10 years of experience). Cases of disagreement—defined as instances where the two primary annotators produced inconsistent lesion boundary delineations at the segmentation level (no disagreements occurred at the subtype classification level)—were discussed jointly among all three annotators, with the senior specialist rendering the final annotation decision. This tiered annotation process ensured that both the DME subtype labels and the lesion boundaries were accurately and consistently established for each image.

To quantify annotation reliability, we randomly selected a representative subset of 100 B-scans (excluded from model training) and computed the inter-annotator Dice coefficient between the two primary annotators’ segmentation masks. The resulting Dice score of 0.9535 indicates excellent inter-annotator agreement, supporting the reliability of the ground truth annotations used in this study.

Ultimately, a total of 5623 instances across the three subtypes—CME, DRT, and SRD—were annotated with class labels, bounding boxes, and segmentation masks. Examples of labeled images are shown in Figure 2.

Figure 2.

Figure 2.

Example of labeled DME images. From left to right: Original OCT image, bounding boxes with category labels, and segmentation masks (red = CME, green = DRT, and blue = SRD). Best viewed in color.

Data Analysis

As detailed in Tables 2 and 3, the mask statistics reveal substantial variability: CME includes 3117 labeled masks with an average area of 10,172 pixels (range = 11–213,369); DRT has 1839 masks with an average area of 37,634 pixels (range = 323–184,364); and SRD comprises 667 masks with an average area of 5797 pixels (range = 88–68,444). Notably, DRT masks are considerably larger on average area. Table 2d shows the breakdown: 1535 images have one DME subtype with a single instance, 320 have one subtype with multiple instances, 1040 feature 2 subtypes, and 225 include all 3 (CME, DRT, and SRD). Co-occurrence analysis (Table 2c) indicates that 871 images display both CME and DRT, 440 have both CME and SRD, and 413 contain both DRT and SRD—revealing that approximately 37% of CME images also include DRT.

Table 2.

DME-Seg Dataset

graphic file with name tvst-15-6-32-fx001.jpg

a and b = mask distributions and c and d = DME image types.

Table 3.

Mask Statistics for Different Subtypes

Subtype Masks Avg Area, Pixel Min Area, Pixel Max Area, Pixel
CME 3,117 10,172 11 213,369
DRT 1,839 37,634 323 184,364
SRD 667 5,797 88 68,444

Compared with existing DME datasets,2124 our DME-seg dataset is sourced from four different OCT devices. It provides more annotated samples and richer multi-label annotations, including subtype categories, bounding boxes, and segmentation masks. This supports advanced medical imaging tasks like DME subtype classification, detection, and lesion region segmentation, as shown in Table 4.

Table 4.

Comparison With Other DME Datasets

Supported Vision Tasks
Dataset Images Devices Classification Detection Segmentation
Minarva and Murugeswari23 8364 One × ×
Vidal et al.22 356 Two ×
Kirik et al.24 695 One ×
De et al.21 262 One
DME-Seg (Ours) 3120 Four

DME Instance Segmentation

Introduction to You Only Look Once

The You Only Look Once (YOLO) series is a set of real-time object detection and segmentation models known for their efficiency and accuracy. Recent studies show that YOLO11 achieves a superior balance of speed and precision, with up to 22% fewer parameters than YOLOv8 while maintaining competitive mean average precision (mAP).25,26

YOLO11-Based DME Subtype Instance Segmentation

Task Definition. Instance segmentation of DME in OCT images includes detecting, classifying, and segmenting lesion regions into subtypes: CME, DRT, and SRD. This complex task demands concurrent localization, precise subtype classification, and pixel-level segmentation to delineate lesion boundaries. The analysis of lesion attributes, including type, size, and spatial distribution, is essential for informing clinical decisions, tracking disease progression, and tailoring treatment strategies. By leveraging AI, such as YOLO11-Seg,27 automated detection and segmentation streamline the diagnostic process, reduce time demands, alleviate ophthalmologist workload, and enhance accuracy, thereby supporting overburdened healthcare systems and improving patient outcomes through timely and reliable interventions.

Model Structure. The YOLO11x-Seg model, a segmentation variant of YOLO11 developed by Ultralytics,27 is tailored for instance segmentation of DME subtypes in this study. Based on the official training documentation, the model architecture consists of a backbone, neck, and head, optimized for precise lesion detection and segmentation in medical imaging, as shown in Figure 3.

Figure 3.

Figure 3.

YOLO11-seg based DME diagnosis. We constructed the DME OCT image dataset DME-Seg from four different OCT devices. Three ophthalmologists performed multi-dimensional label annotation on DME lesions, including subtype categories, bounding boxes, and segmentation masks.

YOLO11x-Seg uses a Cross Stage Partial (CSP)-based backbone with stacked C3k2 modules that use parallel smaller convolutions to efficiently extract features, followed by a Spatial Pyramid Pooling-Fast (SPPF) block for multi-scale context and a Cross Stage Partial with Spatial Attention (C2PSA) module to highlight salient lesion regions. The neck upsamples and fuses deep semantic features with shallow high-resolution details through C3k2 blocks and C2PSA attention, preserving fine-grained lesion boundaries. The YOLO ESegment head unifies detection and segmentation, performing bounding-box regression for DME subtype localization alongside high-resolution instance mask prediction for precise lesion delineation. With approximately 62 million parameters and 320 giga floating-point operations per second (GFLOPS), the model achieves an inference time of only 23.43 ms per image on a single NVIDIA A100 GPU, balancing accuracy and computational efficiency to support real-time clinical diagnosis of DME.

Loss Function. To ensure robust convergence during training, we used a composite loss function tailored for medical instance segmentation. The objective function integrates: (1) box loss (utilizing CIoU) to enforce accurate localization of lesion centers; (2) distribution focal loss (DFL) to refine the boundaries of ambiguous lesions; and (3) pixel-wise segmentation loss to ensure high fidelity in lesion shape reconstruction. This multi-objective optimization ensures that the model balances classification accuracy with morphologic precision.

Results

This section evaluates our novel dataset, DME-Seg, for DME subtype instance segmentation, believed to be the largest of its kind. We describe the dataset, evaluation metrics, and implementation details of a fine-tuned YOLO11x-Seg model. Following this, we analyze the detection and segmentation performance for CME, DRT, and SRD.

Experimental Setting

Datasets. Our dataset comprises 3120 human-segmented medical OCT images. All data splits were performed strictly at the eye level to prevent data leakage. The dataset was split at a ratio of 0.8:0.2, resulting in 2496 training images and 624 validation images. The training data was prepared according to the official YOLO documentation to ensure compatibility and optimal quality. Preprocessing involved annotation formatting and data augmentation tailored to the model's requirements. The original OCT images, sourced from various devices such as Topcon 3D-2000, Topcon Triton swept-source DRI-OCT, BK400K UWF SS-OCT, and VG200 SS-OCT, exhibited resolutions ranging from 512 to 1536 pixels. To improve compatibility with the YOLO model, all images were uniformly resized to 1024 × 1024 pixels.

Evaluation metrics. Model performance was evaluated using standard YOLO metrics complemented by the Dice coefficient (mean ± SD, 95% confidence interval [CI]) for detection and segmentation tasks. For detection (box-based), we report precision (B), recall(B), mean average precision at intersection over union (IoU) = 0.5 (mAP50 (B)), and mean average precision across IoU thresholds from 0.5 to 0.95 (mAP50-95 (B)) and the Dice coefficient. For instance, segmentation (mask-based), the metrics are precision (M), recall (M), mAP50 (M), and mAP50-95 (M) and the Dice coefficient, which was used to quantify the spatial overlap and pixel-level consistency between the predicted regions and manual annotations.

Implementation. We utilized the Ultralytics YOLO11x-Seg for instance segmentation tasks, initialized with the pre-trained YOLO11x-Seg checkpoint from the official Ultralytics repository (pre-trained on the COCO dataset).27 The training was conducted using 2 NVIDIA A800 GPUs, with a dropout rate of 0.2 applied to enhance model generalization. The loss function was composed of four components—box loss, classification loss (cls), DFL, and segmentation loss—with weights set at 7.5, 0.5, 1.5, and 1.0, respectively, to balance their contributions. The training process spanned 100 epochs, including 5 warmup epochs to stabilize initial learning, with a batch size of 128. All other hyperparameters were configured according to the default values specified in the Ultralytics official documentation. The overall training and validation curves are illustrated in Figure 4.

Figure 4.

Figure 4.

Training and validation curves.

Main Experiments

Detection Results. The experimental results demonstrate the fine-tuned YOLO11x-Seg model's effectiveness in detecting DME subtypes in OCT images. Performance metrics in Table 5 vary across CME, DRT, and SRD. SRD exhibits the strongest performance with higher precision and recall, due to its distinct imaging features and clearer instances. In contrast, CME and DRT show moderate outcomes, with lower mAP50-95 (B) values of approximately 0.51 and 0.44 and correspondingly lower Dice coefficients of 0.80 ± 0.23 (95% CI = 0.78–0.82) and 0.83 ± 0.15 (95% CI = 0.81–0.85), respectively, likely from overlapping features, smaller lesion sizes, or complex interactions that challenge precise localization across IoU thresholds. Bolded averages across subtypes are: precision (B) 0.81, recall (B) 0.78, mAP50 (B) 0.82, and mAP50-95 (B) 0.56 and Dice coefficient 0.82 ± 0.20 (95% CI = 0.81–0.83), indicating a solid foundation for DME lesion detection. Detection results, including bounding boxes and class labels, are visualized in Figure 5.

Table 5.

Detection Results

Subtype Instances Precision (B) Recall (B) mAP50 (B) mAP50-95 (B) Dice Coefficient (Mean ± SD, 95% CI)
CME 611 0.74 0.71 0.75 0.51 0.80 ± 0.23, 95% CI = 0.78–0.82
DRT 371 0.74 0.67 0.74 0.44 0.83 ± 0.15, 95% CI = 0.81–0.85
SRD 122 0.96 0.95 0.98 0.73 0.89 ± 0.12, 95% CI = 0.87–0.91
Average 0.81 0.78 0.82 0.56 0.82 ± 0.20, 95% CI = 0.810.83

Average results of DME subtypes are presented in bold.

Figure 5.

Figure 5.

Detection result examples. Green and yellow denote human-annotated (ground truth) and model-predicted bounding boxes with confidence scores indicating prediction certainty. Best viewed in color.

Segmentation Results. Segmentation performance is evaluated using: Precision (M), Recall(M), mAP50 (M), and mAP50-95 (M) and Dice coefficient (mean ± SD, 95% CI), detailed in Table 6, with average results highlighted in bold for clarity.

Table 6.

Segmentation Results

Subtype Instances Precision (M) Recall (M) mAP50 (M) mAP50-95 (M) Dice Coefficient (Mean ± SD, 95% CI)
CME 611 0.76 0.71 0.78 0.47 0.77 ± 0.20, 95% CI = 0.75–0.79
DRT 371 0.77 0.69 0.76 0.42 0.82 ± 0.13, 95% CI = 0.81–0.83
SRD 122 0.96 0.94 0.97 0.72 0.77 ± 0.19, 95% CI = 0.74–0.80
Average 0.83 0.78 0.84 0.54 0.79 ± 0.18, 95% CI = 0.78–0.80

Average results are presented in bold.

The results demonstrate consistent segmentation performance across all subtypes. Notably, SRD exhibits superior metrics (precision = 0.96, recall = 0.94, mAP50 = 0.97, mAP50-95 = 0.72, and Dice coefficient = 0.77 ± 0.19 (95% CI = 0.74–0.80), likely attributable to its distinct morphological characteristics. In contrast, CME and DRT yield slightly lower but comparable metrics, with mAP50 values of 0.78 and 0.76, respectively.

Although the average mAP50-95 (M) of 0.54 and average Dice coefficient of 0.79 ± 0.18 (95% CI = 0.78–0.80) appears modest, it competes well with other instance segmentation tasks (e.g. COCO's 0.44).27 Given the inherent variability and complexity of DME subtype lesion regions, our model's performance is encouraging and indicates a promising approach for this specific medical imaging challenge. Figure 6 illustrates the alignment between human-annotated and model-predicted masks, supporting the quantitative findings.

Figure 6.

Figure 6.

Segmentation result examples. Left to right: Original image, human-annotated masks, model-predicted masks with confidence scores. Best viewed in color.

Generalization Assessment

To evaluate the model's generalizability beyond the training data distribution, we collected an additional 330 independent OCT images from new patients who were entirely excluded from all stages of model development. These images were acquired from the same OCT devices but originated from different patients not represented in the training or validation sets. The generalization test results are presented in Table 7.

Table 7.

Generalization Test Results on 330 Independent Test Images (Dice, Mean ± SD)

Evaluation Task CME DRT SRD Overall
Detection 0.82 ± 0.18 0.66 ± 0.23 0.88 ± 0.09 0.80 ± 0.20
Segmentation 0.82 ± 0.13 0.66 ± 0.21 0.75 ± 0.17 0.77 ± 0.18

The model demonstrated robust performance on the entirely unseen independent test set. For the detection task, the overall Dice reached 0.80, with CME and SRD achieving scores of 0.82 and 0.88, respectively. For the segmentation task, the overall Dice was 0.77, with CME maintaining accuracy of 0.82. The DRT Dice scores were comparatively lower in both detection (0.66) and segmentation (0.66), consistent with the inherent clinical difficulty of delineating diffuse retinal thickening boundaries. Overall, these results support the model's generalization potential while underscoring the need for further multi-center validation.

Discussion

This study established instance segmentation dataset for DME subtypes (DME-Seg). Using it, we developed and validated a YOLO11x-Seg model for automated detection and segmentation of CME, DRT, and SRD in OCT images. Results show robust performance in both tasks, offering a reliable AI tool for automated DME analysis.

Model Performance Analysis

Our model achieved overall mAP50 values of 0.85 and 0.84 in the detection and segmentation tasks, respectively, establishing a solid foundation for the automated identification of DME lesions. Notably, the model exhibited discordant trends between the mAP50 and Dice coefficients across different subtypes. This phenomenon profoundly reflects the distinct focuses of these evaluation metrics, as well as the varying pathomorphological characteristics of the DME subtypes.

At the instance detection level, SRD demonstrated exceptionally high performance (segmentation mAP50 = 0.97). This is attributable to its distinct morphological features: SRD manifests as a clearly demarcated hyporeflective subneural fluid pocket with high contrast, enabling the model to accurately localize it with high confidence. This observation aligns with the report by Wu et al.,17 whose model also achieved a high area under the curve (AUC) of 0.994 for SRD detection. However, in our study, the pixel-level segmentation Dice coefficient for SRD was comparatively moderate (0.77 ± 0.19). This discrepancy indicates that although the model can precisely “lock onto” the main body of the SRD lesion, achieving perfect alignment at the pixel level remains challenging because subretinal fluid often tapers into a conical or wedge-shaped contour at its margins.

In contrast, although DRT yielded a relatively lower segmentation mAP50 (0.76), it achieved the highest segmentation Dice coefficient among the 3 subtypes (0.82 ± 0.13). This primarily stems from the sensitivity of the Dice coefficient to the spatial connectivity and area of the target (i.e. the volume/size effect). DRT presents as a large area, diffuse retinal thickening lacking clear demarcation from normal tissue, which inherently lowers the confidence (mAP) of instance detection. However, precisely because of its massive lesion area, the correctly predicted pixels in the core region dominate the mathematical calculation, thereby largely “diluting” the pixel-level errors caused by blurred boundaries.

The performance of CME (segmentation mAP50 = 0.78 and segmentation Dice = 0.77 ± 0.20) highlights the inherent difficulties in small-object segmentation. Manifesting as multiple, fragmented tiny cystic cavities, CME is highly susceptible to the “spatial penalty”—even a minimal misclassification of edge pixels can lead to a drastic drop in the overall Dice coefficient. This finding is consistent with the report by Moura et al.,21 who similarly found that the Dice score for CME (0.7475) significantly lagged behind that of DRT (0.8913). Integrating our analyses of the marginal morphology of SRD and the micro-cystic nature of CME, our findings highly corroborate the conclusions of Hsu et al.28 They noted that lesions characterized by internal structural disruption, such as intraretinal cysts (IRCs) and subretinal fluid (SRF), are generally more difficult to segment with high pixel-level precision. Our results further substantiate this, indicating that current deep learning models still face shared challenges when tackling segmentation tasks involving small, fragmented targets with overlapping features and indistinct boundaries (Fig. 7).

Figure 7.

Figure 7.

Failure cases in accurately detecting and segmenting DRT lesion regions.

Comparison of Related Work

Early studies, such as those by Minarva and Murugeswari23 and Gan et al.,29 primarily focused on image-level classification of DME (determining whether DME is present in an image). The high accuracy rates reported (97.4% and 93.8%) demonstrate the value of deep learning in screening applications. Subsequent work by Zhu et al.,30 Tang et al.,16 and Padmasini and Umamaheswari31 further refined classification into subtypes or severity levels, achieving excellent area under the receiver operating characteristic curve (AUROC; 0.881–0.975) and accuracy (0.993). However, classification tasks alone cannot provide quantitative information—such as the location, morphology, and volume of lesions—that is critical for clinical diagnosis and treatment. In recent years, research focus has gradually shifted toward pixel-level segmentation. For instance, Hu et al.32 concentrated on the segmentation of retinal fluid, and their DAA-UNet model achieved a notably high Dice coefficient of 91.6%, demonstrating the potential of dedicated network architectures in improving segmentation accuracy. Ye et al.33 further identified up to 7 specific DME-related lesion features, with remarkably high AUC values ranging from 0.887 to 0.999, highlighting the model's powerful capability in recognizing complex features. The work by Moura et al.21 is most closely aligned with our study, as they were the first to report Dice coefficients specifically for the three major subtypes—SRD, CME, and DRT—thereby providing a direct performance benchmark for our research. Our study extends this by reporting both detection and segmentation Dice coefficients (averaging 0.82 and 0.79, respectively), providing a more comprehensive quantitative assessment of model performance under multi-device scenarios.

The key distinction of our study lies in three main aspects. First, we are the first to explicitly formulate the task as instance segmentation for DME subtypes. This means our model is capable not only of delineating precise lesion contours (segmentation), but also of distinguishing multiple lesion instances belonging to different subtypes within the same image (identification)—a critical capability for assessing complex clinical cases with co-existing pathologies. Second, our DME-Seg dataset integrates images acquired from four mainstream OCT devices, encompassing variations in imaging principles and resolutions. Although this diversity may increase the learning difficulty in the short term, it may improve the model's robustness and generalization potential when processing multi-device images to a certain extent, making it better suited for real-world applications across different hospitals and devices. The studies by Vidal et al.22,34 previously highlighted the challenges of cross-device and cross-center validation; our dataset construction strategy represents a direct response to these challenges. Finally, we adopted the latest YOLO11x-Seg model. Compared with classical two-stage segmentation networks, such as U-Net28 or its variants,32 YOLO—as a single-stage model—typically offers faster inference speed while maintaining high accuracy. This provides a significant advantage for developing future real-time diagnostic assistance systems and integrating them into clinical workflows.

Limitation

This study has several limitations. First, the model was developed and internally validated using only a single-center dataset; thus, its generalizability and robustness require further evaluation through multi-center and diverse external datasets. Second, among DME subtypes, DRT posed a substantial segmentation challenge owing to its ill-defined and low-contrast boundaries. Conventional pixel-wise loss functions may be suboptimal for capturing such subtle transitions. Future research should investigate boundary-aware loss functions (e.g. boundary loss or Hausdorff distance loss) to enhance the precision of DRT delineation. Finally, the retrospective nature of the study and the exclusion of non-DME pathologies, such as age-related macular degeneration (AMD) and epiretinal membrane (ERM), mean the model's clinical applicability to other macular diseases remains unestablished. Consequently, prospective multi-center studies with diverse cohorts are imperative prior to formal clinical deployment.

Future work

Building on the findings of this study, we identify four key directions for future research to enhance clinical applicability. First, multimodal data integration—combining serial OCT with fundus fluorescein angiography (FFA), indocyanine green angiography (ICGA), electronic health records, and genomic data—could enable comprehensive disease progression modeling. Second, prospective multi-center validation across diverse geographic regions and OCT platforms is essential to establish true generalizability. Third, image quality stratification is needed to evaluate model performance on real-world, lower-quality clinical images beyond our controlled dataset. Finally, transitioning to dense cube scan protocols would enable three-dimensional volumetric analysis (e.g. mm³ per subtype), providing more precise structural assessments for longitudinal monitoring.

The DME-Seg dataset and associated source code will be made publicly available upon acceptance of this manuscript and will be hosted on GitHub or Huggingface. All data will undergo rigorous de-identification in accordance with medical data privacy regulations prior to release.

In summary, the high-quality DME-Seg dataset constructed in this study, along with the YOLO11x-Seg-based instance segmentation model, demonstrates promising performance in the automated identification and segmentation of DME subtypes. However, prospective, multi-center clinical evaluation is essential before these tools can be integrated into routine clinical practice. We believe this work represents a solid step toward developing computer-aided diagnostic tools that may ultimately contribute to precision medicine in DME.

Supplementary Material

Supplement 1
tvst-15-6-32_s001.pdf (243.6KB, pdf)

Acknowledgments

Supported by the Postdoctoral Scientific Research Foundation, General Hospital of Central Theater Command (20210517KY04 and the 2025BSH070); and General Program of National Natural Science Foundation of China (No. 82471110).

Disclosure: M. Yan, None; X. Zhang, None; R. Li, None; Q. Ding, None; Y. Ye, None; Z. Huang, None; C. Chen, None; W. Zhang, None; L. Tang, None; Y. Song, None

References

  • 1. Parravano M, Cennamo G, Di Antonio L, et al.. Multimodal imaging in diabetic retinopathy and macular edema: an update about biomarkers. Surv Ophthalmol. 2024; 69(6): 893–904. [DOI] [PubMed] [Google Scholar]
  • 2. Yau JW, Rogers SL, Kawasaki R, et al.. Global prevalence and major risk factors of diabetic retinopathy. Diabetes Care. 2012; 35(3): 556–564. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Teo ZL, Tham YC, Yu M, et al.. Global prevalence of diabetic retinopathy and projection of burden through 2045: systematic review and meta-analysis. Ophthalmology. 2021; 128(11): 1580–1591. [DOI] [PubMed] [Google Scholar]
  • 4. Viggiano P, Vujosevic S, Palumbo F, et al.. Optical coherence tomography biomarkers indicating visual enhancement in diabetic macular edema resolved through anti-VEGF therapy: oct biomarkers in resolved DME. Photodiagnosis Photodyn Ther. 2024; 46:104042. [DOI] [PubMed] [Google Scholar]
  • 5. Yanxia C, Xiongyi Y, Min F, et al.. Optical coherence tomography-based grading of diabetic macular edema is associated with systemic inflammatory indices and imaging biomarkers. Ophthalmic Res. 2024; 67(1): 96–106. [DOI] [PubMed] [Google Scholar]
  • 6. Hui V, Szeto S, Tang F, et al.. Optical coherence tomography classification systems for diabetic macular edema and their associations with visual outcome and treatment responses - an updated review. Asia Pac J Ophthalmol (Phila). 2022; 11(3): 247–257. [DOI] [PubMed] [Google Scholar]
  • 7. Panozzo G, Cicinelli MV, Augustin AJ, et al.. An optical coherence tomography-based grading of diabetic maculopathy proposed by an international expert panel: the European School for Advanced Studies in Ophthalmology Classification. Eur J Ophthalmol. 2020; 30(1): 8–18. [DOI] [PubMed] [Google Scholar]
  • 8. Otani T, Kishi S, Maruyama Y. Patterns of diabetic macular edema with optical coherence tomography. Am J Ophthalmol. 1999; 127(6): 688–693. [DOI] [PubMed] [Google Scholar]
  • 9. Hu Y, Wu Q, Liu B, et al.. Comparison of clinical outcomes of different components of diabetic macular edema on optical coherence tomography. Graefes Arch Clin Exp Ophthalmol. 2019; 257(12): 2613–2621. [DOI] [PubMed] [Google Scholar]
  • 10. Wang XN, Cai X, He S, et al.. Subfoveal choroidal thickness changes after intravitreal ranibizumab injections in different patterns of diabetic macular edema using a deep learning-based auto-segmentation. Int Ophthalmol. 2023; 43(12): 4399–4407. [DOI] [PubMed] [Google Scholar]
  • 11. Sharma S, Karki P, Joshi SN, Parajuli S.. Optical coherence tomography patterns of diabetic macular edema and treatment response to bevacizumab: a short-term study. Ther Adv Ophthalmol. 2022; 14: 960339881. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Shahriari MH, Sabbaghi H, Asadi F, Hosseini A, Khorrami Z.. Artificial intelligence in screening, diagnosis, and classification of diabetic macular edema: a systematic review. Surv Ophthalmol. 2023; 68(1): 42–53. [DOI] [PubMed] [Google Scholar]
  • 13. Song W, Kaakour AH, Kalur A, et al.. Performance of a machine-learning computational image analysis algorithm in retinal fluid quantification for patients with diabetic macular edema and retinal vein occlusions. Ophthalmic Surg Lasers Imaging Retina. 2022; 53(3): 123–131. [DOI] [PubMed] [Google Scholar]
  • 14. Lv B, Li S, Liu Y, et al.. Development and validation of an explainable artificial intelligence framework for macular disease diagnosis based on optical coherence tomography images. Retina. 2022; 42(3): 456–464. [DOI] [PubMed] [Google Scholar]
  • 15. Lam C, Wong YL, Tang Z, et al.. Performance of artificial intelligence in detecting diabetic macular edema from fundus photography and optical coherence tomography images: a systematic review and meta-analysis. Diabetes Care. 2024; 47(2): 304–319. [DOI] [PubMed] [Google Scholar]
  • 16. Tang F, Wang X, Ran AR, et al.. A multitask deep-learning system to classify diabetic macular edema for different optical coherence tomography devices: a multicenter analysis. Diabetes Care. 2021; 44(9): 2078–2088. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Wu Q, Zhang B, Hu Y, et al.. Detection of morphologic patterns of diabetic macular edema using a deep learning approach based on optical coherence tomography images. Retina. 2021; 41(5): 1110–1117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Tan TE, Ng YP, Calhoun C, et al.. Detection of center-involved diabetic macular edema with visual impairment using multimodal artificial intelligence algorithms. Ophthalmol Retina. 2025; 9: 955–963. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Meng Z, Chen Y, Li H, et al.. Machine learning and optical coherence tomography-derived radiomics analysis to predict persistent diabetic macular edema in patients undergoing anti-VEGF intravitreal therapy. J Transl Med. 2024; 22(1): 358. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Kulyabin M, Zhdanov A, Pershin A, et al.. Segment anything in optical coherence tomography: SAM 2 for volumetric segmentation of retinal biomarkers. Bioengineering (Basel). 2024; 11(9): 940. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. de Moura J, Samagaio G, Novo J, Almuina P, Fernández MI, Ortega M.. Joint diabetic macular edema segmentation and characterization in oct images. J Digit Imaging. 2020; 33(5): 1335–1351. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Vidal P, de Moura J, Novo J, Ortega M. Multivendor fully automatic uncertainty management approaches for the intuitive representation of DME fluid accumulations in OCT images. Med Biol Eng Comput. 2023; 61(5): 1209–1224. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Minarva DK, Murugeswari S. Enhanced hybrid feature extraction and selection based on OCT images for diabetic macular edema classification. J Innov Image Proc. 2025; 7(2): 315–332. [Google Scholar]
  • 24. Kırık F, Demirkıran B, Ekinci Aslanoğlu C, Koytak A, Özdemir H. Detection and classification of diabetic macular edema with a desktop-based code-free machine learning tool. Turk J Ophthalmol. 2023; 53(5): 301–306. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Hidayatullah P, Syakrani N, Sholahuddin MR, Gelar T, Tubagus R. YOLOv8 to YOLO11: a comprehensive architecture in-depth comparative review 2025. Semantic Scholar Preprint. Available at: https://api.semanticscholar.org/CorpusID:275820077.
  • 26. Jegham N, Koh CY, Abdelatti M, et al.. YOLO evolution: a comprehensive benchmark and architectural review of YOLOv12, YOLO11, and their previous versions. arXiv Preprint. 2024. Available at: 10.48550/arXiv.2411.00201. [DOI]
  • 27. Jocher G. Ultralytics YOLO11. 2024. Available at: Https://github.com/ultralytics/ultralytics.
  • 28. Hsu HY, Chou YB, Jheng YC, et al.. Automatic segmentation of retinal fluid and photoreceptor layer from optical coherence tomography images of diabetic macular edema patients using deep learning and associations with visual acuity. Biomedicines. 2022; 10(6): 1269. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Gan F, Wu FP, Zhong YL. Artificial intelligence method based on multi-feature fusion for automatic macular edema (ME) classification on spectral-domain optical coherence tomography (SD-OCT) images. Front Neurosci. 2023; 17: 1097291. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Zhu H, Ji J, Lin JW, et al.. Development and validation of a 3-D deep learning system for diabetic macular oedema classification on optical coherence tomography images. BMJ Open. 2025; 15(5): e099167. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Padmasini N, Umamaheswari R. Automated detection of multiple structural changes of diabetic macular oedema in SDOCT retinal images through transfer learning in CNNs. IET Image Processing. 2020; 14(16): 4067–4075. [Google Scholar]
  • 32. Hu T, Ding J, Liu Y, Zhang Y, Yang L. DAA-UNet: a dense connectivity and atrous spatial pyramid pooling attention UNet model for retinal optical coherence tomography fluid segmentation. IET Software. 2025; 2025(1): 6006074. [Google Scholar]
  • 33. Ye X, Qiu W, Tu L, et al.. Detection of diabetic macular oedema patterns with fine-grained image categorisation on optical coherence tomography. BMJ Open Ophthalmol. 2025; 10(1): e002037. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Vidal PFL, de Moura J, Díaz M, Novo J, Ortega M. Diabetic macular edema characterization and visualization using optical coherence tomography images. Appl Sci (Basel). 2020; 10(21): 1–23. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplement 1
tvst-15-6-32_s001.pdf (243.6KB, pdf)

Articles from Translational Vision Science & Technology are provided here courtesy of Association for Research in Vision and Ophthalmology

RESOURCES