Skip to main content
Ophthalmology Science logoLink to Ophthalmology Science
. 2025 Jun 16;5(6):100852. doi: 10.1016/j.xops.2025.100852

Automated Segmentation of Subretinal Fluid from OCT: A Vision Transformer Approach with Cross-Validation

Julie Midroni 1,2, Jack Longwell 2, Nishaant Bhambra 3, Sueellen Demian 4,5, Aurora Pecaku 4,5, Isabela Martins Melo 4,5, Rajeev H Muni 4,5,∗
PMCID: PMC12329092  PMID: 40778358

Abstract

Purpose

We present an algorithm to segment subretinal fluid (SRF) on individual B-scan slices in patients with rhegmatogenous retinal detachment (RRD). Particular attention is paid to robustness, with a fivefold cross-validation approach and a hold-out test set.

Design

Retrospective, cross-sectional study.

Participants

A total of 3819 B-scan slices across 98 time points from 45 patients were used in this study.

Methods

Subretinal fluid was segmented on all scans. A base SegFormer model, pretrained on 4 massive data sets, was further trained on raw B-scans from the retinal OCT fluid challenge data set of 4532 slices: an open data set of intraretinal fluid, SRF, and pigment epithelium detachment. When adequate performance was reached, transfer learning was used to train the model on our in-house data set, to segment SRF by generating a pixel-wise mask of presence/absence of SRF. A fivefold cross-validation approach was used, with an additional hold-out test set. All folds were first trained and cross-validated and then additionally tested on the hold-out set. Mean (averaged across images) and total (summed across all pixels, irrespective of image) Dice coefficients were calculated for each fold.

Main Outcome Measures

Subretinal fluid volume after surgical intervention for RRD.

Results

The average total Dice coefficient across the validation folds was 0.92, the average mean Dice coefficient was 0.82, and the median Dice was 0.92. For the test set, the average total Dice coefficient was 0.94, the average mean Dice coefficient was 0.82, and the median Dice was 0.92. The model showed strong interfold consistency on the hold-out set, with a standard deviation of only 0.03.

Conclusions

The SegFormer model for SRF segmentation demonstrates a strong ability to segment SRF. This result holds up to cross-validation and hold-out testing, across all folds. The model is available open-source online.

Financial Disclosure(s)

Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.

Keywords: Deep learning, Machine learning, Retinal detachment, Segmentation, Subretinal fluid


Rhegmatogenous retinal detachment (RRD) is often an acute emergent condition of the eye, whereby the retina tears and becomes separated from the retinal pigment epithelium, leading to vision loss.1 Treatment for RRD often results in a gross reattachment of the retina, but in many cases, there could be residual subretinal fluid (SRF), which could impact postoperative visual acuity if it is located under the fovea.1 This residual SRF most often gradually resorbs sometimes over several months.1

As a complete retinal reattachment is the primary end point of treatment after RRD, serial OCT can be used to assess improvement in SRF over time.2 Currently, residual SRF—as seen in Figure 1—is assessed qualitatively by ophthalmologists on clinical examination and with OCT. This approach, although effective, is limiting in that it prevents quantitative assessment and trending of SRF, resulting in a reduced understanding of a patient’s course, for both clinical and research purposes. Manual quantification of SRF in RRD is time-consuming, but there is an opportunity for automated methods to fill this gap, by segmenting out and calculating the volume of SRF using a computer algorithm. There has recently been increased interest in deep learning-based approaches to segment SRF in patients with RRD. Currently, many such algorithms have been published with performance metrics varying from 0.6 to 0.95.3, 4, 5, 6, 7 However, although these results are promising, not all studies perform additional robustness testing, such as hold-out testing and cross-validation, to assess robustness and generalizability.3,4,6

Figure 1.

Figure 1

Example residual subretinal fluid after retinal detachment repair. Note the 3 fluid blebs underneath the retina. Such fluid is the target of the segmentation task described in this work.

To that end, we present an algorithm to automatically segment SRF from OCT B-scans, using a transformer-based deep learning approach on an original in-house data set of several thousand images. We prioritize robustness, with both hold-out testing and fivefold cross-validation, and maintain strong performance, greatly increasing the strength of our results. We use a model architecture not yet described in SRF segmentation literature. Importantly, in the interest of open science, we release both our model code and final trained weights. This adds value to our work not only for its new data set and different architecture but also for the fact that it is a robust model trained on an original data set, which can therefore be used as a stepping stone or foundation on which iteratively better work can be based, using either other in-house data sets or the common data sets often used and available online.

Methods

Data Set

This study adhered to the Declaration of Helsinki, with informed consent of all participants. Institutional review board approval was obtained at St. Michael's Hospital, Toronto, Ontario, Canada, for all data as described as follows. Two main data sets were used in this project: retinal OCT fluid challenge (ReTOUCH)8 and our in-house data set. ReTOUCH is an open data set of 4532 OCT B-scans, raw and unprocessed, collected from patients with macular edema.8 The use of ReTOUCH was approved by the ReTOUCH institutional review board. Retinal OCT fluid challenge images have been annotated for SRF, intraretinal fluid (IRF), and pigment epithelium detachment. Because these images are raw, of variable size, and can only be extracted from machines with a research license, we obtained and used them for transfer learning, but without the intention of using them for validation or testing, for which we elected to use images compatible with our workflow and machine setup. They were split 80:20 for train/validation, to tune the model for this stage.

Our in-house data set was gathered retrospectively from a subset of the Pneumatic Retinopexy versus Vitrectomy for the Management of Primary Rhegmatogenous Retinal Detachment Outcomes Randomized Trial (PIVOT) trial,9 an interventional clinical trial comparing pneumatic retinopexy versus pars plana vitrectomy for RRD, with approval for the use of these data from the PIVOT institutional review board. This population differs from that of ReTOUCH based on primary pathology (RRD as opposed to macular edema in ReTOUCH8) as well as the postoperative nature of our data set. We reviewed PIVOT participants and selected a subset of patients who met the criteria of adequate image quality (retina visualized, minimal-medium shadowing only). We intentionally selected patients to include those with and without SRF, to train the model to also identify when fluid is absent. We ultimately selected 3819 2D images (slices) from 98 time points, across 45 patients. The Cirrus software-processed scans were manually extracted from the Cirrus OCT viewer software. Importantly, these images were preprocessed by the Cirrus proprietary software, unlike the ReTOUCH images which are raw8 and often protected by a research license that makes them challenging to extract. J.M., a student, was trained by S.D., an ophthalmology fellow, to identify and annotate SRF using a contour-based method, and annotated the SRF on all images, wherever present, whenever present, resulting in an equal number of binary yes or no SRF masks. Each image had a dimension of 844 × 563 pixels. Pixel width was 5.35 μm in the axial direction and 10.6 μm in the lateral direction.

The data were then train–test split in a 66:34 ratio. This irregular division of data was employed, as opposed to a classic 70:30, because the train–test split was divided based on randomizing patients, rather than individual scan slices, to prevent any data leakage between the training and testing sets. The train set was then split 5 ways for fivefold cross-validation, creating 5 train/validation sets. Again, this 5-way split was done by randomizing patients, rather than individual scan slices, to ensure that there was absolutely no data leakage between validation folds. We verified after all data set splits that there was no data leakage. In terms of the randomization method used, we employed stratified random sampling based on the presence/absence of SRF in each patient but quality-checked each fold after randomization to ensure that they were approximately the same size because for some patients there were only a few images and for some there were hundreds. For each of the validation folds, the train fraction was augmented using Gaussian jitter and mirroring along the y-axis, to maximize data utility and decrease overfitting.

Model Design and Training

SegFormer, a semantic segmentation deep learning model,10 was chosen as the appropriate architecture for this task. The SegFormer is a lightweight transformer capable of achieving state-of-the-art accuracy on segmentation compared with comparably sized models.10 Its strong performance on other tasks, while remaining lightweight and able to be run on a laptop, makes it appropriate and convenient for medical imaging and the design of future clinical/research tools. When deciding to use a transformer as opposed to a convolutional neural network (CNN), we considered specifically that SegFormer has dramatically fewer parameters than comparably performing CNN-based methods, while still maintaining strong performance, which was important for our computational limitations.10 As well, there is significant emerging evidence that even on small medical imaging data sets, transformers and segmentation models with transformer backbones are able to achieve state-of-the-art performance, outperforming CNNs.11, 12, 13

As per industry standards for training deep learning models,14 we selected a SegFormer architecture that has been pretrained on general image segmentation classification and segmentation data sets to enhance model stability before training on our specific task.10,15

The ReTOUCH data set was used to train the SegFormer to segment SRF pixel-wise, until adequate, nonstochastic performance was attained based on validation loss stabilization and Dice score stabilization. We did not hyperparameter tuning on the ReTOUCH data set because it was not the focus of this study. The ReTOUCH-trained model was then independently trained on each of the 5 cross-validation folds from our Cirrus-processed, in-house data set. The output of the model was once again a pixel-wise mask of the presence/absence of SRF in the input image. The model and training hyperparameters were configured via manual hyperparameter tuning. An initial learning rate of 6e to 6 with gradient accumulation of size 2 and AdamW optimization16 was used, with combined Dice-binary cross-entropy loss.17 Training was terminated when each k-fold model performance plateaued based on Dice score stabilization and loss stabilization, and each k-fold model was then validated on its respective validation fold. Finally, all 5 k-fold models were tested on the same hold-out test set.

Statistics

Because the output of the model is a pixel-wise classification of yes or no SRF, the appropriate metric to measure model performance is the Dice coefficient.18 The Dice coefficient is calculated as truepositivestruepositives+0.5(falsepositives+falsenegatives) and ranges from 0 to 1, with 1 representing perfect segmentation compared with the ground truth mask.18 For 3-dimensional images, such as OCT scans, the Dice coefficient can be calculated in 2 main ways: total and mean. The total method score treats all slices and patients in the test or validation set as a common pool of pixels and calculates the total precision across all pixels of all slices, as well as the total recall across all pixels of all slices, and then applies the Dice formula.19 The mean method calculates the individual Dice coefficient for each slice and then averages them across the entire test or validation set.19 Both methods are commonly reported for medical segmentation, with the total method being generally preferred for class-imbalanced data sets, such as in the case of SRF segmentation.19 For completeness, we calculate both methods for each of our validation folds. We also present the median Dice because our results were heavily skewed. We also present the average, across all fivefolds, for each of our computed metrics. To compare individual-image performance on the hold-out test set between the 5 model folds, we present a histogram and box plot of the Dice coefficients on a per-image basis for all model folds on the hold-out test set. We also present the standard deviation of individual-image Dice coefficients, for the hold-out test set, as well as the Dice value at the 25th percentile. Finally, we conduct a Wilcoxon signed-rank test on the Dice scores for each individual image in the hold-out test set, for all pairs of model folds, to assess if the distribution of Dice scores for each model is statistically significantly different on the hold-out set.

To analyze model errors, we examined the individual model outputs, ground truth annotation, and original images for every slice that had a Dice score of <0.5. We classify these errors based on the relative size of the SRF, the type of error the model is making (false positives vs. false negatives), and the presence or absence of non-SRF pathology in the image.

Results

The results of the fivefold cross-validation can be seen in Table 1. The median Dice was 0.92, the average mean Dice score was 0.82, and the average total Dice score was 0.92, with a mean Dice standard deviation of 0.02 and 0.03 between folds, respectively. The results of the hold-out testing on each validation fold were similarly strong, with a median Dice score of 0.92, an average mean Dice score of 0.82, and an average total Dice score of 0.94, with between-fold Dice standard deviations of <0.01 for both mean metrics.

Table 1.

Performance of All Validation Folds on their Respective Validation Sets and the Common Hold-Out Test

Total Precision Total Recall Total Dice Average Precision Average Recall Average Dice Median Precision Median Recall Median Dice
Validation fold 1 0.92 0.94 0.93 0.84 0.80 0.82 0.88 0.87 0.86
Validation fold 2 0.93 0.94 0.9 0.84 0.86 0.83 0.92 0.92 0.95
Validation fold 3 0.94 0.94 0.94 0.83 0.81 0.86 0.93 0.91 0.95
Validation fold 4 0.84 0.91 0.87 0.77 0.85 0.80 0.90 0.90 0.92
Validation fold 5 0.93 0.90 0.92 0.81 0.82 0.80 0.92 0.88 0.89
Test fold 1 0.95 0.93 0.94 0.83 0.87 0.82 0.94 0.92 0.92
Test fold 2 0.95 0.94 0.94 0.82 0.88 0.80 0.93 0.93 0.91
Test fold 3 0.95 0.94 0.94 0.86 0.86 0.84 0.94 0.93 0.93
Test fold 4 0.95 0.93 0.95 0.83 0.88 0.83 0.94 0.92 0.92
Test fold 5 0.95 0.94 0.95 0.84 0.87 0.83 0.93 0.93 0.92

Results of hold-out testing and fivefold cross-validation: total/mean and median precision, recall, and Dice for all folds, on both sets.

Figure 2 shows example successful segmentations and failure cases from the hold-out test set, side-by-side with the corresponding manual annotation. Figure 2A depicts a correctly annotated large bleb, Figure 2B depicts a correctly annotated small bleb, and Figure 2C shows the model’s ability to classify SRF in the presence of other fluid anomalies, such as IRF.

Figure 2.

Figure 2

Example results from validation fold model 1 on the hold-out test set. A, Depicts a correctly annotated large bleb. B, Depicts a correctly annotated small bleb. C, Correct SRF segmentation in the presence of other fluid anomalies—specifically, IRF. D, Failed segmentation: false-negative errors. IRF = intraretinal fluid.

Figure 3 shows example annotations for each of the 5 cross-validated models on the hold-out test set, for the same bleb. Figure 3A to E show the annotations for each fold, and Figure 3F shows the common area of overlap for the fivefolds.

Figure 3.

Figure 3

Segmentations of the same hold-out test image from all 5 cross-validation folds. A–E, Segmentations of each of the folds. F, Overlaid segmentations, highlighting the large area of overlap.

Figure 4 shows the histograms of the per-image Dice coefficients for the fold 1 model on the hold-out test set. The distribution of Dice scores for the hold-out test set is strongly skewed to the right, with 25th quantile values of 0.82, 0.84, 0.84, and 0.84 for the fivefolds. Given the strong skew, it seems a small number of low coefficients that may be dragging down the mean average Dice score and increasing the standard deviation. All other folds showed very similar results, with standard deviations of 0.27, 0.27, 0.24, 0.26, and 0.26 for the fivefolds on the hold-out test set.

Figure 4.

Figure 4

Histograms and box plots showing the distribution of Dice coefficients for all models on the hold-out test set. Circles in the box plots represent outliers. Note the strong skew of all 5 distributions.

For the same hold-out test image, the standard deviation between model folds was on average, 0.03. The results of the Wilcoxon signed-rank tests for each pair are seen in Table 2.

Table 2.

Statistics and P Values for the Wilcoxon Signed-Rank Test for Pairwise Dice Score Comparisons between All Model Folds, on the Hold-Out Test Set

Model Fold 1 Model Fold 2 Model Fold 3 Model Fold 4 Model Fold 5
Model fold 1 - W = 243 662
P = 3e–19
W = 261 150
P = 0.0001
W = 253 990
P = 6e–5
W = 257 594.5
P = 0.0001
Model fold 2 - - W = 276 133
P = 2e–13
W = 209 746
P = 4e–35
W = 274 442
P = 2e–13
Model fold 3 - - - W = 228 025.5
P = 3e–9
W = 281 539
P = 0.78
Model fold 4 - - - W = 222 072
P = 1e–9
Model fold 5 - - - - -

Wilcoxon signed-rank test results for the comparison of model fold performance on the hold-out test set.

Table 3 shows the error analysis per-model fold on the hold-out test set. Under 10% of the images in the hold-out test set had a Dice score of <0.5. Of these, the vast majority were false-positive errors. Specifically, for all folds, the most common error was false-positive segmentation of retinal edema as SRF. This occurred most often in images with no SRF. False-negative segmentations, as well as false-positive segmentations of structures other than retinal edema, were much less common. Although a small proportion of the low-Dice images had IRF, misclassification of IRF as SRF was very rare, occurring in only 3 images in model fold 4, and no images in the other threefolds. The average error size was around 120 pixels for each fold, which is equivalent to approximately an 11 × 11 square or a circle with a radius of 6 pixels; these are very small errors because the total pixel area of an image is 475 172 pixels, and many blebs in our data set were well over 1000 pixels large.

Table 3.

Error Analysis Per-Model Fold for the Hold-Out Test Set

Number of Slices with Significant Segmentation Defect (Dice <0.5) False Positives False Positives with Edema False Negatives Concurrent IRF Concurrent IRF Misclassification Average Error (Pixels) Errors <20 Pixels
Test fold 1 122 110 97 20 12 0 125 24
Test fold 2 116 104 98 17 19 0 139 19
Test fold 3 94 69 63 29 15 0 120 17
Test fold 4 112 100 88 19 18 3 127 17
Test fold 5 107 90 83 20 16 0 115 17

Error analysis per-model fold. “False positives” and “false negatives” columns describe the number of slices where there was a major contributing error of that kind. “False positives with edema” describes false-positive errors in which there was significant retinal edema without subretinal fluid. “Concurrent IRF” describes images with IRF, and “concurrent IRF misclassification” is the subset of those images with IRF in which the error was due to the classification of IRF as subretinal fluid.

IRF = intraretinal fluid.

The 5 model folds are available at https://github.com/jui434/Automatic-Segmentation-of-SRF and can be run using the inference code provided.

Discussion

The goal of this article was to provide a high-performing, open-source model for SRF segmentation that has been trained on a new data set. Our work demonstrates that the segmentation of SRF using a fine-tuned SegFormer model results in strong, competitive performance and holds up to hold-out testing, as well as to cross-validation. This is a significant result because generalizability is an important step in clinical and research implementation. It is much easier to fine-tune a model’s hyperparameters to perform well on 1 validation fold than multiple. The use of a different architecture and data set enhances the novelty of our results, and the provision of our code and model checkpoints is a key part of our contribution to open science.

In the literature, the best-performing models attain Dice scores in the mid-0.9s, and most are CNN-based models.4 It is significant that our model consistently attains total and median Dice scores in the mid-90s on the test set, competitive with previously published best-performing models, while using a fundamentally different model architecture, multiple-fold hold-out testing, and with low variance in total and average Dice between folds, as well as a high first quartile value within each fold for the hold-out test set, despite the strong skew in the distribution of hold-out test Dice scores for each model fold. This indicates that our model is robust and consistent in its high Dice. The balanced precision and recall also indicate that our model is not performing well by sacrificing false positives for false negatives or vice versa. The strong performance across all folds on the hold-out test set, with a very low standard average deviation of 0.03, indicates that our model is consistently strong on the hold-out test set. Granted, in Table 1, the results of the Wilcoxon signed-rank test indicate that this consistency is not perfect: 9 of 10 possible pairs of model folds have significantly different individual-image Dice distributions on the hold-out set. However, this does not undermine the consistently strong performance of each individual fold, which is true regardless of differences in Dice score distribution and clearly seen in the case of our model. The large values of the W statistics are not an indicator of poor performance in this case but, rather, reflect the large size of the data set and consistently high performance with low within-fold variance, which makes it such that small changes in performance between folds can significantly change the ranking order of scores. Figure 3 provides a good visual example of the qualitative consistency of our model between folds: the area of overlap is effectively equal to the area of the entire bleb.

The difference between mean and total Dice scores can be explained by examining Figure 4, and the 25th percentile quantile values for all fivefolds, which are all >0.8. All of this indicates that the vast majority of individual images have Dice coefficients >0.8. However, a small number of low-Dice coefficients seem to drag down the mean. Indeed, the median scores are much higher than the mean scores. This is an important result: it indicates that our model performs remarkably well the majority of the time and that a few outliers—not unexpected with such a large data set—pull the score down. The average Dice thus highlights that there are a small number of images the model struggles with significantly; as per our analysis, these are usually images with no blebs at all, in which the model misclassifies a small number of pixels as bleb and thus automatically gets a Dice of 0, or smaller blebs in which the edge of the bleb is a greater proportion of the entire bleb and thus edge errors drag down performance. We therefore interpret average Dice not as a measure of overall performance but as an indicator of the heavy skew of the data.

Speaking specifically to error types, our error analysis and Figure 4 show that most errors are 0-Dice errors, wherein the model mistakenly classified retinal edema as SRF, particularly in images where there was no SRF. These errors were mostly quite small, representing blebs of—on average—a radius of around 6 to 7 pixels only. It is reassuring that the model makes these small-scale errors, rather than misclassifying other types of fluid or making large segmentation errors. These are also challenging images to classify, because retinal edema often seems quite similar to SRF but without the fluid. Given that many of these errors occurred in images without any SRF, even a 1-pixel misclassification would result in a Dice score of 0, explaining the size of the 0-Dice bars in the histograms of Figure 4.

Our study is not, however, without limitations. Although we used 2 important tests to assess for robustness—hold-out testing and fivefold cross-validation—we did not validate on external data sets from other machines. It is possible that Zeiss machines are easier or more difficult to segment than others or that there are other differences in image acquisition that would limit the generalization of our algorithm to other machines. Similarly, although it was not feasible for this study, it would be beneficial to train multiple annotators, rather than just 1, because this would increase the clinical reliability of results. As well, our scans were collected at variable times postsurgery. Although this was important as the intention of our tool is to be used longitudinally postoperatively to monitor SRF resolution, we do recognize that the variable time postsurgery could influence our results. Unfortunately, we were unable to control for this in the statistical analysis because the majority of our patients had complete resolution of SRF by 3 months, and therefore, we were severely limited by data availability and could not conduct the appropriate analyses with adequate power to be included. In the future, a larger data set with SRF present at variable times postsurgery could help assess if time postsurgery influences model performance.

Conclusion

Ultimately, we demonstrate that the segmentation of SRF using deep learning holds up to robustness testing and obtains near-top performance, despite the use of multiple validation folds and a hold-out test set. Critically, we provide all of our code and model weights free online, for use in future research. The next steps should include obtaining data sets from other machines and institutions, to further train on and thus demonstrate generalizability.

Manuscript no.: XOPS-D-25-00060R2.

Footnotes

Disclosure(s):

The Article Publishing Charge (APC) for this article was paid by Unity Health Toronto.

All authors have completed and submitted the ICMJE disclosure form.

The authors made the following disclosures:

R.H.M.: Consultant – AbbVie, Alcon, Bausch+Lomb, Bayer, Novartis, Roche. Research Funding/Grants – AbbVie, Alcon, Bayer, Novartis, Roche; Stocks – Dragonfleye Therapeutics Corp.

Supported by the Silber TARGET Fund, Temerty Faculty of Medicine, University of Toronto, Toronto, Ontario, Canada.

Rajeev H. Muni, MD, MSc, an editorial board member of this journal, was recused from the peer-review process of this article and had no access to information regarding its peer-review.

The other authors have no proprietary or commercial interest in any materials discussed in this article.

HUMAN SUBJECTS: Human subjects were included in this study. Institutional review board approval was obtained at St. Michael's Hospital, Toronto, Ontario, Canada, for all data as described. All research adhered to the tenets of the Declaration of Helsinki. All participants provided informed consent.

No animal subjects were used in this study.

Author Contributions:

Conception and design: Midroni, Longwell, Bhambra, Muni

Data collection: Midroni, Longwell, Bhambra, Demian, Pecaku, Melo

Analysis and interpretation: Midroni, Longwell, Demian, Pecaku, Melo, Muni

Obtained funding: Supported by the Silber TARGET Fund, Temerty Faculty of Medicine, University of Toronto, Toronto, Ontario, Canada

Overall responsibility: Midroni, Longwell, Bhambra, Demian, Pecaku, Melo, Muni

References

  • 1.Blair K., Czyz C.N. StatPearls. Treasure Island, FL. StatPearls Publishing; 2022. Retinal detachment. [PubMed] [Google Scholar]
  • 2.Colucciello M. Serial “en face” optical coherence tomography imaging of slowly resorbing subretinal fluid after pneumatic retinopexy. Retin Cases Brief Rep. 2016;10:286–288. doi: 10.1097/ICB.0000000000000252. [DOI] [PubMed] [Google Scholar]
  • 3.Daanouni O., Cherradi B., Tmiri A. Automated end-to-end architecture for retinal layers and fluids segmentation on OCT B-scans. Multimed Tools Appl. 2024;84:14305–14328. [Google Scholar]
  • 4.Lin M., Bao G., Sang X., Wu Y. Recent advanced deep learning architectures for retinal fluid segmentation on optical coherence tomography images. Sensors (Basel) 2022;22:3055. doi: 10.3390/s22083055. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Akpinar M.H., Sengur A., Faust O., et al. Artificial intelligence in retinal screening using OCT images: a review of the last decade (2013–2023) Comput Methods Programs Biomed. 2024;254 doi: 10.1016/j.cmpb.2024.108253. [DOI] [PubMed] [Google Scholar]
  • 6.Ndipenoch N., Miron A., Li Y. Performance evaluation of retinal OCT fluid segmentation, detection, and generalization over variations of data sources. IEEE Access. 2024;12:31719–31735. [Google Scholar]
  • 7.Karn P.K., Abdulla W.H. Precision segmentation of subretinal fluids in OCT using multiscale attention-based U-net architecture. Bioengineering (Basel) 2024;11:1032. doi: 10.3390/bioengineering11101032. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Bogunovic H., Venhuizen F., Klimscha S., et al. RETOUCH: the retinal OCT fluid detection and segmentation benchmark and challenge. IEEE Trans Med Imaging. 2019;38:1858–1874. doi: 10.1109/TMI.2019.2901398. [DOI] [PubMed] [Google Scholar]
  • 9.Hillier R.J., Felfeli T., Berger A.R., et al. The pneumatic retinopexy versus vitrectomy for the management of primary rhegmatogenous retinal detachment outcomes randomized trial (PIVOT) Ophthalmology. 2019;126:531–539. doi: 10.1016/j.ophtha.2018.11.014. [DOI] [PubMed] [Google Scholar]
  • 10.Xie E., Wang W., Yu Z., et al. SegFormer: simple and efficient design for semantic segmentation with transformers. Adv Neural Inf Process Syst. 2021;34:12077–12090. [Google Scholar]
  • 11.Takahashi S., Sakaguchi Y., Kouno N., et al. Comparison of vision transformers and convolutional neural networks in medical image analysis: a systematic review. J Med Syst. 2024;48:84. doi: 10.1007/s10916-024-02105-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Chen J., Mei J., Li X., et al. TransUNet: rethinking the U-Net architecture design for medical image segmentation through the lens of transformers. Med Image Anal. 2024;97 doi: 10.1016/j.media.2024.103280. [DOI] [PubMed] [Google Scholar]
  • 13.Hatamizadeh A., Tang Y., Nath V., et al. 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) IEEE; Piscataway, NJ: 2022. UNETR: transformers for 3D medical image segmentation; pp. 1748–1758. [Google Scholar]
  • 14.Kim H.E., Cosa-Linan A., Santhanam N., et al. Transfer learning for medical image classification: a literature review. BMC Med Imaging. 2022;22:69. doi: 10.1186/s12880-022-00793-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Wolf T., Debut L., Sanh V., et al. Huggingface’s transformers: state-of-the-art natural language processing. arXiv preprint. 2019 arXiv:1910.03771. [Google Scholar]
  • 16.Loshchilov I., Hutter F. Decoupled weight decay regularization. arXiv preprint. 2017 arXiv:1711.05101. [Google Scholar]
  • 17.Wazir S., Fraz M.M. 2022 12th International Conference on Pattern Recognition Systems (ICPRS) IEEE; Piscataway, NJ: 2022. HistoSeg: quick attention with multi-loss function for multi-structure segmentation in digital histology images; pp. 1–7. [Google Scholar]
  • 18.Jain Y., Godwin L.L., Ju Y., et al. Segmentation of human functional tissue units in support of a Human Reference Atlas. Commun Biol. 2023;6:717. doi: 10.1038/s42003-023-04848-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Forman G., Scholz M. Apples-to-apples in cross-validation studies. SIGKDD Explor Newslett. 2010;12:49–57. [Google Scholar]

Articles from Ophthalmology Science are provided here courtesy of Elsevier

RESOURCES