Skip to main content
Ophthalmology Science logoLink to Ophthalmology Science
. 2025 Dec 18;6(3):101030. doi: 10.1016/j.xops.2025.101030

Federated Learning for Multi-Disease Ophthalmic Diagnostics Using OCT Angiography

Ahammed Sakir Nabil 1, Sina Gholami 1, Theodore Leng 2, Jennifer I Lim 3, Minhaj Nur Alam 1,∗
PMCID: PMC12925154  PMID: 41732778

Abstract

Purpose

To conduct a comprehensive systematic evaluation of federated learning (FL) strategies for multi-disease retinal classification using OCT angiography (OCTA), implementing a 2-part experimental framework to establish foundational feasibility and optimize performance under realistic heterogeneous conditions while ensuring privacy preservation.

Design

Retrospective multi-center FL study using a systematic 2-part experimental design: (1) foundational feasibility evaluation under controlled homogeneous conditions, and (2) comprehensive optimization under realistic heterogeneous conditions using Dirichlet distribution partitioning (α = 0.5).

Participants

A total of 456 OCTA images from patients with 7 retinal pathologies, with diabetic retinopathy (31.1%) and normal cases (25.2%) comprising the majority, sourced from the public OCTA-500 data set (n = 300) and a private collection from the University of Illinois Chicago (n = 156).

Methods

Five FL aggregation strategies (federated averaging [FedAvg], federated proximal [FedProx], federated magnetic resonance imaging [FedMRI], federated Adagrad, and federated Yogi) were systematically evaluated across multiple optimization dimensions: 7 architecture configurations spanning vision transformers, established convolutional neural networks, and hybrid models; 5 transfer learning freezing strategies; 3 local epoch configurations (2, 5, and 10); and scalability analysis across 2, 3, and 5-client federations. Security mechanisms including differential privacy (ε = 1.0–8.0) and secure aggregation were integrated and evaluated. Performance was assessed across 3 classification scenarios: 7-class, 4-class modified, and 4-class streamlined.

Main Outcome Measures

Classification accuracy, receiver-operating-characteristic area under the curve (ROC-AUC), and macro-averaged F1-score with comprehensive privacy-utility analysis and computational efficiency metrics.

Results

Under controlled conditions, FL achieved superior performance in simplified classifications, with FedAvg, FedProx, and FedMRI reaching 72.09% accuracy versus 69.77% centralized training. Comprehensive optimization identified DenseNet121 as optimal architecture (79.55% accuracy, 89.68% ROC-AUC), with “most” freezing strategy (75% frozen layers) providing 60% training time reduction while maintaining superior performance. Federated proximal demonstrated exceptional resilience to heterogeneity (–11.7% degradation). Bonawitz secure aggregation achieved optimal privacy-utility balance (63.64% accuracy with cryptographic guarantees), whereas differential privacy maintained clinical utility under moderate constraints (ε ≈ 4–6).

Conclusions

This systematic evaluation establishes FL as a comprehensive solution for privacy-preserving multi-institutional OCTA-based disease classification, with careful architectural selection, optimization strategies, and security mechanisms enabling performance that matches or exceeds centralized approaches while maintaining regulatory compliance and clinical utility.

Financial Disclosure(s)

The authors have no proprietary or commercial interest in any materials discussed in this article.

Keywords: Age related macular degeneration (AMD), Diabetic retinopathy (DR), Federated learning, Optical coherence tomography angiography (OCTA), Privacy-preserving artificial intelligence


Medical imaging has fundamentally transformed clinical practice by enabling noninvasive, high-resolution visualization of anatomic and pathologic structures. It accounts for almost 90% of all health care data1 and thus plays a pivotal role in diagnosis, treatment planning, and disease monitoring.2 The continuous evolution of imaging technologies, from conventional radiography to advanced molecular3 and functional imaging modalities, has significantly enhanced early disease detection and characterization capabilities, thereby improving diagnostic accuracy and patient outcomes. In ophthalmology, this technological advancement has been particularly transformative, as the complex microarchitecture of ocular tissues necessitates high-resolution, multi-modal imaging for precise diagnosis and optimal treatment planning.

Vision impairment and blindness affect approximately 2.2 billion people globally, with retinal diseases such as diabetic retinopathy (DR) being primary contributors.4 Prevention of irreversible vision loss depends critically on early detection and accurate diagnosis, particularly given the projected increase in disease burden associated with global population aging and the diabetes pandemic.5 However, the subtle morphological variations and overlapping phenotypic presentations of retinal pathologies pose diagnostic challenges even for subspecialty-trained ophthalmologists.6

Traditional ophthalmic imaging modalities such as fundus photography and fluorescein angiography have long served as standard tools for retinal examination. Although fundus photography provides valuable information about the retinal surface, its inability to visualize deeper structures limits its diagnostic capabilities.7 Fluorescein angiography, despite offering detailed vascular imaging, requires intravenous contrast agents and can cause adverse reactions in some patients.8 The clinical introduction of OCT in the 1990s represented a paradigm shift in retinal imaging, enabling noninvasive, cross sectional visualization of retinal microstructure.9 However, conventional OCT’s limitation in visualizing blood flow dynamics created a crucial gap in understanding vascular pathologies.

OCT angiography (OCTA) has revolutionized retinal vascular imaging by providing depth-resolved, motion-contrast visualization of retinal and choroidal microvasculature without requiring intravenous contrast administration, offering superior capillary-level detail compared with conventional fluorescein angiography.10 The growing adoption of OCTA has led to an exponential increase in imaging data complexity and volume, making manual interpretation increasingly time-consuming and challenging for clinicians. A single OCTA scan can generate multiple en face projections and cross sectional images, requiring careful examination of various vascular layers.11 Although conventional image processing approaches using handcrafted features such as vessel density quantification and foveal avascular zone morphometry provided initial automation capabilities, these methods demonstrated limited sensitivity for subtle pathologic changes and poor generalizability across diverse retinal pathologies.12 They often fell short in capturing subtle disease patterns and showed limited generalizability across different pathologies.13 Deep learning, particularly convolutional neural networks (CNNs), has emerged as a powerful tool for automated OCTA analysis, demonstrating remarkable accuracy in disease classification tasks. Recent studies have achieved classification accuracies >90% for specific retinal conditions using deep learning approaches.14,15

Despite these technological advances, current deep learning implementations face substantial barriers to clinical translation and widespread adoption. Most successful implementations have been confined to single-institution settings, where data is centrally collected and processed. This constraint stems from several critical challenges: the scarcity of public data sets (with OCTA-500 being one of few available resources),16 strict medical data-sharing regulations such as Health Insurance Portability and Accountability Act (HIPAA) and General Data Protection Regulation (GDPR), and significant variations in image acquisition parameters across different OCTA devices and imaging protocols.9 Furthermore, the prevalence of retinal diseases varies significantly across populations and geographic regions,17 leading to inherent data imbalances. These factors have created isolated data silos, hampering the development of robust, generalizable classification models that could benefit from diverse patient populations. Despite promising results, federated learning (FL) research in medical imaging has predominantly focused on proof-of-concept studies. Rehman et al18 noted that the research on FL for medical imaging artificial intelligence (AI) is still in its early stages, with most studies evaluating limited client numbers. Sáinz-Pardo Díaz and López García19 conducted exhaustive analysis of a use case where the number of clients varies, highlighting the need for systematic scalability evaluation across varying institutional configurations and optimization strategies.

The imperative to develop robust, generalizable AI models while preserving patient privacy has led to a paradigm shift in medical image analysis. Deep learning research has demonstrated that FL, a distributed learning framework, can effectively address the challenges of cross-institutional model training while maintaining data privacy.20 Federated learning represents a transformative approach to collaborative AI development where instead of centralizing sensitive patient data, each institution trains models locally and shares only model parameters, thereby maintaining compliance with privacy regulations while leveraging the collective knowledge of multiple health care institutions. This approach has shown remarkable success in various medical imaging applications, including chest x-ray analysis,21 brain tumor segmentation,22 and coronavirus disease 2019 detection.23

Beyond FL’s inherent privacy benefits, formal privacy-preserving mechanisms are increasingly required for regulatory compliance. Ziller et al24 demonstrated that FL allows training models without explicitly sharing patient data and thus mitigates some confidentiality and privacy issues associated with clinical data. Differential privacy (DP) supplements this with quantitative bounds on the amount of privacy provided. Recent work by Shukla et al25 shows that FL combined with DP achieves 96.1% accuracy with a privacy budget of ε = 1.9, ensuring strong privacy preservation with minimal performance trade-offs for breast cancer diagnosis; however, comprehensive privacy-utility analysis across diverse medical imaging tasks remains limited. The utility of distributed FL for ophthalmic diagnosis has been explored recently by Gholami et al26 where FL demonstrated exceptional capability to train generalizable models across multiple institutions.

The unique characteristics of OCTA imaging make it an ideal candidate for FL approaches. OCT angiography data exhibits significant heterogeneity due to variations in scanning protocols, device manufacturers, and image quality across institutions.27 Additionally, the complex nature of retinal pathologies and their varying representations across different patient populations28 necessitate diverse training data that no single institution can provide. Federated learning’s ability to leverage distributed data sets while maintaining patient privacy could potentially address these challenges, enabling the development of more robust and generalizable models for OCTA analysis.

Despite the promise of FL in other imaging domains, its application to OCTA-based multi-class disease classification has not been explored. In this work, we demonstrate a comprehensive FL-based pipeline for OCTA-based multi-disease diagnosis. Our investigation utilizes a diverse data set comprising 300 images from the public OCTA-500 data set14 and an additional 156 images from a private collection at the University of Illinois Chicago (UIC). The combined data set covers 7 distinct retinal conditions: age-related macular degeneration (AMD), choroidal neovascularization (CNV), central serous chorioretinopathy (CSC), DR, retinal vein occlusion (RVO), normal, and other retinal conditions.

We conduct a comprehensive study evaluating 5 FL aggregation strategies for retinal disease classification using OCTA: federated averaging (FedAvg),29 federated Adagrad (FedAdagrad),30 federated Yogi (FedYogi),30 federated proximal (FedProx),31 and federated magnetic resonance imaging (FedMRI).32 We chose these aggregation strategies based on their potential to address the unique challenges of OCTA classification, such as data heterogeneity, class imbalance, and the need for stable and robust models. Federated averaging serves as a baseline, whereas FedAdagrad and FedYogi introduce adaptive learning rates to handle the varying contributions of different clients. Federated proximal incorporates proximal terms to improve model stability, and FedMRI utilizes a contrastive loss function to enhance feature representation learning.

Our experimental setup involves training a ResNet50 model, pretrained on ImageNet, using each of the 5 aggregation strategies. Although previous FL studies in medical imaging have primarily focused on single algorithm evaluations, systematic comparative analysis across multiple architectures, optimization strategies, and security mechanisms remains largely unexplored, limiting clinical translation and deployment guidance for health care institutions. Kim et al33 found that transfer learning has been arbitrarily configured in the majority of studies in medical imaging, indicating a broader need for systematic optimization frameworks in federated settings. We evaluate the performance of our FL strategies using various metrics, including accuracy, precision, recall, F1-score, and receiver-operating-characteristic area under the curve (ROC-AUC), both at the overall and per-class levels. Additionally, we analyze the convergence behavior of each strategy to gain insights into their learning dynamics.

The main contributions of this work are as follows:

  • 1.

    Two-part experimental framework: A systematic 2-part experimental framework that establishes foundational FL performance under controlled conditions before progressing to comprehensive heterogeneous evaluation, providing methodologically rigorous assessment of 5 federated aggregation strategies (FedAvg, FedAdagrad, FedYogi, FedProx, and FedMRI) for OCTA-based retinal disease classification.

  • 2.

    Comprehensive optimization and security evaluation: Comprehensive architectural optimization and security evaluation encompassing systematic comparison of 7 distinct deep learning architectures across 3 performance tiers (modern vision transformers [ViTs], established CNNs, and emerging hybrids), coupled with privacy-preserving mechanism assessment including DP and secure aggregation protocols.

  • 3.

    Scalability and deployment optimization: Scalability and deployment optimization analysis investigating FL performance across varying client numbers (2–5 institutions), local training configurations, and transfer learning strategies, providing evidence-based guidance for multi-institutional clinical deployment while addressing practical challenges including data heterogeneity, class imbalance, and regulatory compliance requirements.

Data Set

Data Source and Description

In this study, we used OCTA imaging data from 2 distinct clinical locations—the publicly available OCTA-500 data set (Jiangsu Province Hospital, China; predominantly Asian ethnicity) and a private data set from the University of Illinois at Chicago (Chicago, Illinois; demographics reflecting Cook County’s diverse population including White, Black, and Hispanic populations). This study was conducted in accordance with the Declaration of Helsinki and approved by the Institutional Review Board at UIC and University of North Carolina at Charlotte. OCT angiography provides high-resolution, depth-resolved visualization of retinal and choroidal vasculature. This enables noninvasive assessment of vascular networks across various retinal layers without contrast agents. This advanced imaging modality enables detailed examination of microvascular abnormalities associated with various ocular diseases, including DR and AMD. Thus, it provides valuable biomarkers for early disease detection and progression monitoring.

The OCTA-500 data set contains 3 projection types: OCTA (Full) capturing vascular information throughout the entire retinal depth, OCTA (inner limiting membrane–outer plexiform layer) isolating superficial and deep capillary plexuses, and OCTA (outer plexiform layer–Bruch’s membrane) emphasizing choriocapillaris and choroidal vasculature. We used exclusively the OCTA (Full) projection maps (en face angiography) to ensure comprehensive vascular visualization and maintain consistency with our UIC data processing approach.

The UIC data set consisted primarily of conventional OCT B-scans rather than native OCTA data. To generate comparable angiographic visualizations, we applied the computational angiography methods described by Badhon et al,34 which extract vascular information from OCT intensity variations to create OCTA-equivalent projections. This processing ensured functional equivalence between data sets despite different acquisition modalities.

The combined data set comprised of 456 OCTA images representing 7 distinct retinal pathologies in scenario A, including AMD, CNV, CSC, DR, RVO, normal cases, and other retinal conditions. Some samples from the dataset are shown in Figure 1. The disease distribution reflects clinical prevalence patterns, with DR and normal cases representing the largest subgroups (31.1% and 25.2%, respectively), followed by other retinal conditions (21.1%) and AMD (9.4%). Less common pathologies include CSC (3.1%), CNV (2.4%), and RVO (1.8%).

Figure 1.

Figure 1

Representative OCT angiography (OCTA) image samples from 7 retinal disease classes across OCTA-500 and University of Illinois Chicago (UIC) data sets: age-related macular degeneration (AMD), choroidal neovascularization (CNV), central serous chorioretinopathy (CSC), diabetic retinopathy (DR), retinal vein occlusion (RVO), normal retina, and other retinal conditions.

To facilitate comprehensive analysis, we organized our experiments into 3 distinct scenarios:

  • •

    Scenario A: All 7 classes (AMD, CNV, CSC, DR, normal, others, and RVO) were maintained separately.

  • •

    Scenario B: Four classes were used, with CNV cases incorporated into the AMD class, and both CSC and RVO cases excluded.

  • •

    Scenario C: Four classes were used (AMD, DR, normal, and others), completely excluding CNV, CSC, and RVO cases.

This structured approach to data set organization allows for robust model evaluation across varying levels of class granularity while addressing the inherent challenge of class imbalance in medical imaging data sets. The natural class imbalance in our data set presents both challenges and opportunities for developing robust classification models.

Data Preprocessing

We implemented a standardized preprocessing pipeline to ensure consistent image quality while preserving diagnostically relevant features across both data sets. The workflow began with intensity normalization to scale pixel values to a standardized range that optimizes neural network training. We then resized all images to 224 × 224 pixels through a combination of padding and center cropping, ensuring consistent dimensions across the data set while maintaining central pathologic features.

For the training data set, we implemented a comprehensive augmentation strategy to enhance model generalizability. This included random 90° rotations, controlled zoom variations (±10%), addition of minor Gaussian noise, contrast adjustments, Gaussian smoothing, and histogram shifts. Each augmentation was applied with carefully calibrated probability to maintain clinical relevance while creating sufficient variability for robust learning. Validation and test data sets underwent the same normalization and standardization procedures without augmentations to ensure faithful evaluation of model performance on clinical-quality images.

Data Set Splitting

We used stratified random sampling for data set partitioning to maintain representative distribution across 3 distinct experimental scenarios. For all scenarios, we used a 2-stage stratified splitting strategy to maintain pathologic representation and source distribution across subsets. The first stage allocated 80% of the data to the training set, with the remaining 20% divided equally between validation and test sets (10% each). Crucially, stratification was performed on both disease class and data source (OCTA-500 vs UIC) simultaneously to ensure balanced representation across all splits.

For scenario A, we combined images from both OCTA-500 and UIC data sets, maintaining all original disease labels. Scenario B involved a targeted reclassification approach where CNV cases were relabeled as AMD, which reflects clinical practice given that CNV is a manifestation of exudative AMD rather than a distinct entity. Central serous chorioretinopathy and RVO samples were excluded to focus on more prevalent conditions. Scenario C further refined the disease spectrum by excluding CNV samples entirely, concentrating on a more limited set of well-differentiated pathologies.

In scenario A (shown in Table 1), all 7 disease categories were preserved, resulting in 364 images for training (79.8%), and 46 images each (10.1%) for validation and testing. This comprehensive approach allowed us to evaluate model performance across the full spectrum of retinal pathologies, despite the class imbalance challenges presented by less common conditions such as CNV (2.4%), CSC (3.1%), and RVO (1.8%).

Table 1.

Training, Validation, and Test Set Distribution across 7 Retinal Disease Classes for Scenario A, Showing Sample Counts from OCTA-500 and UIC Data Sets with Stratified Splitting Maintaining 80%-10%-10% Ratio

Disease Training Set
Validation Set
Test Set
OCTA-500 UIC Total OCTA-500 UIC Total OCTA-500 UIC Total
AMD 34 0 34 4 0 4 5 0 5
CNV 9 0 9 1 0 1 1 0 1
CSC 11 0 11 2 0 2 1 0 1
DR 28 95 123 3 12 15 4 12 16
Normal 73 29 102 9 4 13 9 4 13
Others 77 0 77 10 0 10 9 0 9
RVO 8 0 8 1 0 1 1 0 1
Total 240 124 364 30 16 46 30 16 46

AMD = age-related macular degeneration; CNV = choroidal neovascularization; CSC = central serous chorioretinopathy; DR = diabetic retinopathy; OCTA = OCT angiography; RVO = retinal vein occlusion; UIC = University of Illinois Chicago.

In scenario B (Table 2), we implemented a modified 4-class structure where CNV cases were incorporated into the AMD category, whereas CSC and RVO cases were excluded. This resulted in a more balanced class distribution with 345 total images in the training set: 43 AMD (including original CNV cases), 123 DR, 102 normal, and 77 in the “others” category. The training, validation, and test sets contained 345, 43, and 43 images, respectively.

Table 2.

Training, Validation, and Test Set Distribution for Scenario B with 4 Classes Where CNV Cases Are Merged with AMD and CSC/RVO Cases Are Excluded, Maintaining Stratified Data Splitting across OCTA-500 and UIC Sources

Disease Training Set
Validation Set
Test Set
OCTA-500 UIC Total OCTA-500 UIC Total OCTA-500 UIC Total
AMD 43 0 43 5 0 5 6 0 6
DR 28 95 123 4 12 16 3 12 15
Normal 73 29 102 9 4 13 8 4 12
Others 77 0 77 9 0 9 10 0 10
Total 221 124 345 27 16 43 27 16 43

AMD = age-related macular degeneration; CNV = choroidal neovascularization; CSC = central serous chorioretinopathy; DR = diabetic retinopathy; OCTA = OCT angiography; RVO = retinal vein occlusion; UIC = University of Illinois Chicago.

Scenario C (Table 3) used a strict 4-class taxonomy (AMD, DR, normal, and others) by completely excluding CNV, CSC, and RVO cases. This produced a streamlined data set that focused on the most prevalent retinal conditions while maintaining class separation between AMD and CNV. Similar to scenario B, this configuration used 336 total images in the training set, with 34 AMD cases representing 9.9% of the data set.

Table 3.

Training, Validation, and Test Set Distribution for Scenario C with 4 Distinct Classes Completely Excluding CNV, CSC, and RVO Samples, Showing Final Data Set Composition from OCTA-500 and UIC Sources

Disease Training Set
Validation Set
Test Set
OCTA-500 UIC Total OCTA-500 UIC Total OCTA-500 UIC Total
AMD 34 0 34 4 0 4 5 0 5
DR 28 95 123 4 12 16 3 12 15
Normal 73 29 102 9 4 13 8 4 12
Others 77 0 77 9 0 9 10 0 10
Total 212 124 336 27 16 43 27 16 43

AMD = age-related macular degeneration; CNV = choroidal neovascularization; CSC = central serous chorioretinopathy; DR = diabetic retinopathy; OCTA = OCT angiography; RVO = retinal vein occlusion; UIC = University of Illinois Chicago.

The multi-source composition of our data set merits particular attention, as it directly influences model learning dynamics. Diabetic retinopathy cases derived from both OCTA-500 and UIC sources constitute the largest group across all scenarios, followed by normal cases also obtained from both repositories. The remaining categories—AMD, CNV, CSC, other conditions, and RVO—come exclusively from the OCTA-500 data set. This multi-source approach for common conditions (DR and normal cases) enhances model generalizability, while maintaining comprehensive coverage of rarer pathologies.

Our data set curated for the 3 experimental scenarios offers several distinctive advantages. The inclusion of 7 pathologic categories in scenario A represents broader coverage than most comparable studies, which typically focus on 3 to 4 conditions. Furthermore, the multi-center nature of our data, combining images from different institutions and imaging systems, inherently incorporates clinical variability that is crucial for developing robust clinical applications.

Methods

FL Framework

To systematically evaluate FL for OCTA-based multi-disease classification, we designed a comprehensive 2-part experimental framework that progresses from controlled baseline validation to realistic heterogeneous scenarios. This methodological approach addresses both the foundational feasibility of FL using OCTA and its performance under the heterogeneity characteristic of clinical multi-institutional collaborations.

Experiment 1: Foundational Feasibility Study

In this experiment, we aim to establish a baseline for FL performance under controlled, homogeneous data distribution conditions so that we can isolate the effects from data heterogeneity complexities related to data set sources. We use stratified random sampling to ensure similar class distributions across participating clients, creating an idealized federated scenario where institutions have proportionally balanced disease representations. By doing this, we are able to evaluate the inherent capabilities and limitations of different federated aggregation strategies (FedAvg, FedProx, FedMRI, FedYogi, and FedAdagrad) without the confounding effects of statistical heterogeneity, providing a clean baseline for comparison.

Experiment 2: Comprehensive Heterogeneous Evaluation

In this experiment, we focus on systematic optimization and evaluation under realistic nonindependent and identically distributed (IID) conditions that reflect actual multi-institutional medical data scenarios. We use Dirichlet distribution-based data partitioning to create configurable levels of statistical heterogeneity, enabling controlled simulation of practical FL challenges where institutions have varying disease prevalence patterns. Through this approach, we address genuine FL scenarios where health care institutions naturally exhibit different patient demographics, disease distributions, and imaging characteristics—the primary motivation for FL in medical AI.

Data Distribution Strategies and Implementation

Our FL framework implements 2 distinct data partitioning approaches, each serving specific experimental objectives and reflecting different aspects of multi-institutional medical data-sharing scenarios.

Stratified Partitioning (Experiment 1)

For our foundational feasibility study, we used stratified random sampling to create balanced data distributions across clients. This approach ensures that each participating institution has proportionally similar class distributions: P(c | Dk) ≈ P(c | Dⱼ) for any clients k, j and disease class c ∈ C. Thus, we created IID-like conditions where statistical heterogeneity is minimized, allowing us to evaluate federated aggregation strategies without the confounding effects of data distribution skew. This baseline establishes the upper bound of FL performance when data heterogeneity challenges are removed.

Dirichlet-Based Partitioning (Experiment 2)

For comprehensive heterogeneous evaluation, we implement non-IID data partitioning using the Dirichlet distribution, which has become the gold standard for simulating realistic FL scenarios.35,36 In this setup, the concentration parameter α directly controls the level of statistical heterogeneity across clients. Lower values of α produce more skewed distributions, where some clients dominate specific disease classes, whereas higher values of α generate more balanced partitions approaching IID conditions.

In our implementation, we specifically use α = 0.5, which creates a moderately heterogeneous scenario. This value reflects typical multi-institutional variations observed in medical data while still ensuring that each client receives a sufficient number of samples per class to maintain trainability. By choosing α = 0.5, we balance realism with stability, allowing us to evaluate FL under representative but manageable heterogeneity conditions.

Client-Server Architecture

Our FL system implements a client-server architecture as shown in Figure 2 where K health care institutions collaboratively train a disease classification model while preserving patient data privacy. The central server coordinates the learning process, whereas distributed clients {C1, C2, …, Ck} participate with their local OCTA image data sets Dk.

Figure 2.

Figure 2

Federated learning architecture showing 2-client system where health care institutions collaboratively train OCT angiography disease classification models while preserving patient data privacy through parameter-only communication.

The system initializes with the server distributing a global deep learning classification model Mm with parameters θm. During experiment 1, we kept all convolutional layers in a frozen state while training only the final classification layer (parameters θf ⊂ θm), reducing computational complexity from O(|θm|) to O(|θf|) where |θm| >> |θf|. This creates the foundation for experiment 2 where we perform multiple ablation studies to determine the optimal number of frozen layers for our FL scenario.

During each communication round t, participating clients train locally using their respective data distribution strategy—either stratified (experiment 1) or Dirichlet-partitioned (experiment 2) data sets. The approach ensures clients never share actual patient images, transmitting only locally trained model parameters θkt back to the server for aggregation. Clients enhance local training with specialized medical image augmentation techniques T, creating variations through controlled transformations:

T(xi)={T1(xᵢ),T2(xᵢ),…,Tm(xᵢ)}

where xᵢ represents an original image and Tⱼ are transformations including rotations, zooming, and contrast adjustments.

The server then aggregates these parameters using 5 federated averaging strategies, such as FedAvg, FedProx, FedMRI, FedAdagrad, and FedYogi, each offering different advantages for handling the unique challenges of medical data. Throughout the training process, we use a cosine annealing learning rate schedule:

ηt=ηmin+(1/2)(ηmax–ηmin)(1+cos(πt/T))

where ηt is the learning rate at round t, ηmin and ηmax are the minimum and maximum learning rates, and T is the total number of rounds.

This collaborative approach enables multiple institutions to contribute to a more robust, generalizable multi-disease classification model characterized by performance metrics such as accuracy, F1-score, and ROC-AUC score without compromising patient privacy or violating data-sharing regulations.

Communication Protocol

The communication protocol operates through 3 sequential phases: model distribution, decentralized training, and parameter aggregation.

During the model distribution phase, the server broadcasts the current global model parameters θgt to the selected clients {Ck | k ∈ St}, accompanied by the adaptive learning rate ηt and round-specific metadata.

During the decentralized training phase, each institution performs local model optimization using its private OCTA data sets under the assigned data distribution strategy. Clients execute E local epochs using the Adam optimizer with server-specified learning rates, formally expressed as:

θkt+1=θgt–ηt∇L(θgt,Dk)

where L represents the cross-entropy loss function evaluated on local data sets under either stratified or Dirichlet distribution conditions.

The parameter aggregation phase implements weighted parameter aggregation from all participating clients based on relative data set sizes and the selected federated strategy. This ensures that institutions with larger data sets contribute proportionally while maintaining privacy compliance.

Federated Simulation Environment

We built our simulation environment on the Flower framework (version 1.13.1) which allows us to accurately model multiple institutional nodes with distinct patient populations and data distribution characteristics. We designed the environment to support both experimental paradigms. During experiment 1, we simulate idealized institutions with balanced disease representations through stratified sampling, enabling us to conduct a clean evaluation of federated aggregation algorithms. During experiment 2, we model realistic health care networks with configurable statistical heterogeneity using Dirichlet partitioning, reflecting authentic challenges in multi-institutional medical AI collaborations.

We implemented comprehensive evaluation using centralized test data sets accessible only to the server, providing an unbiased assessment of collaborative learning effectiveness. We also used macro-averaged metrics to ensure equal importance for all disease categories regardless of frequency, which is crucial for rare but clinically significant conditions.

This dual-environment approach enables us to systematically progress from controlled algorithmic validation to realistic deployment scenarios, addressing both fundamental feasibility questions and practical implementation challenges in federated medical imaging.

Model Architecture

Architecture Tier Selection Criteria

The selection of appropriate deep learning architectures for federated OCTA classification requires systematic evaluation of diagnostic performance, computational efficiency, and clinical deployment feasibility. We developed a multi-tier evaluation framework organizing architectures based on their validation in medical imaging applications.

Modern ViTs

Recent studies have demonstrated the efficacy of ViTs in retinal disease classification. Vision transformers have shown promising performance in various classification tasks previously dominated by CNNs,37 with specific applications in ophthalmology showing competitive results. For OCTA analysis, ViT, Tokens-To-Token ViT, and Mobile ViT algorithms achieved predictive accuracies of 95.14%, 96.07%, and 99.17%, respectively38 for retinal disease classification using OCT images. By shifting the window partition, the Swin-Poly Transformer constructs connections between neighboring nonoverlapping windows in the previous layer and thus has the flexibility to model multi-scale features,39 making it particularly suitable for analyzing retinal vascular patterns in OCTA.

Established CNN Architectures

Convolutional networks remain highly effective for medical imaging applications. ResNet50 was used as a base network to implement a deep CNN for the detection of early glaucoma, achieving a validation accuracy of 0.9848.40 DenseNet architectures have shown particular promise in retinal applications: DenseNet is much more efficient in terms of parameters and computations to the same degree as ResNet and VGGNet, facilitating the reuse of gradient information and significantly reducing the number of parameters.41 Additionally, dense convolutional networks have demonstrated good applications in medical image analysis.42

Emerging Architectures

Modern hybrid architectures combine advantages of different approaches. ConvNeXt presents itself as a “modernized” convolution network, morphing the foundational principles of a standard ResNet to align more closely with the architectural nuances of ViT.43

Foundational Architecture Configuration

For our controlled baseline evaluation, we selected ResNet50 based on its established performance in medical imaging and FL contexts. Convolutional neural networks are the predominant model used in FL research for imaging, where deep CNNs have demonstrated excellent performance.44

Our implementation employs strategic parameter freezing following established transfer learning practices in medical imaging. The local training process occurs autonomously at each client’s end, enabling iterative updates based on unique data set characteristics.45 We freeze all convolutional layers (23.5 million parameters, approximately 92% of total) while training only the final classification layer (14 343 parameters, 0.056% of total).

This approach leverages several advantages. Deep architectures allow capturing fine-grained details and hierarchical representations, crucial for identifying subtle anomalies in DR images or delineating boundaries in medical scans. Similarly, ResNet’s residual connections help mitigate vanishing gradient issues, allowing effective training of deeper layers without information loss, beneficial for complex medical image analysis tasks.

Comprehensive Architecture Evaluation Framework

Building upon foundational validation, we implemented systematic architectural comparison under realistic federated conditions. Federated learning has remained robust to a variety of machine learning models, with neural networks, especially CNNs, being predominant across medical imaging applications.46

Federated learning in medical imaging requires architectures that can generalize across varying computational environments, data distributions, and model complexities. The systematic comparison spans 7 distinct architecture-freezing combinations, covering representative models from each performance tier with varying parameter optimization strategies.

Architecture selection must consider clinical deployment constraints. Recent results indicate that models trained in a federated approach can compete with those trained in a centralized way and even outperform models trained on separate institutional data.47 Our framework prioritizes configurations suitable for clinical deployment, considering computational constraints typical of health care environments and communication bandwidth limitations affecting multi-institutional collaboration.

Multi-Strategy FL with Secure Aggregation Framework

We implemented a comprehensive FL framework supporting 5 distinct aggregation strategies, each combined with 3 privacy-preserving mechanisms. This modular approach enables systematic evaluation of algorithmic performance across varying security requirements, addressing the dual imperatives of diagnostic accuracy and patient data protection in multi-institutional health care collaborations.

FL Strategy Implementation

Our framework incorporates both established federated optimization algorithms and custom approaches designed for medical imaging applications. Each strategy addresses specific challenges inherent in distributed training of medical AI systems, including statistical heterogeneity, class imbalance, and convergence stability.

FedAvg

We implemented FedAvg29 as our baseline strategy, providing standard weighted parameter averaging:

θt+1=Σk=1K(nk/n)θkt+1

where θt+1 represents the global model parameters, nk denotes the local data set size for client k, and n is the total data set size. This approach establishes performance baselines without additional hyperparameters, facilitating systematic comparison across experimental conditions.

FedProx

To address model divergence in heterogeneous medical data sets, we implemented FedProx,31 which introduces a proximal regularization term constraining local model updates:

Lk(θ)=Lk(θk)+(μ/2)||θk-θ||2

The proximal parameter μ = 0.1 was selected based on established medical imaging applications,31 where this value demonstrated effectiveness in controlling client drift while maintaining convergence stability. This regularization mechanism is particularly relevant for OCTA classification, where institutional variations in imaging protocols and patient demographics create significant data heterogeneity.44

FedAdagrad

We implemented FedAdagrad30 to provide per-parameter learning rate adaptation:

gt=Σk=1t(nk/n)gk,t
θK+1=θK-(η/(vt+ε))gt

where vt accumulates squared gradients. Following established FL practice,30 we configured: η = 0.1 (base learning rate), ηl = 0.316 (local learning rate multiplier, derived as √0.1), β1 = 0.9 (momentum coefficient), and τ = 1 × 10–9 (numerical stability constant). These parameters were selected to handle varying gradient magnitudes across different retinal pathology features.

FedYogi

Federated Yogi30 extends adaptive optimization with enhanced gradient accumulation mechanisms:

vt=vt-1-(1-β2)·sign(vt-1-gt2)·gt2
θK+1=θK-(η/(vt+ε))gt

We configured: η = 0.01 (conservative learning rate for stability), ηl = 0.0316, β1 = 0.9, β2 = 0.99 (second moment decay rate for noise reduction), and τ = 1 × 10–3. These parameters were calibrated to prevent overfitting to dominant disease classes while maintaining responsiveness to rare pathology updates—critical for balanced multi-class OCTA classification.48

FedMRI

We developed FedMRI as a custom strategy incorporating contrastive learning principles for medical imaging applications:

Lcontrastive=–log[exp(sim(zi,zj)/τ)/Σk≠iexp(sim(zi,zk)/τ)]

Our configuration includes: μ = 0.1 (regularization coefficient balancing local and global objectives), contrastive_weight = 0.01 (weight for contrastive learning component), and dual_space = false (single representation space optimization). This approach enforces consistent feature representations across heterogeneous OCTA data sets, addressing variations in vascular pattern manifestation across different imaging centers.32

Hyperparameter Configuration and Optimization

Systematic hyperparameter configuration is critical for ensuring reproducible and fair comparison across FL strategies in medical imaging applications. Our experimental framework employs carefully calibrated parameters based on established FL practices and medical imaging-specific considerations.

Table 4 presents the comprehensive hyperparameter configuration for all FL strategies evaluated in our systematic framework. Base learning rates were selected based on established FL protocols for medical imaging applications, with adaptive methods requiring conservative rates to maintain stability under heterogeneous data conditions.24,49 The proximal regularization parameter μ = 0.1 for FedProx and FedMRI was selected following medical imaging applications demonstrating effectiveness in controlling client drift while maintaining convergence stability.31,32

Table 4.

Comprehensive Hyperparameter Configuration for Federated Learning Strategies Used in Systematic OCTA Disease Classification Experiments

Strategy Key Parameters
FedAvg –
FedProx μ = 0.1
FedMRI μ = 0.1, τ = 0.01
FedAdagrad ηl = 0.316, τ = 1 × 10–9
FedYogi ηl = 0.0316, β2 = 0.99

FedAdagrad = federated Adagrad; FedAvg = federated averaging; FedMRI = federated magnetic resonance imaging; FedProx = federated proximal; FedYogi = federated Yogi; OCTA = OCT angiography.

Adaptive optimization parameters for FedAdagrad and FedYogi were configured following established FL guidelines:30 local learning rate multipliers (ηl) derived as √η for numerical stability, momentum coefficients optimized for medical imaging gradient characteristics, and numerical stability constants (τ) preventing division-by-zero errors during sparse gradient updates.

Privacy-Preserving Aggregation Framework

Our secure aggregation framework implements 3 distinct privacy mechanisms that can be combined with any FL strategy, enabling comprehensive privacy-utility analysis across 15 possible configurations.

Differential Privacy Implementation

We implemented differential privacy using the differentially private stochastic gradient descent methodology,50 applying noise during local training rather than aggregation. This approach provides formal privacy guarantees with quantifiable privacy budgets:

Privacy parameters were configured as: ε = 8.0 (privacy budget providing moderate protection), δ = 1 × 10–5 (probability of privacy violation), and C = 1.0 (L2 sensitivity bound for gradient clipping). These values were selected based on established medical imaging applications demonstrating acceptable privacy-utility trade-offs.51,52

The DP mechanism applies controlled noise during local training through 3 steps: (1) gradient clipping: g˜ᵢ = gᵢ / max(1, ||gᵢ||2/C); (2) noise addition: ĝᵢ = g˜ᵢ + N(0, σ2C2I) where σ is calibrated to achieve (ε, δ)-DP; and (3) standard weighted aggregation on DP-noised updates.

Bonawitz Secure Aggregation

We implemented a simplified version of the Bonawitz secure aggregation protocol,53 providing cryptographic security through pairwise masking. Our implementation addresses key challenges identified in prior work54 while maintaining computational tractability for medical imaging applications.

Configuration parameters include: threshold ratio = 0.8 and dedicated state directory for cryptographic key management. The protocol operates through 3 phases:

Setup Phase

Each client generates Elliptic-Curve Diffie–Hellman keypairs for secure communication using SECP384R1 elliptic curves, following established cryptographic practices.53

Masking Phase

Clients apply pairwise masks using a pseudorandom generator and the AES-CTR derived from ECDH key exchange. For clients i and j, the shared secret sᵢⱼ = ECDH (skᵢ, pkⱼ) generates pairwise masks mᵢⱼ = pseudorandom generator (sᵢⱼ) using AES-CTR as the pseudorandom generator. Masks are applied with cancellation rules: w˜ᵢ = wᵢ + Σⱼ<ᵢ mᵢⱼ - Σⱼ>ᵢ mᵢⱼ.

Aggregation Phase

Server aggregates masked weights where pairwise masks cancel: Σᵢ w˜ᵢ = Σᵢ wᵢ, as the masking terms sum to zero.

Our simplified approach uses smaller mask magnitudes (scale = 1 × 10–8) to prevent numerical overflow while preserving security against honest-but-curious adversaries, following recent advances in secure aggregation for FL.55

Privacy-Utility Evaluation Protocol

For systematic privacy-utility analysis, we vary DP noise levels across clinically relevant ranges. Based on established medical imaging applications,56 we evaluate:

  • •

    High privacy: ε ∈ [1.0, 4.0] (strong privacy protection)

  • •

    Moderate privacy: ε ∈ [4.0, 8.0] (balanced privacy-utility)

  • •

    Low privacy: ε > 8.0 (minimal privacy protection)

  • Each configuration undergoes standardized evaluation using identical experimental protocols: server distributes global ResNet50 model with frozen convolutional layers, clients perform 10 local epochs with strategy-specific optimizations, server applies configured secure aggregation method, and comprehensive metrics are computed at client, federation, and centralized test levels across 20 communication rounds with cosine annealing learning rate schedules.

This systematic framework enables rigorous comparison of FL strategies while quantifying privacy-utility trade-offs essential for clinical deployment in multi-institutional health care environments, addressing regulatory requirements under HIPAA and GDPR while maintaining diagnostic accuracy for patient care.57

Systematic Experimental Design

Our experimental methodology implements a systematic optimization framework designed to establish foundational FL performance while comprehensively evaluating scalability, security, and architectural considerations for OCTA-based multi-disease classification. This progression from baseline establishment to comprehensive optimization ensures reproducible results suitable for clinical deployment.

Foundational Study Configuration

The baseline experimental configuration was established through principled methodological choices addressing computational efficiency and clinical applicability. The foundational setup employs a 2-client federation representing the minimal collaborative scenario between health care institutions, with each client maintaining distinct OCTA data sets following stratified partitioning to setup the foundation for our study.

The baseline model architecture utilizes ResNet50, selected based on established effectiveness in medical imaging applications and demonstrated transfer learning capabilities.58,59 ResNet50 provides an optimal balance between representational capacity and computational requirements in federated environments, with proven performance across diverse medical imaging tasks including ophthalmology applications.18,33 Transfer learning implementation freezes all convolutional layers (full freeze level) and trains only the final classification layer, reducing computational complexity from O(|θg|) to O(|θf|) where |θg| >> |θf|.21

The selection of 3 local epochs represents a systematically determined choice addressing the fundamental tension between local learning effectiveness and communication efficiency in FL. Prior research demonstrates that <3 epochs results in insufficient local adaptation, whereas significantly higher epoch counts lead to client drift and degraded global model performance under heterogeneous data conditions.60,61 Recent comprehensive studies in medical FL confirm that local epoch selection significantly impacts convergence characteristics, with 3 epochs providing sufficient local training for meaningful parameter updates while maintaining convergence stability.18,62

Communication rounds were configured for 20 iterations with full client participation (fraction-fit = 1.0), following established protocols for medical FL evaluation.62,63 The server employs cosine annealing learning rate schedules to optimize convergence characteristics, with comprehensive evaluation conducted on centralized test data sets to provide unbiased performance assessment.

Comprehensive Optimization Framework

Building upon the foundational configuration, our systematic optimization framework evaluates 4 critical dimensions of FL performance through comprehensive ablation studies. Each optimization dimension addresses specific research questions while maintaining experimental rigor through controlled variable manipulation.

Local Epochs Optimization

The local epoch optimization study systematically evaluates training duration impact on FL performance through comprehensive evaluation of 2, 5, and 10 local epochs configurations. This analysis addresses critical questions regarding optimal local training duration while quantifying trade-offs between local adaptation effectiveness and global model coherence, building on recent findings that demonstrate significant sensitivity of FL performance to local epoch selection.30,60

Our experimental design maintains identical network architectures, data partitioning strategies, and communication protocols across all local epoch configurations, isolating the impact of training duration on model performance. Each configuration undergoes evaluation across multiple FL strategies (FedAvg, FedProx, FedMRI, FedAdagrad, and FedYogi) to ensure generalizability of findings across optimization approaches.

Results from 15 distinct experimental configurations measure performance across accuracy, precision, recall, F1-score, and ROC-AUC metrics for each disease class (AMD, DR, normal, and others). This systematic evaluation addresses concerns regarding local epoch selection while providing evidence-based guidance for clinical deployments.61

Freezing Strategy Optimization

The freezing strategy optimization systematically evaluates transfer learning approaches through 5 distinct levels: none (full training), early (freeze first 25% of layers), partial (freeze first 50% of layers), most (freeze first 75% of layers), and full (freeze all convolutional layers). This comprehensive analysis builds on established transfer learning research in medical imaging21,59 while addressing the fundamental question of optimal knowledge transfer in federated medical imaging scenarios.

Each freezing level undergoes systematic evaluation across all FL strategies under identical experimental conditions, enabling direct comparison of transfer learning effectiveness. The experimental design encompasses 12 distinct configurations, measuring both computational efficiency (trainable parameters, training time) and predictive performance across all evaluation metrics, addressing established concerns regarding arbitrary transfer learning configuration in medical imaging studies.64

Client Scalability Analysis

The client scalability study systematically evaluates FL performance across realistic multi-institutional scenarios through comprehensive evaluation of 2, 3, 5, and 10 client configurations. This analysis directly addresses identified limitations regarding scalability assessment while quantifying performance degradation patterns and communication overhead scaling characteristics.62

Each scalability configuration maintains proportional data distribution through Dirichlet partitioning while ensuring minimum sample requirements per client for statistical validity. The experimental design encompasses 9 distinct configurations, systematically measuring performance metrics, convergence characteristics, and communication requirements across varying federation sizes, building on established frameworks for non-IID data evaluation in FL.35,36

Security Evaluation Framework

Our comprehensive security evaluation framework systematically compares 3 distinct privacy-preserving mechanisms: baseline (no security), DP, and Bonawitz secure aggregation.53 This analysis quantifies privacy-utility trade-offs essential for clinical deployment while ensuring compatibility across all FL strategies.

The differential privacy implementation employs differentially private stochastic gradient descent methodology50 with carefully calibrated parameters: ε = 8.0 (moderate privacy protection), δ = 1 × 10–5 (privacy violation probability), and C = 1.0 (L2 sensitivity bound). These parameters were selected based on established medical imaging applications demonstrating acceptable privacy-utility trade-offs while maintaining diagnostic accuracy requirements.63,65

The Bonawitz secure aggregation implementation provides cryptographic security through pairwise masking, using SECP384R1 elliptic curves for key generation and Advanced Encryption Standard in counter mode (AES-CTR) for pseudorandom mask generation.53 Our implementation addresses computational tractability concerns through smaller mask magnitudes (scale = 1 × 10–8) while preserving security guarantees against honest-but-curious adversaries, following recent advances in secure aggregation for FL.66

Security evaluation encompasses 6 distinct experimental configurations, measuring both privacy protection effectiveness and predictive performance degradation. Results demonstrate that privacy-preserving mechanisms can be successfully integrated into FL frameworks with acceptable performance trade-offs, enabling regulatory compliance under HIPAA and GDPR requirements while maintaining clinical utility.63

Results

Foundational Feasibility Study Results

To establish a methodologically rigorous baseline for FL performance in OCTA-based disease classification, we conducted a foundational feasibility study under controlled, homogeneous data distribution conditions. This approach enables isolation of algorithmic capabilities and limitations of different federated aggregation strategies without confounding effects of statistical heterogeneity,35 providing a clean baseline for comprehensive evaluation as recommended for FL medical imaging studies.35,44

We conducted experiments across 3 distinct scenarios using stratified data splitting to ensure similar class distributions across participating clients, creating idealized federated conditions where institutions have proportionally balanced disease representations. The experimental scenarios are:

  • •

    Scenario A: All 7 classes (AMD, DR, normal, others, CNV, CSC, and RVO)

  • •

    Scenario B: 4 modified classes (CNV considered as AMD, and CSC/RVO dropped)

  • •

    Scenario C: 4 classes (CNV, CSC, and RVO completely dropped)

For each scenario, we compare centralized training (using all available data), standalone client models (trained only on individual client data sets), and 5 baseline FL approaches (FedAvg, FedProx, FedMRI, FedYogi, and FedAdagrad) under controlled homogeneous conditions, following established benchmarking protocols for federated medical imaging.67

Global Test Set Performance under Homogeneous Distribution

Scenario A

Under homogeneous data distribution conditions, the centralized model achieved 65.22% accuracy, establishing our target baseline. The standalone models for client 1 and client 2 achieved 60.87% and 58.70% accuracy, respectively. Among the FL approaches under controlled conditions, FedMRI demonstrated the best performance with 60.87% accuracy, closely approaching the centralized baseline and matching client 1’s standalone performance.

This foundational result establishes that under ideal conditions (stratified data distribution), FL can approach centralized performance levels, consistent with findings from recent comprehensive benchmarks of FL in medical imaging.67 The performance gap between the best federated method (FedMRI: 60.87%) and centralized training (65.22%) was only 4.35%, demonstrating the fundamental viability of federated approaches for OCTA-based disease classification.

When evaluating discriminative capability through ROC-AUC (Table 5), FedAvg achieved 79.12%, remarkably close to the centralized model’s 79.55%. For precision-recall balance measured by F1-score (Table 5), FedAvg substantially outperformed all other methods, including the centralized approach, with 46.08% compared with 32.22% for centralized training. This unexpected superiority under controlled conditions suggests that federated aggregation may provide inherent regularization benefits, as observed in other medical imaging applications.35

Table 5.

Comprehensive Performance Evaluation under Stratified Data Splitting (Homogeneous Class Distribution) Showing Classification Accuracy, ROC-AUC, and Macro-Averaged F1-Score (%) across Federated Learning Approaches

Method Scenario A
Scenario B
Scenario C
Acc F1 ROC Acc F1 ROC Acc F1 ROC
Centralized 65.22 32.22 79.55 68.18 60.64 84.50 69.77 53.48 86.21
Client 1 standalone 60.87 28.33 67.44 68.18 63.85 80.33 67.44 52.92 73.10
Client 2 standalone 58.70 26.10 75.00 61.36 45.86 79.64 69.77 54.42 86.88
FedAvg 46.70 46.08 79.12 65.91 57.84 81.36 72.09 62.41 86.12
FedProx 30.25 27.60 76.52 68.18 55.51 80.68 72.09 56.05 86.46
FedMRI 60.87 27.48 70.70 61.36 53.54 81.64 72.09 56.17 87.17
FedYogi 28.26 6.30 72.75 45.45 33.97 69.15 37.21 20.83 72.66
FedAdagrad 45.65 16.00 71.74 29.55 11.40 61.90 30.23 11.61 69.50

Acc = accuracy; FedAdagrad = federated Adagrad; FedAvg = federated averaging; FedMRI = federated magnetic resonance imaging; FedProx = federated proximal; FedYogi = federated Yogi; ROC-AUC = receiver-operating-characteristic area under the curve.

Scenario B

In the 4 modified class scenario with homogeneous distribution, we observed general performance improvement across most methods. The centralized model achieved 68.18% accuracy, and FedProx matched this centralized performance exactly at 68.18% under controlled federated conditions.

For discriminative capability (ROC-AUC), FedMRI achieved the highest score among federated methods at 81.64%, outperforming both FedAvg (81.36%) and FedProx (80.68%) while approaching the centralized baseline of 84.50%. This establishes FedMRI’s superior discriminative capabilities under controlled conditions, consistent with its medical imaging-specific design principles.32

Scenario C

The simplified 4-class scenario under homogeneous distribution yielded the most remarkable foundational results. Although the centralized model achieved 69.77% accuracy, all 3 leading federated methods—FedAvg, FedProx, and FedMRI—achieved identical performance of 72.09% accuracy, surpassing centralized training by 2.32%.

This finding fundamentally challenges the assumption that FL necessarily sacrifices performance for privacy.44 Under controlled conditions with simplified class structures, federated approaches can exceed centralized performance, likely because of the regularization effects of distributed training and parameter averaging observed in medical imaging applications.18

For ROC-AUC, FedMRI achieved the highest score at 87.17%, surpassing both the centralized model (86.21%) and standalone approaches. For F1-score, FedAvg achieved 62.41%, substantially outperforming centralized (53.48%) and standalone approaches.

Client-Specific Performance under Controlled Conditions

Under controlled homogeneous conditions, distinct performance patterns emerged across clients. As can be seen in Table 6, FedMRI imaging consistently performed well for client 1 across all scenarios (50.00%, 57.14%, and 71.43%), whereas client 2 showed varying optimal approaches depending on scenario complexity. This foundational finding suggests that even under idealized conditions, institutional characteristics may influence optimal federated strategy selection, as reported in recent multi-institutional FL studies.49

Table 6.

Client-Specific Validation Set Accuracy (%) under Stratified Data Splitting Showing Individual Institutional Performance Patterns under Homogeneous Federated Learning Conditions

Method Scenario A
Scenario B
Scenario C
Client 1 Client 2 Client 1 Client 2 Client 1 Client 2
Client 1 standalone 66.67 45.83 52.38 77.27 57.14 52.38
Client 2 standalone 58.33 37.50 57.14 63.64 52.38 57.14
FedAvg 28.16 30.60 57.14 77.27 71.43 66.67
FedProx 31.02 23.57 47.62 68.18 76.19 61.90
FedMRI 50.00 45.83 57.14 68.18 71.43 52.38
FedYogi 29.17 25.00 52.38 40.91 33.33 28.57
FedAdagrad 25.00 37.50 28.57 31.82 28.57 28.57

FedAdagrad = federated Adagrad; FedAvg = federated averaging; FedMRI = federated magnetic resonance imaging; FedProx = federated proximal; FedYogi = federated Yogi.

Privacy-Utility Analysis under Controlled Conditions

The privacy-utility analysis under controlled conditions reveals that moderate privacy protection (ε ≈ 4–6) maintains clinically acceptable performance across all scenarios, consistent with findings from recent DP studies in medical imaging.25,65 Notably, in scenario A, controlled noise injection actually improved performance, demonstrating potential regularization benefits under complex classification tasks.52 Table 7 provides a quantitative analysis of our experiments related to privacy-utility tradeoff.

Table 7.

Privacy-Utility Trade-Off Analysis under Controlled Federated Learning Conditions Showing Classification Accuracy Degradation with Differential Privacy Protection

Scenario A
Scenario B
Scenario C
Noise Privacy (ε) Accuracy Noise Privacy (ε) Accuracy Noise Privacy (ε) Accuracy
0 – 46.70 0 – 65.91 0 – 72.09
0.5 17.35 52.17 0.5 17.78 50.00 0.5 17.78 60.47
0.8 6.07 50.00 0.8 6.31 36.36 0.8 6.31 58.14
1.0 3.80 50.00 1.0 3.98 45.45 1.0 3.98 58.14
1.5 1.81 50.00 1.5 1.91 38.64 1.5 1.91 48.84
2.0 1.17 52.17 2.0 1.23 38.64 2.0 1.23 46.51

Figure 3 visualizes these privacy-utility relationships, clearly illustrating the distinct trade-off curves for each scenario. The visualization reveals that scenario C consistently outperforms the other scenarios across all privacy levels, reinforcing our earlier findings about the benefits of focused class structures. Notably, all scenarios demonstrate acceptable utility levels even under moderate privacy constraints (ε ≈ 4-6), suggesting that meaningful privacy protection is achievable without prohibitive accuracy losses.

Figure 3.

Figure 3

Privacy-utility trade-off curves showing classification accuracy versus differential privacy epsilon values for 3 OCT angiography disease classification scenarios, demonstrating varying sensitivity to privacy-preserving noise injection.

Comprehensive Heterogeneous Evaluation Results

Building upon our foundational feasibility study, we conducted systematic optimization under realistic heterogeneous conditions using Dirichlet distribution-based data partitioning with concentration parameter α = 0.5. This comprehensive evaluation addresses 6 critical optimization dimensions identified from our foundational analysis, progressing from controlled baseline validation to realistic deployment scenarios for OCTA-based disease classification.

Data Heterogeneity Impact Analysis

To quantify the impact of realistic data heterogeneity on FL performance, in Table 8 we systematically compared stratified partitioning (homogeneous conditions from section Foundational Feasibility Study Results) against Dirichlet distribution-based partitioning (α = 0.5) representing moderate heterogeneity characteristic of multi-institutional medical data.35,60 The concentration parameter α = 0.5 creates realistic statistical heterogeneity where each client exhibits distinct disease prevalence patterns while maintaining sufficient samples per class for stable training.36

Table 8.

Data Heterogeneity Impact Analysis Comparing Stratified (Homogeneous) versus Dirichlet (α = 0.5) Partitioning with Accuracy (%), F1-Score (%), and ROC-AUC (%)

Strategy Homogeneous (Stratified)
Heterogeneous (Dirichlet)
Performance Change (%)
Acc F1 AUC Acc F1 AUC
FedMRI 72.09 56.17 87.17 61.36 49.45 79.44 –14.9
FedAvg 72.09 62.41 86.12 61.36 49.29 81.00 –14.9
FedProx 72.09 56.05 86.46 63.64 51.72 76.73 –11.7
FedAdagrad 30.23 11.61 69.50 45.45 33.31 73.59 +50.4
FedYogi 37.21 20.83 72.66 45.45 28.57 60.63 +22.1

Acc = accuracy; FedAdagrad = federated Adagrad; FedAvg = federated averaging; FedMRI = federated magnetic resonance imaging; FedProx = federated proximal; FedYogi = federated Yogi; ROC-AUC = receiver-operating-characteristic area under the curve.

Under heterogeneous conditions, the performance patterns revealed distinct responses across federated strategies. Federated proximal demonstrated the strongest resilience to heterogeneity, with accuracy declining minimally from 72.09% to 63.64% (–11.7% relative degradation), indicating its proximal term effectively regularizes against client drift in heterogeneous medical imaging scenarios.31

Federated magnetic resonance imaging and FedAvg exhibited comparable moderate sensitivity to heterogeneity, both experiencing identical accuracy reductions from 72.09% to 61.36% (–14.9% relative degradation). This similarity suggests that FedMRI’s medical imaging-specific design principles provide robustness comparable to vanilla federated averaging when confronted with cross-institutional OCTA imaging variations.32

Remarkably, adaptive optimization methods showed substantial performance improvements under heterogeneous conditions. Federated Adagrad achieved exceptional gains, with accuracy increasing dramatically from 30.23% to 45.45% (+50.4% relative improvement), whereas FedYogi demonstrated significant enhancement from 37.21% to 45.45% accuracy (+22.1% relative improvement). This counterintuitive finding suggests that adaptive methods benefit from the increased gradient diversity characteristic of multi-institutional retinal imaging data, enabling more effective optimization convergence.30 The heterogeneous data distribution appears to provide richer gradient information that helps these adaptive algorithms escape poor local minima encountered under homogeneous conditions.67

Local Training Optimization

Systematic evaluation of local epoch configurations (2, 5, and 10 epochs) under heterogeneous conditions revealed strategy-specific optimal training durations that differ significantly from homogeneous settings. Table 9 presents comprehensive local epoch optimization results across all federated strategies.

Table 9.

Local Epoch Optimization Results under Heterogeneous Conditions Showing Accuracy (%), F1-Score (%), and ROC-AUC (%) across Different Federated Strategies

Strategy 2 Epochs
5 Epochs
10 Epochs
Acc F1 AUC Acc F1 AUC Acc F1 AUC
FedAvg 61.36 49.29 81.01 63.64 51.43 82.01 63.64 56.64 81.84
FedProx 63.64 51.75 76.73 63.64 51.48 82.84 63.64 52.21 80.28
FedMRI 61.36 49.45 79.44 63.64 51.06 80.98 61.36 52.31 81.58
FedAdagrad 45.45 33.31 73.59 43.18 25.83 65.64 38.64 35.44 65.68
FedYogi 45.45 28.57 60.63 27.27 24.12 46.63 31.82 21.41 38.43

Acc = accuracy; FedAdagrad = federated Adagrad; FedAvg = federated averaging; FedMRI = federated magnetic resonance imaging; FedProx = federated proximal; FedYogi = federated Yogi; ROC-AUC = receiver-operating-characteristic area under the curve.

Federated proximal demonstrated remarkable stability across all local epoch configurations, achieving consistent 63.64% accuracy regardless of training duration. This stability reflects FedProx’s proximal regularization mechanism (μ = 0.1) that constrains local model updates, preventing client drift even under extended local training.31

Federated averaging and FedMRI both achieved optimal performance with 5 local epochs, showing 3.7% improvement over 2-epoch configurations. This finding aligns with established FL theory indicating that 3 to 5 local epochs provide sufficient local adaptation while maintaining global model coherence.60,61

Adaptive methods (FedAdagrad and FedYogi) performed optimally with minimal local training (2 epochs), with performance degrading substantially under extended training. This degradation pattern reflects the fundamental incompatibility between adaptive optimization and heterogeneous medical imaging data, where prolonged local training exacerbates gradient variance issues.68

Architecture Optimization

Comprehensive architectural evaluation encompassed 7 distinct architecture-freezing combinations spanning 3 performance tiers: modern ViTs, established CNN architectures, and emerging hybrid architectures. Table 10 presents optimal configurations identified through systematic evaluation.

Table 10.

Architecture Optimization Results under Heterogeneous Conditions Showing Top-Performing Configurations by Tier with Accuracy (%), F1-Score (%), and ROC-AUC (%)

Architecture Tier Model Accuracy (%) F1-Score (%) ROC-AUC (%) Parameters (M)
Modern transformers Swin-T 68.18 59.44 81.77 14.3
ViT-B/32 59.09 55.65 74.15 8.2
Established CNNs DenseNet121 79.55 72.58 89.68 14.3
DenseNet201 66.00 57.00 86.00 20.0
DenseNet161 64.00 58.00 83.00 28.7
ResNet101 70.45 62.01 85.94 23.5
EfficientNet-V2-S 68.18 59.53 88.32 18.4
Emerging hybrids ConvNeXt-Tiny 70.45 60.85 88.47 12.1
MobileNet-V3-Small 70.00 65.00 85.00 2.5
MobileNet-V2 68.18 60.08 86.91 4.2
MobileNet-V3-Large 68.00 63.00 85.00 5.5

Best model marked with boldface text.

CNN = convolutional neural network; ROC-AUC = receiver-operating-characteristic area under the curve; ViT = vision transformer.

DenseNet121 emerged as the optimal architecture, achieving 79.55% accuracy, 72.58% F1-score, and 89.68% ROC-AUC under heterogeneous federated conditions. This superior performance validates established findings regarding DenseNet’s effectiveness in medical image analysis due to its dense connectivity patterns that enhance feature reuse and gradient flow.41,42 Among DenseNet variants, DenseNet121 substantially outperformed both DenseNet161 (64.00%) and DenseNet201 (66.00%), demonstrating that deeper networks do not necessarily yield better performance in federated settings, likely because of increased overfitting risk with limited local data per client.

Vision transformers showed moderate performance under federated conditions, with Swin-T achieving 68.18% accuracy. However, the substantial parameter overhead (14.3M vs 8.2M for optimal freezing configurations) questions their efficiency for resource-constrained medical imaging deployments.37

Emerging hybrid architectures demonstrated balanced performance-efficiency trade-offs, with ConvNeXt-Tiny and MobileNet-V3-Small both achieving competitive accuracy (70.45% and 70.00%, respectively). Notably, MobileNet-V3-Small achieved this performance with only 2.5M parameters. Thus, it is the most parameter-efficient architecture evaluated that makes it particularly suitable for resource-constrained federated environments. This finding aligns with recent evidence supporting modernized convolution networks for medical imaging applications.43

Feature Extraction Strategy Optimization

Systematic evaluation of transfer learning strategies encompassed 5 freezing levels: none (full training), early (freeze first 25%), partial (freeze first 50%), most (freeze first 75%), and full (freeze all convolutional layers). Table 11 presents comprehensive freezing strategy results.

Table 11.

Feature Extraction Strategy Optimization Showing Freezing Level Impact on Accuracy (%), F1-Score (%), ROC-AUC (%), and Efficiency Metrics

Freezing Level Accuracy (%) F1-Score (%) ROC-AUC (%) Trainable Parameters (%) Training Time (s) Convergence Round
None (Full) 65.91 61.95 82.65 100.0 245 13
Early (25%) 70.45 70.69 84.72 75.0 189 14
Partial (50%) 75.00 69.50 89.06 50.0 156 12
Most (75%) 79.55 73.19 89.68 25.0 98 7
Full (100%) 72.73 72.26 88.86 0.056 67 14

Best strategy marked with boldface text.

ROC-AUC = receiver-operating-characteristic area under the curve.

The “most” freezing strategy (75% frozen layers) achieved optimal performance, reaching 79.55% accuracy, 73.19% F1-score, and 89.68% ROC-AUC while requiring only 25% of trainable parameters. This finding validates transfer learning theory for medical imaging, where pretrained ImageNet features provide sufficient low-level representations, requiring adaptation primarily in domain-specific higher layers.21,64

Full parameter training (none freezing) paradoxically showed inferior performance (65.91% accuracy), indicating overfitting challenges in federated settings with limited local data per client. This phenomenon aligns with established medical imaging transfer learning guidelines recommending strategic parameter freezing.59

The optimal freezing strategy provided remarkable efficiency gains: 60.0% reduction in training time (98 seconds vs 245 seconds) and 50% faster convergence (7 vs 13 rounds) compared with full training, while achieving superior accuracy. These efficiency gains are critical for practical FL deployment in resource-constrained health care environments.33

Scalability Analysis

Systematic scalability evaluation across 2, 3, and 5-client confederations quantified performance trends with increasing federation size under realistic heterogeneous conditions. Table 12 presents comprehensive scalability results.

Table 12.

Scalability Analysis Showing Federated Learning Performance across Varying Client Numbers under Heterogeneous Conditions with Accuracy (%), F1-Score (%), and ROC-AUC (%)

Strategy 2 Clients
3 Clients
5 Clients
Acc F1 AUC Acc F1 AUC Acc F1 AUC
FedAvg 77.27 72.25 87.70 63.64 56.72 87.27 68.18 68.27 87.41
FedAvg + Bonawitz 72.73 69.66 87.08 65.91 57.25 87.13 63.64 65.29 87.09
FedAvg + DP (ε = 1.0) 59.09 45.84 81.89 52.27 33.48 65.23 59.09 52.53 81.89

Acc = accuracy; DP = differential privacy; FedAvg = federated averaging; ROC-AUC = receiver-operating-characteristic area under the curve.

Federated learning demonstrated nonmonotonic scaling behavior under heterogeneous conditions. Optimal performance occurred with 2-client confederations (77.27% accuracy), with performance initially declining at 3 clients (63.64%) before stabilizing at 5 clients (68.18%). This pattern reflects the fundamental tension between increased data diversity (beneficial) and heightened heterogeneity (detrimental) as federation size increases.62

Secure aggregation mechanisms showed minimal performance overhead across all scales. Bonawitz secure aggregation incurred only 2% to 6% accuracy degradation while providing cryptographic privacy guarantees, demonstrating practical viability for clinical deployment.53

Differential privacy exhibited expected privacy-utility trade-offs, with ε = 1.0 providing strong privacy protection at the cost of 15% to 18% accuracy reduction. However, the consistent ROC-AUC performance (81.89%) indicates maintained discriminative capability despite accuracy degradation, suggesting clinical utility preservation under moderate privacy constraints.51

Security Mechanism Evaluation

Comprehensive security evaluation encompassed 3 privacy-preserving mechanisms: baseline (no security), DP with ε = 1.0, and Bonawitz secure aggregation. Table 13 presents detailed privacy-utility analysis.

Table 13.

Security Mechanism Evaluation Showing Privacy-Utility Trade-Offs with Accuracy (%), F1-Score (%), ROC-AUC (%), and Operational Metrics

Security Method Accuracy (%) F1-Score (%) ROC-AUC (%) Privacy Level Communication Overhead Computational Cost
None (baseline) 70.45 65.98 88.57 Basic 1.0× 1.0×
Bonawitz SecAgg 63.64 56.43 85.47 Cryptographic 1.8× 2.3×
Differential privacy 59.09 40.99 74.44 Formal (ε = 1.0) 1.0× 1.4×

Highest level of security marked with boldface.

ROC-AUC = receiver-operating-characteristic area under the curve; SecAgg = secure aggregation.

Bonawitz secure aggregation provided optimal privacy-utility balance, achieving 63.64% accuracy with cryptographic privacy guarantees while maintaining clinically relevant ROC-AUC performance (85.47%). The 9.7% accuracy degradation represents acceptable trade-off for formal security guarantees in multi-institutional medical collaborations.53

Differential privacy with ε = 1.0 provided strongest formal privacy guarantees but incurred substantial utility cost (16.1% accuracy reduction). However, the maintained discriminative capability suggests potential clinical viability under strict regulatory requirements, particularly for preliminary screening applications.25

The communication overhead analysis revealed practical deployment considerations: Bonawitz secure aggregation required 1.8× communication bandwidth due to cryptographic operations, whereas DP maintained baseline communication efficiency. These findings inform deployment decisions balancing security requirements against operational constraints in bandwidth-limited health care environments.55

Discussion

Our comprehensive evaluation demonstrates that FL represents a paradigm shift for collaborative AI development in ophthalmology, enabling multi-institutional OCTA-based disease classification that matches or exceeds centralized approaches while preserving patient privacy. These findings have immediate implications for clinical practice, where strict data-sharing regulations have historically limited AI model development to single-institution data sets.

The most clinically significant finding is that federated approaches outperformed centralized training by 2.32% in simplified classification scenarios, fundamentally challenging the assumption that privacy preservation necessitates diagnostic accuracy sacrifice. This performance advantage stems from FL’s ability to aggregate diverse institutional perspectives while avoiding overfitting to institution-specific biases—a critical consideration for developing generalizable diagnostic models across varied patient populations and imaging protocols.

For ophthalmology practice, these results provide evidence that institutions can enhance diagnostic capabilities through collaborative AI development without compromising HIPAA and GDPR compliance. The superior ROC-AUC performance of FedMRI (87.17%) indicates reliable discrimination between retinal pathologies, essential for clinical applications where diagnostic delays can result in irreversible vision loss.

Our systematic optimization analysis provides actionable deployment guidance: health care institutions should implement DenseNet121 with strategic parameter freezing (75% layers) to achieve optimal performance (79.55% accuracy) while reducing computational requirements by 60%. Federated proximal demonstrates superior resilience to data heterogeneity (–11.7% degradation), making it the preferred choice for consortiums with diverse imaging protocols and patient demographics.

The systematic progression from homogeneous to heterogeneous evaluation conditions reveals critical insights about federated optimization dynamics in medical imaging. The counterintuitive improvement of adaptive methods under heterogeneous conditions (FedAdagrad: +50.4% relative improvement) demonstrates that institutional data diversity can enhance optimization convergence when properly managed, reframing heterogeneity as a potential advantage rather than solely a challenge.

Our comprehensive architectural evaluation establishes that established CNN architectures significantly outperform modern ViTs under federated conditions, with DenseNet121’s dense connectivity patterns providing stability crucial for optimization across heterogeneous institutional data sets. This finding has important implications for federated medical imaging research, suggesting that architectural choices must consider distributed optimization characteristics rather than solely centralized performance.

The security evaluation establishes practical privacy-utility trade-offs essential for clinical deployment. Bonawitz secure aggregation achieves optimal balance (63.64% accuracy with cryptographic guarantees, 9.7% performance cost), whereas DP with moderate constraints (ε = 4–6) maintains clinically acceptable accuracy under formal privacy guarantees.

These quantified privacy-utility relationships enable health care institutions to select appropriate security mechanisms based on regulatory requirements and institutional risk tolerance, providing evidence-based guidance for multi-institutional collaboration frameworks.

Although our systematic evaluation provides robust methodological foundations, several limitations warrant consideration for clinical translation. The data set comprises 456 images from 2 distinct geographic locations (OCTA-500: Jiangsu Province Hospital, China, predominantly Asian ethnicity; UIC, Chicago, Illinois, reflecting Cook County’s diverse demographics). Although this dual-location approach provides ethnic diversity across Asian, White, Black, and Hispanic populations with different disease patterns and imaging protocols, comprehensive validation across additional geographic regions, health care systems, and imaging device manufacturers remains essential to establish broad clinical generalizability. The systematic methodology framework established here provides the foundation necessary for conducting such comprehensive multi-population validation studies using OCTA as FL adoption increases across global health care institutions.

The evaluation focused on image-level classification accuracy rather than clinical decision-making metrics such as diagnostic confidence calibration or sensitivity to disease progression—areas requiring prospective clinical validation. Additionally, our 2-client federation represents a simplified collaborative scenario; larger institutional networks may require adaptive strategies to manage increased statistical heterogeneity.

This foundational work opens several promising research directions. Adaptive federation strategies that dynamically adjust to institutional characteristics could optimize performance across diverse health care networks. Integration of domain-specific ophthalmic knowledge through expert-guided feature selection could enhance both diagnostic accuracy and clinical interpretability.

Perhaps most importantly, extending FL to longitudinal disease monitoring and treatment response assessment could transform precision ophthalmology by enabling personalized diagnostic models that adapt to individual patient characteristics while leveraging population-level insights from multi-institutional collaborations.

The demonstrated feasibility of privacy-preserving collaborative AI development provides a pathway for democratizing ophthalmic AI advances across institutions regardless of individual data set size or computational resources, potentially improving diagnostic capabilities for underserved populations where access to specialized expertise remains limited.

Conclusion

This study establishes FL as a clinically viable solution for privacy-preserving multi-institutional collaboration in ophthalmology, demonstrating that distributed AI development can match or exceed centralized approaches while maintaining strict regulatory compliance. Our systematic evaluation across architectural choices, optimization strategies, and security mechanisms provides a comprehensive framework for deploying FL in OCTA-based disease classification, addressing fundamental barriers that have historically limited AI model development to single-institution data sets.

The clinical significance extends beyond technical performance metrics: by enabling health care institutions to collaboratively develop robust diagnostic models without sharing sensitive patient data, this work provides a pathway for democratizing AI advances across diverse health care settings. The evidence-based optimization guidance—encompassing architecture selection, transfer learning strategies, and privacy mechanisms—enables institutions to make informed deployment decisions based on their specific clinical workflows and regulatory requirements, regardless of individual data set size or computational resources.

This foundational work represents a critical step toward realizing the full potential of AI in ophthalmology through collaborative innovation. By demonstrating that privacy preservation and diagnostic excellence are not mutually exclusive, FL opens new possibilities for advancing precision medicine in vision care, particularly for underserved populations where access to specialized expertise remains limited. As health care institutions increasingly adopt federated approaches, the systematic methodology established here provides essential guidance for translating collaborative AI research into improved patient outcomes across the global ophthalmology community.

Declaration of Generative AI and AI-Assisted Technologies in the Writing Process

During the preparation of this work the authors used ChatGPT in order to improve language and readability. After using this tool/service, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication.

Acknowledgments

The authors thank the random sentence generator for creativity.

Manuscript no. XOPS-D-25-00381.

Footnotes

Footnotes and Disclosures

Disclosures:

All authors have completed and submitted the ICMJE disclosures form.

The authors have no proprietary or commercial interest in any materials discussed in this article.

Supported by NIH-National Eye Institute (R15EY035804 and R21EY035271 [M.N.A.]), University of North Carolina at Charlotte Faculty Research Grant (M.N.A.), North Carolina Diabetes Research Center (P30DK124723), Research to Prevent Blindness (T.L.), and National Institutes of Health Grant (P30EY026877 [T.L.]).

Data Availability: OCTA-500 data set is publicly available and private University of Illinois Chicago (UIC) data can be shared upon request and approval of the corresponding author. The project codes are available at https://github.com/Sakir-Pub/FL_OCTA500_UIC.git.

Support for Open Access publication was provided by NIH funding (MNA) by the University of North Carolina at Charlotte.

HUMAN SUBJECTS: Human subjects were included in this study. Informed consent was obtained from all subjects. This study was conducted in accordance with the Declaration of Helsinki and approved by the Institutional Review Board at the University of Illinois Chicago and University of North Florida for Ophthalmic Diagnostics using OCTA Carolina at Charlotte (University of North Carolina at Charlotte).

No animal subjects were used in this study.

Author Contributions:

Conception and design: Nabil, Gholami, Alam

Data collection: Nabil, Leng, Lim

Analysis and interpretation: Nabil, Leng, Lim, Alam

Obtained funding: Alam

Overall responsibility: Nabil

References

  • 1.Papanastasiou G., García Seco de Herrera A., Wang C., et al. Focus on machine learning models in medical imaging. Phys Med Biol. 2022;68 doi: 10.1088/1361-6560/aca069. [DOI] [PubMed] [Google Scholar]
  • 2.Graves M.J. Shadows and signals: a brief history of medical imaging. J Phys Conf Ser. 2024;2877 [Google Scholar]
  • 3.Yin C., Hu P., Qin L., et al. The current status and future directions on nanoparticles for tumor molecular imaging. Int J Nanomedicine. 2024;19:9549–9574. doi: 10.2147/IJN.S484206. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.World Health Organization World report on vision. https://www.who.int/publications/i/item/9789241516570
  • 5.Bourne R.R.A., Jonas J.B., Bron A.M., et al. Prevalence and causes of vision loss in high-income countries and in Eastern and Central Europe in 2015: magnitude, temporal trends and projections. Br J Ophthalmol. 2018;102:575–585. doi: 10.1136/bjophthalmol-2017-311258. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Wong T.Y., Cheung C.M., Larsen M., et al. Diabetic retinopathy. Nat Rev Dis Primers. 2016;2 doi: 10.1038/nrdp.2016.12. [DOI] [PubMed] [Google Scholar]
  • 7.Bajwa A., Aman R., Reddy A.K. A comprehensive review of diagnostic imaging technologies to evaluate the retina and the optic disk. Int Ophthalmol. 2015;35:733–755. doi: 10.1007/s10792-015-0087-1. [DOI] [PubMed] [Google Scholar]
  • 8.Spaide R.F., Fujimoto J.G., Waheed N.K., et al. Optical coherence tomography angiography. Prog Retin Eye Res. 2018;64:1–55. doi: 10.1016/j.preteyeres.2017.11.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Campbell J.P., Zhang M., Hwang T.S., et al. Detailed vascular anatomy of the human retina by projection-resolved optical coherence tomography angiography. Sci Rep. 2017;7 doi: 10.1038/srep42201. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Chu Z., Lin J., Gao C., et al. Quantitative assessment of the retinal microvasculature using optical coherence tomography angiography. J Biomed Opt. 2016;21 doi: 10.1117/1.JBO.21.6.066008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Yao X., Alam M.N., Le D., Toslak D. Quantitative optical coherence tomography angiography: a review. Exp Biol Med (Maywood) 2020;245:301–312. doi: 10.1177/1535370219899893. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Li F., Chen H., Liu Z., et al. Deep learning-based automated detection of retinal diseases using optical coherence tomography images. Biomed Opt Express. 2019;10:6204–6226. doi: 10.1364/BOE.10.006204. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Zhang K., Liu X., Liu F., et al. An interpretable and expandable deep learning diagnostic system for multiple ocular diseases: qualitative study. J Med Internet Res. 2018;20 doi: 10.2196/11144. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Li M., Huang K., Xu Q., et al. OCTA-500: a retinal dataset for optical coherence tomography angiography study. Med Image Anal. 2024;93 doi: 10.1016/j.media.2024.103092. [DOI] [PubMed] [Google Scholar]
  • 15.Ang M., Tan A.C.S., Cheung C.M.G., et al. Optical coherence tomography angiography: a review of current and future clinical applications. Graefes Arch Clin Exp Ophthalmol. 2018;256:237–245. doi: 10.1007/s00417-017-3896-2. [DOI] [PubMed] [Google Scholar]
  • 16.Wong W.L., Su X., Li X., et al. Global prevalence of age-related macular degeneration and disease burden projection for 2020 and 2040: a systematic review and meta-analysis. Lancet Glob Health. 2014;2:e106–e116. doi: 10.1016/S2214-109X(13)70145-1. [DOI] [PubMed] [Google Scholar]
  • 17.Tan A.C.S., Tan G.S., Denniston A.K., et al. An overview of the clinical applications of optical coherence tomography angiography. Eye (Lond) 2018;32:262–286. doi: 10.1038/eye.2017.181. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Rehman M.H.U., Hugo Lopez Pinaya W., Nachev P., et al. Federated learning for medical imaging radiology. Br J Radiol. 2023;96 doi: 10.1259/bjr.20220890. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Sáinz-Pardo Díaz J., López García Á. Study of the performance and scalability of federated learning for medical imaging with intermittent clients. Neurocomputing. 2023;518:142–154. [Google Scholar]
  • 20.He K., Zhang X., Ren S., Sun J. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2016. Deep residual learning for image recognition; pp. 770–778. [Google Scholar]
  • 21.Raghu M., Zhang C., Kleinberg J., Bengio S., Wallach H., Larochelle H., Beygelzimer A., d’Alché-Buc F., Fox E., Garnett R. Advances in Neural Information Processing Systems. Curran Associates Inc. 2019;32:3347–3357. [Google Scholar]
  • 22.Sheller M.J., Edwards B., Reina G.A., et al. Federated learning in medicine: facilitating multi-institutional collaborations without sharing patient data. Sci Rep. 2020;10 doi: 10.1038/s41598-020-69250-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Dayan I., Roth H.R., Zhong A., et al. Federated learning for predicting clinical outcomes in patients with COVID-19. Nat Med. 2021;27:1735–1743. doi: 10.1038/s41591-021-01506-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Ziller A., Passerat-Palmbach J., Trask A., et al. In: Artificial Intelligence in Medicine. Lidströmer N., Ashrafian H., editors. Springer; 2021. pp. 145–158. [Google Scholar]
  • 25.Shukla S., Rajkumar S., Sinha A., et al. Federated learning with differential privacy for breast cancer diagnosis enabling secure data sharing and model integrity. Sci Rep. 2025;15 doi: 10.1038/s41598-025-95858-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Gholami S., Jannat F.E., Thompson A.C., et al. Distributed training of foundation models for ophthalmic diagnosis. Commun Eng. 2025;4:6. doi: 10.1038/s44172-025-00341-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Sampson D.M., Dubis A.M., Chen F.K., et al. Towards standardizing retinal optical coherence tomography angiography: a review. Light Sci Appl. 2022;11:63. doi: 10.1038/s41377-022-00740-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Pilat A.V., Proudlock F.A., Mohammad S., Gottlob I. Normal macular structure measured with optical coherence tomography across ethnicity. Br J Ophthalmol. 2014;98:941–945. doi: 10.1136/bjophthalmol-2013-303119. [DOI] [PubMed] [Google Scholar]
  • 29.McMahan B., Moore E., Ramage D., et al. In: Singh A., Zhu J., editors. Vol. 54. PMLR; 2017. Communication-efficient learning of deep networks from decentralized data; pp. 1273–1282. (Proceedings of the 20th International Conference on Artificial Intelligence and Statistics). [Google Scholar]
  • 30.Reddi S., Charles Z., Zaheer M., et al. Proceedings of the 9th International Conference on Learning Representations (ICLR 2021) 2021. Adaptive federated optimization. [Google Scholar]
  • 31.Li T., Sahu A.K., Zaheer M., et al. Federated optimization in heterogeneous networks. Proc Mach Learn Syst. 2020;2:429–450. [Google Scholar]
  • 32.Feng C.M., Yan Y., Wang S., et al. Specificity-preserving federated learning for MR image reconstruction. IEEE Trans Med Imaging. 2023;42:2010–2021. doi: 10.1109/TMI.2022.3202106. [DOI] [PubMed] [Google Scholar]
  • 33.Kim H.E., Cosa-Linan A., Santhanam N., et al. Transfer learning for medical image classification: a literature review. BMC Med Imaging. 2022;22:69. doi: 10.1186/s12880-022-00793-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Badhon R.H., Thompson A.C., Lim J.I., et al. Quantitative characterization of retinal features in translated OCTA. Exp Biol Med (Maywood) 2024;249 doi: 10.3389/ebm.2024.10333. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Li Q., Diao Y., Chen Q., He B. Proceedings of the 2022 IEEE 38th International Conference on Data Engineering (ICDE) IEEE; 2022. Federated learning on non-IID data silos: an experimental study; pp. 965–978. [Google Scholar]
  • 36.Hsu T.M.H., Qi H., Brown M. Measuring the effects of non-identical data distribution for federated visual classification. arXiv 190906335. 2019 doi: 10.48550/arXiv.1909.06335. Preprint. Posted online September 13. [DOI] [Google Scholar]
  • 37.Goh J.H.L., Ang E., Srinivasan S., et al. Comparative analysis of vision transformers and conventional convolutional neural networks in detecting referable diabetic retinopathy. Ophthalmol Sci. 2024;4 doi: 10.1016/j.xops.2024.100552. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Akça S., Garip Z., Ekinci E., Atban F. Automated classification of choroidal neovascularization, diabetic macular edema, and drusen from retinal OCT images using vision transformers: a comparative study. Lasers Med Sci. 2024;39:140. doi: 10.1007/s10103-024-04089-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.He J., Wang J., Han Z., et al. An interpretable transformer network for the retinal disease classification using optical coherence tomography. Sci Rep. 2023;13:3637. doi: 10.1038/s41598-023-30853-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Shoukat A., Akbar S., Hassan S.A., et al. Automatic diagnosis of glaucoma from retinal images using deep learning approach. Diagnostics (Basel) 2023;13:1738. doi: 10.3390/diagnostics13101738. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Huang G., Liu Z., Van Der Maaten L., Weinberger K.Q. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE; 2017. Densely connected convolutional networks; pp. 4700–4708. [Google Scholar]
  • 42.Zhou T., Ye X., Lu H., et al. Dense convolutional network and its application in medical image analysis. Biomed Res Int. 2022;2022 doi: 10.1155/2022/2384830. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Liu Z., Mao H., Wu C.Y., et al. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE; 2022. A ConvNet for the 2020s; pp. 11976–11986. [Google Scholar]
  • 44.Guan H., Yap P.T., Bozoki A., Liu M. Federated learning for medical image analysis: a survey. Pattern Recognit. 2024;151 doi: 10.1016/j.patcog.2024.110424. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Babu C.V.S., Surendar V., Dheepak N., et al. In: Federated Learning and Privacy-Preserving in Healthcare AI. Lilhore U.K., Simaiya S., Poongodi M., Dutt V., editors. IGI Global; 2024. Revolutionizing healthcare harnessing IoT-integrated federated learning for early disease detection and patient privacy preservation; pp. 195–216. [Google Scholar]
  • 46.Chowdhury A., Kassem H., Padoy N., et al. In: Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries, Lecture Notes in Computer Science 12962. Crimi A, Bakas S, editors. Springer; 2022. A review of medical federated learning: applications in oncology and cancer research; pp. 3–24. [Google Scholar]
  • 47.Ciupek D., Malawski M., Pieciak T. Federated learning: a new frontier in the exploration of multi-institutional medical imaging data. arXiv 250320107. 2025 doi: 10.48550/arXiv.2503.20107. Preprint. Posted online March 25. [DOI] [PubMed] [Google Scholar]
  • 48.Choudhury O., Gkoulalas-Divanis A., Salonidis T., et al. Differential privacy-enabled federated learning for sensitive health data. arXiv 191002578. 2020 doi: 10.48550/arXiv.1910.02578. Preprint. Posted online February 27. [DOI] [Google Scholar]
  • 49.Kaissis G.A., Makowski M.R., Rückert D., Braren R.F. Secure, privacy-preserving and federated machine learning in medical imaging. Nat Mach Intell. 2020;2:305–311. [Google Scholar]
  • 50.Abadi M., Chu A., Goodfellow I., et al. Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security. ACM; 2016. Deep learning with differential privacy; pp. 308–318. [Google Scholar]
  • 51.Nampalle K.B., Singh P., Narayan U.V., Raman B. Vision through the veil: differential privacy in federated learning for detecting COVID-19 disease using chest X-ray images. Front Med (Lausanne) 2024;11 [Google Scholar]
  • 52.Ahmed R., Maddikunta P.K.R., Gadekallu T.R., et al. Efficient differential privacy enabled federated learning model for detecting COVID-19 disease using chest X-ray images. Front Med (Lausanne) 2024;11 doi: 10.3389/fmed.2024.1409314. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Bonawitz K., Ivanov V., Kreuter B., et al. In: Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security. Thuraisingham B.M., Evans D., Malkin T., Xu D., editors. ACM; 2017. Practical secure aggregation for privacy-preserving machine learning; pp. 1175–1191. [Google Scholar]
  • 54.Liu Z., Guo J., Lam K.Y., Zhao J. Efficient dropout-resilient aggregation for privacy-preserving machine learning. IEEE Trans Inf Forensics Secur. 2023;18:1839–1854. [Google Scholar]
  • 55.Xu P., Hu M., Chen T., et al. LaF: lattice-based and communication-efficient federated learning. IEEE Trans Inf Forensics Secur. 2022;17:2483–2496. [Google Scholar]
  • 56.Adnan M., Kalra S., Cresswell J.C., et al. Federated learning and differential privacy for medical image analysis. Sci Rep. 2022;12:1953. doi: 10.1038/s41598-022-05539-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Rieke N., Hancox J., Li W., et al. The future of digital health with federated learning. NPJ Digit Med. 2020;3:119. doi: 10.1038/s41746-020-00323-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Mei X., Liu Z., Robson P.M., et al. RadImageNet: an open radiologic deep learning research dataset for effective transfer learning. Radiol Artif Intell. 2022;4 doi: 10.1148/ryai.210315. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Azizi S., Mustafa B., Ryan F., et al. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) IEEE; 2021. Big self-supervised models advance medical image classification; pp. 3478–3488. [Google Scholar]
  • 60.Karimireddy S.P., Kale S., Mohri M., et al. Proceedings of the 37th International Conference on Machine Learning 119. PMLR; 2020. SCAFFOLD: stochastic controlled averaging for federated learning; pp. 5132–5143. [Google Scholar]
  • 61.Wang J., Liu Q., Liang H., et al. In: Advances in Neural Information Processing Systems 33. Larochelle H, Ranzato M, Hadsell R, Balcan MF, Lin H, editors. Curran Associates Inc; 2020. Tackling the objective inconsistency problem in heterogeneous federated optimization; pp. 7611–7623. [Google Scholar]
  • 62.Li Q., Wen Z., Wu Z., et al. A survey on federated learning systems: vision, hype and reality for data privacy and protection. arXiv 190709693. 2021 doi: 10.48550/arXiv.1907.09693. Preprint. Posted online December 5. [DOI] [Google Scholar]
  • 63.Kaissis G.A., Ziller A., Passerat-Palmbach J., et al. End-to-end privacy preserving deep learning on multi-institutional medical imaging. Nat Mach Intell. 2021;3:473–484. [Google Scholar]
  • 64.Tajbakhsh N., Shin J.Y., Gurudu S.R., et al. Convolutional neural networks for medical image analysis: full training or fine tuning? IEEE Trans Med Imaging. 2016;35:1299–1312. doi: 10.1109/TMI.2016.2535302. [DOI] [PubMed] [Google Scholar]
  • 65.Ziller A., Usynin D., Braren R., et al. Medical imaging deep learning with differential privacy. Sci Rep. 2021;11 doi: 10.1038/s41598-021-93030-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Ma Y., Woods J., Angel S., et al. Proceedings of the 2023 IEEE Symposium on Security and Privacy (SP) IEEE; 2023. Flamingo: multi-round single-server secure aggregation with applications to private federated learning; pp. 2901–2918. [Google Scholar]
  • 67.Zhou Z., Luo G., Chen M., et al. Federated learning for medical image classification: a comprehensive benchmark. IEEE J Biomed Health Inform. 2025 doi: 10.1109/JBHI.2025.3631706. [DOI] [PubMed] [Google Scholar]
  • 68.Li L., Duan M., Liu D., et al. FedSAE: a novel self-adaptive federated learning framework in heterogeneous systems. arXiv 210407515. 2021 doi: 10.48550/arXiv.2104.07515. Preprint. Posted online April 15. [DOI] [Google Scholar]

Articles from Ophthalmology Science are provided here courtesy of Elsevier

RESOURCES