Abstract
Purpose
There has been significant progress in detecting Alzheimer's disease (AD) using retinal imaging. We developed an ensemble learning-based deep learning (DL) model, integrating different inputs from OCT for the detection of AD-dementia and early AD.
Design
A retrospective multicenter case-control study.
Participants
A total of 190 participants with AD-dementia and 623 cognitively normal controls were recruited from 2 cohorts in Hong Kong and Singapore as the training and internal validation sets. A total of 46 participants with AD-dementia, 79 participants with mild cognitive impairment (MCI), and 52 cognitively normal controls from 2 cohorts with amyloid-β status identified from positron emission tomography (PET) available in Hong Kong and Singapore as External-1 and External-2, respectively.
Methods
We developed DL models for identifying AD-dementia versus cognitively normal and also tested the proposed ensemble model for classifying MCI (symptom-based) and AD-MCI (PET-based). Inputs were generated from a commercially available OCT device (Cirrus HD-OCT, Carl Zeiss Meditec, Inc), including optic nerve head (ONH)-centered and macula-centered en face images along with retinal nerve fiber layer thickness and deviation maps, ganglion cell-inner plexiform layer thickness and deviation maps, and macular thickness map. Then, to integrate multiple algorithms and inputs simultaneously, we developed an ensemble model that integrated 2 base DL models—ONH model and the macula model, developed by OCT inputs from the ONH and macula regions, respectively—to provide a unified classification via majority voting.
Main Outcome Measures
Discriminative performance of the ensemble model for detecting AD-dementia, MCI, and AD-MCI.
Results
For detecting AD-dementia, the ensemble model achieved the area under the receiver operating characteristic curve (AUROC) of 0.943 (95% confidence interval, 0.906–0.980), 0.786 (95% confidence interval, 0.673–0.899), and 0.795 (95% confidence interval, 0.716–0.874) in the internal validation, External-1, and External-2, respectively. For detecting AD-MCI defined by PET biomarkers, the ensemble model achieved AUROCs of 0.787 (95% confidence interval, 0.643–0.931) and 0.791 (95% confidence interval, 0.694–0.888) in the External-1 and External-2, respectively.
Conclusions
Our proposed ensemble model, integrating multiple base models and inputs from OCT analysis, demonstrates strong potential for leveraging OCT imaging in detecting both AD-dementia and early-stage AD, enabling opportunistic screening for AD during ophthalmic visits.
Financial Disclosure(s)
Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
Keywords: Alzheimer's disease, OCT, Ensemble learning, Artificial intelligence, Opportunistic screening
Alzheimer's disease (AD) remains the leading cause of dementia worldwide.1 Recent approval of 3 drugs, aducanumab, lecanemab, and donanemab, by the US Food and Drug Administration signifies a paradigm shift in the management of AD from a sole symptomatic treatment approach to the exploration of disease-modifying therapies.2 Importantly, the clinical benefits of these drugs are for individuals with early AD, including mild cognitive impairment (MCI) or mild dementia due to AD.3,4 In addition, there is robust evidence that nearly 50% of dementia cases could be prevented or delayed by optimizing a set of modifiable risk factors, such as physical activities and cardiovascular and metabolic morbidities.5 Thus, there is an increasing need to develop simple, scalable, and feasible screening tools to identify early preclinical disease in the primary care or community settings.
In this regard, the potential of using retinal imaging to screen for dementia has been proposed. There is evidence of neurovascular changes, such as changes in the cortical blood flow,6 decreased power of brain oxygenation oscillations,7 and neuronal loss in the hippocampus and cerebral neocortex,8 before the manifestation of clinical symptoms. These changes can be reflected in the retina, which has similar embryology, anatomy, and physiology to the brain.9, 10, 11 OCT has enabled the detailed morphological assessment and quantification of individual retinal layers, including the retinal nerve fiber layer (RNFL), ganglion cell layer, and inner plexiform layer, noninvasively.12, 13, 14, 15 Studies have already shown that patients with AD had reduced RNFL, macula thickness (MT), and macular ganglion cell-inner plexiform layer (GCIPL) thicknesses as measured by OCT.16, 17, 18, 19, 20, 21 The quantitative segmental analysis of retinal layers that are associated with AD demonstrates its potential use as an indicator of neurodegeneration in the brain. Meanwhile, en face images captured along with OCT would allow assessment of retinal vasculature associated with AD.22, 23, 24, 25
The evolving field of artificial intelligence (AI) and deep learning (DL) presents unique opportunities for detecting AD from retinal imaging.26, 27, 28, 29, 30, 31, 32 A notable field of AI and DL is ensemble learning, which combines multiple DL models, achieving better predictive performance when processing highly diversified data inputs.33,34 This approach can also reduce the overall risk of overfitting and is more robust to handle imbalanced datasets (i.e., the number of images from cognitively normal subjects greatly exceeds those from patients with AD), as well as handling noise and outliers that are commonly observed in medical imaging, which can potentially improve the model performance in clinical settings.
Ensemble learning models for OCT could be developed given the comprehensive information provided by OCT (e.g., structural and vascular changes in macula and optic disc scans). Our study aims to develop a novel ensemble learning DL model by integrating different OCT inputs to maximize AD-related retinal features for identifying AD-dementia and AD-MCI.
Methods
This retrospective, multicenter, case-control study was approved by the human ethics boards of the Joint Chinese University of Hong Kong-New Territories East Cluster and Hong Kong Hospital Authority Kowloon Central Cluster Clinical Research Ethics Committee, Hong Kong, as well as local research ethics committees in each center. All studies were conducted following the Declaration of Helsinki. Written informed consent was exempted for the datasets that involved only retrospective analysis using fully anonymized OCT images. We analyzed and reported our results following the DECIDE-AI guideline.35
Dataset Collection
We developed DL models for identifying AD-dementia, MCI, or AD-MCI. The input was different images extracted from the OCT analysis report generated by the Cirrus HD-OCT device (Carl Zeiss Meditec, Inc), including RNFL analysis report from optic nerve head (ONH)-centered and GCIPL and MT analysis from macula-centered cube scans, respectively. The ONH-centered cube scanning protocol was based on a volumetric scan of a 6 × 6 mm2 area centered on the optic disc, collecting information of 1024 (depth) × 200 × 200 points. The macula-centered cube scanning protocols were based on a volumetric scan of a 6 × 6 mm2 area centered on the macula, collecting information of 1024 (depth) × 512 × 128 points. The RNFL and GCIPL analysis reports contain both thickness and deviation maps, while the MT analysis report only contains a thickness map. The ONH-centered and macula-centered paired en face images were also collected. All the reports and en face images were generated by the reviewer software version 11.5.2.54532 (Carl Zeiss Meditec, Inc).
For model training and internal validation, we retrospectively collected OCT reports and the paired en face images of participants with AD-dementia or who were cognitively normal from 2 studies, the Study of Novel Retinal Imaging Biomarkers for Cognitive Decline in Hong Kong and the Harmonization Cohort Study in Singapore, respectively.
For model external testing, we retrospectively collected OCT reports and the paired en face images from 2 unseen datasets, External-1 (the Chinese University of Hong Kong - Screening for Early AlzhEimer's DiseaSe study in Hong Kong) and External-2 (the Amyloid Brain and Retinal Imaging Study in Singapore), including participants with AD-dementia, MCI, or cognitively normal, also with labels of amyloid-β positive/negative from positron emission tomography (PET).
We also collected data from 100 cognitively normal subjects using another OCT vendor (DRI OCT Triton; Topcon Inc) (External-3) to assess cross-vendor adaptability.
Ground Truth Label
Alzheimer's disease-dementia (symptom-based): individuals with AD-dementia fulfilled Diagnostic and Statistical Manual of Mental Disorders, fourth edition or fifth edition criteria for dementia syndrome (Alzheimer's type) and the National Institute of Neurological and Communicative Disorders and Stroke and the AD and Related Disorders Association criteria for probable or possible AD.
Cognitively normal (symptom-based): subjects without dementia or who had no cognitive impairment were defined as having no objective impairment on the neuropsychological assessment.
Mild cognitive impairment (symptom-based): individuals with MCI were defined as impairment on neuropsychological assessment but did not meet the criteria for dementia according to the Diagnostic and Statistical Manual of Mental Disorders, fourth edition or Diagnostic and Statistical Manual of Mental Disorders, fifth edition.28
AD-MCI (PET-Based): subjects with MCI who were amyloid-β positive
Normal (PET-based): cognitively normal subjects who were amyloid-β negative.
In External-1, PET/CT imaging was performed at the Department of Nuclear Medicine & PET of Hong Kong Sanatorium & Hospital, Hong Kong. Positron emission tomography images were obtained at 35 minutes postinjection. Amyloid-β positivity was defined as (1) increased 11C-PIB uptake was visually observed in regions known to have amyloid-beta deposits in patients with AD-dementia, e.g., frontal lobe, parietal lobe, lateral temporal lobe, posterior cingulate, precuneus, or caudate; or (2) global retention ≥1.42.
In External-2, PET-magnetic resonance imaging was performed on an molecular magnetic resonance synchronous PET/magnetic resonance scanner (Siemens Healthcare GmbH) at the Clinical Imaging Research Centre of the National University of Singapore. Positron emission tomography images were obtained at 40 minutes postinjection. Amyloid-β positivity was defined from visual interpretation by experts from 6 AD-specific regions: frontal lobe, parietal lobe, temporal lobe, anterior cingulate, and precuneus/posterior cingulate.28
Image Preprocessing
The image-based inputs for DL models consisted of 2 types of data: (1) thickness maps and deviation maps extracted from the reports of RNFL analysis, GCIPL analysis, and MT analysis, and (2) ONH-centered and macula-centered en face images (Fig 1). All OCT scans were evaluated by trained graders for image quality assessment, and only gradable images were included. The image-based inputs were paired at the participant level to train our DL models. We preprocessed the inputs with data normalization and data augmentation such as flipping, shifting, scale rotation, distortion, and red-green-blue shift, for all the thickness maps, deviation maps, and en face images. We additionally applied the loop-up table function to enhance the luminance and contrast of the original en face images. Each image was resized to 450 × 450 pixels to ensure reasonable quality of input images while keeping the model complexity relatively low.
Figure 1.
Different image-based inputs generated from OCT conventional reports and en face raw images: (1) ONH-centered input: RNFL analysis report, including RNFL thickness maps (red boxes) and deviation maps (yellow boxes), and ONH-centered en face image; and (2) macula-centered input: GCIPL analysis report, including GCIPL thickness maps (green boxes) and deviation maps (blue boxes); MT analysis report, including MT maps (gray boxes); and macula-centered en face image. GCIPL = ganglion cell-inner plexiform layer; MT = macula thickness; OD = oculus dexter; ONH = optic nerve head; OS = oculus sinister; RNFL = retinal nerve fiber layer.
Development of the Ensemble Model
We first developed 2 base DL models, namely the ONH model and the macula model, to identify AD-dementia from cognitively normal (Fig S2, available at www.ophthalmologyscience.org). The ONH model was trained based on RNFL thickness map, RNFL deviation map, and ONH-centered en face image. The macula model used GCIPL thickness map, GCIPL deviation map, MT thickness map, and macula-centered en face image. We then integrated the 2 base models into an ensemble model using the ensemble learning technique, which harmonized different inputs to provide a unified classification (Fig 2). Specifically, we constructed an ensemble feature layer, which is a summation of multiple features at the channel level, between the 2 base models to select representative features for the classification. After the reconstructed features were generated from the 2 base models, the ensemble feature layer aggregated and transmitted them to the combined classifier as well as their respective classifiers. Finally, the classification results of all classifiers were used to generate the final classification result through majority voting. We trained the ensemble model using pretrained parameters from the 2 base models and froze all layers to optimize only the ensemble feature layer and the classifier. The training process also used the Adam optimization technique36 and ran through 50 epochs with a batch size of 4 and a learning rate of 1 × 10–5 initially, with decay every 10 epochs. Gradient-weighted class activation mapping was used to generate heatmaps. More details are shown in the Supplementary material (available at www.ophthalmologyscience.org).
Figure 2.
Illustrates the structure of the ensemble model. It integrates the 2 base models, the ONH model and the macula model, to provide a single and unified classification. In addition, the embodiment provides an ensemble feature layer between these 2 networks to select representative features for the classification. The ensemble feature layer is the summation of multiple features at the channel level. After reconstructed features are generated from these 2 networks, the ensemble feature layer aggregates and transmits them to the ensemble classifier as well as their respective (e.g., first local and second local) classifiers. In the last step, the results of all (e.g., first local, ensemble, and second local) classifiers are used to generate the final classification of AD-dementia or cognitively normal through majority voting. AD = Alzheimer's disease; DSBN = domain specific batch normalization; GCIPL = ganglion cell-inner plexiform layer; MT = macula thickness; ONH = optic nerve head; RNFL = retinal nerve fiber layer.
Domain Adaptation
When testing the DL models on unseen datasets, potential variances, such as ethnicity, ocular pathologies, and different OCT vendors, could lead to notable disparities in the performance. Thus, we adopted the domain adaptation approach to address the issue of dataset discrepancies and further enhance the generalizability. Specifically, the training dataset was defined as the “source domain,” and the external test dataset was defined as the “target domain.” We first trained the model by supervised learning using labels on the source domain. Subsequently, pseudo-labels of the target domain were generated with the pretrained model. In the fusion network, the convolutional layer extracted shared features from different domains for interdomain correlation feature learning. The batch normalization layer was established as 2 branches that independently accept feature extraction from the source and target domains using domain-specific batch normalization. The domain adaptation technique can minimize the gap between the 2 domains in the feature space and allow the performance of the depth model in the target domain to approximate or even be identical to the original domain.
Statistical Analysis
The statistical analyses were performed by RStudio version 2023.12.1 + 402 (2022 by Posit Software, PBC). One-way analysis of variance and chi-squared tests were performed for numerical and categorical data, respectively, to analyze the demographic characteristics of all the participants and data variances of different datasets.
The area under the receiver operating characteristic curve (AUROC), accuracy, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) with a 95% confidence interval were calculated to evaluate the discriminative performance. The cut-off point was the largest Youden Index in each dataset. The Delong test was used to compare the AUROCs of different models. All the hypotheses tested were 2-sided, and a P value of <0.05 was considered statistically significant.
Results
Characteristics of the Datasets for Model Training, Internal Validation, and External Testing
Table 1 summarizes the datasets and respective demographic information, including 236 participants with AD-dementia, 79 participants with MCI, and 675 cognitively normal participants. For model training, we used 3039 paired RNFL, GCIPL, MT reports, ONH-centered, and macula-centered en face images, respectively, including 1239 images from 170 patients with AD-dementia and 1800 images from 593 cognitively normal participants. Internal validation contained 189 sets of OCT data from 20 patients with AD-dementia and 30 cognitively normal participants. External-1 contained 86 sets of OCT data from 12 patients with AD-dementia and 26 cognitively normal participants. External-2 contained 125 sets of OCT data from 34 patients with AD-dementia and 26 cognitively normal participants. There were no significant differences in gender and laterality of the eye for all the datasets (all P values > 0.05), while participants with AD-dementia were significantly older than cognitively normal participants in the training (76.1 ± 7.5 years vs 48.3 ± 14.3 years, P < 0.001) and the internal validation sets (77.5 ± 7.2 years vs 47.4 ± 15.6 years, P < 0.001).
Table 1.
The Characteristics of the Participants and Images
| MCI | AD-Dementia | Cognitively Normal | P Value | |
|---|---|---|---|---|
| Training Data | ||||
| No. of OCT input | ∖ | 1239 | 1800 | ∖ |
| No. of participants | ∖ | 170 | 593 | ∖ |
| Gender (male/female) | ∖ | 61/109 | 222/371 | 0.787 |
| Age, yrs (mean ± SD) | ∖ | 76.1 ± 7.5 | 48.3 ± 14.3 | <0.001 |
| No. of eyes | ∖ | 306 | 1116 | ∖ |
| Eye (right/left) | ∖ | 156/150 | 556/560 | 0.747 |
| Internal validation | ||||
| No. of OCT input | ∖ | 113 | 76 | ∖ |
| No. of participants | ∖ | 20 | 30 | ∖ |
| Gender (male/female) | ∖ | 9/11 | 6/24 | 0.11 |
| Age, yrs (mean ± SD) | ∖ | 77.5 ± 7.2 | 47.4 ± 15.6 | <0.001 |
| No. of eyes | ∖ | 34 | 54 | ∖ |
| Eye (right/left) | ∖ | 17/17 | 29/25 | 0.828 |
| External-1 | ||||
| No. of OCT input | 33 | 27 | 59 | ∖ |
| No. of participants | 15 | 12 | 26 | ∖ |
| Gender (male/female) | 5/10 | 7/5 | 11/15 | 0.49 |
| Age, yrs (mean ± SD) | 71.7 ± 8.4 | 68.2 ± 7.9 | 67.8 ± 7.5 | 0.88 |
| No. of eyes | 30 | 24 | 52 | ∖ |
| Eye (right/left) | 15/15 | 12/12 | 26/26 | >0.99 |
| External-2 | ||||
| No. of OCT input | 121 | 72 | 53 | ∖ |
| No. of participants | 64 | 34 | 26 | ∖ |
| Gender (male/female) | 33/31 | 7/27 | 10/16 | 0.16 |
| Age, yrs (mean ± SD) | 75.3 ± 6.3 | 76.0 ± 7.4 | 75.3 ± 3.6 | 0.66 |
| No. of eyes | 106 | 59 | 46 | ∖ |
| Eye (right/left) | 50/56 | 29/30 | 21/25 | 0.844 |
AD = Alzheimer's disease; MCI = mild cognitive impairment; SD = standard deviation.
Bold values are the ones significantly higher than the comparator.
To further test the proposed ensemble model’s performance in identifying symptom-based MCI and PET-based AD-MCI, we additionally collected 33 and 121 OCT inputs from 15 and 64 participants with MCI in External-1 and External-2, respectively.
Performance of the Proposed Ensemble Model
Table 2 shows the performance of the proposed ensemble model. For AD-dementia versus cognitively normal, the ensemble model achieved an AUROC of 0.943, an accuracy of 90.5%, a sensitivity of 93.3%, a specificity of 87.7%, a PPV of 91.5%, and an NPV of 90.0% in internal validation; an AUROC of 0.786, an accuracy of 80.3%, a sensitivity of 68.0%, a specificity of 87.5%, a PPV of 69.6%, and an NPV of 85.3% in External-1; and an AUROC of 0.795, an accuracy of 74.2%, a sensitivity of 68.1%, a specificity of 84.3%, a PPV of 85.3%, and an NPV of 65.8% in External-2. Figure 3 illustrates the histogram of the model prediction scores in the 3 datasets. The distributions of AD-dementia and cognitive normal classes were less overlapped in the internal validation than in the external testing, which also explained the performance drops in the external testing.
Table 2.
The Performance of the Proposed Fusion Network-Based Deep Learning Models to Classify Participants with AD Dementia vs Cognitively Normal Using Different Inputs from OCT Analysis Reports and Paired En Face Images in Internal Validation and 4 Unseen Datasets
| Datasets | AUROC (95% CI) | Accuracy, % (95% CI) | Sensitivity, % (95% CI) | Specificity, % (95% CI) | PPV, % (95% CI) | NPV, % (95% CI) |
|---|---|---|---|---|---|---|
| AD dementia vs cognitively normal | ||||||
| Internal validation | 0.943 (0.906–0.980) | 90.5 (86.0–94.4) | 93.3 (84.8–99.1) | 87.7 (76.7–95.9) | 91.5 (86.0–96.8) | 90.0 (80.5–98.4) |
| External-1 | 0.786 (0.673–0.899) | 80.3 (63.0–87.7) | 68.0 (44.0–96.0) | 87.5 (51.8–100) | 69.6 (44.9–100) | 85.3 (78.6–96.4) |
| External-2 | 0.795 (0.716–0.874) | 74.2 (66.7–81.7) | 68.1 (50.7–88.4) | 84.3 (60.8–96.1) | 85.3 (73.7–95.2) | 65.8 (57.1–81.3) |
| MCI vs cognitively normal | ||||||
| External-1 | 0.744 (0.633–0.854) | 76.4 (66.3–84.3) | 60.6 (39.4–87.9) | 87.5 (55.4–98.2) | 74.1 (52.9–94.1) | 78.8 (72.6–89.2) |
| External-2 | 0.787 (0.715–0.859) | 72.7 (61.6–82.0) | 68.6 (49.6–86.8) | 80.4 (58.8–96.1) | 89.5 (82.8–96.7) | 52.9 (42.9–68.6) |
| AD-MCI vs normal | ||||||
| External-1 | 0.787 (0.643–0.931) | 79.1 (69.4–88.7) | 71.6 (46.2–92.9) | 81.2 (69.6–91.8) | 52.6 (30.0–75.0) | 90.8 (81.4–97.8) |
| External-2 | 0.791 (0.694–0.888) | 75.1 (65.0–83.8) | 53.0 (36.4–69.7) | 91.3 (82.2–98.0) | 81.9 (63.6–95.7) | 72.5 (60.7–83.6) |
AD = Alzheimer's disease; AUROC = area under the receiver operating characteristic curve; CI = confidence interval; MCI = mild cognitive impairment; NPV = negative predictive value; PPV = positive predictive value.
AD-MCI was defined as subjects with MCI who were amyloid-β positive, and normal was defined as cognitively normal subjects who were amyloid-β negative.
Figure 3.
The histogram of the model prediction scores for AD-dementia among patients with AD-dementia (red) and cognitively normal participants (green) in the internal validation (left), External-1 (middle), and External-2 (right). The distributions were less overlapped in the internal validation than in the external testing, which indicated that the model could differentiate AD-dementia and cognitively normal better in the internal validation. AD = Alzheimer's disease.
For MCI (symptom-based) versus cognitively normal, the ensemble model achieved an AUROC of 0.744, an accuracy of 76.4%, a sensitivity of 60.6%, a specificity of 87.5%, a PPV of 74.1%, and an NPV of 78.8% in External-1; and an AUROC of 0.787, an accuracy of 72.7%, a sensitivity of 68.6%, a specificity of 80.4%, a PPV of 89.5%, and an NPV of 52.9% in External-2.
For AD-MCI (PET-based) versus normal, the ensemble model achieved an AUROC of 0.787, an accuracy of 79.1%, a sensitivity of 71.6%, a specificity of 81.2%, a PPV of 52.6%, and an NPV of 90.8% in External-1; and an AUROC of 0.791, an accuracy of 75.1%, a sensitivity of 53.0%, a specificity of 91.3%, a PPV of 81.9%, and an NPV of 72.5% in External-2.
In the testing using External-3, the proposed ensemble model can identify 97% of cognitively normal subjects.
Performance of Base Models Using Different Inputs
We compared base models using OCT inputs with or without paired en face images (Table 3). For the ONH models, though without significant differences, adding paired en face images improved the AUROC in all datasets (internal validation 0.898 vs 0.844, P = 0.068; External-1 0.767 vs 0.714, P = 0.38; External-2 0.736 vs 0.664, P = 0.25). For the macula models, adding paired en face images significantly improved the AUROCs in External-1 (0.787 vs 0.613, P = 0.007), while the performance is comparable in the internal validation (0.875 vs 0.842, P = 0.24) and External-2 (0.737 vs 0.757, P = 0.70).
Table 3.
The Comparison of Different Fusion Network-Based Deep Learning Models Using Inputs from OCT Reports with or without Paired En Face Images to Identify AD Dementia from Cognitively Normal
| Input | AUROC (95% CI) | P Value | Accuracy, % (95% CI) | Sensitivity, % (95% CI) | Specificity, % (95% CI) |
|---|---|---|---|---|---|
| ONH-Centered | |||||
| Internal validation | |||||
| OCT reports only | 0.844 (0.781–0.908) | Ref | 81.5 (74.8–86.8) | 77.4 (67.9–86.9) | 86.6 (74.6–94.0) |
| With paired en face images | 0.898 (0.846–0.949) | 0.068 | 85.4 (80.1–90.7) | 90.5 (72.6–97.6) | 82.1 (70.2–95.5) |
| External-1 | |||||
| OCT reports only | 0.714 (0.603–0.826) | Ref | 65.0 (53.8–76.3) | 100 (73.1–100) | 50.0 (33.2–74.2) |
| With paired en face images | 0.767 (0.658–0.876) | 0.38 | 76.3 (56.3–85.0) | 73.1 (50.0–100) | 79.6 (35.2–94.4) |
| External-2 | |||||
| OCT reports only | 0.664 (0.563–0.764) | Ref | 64.9 (56.8–73.9) | 50.8 (27.0–79.4) | 85.4 (54.2–100) |
| With paired en face images | 0.736 (0.640–0.832) | 0.25 | 73.0 (64.0–80.2) | 69.8 (47.6–85.7) | 78.1 (58.3–93.8) |
| Macula-centered | |||||
| Internal validation | |||||
| OCT reports only | 0.842 (0.775–0.909) | Ref | 81.1 (74.5–86.9) | 76.5 (64.7–92.9) | 86.8 (66.2–95.6) |
| With paired en face images | 0.875 (0.821–0.928) | 0.24 | 82.4 (76.5–87.6) | 78.8 (68.2–90.6) | 86.8 (72.1–95.6) |
| External-1 | |||||
| OCT reports only | 0.613 (0.486–0.740) | Ref | 61.7 (44.4–74.1) | 80.8 (42.3–100) | 52.7 (18.2–85.5) |
| With paired en face images | 0.787 (0.682–0.892) | 0.007 | 76.5 (60.5–85.2) | 76.9 (50.0–100) | 76.4 (43.6–96.4) |
| External-2 | |||||
| OCT reports only | 0.757 (0.670–0.845) | Ref | 71.6 (64.7–79.3) | 59.1 (43.9–86.4) | 90.0 (60.0–98.0) |
| With paired en face images | 0.737 (0.646–0.828) | 0.70 | 70.7 (62.9–78.5) | 72.7 (45.5–95.5) | 70.0 (40.0–92.0) |
AD = Alzheimer's disease; AUROC = area under the receiver operating characteristic curve; CI = confidence interval; ONH = optic nerve head.
Bold values are the ones significantly higher than the comparator.
Visualization of the Region of Interest
The heatmap examples of a patient with AD-dementia and a cognitively normal participant, shown in Figure S3 (available at www.ophthalmologyscience.org), indicated that the common regions of interest of the proposed ensemble model include ONH, RNFL thickness, GCIPL thickness, GCIPL deviation, and vessels.
Discussion
We developed an ensemble model using OCT, which integrates and combines 2 base DL models' classification results for the final classification task, which showed the highest accuracy for AD-dementia detection and also showed good performance in identifying participants with MCI only or with AD-MCI.
Current modalities and tools to detect AD have limitations for large-scale screening. For example, while PET (i.e., Aβ-PET and tau PET)-based biomarkers have the highest accuracy in AD diagnosis, the widespread use of PET for screening is limited by the high cost, low accessibility, invasiveness, technical complexity, and the risk of using radioactive tracers.12 Hence, research groups have actively explored other viable modalities for detecting AD in a simpler, noninvasive, and cost-effective approach. Health care from the eye prescreening is a proactive approach that applies validated AI-powered retinal imaging to identify a wide range of ocular and systemic conditions.37 With regard to screening for dementia, previous studies26,32 developed DL models with input from multiple retinal imaging modalities, including OCT, OCT angiography, and ultra-widefield fundus photography, indicating that retinal imaging can provide potential features for AD detection. Though promising, these models may be less feasible for screening due to the requirement of multiple relatively costly devices. Our previous study28 demonstrated that a DL model based solely on retinal photographs could achieve accuracies >80% for differentiating patients with AD-dementia from cognitively normal subjects, which can potentially serve as a population-based screening in communities or lower-resourced settings due to its low cost and high availability. However, our established model only analyzed the top-viewed part of the retina. In this study, we utilized intraretinal data generated from OCT, including both macula-centered and ONH-centered thickness/deviation maps reflecting neural changes and en face images reflecting vascular changes. Though we could not conduct a head-to-head comparison with our previous retinal photograph-based DL model due to the variances in datasets and training strategies, the proposed single OCT device-based model, interpreting comprehensive information and providing more fine-grained classification, showed great potential to drive OCT as an ideal alternative for AD detection. With further development in low-cost or portable OCT38, 39, 40 and increasing evidence on AI-driven Oculomics (i.e., identifying ocular biomarkers of systemic diseases),41,42 it would be promising to incorporate our proposed model with existing eye disease screening programs or utilize it as a risk stratification or triage tool in clinics with OCT available for AD opportunistic screening.
Identifying early AD, such as AD-MCI and mild AD-dementia, becomes more important as the recently approved disease-modifying therapies target patients at this stage. Hence, any potential screening strategies identifying patients with MCI or even AD-MCI might further extend the spectrum of patients who can benefit from any disease-modifying therapies.43 A previous study developed an OCT and anatomical information-based DL model to identify AD/MCI versus cognitively normal and achieved AUROCs of 0.910 and 0.840 in the internal and external validations.30 However, this study grouped AD and MCI together, which could not indicate the model's performance on identifying MCI alone. In this study, we used both symptom-based and PET biomarker–based definitions to separately test the proposed ensemble model’s performance for MCI versus cognitively normal and AD-MCI versus normal, respectively. We found that using only a symptom-based reference standard, the model's performance dropped slightly in identifying MCI than in identifying AD-dementia, while combined with the PET-biomarker as the gold standard to define AD-MCI and normal, the model's performance in identifying AD-MCI was almost the same as identifying AD-dementia. Thus, the proposed model showed great potential as a tool for detecting AD at an earlier stage, especially when synergistically combined with other blood-based or neuroimaging-based biomarkers, leading to more timely treatments or modifiable lifestyle interventions.
Complex data are a common obstacle hindering DL model development when training with imbalanced, high-dimensional, or noisy data. It is challenging for DL models to capture multiple characteristics and the underlying features from different inputs. Given the complexity of input from a single OCT device, we incorporated an ensemble learning technique to harmonize different inputs, as well as ensembled features to provide a unified classification using the majority voting scheme. Ensemble learning techniques have been used in ophthalmic AI for different tasks, such as detecting myopic maculopathy,44 predicting ocular hypertension after Descemet membrane endothelial keratoplasty,45 estimating visual function from fundus photographs of eyes with retinitis pigmentosa,46 detecting multiple retinal conditions from fundus photographs,47 predicting visual field from OCT,48 and predicting imminent exudative age-related macular degeneration conversion from volumetric OCT scans.49 Ensemble learning-based models showed potential to reduce model prediction errors when the base models were diverse and independent. In addition, despite numerous base models, an ensemble learning-based model can operate and perform as a single model, which is promising to be utilized in a screening scenario.50 In our study, we also found that the ensemble model could improve performance, especially when tested externally, indicating that the ensemble learning technique could leverage the benefits of the ONH model and macula model and provide more generalizable results. More importantly, the proposed ensemble model provided flexibility for further implementation in clinical settings, as it can provide the classification result with OCT data available from one or both imaging areas. Additionally, the proposed ensemble model can identify 97% of cognitively normal subjects using another OCT vendor, which demonstrates its cross-device adaptability.
Our study had several strengths. First, we trained the DL models for the detection of AD using OCT data alone but combined both neural and vascular features in the retina that were potentially related to AD-dementia. Our strategies enhanced the feasibility of using OCT as a simple and noninvasive screening tool. Second, we used the ensemble learning technique to prevent potential bias from base DL models and improve the overall generalizability. Third, we included a relatively large and comprehensive OCT dataset to train DL models for AD classification (i.e., 3039 paired OCT inputs including 3 analysis reports and 2 en face images) and externally tested our models on 2 independent datasets. Fourth, we used both symptom-based and PET biomarker–based definitions to further evaluate our model's potential to detect MCI and AD-MCI, respectively, which is essential for timely intervention and can potentially prevent progression into AD-dementia with the emergence of disease-modifying therapies. The model's ability to provide more fine-grained classification will also enable future longitudinal tracking in individual patients, for example, from AD-MCI to AD-dementia, or reversal of dementia after therapy.
Although promising results were demonstrated, there were also inherent limitations in our study. First, we could not conduct a head-to-head comparison with our previous retinal photograph-based DL model28 due to the variances in datasets and training strategies. Second, though the ensemble model had higher generalizability than the 2 base models, its performance still dropped in external testing compared with the internal validation. To address this issue, more advanced domain adaptation techniques are warranted to maintain stable performance when testing on new datasets. Third, due to the limited number of participants who received PET scans, our training data were labeled using clinical diagnoses based on symptoms, rather than based on PET biomarkers. Thus, our model could not realize the differentiation of AD-dementia versus other types of dementia. Recent advanced AI developments, such as foundation models for pretraining51,52 and generative AI for image synthesis,53 show great potential to improve the performance of DL and retinal imaging-based AD detection using a smaller sample size. It would be feasible to train an additional model based on a biomarker-labeled dataset with such an approach. Fourth, we used gradient-weighted class activation mapping to generate explainable heatmaps to visualize the model's learned features qualitatively. Though showing reasonable region of interest, such as vessels, ONH, and thicknesses, more quantitative measurements are warranted to further interpret more specific features that can differentiate AD-dementia/AD-MCI and cognitively normal.
Conclusion
Our proposed ensemble model, integrating multiple base models and inputs from OCT analysis, demonstrates strong potential for leveraging OCT imaging in detecting both AD-dementia and early-stage AD, enabling opportunistic screening for AD during ophthalmic visits.
Acknowledgments
The authors acknowledge BrightFocus Foundation (ref. A2018093S); Health and Medical Research Fund, Hong Kong SAR (ref. 04153506; ref. 12230156); and Gerald Choa Neuroscience Institute start-up research fund.
Manuscript no. XOPS-D-25-00731.
Footnotes
Supplemental material available atwww.ophthalmologyscience.org.
Disclosures:
All authors have completed and submitted the ICMJE disclosures form.
The authors made the following disclosures:
K.H.: Co-founder — i-Cognitio Sciences Ltd.
V.C.T.M.: Co-founder — i-Cognitio Sciences Ltd.
C.Y.C.: Co-founder — i-Cognitio Sciences Ltd.
C.-Y.C.: Consultant — Medi-Whale; Co-founder — Eye AI.
T.Y.A.L.: Financial support — Research to Prevent Blindness fund.
T.Y.W.: Consultant — AbbVie Pte Ltd, Aldropika Therapeutics, Bayer, Boehringer-Ingelheim, Carl Zeiss, Genentech, Novartis, Opthea Limited, Plano, Quaerite Biopharm, Research Ltd, Regeneron Pharmaceuticals Inc, Roche, Sanofi, Shanghai Henlius; Co-founder — EyRIS, VISRE companies.
Funding support was provided by BrightFocus Foundation (A2018093S) and the Health and Medical Research Fund, Hong Kong SAR (04153506; 12230156).
Support for Open Access publication was provided by the Hong Kong SAR.
HUMAN SUBJECTS: Human subjects were included in this study. This retrospective, multicenter, case-control study was approved by the human ethics boards of the Joint Chinese University of Hong Kong-New Territories East Cluster and Hong Kong Hospital Authority Kowloon Central Cluster Clinical Research Ethics Committee, Hong Kong, as well as local research ethics committees in each center. All studies were conducted following the Declaration of Helsinki. Written informed consent was exempted for the datasets that involved only retrospective analysis using fully anonymized OCT images.
No animal subjects were used in this study.
Author Contributions:
Conception and design: Ran, Li-Hsian Chen, Wong, Mok, Cheung
Analysis and interpretation: Ran, Hu, Dai, Sham, Zheng, Liu, Kwok, Hilal, Cheng, Chua, Schmetterer, Y.C. Tham
Data collection: Ran, Hu, Hui, Chan, Ng
Obtained funding: Cheung
Overall responsibility: Ran, Hu, Hui, Dai, Chan, Ho, Au, Ng, Sham, Zheng, Liu, He, Tham, Kwok, Cheng, Chua, Schmetterer, T.Y.A. Liu, Y.C. Tham, Li-Hsian Chen, Wong, Mok, Cheung
Supplementary Data
References
- 1.2024 Alzheimer's disease facts and figures. Alzheimers Dement. 2024;20:3708–3821. doi: 10.1002/alz.13809. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.The Lancet Neurology Treatment for Alzheimer's disease: time to get ready. Lancet Neurol. 2023;22:455. doi: 10.1016/S1474-4422(23)00167-9. [DOI] [PubMed] [Google Scholar]
- 3.van Dyck C.H., Swanson C.J., Aisen P., et al. Lecanemab in early Alzheimer's disease. N Engl J Med. 2023;388(1):9–21. doi: 10.1056/NEJMoa2212948. [DOI] [PubMed] [Google Scholar]
- 4.Sims J.R., Zimmer J.A., Evans C.D., et al. Donanemab in early symptomatic alzheimer disease: the TRAILBLAZER-ALZ 2 randomized clinical trial. JAMA. 2023;330:512–527. doi: 10.1001/jama.2023.13239. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Livingston G., Huntley J., Liu K.Y., et al. Dementia prevention, intervention, and care: 2024 report of the Lancet standing Commission. Lancet. 2024;404:572–628. doi: 10.1016/S0140-6736(24)01296-0. [DOI] [PubMed] [Google Scholar]
- 6.Govindpani K., McNamara L.G., Smith N.R., et al. Vascular dysfunction in Alzheimer's disease: a prelude to the pathological process or a consequence of it? J Clin Med. 2019;8:651. doi: 10.3390/jcm8050651. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Bjerkan J., Meglic B., Lancaster G., et al. Neurovascular phase coherence is altered in Alzheimer's disease. Brain Commun. 2025;7 doi: 10.1093/braincomms/fcaf007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Niikura T., Tajima H., Kita Y. Neuronal cell death in Alzheimer's disease and a neuroprotective factor, humanin. Curr Neuropharmacol. 2006;4:139–147. doi: 10.2174/157015906776359577. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.London A., Benhar I., Schwartz M. The retina as a window to the brain—from eye research to CNS disorders. Nat Rev Neurol. 2013;9(1):44–53. doi: 10.1038/nrneurol.2012.227. [DOI] [PubMed] [Google Scholar]
- 10.Alber J., Bouwman F., den Haan J., et al. Retina pathology as a target for biomarkers for Alzheimer's disease: current status, ophthalmopathological background, challenges, and future directions. Alzheimers Demen. 2024;20:728–740. doi: 10.1002/alz.13529. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Snyder P.J., Alber J., Alt C., et al. Retinal imaging in Alzheimer's and neurodegenerative diseases. Alzheimers Demen. 2021;17:103–111. doi: 10.1002/alz.12179. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Cheung C.Y., Mok V., Foster P.J., et al. Retinal imaging in Alzheimer's disease. J Neurol Neurosurg Psychiatry. 2021;92(9):983–994. doi: 10.1136/jnnp-2020-325347. [DOI] [PubMed] [Google Scholar]
- 13.Cheung C.Y., Ikram M.K., Chen C., Wong T.Y. Imaging retina to study dementia and stroke. Prog Retin Eye Res. 2017;57:89–107. doi: 10.1016/j.preteyeres.2017.01.001. [DOI] [PubMed] [Google Scholar]
- 14.Cheung C.Y., Chan V.T.T., Mok V.C., et al. Potential retinal biomarkers for dementia: what is new? Curr Opin Neurol. 2019;32:82–91. doi: 10.1097/WCO.0000000000000645. [DOI] [PubMed] [Google Scholar]
- 15.Snyder P.J., Alber J., Alt C., et al. Retinal imaging in Alzheimer's and neurodegenerative diseases. Alzheimers Dement. 2021;17:103–111. doi: 10.1002/alz.12179. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Ito Y., Sasaki M., Takahashi H., et al. Quantitative assessment of the retina using OCT and associations with cognitive function. Ophthalmology. 2020;127:107–118. doi: 10.1016/j.ophtha.2019.05.021. [DOI] [PubMed] [Google Scholar]
- 17.Byun M.S., Park S.W., Lee J.H., et al. Association of retinal changes with alzheimer disease neuroimaging biomarkers in cognitively normal individuals. JAMA Ophthalmol. 2021;139:548–556. doi: 10.1001/jamaophthalmol.2021.0320. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Thomson K.L., Yeo J.M., Waddell B., et al. A systematic review and meta-analysis of retinal nerve fiber layer change in dementia, using optical coherence tomography. Alzheimers Dement (Amst) 2015;1:136–143. doi: 10.1016/j.dadm.2015.03.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Chan V.T.T., Sun Z., Tang S., et al. Spectral-domain OCT measurements in Alzheimer's disease: a systematic review and meta-analysis. Ophthalmology. 2019;126:497–510. doi: 10.1016/j.ophtha.2018.08.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Cheung C.Y.L., Ong Y.T., Hilal S., et al. Retinal ganglion cell analysis using high-definition optical coherence tomography in patients with mild cognitive impairment and Alzheimer's disease. J Alzheimers Dis. 2015;45:45–56. doi: 10.3233/JAD-141659. [DOI] [PubMed] [Google Scholar]
- 21.Ueda E., Hirabayashi N., Ohara T., et al. Association of inner retinal thickness with prevalent dementia and brain atrophy in a general older population: the hisayama study. Ophthalmol Sci. 2022;2:100157. doi: 10.1016/j.xops.2022.100157. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Cheung C.Y., Ong Y.T., Ikram M.K., et al. Microvascular network alterations in the retina of patients with Alzheimer's disease. Alzheimers Dement. 2014;10:135–142. doi: 10.1016/j.jalz.2013.06.009. [DOI] [PubMed] [Google Scholar]
- 23.Leung K.H.C., Chan V.T., Lam B.Y.K., et al. Retinal vascular changes are associated with PET-based biomarkers of Alzheimer's disease: a pilot study. J Alzheimer's Dis Rep. 2024;8:1639–1648. doi: 10.1177/25424823241300416. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Williams M.A., McGowan A.J., Cardwell C.R., et al. Retinal microvascular network attenuation in Alzheimer's disease. Alzheimers Dement (Amst) 2015;1:229–235. doi: 10.1016/j.dadm.2015.04.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Barrett-Young A., Reuben A., Caspi A., et al. Measures of retinal health successfully capture risk for Alzheimer's disease and related dementias at midlife. J Alzheimers Dis. 2025;108:S324–S333. doi: 10.1177/13872877251321114. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Wisely C.E., Wang D., Henao R., et al. Convolutional neural network to identify symptomatic Alzheimer's disease using multimodal retinal imaging. Br J Ophthalmol. 2022;106:388. doi: 10.1136/bjophthalmol-2020-317659. [DOI] [PubMed] [Google Scholar]
- 27.Tian J., Smith G., Guo H., et al. Modular machine learning for Alzheimer's disease classification from retinal vasculature. Sci Rep. 2021;11:238. doi: 10.1038/s41598-020-80312-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Cheung C.Y., Ran A.R., Wang S., et al. A deep learning model for detection of Alzheimer's disease based on retinal photographs: a retrospective, multicentre case-control study. Lancet Digit Health. 2022;4:e806–e815. doi: 10.1016/S2589-7500(22)00169-8. [DOI] [PubMed] [Google Scholar]
- 29.Marshall C.R., Uchegbu I. Artificial intelligence for detection of Alzheimer's disease: demonstration of real-world value is required to bridge the translational gap. Lancet Digit Health. 2022;4:E768–E769. doi: 10.1016/S2589-7500(22)00190-X. [DOI] [PubMed] [Google Scholar]
- 30.Chua J., Li C., Antochi F., et al. Utilizing deep learning to predict Alzheimer's disease and mild cognitive impairment with optical coherence tomography. Alzheimers Dement (Amst) 2025;17 doi: 10.1002/dad2.70041. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Hao J., Kwapong W.R., Shen T., et al. Early detection of dementia through retinal imaging and trustworthy AI. NPJ Digit Med. 2024;7:294. doi: 10.1038/s41746-024-01292-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Wisely C.E., Richardson A., Henao R., et al. A convolutional neural network using multimodal retinal imaging for differentiation of mild cognitive impairment from normal cognition. Ophthalmol Sci. 2024;4 doi: 10.1016/j.xops.2023.100355. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Müller D., Soto-Rey I., Kramer F. An Analysis on Ensemble Learning Optimized Medical Image Classification with Deep Convolutional Neural Networks. arXiv. 2022 doi: 10.48550/arXiv.2201.11440. [DOI] [Google Scholar]
- 34.Yim J., Chopra R., Spitz T., et al. Predicting conversion to wet age-related macular degeneration using deep learning. Nat Med. 2020;26:892. doi: 10.1038/s41591-020-0867-7. [DOI] [PubMed] [Google Scholar]
- 35.Vasey B., Nagendran M., Campbell B., et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat Med. 2022;28:924. doi: 10.1038/s41591-022-01772-9. [DOI] [PubMed] [Google Scholar]
- 36.Kingma D.P., Ba J. Adam: a method for stochastic optimization. arXiv. 2015 doi: 10.48550/arXiv.1412.6980. [DOI] [Google Scholar]
- 37.Weinreb R.N., Lee A.Y., Baxter S.L., et al. Application of artificial intelligence to deliver healthcare from the eye. JAMA Ophthalmol. 2025;143:529–535. doi: 10.1001/jamaophthalmol.2025.0881. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Song G., Chu K.K., Kim S., et al. First clinical application of low-cost OCT. Transl Vis Sci Technol. 2019;8:61. doi: 10.1167/tvst.8.3.61. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Song G., Jelly E.T., Chu K.K., et al. A review of low-cost and portable optical coherence tomography. Prog Biomed Eng (Bristol) 2021;3 doi: 10.1088/2516-1091/abfeb7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Price H.B., Song G., Wang W., et al. Development of next generation low-cost OCT towards improved point-of-care retinal imaging. Biomed Opt Express. 2025;16:748–759. doi: 10.1364/BOE.551625. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Wagner S.K., Fu D.J., Faes L., et al. Insights into systemic disease through retinal imaging-based oculomics. Transl Vis Sci Techn. 2020;9:6. doi: 10.1167/tvst.9.2.6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Zhu Z., Wang Y., Qi Z., et al. Oculomics: current concepts and evidence. Prog Retin Eye Res. 2025;106 doi: 10.1016/j.preteyeres.2025.101350. [DOI] [PubMed] [Google Scholar]
- 43.Chan V.T.T., Ran A.R., Wagner S.K., et al. Value proposition of retinal imaging in Alzheimer's disease screening: a review of eight evolving trends. Prog Retin Eye Res. 2024;103 doi: 10.1016/j.preteyeres.2024.101290. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Qian B., Sheng B., Chen H., et al. A competition for the diagnosis of myopic maculopathy by artificial intelligence algorithms. JAMA Ophthalmol. 2024;142:1006–1015. doi: 10.1001/jamaophthalmol.2024.3707. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Kim M.S., Kim H., Lee H.K., et al. Artificial intelligence in predicting ocular hypertension after descemet membrane endothelial Keratoplasty. Invest Ophthalmol Vis Sci. 2025;66:61. doi: 10.1167/iovs.66.1.61. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Nagasato D., Sogawa T., Tanabe M., et al. Estimation of visual function using deep learning from ultra-widefield fundus images of eyes with Retinitis pigmentosa. Jama Ophthalmol. 2023;141:305–313. doi: 10.1001/jamaophthalmol.2022.6393. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Pandey P.U., Ballios B.G., Christakis P.G., et al. Ensemble of deep convolutional neural networks is more accurate and reliable than board-certified ophthalmologists at detecting multiple diseases in retinal fundus photographs. Br J Ophthalmol. 2024;108:417–423. doi: 10.1136/bjo-2022-322183. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Lazaridis G., Montesano G., Afgeh S.S., et al. Predicting visual fields from optical coherence tomography via an ensemble of deep representation learners. Am J Ophthalmol. 2022;238:52–65. doi: 10.1016/j.ajo.2021.12.020. [DOI] [PubMed] [Google Scholar]
- 49.Liu T.Y.A., Liu Y., Gastonguay M.S., et al. Predicting imminent conversion to exudative age-related macular degeneration using multimodal data and ensemble machine learning. Ophthalmol Sci. 2025;5 doi: 10.1016/j.xops.2025.100785. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Mahajan P., Uddin S., Hajati F., Moni M.A. Ensemble learning for disease prediction: a review. Healthcare (Basel) 2023;11:1808. doi: 10.3390/healthcare11121808. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Moor M., Banerjee O., Abad Z.S.H., et al. Foundation models for generalist medical artificial intelligence. Nature. 2023;616:259–265. doi: 10.1038/s41586-023-05881-4. [DOI] [PubMed] [Google Scholar]
- 52.Zhou Y., Chia M.A., Wagner S.K., et al. A foundation model for generalizable disease detection from retinal images. Nature. 2023;622:156–163. doi: 10.1038/s41586-023-06555-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Ktena I., Wiles O., Albuquerque I., et al. Generative models improve fairness of medical classifiers under distribution shifts. Nat Med. 2024;30:1166–1173. doi: 10.1038/s41591-024-02838-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.



