Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 Jul 28.
Published in final edited form as: Asia Pac J Ophthalmol (Phila). 2024 Aug 28;13(4):100095. doi: 10.1016/j.apjo.2024.100095

Development of oculomics artificial intelligence for cardiovascular risk factors: A case study in fundus oculomics for HbA1c assessment and clinically relevant considerations for clinicians

Joshua Ong a,1, Kuk Jin Jang b,1, Seung Ju Baek e, Dongyin Hu b, Vivian Lin b, Sooyong Jang b, Alexandra Thaler c, Nouran Sabbagh c, Almiqdad Saeed c,d, Minwook Kwon e, Jin Hyun Kim f, Seongjin Lee e, Yong Seop Han g, Mingmin Zhao b, Oleg Sokolsky b, Insup Lee b,*, Lama A Al-Aswad b,c,**
PMCID: PMC12303377  NIHMSID: NIHMS2094464  PMID: 39209216

Abstract

Artificial Intelligence (AI) is transforming healthcare, notably in ophthalmology, where its ability to interpret images and data can significantly enhance disease diagnosis and patient care. Recent developments in oculomics, the integration of ophthalmic features to develop biomarkers for systemic diseases, have demonstrated the potential for providing rapid, non-invasive methods of screening leading to enhance in early detection and improve healthcare quality, particularly in underserved areas. However, the widespread adoption of such AI-based technologies faces challenges primarily related to the trustworthiness of the system. We demonstrate the potential and considerations needed to develop trustworthy AI in oculomics through a pilot study for HbA1c assessment using an AI-based approach. We then discuss various challenges, considerations, and solutions that have been developed for powerful AI technologies in the past in healthcare and subsequently apply these considerations to the oculomics pilot study. Building upon the observations in the study we highlight the challenges and opportunities for advancing trustworthy AI in oculomics. Ultimately, oculomics presents as a powerful and emerging technology in ophthalmology and understanding how to optimize transparency prior to clinical adoption is of utmost importance.

Keywords: Artificial intelligence, Reliability, Trustworthy, Ophthalmology, Oculomics, Machine learning

1. Introduction

Artificial Intelligence (AI) has emerged as a transformative force across various domains of healthcare, from disease diagnosis to risk-stratification and clinical decision making, paving the way for more efficient, higher quality, and more personalized patient care.1 Ophthalmology continues to be an attractive field for AI applications given its image-based and data-rich nature. In ophthalmology, AI has the potential to revolutionize identification of multiple systemic diseases with ocular manifestations, such as cardiovascular disease (CVD) and diabetes. Improved efficiency and accuracy in image interpretation enables early detection of life-threatening diseases, and improves quality of care in rural settings without access to highly-specialized ophthalmologic care. At scale, AI has the potential to increase access to high quality, personalized care at low cost.2 For these reasons, it appeals to multiple stakeholders in the healthcare ecosystem, including patients, physicians, payers, and industry. To date, deep learning (DL) algorithms have been developed to diagnose retinal pathologies including but not limited to diabetic retinopathy (DR), age-related macular degeneration (ARMD)3, multiple sclerosis,4 and retinopathy of prematurity (ROP),5 as well as to determine surgical candidacy for complex procedures.68

Oculomics describes integrating macroscopic, microscopic, and molecular ophthalmic features to identify image-based biomarkers of systemic disease beyond retinal pathologies.7 The goal of oculomics is to develop rapid, non-invasive, cost-effective biomarkers for screening, diagnosis, and risk-stratification of systemic disease.9 The predominant image-based biomarker in ophthalmology is optical coherence tomography (OCT), which has revolutionized the understanding, diagnosis, and management of many ocular diseases.10 Its objective ability to predict, assess, and diagnose pathology makes it a powerful biomarker, and coupled with the application of artificial intelligence, an important tool for predicting disease and designing treatment plans. OCT and other imaging modalities, such as fundus photography, continue to play important roles in oculomics development.

CVD, neurodegenerative disease, and diabetes are among the most studied diseases in oculomics due to their significant morbidity and mortality. CVD is the leading cause of death both in the United States and worldwide. Diabetes and elevated HbA1c is both an important risk factor for CVD, as well as leads to devastating micro- and macrovascular sequelae including blindness, neuropathy, and kidney disease.11 Given the chronic nature of these systemic diseases, early diagnosis and intervention is of paramount importance to prevent blindness and minimize morbidity and mortality. Early identification has also been shown to be cost effective, as controlling these diseases early can prevent long-term sequelae, hospitalization, and death.1214 Oculomics presents a promising, non-invasive way to screen for such systemic diseases with the eye and possibly lead to personalized treatment planning.9

Ocular findings associated with CVD and diabetes can be identified with indirect fundoscopic examination, fundus photography, and OCT. There have been multiple landmark studies showing the utility of oculomics in diagnosing both ocular and as well as systemic disease including cardiovascular and neurodegenerative disease.15,16 In 2018, Abramoff et al.17 demonstrated the ability to autonomously detect DR in a primary care setting using fundus photography and OCT.17 This model, IDx-DR, became the first FDA cleared AI diagnostic tool in any field of medicine. In 2019, Yoo et al. developed a DL algorithm to identify candidates for corneal refractive surgery as well as accurately identify subclinical signs of forme fruste keratoconus.6 Other studies have demonstrated the ability of automated vessel segmentation to predict cardiovascular events using retinal arteriolar caliber18,19 and deep neural networks have been shown to have potential in extracting age, sex, HbA1c, relative fat mass, and testosterone.20

Despite the breakthroughs, there remain significant obstacles to widespread adoption of AI models in ophthalmology. Most of these relate to the trustworthiness of an AI system. The results of any AI model are only as robust as the quality of its training data in terms of image quality, diversity in ethnicity, age, and gender of patients, and reference decision-making of expert ophthalmologists in the training dataset.21 To improve trustworthiness in AI for ophthalmology, the FDA and the Collaborative Community on Ophthalmic Imaging (CCOI) have established regulatory frameworks and initiatives to standardize the development and deployment of AI-based medical devices in ophthalmology, in an effort to make them more reliable.22 The FDA has defined principles of Software as a Medical Device SaMD23 to acknowledge software’s critical role in healthcare, as well as potential privacy issues and ambiguity. Moreover, the FDA has established three programs, the Breakthrough Device,24 Early Feasibility Studies (EFS),25 and the Humanitarian Device Exemption Program26 to encourage and facilitate the launch of new AI technology. If granted, the Breakthrough Device classification (given to Idx-DR) can expedite the regulatory process by enabling adaptive clinical trial design and moving pre-market data collection obligations to post-market. These provide substantial advances towards the deployment of AI technologies in ophthalmology.

With the advent of oculomics, which is anticipated to continue to rise, this paper aims to examine the lessons in developing trustworthiness in prior AI techniques, as well as provide an oculomics case study to analyze the new challenges that will be encountered with this novel technology. In this manuscript, we first introduce the field of oculomics and its powerful applications. We then present a pilot study for developing an oculomics model for providing HbA1c levels from retina fundus photos. Through this pilot study, we examine challenges faced by medical AI applications in other fields of medicine and understand how the techniques for overcoming them could be applied to the oculomics case study. We then discuss challenges and specific techniques for developing trustworthy AI for oculomics. We highlight the importance of optimizing transparency and reliability prior to clinical deployment and consider ensuring trustworthy AI during deployment. We conclude with a summary and recommendations for future oculomics development.

2. Pilot study in oculomics for HbA1c estimation

Prior studies have demonstrated reliable measurements of systemic biomarkers, such as HbA1c.18,27 As promising as oculomics may be, many challenges exist in the development of reliable AI models. In order to demonstrate the feasibility and challenges of developing models for oculomics, we present preliminary results from a pilot study for developing an oculomics model for assessing the level of HbA1c from fundus images. The purpose of this study is to not achieve high accuracy of the model as this is still ongoing work to achieve high accuracy; the purpose of the study is to utilize this oculomics model and alter certain aspects to demonstrate the impact of various factors on its performance.

2.1. Diabetes and its importance as a cardiovascular risk factor

Diabetes is one of the fastest growing chronic diseases in the world.28 It is a condition in which the body is unable to regulate blood glucose effectively, which can lead to serious health problems over time. It affects multiple organs in the body and can lead to a variety of complications that can be life-threatening if not managed well. For example, diabetes is a major risk factor for cardiovascular disease, and people with diabetes have a much higher risk of heart disease and stroke than people without diabetes.29,30 Diabetes can also cause a range of long-term health problems, including nerve damage, kidney damage and vision loss.31,32

Early diagnosis of diabetes is key to preventing or delaying complications from the disease.32,33 Patients diagnosed early can effectively manage their blood sugar levels through lifestyle changes, proper diet, regular exercise, and necessary medication.33,34 Early diagnosis can prevent irreversible nerve damage and proper cardiovascular screening. This management can reduce the risk of long-term health problems caused by diabetes and improve the patient’s overall health status.35

Diabetes is commonly diagnosed by testing the HbA1c level. The HbA1c test measures the percentage of glycated hemoglobin (Hemoglobin A1c, HbA1c) in the blood to estimate the average blood sugar level over the past 2–3 months. Glycated hemoglobin is formed when hemoglobin in the blood combines with glucose, and the higher the glucose level in the blood, the more glycated hemoglobin is produced. HbA1c levels below 5.7 % are considered normal, 5.7 % to 6.4 % are pre-diabetes, and levels 6.5 % or higher are diagnosed as diabetes. It is a widely used test for diagnosing diabetes in most health institutions and diabetes specialist organizations, and is performed on many patients in order to diagnose diabetes at an early stage.36

2.2. Utility of an oculomics model for assessing HbA1c

HbA1C development has improved diagnosing and treating diabetes as it reflects serum glucose control over 3 months as most patients are not diligent in reporting their blood sugar to their PCPs. However, HbA1c is not without drawbacks as it is inaccurate in certain situations like recent blood transfusion, sickle cell anemia, pregnancy, etc.

Diabetes treatment and management is multi-disciplinary and multi-faceted requiring insights from many fields of medicine as well as effort from the patient. In a cross-sectional study conducted to assess the most valuable sense, 88 % ranked sight as the most important sense, highlighting the importance of proper care by ophthalmologists. While often accessible in large integrated health systems, HbA1c values may not be readily available in community ophthalmology clinics, often having ophthalmologists rely on patient-reported BG which is often limited to “good” or “bad.” In addition, this technology may help address certain barriers regarding attendance to PCP offices. In the UK for example, over 50 % of the population attended the UK national health service community optometric practice for eye checks in 2016, compared to 12.8 % attendance rate for CVD risk stratification by PCP from 2009–2013. Developing an oculomics tool for assessing HbA1c and other CVD risk factors will provide a means for clinicians to a more holistic assessment of compliance and early detection and prevention of complications. From the patient’s perspective, an oculomics model will be able to connect abstract numbers to detectable changes in a retinal image or changes in vision. This would encourage patients to monitor glucose and take more serious steps to controlling their diabetes and improve their overall compliance.

2.3. Dataset and experimental approaches

The data used in this study was provided by the ophthalmology department of Changwon Gyeongsang National University Hospital. Fundus images were collected from patients who were diagnosed with and without diabetes through HbA1c level testing within two weeks before and after fundus acquisition. A total of 6118 fundus images were collected, of which 1138 were diagnosed as normal (Class 1, HbA1c <5.7 %), 1110 were diagnosed as pre-diabetes (Class 2, HbA1c between 5.7 % and 6.4 %), and 3870 were diagnosed as diabetes (Class 3, HbA1c > 6.4 %). The fundus images were subjected to three preprocessing steps: First, the image was resized to 224 × 224. Second, CircleCrop was applied to extract only the ocular region within the fundus, and Third, normalization was applied.

The objective of this study was to replicate and assess the feasibility of measuring HbA1c levels from fundus images. First, we compared the performance of using a single, monolithic model versus utilizing an ensemble architecture. In addition, our goal was to explore possible factors affecting the performance of the model. Specifically the following experiments were executed and analyzed:

  1. (Reliability) Analysis of model size and ensemble architecture on performance

  2. (Bias) Analysis of the effect of age

  3. (Bias) Analysis of the effect of biological sex

2.3.1. Dataset characterization

When selecting the architecture of the model to perform the analysis, 150 samples were selected from each class as validation data, while the remaining data was used as training data to train different models (see Section 2.3.2). This approach was to reduce the effect of data imbalance on the performance of the model for this preliminary study. Section 2.5 discusses additional considerations needed to be addressed for further development.

In addition to the HbA1c levels of the study participants, the sex and age of the participants was collected and analyzed. As can be seen in Fig. 1, the overall male-to-female ratio in the dataset was 58.8:47.2 with a slight dominance of male participants. Table 1 shows the distribution by sex for each of the classes in the training set. A higher proportion of Class 3 (diabetes) samples can be observed vs Class 1 (normal). The test set was selected to have a uniform distribution by sex for a size of 255 samples for each class. Additionally, Fig. 2 shows that the mean and range of HbA1c values is similar for both groups. This indicates that differences in performance across sexes will most likely not be attributed to differences in HbA1c levels. We analyze this factor in Section 2.4.3.

Fig. 1.

Fig. 1.

Male to female ratio of patients in cohort for pilot study.

Table 1.

Train and test dataset split by sex.

Gender Train Dataset Test Dataset


Class1 Class2 Class3 Class1 Class2 Class3

Male+Female 888 1044 3436 255 255 255
Male 513   625 2026
Female 375   419 1410
Fig. 2.

Fig. 2.

HbA1c level distributions for male and female patients.

From Fig. 3, we can observe that a significant portion of the data is collected from patients in the range of age 50 or higher. In order to analyze the effect of age on model performance, we split the dataset into two groups of age 50 (senior) or above and below 50 (youth). Note that this division was selected arbitrarily to test the effect of age on performance. Table 2 shows the distribution by age group for each of the classes in the training set. Overall, the size of the youth group is lower and as seen in Fig. 4, the HbA1c level of the youth group tends to be higher. This can be attributed to differences in the cause of high levels of HbA1c. Younger individuals tend to be diagnosed with Type I diabetes, additionally, as the effects of high glucose are less pronounced, the management is less controlled. Comparatively, older patients develop Type II diabetes and tend to be more adherent to recommendation in management leading to a lower HbA1c. Similar to before, 150 samples for each class were selected to maintain a balanced distribution of age groups.

Fig. 3.

Fig. 3.

Age distribution of patients in pilot study cohort.

Table 2.

Train and test dataset split by age group.

Age Train Dataset Test Dataset


Class1 Class2 Class3 Class1 Class2 Class3

Seniors+Youth 988 1144 3536 150 150 150
Seniors 668   993 2786
Youth 320   151   750
Fig. 4.

Fig. 4.

Mean and standard deviation of HbA1c level by age group.

2.3.2. Model architectures and methods

For the AI model, we used four different models. The four models were trained under the same conditions to compare their performance, and each model has a different network size and structure. We used VGG16,37 VGG19, ResNet-50,38 which are models designed based on convolutional neural networks (CNN), and Vision Transformer (ViT),39 which is a transformer-based model that has recently shown excellent performance in various fields. By comparing the four models, we aimed to select the best-performing deep learning model for fundus-based diabetes prediction within a limited amount of data. For all models, SGD with a learning rate of 0.01 was used to train the model for 50 epochs.

2.4. Experimental results and considerations for oculomics development

2.4.1. Effect of model size and model architecture

In this section, we present the results of varying the architecture both in terms of size and as well as the overall structure.

2.4.1.1. Performance by varying size of models.

Table 3 depicts the results of utilizing various architectures. The accuracy of the models was 56.67 % for VGG16, 60.87 % for VGG19, 56.22 % for ResNet-50 %, and 44.02 % for Vision Transformer. However, when using the VGG19 model to classify the three stages of diabetes, the most important pre-diabetes (Class 2) and diabetes data (Class 3) had relatively low accuracy. Looking at the confusion matrix of the trained VGG19 in Fig. 5, we can see that the percentage of correct answers for Class 2 and Class 3 is not high compared to the percentage of correct answers for Class 1, which is 0.75. Overall, the VGG19 model demonstrated the best performance across all metrics (see Table 3).

Table 3.

Validation set performance metrics for 4 trained deep learning models. For each class i, Precisioni = TPi/(TPi +FPi), Recalli =TPi/(TPi +FNi), F1i = 2TPi/(2TPi +FPi +FNi). The overall metric is computed as the average of the metric across the classes. Accuracy = # of instances correctly labeled/total # of instances.

Precision Recall F1 Score Accuracy

Vision Transformer 0.4413 0.44 0.4388 0.44
VGG–16 0.5664 0.5666 0.5661 0.5667
VGG–19 0.6125 0.6088 0.6053 0.6089
ResNet–50 0.5746 0.5622 0.5612 0.5622
Fig. 5.

Fig. 5.

Confusion matrix for output of VGG 19 model.

2.4.1.2. Performance with ensemble architecture.

As an alternative architecture, we explored utilizing an ensemble technique which combines instead of using a single model. The architecture is depicted in Fig. 6. Given input fundus data, Binary classifier 1 checks whether the fundus belongs to Class 3 or not. If Binary classifier 1 predicts that the data does not belong to Class 3, then Binary classifier 2 can classify whether the fundus belongs to Class 1 or Class 2, thus classifying the three diabetes stages. The accuracy of Binary classifier 1 was 74.8 % and the accuracy of Binary classifier 2 was 85.33 %. The performance metrics of the ensemble model consisting of the two binary classifiers are shown in Table 4. Overall, the accuracy improved by about 2.7 % to 62.73 % and a slight increase in F1 Score to 61.6 %.

Fig. 6.

Fig. 6.

Three-stage diabetes triage process architecture based on Ensemble Model.

Table 4.

Performance metrics for two binary classifiers and an ensemble model.

Accuracy Precision Recall F1 Score

Binary Classifier 1 0.748
Binary Classifier 2 0.8533
Ensemble Model 0.6273 0.6372 0.6227 0.616

The ensemble model resulted in a performance improvement of about 2 % over the single model. Figs. 7 and 8 show that although there is no significant improvement in overall performance, the percentage of correct answers for Class 2 and Class 3 is significantly improved. We can also see that the rate at which data from Class 1 and Class 2 is misdiagnosed as diabetes (Class 3) is significantly lower. As a result, the ensemble model showed significant improvement over the single model, and we can potentially expect better performance by improving the ensemble model architecture and increasing the amount of training data.

Fig. 7.

Fig. 7.

Ensemble model ROC curve.

Fig. 8.

Fig. 8.

Ensemble model confusion matrix.

2.4.1.3. Considerations for model reliability.

These experimental results demonstrate the feasibility of HbA1c assessment using fundus images. An important factor that needs to be considered when developing a reliable model is the size of the model relative to the dataset availability. A trend in model development is the increasing size and complexity of the model architectures. Our results highlight that various considerations must be made when developing a reliable oculomics model.

First, a larger model does not necessarily lead to better results. In our results, the VGG19 model outperformed smaller models (VGG16) and significantly larger models such as the ViT and Resnet-50 models. This could be due to the dataset size as complex models tend to require a larger number of samples in order to learn the larger number of parameters appropriately. These differences in performance may change significantly as the dataset size increases. In such cases, ViTs may outperform CNN-based models. Even so, our results demonstrate the potential of utilizing ensembles of smaller models to enable better performance when data availability is an issue.

Second, although the ensemble approach leads to better performance, this does lead to additional complexity when determining the reliability and robustness of the performance. In safety-critical applications of AI, as in oculomics, decisions and diagnoses could have profound effects on patient health and safety. Analyzing the reliability in other aspects than just performance on a single testing set and considering the generalizability of the model is important. These issues will be described more in detail in Section 3 and Section 4.

One specific consideration of trustworthiness is to address issues of bias that can affect the performance of the model. The next sections highlight the results of our pilot study when considering factors such as age and sex on model performance.

2.4.2. Effect of age on performance

To analyze the effect of age on performance, we conducted experiments using the VGG19 model, which demonstrated the highest overall performance in the previous set of experiments.

2.4.2.1. Overall model performance when trained by age group.

Table 5 depicts the results of training on datasets of various age groups. In this experiment, the testing set consists of samples from all age groups. As can be seen from the results, the model that was trained on the samples from both seniors and youth demonstrated higher accuracy. In comparison, when the training dataset consisted of a single group (above 50 or less than 50), the overall accuracy decreased. This reflects a degraded generalization capability of the model. Moreover, the mean HbA1c of the ‘youth’ group was higher than the ‘senior’ group. This is reflected in a higher F1 score for class 3 (diabetes) for the model trained on the ‘youth’ group compared to the model trained on the ‘senior’ group. This is further demonstrated in Fig. 9 where the model trained on the ‘youth’ dataset demonstrates a clear tendency to predict Class 3 when compared to the ‘senior’ model.

Table 5.

Performance by age group. Testing set contains samples from both groups.

F1 Acc

Age Class1 Class2 Class3

Senior + Youth 0.53 0.43 0.63 0.54
Seniors 0.45 0.31 0.55 0.45
Youth 0.47 0.34 0.57 0.49
Fig. 9.

Fig. 9.

Confusion matrix of predictions when model is trained on senior and youth groups, respectively.

2.4.2.2. Evaluation across different age groups.

This difference in performance degradation occurs in both directions as demonstrated by the results in Table 6. When a model is trained on a certain age group (‘senior’ or ‘youth’), the accuracy when tested on the same group is higher than when tested on the other group. This is more apparent with the ‘youth’ group which has a 19 % performance gap. Again, this is likely due to the biased distribution of HbA1c in the younger group.

Table 6.

Results from evaluation across various age groups.

Train Test F1 Acc


Class 1 2 3

Seniors + Youth Seniors 0.37 0.38 0.63 0.48
Youth 0.64 0.49 0.64 0.60
Seniors Seniors 0.46 0.40 0.61 0.51
Youth 0.44 0.19 0.48 0.40
Youth Seniors 0.24 0.25 0.51 0.39
Youth 0.62 0.42 0.65 0.58
2.4.2.3. Considerations for model bias.

The results of this experiment highlight the importance of diverse representation in the data used to train and evaluate a deep learning-based model. The ability of a model to demonstrate reliable and robust performance across a highly varied population is an important consideration when developing trustworthy AI solutions. As shown in this experiment, a simple bias in age and HbA1c ranges can lead to degraded performance. However, differences in HbA1c between the age groups do not fully explain the gap in performance and require subsequent analysis and development of methods to provide transparency into the decision-making process of AI models for oculomics. These considerations are further touched upon in the next experiment.

2.4.3. Effect of gender on performance

In this section, we evaluate the differences in model performance stemming from differences in the distribution of sexes in the training set. Similar to before, the VGG19 model was used to train the model.

2.4.3.1. Overall performance when biased by sex.

Table 7 depicts the evaluation results when a model is trained on training sets of various divisions of gender. Similar to the results when the training set is divided by age, the model trained on both sexes demonstrates the highest performance and a 5 % degradation is observed when the model is trained on only one of the sexes. This result is further highlighted by Table 8 which shows the performance differences when evaluated over various train-test combinations. The difference in accuracy when tested on male or female samples is the smallest when the model is trained on both sexes (2 %). In comparison, when trained on only males or only females, the performance difference is significantly higher with a gap of 17 % and 10 %, respectively. As noted in Sec. 3.3.1, various factors could contribute to this difference, however, the range of HbA1c does not seem to be a factor.

Table 7.

Performance by Gender. Testing set contain samples from both groups.

F1 Acc

Sex Class1 Class2 Class3

Male + Female 0.49 0.43 0.58 0.51
Male 0.38 0.39 0.53 0.46
Female 0.46 0.29 0.55 0.46
Table 8.

Performance by Gender with various train-test combinations.

Train Test F1 Acc


Class 1 2 3

Male + Female Male 0.51 0.45 0.58 0.52
Female 0.47 0.41 0.58 0.50
Male Male 0.49 0.52 0.59 0.54
Female 0.27 0.27 0.48 0.37
Female Male 0.37 0.26 0.49 0.41
Female 0.53 0.32 0.61 0.51
2.4.3.2. Prediction of gender from fundus images.

In order to better understand this surprising result, an additional model was trained to output the sex based on the fundus image. The results are in Table 9 and Fig. 10. The model exhibited surprising performance in classifying male and female fundus images with an overall accuracy of 87 %. However, this result does not necessarily signify a difference in the fundus image of a gender, but could potentially be a sign of unseen bias in the training dataset. For example, the model may learn where the optic nerve location is in the photo, or if the dataset and demographics are skewed in one direction, this may lead to these results as well.

Table 9.

Results of sex prediction based on fundus images.

Precision Recall F1-score Acc



Male Female Male Female Male Female

0.88 0.85 0.85 0.88 0.86 0.87 0.87
Fig. 10.

Fig. 10.

Confusion matrix of model predictions of sex based on fundus images.

2.4.3.3. Utilizing explainable AI techniques.

Digging deeper into the decision-making process of trained models, we employ a well-known interpretable AI technique called Grad-CAM to understand the important features in the fundus image to determine the various classification labels. Other methods employed for interpretable and explainable AI are described in Section 3.2.

Fig. 11 illustrates the differences in important image features when classifying male vs. female from the fundus image. Typically, when the model outputs a label of ‘female’, the focus is around the optic disc and cup area of the fundus image. Moreover, the output is more ‘focused’ and less noisy. In comparison, when outputting ‘male,’ the model tends to focus on areas near and around the macula. These characteristics (focused, optic-disc centered vs. dispersed, macula-centered) seem to be observable in the models trained for HbA1c assessment, as shown in Fig. 12.

Fig. 11.

Fig. 11.

Sample of results comparing Grad-CAM output for classifying male vs female.

Fig. 12.

Fig. 12.

Comparison of Grad-CAM output on different classes.

Fig. 12 compares the Grad-CAM output across the different classes and sexes when a model is trained over different subsets of the data. The first row shows that when the model is trained on male-only samples. When this model is tested on male samples of class 1 to 3, which signifies a higher level of HbA1c, the model tends to go from being tightly focused on the optic disc area to being more dispersed and focused in areas near to the macula and beyond. Conversely, for the female-only trained model (last row), the model shows the opposite trend of being more dispersed to being more focused on the optic disc area when going from class 1 to 3 of female samples. These trends hold when the models are tested on the excluded group from which they were trained on. This could potentially be a reason for the degradation in performance.

By comparison, the model trained on samples from both sexes (rows 2 and 3) exhibits a similar trend of focus when moving from class 1 to 3 to the male-only model and female-only model when tested on the male and female samples, respectively. This could mean that the model is learning to process samples from males and females differently, leading to a reduced gap in performance.

2.4.3.4. Considerations for model transparency.

The results utilizing Grad-CAM complement the analysis based on performance alone that the importance of a diverse dataset is needed for reliable and robust performance. Moreover, methods such as Grad-CAM allow for better transparency when understanding the decision-making process of deep learning-based models.

However, extreme caution must be taken when considering the differences exhibited between samples from males and females. In the current dataset, it seems plausible that the fundus image can be used to differentiate male and female samples and separate processes may be needed for assessing HbA1c, the underlying cause may be due to some other systemic bias in the data collection unrelated to the sex of the patient. Additional analysis and further validation over a larger dataset would be needed before drawing conclusions.

2.5. Summary and additional considerations for trustworthy AI in oculomics

The results of this study demonstrate the possibility of using deep learning in ophthalmology to diagnose diabetes without testing HbA1c levels. The results show that despite data limitations such as small amounts of data and class imbalance, the method achieves a certain level of performance. It is important to note that the model listed above is a preliminary model and employed for the use of understanding how changes in a dataset can impact performance in oculomics. HbA1c assessment and monitoring of diabetes through fundus imaging is still in the research phase and requires further validation.

Developing trustworthy oculomics for HbA1C assessment and monitoring presents several challenges, particularly in ensuring reliability, transparency, and combating bias. This study touched upon several issues that must be addressed before oculomics technology can be clinically deployed. In terms of reliability, the experiments demonstrated the data-related challenges which include the need for high-quality, diverse datasets to train models effectively while maintaining robust model performance across different populations and conditions. Transparency involves making model outputs interpretable so that healthcare providers can understand and trust the predictions, and quantifying uncertainty to inform clinical decision-making. Methods used in the study, such as Grad-CAM, aid in improving the transparency of decision-making. Additionally, it is important to address bias and ensure fairness in model predictions in order to avoid perpetuating healthcare disparities and ensuring equitable treatment for all patients. There are also limitations to the model as an example for oculomics. A limitation is that the dataset has a lack of ethnic diversity, thus, the generalizability for this example is weakened. It also did not analyze an ethnically diverse dataset for its validation dataset which could also be a point to analyze regarding oculomics transparency. This limitation can be improved upon by expanding the training and validation dataset to include different ethnic groups in the future. This can be achieved through technical innovations in generative AI, as will be discussed in Section 3.1, and through multi-institutional collaborations to enable diverse representation in datasets.

In the following section, we build upon these observations to further define the challenges facing trustworthy AI development for oculomics.

3. Challenges for trustworthy AI in oculomics development

While oculomics represents a promising opportunity to harness innovation to transform eye care, its potential is moderated by a series of significant challenges from building and training models to integrating validated technology into existing clinical workflows. There are many aspects regarding the challenges and requirements for developing trustworthy AI. Covering all these aspects is outside the scope of this paper. This section reveals the array of key complexities in developing oculomics technologies including reliability, transparency, and bias.

3.1. Artificial intelligence reliability in oculomics

Before AI-based models can be deployed for clinical decision-making and patient care, they need to be rigorously validated from a technical and clinical perspective. When ophthalmologists evaluate the reliability of AI-based models, they will likely require evidence from outputs on high-resolution images, consistency between imaging devices used in training and test data, as well as with their clinic’s own diagnostic imaging devices, and standardization of data collection, to ensure generalizability and reproducibility of the results beyond the original study. Reliability and openness of training data are critical to gaining widespread clinical acceptance.40 The AAO and FDA are aware of this need for standardization, and recent recommendations to promote standardization and transparency in AI suggest including minimally allowable information (e.g. basic demographic and clinical features) consistent with non-AI predictive models, as well as including a standard checklist or model score card for models prior to clinical implementation.41

While there is undoubtedly a balance between full transparency and management of proprietary model design, clinicians must be able to evaluate new models based on their medical education and post-graduate ophthalmology training. To this end, the model scorecard that is provided to the clinician should include its primary intended use, intended users, and scope of use to guide clinical decision-making by ophthalmologists who may be unfamiliar with using AI-based tools. A point to consider is also the level of clinical training and expertise of the individual utilizing the technology. It will be important for less-experienced physicians not to become too dependent on AI. However, the technology itself may also help less-experienced physicians become aware of a mistake. Ultimately, the responsibility for clinical decision-making is that of the physician; this applies when using AI.42 Standardization in model reporting will increase reliability and ensure that physicians are fully aware of the benefits and shortcomings of the models presented.

As demonstrated in the pilot study, overcoming small/limited and/or imbalanced training data presents a significant challenge for obtaining reliable models. Among many notable AI innovations, transfer learning and federated learning have been introduced to overcome such limitations in AI for opthalmology and medical imaging.4346 In these approaches, models pretrained with larger or more diverse data for a different task may be finetuned to similar tasks in opthalmology. Similar techniques could be applied to oculomics.

During the development process, advances in generative AI techniques may provide one of many solutions for improving reliability. Within the last decade, a breadth of literature has shown how generative AI techniques can overcome the challenges of using AI in ophthalmology. As demonstrated in the pilot study, a key challenge in developing reliable models is the availability of diverse medical datasets. Generative AI presents an opportunity to create alternative synthetic datasets for training more trustworthy AI models. A variety of approaches have used GANs4752 and diffusion models5355 to generate synthetic fundus data. Costa et al.,47 Alimanov and Islam,54 Go et al.56 used a two-stage approach, first generating a blood vessel map, then generating the fundus image conditioned on the map. Generative models have also been used to create synthetic data of other modalities. Wu et al.57 and Agharezaei et al.58 used diffusion models to generate OCT images and corneal topographic maps, respectively.

The ability to generate synthetic data, enhance data quality, and analyze the learned latent space has contributed to advancements in reliable AI for ophthalmology. While distinct from their ophthalmic counterparts, AI methods for oculomics will naturally face similar challenges and thus may be similarly aided by generative AI.

3.2. Artificial intelligence transparency in oculomics

The potential of oculomics to revolutionize disease diagnosis, treatment, and patient outcomes is exciting, however this promise may be curtailed by a lack of transparency in model decision-making and interpretation, as the rationale behind its conclusions is often opaque. This opacity constitutes a significant challenge to its clinical adoption, risking the trust of both healthcare professionals and patients, and raising liability concerns.7,59

Methods for improving transparency in AI for ophthalmology are likely to be relevant to oculomics. AI techniques such as deep learning models, which were essential in AI-based achievements in diagnosing retinal conditions such as DR, glaucoma, and macular degeneration, present steep challenges to transparency due to their complexity. Utilizing methods like saliency maps have been a step towards understanding the “black box” issue of AI.60 Methods based on saliency maps explain how AI makes decisions by highlighting important regions in model inputs such as volumetric OCT scans, making diagnoses more understandable and accurate.61 For example, Poplin et al.18 employed the soft attention mechanism62,63 to create a saliency map, highlighting areas in the input retinal fundus images crucial for predicting cardiovascular risk factors. In their work, intermediate activations from a smaller model were upscaled to the dimensions of the original input, which served as the saliency map. In their study, three ophthalmologists reviewed a subset of the generated saliency maps, confirming a consistent emphasis on specific anatomical features in the retinal images for predicting various risk factors.

Similar explanations based on saliency maps are also introduced to characterize DR lesions.64 Niu et al.65 perform a back-propagation-like operation66 by feeding the bottleneck activation from the trained detection network to the reversed version of this detection network. Such operation decodes both the appearance and spatial information of DR lesions in the bottleneck activation to saliency maps, potentially leading to improved diagnostic accuracy and increased uptake in clinical practice.67 All of these examples represent a significant advancement in this field.

An increasing number of studies are incorporating explanations as an integral part of the model.68 Both policymakers and researchers recognize the importance of understanding the decisions made by deep learning models.23,69 However, existing explanation methods used by prior researchers still have limitations. First, models may mistakenly attribute high importance to features that are either irrelevant or lack distinguishing power for the task being addressed.70 Furthermore, explanation methods do not necessarily represent the original model, as they often require the explanation model to be as complex as the original model.70 Deeper collaboration between scientists and clinicians is needed to examine and fully validate the output of explanation methods.

3.3. Bias in oculomics

Models built by artificial intelligence are only as good as their input data, and potential bias potentially embedded within AI training data poses a significant ethical concern, as this greatly affects clinical outcomes.42 This has led to calls for regulation and standardization of AI in clinical practice similar to “Model cards” introduced by Google’s ethical AI team.41 In ophthalmology, models should consider reference standards and database distribution, in addition to incorporating standardized methods for development, testing, and deployment, to ensure the safety, effectiveness, and equity of these systems.59 Similar guidelines should be developed for oculomics.

The evolving field of oculomics requires systems that are both transparent in decision-making while serving diverse patient populations. The validation of these AI tools must be supported by rigorous, peer-reviewed research that ensures their integration into clinical practice without exacerbating biases.59 Access to large datasets during development and evaluation, including images from patients with varied ethnicities, geographic origins, and co-morbidities, is crucial. This multi-faceted approach will improve precision in the model’s outcomes while addressing potential bias.71 Without diverse training data, AI-based models may underperform in their ultimate clinical use, (e.g. in historically underrepresented trial populations) limiting their utility in clinical care and exacerbating ‘health poverty.’ Health poverty describes the situation in which individuals or populations are unable to benefit from or even be harmed by AI due to a scarcity of representative data.72 This risk may be less than the potential implicit bias of a physician, but it is nevertheless important to consider as models are developed and deployed into clinical care.

By construction, AI models adopt the same biases embedded within the training data. For example, such tools will perform sub-optimally on population groups that are not included or underrepresented in the dataset. One recent study found that AI could infer self-reported race from fundus photographs converted to retinal vessel maps that were not thought to contain this information.73 This suggests that AI algorithms may have the potential for bias even if based on biomarkers rather than raw images.

To combat this challenge, generative AI can be used to supplement existing datasets with synthetic data for such populations. Joshi and Burlina74 used a GAN to generate synthetic retinal images from African American populations. On an age-related macular degeneration (AMD) detection task, the authors found that a model trained with the synthetic data predicted AMD more consistently across different racial groups than a model trained without. Generative AI has also been used to lower the barriers of high costs in medicine. Browne et al.75 used a diffusion model to enhance smartphone images of ophthalmic pathology slides, allowing for a lower-cost alternative to the digital slide scans that are commonly required for telepathology.

Additionally, generative AI also presents unique opportunities for oculomics and combating bias. The latent space of generative models has been shown to reveal correlations between patient characteristics (e.g., age, ethnicity, diagnosis) and ophthalmic data.76,77 Similar analyses using generative AI can be leveraged to aid oculomics research, understand factors leading to bias, and develop methods for fairness.

4. Trustworthy deployment of AI for oculomics

The deployment and integration of AI models into medical practice is an area of significant interest and development, aiming to enhance diagnostics, treatment planning, and patient care. Unlike the challenges during development, the deployment of AI technologies in such a sensitive domain mandates thorough consideration of reliability, adaptability to unforeseen data distributions, and the imperative for transparency. Here, we summarize some considerations and methodologies to address challenges surrounding reliability, transparency, and clinician-AI interactions.

4.1. Considerations for reliable deployment of AI

4.1.1. Regulatory guidelines

Ensuring the reliability of AI applications or models during deployment in medical practice requires strict validation and testing protocols.23,78 This embraces extensive cross-validation within diverse and multi-institutional datasets to ascertain the robustness and generalizability of AI models.78 The integration of AI-based applications with clinical decision support systems provides healthcare professionals with AI-driven insights while preserving the autonomy of clinical judgment. Regulatory guidelines, such as those promoted by the FDA for AI/ML-based Software as a Medical Device (SaMD), emphasize the necessity for AI models to follow rigorous safety and efficacy standards prior to clinical integration.23 In addition, clear legal regulations on the accountability and responsibility of AI usage in clinical settings must be established before the widespread usage of AI-based clinical decision-support tools. It is also important to take into account the importance of privacy, particularly with building and deploying trustworthy AI/oculomics. When adopting such technology into the clinical environment, there is a risk of patient data leakage. Securing patient data is of utmost importance and thus should be highly considered in the regulatory guidelines for oculomics development and deployment.

4.1.2. Diverse clinical environments

Ultimately, despite rigorous validation during development, AI models may face performance degradation due to unforeseen clinical scenarios or out-of-distribution (OOD) inputs during deployment.79,80 A systematic approach to mitigating this issue involves the incorporation of continuous learning frameworks, allowing AI systems to evolve and adapt to new data patterns. The employment of anomaly detection algorithms serves as a safeguard,81,82 capturing instances of noteworthy deviation from the model’s training distribution for manual review. Periodic model updates, incorporating novel and diverse data, and continuous evaluation, will be important to sustain the relevance and accuracy of AI applications in dynamic clinical environments.

4.2. Transparency in deployment of AI

As mentioned previously, the transparency of AI-driven decision-making processes presents a considerable barrier to the clinical adoption of AI technologies. Promoting transparency enforces the development of interpretable and explainable AI models, where the logic behind predictions and recommendations can be easily perceived by medical doctors. Model-agnostic explanation frameworks, such as Local Interpretable Model-agnostic Explanations (LIME)83 and SHapley Additive exPlanations (SHAP),84 offer valuable insights into model behavior not only during development but also during deployment. In addition to explainable features, respecting and complying with documents and standards, as outlined in the CONSORT-AI and SPIRIT-AI guidelines,85 could ensure comprehensive transparency in AI research findings and deployment methodologies.

The successful integration of AI into healthcare depends on awareness of these strategies, guided by up-to-date scholarly research and regulatory management. As the research field matures, developers, clinicians, and policymakers must promote AI systems that are not only technologically advanced but also ethically sound, transparent, and aligned with the highest value of providing the best care for the patient.

4.3. Considerations of clinician-AI interaction

4.3.1. Clinician over-reliance

When AI is utilized in clinical settings, it is anticipated that clinicians will initially approach it with caution and skepticism, comparing its suggestions to their own expertise.86 It is anticipated, however, that as the technology becomes further validated, clinicians may gradually gain confidence in the AI’s highly accurate performance, leading to its widespread adoption in clinical practice.87 While AI’s potential for consistency and efficiency is attractive, there is a risk when healthcare professionals become excessively dependent on it.87 Such reliance could be perilous in complex clinical decision-making that necessitates human discernment.88 For instance, should an AI system predict a high likelihood of a particular critical diagnosis and prove incorrect, the patient could endure considerable physical and emotional distress.89 Additionally, there could be legal ramifications if a physician, placing undue trust in the AI, pursues unnecessary treatments.90

Therefore, clinicians must critically evaluate AI recommendations and derive conclusions based on their medical knowledge and experience. Moreover, there is a risk that clinicians may defer to the AI’s judgment even when it contradicts their own and the rationale behind the AI’s decision is unclear to them. Consequently, societal dialogue and consensus are necessary to mitigate overreliance and ensure equitable utilization of AI in healthcare. It may also be necessary to design mandatory courses for education and training in using AI-based tools. Although beyond the scope of this review, this topic plays a pivotal role in responsible deployment.

4.3.2. Clinician-AI workflow design

AI-powered clinical care systems should be designed to enhance healthcare service efficiency without disrupting doctor-patient interactions. These systems require rapid data processing abilities and should be user-friendly, necessitating minimal training due to an intuitive interface. The role of AI should be to assist in minimizing superfluous tasks and expedite decision-making, thereby allowing doctors to allocate more time for patient care.91

4.4. Summary and future directions

This section discussed the integral considerations for the trustworthy deployment of AI in oculomics, emphasizing compliance with regulatory guidelines and adaptability to diverse clinical environments. Key aspects include ensuring transparency in AI deployment and addressing the dynamics of clinician-AI interaction, such as the risks associated with clinicians’ over-reliance on AI and the importance of well-designed workflows.

AI systems for oculomics must be developed as supporting tools that enable clinicians to prioritize patient care.91 The aim of these technological advancements should be to enhance human-centered healthcare by simplifying clinicians’ workflows, elevating care quality, and improving patient health outcomes. Achieving these objectives demands ongoing collaboration and integration among various stakeholders throughout the development and deployment lifecycle. This collaborative effort ensures that AI systems are effective, reliable, and aligned with the real-world needs and challenges clinicians confront in diverse clinical settings. This approach will help realize AI’s full potential in transforming healthcare delivery.

5. Conclusion

The advent of oculomics has revolutionized how we approach systemic and ocular diseases. With each revolutionary artificial intelligence technology that impacts clinical care also comes a myriad of considerations and a need for regulatory and transparent protocols. In harnessing the potential of oculomics for systemic disease detection, it is imperative to maintain accountability at every stage of research and application. A regulated system of accountability ensures that the data collected and the algorithms developed are accurate, reliable, and minimize bias. Rigorous validation of oculomics-based diagnostic tools through robust clinical trials and peer-reviewed studies will be vital to establish their efficacy and safety prior to widespread clinical deployment.

Our oculomics case study highlights the exciting potential of this technology. Although still an ongoing area of research, the oculomics model was able to identify HbA1c levels that have been historically measured through invasive blood draws. As a critical biomarker, this HbA1c oculomics research represents major breakthroughs that can help increase accessibility and address longstanding barriers to optimal healthcare. Our oculomics case study also highlighted how bias can lead to suboptimal results. We saw how deliberate changes to the dataset can lead to large changes in the accuracy and outcomes for certain patient demographics. Thus, transparent reporting of methodologies, including data collection protocols and analysis techniques, will be critical to foster confidence in the scientific and medical community for clinical use of these technologies. It will be important to continuously re-evaluate these methodologies and datasets to ensure that this technology continues to serve all communities, and to avoid any exacerbation of existing inequalities in healthcare delivery. Ultimately, the era of oculomics is likely still in its infancy, and the technology holds promising potential to revolutionize clinical care. We urge all researchers and clinicians in this space to continuously remain cognizant and diligent in ensuring that these technologies remain equally trustworthy and clinically impactful in their development and deployment.

Funding

Research to Prevent Blindness.

Footnotes

Disclosures

The authors have no financial disclosures to disclose.

References

  • 1.Al Kuwaiti A, Nazer K, Al-Reedy A, et al. A review of the role of artificial intelligence in healthcare. J Pers Med. 2023;13(6):951. 10.3390/jpm13060951. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Khanna NN, Maindarkar MA, Viswanathan V, et al. Economics of artificial intelligence in healthcare: diagnosis vs. treatment. Healthcare. 2022;10(12):2493. 10.3390/healthcare10122493. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Yan Q, Weeks DE, Xin H, et al. Deep-learning-based prediction of late age-related macular degeneration progression. Nat Mach Intell. 2020;2(2):141–150. 10.1038/s42256-020-0154-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Naji Y, Mahdaoui M, Klevor R, Kissani N. Artificial intelligence and multiple sclerosis: up-to-date review. Cureus; 15(9): e45412. DOI: 10.7759/cureus.45412. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Gensure RH, Chiang MF, Campbell JP. Artificial intelligence for retinopathy of prematurity. Curr Opin Ophthalmol. 2020;31(5):312–317. doi: 10.1097/ICU.0000000000000680. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Yoo TK, Ryu IH, Lee G, et al. Adopting machine learning to automatically identify candidate patients for corneal refractive surgery. npj Digit Med. 2019;2(1):1–9. 10.1038/s41746-019-0135-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Wagner SK, Fu DJ, Faes L, et al. Insights into systemic disease through retinal imaging-based oculomics. Transl Vis Sci Technol; 9(2): p. 6. DOI: 10.1167/tvst.9.2.6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Montolío A, Martín-Gallego A, Cegoñino J, et al. Machine learning in diagnosis and disability prediction of multiple sclerosis using optical coherence tomography. Comput Biol Med. 2021;133, 104416. 10.1016/j.compbiomed.2021.104416. [DOI] [PubMed] [Google Scholar]
  • 9.Wu JH, Liu TYA. Application of deep learning to retinal-image-based oculomics for evaluation of systemic health: a review. J Clin Med. 2022;12(1):152. 10.3390/jcm12010152. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Aumann S, Donner S, Fischer J, Müller F. Optical coherence tomography (OCT): principle and technical realization. In: Bille JF, ed. High resolution imaging in microscopy and ophthalmology: new frontiers in biomedical optics. Springer; 2019. Accessed August 7, 2024 〈http://www.ncbi.nlm.nih.gov/books/NBK554044/〉. [PubMed] [Google Scholar]
  • 11.Arnold F, Kappes J, Rottmann FA, Westermann L, Welte T. HbA1c-dependent projection of long-term renal outcomes. J Intern Med. 2024;295(2):206–215. 10.1111/joim.13736. [DOI] [PubMed] [Google Scholar]
  • 12.Avci BS, Saler T, Avci A, et al. Relationship between morbidity and mortality and HbA1c levels in diabetic patients undergoing major surgery. J Coll Physicians Surg Pak. 2019;29(11):1043–1047. 10.29271/jcpsp.2019.11.1043. [DOI] [PubMed] [Google Scholar]
  • 13.Anyanwagu U, Mamza J, Donnelly R, Idris I. Relationship between HbA1c and all-cause mortality in older patients with insulin-treated type 2 diabetes: results of a large UK Cohort Study. Age Ageing. 2019;48(2):235–240. 10.1093/ageing/afy178. [DOI] [PubMed] [Google Scholar]
  • 14.Zeng R, Zhang Y, Xu J, et al. Relationship of glycated hemoglobin A1c with all-cause and cardiovascular mortality among patients with hypertension. J Clin Med. 2023;12(7):2615. 10.3390/jcm12072615. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.La Morgia C, Di Vito L, Carelli V, Carbonelli M. Patterns of retinal ganglion cell damage in neurodegenerative disorders: parvocellular vs magnocellular degeneration in optical coherence tomography studies. Front Neurol. 2017;8:710. 10.3389/fneur.2017.00710. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Mutlu U, Colijn JM, Ikram MA, et al. Association of retinal neurodegeneration on optical coherence tomography with dementia: a population-based study. JAMA Neurol. 2018;75(10):1256–1263. 10.1001/jamaneurol.2018.1563. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Abràmoff MD, Lavin PT, Birch M, Shah N, Folk JC. Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices. NPJ Digit Med. 2018;1:39. 10.1038/s41746-018-0040-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Poplin R, Varadarajan AV, Blumer K, et al. Prediction of cardiovascular risk factors from retinal fundus photographs via deep learning. Nat Biomed Eng. 2018;2(3):158–164. 10.1038/s41551-018-0195-0. [DOI] [PubMed] [Google Scholar]
  • 19.Cheung CY, Xu D, Cheng CY, et al. A deep-learning system for the assessment of cardiovascular disease risk via the measurement of retinal-vessel calibre. Nat Biomed Eng. 2021;5(6):498–508. 10.1038/s41551-020-00626-4. [DOI] [PubMed] [Google Scholar]
  • 20.Gerrits N, Elen B, Craenendonck TV, et al. Age and sex affect deep learning prediction of cardiometabolic risk factors from retinal images. Sci Rep. 2020;10(1):9432. 10.1038/s41598-020-65794-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Jin K, Ye J. Artificial intelligence and deep learning in ophthalmology: current status and future perspectives. Adv Ophthalmol Pract Res. 2022;2(3), 100078. 10.1016/j.aopr.2022.100078. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Blumenkranz MS, Tarver ME, Myung D, Eydelman MB. Collaborative Community on Ophthalmic Imaging Executive Committee. The Collaborative Community on Ophthalmic Imaging: accelerating global innovation and clinical utility. Ophthalmology. 2022;129(2):e9–e13. 10.1016/j.ophtha.2021.10.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Health C for D and R. Software as a medical device (SaMD). FDA; 2020. Accessed May 16, 2024. 〈https://www.fda.gov/medical-devices/digital-health-center-excellence/software-medical-device-samd〉. [Google Scholar]
  • 24.Health C for D and R. Breakthrough devices program. FDA; 2024. Accessed June 3, 2024. 〈https://www.fda.gov/medical-devices/how-study-and-market-your-device/breakthrough-devices-program〉. [Google Scholar]
  • 25.Health C for D and R. Early feasibility studies (EFS) program. FDA; 2023. Accessed June 3, 2024. 〈https://www.fda.gov/medical-devices/investigational-device-exemption-ide/early-feasibility-studies-efs-program〉. [Google Scholar]
  • 26.Health C for D and R. Humanitarian device exemption. FDA; 2023. Accessed June 3, 2024. 〈https://www.fda.gov/medical-devices/premarket-submissions-selecting-and-preparing-correct-submission/humanitarian-device-exemption〉. [Google Scholar]
  • 27.Babenko B, Traynis I, Chen C, et al. A deep learning model for novel systemic biomarkers in photographs of the external eye: a retrospective study. Lancet Digit Health. 2023;5(5):e257–e264. 10.1016/S2589-7500(23)00022-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Saeedi P, Petersohn I, Salpea P, et al. Global and regional diabetes prevalence estimates for 2019 and projections for 2030 and 2045: results from the International Diabetes Federation Diabetes Atlas, 9th edition. Diabetes Res Clin Pract. 2019;157, 107843. 10.1016/j.diabres.2019.107843. [DOI] [PubMed] [Google Scholar]
  • 29.Dal Canto E, Ceriello A, Rydén L, et al. Diabetes as a cardiovascular risk factor: an overview of global trends of macro and micro vascular complications. Eur J Prev Cardiol. 2019;26(2_suppl):25–32. 10.1177/2047487319878371. [DOI] [PubMed] [Google Scholar]
  • 30.Mosenzon O, Cheng AY, Rabinstein AA, Sacco S. Diabetes and stroke: what are the connections? J Stroke. 2023;25(1):26–38. 10.5853/jos.2022.02306. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Crabtree GS, Chang JS. Management of complications and vision loss from proliferative diabetic retinopathy. Curr Diab Rep. 2021;21(9):33. 10.1007/s11892-021-01396-2. [DOI] [PubMed] [Google Scholar]
  • 32.Klein R, Klein BE, Moss SE. Visual impairment in diabetes. Ophthalmology. 1984;91(1):1–9. [PubMed] [Google Scholar]
  • 33.Ghartey KN. The importance of early detection of diabetic retinopathy. J Ophthalmic Nurs Technol. 1990;9(5):193–198. [PubMed] [Google Scholar]
  • 34.Kropp M, Golubnitschaja O, Mazurakova A, et al. Diabetic retinopathy as the leading cause of blindness and early predictor of cascading complications—risks and mitigation. EPMA J. 2023;14(1):21–42. 10.1007/s13167-023-00314-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Uusitupa M, Khan TA, Viguiliouk E, et al. Prevention of type 2 diabetes by lifestyle changes: a systematic review and meta-analysis. Nutrients. 2019;11(11):2611. 10.3390/nu11112611. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Lyons TJ, Basu A. Biomarkers in diabetes: hemoglobin A1c, vascular and tissue markers. Transl Res J Lab Clin Med. 2012;159(4):303–312. 10.1016/j.trsl.2012.01.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Simonyan K, Zisserman A. Very deep convolutional networks for large-scale image recognition; 2015. DOI: 10.48550/arXiv.1409.1556. [DOI] [Google Scholar]
  • 38.He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR). IEEE; 2016: p. 770–8. DOI: 10.1109/CVPR.2016.90. [DOI] [Google Scholar]
  • 39.Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16×16 words: transformers for image recognition at scale; 2021. DOI: 10.48550/arXiv.2010.11929. [DOI] [Google Scholar]
  • 40.Lee CS, Brandt JD, Lee AY. Big data and artificial intelligence in ophthalmology: where are we now? Ophthalmol Sci. 2021;1(2), 100036. 10.1016/j.xops.2021.100036. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Mitchell M, Wu S, Zaldivar A, et al. Model cards for model reporting. In: Proceedings of the conference on fairness, accountability, and transparency. FAT* ‘19. Association for Computing Machinery; 2019: p. 220–9. DOI: 10.1145/3287560.3287596. [DOI] [Google Scholar]
  • 42.Abdullah YI, Schuman JS, Shabsigh R, Caplan A, Al-Aswad LA. Ethics of artificial intelligence in medicine and ophthalmology. Asia-Pac J Ophthalmol Philos. 2021;10(3):289–298. 10.1097/APO.0000000000000397. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Kim HE, Cosa-Linan A, Santhanam N, Jannesari M, Maros ME, Ganslandt T. Transfer learning for medical image classification: a literature review. BMC Med Imaging. 2022;22(1):69. 10.1186/s12880-022-00793-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Yan B, Cao D, Jiang X, et al. FedEYE: a scalable and flexible end-to-end federated learning platform for ophthalmology. Patterns. 2024;5(2), 100928. 10.1016/j.patter.2024.100928. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Lo J, Yu TT, Ma D, et al. Federated learning for microvasculature segmentation and diabetic retinopathy classification of OCT data. Ophthalmol Sci. 2021;1(4), 100069. 10.1016/j.xops.2021.100069. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Nguyen TX, Ran AR, Hu X, et al. Federated learning in ocular imaging: current progress and future direction. Diagnostics. 2022;12(11):2835. 10.3390/diagnostics12112835. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Costa P, Galdran A, Meyer MI, et al. End-to-end adversarial retinal image synthesis. IEEE Trans Med Imaging. 2017;37(3):781–791. [DOI] [PubMed] [Google Scholar]
  • 48.Diaz-Pinto A, Colomer A, Naranjo V, Morales S, Xu Y, Frangi AF. Retinal image synthesis and semi-supervised learning for glaucoma assessment. IEEE Trans Med Imaging. 2019;38(9):2211–2218. [DOI] [PubMed] [Google Scholar]
  • 49.Chen H, Cao P. Deep learning based data augmentation and classification for limited medical data learning. In: Proceedings of the IEEE international conference on power, intelligent computing and systems (ICPICS). IEEE; 2019: p. 300–3. [Google Scholar]
  • 50.Balasubramanian R, Sowmya V, Gopalakrishnan E, Menon VK, Variyar VS, Soman K. Analysis of adversarial based augmentation for diabetic retinopathy disease grading. In: Proceedings of the 11th international conference on computing, communication and networking technologies (ICCCNT). IEEE; 2020: p. 1–5. [Google Scholar]
  • 51.Lim G, Thombre P, Lee ML, Hsu W. Generative data augmentation for diabetic retinopathy classification. In: Proceedings of the IEEE 32nd international conference on tools with artificial intelligence (ICTAI). IEEE; 2020: p. 1096–103. [Google Scholar]
  • 52.Choi JY, Ryu IH, Kim JK, Lee IS, Yoo TK. Development of a generative deep learning model to improve epiretinal membrane detection in fundus photography. BMC Med Inf Decis Mak. 2024;24(1):25. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Nderitu P, do Rio JMN, Webster L, et al. Conditional diffusion models and retinal image synthesis in diabetic retinopathy. Investig Ophthalmol Vis Sci. 2023;64(8), 2389–2389. [Google Scholar]
  • 54.Alimanov A, Islam MB. Denoising diffusion probabilistic model for retinal image generation and segmentation. In: Proceedings of the IEEE international conference on computational photography (ICCP). IEEE; 2023: p. 1–12. [Google Scholar]
  • 55.Kim HK, Ryu IH, Choi JY, Yoo TK. Early experience of adopting a generative diffusion model for the synthesis of fundus photographs; 2022.
  • 56.Go S, Ji Y, Park SJ, Lee S. Generation of structurally realistic retinal fundus images with diffusion models. ArXiv Prepr ArXiv230506813; 2023. [Google Scholar]
  • 57.Wu Y, He W, Eschweiler D, et al. Retinal OCT synthesis with denoising diffusion probabilistic models for layer segmentation. ArXiv Prepr ArXiv231105479; 2023. [Google Scholar]
  • 58.Agharezaei Z, Firouzi R, Hassanzadeh S, et al. Computer-aided diagnosis of keratoconus through VAE-augmented images using deep learning. Sci Rep. 2023;13(1), 20586. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Chen DK, Modi Y, Al-Aswad LA. Promoting transparency and standardization in ophthalmologic artificial intelligence: a call for artificial intelligence model card. Asia-Pac J Ophthalmol Philos. 2022;11(3):215–218. 10.1097/APO.0000000000000469. [DOI] [PubMed] [Google Scholar]
  • 60.Siontis KC, Suárez AB, Sehrawat O, et al. Saliency maps provide insights into artificial intelligence-based electrocardiography models for detecting hypertrophic cardiomyopathy. J Electrocardiol. 2023;81:286–291. 10.1016/j.jelectrocard.2023.07.002. [DOI] [PubMed] [Google Scholar]
  • 61.Muntean GA, Marginean A, Groza A, et al. The predictive capabilities of artificial intelligence-based OCT analysis for age-related macular degeneration progression-a systematic review. Diagnostics. 2023;13(14):2464. 10.3390/diagnostics13142464. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Cho K, Courville A, Bengio Y. Describing multimedia content using attention-based encoder-decoder networks. IEEE Trans Multimed. 2015;17(11):1875–1886. 10.1109/TMM.2015.2477044. [DOI] [Google Scholar]
  • 63.Xu K, Ba J, Kiros R, et al. Show, attend and tell: neural image caption generation with visual attention; 2016. DOI: 10.48550/arXiv.1502.03044. [DOI] [Google Scholar]
  • 64.Sheng B, Chen X, Li T, et al. An overview of artificial intelligence in diabetic retinopathy and other ocular diseases. Front Public Health. 2022;10, 971943. 10.3389/fpubh.2022.971943. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Niu Y, Gu L, Zhao Y, Lu F. Explainable diabetic retinopathy detection and retinal image generation. IEEE J Biomed Health Inf. 2022;26(1):44–55. 10.1109/JBHI.2021.3110593. [DOI] [PubMed] [Google Scholar]
  • 66.Zeiler MD, Fergus R. Visualizing and understanding convolutional networks. In: Fleet D, Pajdla T, Schiele B, Tuytelaars T, eds. Computer Vision – ECCV 2014. Springer International Publishing; 2014:818–833. 10.1007/978-3-319-10590-1_53. [DOI] [Google Scholar]
  • 67.Li Z, Wang L, Wu X, et al. Artificial intelligence in ophthalmology: the path to the real-world clinic. Cell Rep Med. 2023;4(7), 101095. 10.1016/j.xcrm.2023.101095. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Chaddad A, Peng J, Xu J, Bouridane A. Survey of explainable AI techniques in healthcare. Sensors. 2023;23(2):634. 10.3390/s23020634. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Gunning D, Vorm E, Wang JY, Turek M. DARPA’s explainable AI (XAI) program: a retrospective. Appl AI Lett. 2021;2(4), e61. 10.1002/ail2.61. [DOI] [Google Scholar]
  • 70.Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat Mach Intell. 2019;1(5):206–215. 10.1038/s42256-019-0048-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Arnould L, Meriaudeau F, Guenancia C, et al. Using artificial intelligence to analyse the retinal vascular network: the future of cardiovascular risk assessment based on oculomics? A narrative review. Ophthalmol Ther. 2023;12(2):657–674. 10.1007/s40123-022-00641-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Evans NG, Wenner DM, Cohen IG, et al. Emerging ethical considerations for the use of artificial intelligence in ophthalmology. Ophthalmol Sci. 2022;2(2), 100141. 10.1016/j.xops.2022.100141. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Coyner AS, Singh P, Brown JM, et al. Association of biomarker-based artificial intelligence with risk of racial bias in retinal images. JAMA Ophthalmol. 2023;141(6):543–552. 10.1001/jamaophthalmol.2023.1310. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Joshi N, Burlina P. AI fairness via domain adaptation; 2021. DOI: 10.48550/arXiv.2104.01109. [DOI] [Google Scholar]
  • 75.Browne AW, Kim G, Vu AN, et al. Deep learning assisted imaging methods to facilitate access to ophthalmic telepathology. Ophthalmol Sci. 2024;4(3), 100450. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Jui-Kai W, Garvin MK, Kupersmith MJ, Kardon RH. Quantifying spatial patterns of OCT total retinal thickness (TRT) in Papilledema over time using a deep learning variational AutoEncoder. Investig Ophthalmol Vis Sci. 2022;63(7), 436–436. [Google Scholar]
  • 77.Mandal S, Jammal AA, Medeiros FA. Assessing glaucoma in retinal fundus photographs using deep feature consistent variational autoencoders. ArXiv Prepr ArXiv211001534; 2021. [Google Scholar]
  • 78.He J, Baxter SL, Xu J, Xu J, Zhou X, Zhang K. The practical implementation of artificial intelligence technologies in medicine. Nat Med. 2019;25(1):30–36. 10.1038/s41591-018-0307-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Liang S, Li Y, Srikant R. Enhancing the reliability of out-of-distribution image detection in neural networks. ArXiv Learn; 2017. Accessed May 22, 2024. 〈https://www.semanticscholar.org/paper/Enhancing-The-Reliability-of-Out-of-distribution-in-Liang-Li/547c854985629cfa9404a5ba8ca29367b5f8c25f〉. [Google Scholar]
  • 80.Schulam P, Saria S. Can you trust this prediction? Auditing pointwise reliability after learning. In: Proceedings of the twenty-second international conference on artificial intelligence and statistics. PMLR; 2019: p. 1022–31. Accessed May 22, 2024. https://proceedings.mlr.press/v89/schulam19a.html. [Google Scholar]
  • 81.Schlegl T, Seeböck P, Waldstein SM. Unsupervised anomaly detection with generative adversarial networks to guide marker discovery. In: Niethammer M, Styner M, Aylward S, et al. , eds. Information processing in medical imaging. Springer International Publishing; 2017:146–157. 10.1007/978-3-319-59050-9_12. [DOI] [Google Scholar]
  • 82.Subbaswamy A, Saria S. From development to deployment: dataset shift, causality, and shift-stable models in health AI. Biostatistics. 2020;21(2):345–352, 10.1093/biostatistics/kxz041. [DOI] [PubMed] [Google Scholar]
  • 83.Ribeiro MT, Singh S, Guestrin C. “Why should I trust you?”: explaining the predictions of any classifier. In: Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. ACM; 2016: p. 1135–44. DOI: 10.1145/2939672.2939778. [DOI] [Google Scholar]
  • 84.Lundberg SM, Lee SI. A unified approach to interpreting model predictions. In: Proceedings of the 31st international conference on neural information processing systems. NIPS’17. Curran Associates Inc.; 2017: p. 4768–77. [Google Scholar]
  • 85.Cruz Rivera S, Liu X, Chan AW, Denniston AK, Calvert MJ. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. Nat Med. 2020;26(9):1351–1363. 10.1038/s41591-020-1037-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Weber S, Wyszynski M, Godefroid M, Plattfaut R, Niehaves B. How do medical professionals make sense (or not) of AI? A social-media-based computational grounded theory study and an online survey. Comput Struct Biotechnol J. 2024;24:146–159. 10.1016/j.csbj.2024.02.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Davenport T, Kalakota R. The potential for artificial intelligence in healthcare. Future Health J. 2019;6(2):94–98. 10.7861/futurehosp.6-2-94. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Alanazi A. Clinicians’ views on using artificial intelligence in healthcare: opportunities, challenges, and beyond. Cureus; 15(9): e45255. DOI: 10.7759/cureus.45255. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Nakagawa K, Moukheiber L, Celi LA, et al. AI in pathology: what could possibly go wrong? Semin Diagn Pathol. 2023;40(2):100–108. 10.1053/j.semdp.2023.02.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Cestonaro C, Delicati A, Marcante B, Caenazzo L, Tozzo P. Defining medical liability when artificial intelligence is applied on diagnostic algorithms: a systematic review. Front Med. 2023;10, 1305756. 10.3389/fmed.2023.1305756. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Alowais SA, Alghamdi SS, Alsuhebany N, et al. Revolutionizing healthcare: the role of artificial intelligence in clinical practice. BMC Med Educ. 2023;23(1):689. 10.1186/s12909-023-04698-z. [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES