Skip to main content
Ophthalmology Science logoLink to Ophthalmology Science
. 2026 Mar 11;6(7):101149. doi: 10.1016/j.xops.2026.101149

Toward Long-Term Visual Field Appearance Forecasting Using Artificial Intelligence for Ophthalmic Education and Diagnosis

Ye Tian 1,∗, Puja Bhavsar 2, Mingyang Zang 1, Zhuoying Gu 4, Ari Leshno 4, Allison Cui 3, Ives A Valenzuela 4, Carlos Gustavo De Moraes 4, Emmanouil Tsamis 4, Kaveri A Thakoor 1,4
PMCID: PMC13235488  PMID: 42256006

Abstract

Objective

To provide state-of-the-art, post hoc-explainable visual field (VF) forecasts to aid in training ophthalmic residents to characterize glaucoma progression (GP), we train artificial intelligence (AI) to take as input VFs to detect GP and forecast VF appearance up to 10 years in the future. We conduct a user study to assess the clinical utility of forecasted VFs.

Design

A retrospective study evaluating the impact of AI-assisted VF forecasts on the clinical decision-making of Columbia University Irving Medical Center ophthalmology residents.

Subjects

For AI training, we utilize the largest available VF dataset to date including minority populations. We invited 8 ophthalmology residents to participate in the user study and 3 faculty to provide ground truth ratings.

Methods

We implemented GenVF, a generative vision transformer for forecasting future VFs up to 10 years in the future. We showcase the promising utility of AI-forecasted VFs (AI-VF) by conducting a user study to evaluate the impact of AI information on the clinical decision-making process. In the user study, we provide residents with VFs generated by AI side-by-side with the actual first 3 VFs for a given patient to compare the outcome difference with versus without AI.

Main Outcome Measures

We assess GP rating performance using mean absolute error (MAE), confidence levels, and reaction time with versus without AI guidance, analyzing all metrics with the nonparametric Wilcoxon signed-rank test.

Results

Globally, MAE for GP ratings with and without AI forecasts show minor differences. Stratification by postgraduate year levels reveals the AI-VF condition demonstrates a uniform MAE, while VF-only demonstrates variability. Patient-specific analysis of resident performance reveals heterogeneous impacts of AI forecasts. Residents make more conservative decisions and report higher confidence levels with AI forecasts, while reaction times remain comparable with or without AI guidance.

Conclusions

Our state-of-the-art, post hoc-explainable Al method using our GenVF model impacts performance, confidence, and time of GP detection, yielding more conservative decisions from user-study participants, which may help expedite glaucoma treatment for broader patient populations. Future work will address challenges in the variability of AI effectiveness across expertise levels and patient-specific contexts.

Financial Disclosure(s)

The authors have no proprietary or commercial interest in any materials discussed in this article.

Keywords: Glaucoma progression, Artificial intelligence, User study, Ophthalmic education


Glaucoma is a progressive disease that damages the optic nerve and may cause blindness. It is estimated to impact >100 million people globally by 2040.1 OCT measurements provide an evaluation of structural damage to retinal layers.2 Playing a complementary role, standard automated perimetry tests, also known as visual field (VF) tests, are used to evaluate the state of functional vision of glaucoma patients.3 The structural and functional alterations in the retina are pivotal biomarkers for tracking glaucoma progression (GP).4

Clinical analysis usually employs 3 consecutive VF tests to identify potential GP and reveal identical points that demonstrate statistically significant changes. The grayscale values in the VF depict the patient's functional sensitivity to light at various positions in their VF. The darker shades correspond to lower visual sensitivity. Clinicians depend on the available VF test series obtained during earlier follow-ups to make an informed decision about glaucoma treatment for a given patient.5 Though timely treatment and monitoring are essential in managing glaucoma, subjective assessment of VF progression can be unreliable due to diagnostic bias and expertise level.6 Also, automated VF analyzers are expensive and not yet available in all regions worldwide.7 Therefore, artificial intelligence (AI)-based methods for forecasting VFs have emerged as a paradigm shift and catalyst for improving GP diagnosis and treatment.8, 9, 10, 11

Within the last 5 years, AI, specifically machine learning algorithms, have achieved equivalent or sometimes better accuracy than existing mathematical approaches at VF forecasting.12,13 Multiple mathematical algorithms are used to detect GP in clinical practice, including the slope of the mean deviation (MD) (measured in decibels [dB/year]) and guided progression analysis. Artificial intelligence approaches such as image-to-image translation algorithms, on the other hand, overcome the problem of overparameterization in numerical models,14 offering a more balanced and robust solution. While maintaining high robustness, machine learning approaches are more generalizable and adaptable to a broader range of datasets after training on diverse data distributions.

In this paper, we adapted and improved existing AI models to provide more accurate VF forecasts than a current state-of-the-art approach15 and extended the predictive power to up to 10 years in the future. The AI-based forecasting of VF appearance >5 years into the future has not been demonstrated in previous studies and remains challenging due to limited input information and complex temporal variations.16,17 Ishikawa et al17 achieve mean absolute errors (MAEs) within 2.75 dB when predicting up to 4 years into the future, and Abbasi et al18 achieve MAEs ranging from 1.5 to 2 when MD and false-negative rate are low in the dataset.

Here, we define our AI-based VF forecasting method as capable of long-term forecasting because it can generate VFs up to 10 years into the future using more advanced transformer-based models. Compared with these previous studies, our work has a larger dataset with minority cases and high MD, uses a more advanced model, and only leverages one baseline VF as input.

We utilized deep learning models to forecast future VF appearances. By comparing the outcomes of both convolutional-based networks and architectures leveraging attention mechanisms (e.g., generative vision transformer), we selected and further implemented the best model for downstream artificial-VF (AI-VF) generation and the subsequent user study.

We used 24-2 VFs, which measure 24° temporally and 30° nasally, and tested 54 points as they are the most common VF test repeated for patients at routine intervals. The 24-2 VFs capture a wider field compared to 10-2 VFs that focus on the central field in greater detail.19 Furthermore, we utilized one of the most extensive available dataset of VFs to-date, including minority populations of African, Hispanic, and Asian descent, to train our models, thus ensuring our models' generalizability to new test patient data from minority populations.

To evaluate the impact of AI information on the clinical diagnosis process, we designed and conducted a user study via PsychoPy20 software and its online experiment platform, Pavlovia. We provided a graphical user interface presenting both ground truth VFs and our AI-forecasted VFs to clinical experts, including glaucoma residents and faculty, to obtain their decision-making speed, accuracy, and confidence level with or without AI-generated VFs. Their qualitative feedback, in conjunction with our quantitative metrics, collectively assessed the trustworthiness of our AI tool and its potential impact when integrated into the clinical workflow. More details of the user study are explained in the Model Architecture of the Methods section.

Methods

Datasets for AI Training

Our study used 24-2 Humphrey visual field (VF) tests represented by total deviation values in dB, as these reflect localized functional loss relative to age-adjusted norms and are widely used in glaucoma research.

We leveraged our Columbia University Irving Medical Center patient population to curate the 62K+ dataset, one of the largest available dataset of VFs to date, including minority populations. The dataset includes raw retrospective VF test data of up to 5167 patients, with 62 197 VFs collected between October 2010 and October 2023 (IRB-AAAU5610, AAAT9362). Our 62K+ dataset consists of 6.7% individuals of African descent, 6.8% of Asian descent, and 3.7% of Hispanic descent. This dataset thus includes the most significant Hispanic representation of existing glaucoma datasets containing ophthalmic imaging. Compared with the University of Washington Humphrey Visual Field (UWHVF) public dataset,21 the 62K+ dataset includes more VFs and patient group diversity. A detailed summary of both datasets is in Table S1 (www.ophthalmologyscience.org).

Preprocessing and Training

Utilizing VFs alone, the AI model cannot fully address the limitations associated with glaucomatous changes that may occur in local VF points, as it primarily focuses on minimizing the global image error instead of optimizing qualitative VF outcome. Thus, we introduced the second modality of retinal nerve fiber layer (RNFL) thickness values to advance performance.

We first excluded patients with comorbid ocular conditions (e.g., macular degeneration and diabetic retinopathy) that could confound VF changes. Only baseline–forecast VF pairs accompanied by OCT examinations within 3 months of the baseline were retained. Then, we included only the VF baseline-forecast pairs that were accompanied by OCT tests conducted within a 3-month timeframe from the baseline VF test. After data filtering and preprocessing, the VF sectors and their corresponding mapping on the optic nerve head are taken from Garway-Heath.22 The corresponding RNFL thickness values (μm) were extracted from OCT images, following the Garway-Heath sector mapping for topographic correspondence between VF and optic disc regions. The number of VFs in training and validation is outlined in Figure S1 (www.ophthalmologyscience.org).

The data preprocessing handled VF points and RNFL thickness in 2 different input channels. Each pixel in the RNFL input channel corresponds to the exact RNFL thickness value in that specific Garway-Heath Zone, as shown in Figure 1A. The RNFL thickness values are mapped on the original VF based on their pixel correspondence on the Garway-Heath zone to generate a weighted sum (fusion map). Then normalization and smoothing are applied to create the final input to the model.

Figure 1.

Figure 1

Overview of our study pipeline, including (A) AI model training and (B) clinical user study. A, The GenVF model was trained to predict future VFs using both baseline VFs and RNFL thickness, mapped according to the 6 Garway-Heath zones. These zones provide a clinically aligned anatomical correspondence between structural and functional VF regions, which opens the “black-box” by mapping AI inputs to ophthalmic knowledge. Each pixel in the input field was modulated by the RNFL thickness of its corresponding zone to provide physiologically informed guidance. The GenVF model backbone includes (from top to bottom): patch embedding, linear projection, attention block, token mapping, and patch reconstruction. B, The clinical user study was implemented via a PsychoPy interface. Glaucoma residents evaluated VF progression based on real and AI-forecasted VFs, while glaucoma faculty provided the ground truth. AI = artificial intelligence; RNFL = retinal nerve fiber layer; VF = visual field.

To ensure consistency, all VF images in both the full 62K+ dataset and its RNFL-paired subset were standardized to right eye. Each 24-2 VF was structured as an 8 × 9 matrix and 0-padded to 9 × 9 for computational convenience. Within this format, 54 points carried informative perimetry data, while the remaining positions were assigned a 0 value as a nonnoise background. Before training, all VF and RNFL values were normalized to [0, 1] using min–max normalization to ensure numerical stability and consistent scale across patients. We also applied a Gaussian smoothing filter (σ = 1) to both modalities. The final inputs thus preserve local structural patterns while minimizing intertest variability. Additionally, the VF test year was recorded to calculate time intervals between tests and assess GP rates.

Figure 1 illustrates the study pipeline, including the multimodal training of the deep learning model and the downstream user study with glaucoma residents. The model takes 1 VF input to generate 1 VF output. We have 3 independent models for the 3 time ranges. There are 318 patients (approximately 3%) whose second VF record occurred >5 years after their baseline VF, and 1164 patients (approximately 11%) whose second VF was recorded >2 years after baseline VF.

We have 3 AI models: a “0∼2-year model”, a “2∼5-year model,” and a “5∼10-year model” that share an identical algorithm but work on forecasting future VFs in the 3 different time ranges in the future. To prove the AI's generalizability, testing was done on the external UWHVF dataset. The GenVF model is trained by Adam optimizer with a batch size of 36 and a learning rate decaying from 1e-4. The training set includes both progressing and nonprogressing cases.

To visualize the model heatmap of our 9 × 9 VF image, we redefine each Garway-Heath zone as an individual feature in a 1-dimensional vector. We then apply a permutation importance algorithm to select features and quantify the relative contribution of values in the 6 zones to the prediction. We then compute the layer output before patch reconstruction in GenVF and normalize it to [0, 1]. This produces a heatmap over the VF grid, and we project it onto the 6 zones' location on VF.

Model Architecture

To achieve point-wise VF-to-VF generation, we used a patient's initial VF to forecast future VFs in 1 to 2 years, 3 to 5 years, and over 5 years, respectively. We utilized and modified the generative vision transformer (GenViT) model, a combination of a ViT and a diffusion model23 that has demonstrated higher output quality than other image generation models, such as generative adversarial networks.24 We named our GenViT adaptation “GenVF” due to its ability to forecast future VFs. See the model architecture of GenVF in Figure 1A, which was originated from our previous work.15 Our GenVF breaks the image into patches and feeds them into vision transformer blocks. At the output stage, instead of having a multilayer perceptron as a classification head, it reconstructs the patches based on their positional encoding, forming a desired output image with the same shape as the input. Without requiring additional demographic information, as required by the previous, state-of-the-art convolution-based approach,13 our model performs better in point-wise MAE, extends forecasting power to >5 years in the future, and generalizes well on the unseen UWHVF dataset. The performance of GenVF is shown in the Results section.

Integrating both modalities caused a significantly smaller scale of usable sample size because the majority of patients did not have an OCT test result within 3 months of the baseline VF test and had to be excluded from our multimodal pairing strategy. Therefore, we also built smaller models with fewer trainable parameters from scratch, including U-Net, a convolutional model for biomedical image segmentation.25 We combined the RNFL map and the VF image, resulting in a 2-channel input fed into the U-Net. In the network, each downsampling step is done via 2 convolutional blocks. The first block has a kernel size of 3, and the second block has a kernel size of 1. The ReLU function is applied after each convolutional block.

To fully make use of our VF and RNFL modalities, we also leveraged the popular multimodal framework, Contrastive Language-Image Pretraining (CLIP) model,26 which efficiently learns visual concepts from textual supervision. We implemented CLIP by vectorizing the RNFL thicknesses from OCT reports as a 1-dimensional input (i.e., treating it as text) to guide the learning of VF images. A depiction of how CLIP model takes VF and RNFL as input is shown in Figure S2 (www.ophthalmologyscience.org). All VFs and their associated RNFLs were passed through their respective encoders, mapping all objects into a 256-dimensional space. The cosine similarity of each (VF and RNFL-thickness-vector) pair was maximized as the training objective of contrastive learning.

We conducted an ablation study and compared the generated VFs with ground truth VFs in terms of MAE. The results are shown in Tables 1 and 2. Due to its superior performance across all metrics, we selected GenVF as our best AI model to generate artificial VFs for the downstream user study.

Table 1.

GenVF Performance for Each Time Bin on Both 62K+ and UWHVF Datasets

Time Bins 62K+ Dataset
UWHVF Dataset
Global MAE VF Deficit MAE (Mild/Moderate/Severe) Global MAE VF Deficit MAE (Mild/Moderate/Severe) Wen et al.
0∼2 yr 2.30 1.82/3.65/5.67 2.54 1.57/3.39/4.76 2.68
2∼5 yr 2.37 1.63/3.41/5.13 2.64 1.63/3.60/4.81 2.71
5∼10 yr 2.75 1.61/3.30/4.45 2.80 1.98/3.79/4.94 NA

MAE = mean absolute error; UWHVF = University of Washington Humphrey Visual Field; VF = visual field.

The time bins include 0∼2, 2∼5, and 5∼10 years in the future. The unit is pixelwise MAE. For each dataset, the MAE is reported both globally and also under different VF deficits (mild, moderate, and severe). The last column of UWHVF provides a direct comparison between previous deep learning approach by Wen et al. and our global MAE.

Bold indicates GenVF performance for both datasets.

Table 2.

Artificial Intelligence Model Performance before and after Employing the RNFL Thickness Map as a Second Input Modality

Model VF Only VF + RNFL
GenVF (ours) 3.43 2.99
UNet 3.03 3.67
CLIP26 NA 3.25

CLIP = Contrastive Language-Image Pretraining; MAE = mean absolute error; NA = not applicable; RNFL = retinal nerve fiber layer; VF = visual field.

The models we are comparing are our GenVF model, convolutional-based UNet model and CLIP visual-text (RNFL-VF) model. Because CLIP requires both 1-dimensional and 2-dimensional input, its VF-only result is not applicable. The evaluation metric is MAE.

Bold indicates the best MAE performance achieved among all models and settings (with/without RNFL).

User Study

Beyond achieving a quantitatively low error rate between forecasted and ground-truth VFs, ensuring qualitatively promising performance of AI-forecasted VFs is essential for their trustworthiness in clinical applications. We addressed this by conducting a user study that examined the influence of additional AI information on the clinical decision-making process. In the user study, each clinician was presented with 24 slides. There were 12 slides in which participants only saw the patient's first 3 real VFs. In another set of 12 slides, 2 AI-forecasted VFs were shown to the participant along with 3 real VFs. These 3 AI-generated VFs were in the time range of 0∼2, 2∼5, and 5∼10 years into the future, respectively.

In instances where a patient had multiple existing real VFs within a given time range, we excluded the corresponding AI-forecasted VFs. For example, if 1 patient's VFs were available at baseline time, then 1.1 years from baseline, and lastly 2.2 years from baseline, the AI forecast for the 0∼2-year interval would be omitted. In contrast, forecasts for the 2∼5-year and 5∼10-year intervals were still provided to clinicians.

We designed our user study's graphical user interface using Psychopy, an open-source Python tool for designing behavioral experiments,23 as shown in the example slide in Figure 2. Online participation and data transmission were done via Pavlovia, a website interface for hosting such experiments on any laptop/desktop computer screen. Visual fields with and without side-by-side AI forecasts were provided to clinicians who participated in this study. For example, in Figure 2, clinicians were asked to forecast VF progression over a 0∼2-year interval, as the real VFs provide information about the actual 0∼2-year progression on the left side of Figure 2. Note that the VF time interval in this example does not apply to all cases, and residents were instructed to respond based on the specific time range indicated by the real VFs in each case. The response data we collected were the user's GP rating score (from 1 to 5, 1 being no progression, and 5 being definitely GP), clinical diagnosis (monitor or treat), and confidence in their diagnosis (on a scale between 1 and 5, 1 being least confident, and 5 being most confident).

Figure 2.

Figure 2

An example slide in the GUI of the user study. Guidance is printed on top; VFs, including real ones and occasionally AI-forecasted ones, are shown side by side in the middle. Buttons and bars for giving responses are at the bottom. There are 24 slides in the whole study. AI = artificial intelligence; GUI = graphical user interface; VF = visual field.

As introduced earlier, each subject was shown 24 sets of patient VF images (retrospective data collected from a previously approved institutional review board protocol: AAAT9362). An equal number of images were shown in randomized order without AI (12 images) and with AI information (12 images). Five of 24 cases were replicas of an earlier-presented patient but with the AI condition altered. For example, if a patient with real and AI-generated VFs was presented as case #1, then that case may have been presented again with real but no AI-generated VFs as #24. This resulted in a total number of 19 distinct patients. Besides the AI information appearing randomly in 12 of the 24 image samples, all other aspects of the graphical user interface input/output collection were exactly the same for all samples. In the 19 patient cases in our user study, the MD is within the range of [–19.1 to –1.4]. The average MD is –9.2. Of 19 total cases, 6 are severe, 8 are moderate, and 5 are mild based on the Parish severity criteria.27 The average number of VFs per patient is 5.8, and the average MD change rate is –1.2 dB/year.

We invited 8 ophthalmology residents ranging in postgraduate years (PGYs) from 1 to 4 as our participants. On average, each participant took 10.2 minutes to submit their response. For ground truth, we sent all real VFs (on average, 5.7 VFs in 7.2 years) of all 24 patients to 3 glaucoma faculties. Their responses included a GP rating and a decision on treatment (monitor vs. treat). Two ground truths were computed: one by determining the average between the 3 faculty members (consensus method) and the second by only utilizing the third faculty member's decision, if there were discrepancies between the first 2 (tie-breaking method). We used the faculty's consensus as our final ground truth to evaluate the resident participants' performance.

This setup allowed us to assess clinician accuracy of diagnosis, time to diagnosis, and confidence in diagnosis, with or without AI forecasts as an aid. Upon analyzing feedback from each subject regarding the same 24 sets of VFs, we quantified the results in terms of correct diagnosis rate (%) and decision error (MAE), decision-making speed (seconds/decision), and self-reported confidence level.

To mitigate individual bias and ensure generalizable conclusions, we analyzed both group-level averaged results and individual outcomes in the Results section. For statistical validity, we powered our study to detect a medium effect size (Cohen d = 0.5) with 80% power. Based on the expected difference in metric, i.e., MAE, between conditions, this calculation led us to recruit 8 subjects, resulting in 96 samples per class (with a minimum of 64 samples per class required to achieve this statistical power). This approach guaranteed a sufficiently powered dataset size for analysis, enabling us to draw meaningful conclusions regarding the impact of AI presence on clinical GP diagnosis and workflow.

This study, Protocol AAAU5610, was approved by the Columbia University Irving Medical Center Institutional Review Board on March 30, 2024, and is in accordance with the tenets set forth by the Declaration of Helsinki. Informed consent was obtained from all study participants.

Results

To show the state-of-the-art performance of our AI, we first present the performance of our GenVF model under the input condition of VF only and VF + RNFL. The results are in the Tables 1 and 2. To comprehensively evaluate the impact of AI-generated forecasts on clinical decision-making, we assessed clinicians' diagnostic performance, ratings (decision rating, GP rating, and confidence rating), and reaction time. The same 8 clinicians were evaluated under both the AI-VF (with AI forecasts) and VF-only (without AI forecasts) conditions. Given the data were paired and not normally distributed, the nonparametric Wilcoxon signed rank test was conducted for each pair of clinicians' responses with/without AI to determine the significance of the results.

GenVF Model Performance

We offer AI-generated VFs as an alternative for GP assessment beyond clinical data-processing/conventional GP analysis approaches. The evaluation metric of VF predictive error is the standard pixel-wise MAE. In Tables 1 and 2, we present the results with VF-only input and VF + RNFL input, respectively.

The Figure S3 (www.ophthalmologyscience.org) shows how the heatmaps change as a function of adding RNFL (left-to-right) and as a function of increasing prediction horizon (top-to-bottom). Adding RNFL information increases the importance of the superior temporal and superior nasal regions, as portrayed by the increasing values in these regions for heatmaps on the right. Increasing forecast time ranges (from 0∼2 to 2∼5 and to 5∼10 years) retains the importance of the inferior nasal and inferior temporal regions, as these regions remain degenerated over progression of the disease. Overall, the inferior nasal and inferior temporal regions are the most contributing to AI's decision, consistent with the clinical observation that early glaucomatous damage and functional loss in these zones are particularly informative for long-term forecasting.

For the VF-only condition, we also tested our model on the external UWHVF21 dataset to prove its generalizability. We also compare our model's capacity for forecasting mild, moderate, and severe VF deficits. We defined mild VF deficits as an MD of –6 dB or better; moderate as a MD of –6 to –12 dB, and severe as a MD of –12 dB or worse. This is based on the Parish Criteria.27 We evaluated our model on subgroups to further showcase that GenVF's generalizable performance extends to minority groups in the data (see Table S2, www.ophthalmologyscience.org).

GenVF performance declines as the forecasting horizon increases and as the number of training samples decreases when deviating from the population mean. This is consistent with other AI forecasting models.16,17

After integrating RNFL input as the second modality, our GenVF (attention-based) model surpassed the convolutional-based UNet and contrastive multimodal CLIP model and thus was selected as the best-performing AI backbone. The structural information from RNFL thickness improved the visualization of output VFs, making them more accurate by exhibiting better grayscale dynamic range in the transition between adjacent subregions. This bimodal training translated the RNFL thickness data as a second input modality to the model with the same spatial dimensions as the VF input.

GP Rating and Decision-Making Performance Stratified by Patients

The MAE for the GP rating under the AI-VF and VF-only conditions only portrayed minor differences at the global level. In Figure 3A, the MAE is slightly higher for the AI-VF condition than for the VF-only condition. This difference was not statistically significant (P = 0.110). When an alternative ground truth was employed in Figure 3B (tie-breaking method), there was a marginal elevation in the P value for the AI-VF condition compared to that for the VF-only condition without reaching statistical significance (P = 0.276).

Figure 3.

Figure 3

Mean GP global rating: mean absolute error under the AI-VF and VF-only conditions across the entire dataset. A, Mean absolute error based on ground truths generated from the consensus method. B, Mean absolute error based on ground truths generated from the tie-breaking method. C, Mean absolute error variation among participating residents from 4 PGY training levels, based on ground truths generated from the consensus method. Pairwise comparisons for AI-VF condition: PGY1 vs. PGY2 (P = 0.920), PGY1 vs. PGY3 (P = 0.879), PGY1 vs. PGY4 (P = 0.674), PGY2 vs. PGY3 (P = 0.862), PGY2 vs. PGY4 (P = 0.759), PGY3 vs. PGY4 (P = 0.560). Pairwise comparisons for VF-only condition: PGY1 vs. PGY2 (P = 0.058), PGY1 vs. PGY3 (P = 0.283), PGY1 vs. PGY4 (P = 0.795), PGY2 vs. PGY3 (P = 0.262), PGY2 vs. PGY4 (P = 0.106), PGY3 vs. PGY4 (P = 0.476). AI = artificial intelligence; GP = glaucoma progression; PGY = postgraduate year; VF = visual field.

The MAE was further stratified and examined across PGY1 to PGY4 training levels to assess the impact of clinical expertise on GP rating error in Figure 3C. The MAE remained slightly higher for the AI-VF condition than for the VF-only condition for all the PGY levels except for PGY2, although there were no statistically significant differences (P values: PGY1 = 0.096, PGY2 = 0.603, PGY3 = 0.231, PGY4 = 0.527) between the 2 conditions across all 4 training levels. This suggests that the inclusion of AI-VFs does not necessarily result in a systematic degradation of performance. Rather, it indicates that decisions made with AI-VF's assistance remain within a similar error range as those made with VF-only assistance, regardless of the method used to obtain ground truth.

Interestingly, in Figure 3C, the VF-only condition displayed variability in MAE along the PGY levels. In contrast, the MAE for the AI-VF condition was uniform across the clinical expertise levels. This pattern suggests that AI integration may assist in standardizing performance across the expertise levels, reducing disparities between junior and senior residents. The reduction in variability suggests AI may provide a standardizing effect, where GP rating assessments are consistent regardless of clinical expertise.

Pairwise comparisons of MAE for each condition (AI-VF across all 4 PGY levels and for VF-only across all 4 PGY levels) support these observations. The AI-VF P values ranged from 0.5∼0.9, indicating no significant difference in performance across clinical expertise levels. The VF-only P values, alternatively, range from 0.05∼0.7. There was a trend toward statistical significance in the difference between the performance of PGY1 and PGY2 residents (P = 0.058, not shown in the figure) for the VF-only condition, suggesting that AI may assist in bringing novice residents (PGY1) to comparable performance levels as senior residents (PGY2).

The results of MAE for GP rating for patient-specific cases are shown in Figure 4. The P values determined by the Wilcoxon signed rank test exceed the conventional P value threshold of P < 0.05, suggesting that the differences in GP rating error between the 2 conditions are not statistically significant for all of the patients. However, substantial interpatient variability exists as displayed by the magnitude of the error bars, suggesting patient-specific effects on AI-VF forecasts. This variability may arise from differences in GP patterns across patients, such as focal versus diffuse loss. Patients with severe disease may progress differently than those with mild disease; due to variability in disease severity, rates of change will also vary. We wanted to include a broad range of disease severity to see the impact of AI on a global spectrum.

Figure 4.

Figure 4

Mean GP rating error under the AI-VF and VF-only conditions across 19 patients, each assessed by 8 residents. Error bars represent the variability of the ratings relative to the ground truth. AI = artificial intelligence; GP = glaucoma progression; VF = visual field.

For some patients, such as patients 1 and 13, the mean GP rating error is higher for the AI-VF condition compared with the VF-only condition. For other patients, such as patients 10 and 19, the mean GP rating error is lower for the AI-VF condition compared with the VF-only condition. Additionally, for many patients, such as patients 3, 7, and 18, there is negligible or minimal difference in GP rating error between both conditions. These observed differences highlight the heterogeneous impact of AI forecasts, in which some benefit from AI integration, whereas others experience reduced performance. Hence, when evaluating the integration of AI in clinical decision-making, this emphasizes the importance of personalized analysis.

For 13 of 19 patients in Figure 5, decisions made with the AI-VF forecast consistently exhibited higher mean decision values than those made without the AI-VF forecasts. For instance, patients 3, 4, 5, 9, 10, 11, and 17 displayed substantial differences between AI-VF and VF-only conditions, with the AI-VF notably higher. Interestingly, for 6 of the 9 patients with ground-truth decisions of “Treat,” the AI-VF cases yielded mean decisions in alignment with the ground truth.

Figure 5.

Figure 5

Mean decision values across 19 patients, each assessed by 8 residents for the “AI-VF” and “VF-only” conditions, with the ground truth (Binary: 1 = Monitor, or 2 = Treat) overlaid as dashed lines for each patient as reference. AI = artificial intelligence; VF = visual field.

On the other hand, for the 10 patients with ground-truth decisions of “Monitor,” none of the AI-VF cases yielded mean decisions precisely in alignment with the ground truth. In 3 of the 10 patients, both the AI-VF and VF-only cases achieved tied/equal mean decisions, neither favoring nor disfavoring the “Monitor” ground truth, while for 7 of the 10 patients, the AI-VF suggested a decision of “treat,” while residents picked a decision of “monitor.” There is a general trend toward more conservative (treatment-oriented) clinical decisions with the availability of AI-VFs. This conservativeness may not improve accuracy in every case but can be useful for flagging possible risk, helping doctors avoid missed cases that could get worse with time (reduce false-negative rate) by detecting potentially concerning patterns that may not be evident in the most recent clinical scans alone. To prevent this bias toward extra conservative guidance, a future step would be using focal loss or weighted VF input to weigh the minor cases more highly and to utilize a more balanced training set.

In the top half of Figure 6, patient #9 has in total 5 longitudinal VFs ①∼⑤ that were shown to glaucoma faculty for providing a ground truth; real VFs ①∼③ without AI were shown to the first resident, while AI-forecasted VFs ⑥∼⑦, plus real VFs ①∼③ were shown to the second resident. Similarly, in the bottom half of Figure 6, we also showed both a series of 3 real VFs and a series of AI-forecasted VFs plus 3 real VFs of patient #12, respectively, to the 2 residents to assess the difference in their responses.

Figure 6.

Figure 6

An example of how 2 residents responded after seeing patients' VFs with either AI-VF or VF-only. The 2 residents were both at PGY4, and their responses included GP rating and suggested treatment plan. Green: agreed with ground truth; red: differed from ground truth. AI = artificial intelligence; GP = glaucoma progression; PGY = postgraduate year; VF = visual field.

In the first example shown in Figure 6, the first resident suggested changing the plan due to observation of an abnormal pattern progressing from VF ② to VF ③. This disagreed with the ground truth provided by glaucoma faculty who saw all existing VFs of the patient. In contrast, the second resident who saw more AI-anticipated pattern changes after 2 to 5 years and 5 to 10 years as a reference, suggested the same treatment plan as the ground truth. For the second example, the first resident was given AI-forecasted VFs plus an original set of real VFs, from which he made the same decision as the faculty's ground truth.

These 2 examples showcase an ideal scenario when AI's presence positively impacts residents' decision-making. It is especially suited for the case when a patient has a limited number of VFs; the AI serves as a “future insight provider” that helps clinicians see more potential for longitudinal shifts. In other scenarios, when a patient has ≥5 VFs available, AI can still act as a solid reference to enhance confidence and confirm diagnosis. Overall, we designed our AI as a knowledgeable "consultant" to assist in treatment planning.

Overall Confidence, Decision, and GP Ratings

We also compared the overall GP rating decisions from residents with or without AI-forecasted VFs present. The decisions included GP treatment (1 = monitor, 2 = treat), GP rating (on a scale of 1–5), and confidence rating (on a scale of 1–5). These 3 ratings are intended to showcase the AI's effectiveness in influencing decision-making and diagnosis.

As shown in Figure 7, the GP decision, GP rating, and confidence rating in the bar plot all show an increase in rating with AI-VF condition. However, the difference between AI-VF and VF-only is not statistically significant.

Figure 7.

Figure 7

A, Clinicians' mean decision, GP, and confidence ratings with and without AI assistance; B, mean reaction time with and without AI forecasts. AI = artificial intelligence; GP = glaucoma progression; VF = visual field.

The overall results indicate that patients are rated as more likely to have GP and more likely to be treated with AI.

Also, residents express a higher average confidence by 0.17 when AI-generated VFs are provided. This is consistent with expectation, because AI's presence adds an extra 2 images of information to help confirm an anticipated diagnosis after seeing 3 actual VF images.

Reaction Time

The reaction time, defined as the time (in seconds) it took for the residents to respond per choice, was analyzed to determine the impact of AI forecasts on decision-making speed. The choices to make included “Decision” (monitor vs. treat), “GP Rating” (on a scale of 1–5), and “Confidence Rating” (on a scale of 1-5).

Based on Figure 7B, clinicians spent approximately an extra 2.4 seconds deciding the treatment plan and an extra 2.7 seconds rating GP per slide with AI-VF. As the P values were 0.38 and 0.12, respectively, there was no statistically significant difference in reaction time between AI-VF and VF-only conditions. While AI added 2 more forecasted VFs for clinicians to read, clinicians did not spend much more time compared with the VF-only cases. This confirms that our strategy did not add an extra screening burden for analyzing AI information.

Discussion

Multimodal Implementation for State-of-the-Art Performance and Post Hoc Explainability

We leveraged RNFL input containing regional information as a guiding factor map to enhance the accuracy and inherent interpretability of VF appearance forecasts. Given the 1-dimensional nature of the RNFL data, we also attempted to incorporate both an RNFL (text) encoder and the VF (image) encoder into a contrastive pretraining process using the CLIP model. However, our GenVF (incorporating Garway-Heath zone patterns into training) exhibited the best performance and hence was picked for the downstream user study. Unlike convolutional neural networks or transformer approaches that rely on class activation maps or attention roll-out for post hoc interpretability, we enable inherent interpretability by capturing Garway-Heath zone patterns into our generated 9 × 9 VF images. Our model exhibits comparable or better performance than existing state-of-the-art methods, notably outperforming a previous convolution-based approach (Wen et al13) by extending its predictive power by 5 years into the future without requiring additional demographic information. Nonetheless, our model has potential for improvement: future work leveraging other vision-language architectures, such as LLaVA (Large Language and Vision Assistant),28 may enhance the reliability of AI-generated VFs. By building a state-of-the-art and post hoc-explainable AI model, we strive to streamline and expedite glaucoma treatment for diverse GP patient populations.

User Study

MAE Assessment

In our user study, the lack of statistical significance for MAE in GP rating between AI-VF and VF-only suggests that the AI, in its current form, does not enhance diagnostic precision. This opens an opportunity to rethink the GP rating process to achieve better outcomes and other quantitative metrics to reflect those outcomes. The MAE may oversimplify the complexity of GP ratings. As physicians themselves had discrepancies on the consensus for the ground truth, this highlights the nuanced nature of GP. Alternative quantitative methods, such as uncertainty quantification, ordinal regression, or composite metrics may better address the multifaceted approach required. In particular, interrater agreement metrics such as Cohen or Fleiss Kappa could provide a complementary perspective on whether AI-VF assists clinicians in making more consistent decisions. Additionally, instead of relying solely on MAE to evaluate clinician performance with and without AI, employing a relative ranking order approach29 to assess relative severity changes in diagnosis with AI versus without AI may offer a more nuanced assessment of clinicians' performance.

Expertise Level

Stratifying the residents into expertise levels allowed for an analysis of the impact of clinical experience and performance in this study. The observed uniformity in MAE across the PGY levels in the AI-VF condition, supported by the insignificant P values between the AI-VF values, indicates that AI may contribute to standardizing performance across varying levels of clinical experience. The VF-only condition displayed variability in MAE along the PGY levels with a statistically significant difference between the performance of PGY1 and PGY2 participants. The disparities between the junior and senior residents during the VF-only condition underscore the potential of AI as an educational tool, especially for inexperienced clinicians.

Patient Variability

Patient-specific effects have an impact on performance of AI conditions, as the AI-VF performance fluctuated based on individual case characteristics and VF MD. The heterogeneous impact emphasizes the need for a nuanced approach: tailoring the AI for patient-specific contexts in the clinical workflow. Future work will aim to design AI to build stage-specific (mild/moderate/severe) models.

Tentative Influence of AI on Resident Decisions

Another critical finding is that AI forecasts may influence residents toward more conservative, treatment-oriented decisions. The mean decision values demonstrate a noticeable skew toward the higher GP decision values for the AI-VF condition. For 13 of the 19 patients, the mean GP decision for the AI-VF condition was higher (closer to “treat” than “monitor”) than that for VF-only condition (Fig 5). In 9 of these cases, the AI-VF forecasts were not aligned with the ground truth, indicating that the AI-VF's influence does not necessarily improve decision accuracy. Although not significant, residents showed a tendency toward increased confidence when they had access to the AI as a guide. This trend suggests that AI may serve as a reference or nudge rather than a definitive guide, highlighting the importance of incorporating AI predictions cautiously. This could be addressed in future developments through the use of AI education programs for all residents that allow them to understand the process of incorporating AI outputs into their knowledge of the clinical context. Education programs that promote skepticism and a critical evaluation of the AI input can mitigate overreliance, encouraging clinicians to be more cautious when using AI forecasts as a reference.

Reaction Time

Analysis of reaction time reveals that AI integration does not significantly alter the decision-making time. The marginal increase in reaction time indicates clinicians can incorporate AI-generated VFs into their clinical decision-making without incurring additional cognitive burden. Future AI systems could further optimize the time spent in decision-making by offering real-time explanations for support. Interactive interfaces can allow clinicians to receive the support necessary.

Future Directions

Future work should address challenges in the variability of AI effectiveness across expertise levels and patient-specific contexts. Future AI implementations for GP necessitate improving the quality and quantity of data. Furthermore, potential input modalities other than RNFL thickness should be explored, allowing AI to learn meaningful context from typically unused or underestimated indicators in the clinic. Trustworthiness and confidence when using AI is always critical in AI in health care. Increasing users' awareness of AI uncertainty or limitations could mitigate this effect. As a future direction, including AI's forecast error or a heatmap visualization may help residents to make better-informed decisions and achieve a more convincing confidence score. Additional efforts should be directed to the development of such aids to ensure AI systems are trustworthy before being integrated into clinical diagnostic workflows.

Acknowledgments

The authors thank George A. Cioffi and Jeffrey M. Liebmann for their guidance.

Manuscript no. XOPS-D-25-00251.

Footnotes

Supplemental material available at www.ophthalmologyscience.org.

Disclosure(s):

All authors have completed and submitted the ICMJE disclosures form.

The authors have no proprietary or commercial interest in any materials discussed in this article.

This study was funded by the Columbia University's Data Science Institute Seed Fund and an unrestricted grant from Research to Prevent Blindness, New York, NY, USA.

The Article Publishing Charge for Open Access was paid by Columbia University.

HUMAN SUBJECTS: Human subjects were included in this study. This study, Protocol AAAU5610, was approved by the Columbia University Irving Medical Center Institutional Review Board onMarch 30, 2024 and is in accordance with the tenets set forth by the Declaration of Helsinki. Informed consent was obtained from all study participants.

No animal subjects were used in this study.

Author Contributions:

Conception and design: Tian, Zang, Valenzuela, De Moraes, Tsamis, Thakoor

Data collection: Tian, Gu, Leshno, Valenzuela, De Moraes, Tsamis, Thakoor

Analysis and interpretation: Tian, Bhavsar, Zang, Cui, Thakoor

Obtained funding: Thakoor

Overall responsibility: Tian, Bhavsar, Valenzuela, De Moraes, Tsamis, Thakoor

Supplementary Data

Tables S1 and S2 and Figures S1–S3
mmc1.pdf (458.6KB, pdf)

References

  • 1.Allison K., Patel D., Alabi O. Epidemiology of glaucoma: the past, present, and predictions for the future. Cureus. 2020;12:e11686. doi: 10.7759/cureus.11686. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Hood D.C. Improving our understanding, and detection, of glaucomatous damage: an approach based upon optical coherence tomography (OCT) Prog Retin Eye Res. 2017;57:46–75. doi: 10.1016/j.preteyeres.2016.12.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Yaqub M. Visual fields interpretation in glaucoma: a focus on static automated perimetry. Community Eye Health. 2013;25:1. [PMC free article] [PubMed] [Google Scholar]
  • 4.Tatham A.J., Medeiros F.A. Detecting structural progression in glaucoma with optical coherence tomography. Ophthalmology. 2017;124:S57–S65. doi: 10.1016/j.ophtha.2017.07.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Nouri-Mahdavi K., Hoffman D., Gaasterland D., Caprioli J. Prediction of visual field progression in glaucoma. Invest Ophthalmol Vis Sci. 2004;45:4346–4351. doi: 10.1167/iovs.04-0204. [DOI] [PubMed] [Google Scholar]
  • 6.Kastner A., King A.J. Advanced glaucoma at diagnosis: current perspectives. Eye. 2020;34:116–128. doi: 10.1038/s41433-019-0637-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Broadway D.C. Visual field testing for glaucoma–a practical guide. Community Eye Health. 2012;25:66. [PMC free article] [PubMed] [Google Scholar]
  • 8.Kumar R., Waisberg E., Ong J., et al. Artificial intelligence-based methodologies for early diagnostic precision and personalized therapeutic strategies in neuro-ophthalmic and neurodegenerative pathologies. Brain Sci. 2024;14:1266. doi: 10.3390/brainsci14121266. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Zhu Y., Salowe R., Chow C., et al. Advancing glaucoma care: integrating artificial intelligence in diagnosis, management, and progression detection. Bioengineering. 2024;11:122. doi: 10.3390/bioengineering11020122. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Djulbegovic M.B., Bair H., Gonzalez D.J.T., et al. Artificial intelligence for optical coherence tomography in glaucoma. Transl Vis Sci Technol. 2025;14:27. doi: 10.1167/tvst.14.1.27. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Zhou Z., Li B., Su J., et al. An artificial intelligence model for the simulation of visual effects in patients with visual field defects. Ann translational Med. 2020;8:703. doi: 10.21037/atm.2020.02.162. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Park K., Kim J., Lee J. Visual field prediction using recurrent neural network. Sci Rep. 2019;9:8385. doi: 10.1038/s41598-019-44852-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Wen J.C., Lee C.S., Keane P.A., et al. Forecasting future Humphrey visual fields using deep learning. PLoS One. 2019;14 doi: 10.1371/journal.pone.0214875. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Pang Y., Lin J., Qin T., Chen Z. Image-to-image translation: methods and applications. IEEE Trans Multimedia. 2021;24:3859–3881. [Google Scholar]
  • 15.Tian Y., Zang M., Sharma A., et al. International Workshop on Ophthalmic Medical Image Analysis. Springer Nature Switzerland; Cham: 2023. Glaucoma progression detection and Humphrey Visual field prediction using discriminative and generative vision transformers; pp. 62–71. [Google Scholar]
  • 16.Asaoka R., Murata H. Predicion of visual field progression in glaucoma: existing methods and artificial intelligence. Jpn J Ophthalmol. 2023;67:546–559. doi: 10.1007/s10384-023-01009-3. [DOI] [PubMed] [Google Scholar]
  • 17.Ishikawa H., Abbasi A., Gowrisankaran S., et al. How far in the future can a deep learning model forecast pointwise Visual Field (VF) data based solely on one VF data input. Invest Ophthalmol Vis Sci. 2024;65:373. [Google Scholar]
  • 18.Abbasi A., Gowrisankaran S., Li W.C., et al. A hybrid deep learning based approach for visual field test forecasting. Ophthalmol Sci. 2025;5 doi: 10.1016/j.xops.2025.100803. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Nishida T., Moghimi S., Liebmann J.M., et al. Comparison of methods for visual field progression in eyes with central VF defects. Invest Ophthalmol Vis Sci. 2024;65:4797. doi: 10.1167/tvst.14.11.3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Peirce J., Hirst R., MacAskill M. Sage; London: 2022. Building experiments in PsychoPy. [Google Scholar]
  • 21.Montesano G., Chen A., Lu R., et al. UWHVF: a real-world, open source dataset of perimetry tests from the Humphrey Field Analyzer at the University of Washington. Transl Vis Sci Techn. 2022;11:2. doi: 10.1167/tvst.11.1.1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Garway-Heath D.F., Hitchings R.A. Quantitative evaluation of the optic nerve head in early glaucoma. Br J Ophthalmol. 1998;82:352–361. doi: 10.1136/bjo.82.4.352. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Müller-Franzes G., Niehues J.M., Khader F., et al. A multimodal comparison of latent denoising diffusion probabilistic models and generative adversarial networks for medical image synthesis. Sci Rep. 2023;13 doi: 10.1038/s41598-023-39278-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Goodfellow I., Pouget-Abadie J., Mirza M., et al. Generative adversarial networks. Commun ACM. 2020;63:139–144. [Google Scholar]
  • 25.Ronneberger O., Fischer P., Brox T. International Conference on Medical image computing and computer-assisted intervention. Springer International Publishing; Munich, Germany: 2015. U-net: convolutional networks for biomedical image segmentation. [Google Scholar]
  • 26.Radford A., Kim J.W., Hallacy C., et al. International Conference On Machine Learning. PMLR; 2021. Learning transferable visual models from natural language supervision; pp. 8748–8763. [Google Scholar]
  • 27.Parish D.H., Sperling G. Object spatial frequencies, retinal spatial frequencies, noise, and the efficiency of letter discrimination. Vis Res. 1991;31:1399–1415. doi: 10.1016/0042-6989(91)90060-i. [DOI] [PubMed] [Google Scholar]
  • 28.Liu H., Li C., Wu Q., Lee Y.J. Visual instruction tuning. Adv Neural Inf Process Syst. 2023;36:34892–34916. [Google Scholar]
  • 29.Jayashree R., Christy A. Enhanced user-driven ranking system with splay tree. TELKOMNIKA (Telecommunication Computing Electronics and Control) 2018;16:432–444. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Tables S1 and S2 and Figures S1–S3
mmc1.pdf (458.6KB, pdf)

Articles from Ophthalmology Science are provided here courtesy of Elsevier

RESOURCES