Skip to main content
Clinical Ophthalmology (Auckland, N.Z.) logoLink to Clinical Ophthalmology (Auckland, N.Z.)
. 2026 Aug 13;20:621514. doi: 10.2147/OPTH.S621514

AI-Assisted Detection of Macular OCT Abnormalities by Optometrists: A Retrospective Reader Study

Adrian Hock Chuan Koh 1,✉
PMCID: PMC13480387  PMID: 42610058

Abstract

Background

Optical coherence tomography (OCT) is a key imaging modality for diagnosing retinal disease, but accurate interpretation may require specialist expertise. Artificial intelligence (AI)-assisted OCT has the potential to enhance detection of retinal abnormalities in clinical practice. However, evidence on how AI decision-support tools influence optometrist diagnostic performance in real-world OCT interpretation remains limited.

Objective

This study aimed to evaluate the impact of an AI-integrated OCT interpretation tool, CIRRUS® PathFinder™ (PF), on optometrists’ ability to detect retinal abnormalities on macular OCT scans compared with an expert ophthalmologist reference standard.

Methods

This retrospective diagnostic accuracy reader study analyzed 200 anonymized macular OCT B-scans from 200 patients at a retinal clinic in Singapore. All scans met pre-specified signal strength criteria. Two optometrists independently classified scans as normal or abnormal under unaided and PF-assisted conditions, and classifications were then compared with an expert ophthalmologist reference standard. Diagnostic performance metrics, including sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and accuracy, were calculated with 95% confidence intervals (CI), and agreement with the expert grader was assessed using Cohen’s kappa (κ). Changes in paired classifications were evaluated using McNemar’s exact test.

Results

Of the 200 scans, 73 (36.5%) were classified as abnormal by the expert. PF assistance increased abnormality detection and substantially reduced false negative classifications. Sensitivity improved from 61.6% (95% CI 49.5–72.8) to 98.6% (95% CI 92.6–100.0) for Optometrist A (p<2.0x10−8) and from 87.7% (95% CI 77.9–94.2) to 95.9% (95% CI 88.5–99.1) for Optometrist B (p=0.031). Agreement with expert grading increased from moderate to substantial for Optometrist A (κ=0.528 to 0.729) and remained almost perfect for Optometrist B (κ=0.817 to 0.842). Overall diagnostic accuracy improved from 79.0% to 86.5% and from 91.5% to 92.5%, respectively, with modest reductions in specificity.

Conclusion

AI-assisted OCT interpretation with PF improved optometrists’ sensitivity for detecting expert-confirmed macular OCT abnormalities, primarily by reducing false-negative classifications. Further prospective studies in broader clinical settings are needed to determine the impact on referral appropriateness, workflow efficiency, reader behavior and patient outcomes.

Keywords: clinical decision support, diagnostic accuracy, retinal imaging, sensitivity and specificity, deep learning

Introduction

Optical coherence tomography (OCT) enables non-invasive, high-resolution, cross-sectional visualization of retinal microstructures, and is a standard tool used in the diagnosis of major ophthalmic diseases. Despite the advancements in OCT technology, non-specialist clinicians still struggle to interpret complex OCT scans consistently. Combining artificial intelligence (AI) with OCT technology has real potential to improve the detection of abnormalities and help clinicians deliver better care.1 Retinal disease management is particularly well suited to AI integration as OCT generates large volumes of high-resolution images, which provides rich datasets for the training of AI algorithms.2

As AI-assisted OCT interpretation becomes more widely used, there has been growing interest in comparing its performance against expert ophthalmologists to evaluate diagnostic accuracy, reproducibility, and suitability for screening or triage. AI-assisted OCT has demonstrated its utility as an accurate, sensitive and specific detection tool for retinal disorders such as pigment epithelial detachment (PED), posterior vitreous detachment (PVD), epiretinal membranes (ERMs), subretinal fluid (SRF), choroidal neovascularization (CNV), drusen, cystoid macular edema (CME), exudation, macular hole (MH), retinal detachment (RD), choroid atrophy, and retinal hemorrhage.3 The DeepMind Health and Moorfield’s Eye Hospital NHS Foundation Trust study reported in 2018 that AI-assisted OCT was able to recommend correct referral decision, with up to 94% accuracy, on a test set of 997 OCT scans, therefore matching the performance of expert clinicians.4 More recently, the PAIR study by Fong et al concluded that an AI-integrated OCT-interpretation tool was valuable in terms of real-time decision-support in resource-limited settings; however, they noted that disease-specific refinements and clinical oversight remain important, particularly for vision-threatening conditions.5 In Singapore, optometrists may be the first healthcare professionals to review patients with ocular signs of disease and can potentially play a key role in detecting new eye diseases.6 As such, gaps in the accuracy of OCT interpretation by non-specialists may not only carry clinical consequences but may also result in delay in referral to specialist care.

This retrospective diagnostic accuracy reader study evaluated 200 OCT brightness scans (B-scans) acquired during routine clinical care. The objective was to determine if CIRRUS® PathFinderTM (PF) assistance was able to improve optometrists’ detection of expert-confirmed macular OCT abnormalities compared with unaided interpretation. The paired-reader design allowed assessment of the incremental effect of PF assistance within the same readers. We hypothesized that PF assistance would improve the sensitivity, diagnostic agreement and simulated referral decisions compared with unaided assessment.

Materials and Methods

Study Dataset

This was a cross-sectional diagnostic accuracy analysis of 200 anonymized OCT images obtained from 200 patients who visited Eye and Retina Surgeons (ERS), Camden Medical, Singapore, between January 2025 and June 2025 as part of routine clinical care at ERS. The date range of January to June 2025 reflects the period during which the clinical images were originally acquired in routine practice, rather than a period of prospective enrolment for research. No images were acquired specifically for research purposes, and the study was conceived as a research evaluation only after the clinical images had been acquired. Patients had provided written informed consent for the use of their anonymized data for scientific purposes. An ethics exception application was submitted to Parkway Independent Ethics Committee (PIEC) after the study was conceived, and approval and exemption was granted by PIEC on 9 June 2025 (reference number PIEC/2024/051). Following the granting of ethics exemption by PIEC, ERS retrospectively retrieved the dataset corresponding to this period for research use. No data was accessed, extracted, or analyzed for research purposes prior to PIEC approval for exemption. The optometrist grading exercises (both unaided and PF-assisted) – which constitute the research-specific activities in this study – were conducted after ethics exemption was granted the PIEC.

The mean age of participants was 58.5 years (± 13.9 years), with a near-equal distribution of sexes (104 male [52.0%], 96 females [48.0%]), and right (Oculus Dexter [OD]) and left (Oculus Sinister [OS]) eyes (OD 103 [51.5%], OS 97 [48.5%]). All scans met predefined image-quality criteria, with a mean Signal Strength (SS) of 8.9 (± 1.0) (Table 1).

Table 1.

Baseline Participant Profile and OCT Scan Characteristics

Variable Overall (n=200)
Age (years) 58.5 ± 13.9
Gender Male 104 (52.0%)
Female 96 (48.0%)
Eye assessed Oculus Dexter (OD) 103 (51.5%)
Oculus Sinister (OS) 97 (48.5%)
Signal strength (0–10) 8.9 ± 1.0

Abbreviations: OD, Oculus Dexter; OS, Oculus Sinister.

PIEC is an independent ethics committee that provides oversight for research conducted at ERS, Camden Medical, Singapore, where the study data were generated. Neither the author nor ERS is affiliated with PIEC. This study was conducted in accordance with the tenets of the Declaration of Helsinki.

Evaluator Information

Three independent readers participated in the study. Optometrist A (Optom A) had 10 years of clinical experience but relatively limited exposure to retinal pathology, while Optometrist B (Optom B) had 6 years of clinical experience with greater specialization in retinal cases and retinal imaging. Both optometrists were based at ERS. Before PF-assisted grading, both optometrists received standardized instruction on the interpretation of PF outputs and the grading workflow. No study images were used for training or calibration.

The reference standard was established by a senior consultant ophthalmologist based at ERS specializing in retinal diseases and OCT interpretation, with approximately 30 years of clinical experience.

OCT Image Acquisition and Statistical Analyses

OCT scans were acquired with an OCT (Cirrus 6000, Carl Zeiss Meditec Inc, USA) equipped with the PF AI Decision Support feature, which uses deep learning algorithms to automatically identify abnormal macular OCT B-scans that may require additional review.7 Only one OCT scan from each patient was included, avoiding within-patient inter-eye correlation in the primary analysis. The primary grading unit was the central foveal B-scan; however, readers could review adjacent B-scans within the same macular cube when needed to clarify structural abnormalities. The final classification was recorded at the scan level as normal or abnormal.

The abnormalities included for assessment in this study were disruption to the inner retinal layer, disruption to vitreoretinal interface (VRI) disruption, inner segment/outer segment (IS/OS) junction disruption, retinal pigment epithelium (RPE) atrophy, RPE elevation, intraretinal and subretinal fluid, and other structural retinal abnormalities (see Supplementary Table 1).

Image Grading and Masking Procedures

All 200 OCT scans were independently reviewed by the two optometrists and the expert ophthalmologist. Each reader classified scans as either normal or abnormal based on structural retinal findings. The optometrists evaluated the OCT scans under two conditions: Unaided interpretation, without AI support; PF-assisted interpretation, with PF decision-support tool activated.

To minimize interpretation bias, masking procedures were implemented throughout the evaluation process. The expert ophthalmologist who established the reference standard was masked to the classifications made by Optom A and Optom B, as well as to their PF-assisted interpretations. Similarly, when performing unaided OCT interpretation, the optometrists were masked to the results obtained during PF-assisted assessment, and vice versa. Each grading session was conducted independently to ensure that interpretations were not influenced by prior assessments or AI outputs. A minimum washout period of four weeks was implemented between the unaided and PF-assisted grading sessions. The unaided grading sessions were completed in month 1, and PF-assisted sessions commenced in month 2, with a minimum of 28 days between the individual readers’ sessions. To further minimize recall bias, images were presented in a different randomized order at each session.

Reference Standard

The reference standard was established by a single senior retinal specialist. While this approach provided consistency across all scans, it did not exclude reference standard misclassification. Therefore, improved agreement should be interpreted as improved concordance with this expert reference standard, rather than definitive proof of diagnostic correctness.

Statistical Analysis

Statistical analyses were performed on retrospectively retrieved clinical data using the expert ophthalmologist’s grading as the reference standard. Continuous variables were summarized via means and standard deviations, and categorial variables were summarized as counts and percentages. Diagnostic performance metrics (sensitivity, specificity, positive predictive value [PPV], negative predictive value [NPV], and accuracy) were calculated with corresponding 95% confidence intervals (CI) for each optometrist under PF-assisted and unassisted conditions. The sample size of 200 scans was determined based on a pre-specified power calculation for the primary endpoint (change in sensitivity with PF assistance). Changes in sensitivity and specificity were evaluated using McNemar’s exact test (two-sided), applied separately for each optometrist, and any statistical significance was defined as a two-sided p-value <0.05. Using a two-sided McNemar’s test with an expected baseline sensitivity of approximately 75%, a clinically meaningful improvement of 15% points (to 90%), a significance level of 0.05, and 80% power, a minimum of 170 paired observations were required. The final sample of 200 was selected to provide a margin above this minimum and to capture sufficient abnormal scans for subtype analysis (at an estimated abnormality prevalence of 35%, n=70 abnormal scans). Agreement between optometrist and expert classifications was assessed using Cohen’s kappa (κ), with standard interpretive thresholds (0–0.20: slight/poor agreement; 0.21–0.40: fair; 0.41–0.60: moderate; 0.61–0.80: substantial; and 0.81–1.0: almost perfect).

Study Endpoints

The primary endpoint was the change in sensitivity for detecting expert-confirmed abnormal OCT scans with PF assistance compared with unaided interpretation. Secondary endpoints were specificity, PPV, NPV, overall accuracy, Cohen’s κ agreement with the expert reference standard, false-positive and false-negative classifications, abnormality subtype detection, and simulated referral recommendations.

After classifying each scan as abnormal or abnormal, readers independently recorded whether specialist referral would be recommended based on OCT findings alone. Referral was treated as a distinct simulated clinical decision and did not automatically follow an abnormal classification.

Results

Primary Endpoints

Percentage of OCT Scans Flagged as Abnormal

Of the 200 OCT B-scans analyzed, 73 (36.5%) were classified by the expert ophthalmologist as abnormal. Without PF assistance, Optom A classified 29.5% of scans as abnormal, while Optom B classified 36.0% as abnormal. With PF assistance, abnormality classification increased to 49.0% for Optom A, and 41.0% for Optom B, indicating a reduction in under-detection and improved alignment with the expert’s classification (Table 2 and Figure 1). PF-assisted readings demonstrated closer alignment with expert classifications across key retinal biomarkers, including VRI disruption, IS/OS disruption, RPE elevation or atrophy, and intraretinal or subretinal fluid. Full subtype distributions are presented in Supplementary Table 1.

Table 2.

Distribution of OCT Scan Results by Optometrist (with and without PF) and Expert

Result Optom A Optom B Expert
Without PF With PF Without PF With PF
Normal 141 (70.5%) 102 (51.0%) 128 (64.0%) 118 (59.0%) 127 (63.5%)
Abnormal 59 (29.5%) 98 (49.0%) 72 (36.0%) 82 (41.0%) 73 (36.5%)
Abnormality Type
Disruption Inner Retinal 2 (1.0%) 2 (1.0%) 11 (5.5%) 16 (8.0%) 5 (2.5%)
VRI Disruption 25 (12.5%) 26 (13.0%) 24 (12.0%) 23 (11.5%) 25 (12.5%)
IS/OS Disruption 3 (1.5%) 4 (2.0%) 14 (7.0%) 30 (15.0%) 12 (6.0%)
RPE Atrophy 2 (1.0%) 1 (0.5%) 16 (8.0%) 32 (16.0%) 3 (1.5%)
RPE Elevation 23 (11.5%) 38 (19.0%) 29 (14.5%) 34 (17.0%) 34 (17.0%)
Intraretinal Fluid 5 (2.5%) 5 (2.5%) 3 (1.5%) 4 (2.0%) 7 (3.5%)
Subretinal Fluid 14 (7.0%) 15 (7.5%) 5 (2.5%) 7 (3.5%) 8 (4.0%)
Others 14 (7.0%) 13 (6.5%) 7 (3.5%) 8 (4.0%) 3 (1.5%)

Abbreviations: IS/OS, inner segment/outer segment; Optom, optometrist; PF, PathFinder; RPE, retinal pigment epithelium; VRI, vitreoretinal interface.

Figure 1 .

A stacked bar graph showing normal and abnormal classifications across optometrist readings and expert.

Proportion of normal and abnormal OCT classifications.

Abbreviations: OCT, optical coherence tomography; Optom, optometrist; PF, PathFinder.

Sensitivity Analysis – Performance in Expert-Confirmed Abnormal Scans

PF assistance reduced false-negative classifications from 28 to 1 for Optom A and from 9 to 3 for Optom B, corresponding to correct identification of 72/73 and 70/73 expert-confirmed abnormal scans, respectively. The magnitude of improvement was greater for Optom A, suggesting that the benefit of AI assistance may depend on baseline retinal imaging experience. Both optometrists detected nearly all expert-identified abnormalities with PF assistance. PF also enhanced the detection of specific pathological subtypes – Optom B’s detection of IS/OS disruption increased from 18% to 37%, and Optom A showed consistent gains across VRI and IS/OS abnormalities – indicating that PF improved recognition of specific pathological subtypes (Table 3).

Table 3.

Optometrist Performance on Expert-Confirmed Abnormal Scans (n=73)

Result Optom A Optom B Expert
Without PF With PF Without PF With PF
Normal 28 (38%) 1 (1%) 9 (12%) 3 (4%)
Abnormal 45 (62%) 72 (99%) 64 (88%) 70 (96%) 73 (100%)
Abnormality Type
Disruption Inner Retinal 2 (3%) 2 (3%) 11 (15%) 15 (21%) 5 (7%)
Disruption VRI 12 (16%) 21 (29%) 23 (32%) 22 (30%) 25 (34%)
IS/OS Disruption 3 (4%) 3 (4%) 13 (18%) 27 (37%) 12 (16%)
RPE Atrophy 2 (3%) 1 (1%) 16 (22%) 30 (41%) 3 (4%)
RPE Elevation 23 (32%) 35 (48%) 24 (33%) 32 (44%) 34 (47%)
Intraretinal Fluid 4 (5%) 4 (5%) 3 (4%) 4 (5%) 7 (10%)
Subretinal Fluid 13 (18%) 13 (18%) 5 (7%) 7 (10%) 8 (11%)
Others 13 (18%) 12 (16%) 5 (7%) 3 (4%) 3 (4%)

Abbreviations: IS/OS, inner segment/outer segment; Optom, optometrist; PF, PathFinder; RPE, retinal pigment epithelium VRI, vitreoretinal interface.

Agreement with Expert Ophthalmologist

Agreement between the optometrists and the expert ophthalmologist improved substantially with PF assistance. Without PF, Optom A demonstrated moderate agreement (κ=0.528) whereas Optom B demonstrated almost perfect agreement (κ=0.817). With PF assistance, agreement increased to 0.729 (substantial agreement) for Optom A, and 0.842 (reinforcing almost perfect agreement) for Optom B. Observed agreement followed the same trend (Optom A: 79% to 86.5%; Optom B: 91.5% to 92.5%), indicating that PF strengthened interpretive consistency, particularly for the less experienced optometrist, Optom A (Table 4).

Table 4.

Diagnostic Performance and Agreement Metrics for Optometrists with and without PF (n=200)

Parameter Optom A Optom B
Without PF With PF Without PF With PF
Sensitivity, % (95% CI) 61.6 (49.5–72.8) 98.6 (92.6–100.0) 87.7 (77.9–94.2) 95.9 (88.5–99.1)
Specificity, % (95% CI) 89.0 (82.2–93.8) 79.5 (71.5–86.2) 93.7 (88.0–97.2) 90.6 (84.1–95.0)
PPV, % (95% CI) 76.3 (63.4–86.4) 73.5 (63.6–81.9) 88.9 (79.3–95.1) 85.4 (75.8–92.2)
NPV, % (95% CI) 80.1 (72.6–86.4) 99.0 (94.7–100.0) 93.0 (87.1–96.7) 97.5 (92.7–99.5)
Accuracy, % 79.0 86.5 91.5 92.5
Cohen’s κ 0.528 0.729 0.817 0.842
Observed Agreement, % 79.0 86.5 91.5 92.5
Δκ* 0.201 0.025

Notes: Δκ* = (Cohen’s κ with PF) − (Cohen’s κ without PF).

Abbreviations: Optom, optometrist; NPV, negative predictive value; PPV, positive predictive value; PF, PathFinder.

Diagnostic Performance Relative to Expert Grading

Overall, PF assistance considerably enhanced the diagnostic performance of both optometrists. While PF substantially improved sensitivity (Figure 2a) for both readers, there was some modest decline in specificity (Figure 2b). For Optom A, sensitivity increased from 61.6% (45/73, 95% CI 49.5–72.8) without PF, to 98.6% (72/73, 95% CI 92.6–100.0) with PF (p<2.0x10−8). For Optom B, sensitivity increased from 87.7% (64/73, 95% CI 77.9–94.2) without PF, to 95.9% (70/73, 95% CI 88.5–99.1) with PF (p=0.031). With regard to specificity, this declined for Optom A from 89.0% (95% CI 82.2–93.8) without PF, to 79.5% (95% CI 71.5–86.2) with PF (p=4.9x10−4). False positives increased from 14 to 26. For Optom B, specificity declined from 93.7% (95% CI 88.0–97.2) without PF, to 90.6% (95% CI 84.1–95.0) with PF (p=0.125). False positives increased from 8 to 12 (Table 4).

Figure 2 .

Two bar charts showing optometrist sensitivity and specificity with and without PF. Image A: Grouped bar chart with y-axis labeled ′Percentage′ (0-120). X-axis: Optom A and B Sensitivity. Bars: ′Without PF′ and ′With PF′ with error bars. Optom A: Without PF 61.6%, With PF 98.6% (p < 2.0 x 10 superscript -8). Optom B: Without PF 87.7%, With PF 95.9% (p = 0.031). Image B: Grouped bar chart with y-axis labeled ′Percentage′ (0-120). X-axis: Optom A and B Specificity. Bars: ′Without PF′ and ′With PF′ with error bars. Optom A: Without PF 89%, With PF 79.5% (p = 4.9 x 10 superscript -4). Optom B: Without PF 93.7%, With PF 90.6% (p = 0.125). Legend: Without PF, With PF.

Sensitivity (a) and specificity (b) for Optoms A and B with and without PF.

Abbreviations: Optom, optometrist; PF, PathFinder.

Secondary Endpoints

Diagnostic Accuracy Compared with Expert

PF assistance led to a redistribution of classification errors. False-negative classifications decreased substantially for both optometrists, particularly for Optom A, while false-positive classifications increased modestly. Without PF support, Optom A correctly identified 45 true-positive and 113 true-negative scans but misclassified 28 abnormal and 14 normal scans. With PF assistance, true-positives increased to 72, false negatives decreased to 1, true-negatives were 101, and false-positives increased slightly to 26, indicating improved performance in detection. Optom B demonstrated similar gains: without PF, Optom B identified 64 true-positives and 119 true-negatives; and misclassified 9 abnormal and 8 normal scans. With PF support, true-positives increased to 70, false-negatives declined to 3, true-negatives were 115 and false-positives increased to 12, demonstrating consistent improvement across readers. These findings confirm that PF substantially enhanced diagnostic accuracy by reducing false-negative errors and increasing correct identification of abnormal scans.

Clinical Outcomes – Referral Status and Final Diagnosis

PF assistance influenced simulated OCT-based referral recommendations. Referral recommendations increased from 29.5% to 49.0% for Optom A, and from 24.0% to 29.5% for Optom B when PF was used, compared with 36.5% by the expert reference grader (Table 5). These findings suggest that PF assistance may influence referral thresholds; however, referral recommendations in this reader study should not be interpreted as actual clinical referrals, as real-world referral decisions would also incorporate symptoms, visual acuity, ocular history, risk factors and clinician judgement.

Table 5.

Referral Recommendations

Parameter Optom A
Without PF
Optom A
With PF
Optom B
Without PF
Optom B
With PF
Expert
Referral Status
No Referral 141 (70.5%) 102 (51.0%) 152 (76.0%) 141 (70.5%) 127 (63.5%)
Refer to Specialist 59 (29.5%) 98 (49.0%) 48 (24.0%) 59 (29.5%) 73 (36.5%)

Abbreviations: Optom, optometrist; PF, PathFinder.

Diagnostic Classification Outcomes (Normal vs Abnormal)

Consistent with referral patterns, PF assistance increased the proportion of scans classified as abnormal by both optometrists. For Optom A, abnormal classifications increased from 29.5% to 49.0% with PF, whereas for Optom B, abnormal classifications increased from 36.0% to 41.0% with PF (Figure 1). These shifts reduced under-classification of abnormal scans relative to the expert abnormality prevalence of 36.5%.

Concordance with the expert’s Final Diagnosis

The expert ophthalmologist’s final diagnosis classified 127 scans (63.5%) as normal and 73 scans (36.5%) as abnormal. PF assistance did not alter the expert reference distribution but improved concordance between the optometrists’ diagnostic classifications and expert determinations, primarily by reducing false-negative classifications and increasing correct identification of expert-confirmed abnormal scans.

Discussion

PF assistance improved optometrists’ sensitivity in detecting expert-confirmed macular OCT abnormalities and reduced false-negative classifications. The magnitude of improvement was greater in the reader who had less exposure to retinal pathology, suggesting that AI assistance may be particularly useful for supporting less specialized OCT readers.

This study differed from the Fong et al study (PAIR) in reader population, design, reference standard and outcomes measured.5 The readers in the PAIR study comprised non-retina specialist ophthalmologists, whereas in this study, the readers were optometrists with distinctly different clinical roles, training backgrounds, scope-of-practice, and access to specialist oversight. The PAIR study assessed diagnostic and referral agreement between non-retinal specialists and retinal specialist gold standards; while this study is a within-reader, paired-design evaluation measuring the differences in diagnostic performance with and without PF assistance in the same readers, using McNemar’s test to directly quantify the incremental value of AI assistance. In the PAIR study, a multi-reader retinal specialist consensus was the gold standard whereas in this study, the reference standard is a single senior retinal specialist with approximately 30 years of experience, a design reflecting many real-world solo-practice settings. In terms of outcomes, this study documents the reduction in false-negative errors and the change in referral behavior, which are particularly relevant to patient safety in primary eye care settings.

Improved Detection Accuracy and Reader Confidence

PF assistance resulted in marked improvements in diagnostic performance across the primary endpoints, particularly for the less experienced optometrist. For both optometrists, sensitivity increased substantially, while overall diagnostic accuracy and agreement with expert grading also improved. The improvement in Cohen’s κ for Optom A from moderate to substantial agreement suggests that PF not only improved detection but also enhanced interpretive consistency in OCT assessment.

These findings are consistent with prior studies demonstrating that AI decision-support tools can reduce inter-reader variability and support non-specialist clinicians in image-based diagnostics, particularly in ophthalmology where OCT interpretation requires specialized training and experience.4,8,9 By flagging structural abnormalities and relevant biomarkers, PF appears to reinforce correct interpretation rather than replace clinical judgement.

Reduction in False Negatives and Improved Diagnostic Safety

A key strength of PF assistance was the reduction in false-negative classifications. Missed retinal pathology represents a critical safety risk, as delayed diagnosis and referral may result in irreversible visual loss, particularly in conditions such as neovascular age-related macular degeneration, diabetic macular edema, and vitreoretinal interface disorders.10–12 In the present study, PF assistance reduced false negatives to near zero for one optometrist and substantially for the other, resulting in exceptionally high negative predictive values.

From a clinical risk perspective, this reduction in missed abnormalities is arguably more important than marginal changes in specificity. A high negative predictive value supports the safe exclusion of pathology and provides reassurance to both clinicians and patients when scans are classified as normal. These findings support PF’s potential as a safety-enhancing tool for non-specialist readers, particularly where access to specialist interpretation may be limited.

Sensitivity-Specificity Trade-Off and Clinical Vigilance

The reduction in scans classified as normal (with PF assistance) for both Optom A and Optom B (Figure 1), may be due to the increase in sensitivity with PF assistance, which comes at the cost of some specificity. This sensitivity-specificity trade-off is more pronounced in Optom A.

The sensitivity-specificity trade-off is well recognised in diagnostic test evaluation and is especially anticipated for AI tools designed to prioritize detection over exclusion.13 In a screening context, increased sensitivity at the expense of some false positives is generally considered acceptable, as the primary objective is to flag potential cases among a larger pool, serving as an initial triage step to flag patients who may require closer evaluation assessment, in order to prioritize care and improve healthcare outcomes.14

In our study, we observed that while specificity declined with PF assistance, most notably for Optom A, where 79.5% specificity corresponds to a 20.5% false positive rate, this meant that approximately 1 in 5 normal scans were flagged as abnormal by Optom A with PF assistance. This pattern suggests increased clinical vigilance and aligns with the intended use of AI-assisted tools as decision support systems rather than definitive diagnostic tools.

Impact on Referral Decisions and Clinical Workflow

The secondary endpoint analyses demonstrated that PF assistance had a positive impact on downstream clinical decisions, particularly referral behavior. Referral rates increased with PF use, reducing under-referral relative to expert judgement, and potentially influencing referral decision-making. The improved alignment between optometrists’ classifications and expert-defined abnormality patterns further suggests that PF supports both detection as well as clinically meaningful interpretation of OCT features.

The referral rate for Optom A (who had 10 years of experience but limited exposure to retinal pathology) rose to 49% with PF, compared with the expert’s 36.5%. This reflects a pattern of over-referral, consistent with the significant improvement in sensitivity (from 61.6% without PF to 98.6% with PF) which was accompanied by a modest decline in specificity (from 89.0% without PF to 79.5% with PF). The higher referral rate may be driven by the increase in false positives (from 14 without PF to 26 with PF). In a screening or primary care context, this trade-off of accepting some over-referral to avoid missed pathology may be clinically appropriate and preferable to under-detection of vision-threatening conditions.

For Optom B (who had 6 years of experience but greater retinal specialization), the referral rate without PF was notably lower (24.0%) than that of the expert (36.5%), indicating under-referral and missed pathology at baseline. With PF assistance, this increased to 29.5%, moving closer to expert behavior. The persistent gap may be due to Optom B’s tendency to rely primarily on OCT structural features and the higher threshold applied for referral, even when abnormalities were flagged with PF assistance. In this study, an abnormal classification does not automatically trigger a referral recommendation. Due to the greater retinal specialization of Optom B, an additional threshold of clinical significance was applied before recommending referral, such that scans which were deemed as mildly abnormal or representing changes deemed as manageable without specialist review were not referred. This explains why Optom B had an abnormality classification rate of 36% and a referral rate of 24% without PF assistance. This is reflective of real-world optometric practice, where clinicians must weigh the clinical severity of a finding against the need for specialist input.

With regard to the pronounced difference in PF-unaided and PF-aided sensitivity for Optom A, the near-complete sensitivity (98.6% with PF vs 61.6% without) does raise a concern about whether the use of an AI-assisted tool augments clinical judgement or substitutes it. Automation complacency is a well-recognized risk with AI-assisted diagnostic tools.15 When an AI tool drives abnormality detection, readers may become uncritical conduits for its outputs rather than independent interpreters.

Both the risk of automation complacency and the divergent referral behavior between readers point to the same conclusion: AI assistance needs to sit within a structured referral framework, not substitute for clinical training in referral decision-making. Longitudinal studies to assess if the use of PF maintains or erodes independent OCT interpretation skills may be an important direction for future research.

Overall, these findings are consistent with emerging evidence that AI-supported OCT analysis can play an important role in enhancing referral accuracy and optimizing care pathways.3,4 By reducing missed expert-confirmed abnormalities, PF may facilitate earlier specialist review in some settings, which may lead to earlier diagnosis and improved outcomes for patients.

Implications for Non-Specialist Use and Scalability

The differential impact of PF on the two optometrists underscores the potential value of an AI-integrated OCT-interpretation tool in supporting non-specialist readers and reducing experience-related performance gaps. This has important implications for scalability, workforce capacity, and equity of care, particularly in areas with limited access to specialists.

Study Limitations

This study has several limitations. First, the reference standard was based on grading by a single expert ophthalmologist rather than a consensus panel. Furthermore, a formal intra-rater reliability assessment was not conducted for the reference standard grader. Although expert adjudication is commonly used in OCT diagnostic studies, reliance on a single grader may introduce the possibility of reference standard misclassification. The grader’s extensive experience (approximately 30 years) and use of established structural criteria likely reduced, but cannot fully exclude, the potential for systematic bias. The use of a single expert grader may also ensure consistent evaluation across all scans. Future studies should incorporate a multi-reader consensus panel or, at minimum, double grade a representative subset to estimate reference standard reliability, to further strengthen the reference standard.

Second, the study included a limited number of readers from a single clinical center, which may limit generalizability. Diagnostic performance improvements associated with PF assistance may vary across clinicians with different levels of OCT experience.

Third, this was a cross-sectional diagnostic accuracy study rather than a prospective clinical workflow evaluation. The retrospective design may introduce selection bias, and the analysis was conducted on a fixed dataset of high-quality OCT B-scans, which may not accurately reflect real-world imaging variability. Future research should examine the real-world impact of AI-assisted OCT interpretation on referral patterns, diagnostic accuracy, and patient outcomes.

Fourth, we acknowledge that residual recall bias cannot be entirely excluded from this study, particularly given that the dataset comprised 200 images, and the possibility that readers may have recalled specific cases despite the washout period.

Conclusion

In this diagnostic accuracy reader study, PF assistance improved optometrists’ sensitivity for detecting expert-confirmed macular OCT abnormalities and reduced false-negative classifications, with a modest reduction in specificity. The magnitude of benefit was greater for the reader with less retinal pathology exposure, suggesting that AI-assisted OCT interpretation may help reduce experience-related variability among non-specialist readers. These findings support integrating AI decision support into routine retinal imaging workflows to improve consistency and safety among non-specialist clinicians and to reduce missed diagnoses. This study should be interpreted in the context of a single-center study, a limited number of readers, and a single expert reference standard. Further prospective studies involving more readers, broader clinical settings, disease-level adjudication, and multi-specialist reference standards are needed to confirm generalizability and determine the impact on referral appropriateness and patient outcomes.

Acknowledgments

I would like to thank the following: Ms. Berlisa Chong Yong Qin (Senior Optometrist, Eye & Retina Surgeons, Camden Medical, Singapore) and Mr. Calvin Foo Wei Ming (Optometrist, Eye & Retina Surgeons, Camden Medical, Singapore), for actively participating in the OCT analysis exercise; Ms. Kalin Siow, Study Manager and Coordinator, Eye & Retina Surgeons, Camden Medical, Singapore, for coordination of the study and provision of all the necessary administrative support; Mr. Terrance Siew, Senior Manager, Global CDM Post Approval Studies of Carl Zeiss, for provision of administrative and technical support related to the investigator-initiated study program; and Ms. Su Ping Chuah for medical writing support.

Funding Statement

This study was funded by Carl Zeiss Meditec, Inc. through an investigator-initiated study program. The funder was not involved in the study design, data collection, analysis or interpretation of data, and did not influence the results of the study.

Disclosure

Dr Adrian Koh reports consulting fees from Heidelberg, Novartis, Apellis, Astellas, Bayer, Roche, Carl Zeiss Meditec; honoraria from Carl Zeiss, Roche, Bayer; leadership or fiduciary roles for Retina Society of Singapore and Asia Pacific Vitreoretina Society, outside the submitted work.

References

  • 1.Alikarami M, Faraj TA, Hama NH, et al. Artificial intelligence in advancing optical coherence tomography for disease detection and cancer diagnosis: a scoping review. Eur J Surg Oncol. 2025;51(9):110188. [DOI] [PubMed] [Google Scholar]
  • 2.Oshika T. Artificial intelligence applications in ophthalmology. JMA J. 2025;8(1):66–12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Bai J, Wan Z, Li P, et al. Accuracy and feasibility with AI-assisted OCT in retinal disorder community screening. Front Cell Dev Biol. 2022;10:1053483. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.De Fauw J, Ledsam JR, Romera-Paredes B, et al. Clinically applicable deep learning for diagnosis and referral in retinal disease. Nat Med. 2018;24(9):1342–1350. doi: 10.1038/s41591-018-0107-6 [DOI] [PubMed] [Google Scholar]
  • 5.Fong KCS, Wong WJ, Samsudin A, et al. PAIR: evaluating the limits of agreement among non-retinal specialist using PathFinder artificial intelligence tool for retinal disease referrals: a prospective observational study. Clin Ophthalmol. 2026;20:584717. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Yeo LYX, Tan CYM, Allen JW, et al. Innovative care models: expanding nurses’ and optometrists’ roles in ophthalmology. Nurs Ethics. 2025;32(6):1900–1910. doi: 10.1177/09697330251317670 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.CIRRUS PathFinder. Available from: www.zeiss.com/meditec/en/products/optical-coherence-tomography-devices/cirrus-6000-performance-oct/cirrus-pathfinder.html. Accessed January 12, 2026.
  • 8.Ting DSW, Pasquale LR, Peng L, et al. Artificial intelligence and deep learning in ophthalmology. Br J Ophthalmol. 2019;103:167–175. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Schmidt-Erfuth U, Sadeghipour A, Gerendas BS, et al. Artificial intelligence in retina. Prog Retin Eye Res. 2018;67:1–29. [DOI] [PubMed] [Google Scholar]
  • 10.Vangipuram G, Li C, Li S, et al. Timing of delayed retinal pathology in patients presenting with acute posterior vitreous detachment in the IRIS® registry (Intelligent research in sight). Ophthalmol Retina. 2023;7(8):713–720. [DOI] [PubMed] [Google Scholar]
  • 11.Zhou C, Li S, Ye L, et al. Visual impairment and blindness caused by retinal diseases: a nationwide register-based study. J Glob Health. 2023;13:04126. doi: 10.7189/jogh.13.04126 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Almazroa A, Almatar H, Alduhayan R, et al. The patients’ perspective for the impact of late detection of ocular diseases on quality of life: a cross-sectional study. Clin Optom. 2023;15:191–204. doi: 10.2147/OPTO.S422451 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Wang Y, Liu C, Hu W, et al. Economic evaluation for medical artificial intelligence: accuracy vs. cost-effectiveness in a diabetic retinopathy screening case. NPJ Digit Med. 2024;7:43. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.AlShawabkeh M, AlRyalat SA, Al Bdour M, et al. The utilization of artificial intelligence in glaucoma: diagnosis versus screening. Front Ophthalmol. 2024;4:1368081. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Saadeh MI, Janhonen J, Beer E, et al. Automation complacency: risks of abdicating medical decision making. AI Ethics. 2025;5:5783–5793. [Google Scholar]

Articles from Clinical Ophthalmology (Auckland, N.Z.) are provided here courtesy of Dove Press

RESOURCES