Abstract
Introduction
Screening diabetic retinopathy (DR) for timely management can reduce global blindness. Many existing DR screening programs worldwide are non-digital, standalone, and deployed with grading retinal photographs by trained personnel. To integrate the screening programs, with or without artificial intelligence (AI), into hospital information systems to improve their effectiveness, the non-digital workflow must be transformed into digital. We developed a cloud-based digital platform and implemented it in an existing DR screening program.
Methods
We conducted the following processes in the platform for prospective DR screening at a community hospital: capturing patients’ retinal photographs, uploading them for grading by AI or trained personnel on alternate weeks for 32 weeks, and referring vision-threatening DR to a referral center. At this center, the platform was applied for the assessment of potential missed referrals via remote over-reading by a retinal specialist and tracking referrals. Implementational outcomes, such as detecting positive cases, agreement between AI and over-reading, and referral adherence were assessed.
Results
Of 645 patients screened by AI, 201 (31.2%) were referrals, 129 (64.2%) of which were true positives agreeable by over-reading; 115 of these true positives (89.1%) had referral adherence. False negatives judged by over-reading were 1.1% (5/444). Of 730 patients in manual screening, 175 (24.0%) were potential referrals, 11 (6.3%) of which were referred at the point-of-screening; eight of these (72.7%) adhered to referral. The remaining 164 cases were appointed for later examination by a visiting general ophthalmologist; 11 of these 116 examined (9.5%) were referred for non-DR-related eye conditions with 81.8% (9/11) referral adherence. No system failure or interruption was found.
Conclusions
The digital platform can be practically integrated into the existing non-digital DR screening programs to implement AI and monitor previously unknown but important indicators, such as referral adherence, to improve the effectiveness of the programs.
Trial Registration
ClinicalTrials.gov. (registration number: NCT05166122).
Keywords: Diabetic retinopathy, Vision-threatening diabetic retinopathy, Artificial intelligence, Digital health, Real-world implementation
Key Summary Points
| Why carry out this study? |
| The transformation of existing non-digital diabetic retinopathy (DR) screening programs and workflows into digital is essential to implement artificial intelligence (AI) and monitor important indicators, which are not possible to be measured without digital workflows. |
| To assess the transformation, a cloud-based digital platform was developed and implemented to prospectively screen DR, using either AI or manual grading of retinal photographs, in clinical workflows of a community hospital and a referral center. |
| What was learned from this study? |
| The platform was found to be an effective tool to digitize the existing non-digital DR screening workflows to be integrated with AI and manual gradings of retinal photographs. |
| Advantages of the platform include monitoring of screening performances, for which we found trends of more referrals and true positives detected by AI compared to manual screening, and tracking of referrals, for which we found more referral adherence while screening by AI. |
Introduction
Screening for diabetic retinopathy (DR) is recommended as one of the key strategies to reduce global blindness [1]. Despite worldwide screening, DR is still the leading cause of global blindness among the working population and projected to be so in the coming decades [2]. The effectiveness of DR screening programs is determined by timely detection, referral, and treatment for vision-threatening diseases [3, 4]. Artificial intelligence (AI), with its robust diagnostic accuracy for DR screening [5], can be applied and has shown benefits in real-world, resource-constrained settings [6].
Thailand is an upper-middle-income country with an estimated 6 million patients with diabetes [7]. The Thai Ministry of Public Health established DR screening programs by training mid-level ophthalmic personnel, such as nurses, to provide manual grading of retinal photographs at primary care settings nationwide. Despite the successful program in improving access to DR screening for a large Thai population, challenges remain.
Firstly, the average screening coverage, 50% of patients with diabetes, falls short of the national target of 60%. Secondly, manual screening by trained personnel has limited accuracy in identifying referrable DR [8]. Thirdly, the referral adherence rate, measured as a proportion of patients screened positive as referrals who attended tertiary care hospitals for management, is unknown. These challenges are not unique to the Thai programs but are commonly observed in other screening programs worldwide.
To address part of these challenges, on the perspective of health systems, there is a need to transform Thailand’s existing conventional DR screening workflows, which are mostly non-digital and standalone without hospital information system integration, into digital workflows. Digitizing these workflows, the screening programs may have the capability of not only being integrated with AI but also monitoring important implementation indicators, such as referral adherence, which have never been measured before.
In this study, we designed, developed, and implemented a cloud-based digital platform in the conventional DR screening workflow for either AI or manual grading of retinal photographs. The main objective of this study is to demonstrate the implementation of the platform in AI and manual screenings. The outcomes of this implementation, which were not aimed for direct comparison between both screenings, were evaluated as real-world results in each screening. These results may be extrapolated to other screening programs worldwide. This is to achieve high-performing DR screening programs to reduce the global burden of diabetic blindness.
Methods
Study Design
A digital application platform was designed and developed for DR screening via a web browser on the Internet. We prospectively conducted DR screening using this platform in a community hospital where the conventional screening was regularly conducted to implement AI and manual modality on alternate weeks between January and August 2022.
Setting
Thailand’s nationwide DR screening programs have been conducted since 2013 in almost 900 primary care centers and community and district hospitals across the country to screen approximately 500,000 to 2,000,000 patients with diabetes yearly. Full coverage of at least a community hospital in each district nationwide is the foundation of primary healthcare systems development in Thailand [9]. Most of the primary care centers and community, district hospitals have diabetes clinics, though they do not have ophthalmologists or eye clinics. DR screenings in these hospitals are mostly conducted in standalone, non-digital workflows, which are separated from hospital information systems. The existing main screening modality is manual grading of retinal photographs by trained mid-level health personnel, who also capture the photographs, to detect referrals.
Participants
A sample population is a cohort of patients with diabetes who underwent yearly DR screening at U-Tai District Hospital in Ayutthaya Province of Thailand. This hospital was purposely chosen for this pilot. All patients with diabetes were included except those who had already been referred for follow-up or treatment of eye diseases, or those who had conditions where retinal photographs could not be taken, such as inability for positioning for image capture. The patients who met referral criteria were referred to Ayutthaya Hospital, a provincial tertiary care center.
Implementation and Variables
In general, the digital platform was designed to support three stages of DR screening workflows: (1) Detecting referrals at primary care centers (PC), where patients had their retinal photographs captured and uploaded for grading by AI providing real-time results, or grading by trained personnel; (2) Over-reading of retinal photographs by remote experts at referral, tertiary care centers (TC); (3) Tracking referral adherence of patients referred at PC and presented at TC. In this study, PC was U-Tai Hospital and TC was Ayutthaya Hospital.
For the implementation at PC, patients with diabetes were registered on the digital platform for screening by either AI or manual modality. Visual acuity (VA) test was provided with pinhole for both eyes to all patients using a Snellen chart. Baseline characteristics and VA of both eyes were also recorded. Those with positive screening results were appointed via the platform for further management at TC. Those with negative results were appointed for re-screening at PC in the following year. This workflow was routine and followed in this study.
In the AI modality, patients were fully informed about AI-based screening. This AI model was developed using Inception-v4 architecture based on datasets of more than 1,600,000 retinal photographs of patients with diabetes from the United States, India, France and tested on the United States data [10]. The architecture of this model was previously described [10]. This cloud-based AI model, called Automated Retinal Disease Assessment (ARDA), could provide outputs as severity of DR, the presence or absence of diabetic macular edema (DME), and image quality with both sensitivity and specificity of > 95% according to the previous studies [5, 6, 8, 10]. For referral positives detected by AI, the platform automatically sent referral reminders to the patients’ phones a few days before the referral date. All retinal photographs graded by AI were accessible on the platform and over-read by a retinal specialist at TC. This over-reading was required once a week. If false positives, which means the over-reading disagreed with AI positive reading, or false negatives, which means missed referrals, were detected, the nurses at PC were notified to inform the patients on changing referral accordingly.
Since the ground truth adjudication to determine the actual true positive and false positive cases was not feasible during this real-world screening period, over-reading by the retinal specialist at TC was used as the benchmark for AI.
In the manual screening, the retinal photographs were graded by trained personnel. The patients with positive results were either referred to TC on the screening day or appointed for in-person retinal examination by a visiting, general ophthalmologist in the monthly eye clinic at PC to determine referrals. This is a routine practice of screening at this PC.
The same retinal camera (Topcon® TRC-NW400, Tokyo, Japan) was used to capture retinal photographs, 45°, single field for each eye with the macula at the center in all patients, by the same trained personnel in both screening modalities. This mid-level practitioner received specific theoretical and practical training, including capturing and grading of retinal photographs, with more than 10 years of experience in DR screening. The training on grading the photographs was a 1.5-day instruction course covered pre-test, understanding of DR including its pathophysiology, normal structures of retina including pathological lesions of DR and how to identify them in retinal photographs, severity scales of DR according to the International Clinical Classification, the presence or absence of DME, and post-test. The accuracy of the trained personnel for detecting referrable DR was maintained at approximately 85% [11]. Their sensitivity and specificity were found at approximately 70% and 95%, respectively [8].
The quality of the retinal photographs was assessed as gradable or ungradable. This gradability was assessed by judging the visibility of normal retinal structures and retinal pathologies. For photographs with obscured detail in half or more of their entire area, ungradable was called. In the manual screening, the nurses who took the photographs were those who made the call. For AI, it was trained from labeling the photographs a similar way.
From January to March 2022, the retinal photographs were taken without mydriasis, whereas from April to August 2022, the photographs were taken with mydriasis, in both screening modalities. The mydriasis was provided by applying a drop of 1% tropicamide eye solution to each eye. Screening performance with and without mydriasis in each modality was also assessed as an additional objective in this study since mydriasis is not required as the routine practice of DR screening in the current system.
In both AI and manual screenings, referrals were defined as patients with referrable conditions in either eye. These conditions were vision-threatening DR (VTDR), which included DME, severe non-proliferative DR (NPDR), and proliferative DR, or ungradable retinal photograph. Patients were also referred if their VA with pinhole in either eye was ≤ 20/70.
Data recorded on the platform based on AI screening were categorized for each eye into (1) DR severity level classified according to the International Clinical Classification of DR [3], (2) Presence or absence of DME, and (3) Gradability of retinal photographs for either DR or DME.
Data recorded based on the manual screening were categorized into only referrable or non-referrable since the personnel were trained to detect referrable cases without classifying them in detail [12].
For the implementation at TC, nurses at the eye clinic registered patients referred from either screening modality on the same platform. They recorded the management of patients, which was categorized into either treatment type, or monitoring, or referral to other hospitals.
Outcome Measures
Outcomes were the number of referrals identified by either screening modality on the platform. Also, the number of patients with true positive (TP), false positive (FP), true negative (TN), and false negative (FN), including referral adherence at TC. These parameters, TP, FP, TN, and FN, reflected the comparison of AI against the retinal specialist over-reader at TC. They did not reflect the true values since the adjudication of gradings of the retinal photographs in both modalities as gold standards was not conducted. The duration of workflow in each modality was also recorded.
Analyses
Student’s t test was used to assess continuous variables, with P value < 0.05 indicating statistical significance. Pearson’s chi-square test was used to assess categorical variables in normal distribution, whereas two-sided Fisher’s exact test was used for variables not in normal distribution, with P value < 0.05 indicating statistical significance. All analyses were conducted using SPSS® version 22.
Ethical Approval
This study was approved by the Research Ethics Committee of Rajavithi Hospital (reference number: 64077) and fully adhered to Thailand’s Personal Data Protection Act 2019. All patient data were de-identified; all patients signed informed consent.
Results
A total of 708 and 746 patients with diabetes were prospectively screened by AI and manual modality, respectively. Their baseline characteristics are shown in Table 1. The results of each screening modality are shown in Table 2 and Fig. 1.
Table 1.
Baseline characteristics of patients in each screening modality
| Screening by artificial intelligence (AI) | Manual screening | P value | |
|---|---|---|---|
| Number of patients | 708 | 746 | |
| Mean age ± SD | 63.01 ± 10.62 | 61.97 ± 10.85 | 0.064 |
| Gender | 0.934 | ||
| Male (%) | 272 (38.4) | 282 (37.8) | |
| Female (%) | 435 (61.4) | 462 (61.9) | |
| Transgender (%) | 1 (0.1) | 2 (0.3) | |
| Types of diabetes | 0.615 | ||
| Diabetes type 1 (%) | 2 (0.3) | 1 (0.1) | |
| Diabetes type 2 (%) | 706 (99.7) | 745 (99.9) | |
| Duration of diabetes, mean in years ± SD (min–max) | 8.37 ± 6.78 (1–25) | 8.81 ± 6.7 (1–35) | 0.224 |
| Number of patients in each duration | |||
| 1–10 years (%) | 533 (75.3) | 542 (72.7) | 0.254 |
| 11–20 years (%) | 133 (18.8) | 165 (22.1) | 0.116 |
| 21–30 years (%) | 23 (3.2) | 23 (3.1) | 0.857 |
| 31–40 years (%) | 2 (0.3) | 3 (0.4) | 1.000 |
| > 40 years (%) | 1 (0.1) | 0 (0.0) | 0.487 |
| Data unavailable by patients (%) | 16 (2.3) | 13 (1.7) | 0.481 |
| Pupillary dilatation | 0.050 | ||
| With dilatation (%) | 328 (46.3) | 384 (51.5) | |
| Without dilatation (%) | 380 (53.7) | 362 (48.5) | |
The P values are from the comparisons between the two screening modalities using descriptive statistical analyses as indicated in text
Table 2.
Screening results of patients in each screening modality after excluding referrals by the visual acuity (VA) criteria (≤ 20/70)
| Screening by artificial intelligence (AI) | Manual screening | P value | |
|---|---|---|---|
| Number of patients screened after exclusion by VA criteria | 645 | 730 | |
| Referral positives identified at the point-of-screening (%) | 201/645 (31.2) | 175/730 (24.0) | 0.003 |
| True positives (TP) (%) | 129/201 (64.2) | 11/116 (9.5) | < 0.001 |
| Adherence to referral for TP (%) | 115/129 (89.1) | 17/22 (77.3) | 0.158 |
| False negatives (FN) (%) | 5/444 (1.1) | NA | NA |
| Adherence to referral for FN (%) | 4/5 (80.0) | NA | NA |
| Number of patients screened with mydriasis after exclusion by VA criteria (%) | 297 (46.1) | 355 (48.6) | 0.338 |
| Referral positives with mydriasis (%) | 55/297 (18.5)* | 55/355 (15.5)* | 0.304 |
| Referral positives without mydriasis (%) | 146/348 (42.0)* | 120/375 (32.0)* | 0.006 |
| TP with mydriasis (%) | 38/55 (69.1)** | 11/38 (28.9)* | < 0.001 |
| TP without mydriasis (%) | 91/146 (62.3)** | 0/78 (0.0)* | < 0.001 |
True or false positives in AI screening were determined by the retinal specialist over-reader at the tertiary care who over-read retinal photographs of all patients in this modality. True or false positives in manual screening were determined by retinal examination to the patients, screened as positives, who re-visited in the monthly general ophthalmologist clinic at the primary care hospital. False negatives were not available in this manual screening since the patients with negative screening results were required to re-screening next year
The P values are from the comparisons between the two screening modalities using either chi-square test or Fisher’s exact test, when indicated
*P values for these comparisons between with and without mydriasis in each modality are < 0.001
**P values for this comparison of true positives between with and without mydriasis in the AI screening is 0.373
Fig. 1.
Screening workflow of each modality. AI artificial intelligence, VA visual acuity; + = screened positive for referral; – = screened negative for referral; TP true positive, FP false positive, FN false negative, TN true negative, RTC return to the monthly ophthalmologist clinic. The number in each box indicates the number of patients
After excluding 63 patients in the AI group and 16 in the manual group because of VA ≤ 20/70, 201 out of 645 (31.16%) and 175 out of 730 (23.97%) were identified as referrals in each respective group. See Fig. 1.
In the AI group, patients determined as TP and TN, which meant cases with good concordance between AI and the over-reader, were found at 88% (568/645). The over-reader agreed with 64.18% of the positives by AI as TP (129/201), of which 89.15% (115/129) presented at TC. For the 72 FP cases, the nurses were able to contact and inform 68.06% (49/72) of them for reverting referral, though two (4.08%) still presented at TC. Of the remaining 23 FP cases the nurses could not contact, 65.22% (15/23) presented at TC. There were only 1.13% FN (5/444); all of them were contacted and 80% (4/5) presented at TC.
In the manual group, at the point-of-screening, the trained personnel were confident to refer 11 out of 175 referral positives (6.29%) and appointed the rest 164 suspected positives to return to the monthly eye clinic for confirmation. A total of 70.73% (116/164) returned and the general ophthalmologist at the clinic identified 9.48% (11/116) as referrals with eye conditions other than DR. Finally, 81.81% (9/11) of these referrals presented at TC.
To observe the trends of the two screening modalities (Table 2), the proportions of patients with referral positive and TP were found to be significantly higher in the AI group. Adherence to referral was higher in the AI screening, though not statistically significant. We did not assess FN in the manual screening since we intended to follow the routine screening practice at the PC in which patients with negative screening results were appointed for re-screening next year. They were not required to return to the monthly eye clinic.
There were 297 (46.05%) and 355 (48.63%) patients screened with mydriasis in the AI and manual group, respectively. Mydriasis significantly decreased the number of patients with referral positive and increased the number of patients with TP in both modalities. Patients with TP in the AI screening increased from 62.33% without mydriasis to 69.09% with mydriasis, whereas the TP in the manual screening increased significantly from none without mydriasis to 28.95% with mydriasis.
Table 3 presents the management of patients who attended their referrals at TC. Slightly less than half of patients with referral adherence in AI screening (48/115, 41.74%) were VTDR; the rest were those with ungradable photographs in either eye. Most of the ungradable photographs were due to cataracts, which were diagnosed using slit-lamp examinations by ophthalmologists at TC; two of the patients with cataract received treatment whereas the rest (63 patients) were monitored. We found three out of four patients with FN in AI screening as severe NPDR, and another ungradable due to cataract. Slightly more than half of those referred due to worse VA (28/52, 53.85%) in AI screening had ungradable retinal photographs, while 44.23% (23/52) of them had eye diseases other than DR, and one had severe NPDR. In the manual group, the workflow did not allow exploring the causes of referrals, such as VTDR, ungradable, or other causes, though management at TC could be tracked.
Table 3.
Management of referred patients presenting at the provincial, tertiary care hospital
| Receive treatment | Monitoring | Referral | No data | |
|---|---|---|---|---|
| Referral adherence in AI screening (n = 115) | ||||
| VTDR (48; DME 45, Severe NPDR or PDR 3) | 18 | 25 | 5 | – |
| Ungradable (67) | 2 | 63 | 1 | 1 |
| False negative in AI screening (n = 4) | ||||
| VTDR (3; Severe NPDR) | – | 3 | – | – |
| Ungradable (1) | – | 1 | – | – |
| Referral by VA criteria in AI screening (n = 52) | ||||
| VTDR (1; Severe NPDR) | – | 1 | – | – |
| Ungradable (28) | 1 | 26 | – | 1 |
| Other causes (23) | 1 | 22 | – | – |
| Referral adherence in manual screening (n = 29) | ||||
| From trained nurses at the point-of-screening (8) | 7 | 1 | – | – |
| From visiting ophthalmologist at the monthly eye clinic (9) | – | 9 | – | – |
| Referral by VA criteria in conventional screening (12) | 3 | 9 | – | – |
Receive treatment means laser photocoagulation or intravitreal injection in case of VTDR, and cataract surgery in case of cataract. Referral means patients are referred to other larger tertiary care hospitals
AI artificial intelligence, VTDR vision-threatening diabetic retinopathy, DME diabetic macular edema, NPDR non-proliferative diabetic retinopathy, PDR proliferative diabetic retinopathy, VA visual acuity
The duration of the screening processes, including retinal photograph capture, uploading and grading the photographs in the AI group, grading without uploading in the manual group, and determination of referral by saving data on the platform, was measured for each modality. This duration excluded the processes of patient data registration prior to photograph capture. The median duration for a patient in the AI and manual screening was 6:30 and 5:38 min, respectively. There was no system failure in implementing the platform in both modalities.
Discussion
With this prospective, real-world implementation research, we demonstrated that a cloud-based digital platform could be integrated with either AI or manual modality for real-world DR screening to effectively measure important parameters to assess the performances of the screening. Recently, there have been more studies on prospective validation of AI models for DR screening though some of them may not be real-world implementation. We summarize these studies as in Table 4 [6, 13–19]. None of these studies assessed AI in integration with a digital platform for real-world implementation.
Table 4.
A summary of recent publications on prospective validation of artificial intelligence for diabetic retinopathy screening
| Country where the study was conducted | Number of patients and clinics | Performance: sensitivity and specificity | Models and strengths | Challenges or requirements |
|---|---|---|---|---|
| United Kingdom [13] | 30,405 consecutive cases across 3 centers of English Diabetic Eye Screening Programme | 95.7% and 54.0% | EyeArt, potentially saves £0.5 million per 10,000 screening episodes | Huge number of patients with diabetes in this study for systematic screening |
| Australia [14] | 236 from 2 endocrine clinics, 3 Aboriginal medical services | 96.9% and 87.7% | Study as opportunistic screening | Small study sample size |
| China [15] | 47,269 patients from 155 diabetes centers | 83.3% and 92.5% | VoxelCloud Retina, qualified images could reach 92.8% | Quality control module had a 63.3% sensitivity and 85.0% specificity |
| Thailand [6] | 7940 patients from 9 primary care centers | 91.4% and 95.4% | ARDA, AI was embedded into real-world workflows without digital platform | Only 2 of 13 nationwide health regions were studied |
| United States [16] | 900 patients from 10 primary care units | 87.2% and 90.7% | IDx-DR, the first US FDA-approved model for autonomous screening | 2 images per eye |
| Africa [17] | 1574 patients from mobile screening unit in Zambia | 92.3% and 89.0% | SELENA, validated in multiethnic groups | 2 images per eye, small study sample size |
| China [18] | 1001 patients from 3 eye hospitals in China | 86.7% and 96.1% | AIDRScreening system, to detect referrable diabetic retinopathy | 2 images per eye, mydriasis is required and no detection of diabetic macular edema |
| United States [19] | 893 patients from 6 primary care centers, 6 general ophthalmology clinic, and 3 retina centers | 95.5% and 85.0% | EyeArt, another US FDA-approved model for autonomous screening | 2 images per eye |
EyeArt, VoxelCloud, ARDA, IDx-DR, SELENA, and AIDRScreening system are the names of AI models
US FDA United States Food and Drug Administration
Implementing the platform, we found a trend that AI could detect significantly more referrals, including patients with VTDR, than the manual screening. This result supported our previous prospective [6] and retrospective [8] validation studies of AI for DR screening in Thailand, though without a digital platform integration. The identification of referrals at about 30% by AI in this study approximated the same percentage found in our previous prospective study [6]. Other benefits of implementation of a digital platform for disease screening include facilitating service provision and appointment, reminder messaging, and providing relevant information for assessment in a timely manner. The implementation outcomes for both screening modalities can be improved further via linking the digital platform with a hospital’s electronic medical record system, retinal cameras, and patient appointments at different levels of care.
The higher referral adherence rate by AI than manual screening could be because patients received real-time screening results by AI [20], or the digital referral reminder notifications [20]. Nevertheless, it was unclear to what extent AI or digital notifications individually contributed to the high referral adherence rate. The significantly higher TP generated by AI could also be another factor contributing to the high adherence rate.
In the manual screening, however, the referral adherence rate was still higher than other publications, which were typically lower than 50% [21, 22]. This high rate could have been influenced by the awareness of participation in this study by both patients and screening personnel. Nevertheless, these data highlight the importance of referral adherence and completion of the patient journey. Screening programs are only useful if a patient with a true positive result completes the “loop”, seeks follow-up, and receives care in a timely manner. This is particularly important for DR screening since patients with VTDR may not experience visual symptoms that influence them to seek further management. AI models for DR screening, which are available as standalone software, may be more useful with integration into existing screening programs to assist in completing the patient journey of referral adherence [23].
The increased TP by both screening modalities with mydriasis highlighted the essential role of mydriasis in DR screening. Previous studies on manual DR screening showed that mydriasis could reduce ungradable retinal photographs from 40–50% to 10–15% [24]. In addition, few adverse events related to mydriasis were encountered [25–27]. In this study, we found no adverse events from mydriasis.
Due to the infeasibility of conducting adjudication gradings in real-world screening of DR, a retinal specialist at the referral center was assigned to over-read all AI results for comparisons in this study. Although the exact TP, TN, FP, and FN cases were not known without adjudication, there might be approximately 35% of cases which AI unnecessarily referred and approximately 1% of cases which AI missed referral, according to the specialist. Nevertheless, the concordance between AI and the specialist was as high as 88%. This may imply that over-reading of AI results by expert human graders may not be necessary in implementing AI for DR screening but a mechanism to lower the unnecessary referral should be set up in a screening program. This mechanism, for example, can be the integration of VA with grading by AI for referral of DME since early DME detected by AI may not be referred for treatment if VA is better than a certain value. This VA threshold may be different based on real-world resources and settings of screening programs.
We believe the overlooking of FN, since cases with screen negative are not checked, is the major weakness of the manual screening which may be found in other similar DR screening programs worldwide. This problem should be addressed. To overcome this shortcoming, a digital platform should be integrated and specifically designed, such as programming the platform for random selection of a proportion of negative cases for validation by over-reading to verify FN in the manual screening.
The trend of lower referral cases and TP in the manual screening may be due to variation of skills among trained personnel who grade the photographs. The slightly longer duration of AI screening compared to manual screening, caused by image uploading for AI processing, can be minimized through improved digitized workflows, such as installing AI models locally or utilizing faster internet connections. User familiarity and further improvement of user interfaces of the platform may also improve the workflow efficiency in the future.
The visiting specialist clinic at primary care is a proposed modality to improve patient eye care in rural communities [28]. The requirement for patients with the risk of being FP in manual screening to return to a specialist clinic for re-examination, however, is time- and resource-consuming for both the patients and the programs. Online over-reading of the manual screening results via digital platform could be an option for not requiring patients to return to the clinic for re-examination. This may allow the clinic to then focus on patient care with other non-DR eye conditions. The varied results of detecting TP and FP at the visiting eye clinic, when compared to AI in this study, might reflect the different skills between general ophthalmologists and retinal specialists.
Like other studies [14, 29], FN was found in only a few patients (1%) from AI screening. We successfully contacted all patients for further management at the TC, although one patient did not attend. Most (three out of four) of the FN were severe NPDR without DME, which AI graded as moderate NPDR without DME. These two severity levels of NPDR are sometimes challenging to discern, even for retinal specialists [11]. Another case of FN was graded as negative by AI, but the over-reader determined as ungradable and required referral. From the follow-up examination at the TC, this patient was diagnosed as having cataract without surgical indications.
Regarding limitations, although we implemented the platform to screen DR using AI and manual screening on alternate weeks, this study was not designed as a randomized controlled trial (RCT). The sample size in each group was not properly estimated as in a RCT. The two modalities also had different workflows. The results from both screenings, therefore, might not be directly compared, though they could provide the trends. This study did not have external experts’ adjudication as the gold standard gradings of the retinal photographs in both screening modalities, therefore, we did not estimate their diagnostic parameters, such as sensitivity and specificity. However, we assumed that these parameters would be like those in our previous studies [6, 8] since the same AI model was used. While the platform could provide further benefits by capturing the long-term, follow-up, treatment outcomes of the patients, this would have required additional manpower and resources, which were not available in this pilot. Lastly, the interobserver agreement between the over-reader at the TC and the visiting general ophthalmologist at PC, was not assessed.
Conclusions
We found benefits and challenges in implementing a cloud-based digital platform for AI and manual screening for DR. Given the capability of the platform to measure some currently unknown parameters, and the possibility of higher yields in both TP and referrals by AI than manual modality, the digitization of a conventional, non-digital workflow via a cloud-based platform is warranted. The platform may be essential for not only implementing AI but also improving the performance of a screening program. This study can be replicated and scaled up in other similar screening settings, particularly in other low- and middle-income countries. In addition, we are implementing this system, the digital platform integrating AI, to cover all health regions in the country. With this scaling up, we should be able to identify more challenges and obstacles to improve the system further in the future.
Acknowledgements
The digital platform deployed in this study was developed and provided by Forus Health Pvt. Ltd. The AI model for DR screening was developed and provided by Google LLC.
Author Contributions
Peranut Chotcomwongse and Paisan Ruamviboonsuk designed the concept and executed this study; prepared, reviewed, edited, and finalized the manuscript. Chaiwat Karavapitayakul, Koblarp Thongthong and Turean Waiwaree executed the study at the tertiary care level. Anyarak Amornpetchsathaporn and Methaphon Chainakul assisted in the proposal draft, execution of this study, data collection, and data analysis. Malee Triprachanath and Eckachai Lerdpanyawattananukul executed this study at the primary care level. Niracha Arjkongharn assisted in study execution, data collection, and data analysis. Varis Ruamviboonsuk assisted in data collection, presentation of data, literature review, and manuscript preparation. Hathaiphan Ruampunpong analyzed data and presented data. Richa Tiwari designed concepts, analyzed data, reviewed, edited, and finalized the manuscript. Nattaporn Vongsa and Pawin Pakaymaskul reviewed and edited the manuscript. Viroj Tangcharoensathien reviewed, commented, edited, and finalized the manuscript. All authors read and approved the final manuscript.
Funding
This project is funded by Health System Research Institute of the Ministry of Public Health of Thailand. No funding or sponsorship was received for the publication of this article.
Data Availability
Data in this study were anonymized and saved in the system of U-Tai Community Hospital in Thailand. The datasets generated during the current study are available from the corresponding author upon reasonable request.
Declarations
Conflict of Interest
Paisan Ruamviboonsuk received consulting fees from and participated in Advisory Boards of Roche and Bayer; honoraria for lectures for Roche, Novartis, and Bayer; support for meeting or traveling from Roche and Bayer; he is a member of the Editorial Board of Ophthalmology and Therapy but was not involved in the selection of peer reviewers for this manuscript or in any subsequent editorial decisions. Richa Tiwari is a Google employee and owns stock at Alphabet. Hathaiphan Ruampunpong is an employee of Persolkelly HR services recruitment subcontracted with Google. The remaining authors (Peranut Chotcomwongse, Chaiwat Karavapitayakul, Koblarp Thongthong, Anyarak Amornpetchsathaporn, Methaphon Chainakul, Malee Triprachanath, Eckachai Lerdpanyawattananukul, Niracha Arjkongharn, Varis Ruamviboonsuk, Nattaporn Vongsa, Pawin Pakaymaskul, Turean Waiwaree and Viroj Tangcharoensathien) declare no competing interest.
Ethical Approval
This study was approved by the Research Ethics Committee of Rajavithi Hospital (reference number: 64077) and fully adhered to Thailand’s Personal Data Protection Act 2019. All patient data were de-identified; all patients signed informed consent forms.
References
- 1.Wong TY, Sabanayagam C. Strategies to tackle the global burden of diabetic retinopathy: from epidemiology to artificial intelligence. Ophthalmologica. 2020;243(1):9–20. [DOI] [PubMed] [Google Scholar]
- 2.Steinmetz JD, Bourne RRA, Briant PS, Flaxman SR, Taylor HRB, Jonas JB, et al. Causes of blindness and vision impairment in 2020 and trends over 30 years, and prevalence of avoidable blindness in relation to VISION 2020: the right to sight: an analysis for the Global Burden of Disease Study. Lancet Glob Health. 2021;9(2):e144–60. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Wong TY, Sun J, Kawasaki R, Ruamviboonsuk P, Gupta N, Lansingh VC, et al. Guidelines on diabetic eye care: the international council of ophthalmology recommendations for screening, follow-up, referral, and treatment based on resource settings. Ophthalmology. 2018;125(10):1608–22. [DOI] [PubMed] [Google Scholar]
- 4.World Health Organization. Regional Office for E. Diabetic retinopathy screening: a short guide: increase effectiveness, maximize benefits and minimize harm. Copenhagen: World Health Organization. Regional Office for Europe; 2020. [Google Scholar]
- 5.Gulshan V, Peng L, Coram M, Stumpe MC, Wu D, Narayanaswamy A, et al. Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA. 2016;316(22):2402–10. [DOI] [PubMed] [Google Scholar]
- 6.Ruamviboonsuk P, Tiwari R, Sayres R, Nganthavee V, Hemarat K, Kongprayoon A, et al. Real-time diabetic retinopathy screening by deep learning in a multisite national screening programme: a prospective interventional cohort study. Lancet Digit Health. 2022;4(4):e235–44. [DOI] [PubMed] [Google Scholar]
- 7.International Diabetes Federation. Thailand diabetes report 2000—2045. Brussels, Belgium: International Diabetes Federation; 2021. Available from: https://diabetesatlas.org/data/en/country/196/th.html.
- 8.Raumviboonsuk P, Krause J, Chotcomwongse P, Sayres R, Raman R, Widner K, et al. Deep learning versus human graders for classifying diabetic retinopathy severity in a nationwide screening program. NPJ Digit Med. 2019;2:25. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Tangcharoensathien V, Witthayapipopsakul W, Panichkriangkrai W, Patcharanarumol W, Mills A. Health systems development in Thailand: a solid platform for successful implementation of universal health coverage. Lancet. 2018;391(10126):1205–23. [DOI] [PubMed] [Google Scholar]
- 10.Gulshan V, Rajan RP, Widner K, Wu D, Wubbels P, Rhodes T, et al. Performance of a deep-learning algorithm vs manual grading for detecting diabetic retinopathy in India. JAMA Ophthalmol. 2019;137(9):987–93. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Ruamviboonsuk P, Teerasuwanajak K, Tiensuwan M, Yuttitham K. Interobserver agreement in the interpretation of single-field digital fundus images for diabetic retinopathy screening. Ophthalmology. 2006;113(5):826–32. [DOI] [PubMed] [Google Scholar]
- 12.Silpa-Archa S, Limwattanayingyong J, Tadarati M, Amphornphruet A, Ruamviboonsuk P. Capacity building in screening and treatment of diabetic retinopathy in Asia-Pacific region. Indian J Ophthalmol. 2021;69(11):2959–67. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Heydon P, Egan C, Bolter L, Chambers R, Anderson J, Aldington S, et al. Prospective evaluation of an artificial intelligence-enabled algorithm for automated diabetic retinopathy screening of 30,000 patients. Br J Ophthalmol. 2021;105(5):723–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Scheetz J, Koca D, McGuinness M, Holloway E, Tan Z, Zhu Z, et al. Real-world artificial intelligence-based opportunistic screening for diabetic retinopathy in endocrinology and indigenous healthcare settings in Australia. Sci Rep. 2021;11(1):15808. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Zhang Y, Shi J, Peng Y, Zhao Z, Zheng Q, Wang Z, et al. Artificial intelligence-enabled screening for diabetic retinopathy: a real-world, multicenter and prospective study. BMJ Open Diabetes Res Care. 2020;8(1): e001596. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Abràmoff MD, Lavin PT, Birch M, Shah N, Folk JC. Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices. NPJ Digit Med. 2018;1:39. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Bellemo V, Lim ZW, Lim G, Nguyen QD, Xie Y, Yip MYT, et al. Artificial intelligence using deep learning to screen for referable and vision-threatening diabetic retinopathy in Africa: a clinical validation study. Lancet Digit Health. 2019;1(1):e35–44. [DOI] [PubMed] [Google Scholar]
- 18.Yang Y, Pan J, Yuan M, Lai K, Xie H, Ma L, et al. Performance of the AIDRScreening system in detecting diabetic retinopathy in the fundus photographs of Chinese patients: a prospective, multicenter, clinical study. Ann Transl Med. 2022;10(20):1088. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Ipp E, Liljenquist D, Bode B, Shah VN, Silverstein S, Regillo CD, et al. Pivotal evaluation of an artificial intelligence system for autonomous detection of referrable and vision-threatening diabetic retinopathy. JAMA Netw Open. 2021;4(11): e2134254. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Liu J, Gibson E, Ramchal S, Shankar V, Piggott K, Sychev Y, et al. Diabetic retinopathy screening with automated retinal image analysis in a primary care setting improves adherence to ophthalmic care. Ophthalmol Retina. 2021;5(1):71–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Jani PD, Forbes L, Choudhury A, Preisser JS, Viera AJ, Garg S. Evaluation of diabetic retinal screening and factors for ophthalmology referral in a telemedicine network. JAMA Ophthalmol. 2017;135(7):706–14. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Keenum Z, McGwin G Jr, Witherspoon CD, Haller JA, Clark ME, Owsley C. Patients’ adherence to recommended follow-up eye care after diabetic retinopathy screening in a publicly funded county clinic and factors associated with follow-up eye care use. JAMA Ophthalmol. 2016;134(11):1221–8. [DOI] [PubMed] [Google Scholar]
- 23.Yuan A, Lee AY. Artificial intelligence deployment in diabetic retinopathy: the last step of the translation continuum. Lancet Digit Health. 2022;4(4):e208–9. [DOI] [PubMed] [Google Scholar]
- 24.Banaee T, Ansari-Astaneh MR, Pourreza H, Faal Hosseini F, Vatanparast M, Shoeibi N, et al. Utility of 1% tropicamide in improving the quality of images for tele-screening of diabetic retinopathy in patients with dark irides. Ophthalmic Epidemiol. 2017;24(4):217–21. [DOI] [PubMed] [Google Scholar]
- 25.Lagan MA, O’Gallagher MK, Johnston SE, Hart PM. Angle closure glaucoma in the Northern Ireland Diabetic Retinopathy Screening Programme. Eye (Lond). 2016;30(8):1091–3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Xiong K, Wang L, Li W, Wang W, Meng J, Gong X, et al. Risk of acute angle-closure and changes in intraocular pressure after pupillary dilation in patients with diabetes. Eye (Lond). 2023;37(8):1646–51. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Tan GS, Wong CY, Wong TY, Govindasamy CV, Wong EY, Yeo IY, et al. Is routine pupil dilation safe among Asian patients with diabetes? Invest Ophthalmol Vis Sci. 2009;50(9):4110–3. [DOI] [PubMed] [Google Scholar]
- 28.Bowling A, Bond M. A national evaluation of specialists’ clinics in primary care settings. Br J Gen Pract. 2001;51(465):264–9. [PMC free article] [PubMed] [Google Scholar]
- 29.Grzybowski A, Rao DP, Brona P, Negiloni K, Krzywicki T, Savoy FM. Diagnostic accuracy of automated diabetic retinopathy image assessment softwares: IDx-DR and Medios artificial intelligence. Ophthalmic Res. 2023;66(1):1286–92. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
Data in this study were anonymized and saved in the system of U-Tai Community Hospital in Thailand. The datasets generated during the current study are available from the corresponding author upon reasonable request.

