Skip to main content
Ophthalmology Science logoLink to Ophthalmology Science
. 2025 Sep 2;6(1):100933. doi: 10.1016/j.xops.2025.100933

Data Duplication and Errors in Large Medical Data Sets: A Case Study in the IRIS® Registry

Eric A Goldberg 1,, Connor J Ross 1, Vivian Paraskevi Douglas 1, Alexander Ivanov 1, Tobias Elze 1, Joan W Miller 1, Alice C Lorch 1; IRIS Registry Analytic Center Consortium1, on behalf of the
PMCID: PMC12550146  PMID: 41140906

Abstract

Purpose

To investigate entry errors and data duplication within the American Academy of Ophthalmology IRIS® Registry (Intelligent Research in Sight) utilizing cataract surgery (CS), neodymium-doped: yttrium aluminum garnet (YAG) capsulotomy, age-related macular degeneration (AMD), and diabetic retinopathy (DR) records.

Design

Retrospective cohort study.

Participants

Patients in the IRIS Registry.

Methods

We collected records of CS and YAG capsulotomy with specified laterality within the IRIS Registry (years 2013–2023), identifying eyes having >1 record and eyes having ≥1 record on a date after the first entry (different date duplication, Dd). Additionally, we identified eyes amongst records of DR and AMD with (1) a diagnosis indicating a more severe stage then reversion to the less severe stage or (2) a transition to a more severe stage before later being diagnosed with the less severe stage, defined as transition errors. We investigated potential predictors of Dd and transition errors among patient and practice characteristics by evaluating the permutation feature importance (PFI) of classification models.

Main Outcome Measures

For CS and YAG capsulotomy, we measure the proportion of eyes having >1 procedure record, having >1 record only on the initial procedure date, and having ≥1 procedure record on a date after the first entry. For DR and AMD, we measure the proportion of eyes reverting to an earlier stage after starting at a later stage and the proportion reverting to an earlier stage after transitioning to a later stage.

Results

Of the 14 718 896 CS-treated eyes, 30.9% had duplicates, with 5.5% having Dd. For YAG capsulotomy, out of 5 113 679 eyes, 29.1% had duplicates, with 4.1% having Dd. For AMD and DR, 13.6% and 12.7% of eyes, respectively, exhibited transition errors. Models captured a relationship between the eye’s first practice on record and the data errors under study, indicated by F1-loss = 0.230 (Dd model), 0.062 (transition error model) on average by PFI.

Conclusions

Data duplication in large medical data sets necessitates caution when analyzing repeated procedures or relapsing conditions. Addressing problematic errors requires transparency and communication amongst stakeholders across organizations. Within the IRIS Registry, the results indicated an association between the first record’s originating practice and data errors, providing an investigative entry point for upstream data stewards.

Financial Disclosure(s)

Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.

Keywords: Age-related macular degeneration, Cataract surgery, Data duplication, Data quality, Diabetic retinopathy, IRIS registry, YAG laser capsulotomy


High-quality data are an indispensable tool for providing a reliable source of information across a broad array of fields. In medicine, data accuracy is crucial in advancing knowledge and fostering innovation, but achieving this accuracy is challenging in large data sets. Assessment of these data sets, with consensus around biostatistical methods for improving data accuracy, is essential to big data work. The American Academy of Ophthalmology IRIS® Registry (Intelligent Research in Sight) is the world’s largest national specialty clinical data registry. It includes more than 72.7 million unique patients of 13 619 eye care clinicians1 and supports research through identifying processes, procedures, and treatments leading to enhanced patient outcomes as well as to advance population health through the advancement of patient care.2 Although the IRIS Registry originated as a source of quality data for ophthalmic practitioners, it has developed into an important and powerful tool for research in the field of ophthalmology.

In all large data sets, data entry inconsistencies, incomplete data, coding errors, and data duplication can impact data quality and affect subsequent analysis.3 These errors can occur at multiple levels, including during electronic health record (EHR) entry by clinicians, during data ingestion from a single EHR into a big data set, during the integration of multiple data sources into a big data set, or at the time of processing of a large data set for dissemination. Because identifying error sources relies on full reconciliation between the original data source and a final, compiled data set, it requires a coordinated effort of stakeholders at various levels (data entry—clinician or coder, EHR software provider, data steward managing extract–transform–load from EHR source, data steward handling combined data set, and downstream users of the combined data set). Data at these levels may be necessarily siloed for privacy and compliance purposes (Health Insurance Portability and Accountability Act, etc.), complicating communication, responsibilities, and reconciliation efforts. As such, it is essential to quantify and understand the limitations that these errors can present to better the research value of large data sets.

We use records of prevalent ophthalmic diseases and treatments to evaluate data duplication rates within the IRIS Registry, serving as a case study for identifying these problems in large medical data sets (evaluating data quality pillar of uniqueness) and proposing strategies for their mitigation. In addition, we explore relations between record metadata at the eye level, including patient demographics and practice-based factors, and 2 types of data errors defined below, discuss how these errors can affect downstream analysis, and offer recommendations to reduce bias introduced by them.

Methods

We studied data errors in the IRIS Registry within 2 categories: treatments and diagnoses. For treatments, we leveraged cataract surgery (CS) and neodymium-doped: yttrium aluminum garnet (YAG) laser capsulotomy due to their high prevalence and single-instance nature (cannot repeat in the same eye). Because these procedures should not be repeated in clinical settings, they can provide insight into duplication in a large data set. Primary CS cannot be repeated in the same eye, and only very rarely should a YAG capsulotomy be repeated for a “revision,” compared to a vitrectomy, for example, which can be repeated several times in 1 eye.

Understanding duplication in diagnoses is more challenging, as a given diagnosis may validly appear in a patient’s history multiple times, for example, at initial onset or attached to a patient’s visit record for treatment or follow-up reasons. We chose to examine age-related macular degeneration (AMD) and diabetic retinopathy (DR) due to their expected progressive transitions. The diseases do not revert to less severe stages (nonexudative AMD and nonproliferative DR) after progressing to a more severe stage of exudative AMD or proliferative DR. We captured invalid transitions, defined as transition errors, within diagnosis records. Additionally, we quantified the proportion of eyes having invalid transitions exhibiting disjoint sets of practices on the date of the transition error and the preceding record date, that is, the proportion of transition error eyes that may be accounted for due to mismatched diagnoses between practices.

No human subjects were included in this study. The IRIS Registry is a centralized data repository and reporting tool that can be used for research purposes. This does not constitute human subject research because data in the IRIS Registry are deidentified and the investigator does not have access to study identifiers. Therefore, institutional board review and informed consent are not required. This study adheres to the Declaration of Helsinki. No animal subjects were included in this study.

Data Extraction

We retrieved all records for CS using Current Procedural Terminology (CPT) codes (CPT = 66982, 66983, 66984) and YAG capsulotomy (CPT = 66821) with an entry date prior to April 14, 2023, the database freeze date, excluding CS (n = 5 192 693) and YAG capsulotomy (n = 1 647 100) entries with unspecified laterality. The retrieval of diagnostic records used the International Classification of Diseases, 10th Revision codes for AMD (International Classification of Diseases, 10th Revision = H35.31 [1-3, 9]∗, H35.32∗) and DR (International Classification of Diseases, 10th Revision = E11.3 [2-5]∗) records, and excluded patients with AMD (n = 592 482) and DR (n = 1 002 828) with unspecified laterality records for the same diagnosis in their medical history. We determined the earliest record per patient, retaining only those with initial entries between 2013 and 2022 to ensure a minimum 1-year observation period and data availability, see Figure S1, available at www.ophthalmologyscience.org. Bilateral records were split into left and right counterparts for tracking purposes prior to exclusions for diagnoses and postexclusions for treatments; see Figure S2, available at www.ophthalmologyscience.org and corresponding Tables S1-S2 for details.

Duplication and Error Metrics

We identified 2 main types of data errors: duplications for procedures, defined as eyes having >1 record, and transition errors for diagnoses, defined as eyes identified with a less severe stage of AMD or DR following a diagnosis of the more severe stage, with subclassifications as described below.

For treatments, we labeled (1) eyes having >1 record confined to the same date as the first entry (same day duplication, Ds), (2) eyes having ≥1 record on a date after the first entry (different date duplication, Dd), and (3) eyes having >1 record (overall duplication, Do) encompassing both Dd and Ds. We calculated the following measures:

Ds=no.eyeshaving>1recordonthefirstdateofentrytotalno.eyes
Dd=no.eyeshaving>1recordbeyondfirstentrytotalno.eyes
Do=no.eyeshaving>1recordtotalno.eyes
errorrate=1totalno.distincteyestotalno.recordsaftersplittingbilateralentries

We defined 2 types of transition errors with 3 metrics for AMD and DR. Less severe stages of AMD and DR (nonexudative and nonproliferative, respectively) were labeled as stage 1, with more severe stages (exudative and proliferative, respectively) as stage 2. We defined reversion error (Er) as eyes with a starting diagnosis at stage 2 and then reverting to stage 1 (2 → 1) normalized by the total number of eyes with a starting diagnosis at stage 2, postprogression error (E p) as eyes with a starting diagnosis at stage 1, progression to stage 2, then reverting to stage 1 (1 → 2→1) at least once in their diagnostic history, normalized by the total number of eyes with progressions from an initial diagnosis of stage 1, and overall transition error (Eo) as eyes having any transition error normalized by the total number of eyes to have reached stage 2.

Er=no.eyeswith21no.eyesstartingatstage2
Ep=no.eyeswith121no.eyesstartingat1andprogressto2
Eo=no.eyeswithanytransitionerrorno.eyesreachingstage2

We quantified eyes with transition errors having disjoint sets of practices on all records of the date prior to the transition error and all transition error records sharing the transition error date in the following manner: for each transition error per eye having at least one transition error, we (1) gathered the set of practice identifiers (IDs) corresponding to transition error records sharing the same date, (2) gathered the set of practice IDs corresponding to records sharing the closest visit date preceding the transition error date, (3) took the intersection of those 2 sets, and (4) marked the eye as a multipractice error if the resulting intersect operation returned an empty set for any of the transition errors present in that eye’s history. We emphasize that eyes counted as having a multipractice error may in fact only have a single case amongst numerous transition errors.

Modeling

We sought any predictive factors in a patient's and practice’s metadata with a given eye that could indicate the likelihood of a Dd or transition error. If errors occurred at random, we could expect model performance similar to random guessing at the overall sample duplication and transition error rates. Within the IRIS Registry, patient demographics have the least number of missing entries while also relating to other factors such as the medical practice via patient geographic location (state or territory).4 Practice- and provider-related metadata are limited to practice type, provider specialty, provider subspecialty, and practice or provider zip code. Provider metadata are not linked to patient condition data in the IRIS Registry version that was analyzed (Chicago_amc_2023_04_14). In the latest version of the IRIS Registry (released May 2025), provider metadata are linked to both patient condition and procedure data. Provider subspecialty is expanded allowing for up to 3 entries per provider. Electronic health record type or software remains unavailable in all records with the exception of visual acuity (VA) measures in the May 2025 release. In this analysis, we modeled the relationship between eye-level demographics, initial practice and provider information, and Dds as well as transition errors using Light Gradient Boosting Machine (LightGBM), a state-of-the-art machine learning framework. We worked at the eye level rather than record level to yield results relevant to common IRIS Registry–based study designs. We chose LightGBM for 3 main reasons: (1) its flexibility in handling nonlinear decision boundaries; (2) superior performance with high dimensionality, categorical variables, and tabular data; and (3) runtime speed on large data sets.5

To form our data set, we retrieved demographics (patient age, race and ethnicity, insurance status, laterality, sex, smoking status, and geolocation information) based on the date of the earliest record per eye in each diagnostic or treatment category. We excluded birth year, division, region, and zip3 (first 3 digits of zip code) from modeling because of more detailed information in other variables and excluded marital status due to a high degree of missingness (∼87% missing in the IRIS Registry). While provider data were available for CS and YAG capsulotomy records, they were not linked to condition entries. We used the practice ID, practice type, provider ID, provider specialty, and provider subspecialty as covariates for modeling when available, omitting zip code in favor of the geolocation information available at the patient-level. For eyes with multiple practices or providers on their first condition or procedure entry, we took the minimum of the practice and provider identifiers present (integer data type) and created an indicator variable to mark the property (multiple identifiers = 1, 0 otherwise). We created two data sets, combining data within the treatment and diagnosis categories, and selected subsets based on unique patient IDs to reduce the correlation between data points (Ntreatments=10,101,947; Ndiagnoses=1,439,425). Each data set was split into training and validation sets at a proportion of 0.75 and 0.25.

We used 3 sets of covariates in treatment data-based model training: (1) using only eye-level demographics; (2) using eye-level demographics and practice metadata; and (3) using eye-level demographics, practice metadata, and provider metadata. Diagnosis data-based models used covariate sets (1) and (2) only. We trained all models with and without random undersampling to mitigate class imbalance, varying hyperparameters to control model complexity and prevent overfitting (number of leaves, regularization [l1, l2], and number of iterations [trees]). We chose the best-performing models based on the F1 score of the target class—eyes with different date duplicates (treatments) or transition errors (diagnoses). F1 score represents the harmonic mean of precision (positive predictive value) and recall (true positive rate) and is often used in binary classification problems. We evaluated permutation feature importance on the validation sets with 10 repeats, assessing the performance loss in the model after destroying a feature’s information through permutation. Permutation feature importance is a model-agnostic technique measuring the contribution of each feature to a model’s performance. The greater the loss to an evaluation metric after the permutation of a variable, the more influential that variable was toward the model’s predictions. In interpreting permutation feature importance, we acknowledge its model-specific quality. Although it does not suffice for causality or prove generalization for a feature’s importance, it can demonstrate the relative strength of relationships between targets and covariates captured by a particular model.6

Results

Out of the 14 718 896 eyes (8 481 432 patients) treated with CS, 4 541 488 (30.9%) had duplicate records, with 3 735 214 (25.4%) having duplicates exclusively on the eye’s earliest treatment date (Ds) and 806 274 (5.5%) having duplicates beyond the earliest treatment date (Dd). The median age at the first record of treatment was 71 (interquartile range [IQR], 65–76) and 71 (IQR, 66–77) years for eyes having Dd and those without, respectively. Patients were predominantly White (Dd = 72.1%, Ds or no duplication = 73.2%) and of female sex (Dd = 59.5%, Ds or no duplication = 58.6%) (Table S3). For YAG capsulotomy, out of the 5 113 679 eyes (3 519 945 patients), 1 490 140 (29.1%) had duplicate records, with 1 279 932 (25.0%) having Ds and 210 208 (4.1%) having Dd, (Table 4). The median age at the first record of treatment was 72 (IQR, 66–78) and 73 (IQR, 67–79) years for eyes having Dd and those without, respectively (Table S5). Like CS, YAG capsulotomy patients were largely White (Dd = 77.1%, Ds or no duplication = 78.4%) and of female sex (Dd = 63.7%, Ds or no duplication = 62.8%). The median number of days beyond the initial record for the first different date duplicate was 13 (IQR, 1–49) for CS and 29 (IQR, 7–329) for YAG capsulotomy (Table 6, Fig 1). Additional demographic characteristics are in Tables S3 and S5.

Table 4.

Duplication and Errors in Treatments

Procedure Type Error Rate Dd Ds Do
Cataract surgery 0.289 0.055 0.254 0.309
YAG capsulotomy 0.325 0.041 0.250 0.291

Dd = different date duplication; Do = overall duplication; Ds = same date duplication; YAG = neodymium-doped: yttrium aluminum garnet.

Error and Duplicate Proportions of the Total Number of Records or Unique Eyes for Each Surgical Group.

Table 6.

Duplication Occurrence Timing

Procedure Type Duplication Type 25th 50th 75th 95th 99th
Cataract surgery Do 0 0 0 21 448
Dd 1 13 49 841 1899
YAG capsulotomy Do 0 0 0 28 916
Dd 7 29 329 1343 2261

Dd = different date duplication; Do = overall duplication; YAG = neodymium-doped: yttrium aluminum garnet.

Describes the Number of Days Beyond the Initial Procedure Record Where x Percentile of Eyes of the Duplicate Category Have Their First Duplicate Record, For Example 50% of Eyes with Different Date Duplicates in the Cataract Surgical Group Have Their First Different Date Duplicate within 13 days of the Initial Procedure Record Date.

Figure 1.

Figure 1

Treatment duplication rates over observation period. Different date duplication (Dd) and overall duplication (Do) rates recorded as the observation period from an eye’s first treatment extends. Different date duplicates occur within a shorter time range in (A) CS compared to (B) YAG capsulotomy. CS = cataract surgery; YAG = neodymium-doped: yttrium aluminum garnet.

We found 13.6% of 1 348 348 eyes (973 271 patients) diagnosed with exudative AMD had transition errors, with 43.2% of transition error eyes (79 441 out of 184 037) having disjoint practice sets on the record dates of and preceding the transition error. Of eyes with an initial diagnosis of nonexudative AMD and progression to exudative (N = 420 839), 19.6% exhibited postprogression error within their records, whereas 11.0% of eyes beginning with an exudative AMD diagnosis (N = 927 509) included Ers, see Table 7. The median age at the time of the first diagnosis record (exudative or nonexudative) for eyes progressing to or starting with exudative AMD was 80 (IQR, 73–87) and 80 (IQR, 74–86) years for eyes without any transition errors and eyes having transition errors within their records, respectively (Table S8). Most patients were of female sex (no Eo = 61.4%, with Eo = 62.3%) and White (no Eo = 78.4%, with Eo = 85.9%). For DR, 12.7% of 858 351 eyes (480 435 patients) diagnosed with proliferative DR held transition errors within their records (38.9% having disjoint practice sets—42 225/108 679). We discovered 25.1% of 188 626 eyes with progressive DR pathology (stage 1 → 2 transition) displayed postprogression errors, with 9.2% of 669 725 eyes beginning with proliferative DR having Ers (Table 7). The median age at first diagnosis for eyes reaching proliferative DR was 61 (IQR, 52–69) and 63 (IQR, 55–69) for those without and with transition errors, respectively (Table S9). For those without and with transition errors in their history, patients were mainly from the South (44.8% and 44.9%, respectively), White (44.7% and 53.0%, respectively), and nonsmokers (56.6%, 58.0%) (Tables S8-S9).

Table 7.

Transition Errors in Diagnoses

Condition Eo Er Ep
AMD 0.136 0.110 0.196
Diabetic retinopathy 0.127 0.092 0.251

AMD = age-related macular degeneration; Eo = overall transition error; Ep = post-progression error; Er = reversion error.

Reversion Error (Stage 2  → Stage 1), Post-progression Error (Stage 1  → Stage 2  → Stage 1), and Overall Transition Error Rates Observed in AMD and Diabetic Retinopathy Records.

When classifying eyes with Dd vs. those without, the model using patient demographics and practice metadata resulted in the highest performance out of the best-performing models using each covariate set (F1 = 0.309 vs. F1 = 0.265 [patient demographics only] and F1 = 0.302 [patient demographics, practice and provider metadata]). The model did not reliably classify eyes with Dd vs. those without yet performed better than random guessing (F1 = 0.309 vs. F1 = 0.048) (Table 10). The originating practice ID contributed the most to classification performance (loss of 0.230 to F1 score on average by permutation feature importance), see Figure S4. The remaining variables (age, race and ethnicity, insurance, laterality, sex, smoking status, patient geolocation [state or territory], treatment [YAG capsulotomy or CS], practice type, and multiple practices indicator) were not influential (≤0.01 loss to F1 score). Model performance in classifying eyes with transition errors vs. those without was similarly superior to random guessing (F1 = 0.303 vs. F1 = 0.128), with the practice identifier leading in feature importance, decreasing F1 score by 0.062 on average when permuted, see Figure S5. Patient insurance status was next in its impact on predictive performance, decreasing F1 score by 0.026 when permuted. All other variables resulted in an average loss <0.01 to the F1 score.

Table 10.

Model Results

Method Dd vs. Rest
Eo vs. No Transition Error
Precision Recall F1 Score Precision Recall F1 Score
Random 0.048 0.048 0.048 0.128 0.128 0.128
LightGBM 0.194 0.765 0.309 0.197 0.657 0.303

Dd = different date duplication; Eo = overall transition error; LightGBM = Light Gradient Boosting Machine.

Validation Set Precision, Recall, and F1 Scores for Best-Performing Models vs. Random Classifiers.

Randomly selected as the positive class at the observed frequency in the training set.

Discussion

Data quality is typically assessed across 6 key dimensions: accuracy, completeness, consistency, timeliness, validity, and uniqueness.7,8 Ensuring quality is a challenge and widespread problem across industries, with a 2017 survey of 75 executives finding only 3% of data meeting basic quality standards.9 As such, in large data sets spanning multiple organizations and systems like the IRIS Registry, data errors may occur at many levels and are likely inevitable. Hence, it is important to identify the extent of data errors so studies may be designed to maximize accuracy and appropriately convey limitations. Large data sets are created with differing objectives, which may affect their infrastructure and content. For example, although the IRIS Registry has now been used for >100 published papers, it was designed to improve the delivery and quality of ophthalmic care and facilitate participation in the Merit-Based Incentive Payment System. The IRIS Registry is evolving over time to also allow research opportunities. In this study, we examined duplication and transition errors in the IRIS Registry through CS, YAG laser capsulotomy, AMD, and DR records. These errors relate to the quality pillars of uniqueness and accuracy with consistency being a latent issue, which we are unable to investigate further because of the deidentification requirements of the IRIS Registry.

We found a significant duplication rate in treatment records (30.9% CS, 29.1% YAG capsulotomy) with roughly a quarter of eyes having duplicate records occurring the same day as the first procedure (Ds) and 5.5% of CS-receiving eyes and 4.1% of YAG capsulotomy-treated eyes having duplicates on dates beyond the initial surgical record (Dd). We also discovered transition errors in the diagnostic records of 12.7% of eyes reaching proliferative DR and 13.6% of eyes with exudative AMD. Notably, 43.2% of 184 037 AMD transition error eyes and 38.9% of 108 679 DR transition error eyes had at least one transition error instance derived from mismatched practices. Although transition errors may occur multiple times per eye, we chose to create mutually exclusive classifications at the eye level for interpretability. For example, an eye with a sequence of record stages following 2 → 1→1 → 2→1 in fact has two transition errors though we classify it based on the initial stage (Er). For this reason, the proportions of transition error eyes with mismatched practices reported above do not reflect the overall proportions of transition errors (record level) with mismatched practices.

A mismatch of diagnostic records between practices may stem from a variety of reasons. A few possibilities are as follows: (1) a data accuracy issue where one practice persists with an earlier stage diagnosis in their records even after the patient’s condition has progressed; (2) a disagreement between providers; and (3) a data pipeline problem leading to issues such as erroneous carryover or duplicate records of the earlier-stage diagnosis in the case of postprogressive errors. Other possibilities exist, though are beyond our visibility as end users of the data set.

To investigate the duplications and errors in the data set, we used LightGBM, a gradient-boosting framework from Microsoft, due to its ability to natively handle categorical variables and fast computation.3 These models had a slight predictive value, outperforming random guessing. Permutation feature importance revealed that classification performance was most influenced by the originating practice (practice ID) for both the model based on treatment records (loss of 0.230 to F1 score) and that based on diagnoses (loss of 0.062 to F1 score). These practice-level differences could arise from several factors, including but not limited to nonstandardized EHRs, bugs in the extract–transform–load process when consolidating source data from particular EHR implementations, faculty training, and local practice procedures. However, we believe these latter two possibilities to be unlikely in many practices, especially those exhibiting transition errors or Dds in over half of the eyes whose treatment or diagnosis originated at that practice, see examples in Figures S4-S5.

At the practice level, clinicians or billers may introduce inaccuracies themselves. At the EHR level, technical issues like network interruptions, system crashes, and software bugs can lead to data duplication or corruption. Errors may also arise during data ingestion from EHRs to data sets when information is collected, transported, and stored for processing. Ultimately, data integrity faces risks spanning the entire information chain.10 Because the IRIS Registry compiles information from participating practices’ EHR systems, the abovementioned challenges and limitations may be introduced in the data set.

Our study has several key limitations. First, we are unsure whether duplications and other data errors occur at similar rates to those of CS, YAG capsulotomy, DR, and AMD for other procedures and conditions within the IRIS Registry. Second, despite an association between the originating practice of a treatment or diagnosis for a given eye and Dds, as well as transition errors, we cannot suggest the sources of these errors with this data—whether the type of EHR used at each practice or how that data was processed. Of note, newer releases (November 2024; May 2025) of the IRIS Registry continue to enrich record metadata such as with additional provider information for both procedure and condition data as well as encoded EHR type for VA measures. Third, some procedure duplications may appear in the data set due to missing modifiers in the global postoperative period for revisions, but, given the span of Dd (Table 6), this likely does not explain all of the identified changes. The November 2024 and later releases of the IRIS Registry included CPT code modifiers. Different date duplicates persisted in these later versions when excluding records with modifiers for adjustments (modifiers = [58, 78, 55, 76, 77, 24]), albeit at a slightly lower proportion (Dd = 3.4% for CS). Finally, because the IRIS Registry is appropriately deidentified, we do not have access to the original data sets to review a patient’s clinical course for quality assurance.

Stakeholders working with big data can minimize the impact of quality issues on downstream analyses through several techniques. First, researchers may choose to restrict the types of analyses they perform or adjust their methodology accordingly. We offer an example of methodological adjustment in the supplementary material from Vu et al.11 Different date duplications can be particularly challenging for time-to-event analysis, as researchers are responsible for adjudicating the true date of a disease or treatment among several replicates in the medical record. An incorrect date may bias studies and lead to erroneous results, especially if the error is systematic to a practice or provider. Repeated measurement studies, such as those where patients can have multiple medications or treatments, have similar issues as it is difficult to identify which treatments are true and which are duplicates. There is a need for a unified methodology of record selection for standard analyses in which multiplicities are present. While using only the first date of a diagnosis or procedure on record is one possible strategy, this does not allow for research on repeated procedures or relapsing conditions. For mitigation of duplication errors with clinical exam data, such as intraocular pressure and VA, averaging date-coarsened, distinct records over a period of interest can be helpful when possible, using methodology put forth by Brant et al.12 We offer examples in the supplementary materials, with implementations and documentation accessible for the Chicago April 14, 2023, version of the IRIS Registry in the public schema.

With results indicating a weak association between practice ID and transition errors or different date duplicates, end users of the IRIS Registry may consider creating filtered subsets of the database to mitigate these problems. However, excluding records from practices with high rates of these problems may have only a narrow application when working with the entire data set. Error rates are identifiable with YAG capsulotomy, CS, AMD, and DR, but an issue with one of these conditions or procedures does not necessarily indicate problems for others at the same practice. We hypothesize that there may be other underlying factors leading to the data quality issues indicated by the practice ID, such as batch-based extract–transform–load complications on a particular day at a given practice, EHR compatibility problems, or others.

For all large data sets, a more comprehensive approach to identify and resolve errors at their source can be promoted by standardized transparency protocols and rigorous quality reporting frameworks. Utilizing a standardized data model such as the Observational Health Data Sciences and Informatics program’s Common Data Model can facilitate quality through shared community collaboration on challenges and solutions surrounding health care data.13 Tools like Observational Health Data Sciences and Informatics's Data Quality Dashboard implement over 1500 automated checks against 3 core data quality categories proposed by Kahn et al14 in 2016—(1) conformance (value formatting or relationships), (2) completeness (missing data patterns), and (3) plausibility (value distributions or temporal coherence)—and facilitate quality benchmark distribution to stakeholders.15 A minimum transparency standard could be published alongside any large data set release. Data Cards—human-centered documentation of data sets originally proposed for artificial intelligence and machine learning—are another potential tool to standardize reporting.16,17 These cards communicate structured summaries of various aspects of the data set’s lifecycle such as upstream sources, data collection and annotation methods, training and evaluation methods (in the case of machine learning–derived data), intended use, and descriptive statistics, acting as a shared knowledge base. With Data Cards, stakeholders have an overview of how the data were put together, the justifications behind design decisions, and importantly, changes made between versions. Transparency transforms quality assurance into a community asset simplifying tasks such as research feasibility assessments and data gap identification. Longitudinal tracking of data quality metrics across data set versions (such as via Data Card versioning) would create quality improvement feedback loops—a practice well-established in clinical trial reporting but absent in observational data pipelines.

In conclusion, Dd and transition errors are likely findings in large data sets, as evidenced by their presence in the IRIS Registry. Extra care should be employed when conducting time-to-event studies, analyses relying upon repeated procedures, and those examining relapsing conditions when duplication is identified. There has continued to be improvement with maturation and subsequent versions of the IRIS Registry. Our work simply demonstrates the importance of identifying and considering these issues in research with any large medical data set. Transparency and communication surrounding data processing, tests, and transformations will broaden confidence in downstream studies and the understanding of their limitations. Large-scale data sets are the foundation to furthering scientific understanding of populational health and building more effective health care systems. We must understand their limitations and work toward resolutions such that high-quality data yield high-quality results.

Manuscript no. XOPS-D-24-00410.

Footnotes

Suzann Pershing, MD (Stanford University, Palo Alto, California); Leslie Hyman, PhD (Wills Eye Hospital, Philadelphia, Pennsylvania); Julia A. Haller, MD (Wills Eye Hospital, Philadelphia, Pennsylvania); Aaron Y. Lee, MD, MSCI (eScience Institute, University of Washington, Seattle, WA, USA; Department of Ophthalmology, University of Washington, Seattle, WA, USA); Cecilia S. Lee, MD, MS (Department of Ophthalmology, University of Washington, Seattle, WA, USA); Flora Lum, MD (American Academy of Ophthalmology, San Francisco, CA, USA); Joan W. Miller, MD (Massachusetts Eye and Ear, Harvard Medical School, Boston, Massachusetts); Alice Lorch, MD, MPH (Massachusetts Eye and Ear, Harvard Medical School, Boston, Massachusetts).

This study was presented at the American Academy of Ophthalmology Annual Meeting, October 18-21, 2024, Chicago, Illinois.

Disclosure(s):

All authors have completed and submitted the ICMJE disclosures form.

The authors made the following disclosures:

T.E.: Consultant – University of Michigan.

J.W.M.: Grants – Lowy Medical Research Institute, Ltd; Mactel Study (no PI salary); Royalties – QLT, Bausch and Lomb, Mass Eye and Ear; Consultant – Sumitomo Pharma America, Inc, ONL Therapeutics, LLC; Honoraria – Connecticut Society of Eye Physicians, Atlantic Coast Retina Conference/Macula 2022, NYU/Langone; Travel expenses – Portuguese Society of Ophthalmology Annual National Congress, Nova Scotia Health Authority; Patents planned, issued or pending – US 7 811 832; US 5 798 349; US 6 225 303; US 6 610 679; CA 2 185 644; CA 2 536 069; Leadership or fiduciary role in other board, society, committee or advocacy group, paid or unpaid – Foundation of the Massachusetts Eye and Ear Infirmary, Massachusetts Eye and Ear Associates, Inc., Aptinyx, Inc., Association of University Professors in Ophthalmology (AUPO), Heed Ophthalmic Foundation, Macula Society, Ophthalmology, Ophthalmology Retina, Drusolv Therapeutics, Harvard Health Publishing; Stock or stock options – Aptinynx, Inc., ONL Therapeutics, LLC, Ciendias Bio.

Financial support was provided by the Massachusetts Eye and Ear Clinical Data Science Fund, Boston, Massachusetts. The sponsor or funding organization had no role in the design or conduct of this research.

Support for Open Access publication was provided by Massachusetts Eye and Ear.

HUMAN SUBJECTS: No human subjects were included in this study. The IRIS Registry is deidentified and the investigator does not have access to study identifiers. Therefore, institutional review board review and informed consent are not required. All research adhered to the tenets of the Declaration of Helsinki.

No animal subjects were used in this study.

Author Contributions:

Conception and design: Goldberg, Ross, Douglas, Elze, Miller, Lorch

Data collection: Goldberg

Analysis and interpretation: Goldberg, Ross, Douglas, Ivanov, Lorch

Obtained funding: Lorch, Elze, Miller

Overall responsibility: Goldberg, Ross, Douglas, Ivanov, Lorch, Elze, Miller

Supplemental material available at www.ophthalmologyscience.org.

Contributor Information

Eric A. Goldberg, Email: egoldberg8@mgb.org.

IRIS Registry Analytic Center Consortium:

Suzann Pershing, Leslie Hyman, Julia A. Haller, Aaron Y. Lee, Cecilia S. Lee, Flora Lum, Joan W. Miller, and Alice Lorch

Supplementary Data

Figure S6
mmc1.pdf (85.5KB, pdf)
Table S1
mmc2.pdf (153.1KB, pdf)
Table S2
mmc3.pdf (132KB, pdf)
Table S3
mmc4.pdf (187KB, pdf)
Table S5
mmc5.pdf (187.2KB, pdf)
Table S8
mmc6.pdf (184.7KB, pdf)
Table S9
mmc7.pdf (184.8KB, pdf)
Figure S1
mmc8.pdf (72.3KB, pdf)
Figure S5
mmc9.pdf (683.4KB, pdf)
Figure S2
mmc10.pdf (92.7KB, pdf)
Figure S4
mmc11.pdf (53.6KB, pdf)
Supplement
mmc12.pdf (94.8KB, pdf)

References

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Figure S6
mmc1.pdf (85.5KB, pdf)
Table S1
mmc2.pdf (153.1KB, pdf)
Table S2
mmc3.pdf (132KB, pdf)
Table S3
mmc4.pdf (187KB, pdf)
Table S5
mmc5.pdf (187.2KB, pdf)
Table S8
mmc6.pdf (184.7KB, pdf)
Table S9
mmc7.pdf (184.8KB, pdf)
Figure S1
mmc8.pdf (72.3KB, pdf)
Figure S5
mmc9.pdf (683.4KB, pdf)
Figure S2
mmc10.pdf (92.7KB, pdf)
Figure S4
mmc11.pdf (53.6KB, pdf)
Supplement
mmc12.pdf (94.8KB, pdf)

Articles from Ophthalmology Science are provided here courtesy of Elsevier

RESOURCES