Summary
Sharing clinical research data is key for increasing the pace of medical discoveries that improve human health. However, concern about study participants' privacy, confidentiality, and safety is a major factor that deters researchers from openly sharing clinical data even after deidentification. This concern is further enhanced by the evolution of artificial intelligence (AI) approaches that pose an ever-increasing threat to the reidentification of study participants. Here, we discuss the challenges AI approaches create that are blurring the lines between identifiable, and non-identifiable data. We present a concept of pseudo-reidentification, and discuss how these challenges provide opportunities for rethinking open data sharing practices in clinical research. We highlight the novel open data sharing approach we have established as part of the AI-READI (Artificial Intelligence Ready, and Exploratory Atlas for Diabetes Insights) project, one of the four Data Generation Projects funded by the National Institutes of Health Common Fund's Bridge2AI Program.
Keywords: Clinical study, Data sharing, Data access, Data reuse, Artificial intelligence, Machine learning
Search strategy and selection criteria.
References for this Review were identified through searches of PubMed, arXiv, and IEEE (Institute of Electrical, and Electronics Engineers) archive with the search terms “data sharing”, “reidentification”, “pseudo-reidentification”, “research participant”, “privacy”, “biometric”, “retinal imaging”, “ECG”, “wearable fitness tracking”, and “continuous glucose monitoring” from 2000 until April, 2025. Articles were also identified through searches of the authors’ files. Only papers published in English were reviewed. The final reference list was generated based on originality, and relevance to the broad scope of this review.
Introduction
Sharing scientific research data is a cornerstone for advancing science, and accelerating discoveries. Clinical research is a prime illustration where broad data sharing can rapidly drive innovations that benefit patient care. For instance, the widespread sharing of genomic data has enabled rapid advancement in cancer research.1 During the COVID-19 pandemic, rapid data sharing enabled a quick understanding of the virus.2 Clinical data sharing efforts have particularly grown over the past decade, driven by the rise of initiatives such as the Findable, Accessible, Interoperable, Reusable (FAIR) principles.3
Despite these efforts, persistent challenges hinder effective data sharing, and reuse. Researchers, and institutions hesitate to make data openly available to external investigators.4, 5, 6 Concerns about participant privacy, data security, and potential misuse limit the sharing of valuable datasets.7 Study participants also have similar concerns. In a 2024 nationwide online survey of adults in the United States, a majority of respondents felt relatively at ease sharing data with healthcare providers.8 The study also highlighted a critical tradeoff between the push for open science to improve clinical outcomes, and public health, and the need to honor patient privacy, autonomy, and trust in data collection, and use. Enhancing the transparency, and governance of data sharing practices is necessary to maintain participant confidence, and willingness to contribute to research.8 Moreover, the advent of increasingly sophisticated artificial intelligence (AI) tools exacerbates the reidentification risk, blurring the line between identifiable, and non-identifiable data types, and raising questions about participant privacy.9
In this paper, we examine the evolving challenges, focusing on the privacy risks introduced by AI. We discuss examples of data types currently not deemed identifiable, in which, through AI-driven analysis, uniqueness can be inferred without actual reidentification, a process we term “pseudo-reidentification”. Finally, we introduce a novel data sharing approach developed within the AI-READI (Artificial Intelligence Ready, and Exploratory Atlas for Diabetes Insights) project, part of the National Institutes of Health (NIH) Common Fund's Bridge2AI Program,10,11 aiming to safeguard participant privacy while preserving the spirit of openness, and collaboration that propels clinical research forward.
Regulatory landscapes for data protection, and privacy
HIPAA, PII, and PHI
Personal identifiable information (PII) refers to any information that links to an individual (Table 1).12,14 A protected health information (PHI) is collected for the provision of healthcare services, and protected by the Health Insurance Portability, and Accountability Act (HIPAA) of 1996.15 The term “deidentification” originates from the HIPAA, which involves the identification, and removal of PHI from data. According to the HIPAA Privacy Rule, if health information is deidentified, it is not considered PHI, and deidentified datasets may be shared more easily. While HIPAA does not regulate deidentified data, other ethical, legal, or institutional constraints may still apply. A dataset can be deidentified under HIPAA following one of two methods: Safe Harbor, or Expert Determination. The Safe Harbor method requires 18 health information elements (listed in Table 2) to be removed from the dataset.16 In Expert determination method, an expert certifies that the risk of reidentification is minimal, regardless of the specific method used.17 In some cases, however, the use, and disclosure of PHI are needed for research, public health, or healthcare operations. In such cases, datasets with PHI can be shared as “limited datasets” with proper restrictions, and security measures, including a data use agreement between the data provider, and the data user.18 HIPAA was designed as flexible guidance rather than strict regulation, enabling it to adapt over time as technology evolves. As a result, the determination of what constitutes PHI (Table 2) can vary among institutions, and organizations.
Table 1.
Glossary of major terms, and concepts relevant in this work.
| Definition | |
|---|---|
| Pseudonymization | The process of replacing private identifiers with fake identifiers, or pseudonyms to protect an individual's identity while retaining data utility. It allows data to be reidentified if necessary using a reidentification key. |
| Anonymization | The process of irreversibally deidentifying data elements, following GDPRb guidelines. |
| Quasi Identifiers | Variables in a research dataset that can not individually identify a participant but, in combination with other variables, identify a record, or participant. |
| Reidentification | The process of matching anonymized, or pseudonymized data with other information, such as a reidentification key, patient ID, publicly available information, and/or other datasets to reestablish the identity of an individual. |
| Deidentification | The process of removing, or altering personal identifiers from data so that the individuals to whom the data pertains cannot be readily identified. Deidentification is often used to maintain privacy in datasets used for research, and analysis. |
| Pseudo-reidentification | The process by which AI, or analytical methods detect unique patterns in deidentified data that suggest individuality without directly linking to a specific person, unlike traditional reidentification, which requires external identifiers, or reference datasets. |
| PII (Personally Identifiable Information) | Any data that could potentially identify a specific individual, including but not limited to names, social security numbers, addresses, phone numbers, and email addresses. |
| PHI (Protected Health Information) | Any health-related information that can be linked to an individual, and is protected under regulations such as HIPAA. PHI includes medical records, insurance information, and other personal health data. |
| HIPAA Covered Entity | Any person, or organization that is authorized to collect, use, and transmit PHI in accordance with HIPAAa regulations. |
Table 2.
List of the 18 HIPAA safe harbor identifiers.
| HIPAA identifiers |
|---|
| Names |
All geographic subdivisions smaller than a State, including street address, city, county, precinct, zip code, and their equivalent geocodes, except for the initial three digits of a zip code if, according to the current publicly available data from the Bureau of the Census:
|
| All elements of dates (except year) for dates directly related to an individual, including birth date, admission date, discharge date, date of death; and all age over 89, and all elements of dates (including year) indicative of such age, except that such ages, and elements may be aggregated into a single category of age 90, or older |
| Telephone numbers |
| Fax numbers |
| Electronic mail addresses |
| Social security numbers |
| Medical record numbers |
| Health plan beneficiary numbers |
| Account numbers |
| Certificate/license numbers |
| Vehicle identifiers, and serial numbers, including license plate numbers |
| Device identifiers, and serial numbers |
| Web Universal Resource Locators (URLs) |
| Internet Protocol address numbers |
| Biometric identifiers, including finger, and voice prints |
| Full-face photographic images, and any comparable images |
| Any other unique identifying number, characteristic, or code |
General Data Protection Regulation (GDPR)
In the European Union (EU), personal data are regulated by the General Data Protection Regulation (GDPR), which came into effect in 2018.19 The GDPR is widely regarded as one of the strictest data protection frameworks globally, and is considered highly “data subject-centric,” protecting the privacy rights of all individuals in the EU-not just patients. The GDPR applies to any organization, whether inside, or outside the EU, that offers goods, or services to, or monitors the behavior of, individuals located in the EU.20 The regulation requires data controllers, and processors to implement robust safeguards to ensure privacy, such as data pseudonymization, or encryption.20 It distinguishes pseudonymization from anonymization (Table 1).21 Pseudonymized data is still considered personal data under the GDPR, and remains subject to its requirements. In contrast, once data are truly anonymized, they are no longer regulated by the GDPR. Under the GDPR, any information that can directly, or indirectly identify an individual, including biometric data, is classified as personal data. Processing personal data is only permitted if there is a lawful basis, such as explicit consent, performance of a contract, compliance with a legal obligation, protection of vital interests, the performance of a task carried out in the public interest, or legitimate interests pursued by the controller, or a third party. The consent must be freely given, specific, informed, and unambiguous, and individuals must be able to withdraw consent at any time. In addition, participants have the right to data portability, allowing individuals to receive their data in a commonly used, machine-readable format, and to transmit that data to another data controller.22 When third parties process data on behalf of a controller, data processing agreements must be established to ensure compliance.23 The approach to de-identification, and secondary use of data may vary across jurisdictions, but the GDPR sets a high standard for privacy, and data security.
Pseudo-reidentification in the era of AI
Reidentification vs. pseudo-reidentification
Traditional reidentification involves linking deidentified data back to a known individual by leveraging external identifiers, or datasets. In studies using deidentified open-source databases, only the primary investigators (PI) who collected the data may possess access to the identifiers linking these records to PHI. Quasi-identifiers (QI) may also put research participants at risk. QIs are data elements that do not directly identify participants but can be used for reidentification when linked to other sources of information.24 Traditional examples of QIs include: birth weight, behavioral data, sex, profession, total income, minority status, locations, spoken languages, ethnicity, education, marital status, criminal history, disability, dates, codes (e.g., diagnosis, procedure, or adverse event codes), and birth plurality.24 When a certain combination of QIs only appears once in a dataset, the record is considered “sample unique” within the dataset or “population unique” in the broader population. These concepts represent dataset-dependent measures of reidentification. While HIPAA's Safe Harbor method removes many identifiers, and some QIs, some QIs may remain in deidentified datasets. Linking QIs to external information that contains PHI may also lead to reidentification.24 Despite advancements in health cybersecurity, and infrastructure, the risk of malicious attacks exposing these identifiers remains a concern.25 While such breaches could expose sensitive patient data, the overall risk remains lower than breaches involving data stored by HIPAA-covered entities and their business associates (e.g., insurance companies, healthcare systems, or EHR providers). Moreover, the likelihood of reidentification by third parties (i.e., other than covered entities, or project PIs) is minimal.26
However, with the advancement of AI techniques, certain modalities that have not traditionally been considered PHI, identifiers, or QIs (e.g., ECG) are now being proposed as potential biometric markers, since they alone can act as a biometric identifier (Table 3) and identify an individual independent of QIs or any tabular data, raising novel concerns about sharing these data types. In contrast, sample and population uniqueness rely on QI (e.g., ZIP code, age, sex) within a tabular dataset, where the risk of reidentification arises only from unique combinations of these variables. AI can identify unique patterns in data elements that, while not explicitly tied to an individual, still imply uniqueness. For example, AI models may detect unique structural, physiological, or behavioral patterns that, while not directly tied to a name, or ID, are unique enough to single out an individual in a dataset.50 We introduce the term ‘pseudo-reidentification’ to describe this identification of unique data patterns. Pseudo-reidentification refers to the identification of unique data patterns that, while not directly linked to an individual, could enable reidentification if external identifiers become available. That do not directly link to an individual, but could be linked if identifiers become available (Fig. 1). The only step between pseudo-reidentification, and reidentification is the linkage of the pseudo-reidentified data elements to external identifiers (e.g., PHI, or PII). While this may be unlikely if only PIs have access to identifiers, pseudo-reidentification remains a concern as technological advances may make this linkage more feasible. These evolving risks challenge current regulatory frameworks, which may not fully account for pseudo-reidentification. Below, we review studies assessing pseudo-reidentification for commonly collected data types in clinical research, and traditionally not considered PHI (Table 3).
Table 3.
A highlight of the recent publications on the approaches to pseudo-reidentification using health data.
| Year | Database | Method | Accuracy (%) | |
|---|---|---|---|---|
| Wearable fitness-tracking devices | ||||
| Spadaccini, et al.27 | 2013 | HSCT-11 | GMM | 86.4 |
| Sancho, et al.28 | 2018 | MIMIC II PRRB |
L2 distance | 78.5 |
| Labati, et al.29 | 2020 | PRRB | SVM | 94.8 |
| Retsinas, et al.30 | 2020 | PersonID | CNN | 55.8 |
| Lee, et al.31 | 2020 | IEEEPPG | CNN | 95.7 |
| Hwang, et al.32 | 2021 | BioSec | CNN + RNN | 87.1 |
| Yadav, et al.33 | 2021 | BioSec DEAP |
LDA | 97.4 |
| Continuous glucose monitoring devices | ||||
| Herrero, et al.34 | 2021 | REPLACE-BG | SVM | 86.8 |
| Electrocardiogram | ||||
| Tan, et al.35 | 2017 | MIT-BIH PhysioNet Mobile ECG |
A two-stage classifier integrating random forest and wavelet distance measure with a probabilistic threshold | 99.52 |
| Arnau-González, et al.36 | 2017 | DREAMER | CNN | 94 |
| Zhao, et al.37 | 2018 | ECG-ID PhysioNet |
Generalized S-transformation with CNN | 99 |
| Patro, et al.38 | 2019 | MIT-BIH ECG-ID |
Feature Extraction, LASSO, KNN | 99.1 |
| Patro, et al.39 | 2020 | PhysioNet ECG-ID PTBDB |
Optimized Feature Selection (GA, PSO, LASSO, EN) with RF | 94.9–95.3 |
| El Boujnouni, et al.40 | 2021 | PTB MIT-BIH |
A combination of CWT, DWT, along with a capsule network | 98.1–100 |
| Parkash, et al.41 | 2022 | ECG-ID PTB CYBHi UofTDB |
A deep learning algorithm based on CNN, and LSTM | 98.2 |
| Parkash, et al.42 | 2023 | ECG-ID | Deep learning | 99.9 |
| Wang, et al.43 | 2023 | ECG-ID MIT-BIH USSTDB |
ECG Feature Vector with Pooling Layer for Variable-Length Signals | 91–97.6 |
| Retinal images | ||||
| Farzin, et al.44 | 2008 | DRIVE STARE |
Blood vessel segmentation, Feature generation, Feature matching | 99 |
| Köse, et al.45 | 2011 | Local data STARE |
Vessel segmentation | 95 |
| Sadikoglu, et al.46 | 2016 | DRIVE | CNN | 97.5 |
| Szymkowski, et al.47 | 2020 | Local data DRIVE STARE Kaggle Retinopathy |
KNN SVM CNN |
96.5 |
| Devi, et al.46,48 | 2022 | VARIA | ANFIS | 97.2 |
| Marappan, et al.49 | 2023 | RIDB VARIA DRIVE STARE |
Multiple feature extraction | 98 |
MIMIC-II: Multiparameter Intelligent Monitoring in Intensive Care II, PRRB: Photoplethysmography Respiratory Rate Benchmark, SVM: Support Vector Machine, CNN: Convolutional Neural Network, IEEEPPG: Institute of Electrical, and Electronic Engineers Photopletysmographic Signals Dataset, RNN: Recurrent Neural Network, DEAP: Database for Emotion Analysis using Physiological Signals, LDA: Linear Discriminant Analysis, HSCT-11: Heart Sounds Catania 2011, GMM: Gaussian Mixture Models, MIT-BIH: MIT–Beath Israel Hospital dataset, ECG-ID: ECG Identification dataset, KNN: K-nearest neighbor, LASSO: least absolute shrinkage, and selection operator, NSRDB: normal sinus rhythm database, STDB: ST change database, PTB: Physikalisch-Technische Bunde-sanstalt, CYBHi: check your bio-signals here initiative, UofTDB: the University of Toronto Database, LSTM: long short term memory, CWT: Continuous Wavelet Transform, DWT: Discrete Wavelet Transform, RF: random forest, EN: elastic net, GA: genetic algorithm, PSO: particle swarm optimization, ANFIS: Adaptive network-based fuzzy inference system. Accuracy: the proportion of all correct classifications, whether positive or negative.
Fig. 1.
Illustration of pseudo-reidentification vs. reidentification. AI: artificial intelligence, PII: personal identifier information.
Practical insights into pseudo-reidentification by modality
Wearable fitness-tracking
Fitness-tracking devices record signals such as heart rate (including variability, and pattern), step count, gait (including fixed time durations, step cycles, and walk cycles), metabolic equivalent of task, energy expenditure, and exercise parameters recorded via the global positioning system (GPS), and accelerometer.51 These wearables use biometric sensors to continuously monitor physiological signals, leveraging each individual's unique baseline to enable accurate authentication and early detection of health changes or unusual health activity. Therefore, it can perform real-time detection of signals in a non-invasive way, making the data acquisition convenient. However, considering the unique individual activity habits and constant monitoring of data elements including photoplethysmography (which detects blood volume changes in the microvascular bed of tissue), heart sounds, movement patterns, and heart rate raise concerns about the reidentifiability of this data type.52 Researchers applied various machine learning methods, such as support vector machines (SVMs), and random forests (RFs), neural networks, and deep learning (DL) strategies.28 All studies reported high accuracy from deidentified wearables information, noting that pseudo-reidentification is possible with small data fragments.53
Continuous glucose monitoring
Continuous glucose monitors (CGMs) track blood glucose levels, enabling improved monitoring and informed diabetes management.54 CGMs allow prompt detection of glycemic changes during acute stress by notifying the user, or healthcare provider, ensuring proactive management of such events.55 CGMs generate a substantial amount of data, which is synchronized, stored, and shared across different platforms. Deidentified CGM data from multiple study groups are accessible for secondary use,56 and are not considered PHI, despite significant cybersecurity implications.34 Some manufacturers may gather PIIs, such as the user's internet protocol (IP) address, network accessibility, internet service, browser, and their activities, which can eventually lead to reidentification.55 Although manufacturers claim to deidentify CGM data, how they perform deidentification is not usually mentioned. This raises privacy, and security concerns for CGM users. CGM data can be used to pseudo-reidentify individuals using ML algorithms (Table 3).34 Reported accuracy may reach as high as 86% in pseudo-reidentifying CGM users. With the growing number of patients with diabetes wearing CGMs, as well as CGM manufacturers, privacy concerns of consumers are increasing. Recognizing this risk, the Institute of Electrical, and Electronics Engineers Standards Association published standards to help stakeholders develop more secure wireless diabetes devices.57
Electrocardiogram
Electrocardiogram (ECG) data is a record of the electrical signals generated by cardiac rhythm, and activity. ECG data is not considered PHI, and several deidentified datasets are publicly available online, with the rationale that their linkage to PII, or PHI is unlikely.58 However, studies reporting on biometric recognition using ECG date back to 2001,59 when Biel et al. applied soft independent modeling by class analogy (SIMCA) on features extracted from 12-lead ECG records to link subsequent ECGs taken in the same visit with their baseline ECG. Nowadays, off-the-person devices (i.e., wearable devices)60 may also be used for this purpose, in addition to the classic on-the-skin (i.e., on-the-person) 12-lead ECG. The accuracy can range between 75 and 100%,61 depending on the used device, test duration, and test intervals. Similar to advancements in databases, and hardware, analytical approaches have evolved. Earlier approaches included manual feature extraction,59 and principal component analysis. Recently, more studies report AI applications,62 including DL-based pseudo-reidentification via uni-, or multimodal modeling63 of ECG along with biometrics such as fingerprints. All studies have used publicly available databases, and PHI is not available in any of the mentioned databases.
Retinal imaging
Retinal imaging involves capturing images of the retina, the light-sensitive tissue at the back of the eye, using advanced technologies such as optical coherence tomography, color fundus photography, and OCT angiography. These techniques provide high-resolution, cross-sectional, or wide-field views, aiding the diagnosis, and management of various retinal conditions. Several publicly accessible retinal imaging datasets are available online.64 The distinct vascular patterns in retinal scans may act as a potential identifier for individuals, posing privacy concerns. The published body of the literature suggests that retinal images are pseudo-reidentifiable (Table 3). The accuracy of pseudoreidentification using retinal images is similar to the previously discussed modalities, ranging from 95 to 99%, suggesting that uniqueness alone can be detected, but can not be linked to PHI without external identifiers. The American Academy of Ophthalmology advises against the consideration of retinal images as biometric identifiers for clinical research.65 Unlike established biometric identifiers, such as fingerprints, or iris scans, retinal imaging quality varies significantly due to differences in equipment, technique, and patient conditions. Moreover, the features detected in retinal images are not static; they can change over time due to aging, disease progression, or treatment, further complicating their reliability as a stable individual identifier.47
Hidden pathways to reidentification: navigating modern data misuse
Although the likelihood of identifying an individual solely from the data types described above may seem relatively low, it is not negligible. The mentioned data elements are not traditionally known as identifiers or QIs. However, their linkage to external sources of information containing PHI may lead to actual reidentification. Advances in AI-driven analytics mean that even fragments of non-traditional biometric data, such as aggregated wearable metrics, ECG signals, or retinal patterns, could be leveraged for malicious purposes.66 Ill-intended individuals may devote significant time, and resources to parsing the web, obtaining additional context from social media, online medical forums, or leaked datasets, and combining these disparate elements to infer a person's identity, sensitive data, or health status.67 These priorities are not irreconcilable, but require tiered, auditable, and governed access protocols with explicit awareness that each reuse incrementally draws down a finite privacy reserve.68 Further, digital health companies are dedicating increasing resources to acquire real-world data from different populations. Data leak, misconduct, or reidentification attempts by companies that already have access to PII can be another form of data misuse.69 Moreover, as data-sharing practices expand, the inevitable presence of data brokers, dishonest third parties, or data enthusiasts increases the risk that reassembled fragments of deidentified information could be used to discriminate, stigmatize, or exploit individuals.70 In other words, the risks, though not prominent, are real enough to demand thoughtful data protection, governance, and oversight.23
Limitations of current data sharing methods
In light of these new risks, it is necessary to reexamine existing data sharing methods. These methods can be broadly classified into three categories: 1) Open sharing relying on deidentification, 2) Controlled access, and 3) Enclave-based access.64
Open data sharing presents the simplest way to maximize data accessibility. Researchers may, or may not be required to register, and share their information to access the data.71 Data is shared under a data reuse license that is typically permissive, such as the Creative Commons Attribution 4.0 International (CC-BY 4.0), which allows reuse for any purpose. Data access is often as easy as clicking a “Download” button. Examples include autonomic nervous system-related datasets available on the NIH SPARC program's repository, and neuroimaging datasets available on the OpenNeuro repository.72,73 They rely on the researchers sharing the data to make sure it does not contain any PHI, through deidentification. These methods may have limited protection against data misuse, as there is little legal framework for tracking, and reinforcement compared to more controlled methods. Depending on the strength of the legal framework, data misuse in this scenario may still have reputational, or legal consequences.
Controlled access methods typically require submitting a data access application, which is reviewed by a committee. Controlled access includes centralized, and decentralized (or federated) approaches. Centralized approaches usually include management of data from multiple sources in a single, centralized repository, and then granting access to the users. In contrast, decentralized approaches distribute data management across multiple sources, each handling its requests, and agreements. This can increase agility, enable more local control, and may foster higher data sharing rates, and faster research outputs, but may require more resources, and coordination, and can introduce inconsistencies in access procedures.74 If access is requested for PHI elements, a Data Usage Agreement (DUA) is established between the sharing, and receiving entities. Accessors may be required to pay for registration, and data access to support the sustainability of the data sharing approach.71 Examples of controlled access methods include datasets from the dbGaP repository, and the UK Biobank.75,76 Controlled access may be viewed as a necessary compromise to protect participants’ privacy, still enabling data reuse.77 While reducing the risk of data misuse through vetting data accessors, these methods may go against the spirit of open science as they could exclude certain individuals from accessing the data (e.g., those not affiliated with a trusted institution). It can also delay data access, and present sustainability risks if the data access committee is unable, or unwilling to continue its duties. NIH recently published a guideline reviewing best practices around controlled access methods.78
Enclave-based access methods consist of granting access to the data within a secure storage called an enclave, where the data accessors must perform their analysis without the ability to take the data out. Getting access to the enclave usually also involves submitting a data access application. Examples of enclave access methods include data from the All of Us Research Program Researcher Workbench, and the N3C COVID Enclave Data.64,79 This may be the most secure approach for protecting participant privacy, and preventing data misuse. However, the drawbacks of this approach may be the exclusion of certain individuals from accessing the data, especially those with limited knowledge of working in enclaves. In addition, this approach could be cost-prohibitive as it may require providing computational resources to users. It can also limit the ability to combine, and analyze data from different studies.
Overall, data openness reduces moving from the open sharing to controlled access, and enclave-based categories, while participants' security, and the cost associated with long-term access to the data correspondingly increase. There is a necessity for a method that protects the participants' privacy, especially against the increasing threats caused by AI, without compromising data openness. Recognizing that current frameworks fall short, and can foster false-security, the AI-READI project has introduced a novel open data sharing approach. This approach is designed to preserve data usability for research while establishing more robust safeguards against misuse, and unauthorized reidentification attempts.
The AI-READI open data sharing method
AI-READI is one of the four Data Generation Projects funded by Bridge2AI, an NIH Common Fund Program aimed at setting the stage for widespread adoption of AI in health research. The project seeks to create a flagship dataset to provide critical insights into Type 2 Diabetes Mellitus (T2DM), including salutogenic pathways to return to health.10,11 Data is collected from individuals with, and without T2DM, and harmonized across three data collection sites in the United States. The composition of the dataset consists of a multimodal array of data, including survey data, physical measurements, cognitive testing, vision testing, laboratory values, retinal imaging, ECG data, continuous glucose monitors, physical activity monitors, and home environmental sensors.11 Participant enrollment for data collection started in the summer of 2023. A total of 4000 participants are planned to be enrolled by the end of the project in November 2026. A major goal is to broadly share this multimodal dataset such that it is ready for AI/ML-related applications. The second version of the dataset, containing data from 1067 participants, was shared in November 2024.80
AI-READI participants provide informed consent emphasizing data sharing practices, and privacy protections. In research, the informed consent process is designed to facilitate understanding of what data will be collected and how data will be shared and used. It explicitly addresses potential risks, including data privacy concerns, and, outlines the measures taken to deidentify data, and maintain confidentiality. Participants are also made aware that while they can withdraw from the study at any time, data that has already been shared may remain in public, or controlled-access databases. This dataset has two sets. The first set is a public set that can only be used for T2DM-related research, and excludes the following data elements: ZIP code, genetic sequence, health records, motor vehicle accident reports, medications, sex, race/ethnicity. The decision to allow use of the public set only for T2DM-related research is meant to align with the consent. The public set is free from PHI, and could be shared under a method from the open access category described previously. Additionally, withholding race/ethnicity, or sex from the public set is intended to prevent findings that may stigmatize certain groups. The second set is a controlled set containing all the collected data, and can be used for any approved purpose. The controlled set requires a data usage agreement (DUA) for access. We took this opportunity to design a novel data access approach, considering pseudo-reidentification risks.
This novel data access process is implemented in FAIRhub, a novel data sharing platform designed to maintain open access while ensuring participants' privacy. We designed this model using a Swiss-cheese approach of open data sharing, and it contains several layers that may not be foolproof to protect participant privacy on their own, but can significantly reduce such risk when put in sequence (Fig. 2). The first layer consists of authenticating with an identity-verified system, which enables getting information about the person accessing the data (name, institutional email, affiliation) in a reliable way. The user is informed that their name, email, and intended use of the dataset are saved in the FAIRhub database, and are visible to the public on the project website. Currently, the authentication process for data access on FAIRhub is done through CILogon, which is an open-source identity, and access management platform operated by the National Center for Supercomputing Applications at the University of Illinois. Users from many institutions across the globe can authenticate using this platform. Its major limitation is the inability to provide attestation for non-academic individuals because it federates with known identity management providers. We are exploring alternative verified identity providers to ensure secure access for researchers with appropriate credentials, and to promote responsible use of the dataset by the end users. The second layer consists of reading the license terms, and agreeing to adhere to them. Identifying a gap in commonly used data sharing licenses such as CC-BY-4.0, we have established a new license that allows reuse of data for research, or commercial purposes but includes certain restrictions in place to protect the privacy of study participants.81 Additionally, the license explicitly prohibits data resharing (excluding with collaborators at the same institution), using models that remember the dataset, and attempting to reidentify the participants in any way. The third layer consists of attesting word-by-word to the major requirements mentioned in the license. This is intended to reinforce the requirements, and create a social contract that targets the individual user (while the License targets institutions). The fourth layer consists of describing the intended use of the data. This description is publicly posted on the dataset's landing page on FAIRhub along with the user's full name to provide full transparency about the use of the dataset, especially to the study participants, who can see what their data is used for. The fifth layer consists of watermarking the data. Watermarking is currently performed at the user level, meaning unique, and traceable watermarks on each file are associated with each user accessing the data (based on their identity obtained through layer 1). This enables tracing of any source of data leaks in the future, and allows individual attribution of any misuse, ensuring that anyone attempting to use the data outside of what the license permits will be held accountable. When ready, the user receives an email with the download instructions at their email address associated with the verified ID system they used to log in to FAIRhub, adding yet another layer of security.
Fig. 2.
Illustration of our new Swiss-cheese method of open data sharing. Each layer is designed to protect participants' privacy from potential risks by targeting primarily the individual user accessing the data, their principal investigator (PI), or their organization/company.
Overall, we have designed a multi-step process that integrates several layers of protection from misuse but remains accessible, rapid, and autonomous, thus maintaining data openness while enhancing the privacy protection of the participants from current, and future threats posed by evolving AI approaches. Our requirement for identity verification, attestation, and public disclosure of user information could still deter some users, or introduce barriers compared to truly open access.
Discussion
Data sharing is critical for advancing science. However, it may introduce inherent risks to study participants. Jurisdictions around the world do not require zero reidentification risk, which is unachievable in the context of data sharing. The line between protected data, and data with the potential for reidentification is not always clear, becoming increasingly blurred with the advent of powerful AI approaches. Therefore, different standards, rules, and policies are implemented at various levels, including research groups, institutional, state, national, and sometimes even continental levels. For example, some data elements may be considered high-risk yet shareable under certain conditions at one institution, while another institution might classify the same data as low-risk, and allow open sharing.
We reviewed the most widely used data-sharing regulations, highlighted their limitations, and explored the concept of pseudo-reidentification. Pseudo-reidentification is possible through ECG, CGM, wearable, or retinal images using advanced AI techniques. While still a step away from reidentification, this shows how evolving technologies represent an ever-increasing risk to participants' privacy. We postulate that new data sharing approaches are required to mitigate these risks. Accordingly, we presented the new approach we have implemented in the AI-READI project, using a Swiss-cheese model for open data sharing. We provided the rationale behind our approach, aiming to reduce risks to participants’ privacy while maintaining data openness.
The Swiss-cheese method of open data sharing is not intended to be a fixed method but to evolve with the addition, or removal of layers, to keep up with evolving privacy risks. As the current implementation is being tested by users accessing the AI-READI dataset through FAIRhub (422 dataset access as of April 2025), we will aim to identify the limitations of the current layers, and address them by exploring advanced technologies, such as blockchain-based audit logs, which offer transparent data tracking. We will also investigate under which circumstances participants are willing to incur different levels of pseudo-reidentification risk, and for which use cases.
We aim for this method to strengthen trust among researchers and study participants in openly shared datasets. We also hope our method will be adopted by other projects, either in its current form or as a foundation for developing new approaches that enhance participant protection while preserving data openness.
Contributors
All authors have contributed equally to the study ideation, design, literature review, manuscript drafting, and critical revision of the manuscript. SH and BP have designed the figures used in the paper. SH, AH, and FGK have generated tables used in the manuscript.
Declaration of interests
SH, NG, MB, JO, ED, BC, JLPT, LZ declare no competing interests. AH, FGPK, EB., NGE., SB., SM have received grants, contracts, or travel funding from NIH. EB is on the board of the Sisters School District School Board. NGE is on the board of NIH/NEI COAST Trial and TopCon. MPS is a cofounder and scientific advisor of Crosshair Therapeutics, Exposomics, Filtricine, Fodsel, iollo, InVu Health, January AI, Marble Therapeutics, Mirvie, Next Thought AI, Orange Street Ventures, Personalis, Protos Biologics, Qbio, RTHM, SensOmics. MPS is a scientific advisor of Abbratech, Applied Cognition, Enovone, Jupiter Therapeutics, M3 Helium, Mitrix, Neuvivo, Onza, Sigil Biosciences, TranscribeGlass, WndrHLTH, Yuvan Research. MPS is a cofounder of NiMo Therapeutics. MPS is an investor and scientific advisor of R42 and Swaza. MPS is an investor in Repair Biotechnologies. CN has received grants, contracts, or travel funding from NIH, NSF, and PCORI. CL has received grants, contracts, or travel funding from NIH, Alzheimer's Disease Drug Foundation, and Gates Ventures. AYL has received grants, contracts, or travel funding from NIH, Amazon, Lantham Vision Science Award, Meta, Regeneron, Santen, Topcon, Zeiss, and Research to Prevent Blindness. AYL has received consulting fees from Sanofi, Boehringer Ingelheim, Genentech, Inc., Gyroscope, Johnson & Johnson, and US FDA. AYL has received non-financial support from Heidelberg, iCareWorld, Microsoft, and Optomed. BP has received grants, contracts, or travel funding from NIH, The Navigation Fund, and The University of California Office of the President. BP has received non-financial support from Microsoft.
Acknowledgements
This work was supported by the NIH through grants OT2OD032644, and T32EY026590. We thank the Microsoft AI for Good Lab for supporting the cloud services needed for the project.
References
- 1.Clinical Cancer Genome Task Team of the Global Alliance for Genomics and Health, Lawler M., Haussler D., et al. Sharing clinical and genomic data on cancer - the need for global solutions. N Engl J Med. 2017;376:2006–2009. doi: 10.1056/NEJMp1612254. [DOI] [PubMed] [Google Scholar]
- 2.Moorthy V., Henao Restrepo A.M., Preziosi M.-P., Swaminathan S. Data sharing for novel coronavirus (COVID-19) Bull World Health Organ. 2020;98:150. doi: 10.2471/BLT.20.251561. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Wilkinson M.D., Dumontier M., Aalbersberg I.J.J., et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3 doi: 10.1038/sdata.2016.18. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Cascini F., Pantovic A., Al-Ajlouni Y.A., Puleo V., De Maio L., Ricciardi W. Health data sharing attitudes towards primary and secondary use of data: a systematic review. eClinicalMedicine. 2024;71 doi: 10.1016/j.eclinm.2024.102551. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Aitken M., de St Jorre J., Pagliari C., Jepson R., Cunningham-Burley S. Public responses to the sharing and linkage of health data for research purposes: a systematic review and thematic synthesis of qualitative studies. BMC Med Ethics. 2016;17:73. doi: 10.1186/s12910-016-0153-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Hutchings E., Loomes M., Butow P., Boyle F.M. A systematic literature review of health consumer attitudes towards secondary use and sharing of health administrative and clinical trial data: a focus on privacy, trust, and transparency. Syst Rev. 2020;9:235. doi: 10.1186/s13643-020-01481-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Data Sharing Concerns Rethinking clinical trials. 2017. https://rethinkingclinicaltrials.org/chapters/dissemination/data-share-top/data-sharing-concerns/
- 8.Niño de Rivera S., Masterson Creber R., Zhao Y., et al. Public perspectives on increased data sharing in health research in the context of the 2023 National Institutes of Health Data Sharing Policy. PLoS One. 2024;19 doi: 10.1371/journal.pone.0309161. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Murdoch B. Privacy and artificial intelligence: challenges for protecting health information in a new era. BMC Med Ethics. 2021;22:122. doi: 10.1186/s12910-021-00687-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.AI-READI Consortium AI-READI: rethinking AI data collection, preparation and sharing in diabetes research and beyond. Nat Metab. 2024;6:2210–2212. doi: 10.1038/s42255-024-01165-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Owsley C., Matthies D.S., McGwin G., et al. Cross-sectional design and protocol for Artificial Intelligence Ready and Equitable Atlas for Diabetes Insights (AI-READI) BMJ Open. 2025;15 doi: 10.1136/bmjopen-2024-097449. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Elliot M., Mandalari A.M., Mourby M., O‘Hara K. Dictionary of Privacy, Data Protection and Information Security. Edward Elgar Publishing Limited; Cheltenham, UK: 2024. Reidentification.https://www.elgaronline.com/view/book/9781035300921/b-9781035300921-R_25.xml Retrieved Sep 17, 2025, from. [Google Scholar]
- 13.Elliot M., Mandalari A.M., Mourby M., O'Hara K. Edward Elgar Publishing Limited; Cheltenham, UK: 2024. Dictionary of Privacy, Data Protection and Information Security. Retrieved Sep 17, 2025, from. [DOI] [Google Scholar]
- 14.Song Z., Ma H., Sun S., Xin Y., Zhang R. Rainbow: reliable personally identifiable information retrieval across multi-cloud. Cybersecur (Singap) 2023;6:19. doi: 10.1186/s42400-023-00146-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Differences Between PII, Sensitive PII, and PHI Municipal websites central help Center. https://www.civicengagecentral.civicplus.help/hc/en-us/articles/1500001543581-Differences-Between-PII-Sensitive-PII-and-PHI
- 16.Department of Health Care Services List of HIPAA identifiers. https://www.dhcs.ca.gov/dataandstats/data/Pages/ListofHIPAAIdentifiers.aspx
- 17.Kayaalp M. Modes of De-identification. AMIA Annu Symp Proc. 2017;2017:1044–1050. [PMC free article] [PubMed] [Google Scholar]
- 18.Office for Civil Rights (OCR) HHS.gov; 2008. Health Information Privacy.https://www.hhs.gov/hipaa/for-professionals/special-topics/research/index.html [Google Scholar]
- 19.Office for Human Research Protections (OHRP) HHS.gov; 2017. Revised Common Rule.https://www.hhs.gov/ohrp/regulations-and-policy/regulations/finalized-revisions-common-rule/index.html [Google Scholar]
- 20.General Data Protection Regulation (GDPR) Compliance Guidelines GDPR.eu. 2018. https://gdpr.eu/
- 21.Chevrier R., Foufi V., Gaudet-Blavignac C., Robert A., Lovis C. Use and understanding of anonymization and de-identification in the biomedical literature: scoping review. J Med Internet Res. 2019;21 doi: 10.2196/13484. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.General Data Protection Regulation (GDPR) – Legal Text General data protection regulation (GDPR) https://gdpr-info.eu/
- 23.Nakayama L.F., de Matos J.C.R.G., Stewart I.U., et al. Retinal scans and data sharing: the privacy and scientific development equilibrium. Mayo Clin Proc. 2023;1:67–74. doi: 10.1016/j.mcpdig.2023.02.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Committee on Strategies for Responsible Sharing of Clinical Trial Data, Board on Health Sciences Policy, Institute of Medicine . Sharing Clinical Trial Data: Maximizing Benefits, Minimizing Risk. National Academies Press (US); 2015. Concepts and methods for de-identifying clinical trial data. [PubMed] [Google Scholar]
- 25.Anthem pays OCR $16 million in record HIPAA settlement following largest U.S. health data breach in history. https://www.hhs.gov/guidance/document/anthem-pays-ocr-16-million-record-hipaa-settlement-following-largest-us-health-data-breach
- 26.Wiepert D., Malin B.A., Duffy J.R., et al. Reidentification of participants in shared clinical data sets: experimental Study. JMIR AI. 2024;3 doi: 10.2196/52054. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Spadaccini A., Beritelli F. Performance evaluation of heart sounds biometric systems on an open dataset. https://ieeexplore.ieee.org/document/6622835
- 28.Sancho J., Alesanco Á., García J. Biometric authentication using the PPG: a long-term feasibility study. Sensors. 2018;18 doi: 10.3390/s18051525. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Labati R.D., Piuri V., Rundo F., Scotti F., Spampinato C. Pattern Recognition ICPR International Workshops and Challenges; 2021. Biometric Recognition of PPG Cardiac Signals Using Transformed Spectrogram Images; pp. 244–257. [Google Scholar]
- 30.Retsinas G., Filntisis P.P., Efthymiou N., Theodosis E., Zlatintsi A., Maragos P. Person identification using deep convolutional neural networks on short-term signals from wearable sensors. https://ieeexplore.ieee.org/document/9053910
- 31.Lee E., Ho A., Wang Y.-T., Huang C.-H., Lee C.-Y. Cross-domain adaptation for biometric identification using photoplethysmogram. https://ieeexplore.ieee.org/document/9053604
- 32.Hwang D.Y., Taha B., Da Saem L., Hatzinakos D. Evaluation of the time stability and uniqueness in PPG-based biometric system. https://ieeexplore.ieee.org/document/9130730
- 33.Yadav U., Abbas S.N., Hatzinakos D. Evaluation of PPG biometrics for authentication in different states. 2017. http://arxiv.org/abs/1712.08583
- 34.Herrero P., Reddy M., Georgiou P., Oliver N.S. Identifying continuous glucose monitoring data using machine learning. Diabetes Technol Ther. 2022;24:403–408. doi: 10.1089/dia.2021.0498. [DOI] [PubMed] [Google Scholar]
- 35.Tan R., Perkowski M. Toward improving electrocardiogram (ECG) biometric verification using Mobile sensors: a two-stage classifier approach. Sensors. 2017;17 doi: 10.3390/s17020410. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Arnau-Gonzalez P., Katsigiannis S., Ramzan N., Tolson D., Arevalillo-Herrez M. 2017 IEEE 17th international Conference on Bioinformatics and Bioengineering (BIBE) IEEE; 2017. ES1D: a deep network for EEG-Based subject identification; pp. 81–85. [Google Scholar]
- 37.Zhao Z., Zhang Y., Deng Y., Zhang X. ECG authentication system design incorporating a convolutional neural network and generalized S-Transformation. Comput Biol Med. 2018;102:168–179. doi: 10.1016/j.compbiomed.2018.09.027. [DOI] [PubMed] [Google Scholar]
- 38.Patro K.K., Jaya Prakash A., Jayamanmadha Rao M., Rajesh Kumar P. An efficient optimized feature selection with machine learning approach for ECG biometric recognition. IETE J Res. 2022;68:2743–2754. [Google Scholar]
- 39.Patro K.K., Reddi S.P.R., Khalelulla S.K.E., Rajesh Kumar P., Shankar K. ECG data optimization for biometric human recognition using statistical distributed machine learning algorithm. J Supercomput. 2020;76:858–875. [Google Scholar]
- 40.El Boujnouni I., Zili H., Tali A., Tali T., Laaziz Y. A wavelet-based capsule neural network for ECG biometric identification. Biomed Signal Process Control. 2022;76 [Google Scholar]
- 41.Allam J.P., Patro K.K., Hammad M., Tadeusiewicz R., Pławiak P. BAED: a secured biometric authentication system using ECG signal based on deep learning techniques. Biocybern Biomed Eng. 2022;42:1081–1093. [Google Scholar]
- 42.Prakash A.J., Patro K.K., Samantray S., Pławiak P., Hammad M. A deep learning technique for biometric authentication using ECG beat template matching. Information. 2023;14:65. [Google Scholar]
- 43.Wang X., Cai W., Wang M. A novel approach for biometric recognition based on ECG feature vectors. Biomed Signal Process Control. 2023;86 [Google Scholar]
- 44.Farzin H., Abrishami-Moghaddam H., Moin M.-S. A novel retinal identification system. EURASIP J Adv Signal Process. 2008;2008 [Google Scholar]
- 45.Köse C., İki˙baş C. A personal identification system using retinal vasculature in retinal fundus images. Expert Syst Appl. 2011;38:13670–13681. [Google Scholar]
- 46.Biometric retina identification based on neural network. Procedia Comput Sci. 2016;102:26–33. [Google Scholar]
- 47.Szymkowski M., Saeed E., Omieljanowicz M., Omieljanowicz A., Saeed K., Mariak Z. A novelty approach to retina diagnosing using biometric techniques with SVM and clustering algorithms. https://ieeexplore.ieee.org/abstract/document/9134747
- 48.Machine Learning for Biometrics. Academic Press; 2022. Retina biometrics for personal authentication; pp. 87–104. [Google Scholar]
- 49.Marappan J., Murugesan K., Elangeeran M., Subramanian U. Human retinal biometric recognition system based on multiple feature extraction. J Electron Imaging. 2023;32 [Google Scholar]
- 50.Wang Z., Kanduri A., Aqajari S.A.H., et al. ECG unveiled: analysis of client Re-identification risks in real-world ECG datasets. 2024. http://arxiv.org/abs/2408.10228
- 51.Shei R.-J., Holder I.G., Oumsang A.S., Paris B.A., Paris H.L. Wearable activity trackers-advanced technology or advanced marketing? Eur J Appl Physiol. 2022;122:1975–1990. doi: 10.1007/s00421-022-04951-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Canali S., Schiaffonati V., Aliverti A. Challenges and recommendations for wearable devices in digital health: data quality, interoperability, health equity, fairness. PLOS Digit Health. 2022;1 doi: 10.1371/journal.pdig.0000104. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.The Lancet Digital Health Wearable health data privacy. Lancet Digit Health. 2023;5 doi: 10.1016/S2589-7500(23)00055-9. [DOI] [PubMed] [Google Scholar]
- 54.Metwally A.A., Perelman D., Park H., et al. Prediction of metabolic subphenotypes of type 2 diabetes via continuous glucose monitoring and machine learning. Nat Biomed Eng. 2024 doi: 10.1038/s41551-024-01311-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Britton K.E., Britton-Colonnese J.D. Privacy and security issues surrounding the protection of data generated by continuous glucose monitors. J Diabetes Sci Technol. 2017;11:216–219. doi: 10.1177/1932296816681585. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Martens T., Beck R.W., Bailey R., et al. Effect of continuous glucose monitoring on glycemic control in patients with type 2 diabetes treated with basal insulin: a randomized clinical trial: a randomized clinical trial. JAMA. 2021;325:2262–2272. doi: 10.1001/jama.2021.7444. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Kleidermacher D., Klonoff D., Nguyen K., Schwartz N., Xu N. Addressing the need for protecting cybersecurity in connected diabetes devices. IEEE Standards Association. https://standards.ieee.org/beyond-standards/addressing-the-need-for-protecting-cybersecurity-in-connected-diabetes-devices/
- 58.Medical biometric databases. BioSec.Lab. https://www.comm.utoronto.ca/∼biometrics/databases.html
- 59.Biel L., Pettersson O., Philipson L., Wide P. ECG analysis: a new approach in human identification. IEEE Trans Instrum Meas. 2001;50:808–812. [Google Scholar]
- 60.Chun S.Y., Kang J.-H., Kim H., Lee C., Oakley I., Kim S.-P. 2016 39th International Conference on Telecommunications and Signal Processing (TSP) IEEE; 2016. ECG based user authentication for wearable devices using short time fourier transform; pp. 656–659. [Google Scholar]
- 61.Pereira T.M.C., Conceição R.C., Sencadas V., Sebastião R. Biometric recognition: a systematic review on electrocardiogram data acquisition methods. Sensors. 2023;23 doi: 10.3390/s23031507. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Wu S.-C., Chen P.-T., Hsieh J.-H. Spatiotemporal features of electrocardiogram for biometric recognition. Multidimens Syst Signal Process. 2019;30:989–1007. [Google Scholar]
- 63.Kim H., Kim H., Chun S.Y., et al. A wearable wrist band-type system for multimodal biometrics integrated with multispectral skin photomatrix and electrocardiogram sensors. Sensors. 2018;18 doi: 10.3390/s18082738. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Gim N., Wu Y., Blazes M., Lee C.S., Wang R.K., Lee A.Y. A clinician's guide to sharing data for AI in ophthalmology. Investig Ophthalmol Vis Sci. 2024;65:21. doi: 10.1167/iovs.65.6.21. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Balancing benefits and risks: the case for retinal images to be considered as nonprotected health information for research purposes - 2024. American Academy of Ophthalmology; 2024. https://www.aao.org/education/clinical-statement/balancing-benefits-risks-case-retinal-images-to-be [DOI] [PubMed] [Google Scholar]
- 66.Price W.N., 2nd, Cohen I.G. Privacy in the age of medical big data. Nat Med. 2019;25:37–43. doi: 10.1038/s41591-018-0272-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Luna R., Rhine E., Myhra M., Sullivan R., Kruse C.S. Cyber threats to health information systems: a systematic review. Technol Health Care. 2016;24:1–9. doi: 10.3233/THC-151102. [DOI] [PubMed] [Google Scholar]
- 68.Vilhuber L. Reproducibility and transparency versus privacy and confidentiality: reflections from a data editor. J Econom. 2023;235:2285–2294. [Google Scholar]
- 69.Thakkar V., Gordon K. Privacy and policy implications for big data and health information technology for patients: a historical and legal analysis. Stud Health Technol Inf. 2019;257 https://pubmed.ncbi.nlm.nih.gov/30741232/ [PubMed] [Google Scholar]
- 70.Bai S., Zheng J., Wu W., Gao D., Gu X. Research on healthcare data sharing in the context of digital platforms considering the risks of data breaches. Front Public Health. 2024;12 doi: 10.3389/fpubh.2024.1438579. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Lin D., McAuliffe M., Pruitt K.D., et al. Biomedical data repository concepts and management principles. Sci Data. 2024;11:622. doi: 10.1038/s41597-024-03449-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Bandrowski A., Grethe J.S., Pilko A., et al. SPARC data structure: rationale and design of a FAIR standard for biomedical research data. bioRxiv. 2021;2021 [Google Scholar]
- 73.Markiewicz C.J., Gorgolewski K.J., Feingold F., et al. The OpenNeuro resource for sharing of neuroscience data. eLife. 2021;10 doi: 10.7554/eLife.71774. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.Alper P., Dĕd V., Herzinger S., et al. DS-PACK: tool assembly for the end-to-end support of controlled access human data sharing. Sci Data. 2024;11:501. doi: 10.1038/s41597-024-03326-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Tryka K.A., Hao L., Sturcke A., et al. NCBI's database of genotypes and phenotypes: dbGaP. Nucleic Acids Res. 2014;42:D975–D979. doi: 10.1093/nar/gkt1211. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Sudlow C., Gallacher J., Allen N., et al. UK biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age. PLoS Med. 2015;12 doi: 10.1371/journal.pmed.1001779. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Sydes M.R., Johnson A.L., Meredith S.K., Rauchenberger M., South A., Parmar M.K.B. Sharing data from clinical trials: the rationale for a controlled access approach. Trials. 2015;16:104. doi: 10.1186/s13063-015-0604-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.NIH security best practices for controlled-access data and repositories. https://sharing.nih.gov/accessing-data/NIH-security-best-practices
- 79.All of Us Research Program Investigators, Denny J.C., Rutter J.L., et al. The ‘All of Us’ research program. N Engl J Med. 2019;381:668–676. doi: 10.1056/NEJMsr1809937. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80.AI-READI Consortium . 2024. Flagship dataset of type 2 diabetes from the AI-READI project. Acccessed October 30 2025. [DOI] [Google Scholar]
- 81.Contreras J, Evans B, Hurst S, et al. 2024. License terms for reusing the AI-READI dataset. Version 1.0. Zenodo.doi:10.5281/zenodo.10642459 [Google Scholar]


