Summary
Large national, integrated biobanks have revolutionized how genetics-linked healthcare data can be scaled, providing access to massive databases to researchers globally. Recognizing the importance of integrated biobanks for public health and national scientific advancement, countries around the world have launched similar initiatives. Despite comprising a quarter of the world’s population, South Asia accounts for only 1·8% of EHR-indexed publications and 0·2% of GWAS participants. We argue for a South Asia Biobank Consortium: (1) a regional governing body overseeing interoperability across (2) national-level integrated biobanks that adapts the UK Biobank model to regional contexts (3) supported by federated analytics infrastructure with global access. If enacted, the consortium represents a scientific imperative and a pathway to digital health equity for nearly two billion people living in South Asia. We present a framework based on hallmarks of successful integrated biobanks and critical success factors. We propose a timeline for its establishment. Without decisive action, current disparities will worsen, leaving South Asia’s population marginalized as the transformative revolution continues. With a federated, equitable strategy, South Asia can transform from a peripheral participant into a central driver of biomedical discovery – strengthening health systems, advancing equity, and realizing the global promise of precision medicine.
Keywords: Biobank, Electronic health records, Federated analytics, Genomics, Health equity, Precision medicine, South Asia
Large national biobanks, like the UK Biobank1 and the All of Us Research Program (USA),2 have revolutionized how research and healthcare can be scaled by creating integrated, multi-modal data warehouses that provide secure and ethical data access to researchers worldwide. The UK Biobank’s 500,000 participants have already transformed biomedical research,3 while South Asia – with eight times the population of the UK – lacks a comparable resource. When researchers investigate diabetes in the UK, they can analyze genetic data from hundreds of thousands of people; when South Asian scientists study diabetes, which affects over 100 million Indians alone,4 they work with fragmented datasets of a few thousand or rely on knowledge derived primarily from non-South Asian populations. This skewness in data distribution represents one of the most consequential blind spots in global biomedical research.
Integrated biobanks are repositories that merge biospecimens (and associated DNA) with electronic health records (EHRs) and other linkable data like environmental exposures or lifestyle exposures, enabling disease-agnostic, multimodal research at population scale.5, 6, 7 Integrated biobanks can be institution-based (e.g., Tata Medical Centre Biorepository (TiMBR) focused on cancer,8,9 National Liver Disease Biobank focused on liver disease10), often focused on specific diseases and patient populations, or national (e.g., UK Biobank, All of Us), which are often population-based, disease-agnostic, and government-supported. A third model comprises consortia, linking institutional or national biobanks to enable cross-border analyses, as exemplified by the Global Biobank Meta-Analysis Initiative (GBMI).11 In contrast, the classical epidemiological approach involves curating an investigator-initiated, proprietary cohort study or detailed database around a single disease or exposure.
While undoubtedly crucial for advancing population health, traditional cohort studies cannot be the only tool for population-health research. As integrated biobanks reshape global medical research, South Asia’s nascent and fragmented biobanking landscape risks perpetuating data inequities. When resources are scarce, creating a massive national data resource that is disease- and exposure-agnostic is a more forward-thinking, inclusive, and cost-effective strategy. Specifically, integrated biobanks can leverage shared infrastructures typically scattered across narrow, bespoke cohort studies to achieve economies of scale in recruitment, consent, data collection, and data management.12 We argue that establishing coordinated national biobanks across South Asia, inspired by the UK Biobank model but adapted to regional contexts, boosted by federated governance and learning, represents a scientific imperative and a pathway to digital health equity for nearly two billion people living in South Asia.
Launched in 2003, the UK Biobank exemplifies the transformative potential of integrated biobanks,1,13 linking genetic, environmental, and lifestyle data, enabling researchers to elucidate disease causes and establish robust genotype-phenotype linkages while contributing over 9000 research publications informing medical care worldwide.1 Recognizing the strategic importance of integrated biobanks for public health and national scientific advancement,14 countries around the world have launched similar initiatives (Table 1), including All of Us (USA),2 China Kadoorie Biobank,15,61 and Biobank Japan,25 each demonstrating how broad consent models and open data access policies accelerate discovery while maintaining privacy protections.62 Research using these biobanks has had a wide-ranging impact, from drug and clinical guideline development to diagnostics and public health policy.
Table 1.
Summary of selected national biobanks and South Asian biobanking initiatives.
| National biobanks | |||||
|---|---|---|---|---|---|
| Biobank name | Country | Key features (Participant Count, Data Types, Focus) | Data access policy overview | Refs. | Link |
| China Kadoorie Biobank (CKB) | China | ∼500,000 population-based subjects. Aims to couple to national EHR systems. Collects genetic, lifestyle, and health outcome data. | Data shared only with bona fide researchers from recognized academic/health organizations. Restricted access to samples. Open access to CKB resources after a period of exclusive use by collaborators. Safeguards ensure anonymity; legal agreement not to identify participants. Access charge for non-China/HK applicants. Researchers must publish findings and return derived data to CKB. CKB owns data/samples. | 15,16 | https://www.ckbiobank.org/front-page |
| Copenhagen Hospital Biobank | Denmark | ∼423,000 participants recruited since 2009. Aim to facilitate research in health and disease by enabling researchers’ access to a large resource of well-defined patient samples. Collects blood samples. | Data available to use in research through collaboration. An application must be prepared in collaboration with CHB and relevant clinicians from the Capital Region. The project subsequently needs clearance from the Copenhagen Hospital Biobank Committee to avoid overlapping with existing research efforts. In addition, according to the Danish legislation, all projects should be notified to and receive permission from the Danish Health Research Ethics Committee System, as well as from Knowledge Centre on Data Protection Compliance, the current personal data protection supervision authority in the Capital Region of Denmark. In cases where direct linking to electronic patient records is desirable, access must also be approved by the Danish Patient Safety Authority. | 17 | https://doi.org/10.1093/ije/dyaa157 |
| Danish Cancer Biobank | Denmark | ∼24,500 participants with a cancer diagnosis recruited and ∼40,400 blood samples collected. The goal for the Danish Cancer Biobank is to strengthen the infrastructure for clinical cancer research and transfer biospecimens for supplementary diagnostics. | Data available upon application for research purpose. The applications will be reviewed by a research ethics committee. | 18 | https://www.rbgb.dk/en/om-rbgb/ |
| Danish National Biobank | Denmark | >10 million Danish biological samples. Biobank based in a country with EHRs and single-payer system, enabling large datasets. | Requires application, full project description, scientific ethics committee approval, institutional approval, and Data Protection Agency approval. Reviewed by Scientific Board. Quantity limits apply (e.g., 100 μl serum/plasma or 1 μg DNA). | 19 | https://www.danishnationalbiobank.com/ |
| Danish National Birth Cohort | Denmark | >180,000 women early in pregnancy recruited. The Danish National Birth Cohort (DNBC) was established to investigate the causal link between exposures in early life and disease later on and the possibilities for disease prevention. Data available include survey results and blood samples | Access to data follows an open access policy. Access to personalised data requires that your DNBC project is listed on your institution's record of data processing activities as well as a permission from the DNBC Reference Group. The access policy for biological specimens will be restrictive. The Scientific Ethical Committee, the Danish Data Protection Agency and the DNBC Management must sanction any access. | 20 | https://www.dnbc.dk/ |
| Estonian Biobank | Estonia | >200,000 participants from population of 1·3 million (15% participation rate). Genetic data linked to EHRs and national health registries. Longitudinal follow-up through national health system integration. Focus on population genetics and precision medicine. | Open access model for academic research. Commercial access available through partnerships. Participants can access their own genetic data and health reports. Strong emphasis on data return and participant engagement. Requires research proposals and ethical approval. | 21,22 | https://genomics.ut.ee/en/content/estonian-biobank |
| FinnGen | Finland | 500,000+ Finnish participants with genetic data linked to comprehensive health registry data spanning decades. Leverages Finland's unique population genetics and extensive health records. Focus on drug discovery and precision medicine applications. | Data available to academic researchers and pharmaceutical partners. Requires detailed research proposals and ethical review. Results must be made publicly available. Strong public-private partnership model with data sharing agreements. | 23,24 | https://www.finngen.fi/en |
| BioBank Japan (BBJ) | Japan | Disease-enriched biobank. ∼270,000 individuals with DNA samples and clinical information (51 diseases). ∼200,000 with serum samples. | Samples and data available to broad research community (institutions, companies). Projects require rigorous screening process based on Japanese rules/guidelines. Data stored under strict security control (ISO standards). | 25,26 | https://biobankjp.org/en/#gsc.tab=0 |
| LifeLines Cohort Study | Netherlands | >167,000 participants (10%) from the northern population of the Netherlands, included participants from three generations, who are followed with a lifespan perspective, to obtain insight into healthy ageing since 2006. Collects urine, serum, plasma, buffycoats, and DNA samples. The data also include questionnaires, measurements, and lab analyses. | Data available upon application for research purpose. The applications will be reviewed by LifeLines and co-reviewers. A data (and material) transfer agreement will be prepared if an application is approved. | 27 | https://www.lifelines-biobank.com/ |
| Netherlands Twin Register (NTR) | Netherlands | ∼120,000 twins and a roughly equal number of their relatives since 1987, focusing on behavioral genetic research. Collecting data and samples, including blood, urine, and clinical data, on twins and multiples and their families | Researchers that contribute unique expertise or unique facilities are specifically welcomed. To facilitate scientific collaboration, all NTR data has been reorganized in a uniform NTR data repository, annotated using the FAIR (Findable, Accessible, Interoperable, Reproducible) principles. Regardless of the type of research interest, all potential collaborations on data contained in the NTR repository must be first reviewed by the NTR Data Access Committee (NTR-DAC). For each project, a Data Sharing Request (DSR) should be submitted by an applicant with a PhD, MD or equivalent degree embedded in a reputable research institute/organization. | 28,29 | https://tweelingenregister.vu.nl/nl |
| Rotterdam Study | Netherlands | ∼15,000 participants (aged ≥ 45) recruited between 1990 and 2008. The main objectives were to investigate the risk factors of cardiovascular, neurological, ophthalmological and endocrine diseases in the elderly. Participants were interviewed at home and went through an extensive set of examinations, bone mineral densitometry, including blood sample collections for in-depth molecular and genetic analyses. | Data can be obtained upon request. Requests should be directed towards the management team of the Rotterdam Study (datamanagement.ergo@erasmusmc.nl), which has a protocol for approving data requests. Because of restrictions based on privacy regulations and informed consent of the participants, data cannot be made freely available in a public repository. | 30 | http://www.erasmus-epidemiology.nl/research/ergo.htm |
| Taiwan Biobank (TWB) | Taiwan | >150,000 individuals (aged 30–70) recruited since 2012. Government-supported prospective cohort with wide phenotypic measurements and genomic data. Aims to link to National Health Insurance Research Database (NHIRD). | Data available upon application for research purposes, requiring detailed research proposal and IRB approval. International data transfer agreement needed for non-Taiwanese researchers. Personal information protected with firewalls, antivirus, encryption; access limited to authorized personnel. Donors have rights to inquire, access, stop collection/use, or delete personal info. | 31,32 | https://www.twbiobank.org.tw/ |
| Genomics England (100,000 Genomes Project) | UK | 100,000+ genomes from NHS patients with rare diseases and cancer. Integrates genomics into routine clinical care. Whole genome sequencing linked to detailed clinical phenotypes and family history data. Focus on clinical implementation and healthcare integration. | Access through approved research environment for academic and commercial researchers. Requires research proposals, ethical approval, and data security compliance. Clinical-grade data with consent for research use. Strong focus on returning actionable findings to patients and clinicians. | 33,34 | https://www.genomicsengland.co.uk/initiatives/100000-genomes-project |
| UK Biobank | UK | ∼500,000 participants (aged 40–69). Extensive genetic, clinical, lifestyle, and imaging data (brain, heart, abdomen). Longitudinal follow-up. Focus on common diseases (cancer, cardiovascular, neurodegenerative). | Open-access resource for public health research. Data is widely available, equitable, and transparent. Identifiable information (name, address, NHS number) is never shared with researchers. Researchers undergo vetting, work for credible organizations, and must demonstrate legitimate scientific purpose. Fees apply, with a fund for low/middle-income countries. Findings must be public. | 1,35 | https://www.ukbiobank.ac.uk/ |
| All of Us Research Program | US | Aims for 1 million diverse participants. Collects health (EHR, questionnaires, physical measurements), genetic (whole-genome sequences), and lifestyle data. High representation of underrepresented minorities (77% from historically under-represented in biomedical research, 46% from under-represented racial/ethnic minorities). | Three tiers: Public (no login), Registered (login required), Controlled (additional approval). User-based authorization. Identifiable information removed, replaced with a code. Researchers sign contracts not to identify participants and use data only for health research. Data stored securely on cloud servers, not downloadable to personal computers. | 2,36 | https://www.joinallofus.org/ |
| South Asian Biobanking Initiatives | |||||
|---|---|---|---|---|---|
| Biobank name | Country | Key features (Participant Count, Data Types, Focus) | Data access policy overview | Refs | Link |
| AMANHI Biobank | Bangladesh Pakistan |
A large cohort of pregnant women and their babies in sub-Saharan Africa and South Asia. The population level biorepository has recruited 5500 pregnant women in South Asia to study interactions between genes, multi-omics and a wide range of varying environmental exposures on key pregnancy and birth outcomes. Collects blood, urine, stool, saliva, and tissue. | Data available upon request by submitting a formal application that includes details of the study team, scientific rationale, research question, study design, population, proposed variables, and analysis plan. Requests are reviewed by the AMANHI group in consultation with site-specific sample utilization committees, to ensure alignment with participant consent and ethical approvals. Data will be shared in accordance with WHO policies and subject to data use agreements. | 37 | https://academic.oup.com/ije/article/50/6/1780/6356791 |
| BANGABANDHU Study | Bangladesh | Ischemic heart disease (IHD) focused study recruiting 750 IHD cases and 750 controls. Collects detailed phenotype, clinical, demographic, and environmental risk factor data with genetic analysis. | Data available upon request from corresponding authors. Patient confidentiality protected with encrypted data. Investigators with genetic access blinded to personal identifiers. Future applications require BSMMU ethical guidelines compliance. | 38 | https://doi.org/10.2147/IJGM.S466706 |
| MAGPIE study | Bangladesh | ∼3000 Bangladeshi stroke and control sample. Collects extensive phenotypic data as well as blood sample for genetic analysis | Data are available upon request from the corresponding authors. Patient confidentiality was ensured through encrypted data, and investigators with genetic access were blinded to personal identifiers. | 39 | https://onlinelibrary.wiley.com/doi/10.1002/hsr2.70227 |
| Bhutan Biobank | Bhutan | ∼3000 participants for pilot activities. Collects blood samples for genomic sequencing and interpretation. Focus on non-communicable diseases, particularly hypertension | Data access policy not yet established. The biobank is in the early planning stages and currently does not have formal procedures for data sharing or research collaboration. | 40,41 |
https://openjicareport.jica.go.jp/pdf/12348769.pdf https://bhutanhrp.moh.gov.bt/index.php/hrp/search/viewProposal/79 |
| GARBH-INi Pregnancy Cohort | India | 12,211 pregnant women early in pregnancy enrolled since 2015. Collects blood, urine, saliva, tissue, and breast milk samples. Focus on finding solutions for better birth outcomes utilizing a multi-pronged approach of integrating clinical epidemiology, multi-omics biomarkers and AI-driven tools for personalized predictions | Data and samples available upon request following user registration and submission of a formal application. Proposals are reviewed by the GARBH-Ini Access Control Committee for scientific merit, feasibility, and ethical compliance. Approved users must sign Data/Material Transfer Agreements; data are shared via secure servers and samples shipped under regulatory guidelines. | 42 | https://garbhinidrishti.thsti.in/vizgarbh/ |
| Human Brain Tissue Repository (HBTR) | India | Brain tissue biobank at NIMHANS focusing on neurological and psychiatric disorders. Collects post-mortem brain tissues with detailed clinical and pathological data. Supports neuroscience research. | Access requires research proposal review and ethical approval. Strict protocols for brain tissue distribution. Collaborative research agreements with approved institutions. | 43 | https://thenimhansbrainbank.in/brain-bank/ |
| Indian Council of Medical Research–India Diabetes (ICMR–INDIAB) study | India | ∼120,000 participants (aged > 20) between Oct 18, 2008, and Dec 17, 2020. A cross-sectional, population-based survey on diabetes. Collect blood samples. | Data access available upon request from corresponding authors. Proposals may require review and approval by institutional ethics committees. | 44 | https://mdrf.in/department/grants.html |
| Indian Study of Healthy Ageing (ISHA-Barshi) | India | ∼39,000 participants (aged 30–69) recruited between 2015 and 2020 from towns and villages around Barshi, Solapur district, Maharashtra state, India. Collects health (questionnaires, physical measurements), lifestyle, and clinical data, as well as nail clippings and blood samples. Focus on the burden, causes and consequences of chronic diseases. | Although data are not currently publicly available, proposals for collaborative research are welcome following completion of data collection. Requests should be directed to the study’s International Steering Committee. Research use may require ethics approval and data use agreements. | 45 | https://doi.org/10.1093/ije/dyae079 |
| National Cancer Tissue BioBank (NCTB) | India | Disease-specific biobank focusing on cancer research. Collects tissue samples, blood, and associated clinical data. Part of IIT Madras initiative for translational cancer research. | Access requires institutional approval and ethical clearance. Data sharing governed by Indian privacy guidelines. Research proposals reviewed by institutional committees. | 46 | https://nctb.iitm.ac.in/about/index.html |
| National Liver Disease Biobank (NLDB) | India | Specialized biobank for liver diseases including hepatitis, cirrhosis, and liver cancer. Collects serum, plasma, tissue samples, and associated clinical phenotype data. Supports hepatology research. | Data access requires ethical clearance and institutional approval. Samples provided for legitimate research purposes. Privacy protection through de-identification protocols. | 47 | https://nldb.in/ |
| Phenome India-CSIR Health Cohort Knowledgebase (PI–CHeCK) | India | ∼10,000 participants since April 2022. Collects clinical questionnaire, lifestyle and dietary habits, anthropometric parameters, assessment for lung function, liver elastography, ECG, biochemical data, and molecular assays, including genomics, plasma proteomics, metabolomics, and faecal microbiome. Focus on developing clinically relevant personalized risk prediction scores of cardiometabolic diseases for the Indian population. | Access restricted to authorized personnel with approved projects and predefined objectives. Patient data are encrypted and confidentiality protected. Requests are reviewed by a data access committee and granted for a defined period if objectives do not overlap with existing studies. | 48 | https://doi.org/10.1101/2024.10.17.24315252 |
| Precision Cardiovascular Diseases Phenotyping and Pathophysiological Pathways in the cArdiometabolic Risk Reduction in South Asia Cohort (Precision-CARRS) | India | ∼14,300 participants (aged ≥ 20) since January 2023. Collects anthropometric and blood pressure measurements, bio-sample collection, and specialized imaging tests, which include cardiac echocardiography, carotid ultrasound, CT scan of the heart and liver, arterial stiffness, and electrocardiogram (ECG). Focus on transforming the traditional paradigm of cardiovascular disease (CVD) prevention/treatment to one of precision medicine and early-stage detection, to promote cardiovascular health for South Asians. | Access requires submission of a proposal to the P-CARRS Publications, Presentations, and Ancillary Studies (PP&A) Committee. Proposals are reviewed for scientific merit and alignment with project goals. Access is granted for approved ancillary research related to the parent P-CARRS study. | 49 |
https://www.carrsprogram.org/precisioncarrs-1 https://ccdcindia.org/projects/precision-carrs/ |
| Rajiv Gandhi Cancer Institute & Research Centre | India | Cancer-focused biorepository collecting tumor tissues, blood samples, and clinical data. Supports oncology research and treatment optimization. Established infrastructure for sample processing and storage. | Access restricted to approved research projects. Requires ethical approval from institutional review board. Samples and data de-identified for research use. | 50 | https://www.rgcirc.org/biorepository/about-us/ |
| Registry of people with diabetes in India with young age at onset (ICMR–YDR) | India | 5546 participants (aged < 25) in Phase I and 709 serum samples in Phase II since 2006. Phase I collects demographics, clinical history, family history, biochemical details, treatment, and complications. Phase II collects blood samples for serum and genetic analysis. Focus on youth-onset diabetes | Data access available upon request from corresponding authors. May require institutional ethics approval for research use. | 51 | https://mdrf.in/department/ydr%20project.html |
| Tata Medical Center Biorepository (TiMBR) | India | Comprehensive cancer biorepository collecting diverse biological samples including fresh frozen tissues, blood, and derivatives. Links samples with detailed clinical annotations and follow-up data. | Access through formal application process. Requires scientific merit review and ethical approval. Collaborative agreements for data and sample sharing. | 9 | https://ttcrc.org/TiMBR.html |
| Translational Health Science and Technology Institute (THSTI) Biorepository Facility | India | ∼1,600,000 biospecimen, including samples of blood, urine and more, from ∼50,000 participants since 2015. Focus on pediatrics and COVID-19. | Access available through formal application process. Use of samples may require project approval and compliance with institutional guidelines. | 52 | https://biorepository.thsti.in/ |
| Community Based Intervention for Control of Hypertension in Nepal (COBIN) trial | Nepal | Initiated in 2015 and still ongoing, the COBIN cohort has completed four waves of data collection with approximately 3000 adult participants. The study collects detailed clinical and biochemical data, including blood pressure (systolic, diastolic, and heart rate), diabetes markers (fasting blood glucose, HbA1c), lipid profile (LDL cholesterol, HDL cholesterol, total cholesterol), thyroid function (TSH, T4, T3), serum creatinine, sodium, and potassium, together with medication adherence and socio-demographic variables. This community-based adult cohort provides a unique opportunity to study biomarkers of non-communicable diseases and their determinants, while fostering collaboration among Nepalese researchers and international scholars. | COBIN data are not publicly available. Participant consent allows for data to be shared for future analyses with appropriate ethics approval. Non-identifiable data and analysis code can be made available to researchers on submission of a reasonable request to the corresponding author. | 53, 54, 55 |
https://www.thelancet.com/journals/langlo/article/PIIS2214-109X(23)00214-0/fulltext https://www.thelancet.com/journals/langlo/article/PIIS2214-109X(17)30411-4/fulltext https://trialsjournal.biomedcentral.com/articles/10.1186/s13063-016-1412-3 |
| Nepal Biobank | Nepal | ∼3000 participants between 2014 and 2017. Collects blood pressure (SBP, DBP, heart rate) and diabetes (blood sugar). Focus on capturing existing discrete data on a broad range of noncommunicable disease risk factors and determinants, and fostering synergy and collaboration between Nepalese researchers locally, as well as scholars around the globe. | Access available through formal application process. Data are de-identified and password protected. Only collaborators may access and download data but are not permitted to modify the structure. | 56 | https://nedsnepal.org.np/nepal-biobank/ |
| Indus Hospital & Health Network Pediatric Cancer Biobank | Pakistan | Pakistan's first leukemia biobank focusing on pediatric cancers. Collects blood samples, bone marrow, and clinical data from pediatric patients. Addresses regional pediatric cancer research needs. | Patient confidentiality protected through encryption and coding. Investigators blinded to personal identifiers. Access requires strict ethical guidelines compliance and future use agreements. Data not publicly available due to privacy restrictions. | 57 | https://indushospital.org.pk/impact/newsroom/pakistans-firsts-leukemia-biobank/ |
| Sri Lankan Twin Registry Biobank (SLTR-b) | Sri Lanka | South Asia's first twin biobank established in 2015. Houses DNA and serum samples linked to longitudinal questionnaire data, clinical investigations, and anthropometric measurements. Part of Colombo Twin and Singleton Study. | Specimens linked to longitudinal health data. Designed to provide opportunities for academic collaborations. Aims to understand genetic and environmental contributions to health outcomes in South Asian populations. | 58 | https://ird.lk/twin-registry/ |
| South Asia Biobank (LOLIPOP cohort) | UK (South Asian focus) | ∼100,000 South Asian participants from LOLIPOP 2003 and 2020 cohorts. Collects lifestyle, environmental, sociodemographic, clinical, biochemical data, and biological samples (blood, urine). Focus on cardiovascular disease, type 2 diabetes, and obesity risk in South Asian populations. | Participants gave permission for long-term follow-up, including linkage to medical and health-related records, genomic studies, and recall by genotype/phenotype. Data captured electronically and stored on secure cloud-based server. Designed to understand high disease risk among South Asians globally. | 59,60 | https://www.sabiobank.org/ |
Critical lessons from these initiatives include the importance of standardized data structures and phenotyping protocols enabling meta-analyses across populations. Through public-private funding partnerships, sustainability is secured while community engagement strategies anchor them in lasting public trust.63 Scale and representativeness remain essential – statistical power to identify genetic variants or environmental risk factors with relatively small effect sizes requires hundreds of thousands of participants for common disease research.64
Integrated biobanks have become critical components of national research ecosystems, underpinning a nation’s health, scientific competitiveness, and economic development by contributing to the burgeoning bio-economy and realizing the transformative potential of data and personalized medicine.14,65,66 The undeniable global shift from individual or consortium-led collections to sophisticated, government-supported infrastructure enables the world to conduct data experiments and develop methods beyond geographical boundaries, a testament to global data democracy.
We surveyed geographic and ancestral representation in EHR and genetic studies – the two core data elements of integrated biobanks – revealing the stark quantitative reality of South Asia’s tepid presence in global biomedical infrastructure. Between 2014 and 2024, South Asia contributed just 1·8% of EHR publications indexed in Embase (Fig. 1), reflecting infrastructure limitations and constrained access to digitized health records essential for integrated biobanking studies. Simultaneously, analysis of genetic studies from 2005 to 2025 using NHGRI-EBI GWAS Catalog shows that while South Asians comprise 25% of the world’s population, they represent just 0·2% of participants in major genome-wide association studies – a 125-fold underrepresentation that persists despite growing awareness of diversity gaps.67
Fig. 1.
Limited Representation of South Asia in Health Data Research: EHR and GWAS Disparities. Panel A: Proportional representation of EHR publications indexed in Embase (Elsevier, 2025) by population region from 2014 to 2024, based on keyword search. Panel B: Cumulative number and proportional representation of individuals included in genome-wide association studies (GWAS) by ancestry group from March 10, 2005, to April 25, 2025, based on data from the NHGRI-EBI GWAS catalog.67
The global rise in integrated biobanks, because of their scientific successes and strategic importance, alongside stark data paucity in and underrepresentation of South Asia, points to a single, inevitable conclusion: South Asia needs coordinated intervention to establish population-scale biobanking infrastructure following blueprints that have repeatedly proven to be feasible and transformative (Table 1, top panel).
South Asia’s biobanking infrastructure sharply contrasts with its disease burden and population scale, with some areas experiencing an epidemiologic transition to increased prevalence of chronic conditions and others dealing with the dual burden of chronic and infectious diseases. While valuable, current efforts (Table 1, bottom panel) either operate at insufficient scale or lack harmonized mechanisms for broad and secure access. Pakistan’s pediatric cancer biobank68 and Bangladesh’s BANGABANDHU38 (atherosclerosis and ischemic heart disease) and MAGPIE39 (stroke) studies collect samples in the thousands rather than the hundreds of thousands required for robust genetic discovery. Due to sustained research investments, India has large prospective cohorts with rich data and biospecimens for pregnancy outcomes (GARBHINI,42 Pune Birth Cohort,69 New Delhi Birth Cohort70,71), aging (CBR-SANSCOG72), common non-communicable diseases (CARRS,49 Phenome India48), as well as disease-specific biobanks focusing on cancer,8,73,74 diabetes (InDiab44), liver diseases,10 and neurological disorders.75 These reach the necessary scale, some being amongst the largest cohorts of their type globally, but harmonized mechanisms for providing access to external investigators are yet to be implemented. Moreover, these collections are often not disease or exposure-agnostic like the UK Biobank.
Unique regional challenges compound this fragmentation:
-
•
Limited EHR adoption/penetration constrains phenotype data availability.76
-
•
Diverse consent practices across religious and ethnic communities require culturally adaptive approaches.77
-
•
Varying healthcare systems (single-payer vs. private vs. informal providers) affect data standardization.78
-
•
Language diversity presents technical challenges for data harmonization.79
-
•
Cross-border data sharing restrictions complicate regional coordination, particularly given geopolitical tensions,80 complicating data misuse and data colonialism concerns.
-
•
Weaknesses in biological specimen collection, transportation, and storage systems are challenges for routine clinical procedures in some parts of South Asia, likely making them difficult to operationalize for research purposes.
-
•
Concerns around data misuse, privacy breaches, or exploitation could affect public trust and willingness to participate.
Current regulatory and ethical frameworks could further exacerbate these challenges. India operates under general privacy guidelines rather than biobank-specific legislation, while harmonized networks for regional collaboration remain absent.81 Recent efforts in India have launched national data repositories that provide access to external researchers, including the Indian Biological Data Centre,82 the ICMR Health Research Data Repository,83 and MIDAS,84 alongside initiatives to aggregate and share diverse data in AI Kosh.85 These efforts do not obviate the need for a transnational or regional attempt; instead, they highlight the opportunity to leverage existing research efforts rather than starting afresh.
At the same time, the obstacles are not only infrastructural. Political economy shapes the region’s scientific marginalization. Concerns over extractive practices – where samples are exported, analyzed abroad, and monetized without fair returns of intellectual property or health benefits – remain widespread. India’s size and infrastructure risk overshadowing smaller nations within the region, raising fears of inequitable benefit distribution. Further, substantial regional diversity exists across the urbanicity gradient within a country, requiring vigilant determination to avoid opportunistic recruitment and maintain representation. Finally, proactive steps must be taken so biobanking efforts remain politically neutral and stable as political power shifts. Without explicit safeguards, a South Asian Biobank Consortium could exacerbate rather than resolve inequities.
Together, these realities highlight the need for technical expansion and a political and ethical architecture that ensures sovereignty, fairness, and inter-regional equity — a need the South Asia Biobank Consortium is designed to meet.
Rather than pursuing a single multinational mega-project, we propose a federated model comprising three interdependent layers: coordinated national biobanks operating under harmonized protocols for sample collection, data capture, consent, and quality control; a regional coordinating and governing body, the South Asia Biobank Consortium, serving as the institutional anchor for governance, benefit-sharing, and equitable representation; and a federated analytics infrastructure enabling cross-border research without transferring raw individual-level data across national boundaries (Fig. 2). This architecture balances regional cooperation with national sovereignty, recognizing political realities and diverse regulatory environments.
Fig. 2.
South Asian Biobank Consortium Implementation Roadmap (10-Year Timeline). Abbreviations: AI, artificial intelligence; IP, intellectual property; ISO, International Organization for Standardization; LIMS laboratory information management system; ML, machine learning; MPC, multi-party computation.
The South Asia Biobank Consortium would be the regional coordinating body to harmonize protocols for sample collection, processing, storage, and data management while allowing individual countries to maintain operational control of their biobanks.86 As the institutional anchor and governing body, the South Asia Biobank Consortium must guarantee meaningful representation to prevent potential domination from larger countries and economies and equitably earmark capacity-building funds. Explicit benefit-sharing agreements – intellectual property, licensing, and authorship – would help counter fears of data colonialism.
Operational benefit-sharing agreements are foundational, not optional. These should include co-authorship requirements for all publications using South Asian biobank data; joint intellectual property agreements with explicit provisions for any commercialized discoveries; mandatory capacity-building deliverables for international partners, including training, technology transfer, and infrastructure support; and a tiered data access model in which South Asian researchers retain a priority access window before broader international release, analogous to the exclusive-use period afforded to collaborators under the China Kadoorie Biobank model.87 Without these safeguards, the initiative risks replicating the inequities it aims to redress.
This approach is feasible. The GBMI demonstrates that diverse biobanks can collaborate successfully through harmonized standards and federated analytics without sharing raw data.11 Established cohorts that have successfully recruited South Asian populations, like the LOLIPOP cohort in the UK and several of India’s emerging repositories, can provide technical anchors and best practices for regional collaboration.
Five core elements would underpin implementation:
-
•
Harmonized regulatory frameworks that address broad consent, privacy protection, and data ownership while remaining culturally sensitive and adaptable to diverse legal systems across the region.
-
•
Advanced infrastructure investment, including state-of-the-art biobanking facilities, automated sample management systems, and high-performance computing platforms for genomic analysis.88
-
•
Human capital development through comprehensive training programs for bioinformaticians, genetic counselors, biobank managers, and research ethicists.89
-
•
Rigorous data security frameworks, including federated learning approaches in which only aggregated model parameters, never individual-level data, are shared across borders; differential privacy techniques for genomic summary statistics; and secure multi-party computation for analyses requiring shared inputs. ISO-standard security requirements would apply to all participating nodes, consistent with the approach taken by BioBank Japan.90
-
•
Sustainable funding mechanisms that ensure long-term viability and continuous technological advancement.
We propose a three-phase implementation plan to realise the South Asia Biobank Consortium (Fig. 2). Using the All of Us Research Program target sample size as a proportion of the US population as a guide (∼0·28%), we propose half of this proportion as a conservative target recruitment goal (i.e., averaging ∼0·14% of the country’s population). Across all South Asian Biobanks, we propose the target recruitment goal is 2,100,000 by the end of Phase 3 (i.e., 10 years). During the initial recruitment phase (Phase 2), the target sample size represents 10–20% of the final target, totaling 310,000 across all biobanks by the end of Phase 2 (i.e., 5 years).
-
•
Phase 1 (1–2 years): Secure political commitments and develop standardized regional frameworks for consent, ethics, IP, and benefit-sharing.
-
•
Phase 2 (3–5 years): Launch pilot biobanks – India (200,000 participants), Pakistan (50,000), Bangladesh (40,000), Nepal (10,000), and Sri Lanka (10,000). These should test equity safeguards such as co-authorship rules and joint intellectual property.
-
•
Phase 3 (5–10 years): Establish a federated regional network with secure cross-border analytics and structured global partnerships. Target enrollment attained: India (1,500,000), Pakistan (300,000), Bangladesh (200,000), Nepal (50,000), Sri Lanka (50,000).
Concrete research opportunities illustrate the value of this framework. A federated genome-wide association study of type 2 diabetes across Bangladeshi, Indian, Nepali, Pakistani, and Sri Lankan biobanks could identify South Asia-specific risk variants undetectable in European-ancestry cohorts, without transferring raw data across borders. Existing pregnancy cohorts such as GARBHINI and AMANHI could be harmonized under shared protocols to enable cross-border analyses of adverse birth outcomes at a scale sufficient for genetic, environmental, and gene-by-environment interaction discovery. A pharmacogenomics study of cardiovascular drug response could leverage the region's genetic diversity and broad exposure ranges to detect associations requiring far larger sample sizes in more homogeneous populations. These examples represent research that is currently infeasible due to fragmentation but would become tractable under the proposed framework.
Public trust represents the cornerstone of any successful biobank. Building and sustaining it across South Asia's diverse cultural, linguistic, and socioeconomic contexts requires specific mechanisms: community advisory boards with representation from religious, ethnic, and linguistic minority communities; dynamic consent models adapted to low-literacy and low-connectivity settings; community health worker-based recruitment to reach rural and informal-sector populations; multilingual outreach supported by frugal AI tools and platforms, such as India’s BHASHINI language translation initiative to enable participant communication across the region's hundreds of spoken languages; and the return of clinically actionable findings to participants, following models such as the Estonian Biobank.91,92 Empirical evidence from South Asia reinforces this point, as a qualitative study in Sri Lanka found that trust functioned as one of the key mediators of participants' consent to biobanking.93 Strategic international partnerships with established biobanks can provide invaluable knowledge transfer and technical assistance, structured to ensure equitable benefit-sharing and capacity building within South Asia. The initiative must also prioritize diseases with significant regional impact – adverse pregnancy outcomes, diabetes, cardiovascular disease, sickle cell disease, and specific cancers – to demonstrate early, tangible benefits, reinforcing public confidence and attracting continued investment.94
Building trust at recruitment is necessary but not sufficient - sustained engagement over time is equally essential. Digital platforms that enable participants to track summary findings, view their individual data in relation to regional trends, and follow publications arising from their contributions can reinforce the reciprocal relationship between biobank and community. Participant representation at annual meetings and in the Consortium's executive decision-making body should be considered a governance requirement, not an afterthought.
Grounded in these principles, we offer four concrete recommendations spanning community engagement, financing, international partnerships, and domestic funding mechanisms:
-
•
Governments and the South Asia Biobank Consortium should establish community advisory boards, implement dynamic consent, and mandate the return of clinically relevant results to participants.
-
•
National governments, in partnership with regional development banks and the private sector, should develop sustainable financing mechanisms, including researcher access fees (e.g., as practiced by the UK Biobank) and structured public-private partnerships (e.g., demonstrated by FinnGen), tied to Phase 1 milestone commitments.
-
•
International partners should operate under formal benefit-sharing agreements covering authorship, intellectual property, and capacity-building obligations, reviewed and enforced by the Consortium's Data Access Committee.
-
•
In India specifically, Corporate Social Responsibility mandates under Section 135 of the Companies Act (2013)95 represent an additional augmentation mechanism, with research and public health promotion among potentially eligible categories.96
This coordinated approach promises to transform South Asia from a peripheral participant to a central contributor in global biomedical research, ultimately enabling precision medicine tailored to the region’s unique genetic and environmental contexts while establishing the scientific infrastructure necessary for addressing the health challenges facing nearly two billion people. Moreover, this effort will allow global researchers to discover new knowledge and solutions based on data from a unique collection of ancestry, social, and cultural heritage that has remained under-analyzed due to data paucity.
Establishing coordinated national biobanks across South Asia represents far more than a regional scientific endeavor – it is a critical intervention for global health equity and economic development. The current genomic data imbalance perpetuates a system where precision medicine advances benefit primarily European-ancestry populations while systematically excluding nearly two billion South Asians from therapeutic innovations.97 This disparity undermines the fundamental promise of precision medicine: that genetic insights should inform better health outcomes for all populations, not just those historically privileged in research participation.
The economic implications are equally compelling. While nascent, South Asia’s precision medicine market represents significant economic potential, as the global market is projected to reach $250 billion by 2030.98 Countries investing in biobanking infrastructure are consumers of precision medicine technologies and innovators capable of developing therapeutics tailored to their populations’ unique genetic profiles. The broader bio-economy benefits extend beyond pharmaceuticals to encompass diagnostic development, health technology innovation, and the attraction of international research investments.
From a scientific discovery perspective, South Asian biobanks would unlock research opportunities that are unavailable elsewhere. The sheer population scale – nearly two billion people – transforms the landscape for genetic discovery: rare diseases that affect dozens of individuals elsewhere occur in the thousands here, while the extraordinary variation in environmental exposures provides statistical power to detect associations that would require much larger sample sizes in a population with more limited exposure ranges. The region’s genetic diversity, environmental exposures, and disease patterns offer unique windows into human health and disease mechanisms.99 Just as open-access biobank models, like the UK Biobank and the Lifelines Biobank27 (Netherlands), accelerated global scientific discovery by making the data available to the international research community, South Asian biobanks could contribute distinctive insights that advance understanding of complex diseases affecting populations worldwide.
In conclusion, the path forward requires coordinated action from multiple stakeholders. Governments must commit to sustained funding and regulatory harmonization. International development organizations should recognize biobanking as an essential health infrastructure that deserves support. The global research community must embrace partnerships that ensure equitable benefit-sharing and knowledge transfer. Without decisive action, the current disparities will only deepen as precision health advances, leaving South Asia’s population increasingly marginalized in the genomic revolution that promises to transform human health.
With a federated, equitable strategy, the region can become a central driver of biomedical discovery — and a model for how the rest of the world closes its data equity gaps while strengthening health systems and realizing the promise of precision medicine for nearly two billion people.
Contributors
MS: Writing – Original Draft, Writing – Review & Editing, Visualization. YW: Writing – Original Draft, Writing – Review & Editing, Visualization. PS: Writing – Original Draft, Writing – Review & Editing. BW: Writing – Original Draft, Writing – Review & Editing. HP: Writing – Original Draft, Writing – Review & Editing. VD: Writing – Original Draft, Writing – Review & Editing. NKB: Writing – Original Draft, Writing – Review & Editing. KJ: Writing – Original Draft, Writing – Review & Editing. NR: Writing – Original Draft, Writing – Review & Editing. RR: Writing – Original Draft, Writing – Review & Editing. FJ: Writing – Original Draft, Writing – Review & Editing. AY: Writing – Original Draft, Writing – Review & Editing. KA: Writing – Original Draft, Writing – Review & Editing. MD: Writing – Original Draft, Writing – Review & Editing. AA: Writing – Original Draft, Writing – Review & Editing. BM: Conceptualization, Writing – Original Draft, Writing – Review & Editing, Visualization. All authors: All authors have read and approved the final manuscript.
Declaration of interests
We declare no competing interests.
Acknowledgements
Kaushalya Jayaweera is fully supported by the Australian Government through the Australian Research Council’s Centre of Excellence for Children and Families over the Life Course (Project ID CE200100025).
Funding: We declare no funding for this work.
References
- 1.Sudlow C., Gallacher J., Allen N., et al. UK Biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age. PLoS Med. 2015;12(3) doi: 10.1371/journal.pmed.1001779. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.All of Us Research Program Investigators, Denny J.C., Rutter J.L., et al. The “All of Us” research Program. N Engl J Med. 2019;381(7):668–676. doi: 10.1056/NEJMsr1809937. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Offord C. UK biobank releases half a million whole-genome sequences for biomedical research. https://www.science.org/content/article/uk-biobank-releases-half-million-whole-genome-sequences-biomedical-research
- 4.Anjana R.M., Unnikrishnan R., Deepa M., et al. Metabolic non-communicable disease health report of India: the ICMR-INDIAB national cross-sectional study (ICMR-INDIAB-17) Lancet Diabetes Endocrinol. 2023;11(7):474–489. doi: 10.1016/s2213-8587(23)00119-5. [DOI] [PubMed] [Google Scholar]
- 5.Beesley L.J., Salvatore M., Fritsche L.G., et al. The emerging landscape of health research based on biobanks linked to electronic health records: existing resources, statistical challenges, and potential opportunities. Stat Med. 2020;39(6):773–800. doi: 10.1002/sim.8445. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Mohanty A. The Scientist. 2024. The evolution of biobanking and its role in precision oncology.https://www.the-scientist.com/the-evolution-of-biobanking-and-its-role-in-precision-oncology-71977 [Google Scholar]
- 7.Yuille M., Van Ommen G.J., Brechot C., et al. Biobanking for Europe. Brief Bioinform. 2007;9(1):14–24. doi: 10.1093/bib/bbm050. [DOI] [PubMed] [Google Scholar]
- 8.University of British Columbia Office of Biobank Education and Research Tata Medical Center Biorepository (TiMBR), India. https://biobanking.org/biobanks/view/867
- 9.Tata Translational Cancer Research Centre TiMBR: Tata medical center biorepository. https://ttcrc.org/TiMBR.html
- 10.Institute for Liver & Biliary Sciences National liver disease biobank (NLDB) https://nldb.in/
- 11.Zhou W., Kanai M., Wu K.H.H., et al. Global Biobank Meta-analysis initiative: powering genetic discovery across human disease. Cell Genom. 2022;2(10) doi: 10.1016/j.xgen.2022.100192. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Watson P.H., Nussbeck S.Y., Carter C., et al. A framework for biobank sustainability. Biopreserv Biobanking. 2014;12(1):60–68. doi: 10.1089/bio.2013.0064. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Allen N.E., Sudlow C., Peakman T., Collins R., on behalf of UK Biobank UK biobank data: come and get it. Sci Transl Med. 2014;6(224) doi: 10.1126/scitranslmed.3008601. [DOI] [PubMed] [Google Scholar]
- 14.Gottweis H., Petersen A.R., editors. Biobanks: governance in comparative perspective. Routledge; 2008. First publ. [Google Scholar]
- 15.Walters R.G., Millwood I.Y., Lin K., et al. Genotyping and population characteristics of the China Kadoorie Biobank. Cell Genom. 2023;3(8) doi: 10.1016/j.xgen.2023.100361. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.China Kadoorie Biobank Oxford population health. https://www.ckbiobank.org/front-page China Kadoorie Biobank.
- 17.Sørensen E., Christiansen L., Wilkowski B., et al. Data resource profile: the Copenhagen Hospital Biobank (CHB) Int J Epidemiol. 2021;50(3):719–720e. doi: 10.1093/ije/dyaa157. [DOI] [PubMed] [Google Scholar]
- 18.Von Wowern F., Høgdall E. Fast processing of gynecologic cancer tissue in Danish Cancer Biobank makes them well-suited for biomarker studies. APMIS. 2025;133(1) doi: 10.1111/apm.13481. [DOI] [PubMed] [Google Scholar]
- 19.Danmarks Nationale Biobank Danish National Biobank. https://www.danishnationalbiobank.com/ Statens Serum Institut.
- 20.Olsen J., Melbye M., Olsen S.F., et al. The Danish National Birth Cohort - its background, structure and aim. Scand J Public Health. 2001;29(4):300–307. doi: 10.1177/14034948010290040201. [DOI] [PubMed] [Google Scholar]
- 21.Leitsalu L., Haller T., Esko T., et al. Cohort profile: estonian biobank of the Estonian genome Center, university of Tartu. Int J Epidemiol. 2015;44(4):1137–1147. doi: 10.1093/ije/dyt268. [DOI] [PubMed] [Google Scholar]
- 22.University of Tartu Institute of Genomics . 2024. Estonian Biobank.https://genomics.ut.ee/en/content/estonian-biobank [Google Scholar]
- 23.Kurki M.I., Karjalainen J., Palta P., et al. FinnGen provides genetic insights from a well-phenotyped isolated population. Nature. 2023;613(7944):508–518. doi: 10.1038/s41586-022-05473-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.FinnGen FinnGen: an expedition into genomics and medicine. https://www.finngen.fi/en
- 25.Nagai A., Hirata M., Kamatani Y., et al. Overview of the BioBank Japan project: study design and profile. J Epidemiol. 2017;27(3S):S2–S8. doi: 10.1016/j.je.2016.12.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.BioBank Japan BioBank Japan. https://biobankjp.org/en/#gsc.tab=0
- 27.Scholtens S., Smidt N., Swertz M.A., et al. Cohort profile: Lifelines, a three-generation cohort study and biobank. Int J Epidemiol. 2015;44(4):1172–1180. doi: 10.1093/ije/dyu229. [DOI] [PubMed] [Google Scholar]
- 28.Boomsma D.I., Geus E.J.C.D., Vink J.M., et al. Netherlands twin register: from twins to twin families. Twin Res Hum Genet. 2006;9(6):849–857. doi: 10.1375/twin.9.6.849. [DOI] [PubMed] [Google Scholar]
- 29.Ligthart L., Van Beijsterveldt C.E.M., Kevenaar S.T., et al. The Netherlands twin register: longitudinal research based on twin and twin-family designs. Twin Res Hum Genet. 2019;22(6):623–636. doi: 10.1017/thg.2019.93. [DOI] [PubMed] [Google Scholar]
- 30.Hofman A., Murad S.D., Van Duijn C.M., et al. The Rotterdam Study: 2014 objectives and design update. Eur J Epidemiol. 2013;28(11):889–926. doi: 10.1007/s10654-013-9866-z. [DOI] [PubMed] [Google Scholar]
- 31.Taiwan Biobank 臺灣人體生物資料庫 官方網站. https://www.twbiobank.org.tw/
- 32.Feng Y.C.A., Chen C.Y., Chen T.T., et al. Taiwan Biobank: a rich biomedical research database of the Taiwanese population. Cell Genom. 2022;2(11) doi: 10.1016/j.xgen.2022.100197. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.The 100,000 Genomes Project Pilot Investigators 100,000 genomes pilot on rare-disease diagnosis in health care — preliminary report. N Engl J Med. 2021;385(20):1868–1880. doi: 10.1056/nejmoa2035790. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.100,000 genomes Project. https://www.genomicsengland.co.uk/initiatives/100000-genomes-project Genomics England.
- 35.UK Biobank UK biobank. 2025. https://www.ukbiobank.ac.uk/ UK Biobank.
- 36.National Institutes of Health . 2025. All of Us Research Program.https://www.joinallofus.org/ [Google Scholar]
- 37.Aftab F., Ahmed S., Ali S.M., et al. Cohort profile: the alliance for maternal and newborn health improvement (AMANHI) biobanking study. Int J Epidemiol. 2022;50(6):1780–1781i. doi: 10.1093/ije/dyab124. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Ranjan R., Hasan M.K., Adhikary A. Bangladeshi atherosclerosis biobank and hub: the BANGABANDHU study. Int J Gen Med. 2024;17:2507–2512. doi: 10.2147/ijgm.s466706. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Ranjan R., Adhikary D., Barman S., et al. Multidimensional approach of genotype and phenotype in stroke etiology: the MAGPIE study. Health Sci Rep. 2024;7(12) doi: 10.1002/hsr2.70227. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Japan International Cooperation Agency . Kingdom of Bhutan Ministry of Health; 2023. Project for strengthening government capacity for using digital technology and data in Bhutan.https://openjicareport.jica.go.jp/pdf/12348769.pdf [Google Scholar]
- 41.Proposal Summary Bhutan health research portal. https://bhutanhrp.moh.gov.bt/index.php/hrp/search/viewProposal/79
- 42.Bhatnagar S., Majumder P.P., Salunke D.M., Interdisciplinary Group for Advanced Research on Birth Outcomes—DBT India Initiative (GARBH-Ini) A pregnancy cohort to study multidimensional correlates of preterm birth in India: study design, implementation, and baseline characteristics of the participants. Am J Epidemiol. 2019;188(4):621–631. doi: 10.1093/aje/kwy284. [DOI] [PubMed] [Google Scholar]
- 43.National Institute of Mental Health and Neurosciences Human brain tissue repository (HBTR) https://thenimhansbrainbank.in/brain-bank/
- 44.Anjana R.M., Pradeepa R., Deepa M., et al. The Indian council of medical research-India diabetes (ICMR-INDIAB) study: methodological details. J Diabetes Sci Technol. 2011;5(4):906–914. doi: 10.1177/193229681100500413. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Mhatre S.S., Bragg F., Panse N., et al. Cohort profile: indian study of healthy ageing (ISHA-Barshi) Int J Epidemiol. 2024;53(4) doi: 10.1093/ije/dyae079. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Indian Institute of Technology, Madras About: national cancer tissue BioBank. https://nctb.iitm.ac.in/about/index.html
- 47.National Liver Disease Biobank National liver disease biobank. https://nldb.in/
- 48.Phenome India Consortium. Sengupta S. Study research protocol for phenome India-CSIR Health Cohort Knowledgebase (PI-CHeCK): a prospective multi-modal follow-up study on a nationwide employee cohort. Public Glob Health. 2024 doi: 10.1101/2024.10.17.24315252. Preprint posted online October 19. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Nair M., Ali M.K., Ajay V.S., et al. CARRS surveillance study: design and methods to assess burdens from multiple perspectives. BMC Public Health. 2012;12(1):701. doi: 10.1186/1471-2458-12-701. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Rajiv Gandhi Cancer Institute & Research Centre About us: RGCIRC biorepository. https://www.rgcirc.org/biorepository/about-us/ Rajiv Gandhi Cancer Institute & Research Centre.
- 51.Praveen P.A., Madhu S.V., Mohan V., et al. Registry of youth onset diabetes in India (YDR): Rationale, recruitment, and current status. J Diabetes Sci Technol. 2016;10(5):1034–1041. doi: 10.1177/1932296816645121. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.THSTI Biorepository Facility. https://biorepository.thsti.in/
- 53.Neupane D., McLachlan C.S., Christensen B., Karki A., Perry H.B., Kallestrup P. Community-based intervention for blood pressure reduction in Nepal (COBIN trial): study protocol for a cluster-randomized controlled trial. Trials. 2016;17(1):292. doi: 10.1186/s13063-016-1412-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Neupane D., McLachlan C.S., Mishra S.R., et al. Effectiveness of a lifestyle intervention led by female community health volunteers versus usual care in blood pressure reduction (COBIN): an open-label, cluster-randomised trial. Lancet Glob Health. 2018;6(1):e66–e73. doi: 10.1016/S2214-109X(17)30411-4. [DOI] [PubMed] [Google Scholar]
- 55.Thapa R., Zengin A., Neupane D., et al. Sustainability of a 12-month lifestyle intervention delivered by community health workers in reducing blood pressure in Nepal: 5-Year follow-up of the COBIN open-label, cluster randomised trial. Lancet Glob Health. 2023;11(7):e1086–e1095. doi: 10.1016/S2214-109X(23)00214-0. [DOI] [PubMed] [Google Scholar]
- 56.Nepal Biobank - NEDS. 2021. https://nedsnepal.org.np/nepal-biobank/ [Google Scholar]
- 57.Indus Hospital & Health Network . 2025. IHHN Establishes Pakistan's First pediatric Acute Leukemia Biobank.https://indushospital.org.pk/impact/newsroom/pakistans-firsts-leukemia-biobank/ [Google Scholar]
- 58.Institute for Research & Development Sri Lankan twin registry. https://ird.lk/twin-registry/
- 59.South Asia Biobank South Asia biobank. https://www.sabiobank.org/
- 60.Song P., Gupta A., Goon I.Y., et al. Data resource profile: understanding the patterns and determinants of health in south Asians—The South Asia Biobank. Int J Epidemiol. 2021;50(3):717–718e. doi: 10.1093/ije/dyab029. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Chen Z., Chen J., Collins R., et al. China Kadoorie Biobank of 0.5 million people: survey methods, baseline characteristics and long-term follow-up. Int J Epidemiol. 2011;40(6):1652–1666. doi: 10.1093/ije/dyr120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Kaye J., Whitley E.A., Lund D., Morrison M., Teare H., Melham K. Dynamic consent: a patient interface for twenty-first century research networks. Eur J Hum Genet. 2015;23(2):141–146. doi: 10.1038/ejhg.2014.71. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Collins F.S., Manolio T.A. Merging and emerging cohorts: necessary but not sufficient. Nature. 2007;445(7125):259. doi: 10.1038/445259a. [DOI] [PubMed] [Google Scholar]
- 64.Visscher P.M., Wray N.R., Zhang Q., et al. 10 years of GWAS discovery: biology, function, and translation. Am J Hum Genet. 2017;101(1):5–22. doi: 10.1016/j.ajhg.2017.06.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Organisation for Economic Cooperation and Development . OECD; 2009. OECD Guidelines on Human Biobanks and Genetic Research Databases.http://www.rettdatabasenetwork.org/Guidelines%20databases.pdf [Google Scholar]
- 66.European Commission Directorate General for Research and Innovation . Publications Office; 2012. Biobanks for Europe: a Challenge for Governance.https://data.europa.eu/doi/10.2777/68942 [Google Scholar]
- 67.Cerezo M., Sollis E., Ji Y., et al. The NHGRI-EBI GWAS catalog: standards for reusability, sustainability and diversity. Nucleic Acids Res. 2025;53(D1):D998–D1005. doi: 10.1093/nar/gkae1070. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Aijaz J., Raza M.R., Sajid K.N., et al. From blueprint to biobank: leveraging expert recommendations for implementing change (ERIC) to pediatric cancer biobanking in Pakistan. PLoS One. 2025;20(5) doi: 10.1371/journal.pone.0321316. Sergi CM, ed. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Yajnik C.S., Fall C.H.D., Coyaji K.J., et al. Neonatal anthropometry: the thin–fat Indian baby. The Pune Maternal Nutrition Study. Int J Obes. 2003;27(2):173–180. doi: 10.1038/sj.ijo.802219. [DOI] [PubMed] [Google Scholar]
- 70.Richter L.M., Victora C.G., Hallal P.C., et al. Cohort profile: the consortium of health-orientated research in transitioning societies. Int J Epidemiol. 2012;41(3):621–626. doi: 10.1093/ije/dyq251. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Bhargava S.K., Sachdev H.S., Fall C.H.D., et al. Relation of serial changes in childhood body-mass index to impaired glucose tolerance in young adulthood. N Engl J Med. 2004;350(9):865–875. doi: 10.1056/NEJMoa035698. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Ravindranath V., SANSCOG Study Team Srinivaspura aging, neuro senescence and COGnition (SANSCOG) study: study protocol. Alzheimers Dement. 2023;19(6):2450–2459. doi: 10.1002/alz.12722. [DOI] [PubMed] [Google Scholar]
- 73.National Cancer Tissue Biobank National Cancer Tissue BioBank (NCTB) https://nctb.iitm.ac.in/about/index.html About.
- 74.Rajiv Gandhi Cancer Institute & Research Centre Biorepository. https://www.rgcirc.org/biorepository/
- 75.National Institute of Mental Health and Neurosciences Human brain tissue repository (HBTR) https://thenimhansbrainbank.in/brain-bank/
- 76.Yadav P. Health product supply chains in developing countries: diagnosis of the root causes of underperformance and an agenda for reform. Health Syst Reform. 2015;1(2):142–154. doi: 10.4161/23288604.2014.968005. [DOI] [PubMed] [Google Scholar]
- 77.Joly Y., Dalpé G., So D., Birko S. Fair shares and sharing fairly: a survey of public views on open science, informed consent and participatory research in biobanking. PLoS One. 2015;10(7) doi: 10.1371/journal.pone.0129893. Marsh V, ed. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Prinja S., Bahuguna P., Pinto A.D., et al. The cost of universal health care in India: a model based estimate. PLoS One. 2012;7(1) doi: 10.1371/journal.pone.0030362. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Qin L., Chen Q., Zhou Y., et al. A survey of multilingual large language models. Patterns. 2025;6(1) doi: 10.1016/j.patter.2024.101118. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80.Mittra J. Palgrave Macmillan US; 2016. The new health bioeconomy: R&D policy and innovation for the twenty-first century. [DOI] [Google Scholar]
- 81.Akash, Sarker S.P. Discrimination based on genetic information in south Asia: an exploratory study of constitutions and relevant laws. Int J Leg Inf. 2023;51(3):183–196. doi: 10.1017/jli.2024.3. [DOI] [Google Scholar]
- 82.Indian biological data centre. https://www.rcb.res.in/indian-biological-data-center-ibdc Regional Centre for Biotechnology.
- 83.Indian Council of Medical Research ICMR health research data repository. https://data.icmr.org.in/ Indian Council of Medical Research.
- 84.Maity D., Satish R., Jadeja D.A., et al. MIDAS: a new platform for quality-graded health data for AI-enabled healthcare in India. Nat Med. 2024;30(10):2704–2705. doi: 10.1038/s41591-024-03198-x. [DOI] [PubMed] [Google Scholar]
- 85.AIKosh. https://aikosh.indiaai.gov.in/
- 86.Knoppers B.M., Zawati M.H., Kirby E.S. Sampling populations of humans across the world: ELSI issues. Annu Rev Genomics Hum Genet. 2012;13(1):395–413. doi: 10.1146/annurev-genom-090711-163834. [DOI] [PubMed] [Google Scholar]
- 87.China Kadoorie Biobank CKB data access and sample preservation policy. 2024. https://www.ckbiobank.org/data-access/ckb-data-access-and-sample-preservation-policy
- 88.Vaught J., Rogers J., Carolin T., Compton C. Biobankonomics: developing a sustainable business model approach for the formation of a human tissue biobank. JNCI Monogr. 2011;2011(42):24–31. doi: 10.1093/jncimonographs/lgr009. [DOI] [PubMed] [Google Scholar]
- 89.De Souza Y.G., Greenspan J.S. Biobanking past, present and future: responsibilities and benefits. AIDS Lond Engl. 2013;27(3):303–312. doi: 10.1097/QAD.0b013e32835c1244. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.BioSample storage facilities. https://biobankjp.org/en/about/1996 BioBank Japan.
- 91.Critchley C.R., Nicol D., Otlowski M.F.A., Stranger M.J.A. Predicting intention to biobank: a national survey. Eur J Public Health. 2012;22(1):139–144. doi: 10.1093/eurpub/ckq136. [DOI] [PubMed] [Google Scholar]
- 92.Milani L., Alver M., Laur S., et al. The Estonian Biobank's journey from biobanking to personalized medicine. Nat Commun. 2025;16(1):3270. doi: 10.1038/s41467-025-58465-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 93.Jayasinghe K., Chamika W.A.S., Jayaweera K., et al. All you need is trust? Public perspectives on consenting to participate in genomic research in the Sri Lankan district of Colombo. Asian Bioeth Rev. 2024;16(2):281–302. doi: 10.1007/s41649-023-00269-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 94.Daar A.S., Singer P.A., Leah Persad D., et al. Grand challenges in chronic non-communicable diseases. Nature. 2007;450(7169):494–496. doi: 10.1038/450494a. [DOI] [PubMed] [Google Scholar]
- 95.Ministries/Departments in the Government of India India code: section 135. https://www.indiacode.nic.in/show-data?actid=AC_CEN_22_29_00008_201318_1517807327856§ionId=1326§ionno=135&orderno=139 Corporate Social Responsibility.
- 96.Ministries/Departments in the Government of India Schedule VII (CSR section 135) https://upload.indiacode.nic.in/schedulefile?aid=AC_CEN_22_29_00008_201318_1517807327856&rid=79
- 97.Martin A.R., Kanai M., Kamatani Y., Okada Y., Neale B.M., Daly M.J. Clinical use of current polygenic risk scores may exacerbate health disparities. Nat Genet. 2019;51(4):584–591. doi: 10.1038/s41588-019-0379-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 98.Precision medicine market size & outlook, 2030. https://www.grandviewresearch.com/industry-analysis/precision-medicine-diagnostics-therapeutics-market
- 99.Dokuru D.R., Horwitz T.B., Freis S.M., Stallings M.C., Ehringer M.A. South Asia: the missing diverse in diversity. Behav Genet. 2024;54(1):51–62. doi: 10.1007/s10519-023-10161-y. [DOI] [PMC free article] [PubMed] [Google Scholar]


