Skip to main content
Human Genomics logoLink to Human Genomics
. 2026 Jun 30;20:134. doi: 10.1186/s40246-026-01007-9

National genomic projects in Asia and Africa: a review

Aisha Hanaya Alsuwaidi 1,2,#, Mira Mousa 1,3,#, Michael Olbrich 1, Nour Al Dain Marzouka 1, Inken Wohlers 4, Saleh Ibrahim 1,5, Habiba Alsafar 1,2,✉
PMCID: PMC13543545  PMID: 42381083

Abstract

National genome projects (NGPs) are increasingly shaping precision medicine by improving representation of population-specific genetic diversity. This review compiles findings from NGPs across Asia and Africa, regions that remain underrepresented in global genomic databases despite their extensive demographic and genetic diversity. A total of 53 studies from 24 countries were identified to understand (1) the genomic approach utilized, (2) novel findings that have emerged, and (3) strategies for improving research in these regions. The NGPs implement population-based variome databases (20 NGPs), linear reference genome assemblies (8 NGPs), and graph-based pangenome assemblies (1 NGP). Novel variants ranged between 0.28% (China) and 19.6% (Iran), whereas rare variants accounted for up to 88.9% of the detected variants in the Chinese population. Each NGP documents its country’s evolutionary and migration history, which impacts disease frequency and pharmacogenomic variants. Clinically, NGPs revealed strong population stratification in disease-associated and pharmacogenomic variants. For example, the GJB2 rs72474224 hearing-loss variant ranged from 13% in Vietnam and 12% in Hong Kong to 0.0894% in Turkey, while the VKORC1 rs9923231 pharmacogenomic variant reached 89.2% in Taiwan but was 20%–25% in European-related Russian subpopulations. These findings demonstrate that clinically relevant allele frequencies, pathogenicity assessments, and drug-response markers differ substantially across ancestries. This review highlights ongoing efforts and strategies to enhance the representativeness of genomic data through NGPs in Asia and Africa. We also suggest future directions for national projects, including integrating family-based studies, multi-omic data, and standardized pipelines to accelerate discovery and support the equitable implementation of precision medicine.

Supplementary Information

The online version contains supplementary material available at https://doi.org/10.1186/s40246-026-01007-9.

Keywords: Pangenome, National genome projects, Variome, Pharmacogenomics, Imputation

Introduction

The rapid advancements in genomic research have been supported by the establishment of foundational variant databases, such as dbSNP, 1KGP database [1], ExAC [2], ClinVar [3], DGVa [4, 5], TCGA database [6], GME Variome Project [7], and gnomAD. These resources enable cataloging of genetic variants, creation of population-specific reference genomes, and facilitation of data sharing across populations. However, most of these efforts have focused on European populations, leading to gaps in our understanding of genetic variation in other regions.

National Genome Projects (NGPs) in Asia and Africa are working to close these gaps by building population-specific reference panels that reflect the genetic diversity of their communities. Developing tailored genomic panels often involves constructing variome databases to organize genetic variants for research and clinical applications. Larger and more diverse panels improve imputation accuracy, which can be assessed using metrics such as correlation, Imputation Quality Score (IQS), and concordance rate [8]. Countries such as China, India, and Korea have implemented population-specific imputation panels to improve accuracy, particularly for rare variants. Population-specific major allele reference genomes are also being developed by aligning genomic data with standard references and editing them to reflect the target population’s genetic profile, reducing disparities in genomic studies.

Linear reference assembly offers an alternative by constructing independent reference genomes through de novo assembly and scaffolding using paired-end read information or optical mapping, thereby minimizing biases inherent in standard reference [9]. Recent advances in genome assembly techniques, including de novo assembly and graph-based pangenomes, offer new ways to build more complete and accurate reference genomes. These methods can capture complex genetic variations and better represent populations that have been underrepresented in global studies [10].

Collectively, these genomic resources are transformative for addressing disparities in genomics, enhancing disease diagnosis, advancing personalized medicine, and contributing to rare disease research and drug development. This review aims to identify countries in Asia and Africa that have established NGPs and to examine how these initiatives contribute to the development of population-specific genomic resources, including variome databases, imputation panels, and reference assemblies. It also examines how NGPs support the discovery of population-specific genetic variants and their applications in fields such as pharmacogenomics. Finally, the review highlights future directions and key challenges in promoting equitable and inclusive genomic research across these regions.

National genome projects

A total of 53 studies from 24 countries (Supplementary Table 1; Supplementary Material for search strategy and eligibility criteria) in Asia and Africa were identified. Of these, 24 NGPs met the predefined inclusion criteria: (1) the establishment of a national genome project, (2) the publication of an associated peer-reviewed manuscript, and (3) geographic location within the Asian or African region. In addition to established NGPs, we included 2 early-stage or pilot studies (Algeria and Indonesia) to highlight gaps in genomic representation.

The NGPs are categorized into three categories based on their primary genomic output: (1) population-scale variome databases, (2) linear reference genome assemblies, and (3) graph-based pangenome assemblies. These categories reflect a progressive evolution in genomic research, moving from cataloging genetic variation toward constructing comprehensive and representative genomic references. Variome-based projects focus on the systematic identification and cataloging of genetic variants within populations, providing essential resources for allele frequency estimation, genotype imputation, and clinical variant interpretation. In contrast, linear reference genome assemblies aim to construct population-specific genomic references, often derived from one or a few individuals, to reduce reference bias inherent in global standards such as GRCh38. More recently, graph-based pangenome assemblies have emerged as a transformative framework, integrating multiple genomes into a unified structure that captures both shared and population-specific variation. Unlike linear references, pangenomes represent genomic diversity more comprehensively, improving variant discovery and enabling more accurate analysis across diverse populations.

While most studies employed a similar framework for participant selection, considering variables such as age, gender, and health status, there are variations in the scale of participant groups. Specific projects focus on ethnically homogeneous samples, while others enrolled diverse cohorts, and in some cases, participants were required to have national ties within a specified number of generations. The details on country-level variant statistics, sequencing methodologies, and bioinformatics tools are provided in Supplementary Tables 1–3. The countries included are further distinguished by their project outcomes (Fig. 1A) and sample selection methods (Fig. 1B). However, projects that did not report results and were limited to proposed frameworks or study designs were excluded from the supplementary tables.

Fig. 1.

Fig. 1

Overview of NGPs. a Overlap between variome, linear genome assembly, and graph-based pangenome studies. This Venn diagram illustrates the overlap between three major categories of genomic research: variome, linear genome assemblies, and graph-based pangenome approaches. b Geographic Distribution and Sampling Diversity of NGPs. This map illustrates the geographic distribution of national genomic initiatives across Asia and Africa, categorized by sampling diversity and study design. Countries are color-coded based on the nature of participant selection: diverse selection, random selection, and limited participant cohorts. ‘Diverse selection’ countries aim to include participants from various subpopulations within their borders. ‘Random selection’ countries include those that either randomly sample participants or focus on specific subgroups. ‘Limited participants’ countries involve at most 10 participants in their projects

Africa

Africa is home to the highest level of human genetic diversity worldwide, as it’s the world’s second-most populous continent [11]; therefore, large-scale genomic studies are essential to ensure these cohorts are accurately represented in precision medicine. Recent years have seen an increase in genomic research across the region, characterized by a dual approach: multiple independent national initiatives and large-scale collaborative efforts, such as the Human Heredity and Health in Africa (H3Africa) consortium [12–14], established in 2012. H3Africa has addressed global health inequalities by investigating the genomic and environmental determinants of both non-communicable disorders, such as stroke and kidney disease, and infectious diseases like HIV and tuberculosis in 30 African countries.

A central pillar of H3Africa is the development of robust scientific infrastructure and human capital within the continent. Key achievements include the creation of H3ABioNet [15], a pan-continental bioinformatics network, and the establishment of major regional biorepositories in Nigeria, Uganda, and South Africa. Furthermore, the consortium developed the H3Africa SNP array to better capture genetic variation specific to African populations [16]. In addition to other key projects such as AWI-Gen [17] and TrypanoGEN [18], which have generated high-quality datasets that improve regional representation in global research and lay the groundwork for future precision medicine tailored to African ancestries.

Complementing these efforts, the AGenDA project [19] is a multinational collaboration across nine African countries aiming to sequence at least 1,000 genomes. While AGenDA provides a vital foundational framework for genomic sovereignty and addresses the complexities of cross-border data sharing among Africa's over 2,000 ethnolinguistic groups, its current sample size remains a small fraction of the continent's immense diversity. Similarly, the African Genome Variation Project (AGVP) [20] offers a dense characterization of sub-Saharan genetic variation, demonstrating that local reference panels are essential for accurate genotype imputation, a capability that global panels like the 1000 Genomes Project often lack.

Northern Africa

Algeria

Algeria has not published a national genetic database; however, a nationwide study investigated Algeria's genetic landscape by analyzing the Y chromosome, mitochondrial DNA (mtDNA), and autosomal genome-wide markers across the four subpopulations within Algeria [21]. The researchers found that genetic heterogeneity in the region does not strictly correlate with geography or linguistic affiliation, as evidenced by some Berber groups showing high diversity while certain Arab groups appear more genetically isolated. The analysis reveals a complex mosaic of ancestral components, with a notable sexual bias: external gene flow into the region was predominantly female-driven, whereas the indigenous North African component is more frequent in paternally transmitted regions. Although this research focuses primarily on population history rather than on clinical database construction, it represents a foundational effort to map the Algerian genetic heritage and provides a critical starting point for future precision medicine initiatives in the region.

Egypt

Egypt has taken an important step in representing North African genetic diversity through the EgyptRef project, which established the first Egyptian de novo genome assembly [22]. This high-quality reference genome was derived from a male individual of Egyptian ancestry, and complemented by WGS data from 110 additional Egyptians to serve as a population-specific genomic resource for Egypt and the broader North African region. The project identified 143 subpopulations and 49 population-specific SNPs with potential clinical implications. Further analysis of 461 protein-coding genes associated with these SNPs revealed an enrichment in genes linked to obesity and metabolic traits. The authors further explored the influence of non-coding variants on gene expression and regulatory mechanisms, revealing 1,180 genes with haplotype-dependent expression patterns involving both common and rare variants. Finally, 22 Runs of Homozygosity (ROH) were identified in the Egyptian reference assembly individual, with 16 larger than 5 Mb, suggesting potential founder effects or consanguinity patterns relevant to disease mapping.

Morocco

The Moroccan Genome Project (MGP) represents the first large-scale WGS effort focused on the Moroccan population [23]. In its initial phase, the project sequenced the genomes of 109 healthy individuals from various regions across Morocco, identifying over 27 million genetic variants. The primary aim was to develop the Moroccan Major Allele Reference Genome (MMARG) to reliably identify genetic variants unique to individuals and common within the population.

Using MMARG for variant calling significantly improved accuracy over the standard GRCh38 reference, notably reducing false positives and improving the representation of Moroccan genetic diversity. The analysis revealed a complex population structure shaped by centuries of admixture, with substantial contributions from North African, European, Middle Eastern, and West African ancestries. The analyses identified pathogenic and likely pathogenic variants related to diseases common in Morocco, such as diabetes, cardiovascular conditions, and hereditary hearing loss. Additionally, the study highlighted a high rate of consanguinity through long runs of homozygosity and provided insights into mitochondrial and Y-chromosome haplogroups, reinforcing the region's unique genetic landscape.

Tunisia

The Genome Tunisia Project is designed to establish the first reference sequence of the Tunisian population and integrate precision medicine into the national healthcare system, targeting 10,000 genomes by 2035 [24]. The project aims to address the global underrepresentation of African genomic data by sequencing an initial pilot group of 100 healthy individuals. However, as the current paper focuses primarily on the strategic and ethical framework, it lacks the genomic results necessary to assess the actual state of genetic variation in Tunisia. While this foundational roadmap is essential, the project’s ultimate success will depend on its transition from a high-level research plan to an actual implementation of this framework.

Eastern Africa

Uganda

The Uganda Genome Resource (UGR) represents a foundational effort to capture the extensive genetic diversity within Uganda [25, 26]. The database included WGS data from ~ 2,000 individuals and genotyping from an additional ~ 5,000, and identified 9.5 million novel variants not previously documented in global databases such as the 1000 Genomes Project. In addition, the study identified 43 distinct genetic signals associated with 34 cardiometabolic traits. Notably, it also revealed differences in trait heritability between African and European populations, such as higher heritability for LDL-cholesterol in Ugandans (54%) compared to estimates from European studies (20–43%), likely reflecting differences in genetic architecture and environmental influences.

The UGR highlights the value of conducting genomic studies in underrepresented populations, demonstrating that such efforts can uncover functionally important, population-specific variants that are often missed in European-focused research. However, within the UGR, a large proportion of participants belong to the Baganda ethno-linguistic group, which may limit how well the findings represent the broader Ugandan population, which includes more than 40 distinct ethnic groups. In addition, the WGS was performed at relatively low coverage (~ 4×), which may reduce sensitivity for detecting rare or low-frequency variants.

Western Africa

Nigeria

The Nigerian 100K Genome Project, led by the NCD-GHS consortium, is a large-scale initiative designed to address the critical lack of African genomic data [27]. By leveraging Nigeria’s immense diversity of over 300 ethnic groups, the project aims to create a comprehensive catalog of human genetic variation from 100,000 adults. This research specifically targets the rising burden of non-communicable diseases (NCDs), such as cardiovascular and metabolic disorders. To date, the researchers have recruited 75,000 individuals and completed the initial sequencing of 1,000 samples to describe the genetic architecture and to develop a custom, low-cost genotype array optimized for African populations.

While the project represents an important step toward strengthening genomic infrastructure in Nigeria, the current publication primarily outlines its strategic framework rather than presenting detailed analytical results. It focuses on building systems for ethical sample collection, biobanking, and local genomic analysis, but does not yet report findings from the sequenced individuals. Instead, the authors highlight a broader five-point agenda centered on sustainable capacity building and enabling African researchers to lead future advances in genomics.

Asia

Genomic research in Asia has rapidly expanded over the past decade, driven by the recognition that global genomic datasets remain heavily biased toward European populations. To address this, large-scale initiatives such as the GenomeAsia 100K Project have been established to capture the extensive genetic diversity across the continent. The pilot phase of this project generated whole-genome sequencing data from 1,739 individuals across 219 population groups in 64 countries, highlighting the remarkable heterogeneity and complex population structure within Asia [28].

One of the key contributions of these initiatives is the identification of novel genetic variants and population-specific allele frequency patterns that are not represented in global reference databases. For example, the GenomeAsia dataset revealed that approximately 23% of protein-altering variants were absent from widely used resources such as gnomAD and the 1000 Genomes Project, emphasizing the critical need for region-specific genomic resources. In addition, hundreds of thousands of variants were identified at frequencies sufficient to affect disease association studies, demonstrating the extent of previously uncharacterized genetic variation.

Complementing these efforts, regional initiatives such as the Greater Middle East (GME) Variome Consortium provide deeper insights into specific subpopulations characterized by unique demographic and genetic features [7]. The GME variome, based on whole-exome sequencing of over 1,100 individuals, revealed extensive genetic diversity shaped by historical migration, admixture, and high rates of consanguinity. Notably, the study demonstrated significantly elevated levels of runs of homozygosity and an increased burden of recessive variants, reflecting the genetic impact of consanguineous marriage practices. Importantly, incorporating this population-specific dataset reduced candidate disease variants by 4–sevenfold in rare disease studies, illustrating its direct clinical utility. Together, these large-scale and region-specific efforts highlight both the progress and the remaining gaps in genomic research across Asia.

Central Asia

Kazakhstan

The Kazakh genome project represents the first WGS effort focused on individuals of ethnic Kazakh descent, aiming to build a foundational genomic reference for Central Asia [19]. The initiative sequenced the genomes of five healthy ethnic Kazakh individuals at high coverage (26 × to 32 ×) to ensure comprehensive representation, and the reads were subsequently aligned to the GRCh37 human reference genome. The resulting genomes were fully assembled, annotated, and made publicly available to support future research. As the first whole-genome resource for the Kazakh population, this study provides a critical foundation for exploring genetic diversity, population structure, and disease-associated variants in the region.

South-eastern Asia

Indonesia

Indonesia has initiated efforts, such as the Indonesian Genome Diversity Project, to address the underrepresentation of its population in global genomic datasets [29]. This initiative has generated high-coverage (∼30×) whole-genome sequencing data from 161 healthy individuals sampled across multiple islands, capturing a broad spectrum of genetic diversity within the archipelago. However, this effort currently represents a research dataset rather than a fully established national genomic program, and, to date, lacks comprehensive peer-reviewed publications detailing its findings.

Singapore

Singapore’s national genomic efforts have established an advanced framework for precision medicine by leveraging the nation's unique multi-ethnic composition. Teo et al. [30] established an early foundation with the Singapore Genome Variation Project (SGVP), which created a publicly available haplotype map of 268 individuals from Chinese, Malay, and Indian ethnic groups in Singapore. This initiative cataloged 1.6 million SNPs to characterize genomic variation, linkage disequilibrium, and recombination rates in Singapore, providing a resource similar in format to the International HapMap Project. The study’s results highlight significant population structures within Singapore and identify signatures of positive natural selection in several gene clusters, including those involved in skin pigmentation (SLC24A5), alcohol metabolism (ADH), and brain development (CENPJ and MCPH1).

These efforts were significantly expanded by Wu et al. [31] through the SG10K project, which performed large-scale WGS on 4,810 Singaporeans. The researchers identified ~ 52 million novel variants and provided a comprehensive snapshot of genetic diversity across East, Southeast, and South Asia. In addition, 20 candidate loci under natural selection were identified, with 14 harboring robust associations with complex traits and diseases. Moreover, the researchers developed an imputation panel and demonstrated substantial improvements in accuracy for all ethnicities within Singapore, Asian populations, and Oceanian populations.

Building on these resources, Tan et al. developed a detailed catalog of structural variation across 8,392 ancestrally diverse Asian genomes as part of the SG10K_Health initiative. The study identified 73,035 structural variants (SVs), of which approximately 65% (47,770 SVs) were novel, highlighting the vast amount of genetic diversity that remains uncaptured by existing global reference databases. Importantly, researchers investigate the functional and clinical significance of these variants by identifying SVs affecting clinically actionable loci using a specialized in-house algorithm that showed superior sensitivity for detecting tandem duplications and insertions compared to standard tools. This approach enabled a deeper exploration of how SVs affect protein-coding integrity, identifying variants with direct, predicted impacts on essential genes. Furthermore, the study established that many of these SVs are in LD with SNPs, suggesting that some previously identified disease-associated SNP signals in Asian populations may reflect underlying structural variants rather than the SNPs themselves.

Thailand

The Genomics Thailand (GeTh) initiative represents a national effort to integrate genomics into public healthcare, targeting five key domains: rare diseases, cancer, non-communicable diseases, pharmacogenomics, and infectious diseases [32]. A central component supporting this initiative is the development of population-specific reference resources, such as the Thai Reference Exome (T-REx) database, which was created to address the underrepresentation of Thai individuals in global genomic datasets [33]. Together, these efforts highlight a coordinated strategy that combines large-scale data generation with direct clinical application in a middle-income setting. Shotelersuk et al. reported completing WGS for its target of 50,000 participants in the GeTh project, capturing diverse representation across Thailand and enabling clinically actionable findings in areas such as rare diseases and cancer.

Complementing this, the T-REx database includes exome sequencing data from 1,092 unrelated Thai individuals and reports 345,681 variants, of which approximately 18.27% are novel. The researchers performed population genetic analyses, which revealed clear substructure within the Thai population. In addition, Functional annotation identified thousands of variants with potential biological impact, including high-impact variants affecting protein function and population-specific pathogenic variants in genes. Based on the high frequencies of certain pathogenic variants in the database, such as rs77358650 (p.V444I) in the SPTAN1 gene and p.V37I in the GJB2 gene, the researchers question their actual pathogenicity. Moreover, the researchers focused on the prevalence of G6PD-deficiency and thalassemia variants in the cohort.

Boonin et al. present a trio-based whole-genome sequencing study aimed at characterizing genetic variation within the Thai population by analyzing 40 family trios (120 individuals) [34]. The study identified over 20.2 million variants, including approximately 1.1 million novel variants not previously reported in dbSNP. The trio-based design enabled more accurate detection of de novo mutations, resulting in the identification of 19,710 such variants, including a pathogenic mutation in the SF3B2 gene associated with craniofacial microsomia. Additionally, the study reported 169 pathogenic variants, many of which are rare or absent from existing databases such as ClinVar, highlighting population-specific genetic risks. Functional analyses revealed that many variants are associated with diseases, including cancer, metabolic disorders, and hematological conditions such as thalassemia. Overall, this study provides a valuable population-specific genomic resource and demonstrates the utility of trio-based sequencing for improving variant detection and informing future genetic screening and precision medicine efforts in Thailand.

While the projects in Thailand show significantly advanced genomic research, the initiative faces several structural and technical limitations. The GeTh paper primarily presents a framework and early outcomes rather than detailed analytical results, particularly for complex and infectious diseases, which remain in exploratory stages. Furthermore, the T-REx cohort may be subject to selection bias, as it was established using exome data from individuals recruited through clinical settings, including patients with rare diseases and their family members. This recruitment strategy can lead to an enrichment of pathogenic variants and introduce bias in allele frequency estimates when compared to the general population. Additionally, the initial reliance on exome sequencing for these resources inherently limits the detection of non-coding variants and complex structural rearrangements, a gap that the project aims to address in its second phase through the increased use of whole-genome and long-read sequencing.

Vietnam

Vietnamese genomic efforts have primarily leveraged KHV individuals from the 1KGP and additional population-specific datasets to enhance reference resources and understand functional genetic variation. Thanh et al. analyzed 100 unrelated KHV individuals, dividing them equally into training and testing sets to assess the performance of a Vietnamese-specific reference genome [35]. The results, particularly from chromosome 20, demonstrated improved read mapping and genotype calling accuracy when using the Vietnamese reference. These findings were validated across multiple genomes, indicating the pipeline’s applicability beyond the Vietnamese population.

Hai et al. focused on a parent–child trio of KHV individuals without known genetic disorders [36]. Starting with high-quality de novo contig assemblies, they identified ‘Mendelian-supported’ contigs by aligning the child’s genome to the parents’ genomes. A functional analysis of nonsynonymous SNPs in the trio revealed a subset of potentially damaging missense mutations across specific genes [36]. Among 1,955 genes linked to GO terms, 20 enriched biological processes were identified, with emphasis on “transcription, DNA-templated”, “RNA metabolic process”, “RNA biosynthetic process”, and “cellular nitrogen compound biosynthetic process”. Twelve genes were associated with all 20 enriched GO terms, highlighting core pathways affected by Vietnamese-specific nonsynonymous SNPs.

Tran et al. analyzed 2,683 pregnant Vietnamese women who underwent noninvasive prenatal testing between 2018 and 2019 [37]. By cross-referencing the NIPT call set with the ClinVar database, the study explored the associations between Vietnamese genomic variants and prevalent genetic diseases. The genetic variation analysis revealed close clustering of the NIPT group with the KHV population, reflecting the Vietnamese demographic and geographic characteristics. This study emphasizes the potential of NIPT-derived data for population-scale genetic disease surveillance and precision health initiatives.

Southern Asia

Bangladesh

A preliminary study of the Bangladeshi population represents an initial step toward developing a genomic reference through WGS of four individuals to identify potentially relevant variants and assess their functional and disease-related impact [38]. The manuscript provides a detailed technical summary of the sequencing results, including read mapping and coverage metrics for each sample. A significant portion of the analysis focuses on identifying protein-coding genes with high variant density. Furthermore, functional annotation and pathway analysis suggest that many of these variants are associated with biological processes related to metabolism, immune response, and disease pathways.

However, these findings should be interpreted with caution, as statistical support for many associations is limited, and the small sample size restricts broader conclusions. Overall, the study offers preliminary insights into genomic variation in the Bangladeshi population, but larger and more representative cohorts will be necessary to validate these observations and establish robust population-level genetic associations.

India

India’s national genomics initiative has advanced significantly with the launch of the IndiGen program, which aims to catalog genetic variation in the highly diverse Indian population [39]. This initiative performed WGS of 1,029 healthy individuals. The database includes 55,898,122 single-allelic genetic variants, of which over 18 million are unique.

The Western-Indian Reference Panel (WIP) was established using genome-wide SNP data from 407 individuals from Western India [8]. WIP merged SNPs from genome-wide data on 407 individuals from WIP with the 1KGP3 reference panel, creating a comprehensive dataset comprising 931,371 high-quality autosomal SNPs. The WIP demonstrated improved imputation accuracy, especially for low-frequency and rare variants, highlighting its value for enhancing GWAS and fine-mapping studies in Indian populations.

Iran

Iran’s national genomic efforts have been advanced by the creation of Iranome, a population-specific genetic database representing eight major Iranian ethnic groups. Developed by Fattahi et al., the project used WES from 800 healthy individuals [40] to catalog genetic diversity across the country. Population structure analysis positioned Iranians as a genetically distinct cluster intermediate between Europeans and South Asians, with specific populations such as Baluchs, Persian Gulf Islanders, and Turkmen linking to other populations. Additionally, four Iranian subpopulations exhibited 1,446 longer ROHs associated with consanguinity, while others showed more ROHs, reflecting higher inbreeding in these subpopulations. Variant-level analysis of 50,030 ClinVar-annotated sites identified 721 pathogenic variants, of which 92.6% were rare (AF ≤ 1%) in the Iranian population. Importantly, 12 pathogenic variants previously linked to rare Mendelian disorders were identified in healthy Iranians, underscoring the need to re-evaluate variant pathogenicity across different population contexts. Using the Iranome dataset, several variants in genes KDM5C, ALDOB, KIF1A, SCN8A, WFS1, PAH, GPR161, ZC3H14, and CHD4 were reclassified from Pathogenic to Variants of Uncertain Significance (VUSs) based on updated ACMG guidelines. This reclassification demonstrates the importance of incorporating ethnically diverse genomic data into clinical variant interpretation to avoid misclassification and overdiagnosis.

Sri Lanka

The Sri Lankan Genome Variation Database (SLGVD) catalogs SNPs across the three major ethnic groups of Sri Lanka [41]. The database serves as a repository for genetic data from published and unpublished research on Sri Lankan populations conducted both locally and abroad. The SLGVD contained genotype frequencies for 34 genomic variants across 14 medically important genes, including those associated with conditions such as preeclampsia and cardiovascular disease.

Furthermore, the Sri Lankan Personal Genome Project provides an overview of a national effort to establish baseline data for whole-genome sequencing in Sri Lanka [42]. As a proof of concept, researchers sequenced the complete genome of an anonymous Sinhalese male and identified approximately 2.8 million SNPs, of which 7.9% (222,739 variants) were novel compared to the dbSNP build 131 database. While this initiative represents a landmark entry for Sri Lanka into the era of genomics, the current scope of the published work is primarily descriptive and conceptual. The project relies on a single personal genome, which limits the generalizability of its findings on specific variants to the broader, ethnically diverse population.

Western Asia

Kuwait

As part of the Kuwaiti genome project (KWP1), Thareja et al. analyzed genomic data from a Kuwaiti individual of Persian ancestry, providing a reference genome for Kuwaiti individuals of Persian ancestry [43]. The study demonstrated strong genotype–phenotype concordance, with several identified variants aligning with the individual's medical history, including Type 2 diabetes and β-thalassemia. The analysis revealed multiple disease-associated genetic variants, including alleles linked to decreased birth weight, metabolic syndromes, and neurological traits. One of these variants, rs5758511, found in the CENPM gene, has implications for metabolic regulation. Further, three out of 28 deleterious SNPs linked to Type 1 diabetes, cognitive functions, and triglyceride levels were homozygous for the associated allele. The genome also carried three β-Thalassemia-associated alleles, 76 of 169 alleles associated with migraines, and 72 alleles related to type 2 diabetes, all commonly found in South and East Asian populations. This pattern reflects a mixed ancestral background, suggesting a blend of Asian and European components among disease-associated SNPs. Lastly, the study introduced a user-friendly genome browser for detailed data exploration.

Qatar

The Qatar Genome Program (QGP) analyzed whole-genome data from 6,218 Qatari participants enrolled through the Qatar Biobank (QBB) population cohort [44]. Population genetic analysis of 6047 individuals revealed cluster-specific variants, including disease-causing recessive alleles associated with structural and developmental disorders, particularly in populations with high levels of consanguinity. However, some common alleles appear non-pathogenic and exceed disease prevalence.

The study highlighted key disease-associated variants, such as NM_025000.4(DCAF17):c.436delC (p.Ala147fs), a variant linked to Woodhouse–Sakati syndrome and enriched in the Peninsular Arabs subcluster. This variant was present in 88 individuals and correlated with decreased insulin levels and increased diabetes risk. Remarkably, its carrier frequency in Qatar (2.5%) is the highest globally, suggesting a strong founder effect. OMIM genes were more frequent in these ROH regions than non-OMIM genes, with Peninsular Arabs having significantly more OMIM genes than other populations.

Further expanding on these results, Razali et al. developed an imputation panel from the QGP dataset to improve the accuracy of genotype imputation for the Arab and Middle Eastern population [45]. The panel showed an increase in imputed variants, especially among rare and common alleles.

United Arab Emirates

The United Arab Emirates (UAE) has launched several foundational genomic efforts to establish a population-specific reference genome, including what is potentially one of the largest population-level initiatives globally: sequencing the entire Emirati population of one million citizens. AlSafar et al. presented the first whole-genome sequences of two Emirati nationals, each sequenced at over 27 × coverage, identifying over 4 million variants per individual, including SNPs and indels [46]. The study identified clinically relevant variants associated with metabolic disorders, including diabetes, hypertension, and obesity, providing a valuable foundation for developing a UAE-specific reference panel to support precision medicine initiatives.

Daw Elbait et al. established a population-specific major allele reference genome (UAERG) [47] using data from the 1,000 Arab Genome Project [48]. The team selected 129 representative Emirati individuals, with an additional 33 samples sequenced by WES, and retained a total of 153 high-quality samples. The study identified 1,669 variants with AF greater than 1%, resulting in loss of gene function. Among these, 1,033 were indels, and 625 were SNPs. Some of these LoF variants are found in genes (MUC6, ZNF717, SPEN, and STK33) linked to different cancer types. Additionally, three other genes were associated with inborn genetic disorders: NRP2 (Hirschsprung disease 1), STAG (Stag2-related disorder), and HTT (Huntington’s Chorea). Notably, 15 variants with AFs > 5% were identified as common in the UAE population, while infrequent in gnomAD populations.

The set identified 137,713 SVs across various types present in at least five individuals. Furthermore, pathway enrichment analysis of population-specific variations found significant associations with metabolism-related pathways. LoF variants showed enrichment in pathways related to stem cell regulation and MAPK activation. The study revealed diversity within the samples, with a significant portion showing a strong Middle Eastern genetic component. In contrast, others exhibited admixture with populations from Central/South Asia and Sub-Saharan Africa. This diversity was reflected in mitochondrial and Y-DNA haplogroups, demonstrating varied maternal and paternal influences.

Saudi Arabia

Saudi Arabia constructed a telomere-to-telomere (T2T) in de novo genome assembly for a Saudi individual, known as KSA001, to better capture the genetic architecture of the Saudi population, led by Kulmanov et al. [49]. The assembly process involved utilizing PacBio HiFi and ONT reads, resulting in multiple assemblies. The best assembly was generated from HiFi reads using Hifiasm [50]. Moreover, different scaffolding methods were employed to correct misassemblies, place contigs within their respective chromosomes, and fill gaps.

Furthermore, contigs were aligned to T2T-CHM13 to refine the assembly. Most of the gaps were found in centromeres, which contain highly repetitive regions and are generally challenging to align and assemble. These centromeric regions were assembled separately, and the gaps in the non-centromeric regions were successfully closed, with exceptions.

Moreover, the ONT-based assembly alignment revealed the gaps in other chromosome regions. These regions were merged with the main assembly. Further refinement involved polishing regions from the ONT-based assembly with quality-filtered Illumina reads to reduce base errors. The study performed alignment and comparison against T2T-CHM13, revealing differences, including SNPs, indels, and structural variants. The results show that variant calling using KSA001 as a reference showed fewer variants than T2T-CHM13, particularly for Saudi individuals with consanguineous parents. Further, KSA001 captured major alleles in the Saudi population better than T2T-CHM13, paving the way for improved genetic screening, disease risk prediction, and precision healthcare in Saudi Arabia.

Turkey

The Turkish (TR) Variome Project, led by Kars et al., analyzed 3,362 unrelated individuals, incorporating WES for 2,589 and WGS for 773 participants within the Turkish population [51]. Turkish variants were categorized by functional impact, resulting in seven main groups. Additionally, admixture was employed to dissect Turkey’s genetic structure, identifying four major ancestral components.

A maximum-likelihood phylogenetic tree was constructed to comprehensively evaluate genetic relationships, linking the TR population to global populations. This approach visually represented global connections and migration patterns, contextualizing the TR population's genetic history within the broader human genetic landscape. Finally, Y chromosome and mtDNA haplogroup analyses were conducted to trace paternal and maternal genetic lineages. These analyses identified Central Asian-specific haplotypes, providing insights into the historical gene flow and ancestral origins of the Turkish population.

Eastern Asia

China

China has led multiple large-scale genomic initiatives aimed at capturing the genetic diversity of its population and improving population-specific resources. Du et al. conducted WGS of 597 healthy individuals and produced a high-quality de novo assembly of a Northern Han genome (NH1.0) [52]. Similarly, Zhang et al. developed the NyuWa genome resource by deep sequencing 2,999 Chinese individuals, resulting in a reference panel with 5,804 haplotypes and over 19 million variants, supporting imputation and disease-variant mapping efforts [53].

Li et al. expanded the scope with the Chinese Millionome Database (CMDB), which is based on WGS data from 141,431 unrelated individuals across 31 administrative divisions [54]. CMDB includes over 9.04 million SNVs and enables researchers to efficiently search for variants, genes, or genomic regions to retrieve information on mutation characteristics, allele frequencies, genic annotations, and frequency distributions across global populations. Yang et al. presented a fully phased, telomere-to-telomere (T2T) diploid human genome assembly (CN1) from a Han Chinese male, revealing 11,413 structural variations (SVs), many of which were previously unreported [16]. Complementing this, Lan et al. performed deep sequencing (~ 80×) of 90 unrelated individuals, identifying 26,000 SVs, surpassing the ~ 7700 SVs reported by the 1000 Genomes Project (1KGP)—thereby enriching the catalog of low-frequency and novel variants in Chinese genomes [55].

The Westlake BioBank for Chinese (WBBC), presented by Cong et al., 2022, developed a population-specific reference panel and an imputation server [56]. The study demonstrated that the WBBC panel improved imputation accuracy for low-frequency and rare variants, especially when merged with the EAS subset of the 1KGP panel. WBBC also identified four Han Chinese subgroups and reported 1,842 ClinVar-listed pathogenic variants—most of which were rare (97.4%). A pathogenic variant in the FECH gene (c.315-48T > C), rare in Europeans ( MAF 0.06), was found to be common in Chinese populations (MAF > 0.3). The study also investigated the genetic makeup of 1,151 healthy individuals, revealing 732 pathogenic and likely pathogenic variants. Additionally, the study showed that alleles associated with alcohol-metabolism genes (ADH1A and ADH1B) in East Asia have become more common over the past 4,000 years.

Gao et al. constructed a Chinese pangenome using 58 core samples from 36 minority ethnic groups and eight linguistic groups [10]. By generating high-quality phased diploid assemblies aligned to the T2T-CHM13 reference, the study improved alignment for East Asian samples and uncovered ethnic group-specific variations. The pangenome revealed novel sequences and facilitated the annotation of functional elements. This resource enhances understanding of genomic regions linked to disease susceptibility, immune function, and phenotypic traits.

Hong Kong

The Hong Kong Genome Project (HKG) represents a key milestone in characterizing the genetic landscape of the Hong Kong Cantonese population [57]. As the first fully accessible variant database for this group, the project used whole-exome sequencing (WES) data from 205 individuals residing in the Hong Kong SAR. HKG identified 799 high-impact novel variants in 731 coding genes that were unique to the Hong Kong population. Genetic structure analysis demonstrated the distinctiveness of the Hong Kong Cantonese population relative to other Chinese subgroups, aligning with geography and confirming Hong Kong’s place within the East Asian genetic landscape. Finally, the study assessed imputation accuracy and documented significant improvements when using HKG data as a local reference panel alongside 1KGP samples.

Japan

Japan has established a comprehensive and multifaceted approach to developing population-specific genomic resources. One of the key milestones was the development of the 1KJPN reference panel, created by Kawai et al. through WGS of 1070 individuals [58]. Based on this dataset, they developed the ‘Japonica array,’ a custom SNP genotyping array designed to improve imputation accuracy in Japanese populations. Their results showed that genotype imputation using the 1KJPN panel outperformed the 1KGP panel, particularly for low-frequency variants.

In a complementary effort, Higasa et al. developed the Human Genetic Variation Database (HGVD), a publicly accessible resource generated from WGS data from 1208 healthy Japanese individuals [59]. The database catalogs over 21 million single-nucleotide variants (SNVs), providing detailed allele frequency data and supporting both genetic and clinical applications by improving variant interpretation in Japanese individuals.

To address limitations of reference genome bias, Takayama et al. constructed the JG1 reference genome, a fully de novo assembly from three unrelated Japanese male individuals [60]. To avoid ethnic biases, the JG1 was generated independently of the Genome Reference Consortium (GRC) [61]. The JG1 genome was generated through a thorough process that included deep sequencing, optical mapping, and meta-assembly strategies. JG1demonstrated high sequence similarity with GRCh38 but carried population-specific variants that better reflect the Japanese genetic landscape.. The genome captured numerous SNP sites with an AF of 1.0 across the three Japanese individuals, suggesting that JG1 may often carry the major alleles in the Japanese genetic makeup.

SV analysis using JG1 revealed distinct patterns unique to Japanese individuals, and the team identified novel sequences absent in GRCh references. Clinical relevance was assessed by analyzing exomes from 22 individuals across seven Japanese families affected by rare diseases. JG1-based analysis successfully identified all known causal variants, reduced false positives, and improved overall accuracy compared to GRCh37. Finally, the potential of JG1 as a reference for WGS analysis was assessed. This entailed mapping WGS short reads from 1,070 Japanese individuals to JG1 and comparing allele frequencies with those of GRCh37, revealing a high degree of equivalence between the two references for autosomal SNP sites across a broad range of allele frequencies.

Flanagan et al. developed and evaluated four Japanese-specific reference panels by merging the 1KGP dataset with data from the Biobank Japan project [62]. These panels aimed to improve imputation accuracy for GWAS on the Japanese population, particularly for rare variants with low minor allele frequency. Among these, the augmented panel with the highest sample size exhibited the highest imputation quality, significantly outperforming the widely used reference panel, especially for rare variants. Overall, Japan’s genomic initiatives, including the creation of reference genomes, imputation panels, and variant databases, represent a robust national effort to improve precision medicine and disease gene discovery within a well-defined population context.

Republic of Korea

Korea’s genomic efforts have evolved into a framework of population-specific resources to improve precision medicine and genomic representation. The first Korean genome, “SJK,” sequenced by Ahn et al., revealed 420,083 novel SNPs and 5.77% unmapped regions, suggesting the presence of uncharacterized genomic sequences not represented in the global database [63]. Cho et al. developed a consensus Korean reference genome, termed KOREF, by sequencing and assembling a high-coverage de novo genome using data from 40 Korean genomes [64]. KOREF captures population-specific variants, reduces mapping bias, and improves the accuracy of Korean variant detection.

The KRG project, in its pilot phase, analyzed WGS data from 1,490 individuals and demonstrated that the KRG reference panel outperforms other reference panels in imputation accuracy for the Korean population [65]. The study combined multiple datasets (TOPMed, GAsP, and KRG) for meta-imputation but excluded ChinaMAP and NARD due to technical constraints. Moreover, the researchers assessed imputation using genotype panels mirroring the Korea Biobank Array and the UKB array for Korean and European populations. Population structure was studied using admixture analysis with 5,388 variants, and the results were visualized alongside subpopulation information from the KRG and 1KGP3 datasets. Interpopulation differences in LD were investigated, with 5,635,734 shared variants between KRG and European samples from 1KGP3. High LD variation regions were identified by selecting regions with varLD scores exceeding the top 1% of standardized varLD scores. PCA analysis of Korean KRG samples compared to multi-ethnic 1KGP3 samples revealed their close genetic relationship with nearby East Asian populations like the Japanese and Chinese. While they shared genetic similarities, Korean samples exhibited greater genetic homogeneity than other East Asian subpopulations, as confirmed by admixture analysis and consistent with previous research [66, 67].

Further refinement was achieved through de novo assembly and haplotype phasing by Seo et al., who resolved 89% of genes and highlighted key genomic regions, including the MHC and clinically relevant genes like CYP2D6 [68]. Their approach generated a remarkably contiguous assembly, closing numerous gaps in the human reference genome GRCh38 and revealing 18,210 structural variants, many of which were previously unreported and specific to the Asian population. Haplotype phasing analysis covered 89% of genes and highlighted key genomic regions, such as the MHC, and clinically relevant genes, including CYP2D6.

Jeon et al. introduced Korea1K, a dataset of 1,094 whole genomes (~ 31 × depth) paired with 79 clinical traits, further advancing genotype–phenotype correlations and imputation accuracy [69]. Furthermore, Korea1K emerges as a promising reference resource, demonstrating improved imputation accuracy for Koreans compared with existing panels and filtering germline variants in cancer samples. Similarly, Jung et al. developed the KRGDB, which includes data from 1,722 individuals and over 32 million unique genetic variants [70], as well as GWAS analyses identifying associations with conditions such as diabetes and hypertension. The KRGDB, available online, aids in understanding Korean genetic diversity, supporting research on diseases, and aiding in primer design. Kim et al. expanded this effort through KoVariome, a detailed catalog of Korean-specific genetic variants, enabling cross-population comparisons and identification of rare pathogenic mutations [71].

Taiwan

The Taiwan Biobank (TWB) integrates high-coverage WGS data from over 1,400 individuals of Han Chinese descent with SNP array genotyping data from a large cohort of 103,106 participants. In silico validation using an additional 137 Han Chinese individuals assessed imputation accuracy, focusing on variants with an MAF exceeding 0.01. The TWB-phased reference panel has improved accuracy over existing reference panels, including the 1KGP EAS and a combined TWB-EAS reference panel. The demographic analysis revealed insights into the TWB cohort’s population structure and historical changes. The study indicated isolation-by-distance trends, with individuals from distant Chinese provinces showing greater genetic differentiation from Taiwanese Minnan. The TWBv2 SNP array was designed to identify functional genetic variants, including Mendelian disorders, complex disease markers, drug metabolism, and HLA region variants.

Moreover, the imputation of ABO blood groups and HLA types demonstrated high accuracy and revealed associations with clinical phenotypes, providing valuable information for healthcare and research. Also, regional variation in HLA allele frequencies was observed, reflecting genetic diversity within the population. The study found that 21.2% of the TWB cohort are carriers of at least one Mendelian recessive disorder, highlighting the prevalence of rare genetic conditions, cancer-susceptibility mutations, and pharmacogenomic variants that affect drug responses.

Further enriched by linkage to Taiwan’s National Health Insurance Database and other national registries, the TWB, as described by Feng et al., includes data from over 150,000 individuals, offering a powerful resource for biomedical and public health research in East Asian populations, particularly those of Han Chinese ancestry [72].

Northern Asia

Russia

The Russia Genome Project aimed to construct a comprehensive, population-specific genomic resource based on WGS data from 264 individuals across the Russian Federation. Zhernakova et al. conducted the study, which included 204 individuals from 52 diverse populations encompassing Russian and other ethnic groups [73], of which 31 were from Mallick et al. [74], 173 were from Pagani et al. [75], and 60 newly sequenced individuals represented three distinct populations in western Russia and eastern Siberia. The samples were collected from family trios, ensuring at least three generations of homogeneous ancestry within the same ethnic group and region. The analysis identified clear geographic clustering: Yakuts aligned genetically with other Siberian populations, whereas Novgorod and Pskov resembled Northeastern Europeans. Ancestry-informative markers and F3 statistics highlighted admixture patterns, including both European and East Asian ancestry components, and identified genetic barriers to gene flow in western and northeastern Siberia, as well as at Russia’s Far Eastern border.

Clinically, the project identified 894 medically relevant gene variants, including 31 unique pathogenic variants associated with diseases such as age-related macular degeneration and Charcot-Marie-Tooth disease. The study also identified 758 high-confidence LoF SNPs, of which 101 were novel. Some LoF variants exhibited significant differences in allele frequency across Russian populations compared to others. Additionally, the study examined variants associated with lactose tolerance, warfarin response, skin pigmentation, and retinitis pigmentosa, revealing distinct allele frequencies among populations impacting personalized medicine.

Variants with medical implications

One of the most consistent findings across National Genomic Projects (NGPs) in Asia and Africa is the high prevalence of population-specific genetic variants, many of which have direct clinical relevance. The effective integration of these variants into clinical practice requires a comprehensive understanding of their allele frequencies, penetrance, and genotype–phenotype relationships across diverse populations. In parallel, pharmacogenomics has emerged as a key pillar of precision medicine, examining how genetic variation influences individual responses to therapeutic agents. A major advance in this field is the identification of ancestry-specific markers and stratified population profiles that shape drug metabolism and efficacy. Accordingly, this section highlights clinically relevant genetic variants reported across multiple NGPs in Asia and Africa, focusing on both disease-associated variants and pharmacogenomic markers. Supplementary Table 4 provides the list population-specific variant–phenotype associations, and demonstrated in Fig. 2.

Fig. 2.

Fig. 2

Population-specific variation in allele frequencies of clinically relevant genetic variants across national genome projects in Asia and Africa. Heatmap illustrating the allele frequency distribution of selected disease-associated and pharmacogenomic variants across multiple populations. Each row represents a genetic variant, and each column represents a country or population-specific dataset. Color intensity corresponds to allele frequency, with darker shades indicating higher frequencies and lighter shades indicating lower frequencies; missing data are shown in light blue. Substantial heterogeneity is observed across populations, with several variants demonstrating marked enrichment in specific regions

Disease-associated variants

Non-syndromic hearing loss

Hearing loss is one of the most common sensory disorders worldwide, and genetic factors account for approximately 50% of cases [76]. Within this spectrum, autosomal recessive non-syndromic hearing loss is predominantly driven by mutations in the GJB2 gene [76, 77]. Here, we report four variants associated with this gene reported in multiple NGPs. The frequency of the rs72474224 (c.109G > A/p.V37I) variant shows remarkable diversity across regions, with prevalence rates of 13% in Vietnam (NIPT), 12% in Hong Kong, 9.2% in Thailand (T-REx), 4.5% in Morocco, and only 0.0894% in Turkey. These disparities demonstrate the complex relationship between genetics, geography, and population history.

Metabolic disorders, hematological, and rare disease variants

Variants associated with metabolic regulation and organ-specific disease risk show clear population stratification across NGPs. A key example is the HFE rs1799945 variant, which is associated with hereditary hemochromatosis. This variant leads to increased intestinal iron uptake, resulting in progressive iron overload and potential damage to the liver, pancreas, and heart. Its allele frequency varies substantially across populations, with a higher prevalence in Turkey (12%) than in Vietnam (5.1%), reflecting differences in disease burden and historical genetic drift.

Another important variant is TBC1D31 rs10101626, associated with diabetic kidney disease and modulation of urinary uromodulin levels, a biomarker of renal function. This variant shows variability across Russian subpopulations, with a high frequency in the Yakut population (71.4%) compared to Novgorod (10%) and Pskov (17.9%), and at an intermediate level in Turkey (25%). Such differences may reflect both environmental pressures and genetic adaptations that influence susceptibility to metabolic and renal disorders.

Hematological and enzymatic disorders

Variants affecting red blood cell physiology and enzymatic pathways are highly enriched in certain populations, particularly in Southeast Asia, where selective pressures such as malaria have historically shaped allele frequencies [78]. For instance, HBA2 rs41464951, associated with alpha thalassemia and Hemoglobin H disease, affects hemoglobin synthesis and can lead to varying degrees of anemia depending on zygosity. This variant is more frequent in Thailand (2.7%) than in China (0.0446%), reflecting regional enrichment of hemoglobinopathies. Similarly, HBB rs33950507, associated with beta thalassemia and hemoglobin E disease, demonstrates elevated frequencies in Thai populations (11.8%) but is very rare in the Qatari dataset (0.05%). These variants significantly impact oxygen transport and are major contributors to inherited anemia in Southeast Asia.

Variants in the G6PD gene (rs72554664 and rs72554665) are also of major clinical importance, as they impair the pentose phosphate pathway and reduce red blood cells' ability to manage oxidative stress. This can lead to hemolytic anemia, particularly after exposure to certain drugs or infections. These variants are more prevalent in Thailand (1.3%) and Hong Kong (2.4%) compared to Turkey (0.0189%).

Rare Mendelian and developmental disorders

The FAH rs11555096 variant is linked to tyrosinemia type I, a metabolic disorder caused by impaired tyrosine catabolism, leading to the accumulation of toxic metabolites and liver dysfunction. This variant shows a higher frequency in Russia–Pskov (14%) compared to Turkey (0.9%), suggesting localized enrichment. Similarly, the PDE11A rs76308115 variant is associated with adrenocortical hyperplasia and endocrine dysregulation. Its higher prevalence in Russia–Pskov (9.1%) compared to Turkey (0.1%) indicates population-specific risk patterns. Additionally, SLC22A18 rs78838117, linked to rhabdomyosarcoma, is more frequent in Thailand (11.2%) and Hong Kong (8.5%) than in Turkey (0.5%), suggesting regional differences in susceptibility to certain malignancies.

Pharmacogenomic variants

Pharmacogenomics explores the complex link between genetics and drug responses. A key development in this field is the identification of unique ancestry markers. These markers reveal genetic variations in specific populations, providing insight into how individuals from diverse ancestries respond to medications. By delving into ancestry-specific genetic profiles, pharmacogenomics aims to improve drug prescriptions, tailoring treatments to individual genetics. This approach optimizes therapeutic outcomes and emphasizes the need to account for diverse genetic influences to achieve more equitable healthcare solutions.

For instance, knowledge of the distribution of variants associated with drug metabolism, such as those in the VKORC and UGT1A1 genes, can facilitate medication regimens tailored to specific populations. The VKORC gene influences drug metabolism and responsiveness, and we observe intriguing population-specific trends for its variants rs9923231 and rs2884737. The AF of the rs9923231 variant in Taiwan stands at 89.2%. However, in Russia, distinct patterns emerge within subpopulations: the Pskov and Novgorod groups, closely tied to Europe, exhibit AFs of 25% and 20%, respectively, while the Yakuts, sharing borders with Asia, show a substantially higher AF of 86%. These disparities emphasize that genetic diversity is not uniformly distributed globally, which may affect drug metabolism [79].

Discussion

Regional comparison of genomic initiatives: scale, maturity, and translation

National genomic initiatives across Asia and Africa reveal substantial heterogeneity in scale, technological maturity, and translational impact. Several Asian programs, particularly in East Asia (China, Japan, and Korea) and selected high-resource settings such as the UAE and Qatar, are characterized by large cohort sizes, often exceeding tens of thousands of individuals, coupled with high-coverage whole-genome sequencing and rapid adoption of advanced genomic frameworks. This includes telomere-to-telomere (T2T) assemblies (led by China and Saudi Arabia), population-specific imputation panels, and graph-based pangenomes. These approaches enable not only comprehensive variant discovery but also improved resolution of complex genomic regions and structural variation. In contrast, many African initiatives, despite representing populations with the highest global genetic diversity, remain in earlier stages of large-scale implementation. These efforts are frequently focused on establishing foundational infrastructure, building local capacity, and generating initial datasets, as exemplified by H3Africa and the Nigerian 100K Genome Project, while also providing unique insights into human genetic diversity and population history.

Differences in study design further contribute to variability in outcomes across countries. Large-scale population cohorts, such as those established in China, Korea, Taiwan, the UAE, and Qatar, provide robust allele frequency estimates and support genome-wide association studies, imputation accuracy, and disease mapping. In contrast, smaller pilot studies in countries such as Bangladesh, Kazakhstan, and Sri Lanka offer preliminary insights but remain limited in statistical power and generalizability. Additionally, the choice of sequencing technology significantly influences the scope of discovery. Projects relying on whole-exome sequencing are constrained to coding regions, limiting the identification of regulatory variants and structural rearrangements. In contrast, high-coverage whole-genome and long-read sequencing approaches enable more comprehensive characterization of rare variants, copy number variations, and previously inaccessible genomic regions. In addition, some countries, such as Egypt, Iran, and India, have implemented geographically distributed sampling strategies across different regions to improve the national representativeness of genomic datasets.

Another key distinction lies in the degree of clinical integration. Several Asian initiatives have successfully translated genomic findings into healthcare applications. For example, Thailand’s Genomics Thailand initiative and the Qatar Genome Program demonstrate the integration of genomic data into national healthcare systems, particularly in pharmacogenomics, rare-disease diagnostics, and population screening. In contrast, many initiatives in Africa and parts of South Asia remain primarily research-driven, with limited integration into clinical practice, reflecting broader systemic challenges, including infrastructure limitations, regulatory frameworks, and resource availability. Collectively, these examples demonstrate how NGPs are beginning to influence three key domains: improving diagnostic yield, enabling ancestry-specific risk prediction, and supporting the implementation of pharmacogenomics in clinical care.

Despite these differences, common themes emerge across regions. Nearly all initiatives emphasize the importance of population-specific reference datasets to address the limitations of global databases that are heavily biased toward European populations. The development of variome databases, imputation panels, and reference genomes consistently improves variant interpretation, reduces false-positive findings, and enhances the accuracy of genetic association studies. Moreover, several projects have demonstrated that variants previously classified as pathogenic in global databases may be benign in specific populations, highlighting the need to incorporate ethnically diverse data into clinical variant classification frameworks.

Building genomic capacity in LMICs: barriers and opportunities

Genomic research in LMICs faces a range of systemic and structural barriers, including limitations in infrastructure, human capital, ethical and legal frameworks, and long-term funding sustainability. Many LMICs lack essential research infrastructure, such as advanced sequencing facilities, large-scale biobanks, and the high-performance computing required for bioinformatics analyses. As a result, several countries either lack active NGPs or rely on external institutions for sequencing and genotyping services.

In parallel, there is a critical shortage of trained professionals capable of leading and sustaining genomic initiatives. This challenge is further compounded by the migration of skilled researchers to higher-income countries, limiting local capacity development. Additionally, many LMICs lack comprehensive regulatory frameworks for the collection, storage, and sharing of sensitive genomic data, raising concerns about data governance, privacy, and equitable access. Although initiatives such as H3Africa provide foundational support, long-term sustainability often depends on increased national investment and the development of public–private partnerships.

Despite these challenges, NGPs present significant opportunities to transform healthcare systems in LMICs. Global genomic databases such as GWAS catalogs remain heavily skewed toward European populations (estimated at ~ 87%). In contrast, African populations are represented by only ~ 3% and Asian populations by ~ 5 [80], highlighting a substantial gap in diversity. Emerging NGPs in countries such as Nigeria, Uganda, and Tunisia are beginning to address this imbalance by capturing region-specific genetic variation and expanding representation in global datasets.

NGPs in LMICs are increasingly bridging the gap between large-scale genomic data generation and real-world clinical application by aligning research priorities with national healthcare needs. In diagnostics, these projects enhance accuracy by developing population-specific reference variomes, reducing reliance on European-derived datasets, and enabling the identification of previously uncharacterized pathogenic variants in rare and undiagnosed diseases. By addressing key barriers, these initiatives support the gradual integration of precision medicine into public health systems.

From data generation to clinical translation: the next phase of NGPs

Genomic research across Asia and Africa is progressing mainly by building population-specific reference genomes and expanding large-cohort studies that combine genetic and phenotypic data. While these efforts align with global trends, there are still clear gaps. Advances such as T2T genome assemblies (China and Saudi Arabia), the application of long-read sequencing (Saudi Arabia and UAE), high-coverage WGS (Vietnam, UAE, Qatar), and machine learning-based imputation models (Qatar, Taiwan) are improving how we discover variants and understand structural changes in the genome. As sequencing costs drop and technology improves, more diverse, high-quality reference genomes will become available, enabling better, more personalized analyses across ethnic groups [10, 81].

The next phase of NGPs lies in transitioning from linear reference genomes to graph-based pangenomes and T2T assemblies. Using graph-based pangenomes combines multiple reference genomes and variant datasets, offering a fuller picture of genetic diversity and being important for population-specific studies. Emerging projects, such as China’s T2T human pangenome initiative and the MMARG, demonstrate that these high-resolution approaches significantly improve variant calling accuracy. Importantly, T2T assemblies enable the characterization of previously inaccessible regions of the genome, including centromeres and telomeres, which may harbor population-specific variants with clinical relevance.

Population characteristics in many regions further provide unique opportunities for genomic discovery. High rates of consanguinity observed in several Middle Eastern and North African populations are associated with an increased prevalence of homozygous loss-of-function variants, offering valuable insights into Mendelian and rare diseases. In parallel, national biobanks with linked clinical and health data serve as powerful platforms for long-term investigations into aging, resilience, and gene–environment interactions. Emerging frameworks, such as the genome–exposome approach adopted in Kuwait, represent an important step toward more comprehensive models of human health [82].

Pharmacogenomics is a critical application of population-specific genomics, as several national projects have identified variants that influence drug metabolism, which vary across populations. Future directions include building robust, ancestry-informed drug response databases, improving predictive models for adverse drug reactions, and supporting the clinical implementation of pharmacogenetic testing. As more diverse populations are studied, we can expect to uncover new variants that will help guide safer, more effective prescribing, particularly in underrepresented populations where current clinical guidelines fall short [15, 40].

Despite these advances, several critical gaps persist. Family-based studies, such as trios, twin, and multi-generational designs, are largely lacking in the region, despite their importance for understanding inheritance and rare mutations; hence, further effort is needed in this area. Regional genomic research should prioritize building rare disease atlases, expanding multi-omic and longitudinal phenotypic studies, and validating polygenic risk scores in diverse populations. In addition, integrating AI and digital tools will enhance variant interpretation and risk prediction, but requires interoperable data systems, regulatory support, and economic evaluation. Precision medicine remains in its early stages of rollout, with limited clinical uptake and few cost-effectiveness studies. Emerging technologies, such as single-cell and spatial genomics, CRISPR functional screens, and synthetic long reads, hold promise for uncovering rare and regulatory variants but are not yet widely used. Deeper collaboration, data standardization, and equitable access will be essential to translate research into meaningful healthcare outcomes.

Ultimately, stronger international collaboration, standardized data-sharing practices, and equitable access to genomic technologies will be essential to translate these advances into meaningful healthcare outcomes. By addressing these challenges, NGPs in Asia and Africa can play a central role in advancing a more inclusive and globally representative precision medicine framework.

Supplementary Information

Supplementary Material 1. (43.1KB, xlsx)
Supplementary Material 2. (23.2KB, docx)

Acknowledgements

This research was supported by the Center for Biotechnology, Khalifa University of Science and Technology (KU-BTC).

Abbreviations

T2T

Telomere-to-telomere

HGP

Human Genome Project

1KGP

1000 Genomes Project

1KGP3

1000 Genomes Phase 3

GRC

Genome Reference Consortium

CPC

Chinese Pangenome Consortium

ExAC

Exome Aggregation Consortium

DGVa

Database of Genomic Variants Archive

TCGA

Cancer Genome Atlas

NGP

National genome program

IQS

Imputation quality score

WES

Whole exome sequencing

WGS

Whole genome sequencing

AF

Allele frequency

MAF

Minor allele frequency

ROH

Runs of homozygosity

VUS

Variants of uncertain significance

PTV

Protein-truncating variants

SV

Structural variants

BAC

Bacterial artificial chromosome

MHC

Major histocompatibility complex

LoF

Loss of function

WSS

Woodhouse–Sakati syndrome

mtDNA

Mitochondrial DNA

PCA

Principal component analysis

Fplink

Inbreeding coefficient

NIPT

Non- invasive prenatal testing

WBBC

Westlake BioBank for Chinese

EgyptRef

Egyptian reference genome

HKG

Hong Kong

SAR

Special Administrative Region

WIP

Western India-specific panel

KRG

Korean Reference Genome

KRGDB

Korean Reference Genome Database

Korea1K

Korean Genome Project

QGP

Qatar Genome Project

QBB

Qatar Biobank

TWB

Taiwan Biobank

UAERG

UAE reference genome

GME

Greater Middle East

CHN

Chinese

TR

Turkish

EUR

Europe

BLK

Balkan

CAU

Caucasus

KHV

Kinh Vietnamese

Author contributions

HA, AHA, NADM, RMO, and MM conceptualized the project. AHA extracted the papers, curation, data analysis, and wrote the original draft. MM contributed to the main manuscript text. All authors reviewed and edited the manuscripts.

Funding

This study was funded by the Abu Dhabi Executive Office granted to Khalifa University under budget ID:KU-EXT-2022-8434000474, and supported by the Center for Biotechnology, Khalifa University of Science and Technology (KU-BTC).

Data availability

No datasets were generated or analysed during the current study.

Declarations

Ethics approval and consent to participate

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Aisha Hanaya Alsuwaidi and Mira Mousa contributed equally to this work.

References

  • 1.The 1000 Genomes Project Consortium, et al. A global reference for human genetic variation. Nature. 2015;526:68–74. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Exome Aggregation Consortium, et al. Analysis of protein-coding genetic variation in 60,706 humans. Nature. 2016;536:285–91. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Landrum MJ, et al. ClinVar: improving access to variant interpretations and supporting evidence. Nucleic Acids Res. 2018;46:D1062–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Database of Genomic Variants archive <EMBL-EBI. https://www.ebi.ac.uk/dgva/.
  • 5.Lappalainen I, et al. dbVar and DGVa: public archives for genomic structural variation. Nucleic Acids Res. 2012;41:D936–41. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.The Cancer Genome Atlas Research Network, et al. The Cancer Genome Atlas Pan-Cancer analysis project. Nat Genet. 2013;45:1113–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Greater Middle East Variome Consortium, et al. Characterization of Greater Middle Eastern genetic variation for enhanced disease gene discovery. Nat Genet. 2016;48:1071–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Ahmad M, et al. Inclusion of population-specific reference panel from India to the 1000 Genomes phase 3 panel improves imputation accuracy. Sci Rep. 2017;7:6733. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Pop M, Kosack DS, Salzberg SL. hierarchical scaffolding with bambus. Genome Res. 2004;14:149–59. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Gao Y, et al. A pangenome reference of 36 Chinese populations. Nature. 2023;619:112–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Population by Continent 2026.
  • 12.Mulder N, et al. H3Africa: current perspectives. Pharmacogenomics Pers Med. 2018;11:59–66. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Sengupta D, Choudhury A, Ramsay M. H3Africa: a model for implementing biobank-based genomic research in resource-constrained settings. Hum Mol Genet. 2025. 10.1093/hmg/ddaf113. [DOI] [PubMed] [Google Scholar]
  • 14.The H3Africa Consortium, et al. Enabling the genomic revolution in Africa. Science. 2014;344:1346–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Mulder NJ, et al. H3ABioNet, a sustainable pan-African bioinformatics network for human heredity and health in Africa. Genome Res. 2016;26:271–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Infinium H3Africa Consortium Array v2.
  • 17.Tluway F, et al. Cohort profile: Africa Wits-INDEPTH partnership for Genomic studies (AWI-Gen) in four sub-Saharan African countries. Int J Epidemiol. 2024;54:dyae173. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Ilboudo H, et al. Introducing the TrypanoGEN biobank: A valuable resource for the elimination of human African trypanosomiasis. PLoS Negl Trop Dis. 2017;11:e0005438. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Ramsay M, et al. Enriching African genome representation through the AGenDA project. Nature. 2026;649:565–73. [DOI] [PubMed] [Google Scholar]
  • 20.Gurdasani D, et al. The African Genome Variation Project shapes medical genetics in Africa. Nature. 2015;517:327–32. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Bekada A, et al. Genetic heterogeneity in Algerian human populations. PLoS ONE. 2015;10:e0138453. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Wohlers I, et al. An integrated personal and population-based Egyptian genome reference. Nat Commun. 2020;11:4719. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Fahime EE et al. Moroccan genome project: Genomic insight into a North African population. 2024. 10.21203/rs.3.rs-4904843/v1. [DOI] [PMC free article] [PubMed]
  • 24.Hamdi Y, et al. Genome Tunisia Project: paving the way for precision medicine in North Africa. Genome Med. 2024;16:104. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Fatumo S, et al. Uganda Genome Resource: A rich research database for genomic studies of communicable and non-communicable diseases in Africa. Cell Genomics. 2022;2:100209. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Gurdasani D, et al. Uganda genome resource enables insights into population history and genomic discovery in Africa. Cell. 2019;179:984-1002.e36. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Fatumo S, et al. Promoting the genomic revolution in Africa through the Nigerian 100K Genome Project. Nat Genet. 2022;54:531–6. [DOI] [PubMed] [Google Scholar]
  • 28.GenomeAsia 100K Consortium, et al. The GenomeAsia 100K Project enables genetic discoveries across Asia. Nature. 2019;576:106–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Indonesian Genome Diversity Project (EGAD00001004156). European Genome-phenome Archive (EGA).
  • 30.Teo Y-Y, et al. Singapore Genome Variation Project: A haplotype map of three Southeast Asian populations. Genome Res. 2009;19:2154–62. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Wu D, et al. Large-scale whole-genome sequencing of three diverse asian populations in Singapore. Cell. 2019;179:736-749.e15. [DOI] [PubMed] [Google Scholar]
  • 32.Shotelersuk V, et al. Advancing precision public health at Genomics Thailand. Nature Health. 2026. 10.1038/s44360-026-00059-4. [DOI] [Google Scholar]
  • 33.Shotelersuk V, et al. The Thai reference exome (T-REx) variant database. Clin Genet. 2021;100:703–12. [DOI] [PubMed] [Google Scholar]
  • 34.Boonin P, et al. Detection of genetic variants in Thai population by trio-based whole-genome sequencing study. Biology. 2025;14:301. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Thanh ND et al. Building population-specific reference genomes: a case study of Vietnamese Reference Genome. In 2015 Seventh International Conference on Knowledge and Systems Engineering (KSE) 97–102 (IEEE, Ho Chi Minh City, 2015). 10.1109/KSE.2015.49. [DOI]
  • 36.Hai DT, et al. Whole genome analysis of a Vietnamese trio. J Biosci. 2015;40:113–24. [DOI] [PubMed] [Google Scholar]
  • 37.Tran NH, et al. Genetic profiling of Vietnamese population from large-scale genomic analysis of non-invasive prenatal testing data. Sci Rep. 2020;10:19142. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Khan S, et al. Whole genome mapping and identification of single nucleotide polymorphisms of four Bangladeshi individuals and their functional significance. BMC Res Notes. 2021;14:105. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Jain A, et al. IndiGenomes: a comprehensive resource of genetic variants from over 1000 Indian genomes. Nucleic Acids Res. 2020. 10.1093/nar/gkaa923. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Fattahi Z, et al. Iranome: A catalog of genomic variations in the Iranian population. Hum Mutat. 2019;40:1968–84. [DOI] [PubMed] [Google Scholar]
  • 41.Samarakoon PS, Jayasekara RW, Dissanayake VHW. The Sri Lankan Genome Variation Database. Sri Lanka J Bio-Med Inform. 2011;2:9. [Google Scholar]
  • 42.Dissanayake VHW, et al. The Sri Lankan Personal Genome Project: an overview. Sri Lanka J Bio-Med Inform. 2011;2:4. [Google Scholar]
  • 43.Thareja G, et al. Sequence and analysis of a whole genome from Kuwaiti population subgroup of Persian ancestry. BMC Genomics. 2015;16:92. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Mbarek H, et al. Qatar genome: Insights on genomics from the Middle East. Hum Mutat. 2022;43:499–510. [DOI] [PubMed] [Google Scholar]
  • 45.Razali RM, et al. Thousands of Qatari genomes inform human migration history and improve imputation of Arab haplotypes. Nat Commun. 2021;12:5929. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.AlSafar HS, et al. Introducing the first whole genomes of nationals from the United Arab Emirates. Sci Rep. 2019;9:14725. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Daw Elbait G, Henschel A, Tay GK, Al Safar HS. A population-specific major allele reference genome from the United Arab Emirates population. Front Genet. 2021;12:660428. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Al-Ali M, Osman W, Tay GK, AlSafar HS. A 1000 Arab genome project to study the Emirati population. J Hum Genet. 2018;63:533–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Kulmanov M, et al. A reference quality, fully annotated diploid genome from a Saudi individual. Sci Data. 2024;11:1278. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Cheng H, Concepcion GT, Feng X, Zhang H, Li H. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat Methods. 2021;18:170–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Kars ME, et al. The genetic structure of the Turkish population reveals high levels of variation and admixture. Proc Natl Acad Sci. 2021;118:e2026076118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Du Z, et al. Whole genome analyses of chinese population and de novo assembly of a northern han genome. Genomics Proteomics Bioinform. 2019;17:229–47. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Zhang P, et al. NyuWa Genome resource: A deep whole-genome sequencing-based variation profile and reference panel for the Chinese population. Cell Rep. 2021;37:110017. [DOI] [PubMed] [Google Scholar]
  • 54.Li Z, et al. CMDB: the comprehensive population genome variation database of China. Nucleic Acids Res. 2023;51:D890–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Lan T, et al. Deep whole-genome sequencing of 90 Han Chinese genomes. GigaScience. 2017;6:gix067. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Cong P-K, et al. Genomic analyses of 10,376 individuals in the Westlake BioBank for Chinese (WBBC) pilot project. Nat Commun. 2022;13:2939. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Ou M, et al. HKG: An open genetic variant database of 205 Hong Kong Cantonese exomes. Preprint at. 2022. 10.1101/2021.06.15.448515. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Kawai Y, et al. Japonica array: improved genotype imputation by designing a population-specific SNP array with 1070 Japanese individuals. J Hum Genet. 2015;60:581–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Higasa K, et al. Human genetic variation database, a reference database of genetic variations in the Japanese population. J Hum Genet. 2016;61:547–53. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Takayama J, et al. Construction and integration of three de novo Japanese human genome assemblies toward a population-specific reference. Nat Commun. 2021;12:226. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Genome Reference Consortium. https://www.ncbi.nlm.nih.gov/grc.
  • 62.Flanagan J, et al. Population-specific reference panel improves imputation quality for genome-wide association studies conducted on the Japanese population. Commun Biol. 2024;7:1665. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Ahn S-M, et al. The first Korean genome sequence and analysis: full genome sequencing for a socio-ethnic group. Genome Res. 2009;19:1622–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Cho YS, et al. An ethnically relevant consensus Korean reference genome is a step towards personal reference genomes. Nat Commun. 2016;7:13637. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Hwang MY, Choi N-H, Won HH, Kim B-J, Kim YJ. Analyzing the Korean reference genome with meta-imputation increased the imputation accuracy and spectrum of rare variants in the Korean population. Front Genet. 2022;13:1008646. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Kim YJ, Jin HJ. Dissecting the genetic structure of Korean population using genome-wide SNP arrays. Genes Genomics. 2013;35:355–63. [Google Scholar]
  • 67.Yoo S-K, et al. NARD: whole-genome reference panel of 1779 Northeast Asians improves imputation accuracy of rare and low-frequency variants. Genome Med. 2019;11:64. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Seo J-S, et al. De novo assembly and phasing of a Korean human genome. Nature. 2016;538:243–7. [DOI] [PubMed] [Google Scholar]
  • 69.Jeon S, et al. Korean Genome Project: 1094 Korean personal genomes with clinical information. Sci Adv. 2020;6:eaaz7835. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Jung KS, et al. KRGDB: the large-scale variant database of 1722 Koreans based on whole genome sequencing. Database. 2020;2020:baz146. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Kim J, et al. KoVariome: Korean National Standard Reference Variome database of whole genomes with comprehensive SNV, indel, CNV, and SV analyses. Sci Rep. 2018;8:5677. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Feng Y-CA, et al. Taiwan Biobank: A rich biomedical research database of the Taiwanese population. Cell Genom. 2022;2:100197. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Zhernakova DV, et al. Genome-wide sequence analyses of ethnic populations across Russia. Genomics. 2020;112:442–58. [DOI] [PubMed] [Google Scholar]
  • 74.Mallick S, et al. The Simons Genome Diversity Project: 300 genomes from 142 diverse populations. Nature. 2016;538:201–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Pagani L, et al. Genomic analyses inform on migration events during the peopling of Eurasia. Nature. 2016;538:238–42. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Yuan L, et al. Genotypic and allelic frequencies of GJB2 variants and features of hearing phenotypes in the Chinese Population of the Dongfeng-Tongji Cohort. Genes (Basel). 2023;14:2007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Kiseleva AV, et al. A data-driven approach to carrier screening for common recessive diseases. J Pers Med. 2020;10:140. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Degechisa ST, Mersha TB. Unlocking Ethiopia’s genomic landscape and its global significance: a call for inclusive genomics research. Hum Genom. 2026;20:24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Aka I, et al. Clinical pharmacogenetics of cytochrome P450-associated drugs in children. J Pers Med. 2017;7:14. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Mills MC, Rahal C. The GWAS Diversity Monitor tracks diversity by disease in real time. Nat Genet. 2020;52:242–3. [DOI] [PubMed] [Google Scholar]
  • 81.Wang T, et al. The Human Pangenome Project: a global resource to map genomic diversity. Nature. 2022;604:437–46. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Ali H, et al. Integrating the genome and exposome for precision health in Kuwait. Nat Rev Genet. 2025. 10.1038/s41576-025-00883-6. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1. (43.1KB, xlsx)
Supplementary Material 2. (23.2KB, docx)

Data Availability Statement

No datasets were generated or analysed during the current study.


Articles from Human Genomics are provided here courtesy of BMC

RESOURCES