Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Jul 11.
Published in final edited form as: Cell. 2026 Mar 27;189(9):2533–2555.e31. doi: 10.1016/j.cell.2026.03.007

Advancing Precision Health Discovery in a Genetically Diverse Health System

Roni Haas 1,2,3,4,19,*, Michael P Margolis 1,5,6,19, Angela Wei 1,7,8,9,19, Takafumi N Yamaguchi 1,2,3,4,10, Jeffrey Feng 11, Thai Tran 6, Veronica Tozzo 9, Katelyn J Queen 12, Mohammed Faizal Eeman Mootor 1,2,3,4, Vishakha Patil 6, Michael E Broudy 13, Paul Tung 13, Shafiul Alam 13, Danielle B Martinez 13, Yash Patel 1,2,3,4,10, Christa Caggiano 14, Nicole Zeltser 1,2,3,4, Rupert Hugh-White 1,2,3,4,10, Jaron Arbet 1,2,3,4, Ruhollah Shemirani 14, Mao Tian 1,2,3,4,10, Prapti Thapaliya 13, Lora Eloyan 13, Lawrence O Chen 1,5,6, Sandra Lapinska 15, Maryam Ariannejad 4, Clara Lajonchere 4,6; UCLA Precision Health Data Discovery Repository Working Group, UCLA Precision Health ATLAS Working Group, UCLA Health IT HPC Team, Regeneron Genetics Center, Eimear E Kenny 14,16,17,18, Bogdan Pasaniuc 15, Alex A T Bui 3,4,11, Valerie A Arboleda 1,3,8,9, Timothy S Chang 6, Noah Zaitlen 1,6,9, Paul T Spellman 1,3,4,12,20,*, Paul C Boutros 1,2,3,4,10,20,*, Daniel H Geschwind 1,4,5,6,20,21,*
PMCID: PMC13353064  NIHMSID: NIHMS2161011  PMID: 41903539

Summary

Linking genetic data with electronic health records in hospital biobanks promises to advance precision medicine, but limited ancestral diversity constrains discovery and generalizability. We analyzed 93,936 participants from the UCLA ATLAS Community Health Initiative to inform disease prevalence and genetic risk across five continental and 36 fine-scale ancestry groups. We discovered numerous unreported gene-phenotype associations, including FN3K with intestinal disaccharidase deficiency in Europeans and admixed Americans. Polygenic scores (PGS) robustly predicted common diseases, with effects markedly diminished in non-Europeans. Furthermore, we reduced the pronounced European bias in curated clinical variants using computational predictors, uncovering unreported disease-gene associations, including ANKZF1 and peripheral vascular disease in African Americans. Longitudinal data revealed that semaglutide efficacy varies across ancestries, is associated with PGS for type 2 diabetes, and is modulated by genetic variation in PTPRU. These findings illustrate how ancestrally diverse biobanks from a single health system yield robust disease associations and pharmacogenomic insights.

In brief:

The UCLA ATLAS biobank was utilized to integrate genetic data with electronic health records from individuals representing five continental and 36 fine-scale ancestries, identifying genotype-phenotype associations and highlighting populations at risk for disease. Computational predictors mitigated European bias in clinical variant curation predictors, uncovering ancestry-specific disease-gene links and identify genetic factors influencing semaglutide-induced weight loss. These findings demonstrate the value of well-curated, ancestrally diverse biobanks in advancing precision medicine.

Graphical Abstract

graphic file with name nihms-2161011-f0007.jpg

Introduction

The integration of electronic health records (EHRs) with genetic data is transforming biomedical research, offering unprecedented opportunities for preventing and managing common medical conditions1. Longitudinal sampling, linked with genetic and environmental data, provides advantages over standard cohort-driven research2. This has fueled the creation of nationwide biobanks, including the UK Biobank (UKBB) 3,4, All of Us (AoU)5, FinnGen6 and Taiwan Biobank7, as well as several large-scale academic biobanks, such as Mt. Sinai’s BioMe8, Vanderbilt’s BioVU9, Geissinger’s MyCode10, and the Michigan Genomics Initiative11. Integration of these biobanks has permitted innovative collaborative efforts, such as the eMERGE consortium12, COVID-19 host genomics initiative13, and the Global Biobank Initiative1.

Although these efforts have substantially advanced genetic and biomedical discovery, their concentration on participants of European (EUR) ancestry limits generalizability14–18. As PGS are validated for clinical use, the importance of measuring their accuracy in diverse populations grows; recent analyses show a continuous relationship between ancestral distance from the reference population and the utility of PGS17. Similarly, rare genetic variation has substantial ancestry-specific effects and distribution; for example, in African Americans, APOE4 alleles have reduced impact on Alzheimer’s disease risk19 and rare protective PCSK9 variants are more prevalent20. The interpretation of clinically-relevant rare variation is further hampered by a bias towards EUR variants in genetic databases21,22. Including non-EUR populations reveals substantial disparities in clinically-relevant rare variant frequencies23 and increases statistical power for discovery24,25. Thus, greater ancestral diversity in biobanks with detailed medical records strengthens efforts to advance precision health14,26,27.

Here, we analyzed data from the UCLA ATLAS Community Health Initiative, linking EHR with genomic information for 92,164 participants with array genotyping and 61,797 with whole-exome sequencing (WES). ATLAS reflects Los Angeles’ ancestral diversity within a single health system28–30, which reduces confounding from differing clinical practices and enables robust cross-population comparisons. We performed phenome-wide associations using common and rare variants within continental (broad-scale) and sub-continental (fine-scale) ancestral groups, and quantified disease diagnoses and genetic risk across these groups. Using longitudinal EHR data, we employed semaglutide as a case study to identify genetic factors that alter drug efficacy. We identified multiple unreported ancestry-specific risk associations, further demonstrating the utility of ancestral diversity for more equitable, personalized medicine research (Figure 1a, schematic overview).

Figure 1. Overview of the UCLA ATLAS biobank.

Figure 1.

a. Schematic of workflow and analysis. b. Choropleth map of ATLAS participants within Los Angeles County, with major highways and landmarks labelled. c. Distribution of ATLAS participants by age and sex. d. Prevalence of phecode groups in the UCLA ATLAS population at the time of collection and within one year of the ATLAS launch date. e. Genetic ancestry sample sizes. f. Genetic ancestry fractions of non-European populations. Percentages were calculated using all ATLAS populations, including EUR. g. Genetic principal components of ATLAS participants. h. Mean yearly encounters vary across genetic ancestries. i. Comorbidity index varies across genetic ancestries. In h-I, ANCOVA was used to obtain P-values; the adjusted means and 95% confidence intervals are presented. j. Associations between clinical phenotypes and broad-scale ancestries using logistic regression.

Results

The demographic and clinical landscape in ATLAS

UCLA Health serves Los Angeles County, one of the world’s most ancestrally diverse metropolitan areas, with a population of 9.6 million. To date, the UCLA ATLAS initiative has consented ~250,000 UCLA Health patients and has collected biomaterials from ~130,000. ATLAS enrollment largely reflects the composition of UCLA Health patients, who are concentrated across Los Angeles (Figure 1b). The EHR, initiated in 2013, enables continuous longitudinal stratification of participants by disease states, with a mean and median of 8.6 and 7.9 years of participation per individual. As of November 2024, ATLAS genomic data included WES data from 61,797 participants and custom array genotyping using the Illumina Global Screening Array from 92,164 participants. Extensive quality control (QC) indicated high data quality (Figure S1). We leveraged these data to interrogate social and genetic factors that affect disease risk and health outcomes (Figure 1a).

Genotyped cohort demographics are summarized in Table 1 and the Methods. Biobank participants were older in age and had a higher comorbidity index than non-biobank patients (2.9 versus 1.7 mean Elixhauser index; 1-year post-collection), consistent with a higher number of clinical visits, which favors enrollment (controlled for data-completeness, Methods; Figure 1c). EHR-based phenotypes, including vital signs, disease diagnoses, and lab tests, were defined through harmonization procedures and subsequently validated (Methods; Figure S2a–l). Common disease domains were led by endocrine/metabolic, cardiovascular, gastrointestinal disorders and neoplasms, reflecting known global health challenges (Figure 1d; Figure S2m–p; Methods)31–33. The ATLAS EHR data includes a total of 71,739,582 lab tests (considering complete blood count [CBC], lipid, metabolic, hemoglobin A1c [HbA1c], and 25-hydroxyvitamin D panels), with a mean of 540 lab test results per participant (Table S1). Detailed prescription data show a total of 5,952,958 prescriptions; the 50 most frequently prescribed medications are listed in Table S1.

Table 1.

Baseline demographic information on the UCLA ATLAS population

Variable Category Overall Alive Deceased
n 92164 88239 3925
Age at collection date, mean (SD) 54.0 (17.3) 53.4 (17.2) 67.4 (15.2)
Self-reported sex, n (%) Female 51362 (55.7) 49666 (56.3) 1696 (43.2)
Male 40765 (44.2) 38536 (43.7) 2229 (56.8)
Unspecified, X 37 (0.0) 37 (0.0)
Self-reported race, n (%) American Indian, Alaska Native 820 (0.9) 798 (0.9) 22 (0.6)
Asian 11070 (12.0) 10700 (12.1) 370 (9.4)
Black, African American 4230 (4.6) 4004 (4.5) 226 (5.8)
Caribbean/West Indian 165 (0.2) 165 (0.2)
Middle Eastern or North African 2358 (2.6) 2309 (2.6) 49 (1.2)
Native Hawaiian, Guamian or Chamorro, Samoan, Other Pacific Islander 267 (0.3) 260 (0.3) 7 (0.2)
Other Race 3414 (3.7) 2967 (3.4) 447 (11.4)
Unknown, declined to specify 12059 (13.1) 11772 (13.3) 287 (7.3)
White 57781 (62.7) 55264 (62.6) 2517 (64.1)
Self-reported ethnicity, n (%) Hispanic/Latino, Cuban, Hispanic/Spanish origin, Mexican, Mexican American, Chicano/a, Puerto Rican 13645 (14.8) 12907 (14.6) 738 (18.8)
Non-Hispanic/Latino 71288 (77.3) 68262 (77.4) 3026 (77.1)
Unknown, declined to specify 7231 (7.8) 7070 (8.0) 161 (4.1)
Self-reported smoking, n (%) Former 24522 (26.6) 22937 (26.0) 1585 (40.4)
Never 63281 (68.7) 61124 (69.3) 2157 (55.0)
Passive Smoke Exposure - Never Smoker 118 (0.1) 112 (0.1) 6 (0.2)
Smoker 3576 (3.9) 3448 (3.9) 128 (3.3)
Unknown, declined to specify 667 (0.7) 618 (0.7) 49 (1.2)
Types of Encounters, n (%) Inpatient and Outpatient Encounters 29491 (32.0) 26275 (29.8) 3216 (82.0)
Only Outpatient Encounters 62626 (68.0) 61918 (70.2) 708 (18.0)
Total Encounters, mean (SD) 115.7 (145.4) 108.9 (136.5) 266.6 (231.2)
Total Inpatient Encounters, mean (SD) 0.8 (2.2) 0.7 (1.9) 3.8 (5.1)
Total Inpatient Days, mean (SD) 16.3 (33.6) 13.6 (28.9) 38.6 (54.5)

Although we primarily consider genetically determined ancestry in our analyses, rather than the social constructs of race and ethnicity34,35, we note substantial diversity based on participant self-reports. For instance, 62.7% of participants self-identify as White, 12% as Asian, 4.6% as Black or African American, 2.6% as Middle Eastern or North African, 0.9% as American Indian or Alaska Native and 0.3% as Pacific Islander. A substantial proportion (14.8%) report Hispanic or Latino ethnicity (Table 1).

Disease burden varies by broad-scale genetic ancestry

Self-identified race/ancestry, a cultural-societal construct, and genetic ancestry are conceptually distinct29,34,36 and genetic ancestry must be considered to prevent confounding in genetic association studies35,37. We classified biobank participants into six broad-scale ancestry populations: European (EUR), African (AFR), South Asian (SAS), East Asian (EAS), Admixed American (AMR) and an Unclassifiable (UNC) population, aligning with the 1000 Genomes Project super-populations38 (Figure 1e–g; Methods). ATLAS is diverse relative to most other large biobanks (Methods). Consistent with self-reported race/ethnicity, almost a third of the biobank participants (32%) were assigned to non-EUR genetic ancestries. There was significant agreement between self-reported race and broad-scale ancestry – 99% of self-identified White participants were assigned to EUR or AMR populations (Figure S3a). Of those with “unknown” self-reported race, 46% were assigned to AMR ancestry and 45% to EUR.

We next asked how health system usage or disease burden varied by broad-scale ancestry, observing that the mean number of encounters varied significantly across ancestries (P-value = 7.5×10−138, ANCOVA, adjusted for genetic sex, age and Barriers to Accessing Services [BAS]39), with the highest numbers of total encounters in participants from AMR ancestry (adjusted mean = 16.2), followed by AFR (15.8) and SAS (14.2), with the lowest values in EUR (12.2) and EAS (12.9) (Figure 1h). We also assessed health burden using the comorbidity index score40,41, which was higher in EAS (adjusted mean = 3.7) and AMR (3.6) compared to others (all other populations 2.6-2.9; Figure 1i; P-value = 2.6×10−5, ANCOVA, adjusted for genetic sex, age, BAS and Area Deprivation Index [ADI]42). Similar results for disease distribution across ancestries were obtained after applying inverse probability weighting (IPW) to adjust for participation bias relative to the broader UCLA Health population (Figure S3b; Methods). Increased encounter numbers were only partially explained by elevated comorbidity index scores, indicated by modest correlations between the two (mean total encounters vs. Elixhauser Comorbidity Index: r = 0.32, P-value <2.2×10−16; mean hospital encounters vs. Elixhauser Comorbidity Index: r = 0.26, P-value <2.2×10−16; Figure S3c–d).

The population diversity of ATLAS enables the interrogation of the combined genetic and social/environmental effects on the risk of disease diagnoses. To illustrate this, we assessed variation in medical conditions across populations, replicating multiple known associations (Methods; Table S2; Figure S3e). Further, we identified previously unreported associations (Figure 1j), including a lower risk of epilepsy in EAS compared to EUR participants (odds ratio [OR] = 0.46 with 95% confidence interval [0.32, 0.65], PBonferroni = 2.3×10−4), not previously detected43,44. Similarly, we found significantly reduced risk for bipolar disorder in those with AMR ancestry, clarifying conflicting findings in previous smaller studies45,46 (OR = 0.47 [0.40, 0.56], PBonferroni = 1.5×10−17). We also show a significantly lower risk of sleep apnea in EAS participants relative to EUR (OR = 0.67 [0.59, 0.75], PBonferroni = 4.1×10−10), which has been controversial47–50, with a weaker, but significant effect after correcting for body mass index (BMI) (OR = 0.84 [0.74, 0.95], P-value = 4.6×10−3; Table S2). These associations remained significant after adjusting for socioeconomic status (SES), using both the BAS and ADI measures, and with IPW (Table S2).

Disease prevalence varies across fine-scale clusters

Fine-scale ancestries37,51–53, many of which are understudied14–16, reflect recent geographic or demographic stratification and can reveal important contributions to health disparities and disease risk14. We quantified fine-scale ancestry using identity-by-descent (IBD)53,54 in a population nearly three times larger than any previously published study52,53. This approach identifies genomic regions shared between individuals due to a common ancestor and defines fine-scale ancestries, which we refer to as “clusters” (Methods)52. We identified 36 fine-scale ancestry clusters with at least 30 participants (Figure 2a–b; Figure S4a–c; Methods), labeled with both a numeric identifier (e.g., IBD-01) sorted based on the cluster sample size, and a corresponding cluster name to facilitate interpretation (Methods). Because PCA captures ancient population structure and IBD segments reflect recent shared ancestry, full agreement between these approaches is not expected, as is observed for some of our IBD clusters, such as IBD-02 (Figure 2a). We replicated previously identified clusters and added unreported ones, such as a Native Hawaiian cluster (nIBD-29 = 64) and a Bantu cluster (nIBD-35 = 32). The largest clusters consisted of Northern Europeans (nIBD-01 = 33,675) and Southern Europeans (nIBD-02 = 14,841). The remaining clusters represent the heterogeneity of ancestral origins in Los Angeles, including Ashkenazi Jewish (nIBD-03 = 14,262), Filipino (nIBD-09 = 1,438), Iranian Jewish (nIBD-11 = 707), and Armenian (nIBD-16,IBD-23,IBD-31 = 560) populations. The diversity of ATLAS is reflected not only by a relatively large portion of non-EUR individuals, but also by substantial diversity within the EUR broad-scale ancestry population. Several EUR clusters are large relative to published datasets, including Southern European, Ashkenazi Jewish, Armenian and Iranian Jewish3,53,55–57 (Table S3). Similarly, for EAS, we identified multiple clusters, including Filipino, representing the largest published Filipino genotyped cohort of which we are aware58. Despite this strength, for some populations, sample sizes are limited, which leads to relatively small clusters.

Figure 2. Identity-by-descent (IBD) mapping and fine-scale ancestry disease risk.

Figure 2.

a. Genetic ancestry assignments from broad- to fine-scale populations. b. Enrichment of reference populations, self-reported language, religion, and race across fine-scale ancestries. c. The risk of phenotype-defined phecodes across fine-scale genetic ancestries. Each fine-scale cluster was tested against all ATLAS participants outside of this cluster using Firth’s logistic regression across 1,253 phecodes.

For a nuanced understanding of diagnostic variation, we tested the prevalence of 1,253 phecodes59–61 across 24 fine-scale clusters with at least 100 participants. Table S3 shows all associations and Figure 2c highlights selected associations (also available at https://atlas-phewas.mednet.ucla.edu/ancestry). We replicated known findings (Methods) and uncovered numerous associations between fine-scale clusters and diseases. For instance, Filipinos, an understudied population in genetics research, exhibited the lowest risk for vitamin B-complex deficiency (ORIBD-09 = 0.50 [0.36, 0.67], FDR = 2.1×10−4), and the highest risk for cholesterolosis of the gallbladder (ORIBD-09 = 4.3 [2.9, 6.2], FDR = 5.6×10−12), both previously unreported. In Iranian Jewish subjects, another group with low research inclusion, we detected a previously unreported elevated risk for glaucoma (ORIBD-11 = 1.9 [1.5-2.4], FDR = 2.9×10−6), a condition known to be more common in individuals of African and East Asian ancestry. We also observed several shared health risk patterns among Jewish ancestry clusters. Both Iranian and Ashkenazi Jewish clusters showed the highest risk for bladder cancer (Iranian Jewish: CIIBD-11 = 1.4-5.6, FDR = 0.008; Ashkenazi Jewish: ORIBD-03 = 1.5 [1.1, 1.9], FDR = 0.02) and a high hyperplasia of prostate risk (Iranian Jewish: ORIBD-11 = 2.3 [1.8, 2.8], FDR = 1.1×10−13; Ashkenazi Jewish: OR IBD-03 = 1.7 [1.6, 1.8], FDR = 4.6×10−66), and the lowest risk for cirrhosis of liver (Iranian Jewish: CIIBD-11 = 0.03-0.4, FDR = 3.1×10−2; Ashkenazi Jewish: ORIBD-03 = 0.4 [0.3, 0.5] , FDR = 6.8×10−24). Mexicans and South Americans suffered consistently more from hormones’ adverse effects in therapeutic use, a previously unreported association (Mexican American clusters: ORIBD-04,-07,and-12 = 2.0-2.8 [1.3, 3.4], FDR = 1.6×10−2 - 2.1×10−24; combined South Americans: ORIBD-08 = 1.9 [1.4, 2.5], FDR = 1.3×10−4). These associations persisted after SES adjustment, using the BAS and ADI ranks. In most cases, associations held after applying IPW and also adjusting for SES, except for hormone adverse effects, where only the largest Mexican American group remained highly significant, likely due to smaller sample sizes resulting from the IPW analyses (Table S3). These findings demonstrate the utility of fine-scale ancestry in identifying disease risk profiles in underrepresented populations.

We next focused on cardio-metabolic diseases due to their global impact on public health. We selected a few targeted cardio-metabolic phenotypes and compared disease risk across fine-scale clusters within the same broad-scale ancestry (Methods; Figure S4d). Strikingly, among Asian clusters, the Filipino cluster had an elevated risk for all tested cardio-metabolic conditions, which remained after adjusting for BMI. Across fine-scale EUR ancestries, the larger Armenian cluster showed high cardiometabolic disease risk. Iranian clusters, both Jewish and non-Jewish, had a relatively higher risk of coronary atherosclerosis, hyperlipidemia and type 2 diabetes, but not hypertension. Adjusting for SES factors and adding IPW maintained these significant patterns (Table S3), except abdominal aortic aneurysm in Filipinos, which lost significance, and hyperlipidemia in Iranian Jewish participants, which did not remain significant only when both SES adjustment and IPW were applied- likely due to reduced sample size, although the OR remained elevated.

Ancestry-specific common variants associated with disease risk

To examine the robustness of ATLAS in predicting genetic risk and to assess the quality of our dataset, we first tested the utility of PGS in stratifying risk for common disorders. We observed high PGS performance in EUR individuals–on average, 18.6% of patients diagnosed with major disorders were in the top PGS decile, rising to 41% for type 1 diabetes (Figure 3; STAR Methods). Consistent with prior reports62–65, this predictive power was diminished in non-European ancestries (Figure 3; Methods).

Figure 3. Polygenic risk and disease diagnoses.

Figure 3.

The top and bottom deciles of PGS were compared to the 5th decile to calculate the odds ratio (OR) using logistic regression (shown as dotmaps) in Europeans. The number of diagnosed patients assigned to the top or bottom PGS bins is shown in bar plots. Red represents the top PGS decile, and blue the bottom decile. Case counts are shown on the right. Diseases are divided into categories (top to bottom): cancer, cardiovascular, metabolic, neuropsychiatric, and immune/autoimmune. FDR was used for multiple testing correction on p values derived from logistic regression for tested bins.

Next, we performed phenome-wide association studies (PheWAS) using Regenie66 to map the common variant architecture of clinical diagnoses and laboratory measurements, conducting separate analyses within each of our broad- and fine-scale ancestral cohorts (Methods; n = 84,110 total participants). After linkage disequilibrium (LD) pruning, we identified 19,431 unique variant-phenotype associations passing a genome-wide significance threshold of 5×10−8 (Table S4), with 5,772 passing the most conservative Bonferroni threshold of 6.4×10−11 (5×10−8 conditioned on 776 tested phenotypes). The majority (83.8%) of associations were detected in only one ancestry at genome-wide significance, though utilizing a more permissive threshold of 1×10−4 showed cross-ancestry support for a substantial portion of associations (54.7%; Figure S6a–b). Supporting the validity of the data used for PheWAS, we confirmed ancestry-specific differences in APOE allele frequencies and attenuated Alzheimer’s risk associated with the ε4 allele in African Americans and replicated known associations such as PNPLA3 rs738409-G with non-alcoholic fatty liver in AMR groups67 (Methods; Figure S6c–e). We identified numerous ‘previously unreported findings’, which we defined as associations that have not been reported at genome-wide significance in the existing literature, the European Molecular Biology Laboratory-European Bioinformatics Institute (EMBL-EBI) GWAS Catalog68, or summary statistics from UKBB69, Taiwan BioBank7, and AoU5. Notably, 39.5% of associations were unreported in the European Molecular Biology Laboratory-European Bioinformatics Institute EMBL-EBI GWAS Catalog68 (Methods). We performed replication analyses for the top previously unreported findings in independent biobanks, including AoU5, UKBB69, Taiwan BioBank7, and BioMe8 (Table S4; Methods). Summary statistics are available at atlas-phewas.mednet.ucla.edu.

We identified a previously unreported association between FN3K rs7208565-T and increased risk for intestinal disaccharidase deficiency across EUR populations (Figure 4A; PEUR = 3.8 × 10−41; OREUR = 1.3; MAFEUR = 33.3%; MAFIBD-01 = 32.3%; MAFIBD-02 = 34.3%; MAFIBD-03 = 35.6%) and the AMR population (PAMR = 4.9 × 10−11; ORAMR = 1.3; MAFAMR = 42.3%; Table S4). This variant was additionally associated with increased risk of the “other abnormal glucose” phecode (PEUR = 1.8 × 10−42; OREUR = 1.2), whereas a variant in close LD (rs113373052-T) was associated with increased HbA1c (PEUR = 5.4 × 10−97; βEUR = 0.13), reflecting additional consequences on glucose homeostasis. Intestinal disaccharidase deficiency is the inability to completely digest sugars such as lactose and sucrose, leading to irritable bowel syndrome (IBS)-like symptoms70. FN3K phosphorylates glycated proteins, preventing the formation of advanced glycation end-products71. Supportive of these findings, rs7208565 has been previously associated with type 2 diabetes72.

Figure 4. Selected PheWAS results and pharmacogenetic variant enrichment.

Figure 4.

a-d. Locus plots for unreported PheWAS associations between a. rs7208565-T and intestinal disaccharidase deficiency b. rs17419569-C and asthma. c. rs74744741-C and gastrointestinal reflux disease (GERD) and d. rs112680741-C/rs2744548-C and chronic renal failure. Point shading (blue to orange to red) indicates the level of ancestry-specific linkage disequilibrium (LD) with the lead single-nucleotide polymorphism (SNP), while gene shading (red) indicates prioritized risk genes. In a-d, Regenie’s Firth logistic regression was used for variant-trait associations. e. Pharmacogenetic variant enrichment across fine-scale ancestries. A bold border indicates FDR ≤ 0.05. Pharmacogenetic variants were aggregated for genes with multiple variants. Previously unreported enrichments highlighted in the text were for: SLCO1B1 in Ashkenazi and Iranian Jewish (ORIBD-03 = 1.4, [1.3, 1.5], FDR = 1.2 × 10−35; ORIBD-11 = 1.5, [1.2, 1.9], FDR = 5.0 × 10−4), CYP4F2 in Ashkenazi Jewish (ORIBD-03 = 1.4, [1.3, 1.5], FDR = 1.4 × 10−45), Lebanese (ORIBD-21 = 2.0, [1.3, 3.1], FDR = 2.7 × 10−3), Armenian 1 (ORIBD-16 = 2.1, [1.6, 2.8], FDR = 9.1 × 10−8), Egyptian Christian (ORIBD-25 = 2.6, [1.4, 4.8], FDR = 3.0 × 10−3), and South Asian 2 (ORIBD-27 = 2.9, [1.5, 5.8], FDR = 1.6 × 10−3), and DPYD in Mexican American, African American, and South American participants (ORIBD-04 = 1.4, [1.3, 1.5], FDR = 3.7 × 10−16; ORIBD-06 = 1.3, [1.1, 1.4], FDR = 2.5 × 10−5; ORIBD-07 = 1.2, [1.1, 1.4], FDR = 5.3 × 10−4; ORIBD-08 = 1.6, [1.4, 1.9], FDR = 3.4 × 10−9).

Furthermore, we discovered several previously unreported associations driven by low-frequency variants (MAF ~1-2%) in non-EUR cohorts. For instance, we linked rs115750084-G in DPP6, a gene implicated in synaptic signaling73 and neurological disorders74,75, to an increased risk of major depressive disorder in the AFR cohort (PAFR = 4.6 × 10−8; ORAFR = 4.2; MAFAFR = 1.1%; Figure S6F). In the EAS cohort, we associated rs77742325-G in MAS1, a key component of the bone-protective ACE2/Angiotensin-(1-7)/Mas axis,76 with osteoporosis (PEAS = 1.5 × 10−8; OREAS = 3.5; MAFEAS = 1.8%; Figure S6G). We also linked an intronic deletion (rs202215133-ATATCATAG>A) in UBR5, an E3 ubiquitin ligase implicated in neuroinflammatory and neurodevelopmental syndromes,77,78 with migraine headache in the Mexican American cluster (PIBD-04 = 1.7 × 10−8; ORIBD-04 = 5.6; MAFIBD-04 = 1.1%; Figure S6H). These associations were independently replicated, albeit with more modest effect sizes in the replication cohort (Table S4). Given that the variants are low MAF, further validation through direct sequencing in large cohorts will be essential to confirm their generalizability.

Similarly, our large AMR discovery cohort revealed several unreported associations. For instance, we linked STARD7 rs17419569-C with increased risk for asthma in the Mexican American cluster (Figure 4B; PIBD-04 = 1.9 × 10−9; ORIBD-04 = 2.6; MAFIBD-04 = 2.5%). Prior work suggests that decreased STARD7 expression is associated with enhanced allergic responses in the human lung and significant increases in airway hyperresponsiveness in haploinsufficient Stard7 mice,79 whereas rare variants within STARD7 were linked with asthma (PUKBB = 5.1 × 10−3; Table S4). In the same cluster, GPX7/SHISAL2A rs74744741-C was associated with gastrointestinal reflux disease (GERD) (Figure 4C; PIBD-04 = 5.8 × 10−10; ORIBD-04 = 1.6; MAFIBD-04 = 8.9%). GPX7 has been associated with carcinogenesis in the context of GERD-associated Barrett’s esophagus,80 whereas SHISAL2A has an unknown function, but is highly expressed in the small intestine and in lymphoid tissues.81 Additionally, we identified two low-frequency AMR variants associated with chronic renal failure (CRF; Figure 4D), rs112680741-C (PAMR = 2.9 × 10−11; ORAMR = 3.8; MAFAMR = 1.3%) and rs2744548-C (PAMR = 9.6 × 10−11; ORAMR = 3.2; MAFAMR = 1.7%), nominating GPLD1, ALDH5A1, and KIAA0319 as potential CRF risk genes. These associations did not replicate in available biobank populations (Table S4), possibly because the ATLAS AMR populations from the Los Angeles area represent distinct Mexican and South American ancestries82,83, whereas those from other biobanks (BioMe, for example) are comprised of individuals with distinct ancestral cluster assignments to Puerto Rico and the Dominican Republic. Other cohort-specific characteristics may also contribute to these differences, so independent replication in additional cohorts is needed.

Frequencies of known clinically relevant rare variants vary across ancestries

This first wave of the ATLAS WES catalog comprises more than 11.5 million autosomal variants, with a median of 9,821 missense and 159 loss-of-function (LOF) variants per individual, closely matching expectations from other studies84. Most (98%) WES variants were rare (MAF < 1% in any broad-scale ancestry), including 2,873,731 rare missense and 172,462 rare LOF variants, with a median of 15 rare LOF and 403 rare missense variants per participant (Table S5). AFR participants showed the greatest number of rare LOF and missense variants, matching prior results85. To evaluate our WES, we analyzed a predefined set of clinically relevant rare variants known to have elevated frequencies in specific populations and demonstrated that ATLAS captures these established patterns (Methods; Figure S7a–e). Then, we examined differences in 69 ClinPGx pharmacogenomic variants86,87 across fine-scale ancestries (Figure 4e; Table S4). We identified several unreported enrichments in specific clusters, including a SLCO1B1 variant, rs4149056, which affects the metabolism and toxicity of statins88 in Ashkenazi and Iranian Jews; a CYP4F2 variant, rs2108622, that can increase the required dosage of warfarin89, enriched in Ashkenazi Jewish, Lebanese, Armenian 1, Egyptian Christian, and South Asian 2 individuals; and PYD variants, which affect the toxicity of chemotherapy drugs90, abundant in Mexican American, African American and South American clusters. This suggests the importance of considering fine-scale ancestry in precision drug prescribing.

Next, we used ATLAS’ ancestral heterogeneity to identify enrichments of known rare clinically relevant variants, aggregated by gene across fine-scale clusters to illustrate the utility of fine-scale ancestry in identifying populations at higher risk of rare, monogenic disease. We queried the full set of the curated rare pathogenic/likely-pathogenic (P/LP) ClinGen91 variants that have strong clinical and genetic evidence (n = 643), identifying 5,223 unrelated carriers (Figure 5a). Alongside known associations (STAR Methods), we detected a previously unreported elevated carrier frequency for variants in the Centers for Disease Control and Prevention (CDC) Tier 1 gene, LDLR, which causes familial hypercholesterolemia (FH) in the Filipino cluster (ORIBD-09 = 3.9 [1.4, 8.8], FDR = 0.03). This finding is notable given the high prevalence of dyslipidemia in Filipinos92 (Figure S4D). Indeed, LDLR FH variants accounted for 3.1% (population attributable risk [PAR]) of high LDL cases (LDL > 190 mg/dl) in the Filipino cluster. However, due to the low number of carriers in this cluster (<10), the PAR CI was wide [0.3, 26.1], and independent replication is required to refine this finding. Additionally, we identified previously unreported elevated carrier frequency of variants that cause non-syndromic genetic deafness within two different genes in two separate clusters: (1) MYO15A in the Mexican American 1 cluster (ORIBD-04 = 6.9 [2.6, 17.4], FDR = 8.44×10−4) and (2) CDH23 in the Chinese + Korean (ORIBD-05 = 6.0 [2.0, 15.6], FDR = 6.04×10−3) and Japanese (ORIBD-10 = 28.1 [8.9, 77.1], FDR = 7.4 × 10−6) clusters. We also identified previously unreported elevated carrier frequencies in the Ashkenazi Jewish cluster for Pendred syndrome variants in SLC264A (ORIBD-03 = 3.8 [2.9, 5.0], FDR = 8.6 × 10−19) and recombinase activating gene 2 deficiency in RAG2 (ORIBD-03 = 2.9 [1.4, 5.8], FDR = 0.02). Finally, we focused on the clinically actionable American College of Medical Genetics and Genomics (ACMG)93 genes, where ClinGen-curated variants within these genes manifested a strong EUR bias (Figure 5b; Figure S7f; Methods). This bias was recently shown in AoU23 and we provide additional supporting evidence by demonstrating this bias in a single health system.

Figure 5. Rare variant findings.

Figure 5.

a. ClinGen pathogenic/likely-pathogenic (P/LP) variant enrichment across fine-scale ancestries. b. Differences across populations in the total numbers of rare ClinGen P/LP variants in American College of Medical Genetics (ACMG) genes. c. Differences across populations in the total numbers of rare predicted damaging missense and loss of function (LOF) variants in ACMG genes. In b-c, ‘Ref’ is the total number of reference alleles, and ‘Alt’ of alternative alleles, which are the rare P/LP ClinGen variants in b, and the rare, predicted LOF/damaging missense variants in c. In a-c, Fisher’s exact tests were used. In b-c, the top panels show broad-scale ancestries, and the bottom panels show fine-scale ancestries. Filled points represent significant results (PBonferroni ≤ 0.05). d. Exome-wide association study significant results for selected traits/ancestries. Regenie’s omnibus “GENE_P” test was used for gene-trait associations.

Predicted rare deleterious variants identify population-specific genetic risk

As clinical variant databases are primarily ascertained from EUR populations, we took a conservative approach to call rare predicted deleterious LOF (dLOF) and missense coding variants (dMIS). High-confidence dLOF were defined using the loss-of-function transcript effect estimator (LOFTEE)94, which accounts for biological context beyond protein truncation. dMIS variants were defined using a consensus of a majority (5 of 9) of state-of-the-art computational methods based on the Critical Assessment of Genome Interpretation (CAGI) project95 for pathogenicity prediction (Methods; Figure S7g–h). This approach is distinct from the one taken with ClinGen variants – variants identified in this manner may not have literature evidence for their impact on gene function but are potentially less biased by EUR oversampling.

First, we kept our focus on ACMG genes93, repeating the ClinGen-based analysis described above using predicted rare dLOF and dMIS variants. We examined the frequency of 1,602 rare dLOF and 6,487 rare dMIS variants within ACMG variants across populations (Figure 5c; Figure S7i). The results were consistent with previous work showing that AFR participants harbor substantially more dLOF compared to other ancestries85, and highlighted other groups with previously unreported low or high ACMG putative damaging variants for further investigation. Overall, these results contrasted in comparison with rare variants identified in ClinGen (Figure 5b), consistent with the interpretation that ClinGen variant frequencies are biased toward EUR participants, as is the case for most curated clinical datasets.

Second, to examine the influence of dLOF and dMIS on ATLAS phenotypes, we performed exome-wide association studies (ExWAS) across 17,537 protein-coding genes (Methods). Using Regenie’s unified gene burden association strategy66, we identified 1,099 unique Bonferroni-significant gene-trait associations. Within the EUR population, we replicated numerous known associations; 45 of the top 50 associations had been previously reported84 (Table S6). These include PKD1 with cystic kidney disease and TTN with primary intrinsic cardiomyopathy (Figure 5e)84. We also identified numerous previously unreported associations (Table S6). In the Northern European cluster, we detected gene-level associations between HNRNPA1L2 and acquired absence of the breast (GENE_PIBD-01 = 1.0×10−25), breast cancer (GENE_PIBD-01 = 2.2×10-8), and malignant neoplasm of the female breast (GENE_PIBD-01 = 1.5 × 10−7). Notably, HNRNPA1L2 is on chromosome 13, ~20 megabases from BRCA2, representing a different locus. Though HNRNPA1L2 has not been previously associated with breast cancer, it is highly expressed in breast invasive carcinoma81 and a paralog HNRNPA1 was previously implicated in breast cancer progression96. In the Ashkenazi Jewish (IBD-03) cluster, we observed associations between EPG5 and HDL cholesterol level (PLOF_MIS = 3.6×10−10; βLOF_MIS = 1.8 [1.2, 2.3]) and triglyceride level (PLOF_MIS = 1.33×10−8; βLOF_MIS = −1.81 [−2.4, −1.2]). EPG5 is an autophagy tethering factor classically associated with Vici syndrome97, a severe developmental disorder, though our findings suggest an additional role in lipophagy and lipid metabolism98.

In the AMR population, we identified an unreported association between dLOF and dMIS variation in CLN3 and cystic kidney disease (GENE_PAMR = 8.7×10−9), supported by experimental evidence that CLN3 is highly expressed in medullary collecting duct principal cells and plays a role in osmoregulation99. We additionally identified unreported associations between PPARG and viral pneumonia (PAMR,MIS = 1.1×10−6, ORAMR,MIS = 53.0 [13.5, 208.5]), as well as NADSYN1 and abnormal lung examination findings (PAMR,MIS = 1.2×10−5, ORAMR,MIS = 3.7 [2.1, 6.5]). PPARG is known to regulate macrophage response to pulmonary inflammation100, while NADSYN1 knockout in mouse models was shown to cause abnormal lung development101, though neither has been previously associated with respiratory traits in humans.

Within the AFR population and African American cluster, we uncovered unreported associations between DDHD2 and dysphagia (PIBD-06,LOF = 1.2×10−6, ORIBD-06,LOF = 8.3 [3.7, 18.7]) and EFCAB13 and essential hypertension (PAFR,MIS = 9.3×10−7, ORAFR,MIS = 13.1 [4.5, 37.7]). These findings align with prior studies; variants in DDHD2 have been shown to cause hereditary spastic paraplegia and symptoms of dysphagia102, though not in African Americans, and methylation studies have prioritized EFCAB13 as a risk factor for heart failure103. We detected an additional association between ANKZF1 and peripheral vascular disease (PAFR,LOF_MIS = 1.6×10−6, ORAFR,LOF_MIS = 58.8 [12.6, 275.1]), which is supported by recent experimental findings104.

Next, in the EAS population, we observed a previously unreported association between AK7 and non-rheumatic mitral valve disorders (PEAS,LOF = 7.6×10−7, OREAS,LOF = 85.7 [39.2, 187.2]). While AK7 has never been associated with cardiovascular phenotypes, it has a well-established role in primary ciliary dyskinesia106 and mutations in other ciliary genes have been associated with mitral valve prolapse107. We further identified an unreported association between KIF2B and nephritis/nephropathy in the Chinese/Korean cluster (PIBD-05,MIS = 7.2×10−7, ORIBD-05,MIS = 26.2 [9.0, 76.4]). Though KIF2B has not been linked to renal traits, other genes in the kinesin family (KIF) have been implicated in the pathogenesis of renal cancer108,109. All previously unreported findings described above were independently replicated in AoU5, UKBB69 or BioMe8 at a nominal P-value < 0.05 (Table S6).

Finally, we examined known risk associations to identify ancestry-specific patterns of deleterious rare variation. We tested two known GBA1 rare missense variants, p.Glu365Lys and p.Thr408Met, in the AMR and EUR populations; both variants are reported to increase risk for Parkinson’s disease110. Despite similar allele frequencies within each population, p.Glu365Lys was only associated with increased risk in EUR (PEUR,E365K = 0.01; PAMR,E365K = 0.8), whereas p.Thr408Met was only associated with increased risk in AMR participants (PEUR,T408M = 0.2; PAMR,T408M = 0.003), displaying evidence of ancestry-specific risk stratification (Figure S7j). While we hypothesize that penetrance is modulated by population-specific genetic backgrounds, further studies are required to identify the causes of these differential associations. In total, We performed replication analyses for 21 of our top previously unreported PheWAS and ExWAS findings in independent biobanks and replicated 16 (76%) at a nominal p value of 0.05.

Longitudinal EHR data reveal genetic factors that shape GLP-1 receptor agonist efficacy

One of the advantages of EHR data is the ability to consider dynamic changes over time. To illustrate this, we integrated all main study components (genetic ancestry, common and rare genetic variants) to study the efficacy of glucagon-like peptide-1 (GLP-1) receptor agonists (GLP1-RAs) for weight loss. We focused on semaglutide, which had the greatest number of prescriptions in ATLAS (Figure S8a–b; 7,340 participants; Figure 6a). We observed a steady decrease in average weight up to ~60 weeks of semaglutide treatment, consistent with previous findings111–113 (Figure S8c). We tested if the medication dose, route, sex, age, and initial BMI affected semaglutide efficacy, corroborating that dose and subcutaneous delivery versus oral delivery were positively correlated with weight loss (Figure 6b; linear mixed-effects model [LMM]: dose: PBonferroni = 3.3×10−97, semipartial R² (explained variance) = 1.8%, effect size = −1.0; route oral vs. subcutaneous: PBonferroni = 7.2×10−37, semi-partial R2 = 1.1%, effect size = 2.2). We did not detect a significant effect of age and sex on efficacy.

Figure 6. Semaglutide efficacy.

Figure 6.

a. Distribution of age by sex in semaglutide users in ATLAS. b. The effect of baseline factors on weight loss in response to semaglutide. c. Weight loss patterns across genetic ancestry groups. d. Differences in weight loss patterns between AMR and EUR populations (left) and EAS and EUR (right). e. Relationships of body mass index (BMI) polygenic scores (PGS) (left) and type 2 diabetes PGS (right) with weight loss on semaglutide in EUR participants. Scaled PGS were divided into high, intermediate and low. In b,d and e, Bonferroni-adjusted P-values were obtained from a linear mixed-effects model that included covariates. In c-e, linear mixed model fitted values are plotted with 95% CI, based on longitudinal data with repeated weight measurements. f. Gene-level test results. Regenie (additive model) was applied to test the genetic association between semaglutide-affected proteins and weight loss on semaglutide. The horizontal line marks the Bonferroni cutoff. The top 5 hits are labeled by gene names. For plotting, gene positions were defined based on the first base pair.

Next, we asked if semaglutide’s effects varied by broad-scale ancestry, finding that ancestry significantly influenced weight loss over 60 weeks of treatment (Figure 6c; ANOVA on an LMM; ancestry: PBonferroni = 0.002; ancestry×time: PBonferroni = 0.0006). Comparisons between EUR and the other populations revealed less weight loss in AMR, and a slower rate of weight loss in AMR and EAS (Figure 6d; LMM; ancestryAMR: PBonferroni = 2.9×10−3, β = 0.8; ancestryAMR×time: PBonferroni = 5.0×10−3, β = 45.2; ancestryEAS ×time: PBonferroni = 3.2×10−3, β = 76.4), consistent with recent analysis based on a single time point, which showed reduced efficacy in AMR and AFR cohorts (here, in AFR, the result was not significant but the direction of effect was consistent).113

Given limited evidence on the role of inherited factors in GLP1-RAs effectiveness114, we used ATLAS to explore the contribution of common genetic variation to semaglutide-related weight loss. We found no correlation between BMI PGS and semaglutide efficacy. However, weight loss was negatively correlated with DM2 PGS (Figure 6e; LMM, PGSHigh vs. Low: PBonferroni = 2.8×10−4, β = 0.96; PGSMed vs. Low: PBonferroni = 8.8×10−3, β = 0.73). The same relationship was observed in a simplified model using only maximum weight loss recorded114 (linear regression, PBonferroni = 1.3×10−2, β = 0.31; Figure S8d). We could not find an equally sized cohort for replication, but we used EUR AoU, which had a sample size 50% smaller than the ATLAS cohort used in this analysis with array genotyping (1,578 vs. 3,165 in AoU and ATLAS). In AoU EUR participants, we observed a concordant direction of effect, supporting replication of the finding, although statistical significance was not reached. This may reflect the smaller sample size or other population differences, such as in SES and compliance (Figure S8e). Larger samples are needed for more formal replication. As a second step, we conducted a genome-wide association study (GWAS) meta-analysis across ancestries, but did not identify any significant loci (Methods; Figure S8f).

Finally, we used WES to test gene-level associations with semaglutide response (Regenie66; Methods). Given the modest sample size, we focused on EUR individuals and limited our test to the subset of proteins whose plasma abundance in humans was recently shown to be altered by semaglutide115. We identified a significant association of weight loss on semaglutide with PTPRU (PBonferroni = 7.6×10−3, β = −0.87) (Figure 6f; Figure S8g). For replication, we used AoU, with a cohort that was 21% smaller than the ATLAS cohort used (1,581 vs. 2,012 in AoU and ATLAS with WES). This analysis supported our original finding by showing a consistent direction of effect, although significance was not reached, possibly due to the smaller sample size or differences between the tested populations (AoU: β = −0.27, P-value = 2×10−1). Therefore, larger sample sizes in future studies will be needed to formally replicate this finding. However, the Bonferroni adjusted combined P from ATLAS and AoU remained significant, supporting this finding (PBonferroni = 2.3×10−2). This association involves 37 variants, of which one is common (rs2235937; nominally associated with weight loss; P = 9.7×10−3, β = −0.06 in EUR), and the rest are rare, the frequency of which varied across ancestries (Figure S8h; Methods). The direction of effect indicates that PTPRU activity is negatively associated with weight loss in patients taking semaglutide (variants in PTPRU with a predicted high or moderate functional consequence contribute to weight loss on semaglutide). Maretty et al. highlighted PTPRU as a protein whose abundance is elevated in individuals with higher genetic risk for increased BMI or type 2 diabetes and is downregulated by semaglutide treatment115. Although direct causality has not been established, these observations, together with our findings, support a model in which semaglutide-induced downregulation of PTPRU and genetically reduced PTPRU function promote greater weight loss during semaglutide treatment. This gene has no previous known functional relationship to weight loss or related metabolic functions. However, combined with the data from serum proteomics115, our genetic analysis nominates this protein kinase and its pathways as candidates for further investigation.

Discussion

Large-scale, longitudinal population studies have accelerated our understanding of the causes and consequences of a wide variety of biomedical conditions. The Framingham study116,117 laid the foundation for population-based cohorts, followed by efforts like the UKBB and health system biobanks3,8,9,11,53,72,84,118,119. Despite the known sources of errors and incomplete phenotyping in health system EHRs, multiple studies have shown that many of these factors can be mitigated, allowing robust analyses3,9,10,120. Indeed, we were able to validate dozens of previously identified genetic associations in ATLAS based on EHR-derived phenotypes alone. We also leverage the heterogeneity of genetic ancestries in our population to validate and extend our knowledge of disease burden and genetic risk factors.

Prior work has mostly relied on broad-scale genetic ancestry to estimate health risk.5,30,117 Fine-scale ancestry discerns more recent genetic variation via shared IBD,120 but smaller clusters can limit power. Through PheWAS and ExWAS across broad- and fine-scale ancestries, we identified numerous unreported genetic associations, many of which were subsequently replicated. We leveraged fine-scale ancestry definitions to benefit populations that are underrepresented in genetics and medical research122. For example, to our knowledge, ATLAS consists of the largest genetic dataset of Filipino individuals (n = 1,438 vs. 1,028 in the next largest dataset58), which can be used to improve precision care for this population and fuel discovery. Similarly, our Ashkenazi Jewish cluster of 14,261 participants, which is 2.8 times larger than any prior single-study genetic cohort57, enabled previously unreported associations and enhanced risk quantification. Future work should assess the transferability of our results across populations123.

Discerning genetic variant risk frequency and effect across fine-scale ancestries is powerful for optimized risk stratification and tailoring therapeutic intervention. It has been well demonstrated that diagnostic misclassification due to differences in variant frequencies across ancestries is a serious risk22. We emphasize that leveraging diversity within a single regional biobank can reduce confounding by ancestral genetic differences and non-genetic variables that differ across countries or geographically distinct biobanks. We demonstrated the application of longitudinal EHR to gain insights into health outcomes over time, by linking genetic factors to weight loss on semaglutide. This involved integrating dynamic EHR changes with detailed prescription information, including the medication dose and route, which we identified as major confounders. We considered only a single GLP-1RA drug, semaglutide, to avoid the complexity introduced by different GLP-1RAs, which have differing efficacy124. Lastly, our study highlights fine-scale populations with unreported low or high ACMG putative damaging variants for further investigation. We expect ATLAS to grow and become an engine for clinical intervention, allowing researchers to rapidly identify individuals for clinical studies and to implement precision medicine as the field evolves.

Limitations of the study

Certain limitations should be considered. First, due to the nature of EHRs, our study captures a partial view of patient phenotypes and care. Second, our analysis focuses on the risk of receiving a diagnosis for a disease, which is related, but not equivalent to the underlying risk of disease development. This may influence how our findings should be interpreted in the context of true disease incidence. Third, we relied on computational predictors for comparing the abundances of predicted damaging variants in ACMG genes across ancestries. Some predictors may introduce ancestry-related biases125,126, potentially affecting the accuracy of the results. To mitigate these potential biases, we applied rigorous filtering and only retained variants upon concordant agreement in a majority of nine tools. Additionally, comparing the numbers of predicted damaging variants across ancestries may be less stable when small ancestry groups are involved. We therefore excluded groups with fewer than 400 individuals and reflected uncertainty using error bars. Lastly, while our study introduces multiple previously unreported associations, findings that could not be replicated should be further evaluated in independent cohorts as they become available.

Resource Availability

Lead contact

Questions should be directed to the lead contact, Daniel H. Geschwind (dhg@mednet.ucla.edu).

Materials availability

No materials were generated in this study.

STAR Methods

EXPERIMENTAL MODEL AND STUDY PARTICIPANT DETAILS

The Institutional Review Board of the University of California, Los Angeles gave ethical approval for this work. The IRB Number for the ATLAS Initiative protocol is IRB#17-001013.

Human participants included in this study were from the UCLA ATLAS Community Health Initiative. Enrollment procedures, consent, and data collection are described in Johnson et al30. Participants provided informed consent for research under IRB#17-001013. The cohort description is provided below and summarized in Table 1 and Figure 1.

Cohort description

ATLAS enrollment reflects a modest over-representation of females by self-reported sex, especially among patients aged 20-60 (Figure 1c; Table 1). Among ATLAS participants, the median age at the most recent time point is 61 years for males and 56 years for females. Medical morbidity was notably more widespread in the biobank population compared to all adult UCLA patients. This point was evidenced by a significantly larger number of diagnoses in the biobank population compared to all UCLA patients (15.6 vs. 10.5 mean ICD codes per patient, one year after collection). The table “Baseline demographic information in ATLAS and in non-biobank UCLA patients” shows baseline demographic information of the UCLA ATLAS population, and a subset of all other adult UCLA patients (Rest of DDR) seen within one year of the ATLAS launch date. All table variables differed significantly between the ATLAS and non-ATLAS populations (Mann-Whitney test for continuous variables and chi-square test for categorical variables) after multiple-testing correction. This included a significant difference overall, as well as in the proportions of alive and deceased individuals.

The biobank population experienced a substantially higher number of diagnoses across the most prevalent medical condition categories, including endocrine, cardiovascular, musculoskeletal, gastrointestinal, neoplasms, neurological, respiratory, mental, genitourinary and infections (Figure S2o–p). For instance, endocrine/metabolic and cardiovascular diseases were diagnosed in 52.0% and 44% of the biobank participants versus 38% and 36% of non-biobank adults, respectively. Biobank participants also experienced substantially more neoplasms (30% vs. 19%). This likely led to significantly more medical encounters in the biobank population (119 vs. 64 mean total medical encounters per patient, and 16 vs. 12 mean inpatient days), which is consistent with expectations, as greater numbers of clinical visits facilitate the passive blood collection that powers our biobank. The most common diagnoses included hypertension (30% of participants), hyperlipidemia (25%), GERD (18%), anxiety disorder (17%), and depression (15%) (Figure S2m).

As time progresses, more diagnoses have an opportunity to be added, and the rate of diagnoses is consistent across a wide spectrum of organ systems (1-2 years after collection; Figure S2n). The table “Characteristics over time” shows time-dependent characteristics of the UCLA ATLAS population (n=59,949) at the time of collection, and a subset of all other adult UCLA patients seen within one year of the ATLAS launch date (“Rest of DDR”, within 1 year of ATLAS launch date; n=186,895). Both cohorts were limited to a subset of individuals who had an encounter between 1-2 years after the initial encounter. Encounter statistics (total encounters, total inpatient encounters, total inpatient days, types of encounters) summarize the encounters within the first year after the initial encounter. The changes within each cohort were tested for a non-zero difference using a Mann-Whitney test for variables available at two times, and a one-sample t-test for the encounter statistics. * denotes a significant difference in the change. The cohorts were also compared at each time point using the Mann-Whitney test (continuous) and the Chi-square test (categorical). After multiple-testing correction, ** denotes a significant difference between the two populations at the respective time (including change).

Baseline demographic information in ATLAS and in non-biobank UCLA patients

ATLAS-Overall ATLAS-Alive ATLAS-Deceased Rest of DDR-Overall Rest of DDR-Alive Rest of DDR-Deceased
n (%) 88,436 84,593 (95.7) 3,843 (4.3) 410,041 392,915 (95.8) 17,126 (4.2)
Age, mean (SD) 54.6 (17.0) 54.1 (16.8) 67.6 (14.8) 51.8 (18.5) 50.9 (18.1) 72.8 (14.8)
Self-reported sex, n (%): Female 49,346 (55.8) 47,688 (56.4) 1,658 (43.1) 234,189 (57.1) 225,735 (57.5) 8,454 (49.4)
Self-reported sex, n (%): Male 39,058 (44.2) 36,873 (43.6) 2,185 (56.9) 175,841 (42.9) 167,169 (42.5) 8,672 (50.6)
Self-reported sex, n (%): Unspecified, X 32 (0.0) 32 (0.0) 11 (0.0) 11 (0.0)
Self-reported race, n (%): American Indian, Alaska Native 785 (0.9) 762 (0.9) 23 (0.6) 2,233 (0.5) 2,175 (0.6) 58 (0.3)
Self-reported race, n (%): Asian 10,537 (11.9) 10,180 (12.0) 357 (9.3) 42,307 (10.3) 40,619 (10.3) 1,688 (9.9)
Self-reported race, n (%): Black, African American 4,079 (4.6) 3,857 (4.6) 222 (5.8) 23,420 (5.7) 22,256 (5.7) 1,164 (6.8)
Self-reported race, n (%): Caribbean/West Indian 161 (0.2) 161 (0.2) 335 (0.1) 332 (0.1) 3 (0.0)
Self-reported race, n (%): Middle Eastern or North African 2,276 (2.6) 2,227 (2.6) 49 (1.3) 5,086 (1.2) 5,017 (1.3) 69 (0.4)
Self-reported race, n (%): Native Hawaiian, Guamanian or Chamorro, Samoan, Other Pacific Islander 255 (0.3) 247 (0.3) 8 (0.2) 1,050 (0.3) 1,011 (0.3) 39 (0.2)
Self-reported race, n (%): Other Race 3,184 (3.6) 2,737 (3.2) 447 (11.6) 42,188 (10.3) 40,137 (10.2) 2,051 (12.0)
 Self-reported race, n (%): Unknown, declined to specify 11,329 (12.8) 11,050 (13.1) 279 (7.3) 62,735 (15.3) 61,692 (15.7) 1,043 (6.1)
 Self-reported race, n (%): White 55,830 (63.1) 53,372 (63.1) 2,458 (64.0) 230,687 (56.3) 219,676 (55.9) 11,011 (64.3)
Self-reported ethnicity, n (%): Hispanic/Latino, Cuban, Hispanic/Spanish origin, Mexican, Mexican American, Chicano/a, Puerto Rican 12,966 (14.7) 12,235 (14.5) 731 (19.0) 50,229 (12.2) 48,100 (12.2) 2,129 (12.4)
 Self-reported ethnicity, n (%): Non-Hispanic/Latino 68,784 (77.8) 65,829 (77.8) 2,955 (76.9) 301,251 (73.5) 287,269 (73.1) 13,982 (81.6)
 Self-reported ethnicity, n (%): Unknown, declined to specify 6,686 (7.6) 6,529 (7.7) 157 (4.1) 58,561 (14.3) 57,546 (14.6) 1,015 (5.9)
Self-reported smoking, n (%): Former 24,054 (27.2) 22,504 (26.6) 1,550 (40.3) 92,364 (22.5) 85,576 (21.8) 6,788 (39.6)
Self-reported smoking, n (%): Never 60,477 (68.4) 58,366 (69.0) 2,111 (54.9) 275,883 (67.3) 266,871 (67.9) 9,012 (52.6)
Self-reported smoking, n (%): Passive Smoke Exposure - Never Smoker 112 (0.1) 105 (0.1) 7 (0.2) 826 (0.2) 769 (0.2) 57 (0.3)
Self-reported smoking, n (%): Smoker 3,412 (3.9) 3,286 (3.9) 126 (3.3) 27,509 (6.7) 26,796 (6.8) 713 (4.2)
 Self-reported smoking, n (%): Unknown, declined to specify 381 (0.4) 332 (0.4) 49 (1.3) 13,459 (3.3) 12,903 (3.3) 556 (3.2)
Types of Encounters, n (%): Inpatient and Outpatient Encounters 28,879 (32.7) 25,725 (30.4) 3,154 (82.1) 72,278 (17.6) 62,279 (15.9) 9,999 (58.4)
 Types of Encounters, n (%): Only Outpatient Encounters 59,557 (67.3) 58,868 (69.6) 689 (17.9) 337,763 (82.4) 330,636 (84.1) 7,127 (41.6)
Total Encounters, mean (SD) 118.8 (146.3) 112.1 (137.5) 267.4 (230.9) 64.2 (95.8) 59.7 (88.1) 166.7 (174.3)
Total Inpatient Encounters, mean (SD) 0.8 (2.2) 0.7 (1.9) 3.8 (5.0) 0.4 (1.3) 0.3 (1.1) 2.2 (3.7)
Total Inpatient Days, mean (SD) 16.2 (33.2) 13.4 (28.6) 38.3 (53.6) 12.3 (25.8) 9.4 (18.9) 29.7 (47.2)

Characteristics over time

Variable ATLAS- At Collection ATLAS- 1 Year after Collection ATLAS-Change Rest of DDR-At Collection Rest of DDR-1 Year after Collection Rest of DDR-Change
Age, mean (SD) 56.4 (16.4) 57.6 (16.3) 1.2 (0.4)* 53.5 (17.9)** 54.8 (17.8)** 1.3 (0.5)*,**
BMI, mean (SD) 27.3 (6.2) 27.3 (28.2) 0.1 (28.7) 26.8 (24.0)** 26.7 (14.8)** −0.1 (27.6)
Diastolic BP, mean (SD) 75.5 (9.7) 75.4 (9.4) −0.1 (10.4) 75.4 (10.2)** 75.6 (10.1) 0.2 (10.5)*,**
Systolic BP, mean (SD) 125.0 (16.7) 124.8 (16.8) −0.2 (17.1) 125.3 (17.6) 125.3 (17.6) 0.0 (16.8)**
Number of ICD Codes, mean (SD) 14.2 (11.2) 15.6 (11.9) 1.5 (6.5)* 9.4 (8.4)** 10.5 (9.1)** 1.1 (5.1)*,**
Number of Phecodes, mean (SD) 7.1 (5.7) 7.8 (6.0) 0.7 (3.2)* 4.7 (4.3)** 5.2 (4.6)** 0.5 (2.6)*,**
Charlson Comorbidity Index, mean (SD) 1.0 (1.8) 1.1 (1.9) 0.1 (1.0)* 0.7 (1.5)** 0.7 (1.5)** 0.1 (0.8)*,**
Elixhauser Comorbidity Index, mean (SD) 2.6 (8.3) 2.9 (8.7) 0.3 (4.4)* 1.6 (6.8)** 1.7 (7.1)** 0.2 (3.6)**
Total Encounters, mean (SD) 23.2 (27.2)* 12.0 (15.8)*,**
Total Inpatient Encounters, mean (SD) 0.2 (0.7)* 0.1 (0.4)*,**
Total Inpatient Days, mean (SD) 1.4 (7.8)* 0.4 (4.2)*,**
Inpatient and Outpatient Encounters (%) 8,413 (14.0) 9,869 (5.3)**
Only Outpatient Encounters (%) 51,536 (86.0) 177,026 (94.7)**

We retrieved the latest results of the 39 most common laboratory tests: blood count, lipid panel, metabolic panel, HbA1c, and vitamin D,25-Hydroxy to explore test frequencies and variation in results. The largest variability in adults was observed for bilirubin (median = 0.4, IQR = 0.3), followed by triglycerides (median = 90, IQR = 66), alanine aminotransferase (median = 22, IQR = 14), neutrophils (median = 3.73, IQR = 2.07), and LDL cholesterol (median = 97, IQR = 50). Similarly, we retrieved information on the 50 most abundant filled prescriptions to learn about treatment and prescribing tendencies. (Table S1). As hospital visits and surgeries were the most common types of encounters in the biobank population, it is not surprising that among the most prescribed generic medications were Acetaminophen (n = 738,061 prescriptions), Ondansetron HCl (n = 636,165), Propofol (n = 633,830), and Lidocaine HCl (n = 492,838) (Table S1).

Ancestral diversity compared to other biobanks

ATLAS consists of 32% participants of non-European ancestry and a representation of five broadscale and 36 fine-scale genetic ancestry groups within a single health system (see METHOD DETAILS to understand how these were defined). By comparison, the vast majority of the UK biobank4 participants are EUR (95.8% EUR, 2.1% South Asian, and 2.1% AFR). FinnGen6 is ~98% Finnish EUR, Taiwan Biobank7 is over 99% EAS Han Chinese, Michigan Genomics Initiative11 is 88% EUR, and Geisinger MyCode169 is approximately 97% EUR. While others like All of Us, Mt. Sinai’s BioMe, and Vanderbilt’s BioVU reflect diverse urban populations, ATLAS captures a wider and more detailed range of genetic ancestries in the Los Angeles population. All of Us5 includes a larger portion of Black participants, but a smaller portion of Asian and Middle Eastern/North African self-identified individuals (under a combined race and ethnicity category: White 53.3%, Black or African American 21.2%, Hispanic or Latino 17.8%, Asian 3.1%, Middle Eastern or North African 0.6%). Mt. Sinai’s BioMe53, the only biobank with reported fine-scale ancestries, included 17 fine-scale clusters vs. 36 in ATLAS, and substantially fewer Asian individuals. Vanderbilt’s BioVU170 and Penn Medicine Biobank119 include small Admixed American populations, and these biobanks, along with the VA Million Veterans Program72 and the Colorado biobank include small Asian populations. The tables below show comparisons of ancestry numbers and ratios across biobanks. Unclassified participants are not included. Some biobank studies have larger sample sizes but lack inferred genetic ancestry and are thus not included.

Numbers of biobank participants by genetic ancestry

Biobank AMR AFR EUR EAS SAS Total
UK Biobank (National) - 9,633 431,805 - 9,252 450,690
All of Us (National) 28,901 34,037 101,613 3,255 - 167,806
Taiwan Biobank (National) - - - 108,955 - 108,955
Mount Sinai BioMe (Academic) 10,638 6,983 8,477 780 617 27,495
Vanderbilt BioVU (Academic) 2,466 15,597 69,810 896 414 89,183
Geisinger MyCode (Academic) - 1,377 44,522 - - 45,899
Michigan Genomics (Academic) 843 5,962 75,943 2,172 1,435 86,355
VA Million Veterans Program (Academic) 59,048 121,117 449,042 6,702 - 635,909
Penn Medicine Biobank (Academic) 711 11,300 30,360 680 573 43,624
Colorado Biobank (Academic) 18,137 7,466 145,070 4,080 - 174,753
UCLA ATLAS 12,822 4,061 62,902 8,080 1,761 89,626

Below, percentages exclude unclassified participants and therefore differ from Figure 1.

Percentages of biobank participants by genetic ancestry

Biobank AMR AFR EUR EAS SAS
UK Biobank (National) 0.0% 2.1% 95.8% 0.0% 2.1%
All of Us (National) 17.2% 20.3% 60.6% 1.9% 0.0%
Taiwan Biobank (National) 0.0% 0.0% 0.0% 100.0% 0.0%
Mount Sinai BioMe (Academic) 38.7% 25.4% 30.8% 2.8% 2.2%
Vanderbilt BioVU (Academic) 2.8% 17.5% 78.3% 1.0% 0.5%
Geisinger MyCode (Academic) 0.0% 3.0% 97.0% 0.0% 0.0%
Michigan Genomics (Academic) 1.0% 6.9% 87.9% 2.5% 1.7%
VA Million Veterans Program (Academic) 9.3% 19.0% 70.6% 1.1% 0.0%
Penn Medicine BioBank (Academic) 1.6% 25.9% 69.6% 1.6% 1.3%
Colorado BioBank (Academic) 10.4% 4.3% 83.0% 2.3% 0.0%
UCLA ATLAS 14.3% 4.5% 70.2% 9.0% 2.0%

METHOD DETAILS

Our study moves from characterizing the ATLAS cohort to ancestry-stratified analyses of disease burden – first at the broad-scale ancestry level, then at the fine-scale cluster level- followed by evaluations of common and rare genetic risk, and finally, an integrative case-study using longitudinal EHR data.

Array data and imputation

Array genotypes were obtained using the Global Screening Array. All data was mapped to GRCh38 and dbSNP, build 147171. Common haplotypes and variants were imputed using the TOPMed Freeze 5 panel29,30 using 668,127 observed single-nucleotide polymorphisms (SNP), resulting in a total of 50,757,223 high-quality called genotypes variants following imputation, an average of 2,048,050 per individual. Details regarding QC and the imputation procedure were described before29. Minimal QC was applied to the imputed genotypes. We retained only non-duplicated, bi-allelic variants with high imputation quality (R2 > 0.7), a minor allele frequency (MAF) > 0.1%, and a missing rate < 5%. Genotypes with a missing rate > 5% were excluded from further analysis. Concordance between genotypes determined by observed array variants and targeted Illumina sequencing was approximately 99.6%, determined using the vcf-compare command in VCFtools (v0.1.16)172.

Retrieving phenotype data

Demographic information, vitals, lab tests, International Classification of Diseases (ICD) codes, and medication prescriptions were retrieved from the UCLA Data Discovery Repository (DDR), established on March 2, 201329, containing deidentified participant EHR data from our health system.

Demographic, vital, and lab data were up to date as of Aug 24, 2024. For vitals and encounter counts, only records from in-person encounters (hospital visits, appointments, surgeries, office visits, and walk-ins) were considered. For vital signs, median values across all encounters were used to reduce the impact of potential recording errors or error-prone self-reported values. For lab data, the most recent values were considered. For Figure 1c, self-reported sex and age at the most recent time point at the time of analysis were used.

Yearly encounter numbers were retrieved in 2024. We included only complete yearly data up to 2023 to ensure consistency and avoid partial data from 2024. Hospital encounters during the COVID pandemic years (2019-2022) were excluded.

For associations involving clinical phenotypes, both ICD-9 and 10 were extracted. To ensure harmonized and comparable phenotypes across data sources, we adopted a structured, standard approach using phecodes59–61. We utilized the existing PheWAS catalog maps and applied standardized rules for case–control definitions. The ICD-9 codes were mapped to phecodes using phecode Map v1.259, while ICD-10 codes were mapped using phecode Map v1.2b160, both of which were obtained through the PheWAS catalog61. Cases were defined using the widely used “Rule of Two” approach61,173–176. For each phecode, participants were classified as cases if they had at least two occurrences of the same phecode, with occurrences spaced at least 30 days apart to ensure persistence of diagnosis. In total, there were 1,308 phecodes with at least 100 cases in ATLAS. Controls were defined as participants who had no record of the phecode in their EHR data and at least two recorded encounters in the system, following similar established rules for minimum data content61,177,178. Undecidable participants were excluded.

Phenotype validation

We used phecodes59–61 to define phenotypes and validated the generalizability of these phenotypes using several approaches. First, we tested cross-checked information between a variety of vital signs and lab tests and specific case-control groups defined by phecodes, using Wilcoxon tests and density plots, which confirmed the expected differences for these phenotypes. This included the following: 1) phecode-lab test pairs: Type 2 diabetes-Hemoglobin A1c, Hyperlipidemia-Trygllicoride, Vitamin D deficiency-Total 25-Hydroxy vitamin D, Deficiency anemias-Mean Corpuscular Volume (MCV), Deficiency anemias-Hemoglobin, and 2) the phecode-vital sign pairs: Essential hypertension-Blood pressure diastolic, Essential hypertensionBlood pressure systolic, Obesity-BMI, Anorexia nervosa-BMI, Short stature-Height (Figure S2a–j). Second, we replicated many known associations, including ancestry-related differences in disease prevalence for 28 major phenotypes. For example, across cancer types, we observed the highest risk for prostate cancer in AFR ancestry, stomach cancer in EAS, bladder cancer in AMR, and breast cancer in EUR179 (see more below and in Table S2). Across fine-scale ancestries, for example, we showed increased breast cancer and Crohn’s disease diagnoses in the Ashkenazi Jewish (IBD-03) cluster53,180 (see more below). We also replicated 11,756 variant–phenotype associations, such as a lower risk of Alzheimer’s disease for African American participants with the ε4/ε4 haplotype (see more below). PGS prediction of 28 traits revealed significant enrichment in the top PGS decile for 89% of these traits in EUR (see more below). Collectively, the replication of known associations using ATLAS-defined phecodes indicates high-quality, well-defined phenotypes that match external definitions. Third, we used EHR data on cancer diagnoses to validate prostate and breast cancer cases (the two most diagnosed cancers) against phecode-defined case-control groups. We found a high agreement of 95%, for both cancers, considering decidable patients based on phecodes (Figure S2k–l). The following sections describe replication results conducted to validate the quality of the data and the phenotype definitions.

Replication of variation in medical conditions across broad-scale genetic ancestries.

We assessed variation in medical conditions and disease prevalence across populations across preselected prevalent major conditions (Table S2). This confirmed known ancestry-disease associations. Across cancer types, we observed the highest risk for prostate cancer in AFR ancestry, stomach cancer in EAS, bladder cancer in AMR, and breast cancer in EUR179. We also replicated several known cardiovascular disease associations. For instance, AFR participants were more affected by hypertension and myocardial infarction, while those with SAS ancestry had the lowest risk of atrial fibrillation, despite the highest incidence of coronary atherosclerosis, an established but seemingly contradictory risk profile181. Major metabolic disorders were more prevalent in AFR than in other ancestries. This finding included diagnostic codes (9- and 10-ICD codes) related to Type 2 diabetes, hypercholesterolemia, and hyperlipidemia, supported by increases in metabolic-related laboratory measurements and vital signs (Figure S3e). Type 2 diabetes was more common in all non-EUR groups relative to EUR. Previous work has suggested that some of these differences reflect social and environmental, rather than genetic, factors182. With regard to neurological conditions, we find a higher risk of dementia in AFR participants, but a lower risk of migraines and Parkinson’s Disease in AFR compared to EUR patients, consistent with prior analyses183–185. Among neuropsychiatric disorders, anxiety and major depressive disorders were most frequent in EUR participants, and least frequent in the continental Asian cluster, as previously reported186,187. These findings demonstrate the utility of ATLAS in robustly replicating known associations within a single health system, reducing potential confounding due to geographic or health system-related effects.

Replication of variation in medical conditions across fine-scale genetic ancestries.

We tested the prevalence of 1,253 phecodes59–61 across fine-scale clusters with at least 100 participants (Table S3; selected tests, Figure 2c). As a measure of integrity, we tested for known increases in prevalence in these fine-scale ancestries, for example, replicating an increase in breast cancer and Crohn’s disease in the Ashkenazi Jewish (IBD-03) cluster. Similarly, we observed the known increase in gout in the Filipino cluster (IBD-09)188 and Alzheimer’s disease and dementias in the Puerto Rican cluster (IBD-15)189. Consistent with the global ancestry findings, African American (IBD-06) had the highest prevalence of hypertension.

Validation of PGS performance.

We leveraged genotyping in our cohort to calculate individual PGS for a range of cancer, cardiovascular, metabolic, neuro-psychiatric and autoimmune diseases and tested their relationship to disease risk. We focused on the top end of the PGS distribution (10%), compared with the 5th decile, in EUR participants (Figure 3; see the table below). For Type 1 diabetes, the top PGS decile was the most enriched (OR = 11.6 [7.6, 19.0], FDR = 4.1×10−25), with 41% of the diagnosed participants in ATLAS assigned to the top PGS decile. The second most enriched trait was Crohn’s disease, with 33% of diagnosed participants within the top PGS decile (OR = 5.5 [4.0, 7.7], FDR = 7.9×10−24), followed by gout with 27% (OR = 5.1 [4.1, 6.4], FDR = 2.0×10−45), testicular cancer with 25% (OR = 3.7 [1.2, 7.3], FDR = 1.2×10−4), and prostate cancer with 21% (OR = 3.4 [2.8, 4.0], FDR = 2.0×10−45) (Figure S5a–e). On average, the top decile of risk accounted for 18.6% of diagnosed participants across 28 preselected prevalent disorders. Twenty-five of 28 (89%) showed significant enrichment of participants within the top PGS decile (mean OR for significance tests = 2.9). Performance declined when we applied PGS to non-EUR populations, as expected55–58. This was due to both reduced sample sizes and the model fit (Figure S5f–i), identifying only 17, 27, 36, and 38% significant associations of diseases with the top PGS decile, for SAS, AFR, EAS, and AMR, respectively. Concordantly, the top PGS decile accounted for fewer cases (on average, 13, 15, 16 and 16%, for SAS, AFR, EAS and AMR). This further supports the need for larger, more ancestrally heterogeneous cohorts for clinical development of PGS55–58.

The OR and prevalence of cases for the top and bottom PGS deciles across traits in EUR

Trait OR top OR bottom Cases top (%) Cases bottom (%) Cases sample size FDR top FDR bottom
Atrial fibrillation 3.2 0.5 19.5 4.8 4565 6.60×10−66 1.90×10−13
Bipolar 1.6 0.7 15.5 7.4 1147 9.60×10−5 0.011
Bladder cancer 1.8 0.8 15.2 6.9 698 0.00028 0.3
Breast cancer 2.4 0.5 18.0 4.9 3016 2.20×10−26 2.40×10−08
Cerebrovascular disease 1.3 1.1 9.9 10.5 334 0.34 0.71
Colorectal cancer 1.6 0.6 15.0 5.9 842 0.0026 0.0064
Coronary atherosclerosis 2.5 0.5 16.0 6.1 7951 2.20×10−59 2.70×10−24
Gout 5.1 0.4 27.2 2.7 1638 2.00×10−45 1.10×10−05
Hypercholesterolemia 1.4 0.6 12.7 6.8 10836 1.50×10−14 5.70×10−24
Hyperlipidemia 1.6 0.6 12.2 7.8 21829 6.70×10−29 1.30×10−30
Hypertension 2.1 0.6 12.6 7.8 21517 2.50×10−63 1.10×10−28
Hypertrophic obstructive cardiomyopathy 1.7 0.4 17.8 3.9 129 0.13 0.11
Lung cancer 1.1 0.7 10.3 7.2 976 0.65 0.062
Major depressive disorder 1.6 0.7 13.0 7.1 10793 3.00×10−22 1.50×10−10
Malignant neoplasm of testis 3.7 0.4 24.9 2.9 173 0.00012 0.13
Melanoma 2.3 0.5 19.6 3.4 1196 8.80×10−12 1.00×10−4
Multiple sclerosis 2.8 0.3 26.1 2.9 379 6.00×10−7 0.0024
Myocardial infarction 1.6 0.8 12.9 7.2 1744 2.30×10−5 0.11
Ovarian cancer 2.1 0.8 17.9 6.7 403 0.00066 0.35
Pancreatic cancer 1.6 0.7 12.6 5.2 382 0.048 0.29
Prostate cancer 3.4 0.3 21.5 2.9 2703 2.00×10−45 1.10×10−17
Psoriatic arthropathy 1.5 0.9 14.1 8.0 503 0.029 0.53
Crohn’s disease 5.5 0.6 33.0 3.6 758 7.90×10−24 0.11
Schizophrenia 2.9 0.3 21.9 2.3 128 0.01 0.11
Systemic lupus erythematosus 2.9 0.9 22.0 6.3 569 4.40×10−9 0.56
Thyroid cancer 3.0 0.6 21.4 4.1 786 2.60×10−12 0.022
Type 1 diabetes 11.6 0.4 40.8 1.5 537 4.10×10−25 0.05
Type 2 diabetes 2.3 0.4 16.9 4.3 6302 2.20×10−48 8.60×10−26

Replication of known phenome-wide associations across ancestries.

Replication of genotype-phenotype associations is important to assess the quality of both genetic and phenotype data, as well as to evaluate the power and effectiveness of the biobank in detecting true genetic associations. To address this, we first examined the distribution of APOE alleles across fine-scale clusters, highlighting the increased frequency of ε4 risk alleles in the African American (IBD-06) and Bantu (IBD-35) clusters (Figure S6c) and replicating the finding that African American participants with the ε4/ε4 haplotype have a lower risk of Alzheimer’s disease compared to other populations (see in the table below FDR-significant results). Next, we replicated two well-known genetic associations190 in the African American cluster, between HBB rs334-A (Figure S6d; MAFIBD-06 = 5.2%) and a diagnosis of sickle cell anemia (PIBD-06 = 2.0×10−78; ORIBD-06 = 15.1) and between the Duffy null ACKR1 rs2814778-C polymorphism (Figure S6d; MAFIBD-06 = 77.1%) and a decrease in neutrophil count (PIBD-06 = 4.0×10−31; ORIBD-06 0.7). Using ATLAS blood work results, we further quantified the impact of rs334-A on mean corpuscular hemoglobin concentration (PIBD-06 = 5.2×10−15; βIBD-06 = 0.4), nucleated red blood cell count (PIBD- 06 = 1.6×10−10; βIBD-06 = 0.2), and mean corpuscular volume (MCV) (PIBD-06 = 3.3×10−8; βIBD-06 = − 0.3). Across multiple Asian clusters, we confirmed the association between a variant in high LD with the --SEA deletion, which causes inherited alpha-thalassemia191, and microcytic anemia. We observed a significant decrease in MCV in response to LUC7L rs372755452-A in the broad-scale EAS ancestry (PEAS = 6.1×10−88; βEAS = −1.6; MAFEAS = 0.9%), and the fine scale Chinese + Korean (PIBD-05 = 2.7×10−52; βIBD-05 = −1.6; MAFIBD-05 = 0.8%) and Filipino (PIBD-09 = 2.2×10−24; βIBD-09 = −1.8; MAFIBD-09 = 1.2%) clusters.

Next, we replicated the finding that AMR participants are twice as likely to carry a PNPLA3 rs738409-G missense variant that greatly increases the risk for non-alcoholic fatty liver disease (Figure S6d), a major cause of cirrhosis that often necessitates liver transplant67. Using fine-scale ancestral mapping, we studied the association between rs738409-G and non-alcoholic cirrhosis of liver in two Mexican American clusters (Figure S6e), IBD-04 (PIBD-04 = 2.7×10−14; ORIBD-04 = 1.90) and IBD-07 (PIBD-07 = 1.5×10−8; ORIBD-07 = 1.9), identifying similar risk but differing minor allele frequencies across these clusters (MAFIBD-04 = 45.0%; MAFIBD-07 = 52.9%). The same variant in Northern Europeans had a smaller effect (PIBD-01 = 3.6×10−9; ORIBD-01 = 1.5) and was half as frequent (MAFIBD-01 = 22.9%).

Ancestry-specific effects of APOE ε4 on Alzheimer’s disease risk

Haplotype Ancestry OR P-value CI low CI high
e4e4 All 2.87 1.25×10−41 2.45 3.28
e4e4 EUR 2.86 1.40×10−31 2.38 3.34
e3e4 All 1.06 3.88×10−19 0.82 1.29
e3e4 EUR 1.01 1.43×10−13 0.74 1.28
e4e4 AMR 4.18 5.99×10−6 2.37 5.99
e3e4 EAS 1.93 1.23×10−4 0.95 2.92
e4e4 AFR 1.76 3.10×10−3 0.59 2.93
e4e4 EAS 4.89 4.59×10−3 1.51 8.26
e3e4 Unclassified 1.80 8.83×10−3 0.45 3.15
e4e4 Unclassified 5.66 9.62×10−3 1.37 9.94

Replicating ancestry-specific rare variant associations.

Analyzing clinically relevant rare variants, known to have elevated frequencies in specific populations, replicated known ancestry-associated patterns, assessing the quality of the data and the robustness of ATLAS in capturing ancestry-associated allele frequency patterns. Focusing on Familial Mediterranean Fever (FMF) variants, we identified 753 carriers with the highest risk in the Armenian clusters192,193 (Figure S7a). Carriers of these variants also had a higher risk of the amyloidosis phecode, which is known to be associated with FMF (OR = 3.7 [2.2, 5.8]). The HBB:p.E7V variant, which is responsible for most sickle cell anemia cases, was carried by 273 participants, with elevated frequency in the African American cluster (ORIBD-06 = 51.4 [39.7, 67.0]; Figure S7b), in line with previous findings194. Finally, we explored protective variants in the PCSK9 gene causing lowered LDL levels195 and identified 49 carriers of loss-of-function variants within the African American cluster (ORIBD-06 = 9.5 [4.8, 17.8]) (Figure S7c), as is known20.

While the above variants were selected via a literature search, we also queried the entire ClinGen pathogenic or likely pathogenic (P/LP) variant catalog, aggregated by gene, to identify enrichment in fine-scale populations. Among replicated associations, variants in BRCA1 and BRCA2 were our first candidates of interest. Ashkenazi Jewish had the primary risk for carrying either ClinGen P/LP variants in BRCA1 or BRCA2 (BRCA1, ORIBD-03 = 47.1 [20.6, 133.0]; BRCA2, ORIBD-03 = 48.2 [23.6, 114.2]) (Figure 5a). Of note, these observations are based on six P/LP BRCA ClinGen variants found in ATLAS, which include only two out of three of the Ashkenazi Jewish founder alleles (at the time of the analysis, the Ashkenazi Jewish founder allele NM_007294.3:c.5266dup was not included in the ClinGen Evidence Repository). The population attributable risk (PAR) was 3.0 [1.8, 5.0]% and 3.2 [1.9, 5.3]% for BRCA1 and BRCA2 variants, respectively, in the Ashkenazi Jewish cluster for breast cancer. Independent of ClinGen, we considered the three BRCA Ashkenazi Jewish founder alleles and found that they were most prevalent in EUR among broad-scale ancestries, and in Ashkenazi Jewish among fine-scale clusters (Figure S7d–e). In Ashkenazi Jewish, the PAR for breast cancer, associated with the three Ashkenazi Jewish founder alleles combined, was 8.2 [5.9, 11.2]% when all ages were included and increased to 12.5 [8.5, 18.1]% in participants under 70. This increase is expected, since BRCA founder alleles are associated with early-onset breast cancer. In Northern Europeans, we observed enrichments of P/LP variants in MYOC (ORIBD-01 = 2.7 [1.7, 4.2]) that are associated with juvenile open angle glaucoma, with a PAR of 0.7 [0.2, 3.4]%.

Replication of EUR bias in ClinGen variants.

A bias toward EUR in ClinGen variants was recently shown in AoU23. We tested whether we could replicate this bias using a single health system. In this part, we asked if the total allele frequencies of the clinically actionable ACMG ClinGen P/LP variants vary across broad- and fine-scale ancestries. Overall, 17 ACMG genes had at least one P/LP variant in ClinGen present in ATLAS participants. For a more nuanced biological understanding, we divided the ACMG variants into two groups of rare LOF (n = 53) and rare P/LP missense (n = 131) variants. We defined ‘LOF’ for variants ranked as high-confidence LOF by LOFTEE94 and ‘missense’ based on the VEP137 “missense_variant” annotation (see below, Differences in total ClinGen allele frequency within ACMG genes across ancestries). We observed the highest frequency of rare P/LP LOF variants in EUR participants among broad-scale populations, and in the Ashkenazi Jewish cluster (IBD-03) among fine-scale clusters (Figure 5b). This number was primarily driven by BRCA P/LP alleles, which accounted for 63% of all ClinGen LOF P/LP variants in EUR and 94% in the Ashkenazi Jewish cluster (OREUR = 3.7 CIEUR = 2.6-5.4, PBonferroni = 6.3×10−17; ORAshkenazi Jewish = 6.5, CIAshkenazi Jewish = 5.3 - 8.1, PBonferroni = 3.5×10−62). Still, when Ashkenazi Jewish individuals are removed from the broad-scale EUR category, a substantial enrichment of ClinGen missense variants is observed for this group (OREUR (non-Ashkenazi Jewish) = 1.8 [1.2, 2.7], P-value = 0.002, PBonferroni = 0.01; Figure S7f), demonstrating that the signal is not derived from Ashkenazi Jewish individuals alone. No differences in rare P/LP missense total counts were identified across broad-scale ancestries, but across fine-scale ancestry clusters, Northern EUR participants (IBD-01) had significantly more (ORNorthern EUR = 1.4 [1.2, 1.8], PBonferroni = 0.02), with a similar trend observed in Southern Europeans that did not reach statistical significance. This suggests that clinical datasets are particularly biased toward Northern EUR driven by the composition of large contributing cohorts such as the UK Biobank. Missense variants were depleted in Ashkenazi Jewish, possibly since they are underrepresented in most EUR large scale cohorts, or due to technical differences in how studies or countries classify variants as P/LP. Excluding Ashkenazi Jewish as a sensitivity analysis resulted in a similar, albeit weaker trend, of enrichment of missense variants in Northern EUR (OR = 1.3 [1.01, 1.6], P-value = 0.04, PBonferroni = 0.5; Figure S6f).

ATLAS EHR baseline characteristics and their comparison to non-biobank UCLA patients

As all samples were collected incidentally, there was interest in characterizing the population, especially in comparison to non-biobank UCLA patients. To characterize the non-biobank UCLA patients while mitigating time-dependent confounding, we included a subpopulation with an encounter within one year of the ATLAS launch date. Encounters could be either inpatient or outpatient, and we summarized the relative proportion of patients with only outpatient encounters or at least one inpatient encounter (there were no patients with only inpatient encounters).

To identify clinical phenotype patterns in ATLAS and to compare these to patterns of non-biobank UCLA patients, we used all ICD codes from each patient to encapsulate past and present conditions, subject to practical challenges196. For the ATLAS population, we captured a snapshot using their encounter closest to their biobank sample collection date. For the non-biobank patients, we used their closest encounter to the ATLAS launch date. These encounters represented the baseline encounters for both populations. The ICD codes at these encounters were converted to phecodes and phecode groups59–61,197 to represent meaningful categories of disease. Phecodes were extracted using pandas v2.2.2 with Python v3.9.19. We reported the unique phecodes with a prevalence ≥5% in the UCLA ATLAS population. ICD codes at the baseline encounters were also used to calculate the Charlson and Elixhauser Comorbidity Indices 40,41,198, both measures that predict mortality. The scores were derived using the comorbidity R function v1.0.7154 with R v4.1.2. To compare the prevalence of disease categories between UCLA Biobank participants and non-biobank UCLA Health patients, we used logistic regression to test the association of each category with participant group, adjusting for age and sex.

In addition to characterizing patient populations at a baseline time, we also described differences in how they changed over time, an important consideration when assessing relative disease burden across populations. Participants were enrolled in UCLA ATLAS and were encountered within the health system at different times; therefore, we controlled the interval over which we measured change. We identified patients with at least one encounter between one and two years after their baseline encounter. From this subpopulation, we obtained patients’ first encounter within this window and summarized the cumulative encounters that occurred between the two times. We also used the ICD codes at the second encounter to derive new comorbidity scores, as well as extract updated phecodes. The change in comorbidity scores and increases in phecode prevalence provided insight into the evolving disposition of the UCLA ATLAS population over time.

Broad-scale genetic ancestry

The genetic ancestry of participants in the ATLAS dataset was estimated by assessing their proximity to the centroids of 1000 Genomes superpopulations in principal component (PC) space. PCA across common genetic variants demonstrates granular relationships and provides a quantitative basis for assessing relationships between ancestry and disease29,36. The top 20 PCs were calculated using the bed_projectPCA function (the bigsnpr155 R package v1.12.2) with default parameters. For every individual, the Euclidean distance to the centroids of the five broad-scale populations (AMR, AFR, EUR, EAS, SAS) was computed. Participants were assigned AMR or AFR ancestry if the nearest centroid corresponded to one of these populations, as these groups are well-separated in PC space. For EUR, EAS, and SAS ancestries, which exhibit more genetic overlap, a stricter distance threshold was enforced to minimize misclassification. Specifically, an individual was assigned to one of these ancestries if their distance to the nearest centroid was less than a scaled threshold, calculated as: Threshold=max(dist)×min(FST)/max(FST)×0.5, where max(dist) is the largest squared distance among centroids and min(FST)/max(FST) accounts for genetic differentiation. Participants who could not be assigned to any ancestry cluster under these criteria were labeled as “admixed/unknown.” Visualization was performed using the Boutros Plotting General (BPG) R package v.7.1.0156.

Broad-scale ancestry related associations

Variation in encounter numbers across broad-scale ancestries was tested using ANCOVA, adjusted for genetic sex, age, and BAS rank to account for healthcare access differences. Comorbidity index values across broad-scale ancestries were compared using ANCOVA, adjusted for genetic sex, age, and both BAS and ADI ranks, which provide complementary measures of socioeconomic status. Since not all patients had BAS and ADI values, the sample sizes for these analyses were reduced (numbers are specified in the body of each figure). Adjusted means and CI were calculated using the R emmeans package, version 1.10.5132.

To test differences in disease diagnosis across ancestries, phecodes (retrieved as described in Retrieving phenotype data) were associated with genetic ancestry populations using logistic regression, adjusting for age (age at diagnosis for cases and the latest age for controls) and genetic sex if applicable. The following disease-phecode pairs were used to define disease diagnosis: Type 2 diabetes-Type 2 diabetes, Sleep apnea-Sleep apnea, Essential hypertension-Essential hypertension, Hyperlipidemia-Hyperlipidemia, Anxiety disorders-Anxiety disorders&Anxiety disorder&Generalized anxiety disorder, Asthma-Asthma, Parkinson’s disease-Parkinson’s disease, Type 1 diabetes-Type 1 diabetes, Schizophrenia-Schizophrenia, Crohn’s disease-Regional enteritis, Chronic kidney disease-Chronic kidney disease, Stage I or II, Multiple sclerosis-Multiple sclerosis, Major depressive disorder-Major depressive disorder, Cerebrovascular disease-Cerebrovascular disease, Atrial fibrillation-Atrial fibrillation, Hypercholesterolemia-Hypercholesterolemia, atherosclerosis-Coronary atherosclerosis, Hyperlipidemia-Hyperlipidemia, Coronary Hypertrophic obstructive cardiomyopathy-Hypertrophic obstructive cardiomyopathy, Myocardial infarction-Myocardial infarction, Systemic lupus erythematosus- Systemic lupus erythematosus , Gout-Gout, Bipolar-Bipolar, Psoriatic arthropathy-Psoriatic arthropathy, Epilepsy-Epilepsy, Neurofibromatosis-Neurofibromatosis, Dementias-Dementias, Obesity-Obesity, Obsessive-compulsive disorders-Obsessive-compulsive disorders, Autism- Autism, Migraine- Migraine, Alzheimer’s disease- Alzheimer’s disease, Coronary atherosclerosis-Coronary atherosclerosis, Posttraumatic stress disorder-Posttraumatic stress disorder. Patients under 18 years old, with ambiguous sex or with “unclassified” genetic ancestry were excluded. Visualization was performed using the BPG R package v.7.1.0156.

PGS analysis

PGS were calculated using array data after imputation in EUR participants. Related participants based on their genetic similarity were excluded (defined using PLINK v2.0a109 with the relatedness coefficient --king-cutoff 0.05). PGS were calculated using pgsc_calc158 with the default settings and --min_overlap of 0.65. Logistic regression was used to associate every PGS with the corresponding trait based on phecodes (see Retrieving phenotype data). The following PGS model IDs from the PGS catalog199 and their phecode pairs were tested: PGS002250-Malignant neoplasm of ovary, PGS003766-Cancer of prostate, PGS000004-Malignant neoplasm of female breast&Breast cancer, PGS001794-Thyroid cancer, PGS000079-Melanomas of skin, PGS004884-Cancer of bronchus; lung, PGS002264-Pancreatic cancer, PGS003395-Colorectal cancer&Colon cancer, PGS000729-Type 2 diabetes, PGS002025-Type 1 diabetes, PGS004254-Regional enteritis, PGS004699-Multiple sclerosis, PGS000134-Schizophrenia, PGS004760-Major depressive disorder, PGS001806-Malignant neoplasm of testis, PGS004687-Malignant neoplasm of bladder & Cancer of bladder, PGS000039-Cerebrovascular disease, PGS004526-Essential hypertension, PGS004706-Atrial fibrillation, PGS004784-Hypercholesterolemia, PGS002029-Hyperlipidemia, PGS003726-Coronary atherosclerosis, PGS000739-Hypertrophic obstructive cardiomyopathy, PGS004528-Myocardial infarction, PGS000803-Systemic lupus erythematosus, PGS001789-Gout, PGS002786-Bipolar, PGS000198-Psoriatic arthropathy. For prostate and testicular cancer, only males were included, and for breast and ovarian cancer, only females. An adjustment was made for age at diagnosis for cases and the latest age for controls, genetic sex when both sexes were included, and the first ten genetic PCs. For Figure 3, the top and bottom PGS deciles compared with the 5th decile were considered, testing only EUR participants. FDR was used for multiple testing correction. The same process was repeated for other ancestries as presented in the supplementary material. Visualization was generated using the BPG R package v.7.1.0156.

Fine-scale ancestry pre-processing and quality control

Data.

ATLAS array data were merged with genotyping data from the 1000 Genomes Project38, the Simons Genome Diversity Project130, and the Human Genome Diversity Project131. BCFtools128 annotate was used to harmonize variant reference SNP ID (RSIDs), and BCFtools128 norm with a GRCh38 genome reference was used to standardize the genotyping data. Sites or individuals with more than 1% missing were removed using PLINK133 --mind and --geno. Only SNPs with MAF > 1% across all participants were kept.

Phasing.

SHAPEIT5134 with default parameters and the distributed GRCh38 map files were used to phase genotyping data, one chromosome at a time.

Identity-by-descent calling and processing.

A custom Python script that converts PLINK bed files to PLINK ped/map133 files while conserving phasing information was used to convert genotyping data. Centimorgan data for the map files were generated using the same genetic map data in SHAPEIT5.

Identity-by-descent segments were called using iLASH135 with the following parameters: slice_size 350, step_size 350, perm_count 20, shingle_size 15, shingle_overlap 0, bucket_count 5, max_thread 20, match_threshold 0.99, interest_threshold 0.70, min_length 2.9, auto_slice 1, slice_length 2.9, cm_overlap 1 and minhash_threshold 55. Identity-by-descent was called one chromosome at a time.

Identity-by-descent quality control.

Identity-by-descent segment outliers were removed as described in Belbin et al.53 and Caggiano et al.52. Segments overlapping centromeres or the human leukocyte antigen (HLA) region were removed. Regions that may have false positive identity by descent were identified using the following process and removed: total identity by descent per each SNP was identified by summing across all identity-by-descent segments that overlapped each SNP; SNPs with a total identity by descent greater than or less than three standard deviations from the genome-wide mean were removed.

Cluster identification.

To identify clusters, we followed the approach of Caggiano et al.52 and Dai et al.54 and applied Louvain clustering200. An undirected network is generated based on pairs of individuals who share identity-by-descent segments: nodes are the individuals, edges are the total, genome-wide identity-by-descent shared as the edges. We used the Python package, NetworkIt201, to iteratively run Louvain clustering four times to detect fine-scale clusters.

Cluster merging.

To avoid redundancy and maximize sample size, clusters were merged in two stages to produce a final set of consensus fine-scale clusters. First, following Caggiano et al.52and Dai et al.54, we computed pairwise Hudson’s FST using PLINK v2.0a136 across 378 clusters identified from the fourth layer of Louvain clustering. After removing clusters with fewer than 10 participants, to avoid unreliable FST estimates, we merged the remaining 356 clusters into 67 clusters if the pairwise FST was less than 0.001.

In the second stage, we refined clusters using IBD sharing. For each cluster pair, we examined all inter-cluster individual pairs to calculate: (1) IBDmean, the average cM shared, and (2) IBDprop, the proportion of pairs sharing at least 3cM of IBD as detected by iLash. We defined a composite metric for cluster merging, IBDweighted = IBDmean×IBDprop, that captures both the extent and prevalence of genomic IBD sharing. Next, clusters were sorted from the smallest to the largest number of ATLAS participants. For each cluster, we computed the IBDweighted score with all larger clusters and calculated the mean of these pairwise values. The cluster was then merged into the most similar larger cluster – i.e., the one with the highest IBDweighted score-only if that score exceeded the mean. Otherwise, the smaller cluster was retained independently. This process merged 14 small clusters into larger parent clusters, resulting in a total of 36 fine-scale clusters with ≥ 30 participants each for downstream analyses52. These fine-scale clusters were assigned unique identifiers (IBD-01 through IBD-36) and manually annotated with labels to ease interpretation. Due to filtering out clusters of small sample sizes before and after merging, not all individuals were assigned to a fine-scale ancestry cluster.

Cluster labeling.

We primarily relied on reference individuals, described in “Fine-scale Ancestry Pre-processing and Quality Control”, to add labels to clusters. When clusters did not contain reference individuals or were heterogeneous, we utilized patient-reported race, ethnicity, preferred language, and religious affiliation (in order of priority) to inform our cluster labeling. These aspects are not caused by identity-by-descent segment sharing but can be indicative of a shared culture for individuals within a cluster; these shared practices can influence a group’s demography, environment, and disease risk. Our labels are not definitive and are our best attempt to generate informative assignments for each cluster.

Fine-scale ancestry and clinical phenotype associations

We used logistic regression to model the association between fine-scale ancestry assignment and phecode prevalence, estimating OR with 95% confidence intervals using the logistf R package v1.26.0 for Firth’s bias-reduced penalized-likelihood logistic regression157. Differences were tested between every fine-scale ancestry cluster with at least 100 participants (with no ambiguous genetic sex and over the age of 18) and all other ATLAS participants. For phecodes that are sex-specific, only participants of the corresponding sex were included in the analysis. Adjustment was made for age at diagnosis for cases, and the latest age for controls, and for sex when applicable. Phecodes with defined categories according to the PheWAS catalog61, and that were diagnosed in at least 100 participants with no ambiguous genetic sex and over the age of 18 in total in ATLAS were tested, resulting in 1,253 phecodes. In situations where the number of cases in a cluster was small, exact ORs were not reported in the text to protect patient privacy. For visualization, fine-scale clusters were grouped into four panels (EUR, Asian, AMR, and AFR) based on the predominant broad-scale ancestry of participants. If the predominant ancestry was “Unclassified” (as in the Japanese and Egyptian Christian clusters), the second most prevalent broad-scale ancestry was used. The full phecode names presented in Figure 2c were: Hormones and synthetic substitutes causing adverse effects in therapeutic use, Cardiomegaly, Allergic reaction to food, Cirrhosis of liver without mention of alcohol, Vitamin B-complex deficiencies, Anemia of chronic disease, Cholesterolosis of gallbladder, Chronic renal failure [CKD], Dementias, Cancer of bladder, Cataract, Leukemia, Hereditary hemochromatosis, Attention deficit hyperactivity disorder, Glaucoma, Hyperplasia of prostate, Amyloidosis, Chronic pulmonary heart disease. These names were shortened in the figure and text for easier reading.

The same model was used to test differences in cardio-metabolic diseases using appropriate phecodes, for each fine-scale cluster within the same broad-scale continental ancestry. For this goal, clusters were assigned to broad-scale ancestries based on the predominant ancestry match among participants within each cluster. The largest population cluster within each broad-scale ancestry was used as the reference level for associations, adjusting for BMI, sex and age at diagnosis for cases, and the latest age for controls. As a sensitivity analysis, we tested whether the observed patterns persisted when correcting for SES factors (adjusting for ADI and BAS ranks) in addition to BMI, age, and sex. As a second step, we also added IPW (see IPW analysis).

FDR was used for multiple testing correction. The data were visualized using the BPG R package v.7.1.0156.

Comparison of ancestry sample sizes with published cohorts

The following studies were used to compare ATLAS’s broad-scale genetic diversity with other large scale biobanks: Halldorsson et al. (UKBB)4, Kurki et al. (FinnGen)6, Feng et al. (Taiwan Biobank)7, Zawistowski et al. (Michigan Genomics Initiative)11, Verma et al. (Geisinger MyCode)169, The All of Us Research Program Genomics Investigators et al.5, Shaw et al. (Vanderbilt’s BioVU)170, Verma et al. (Penn Medicine Biobank)119, Verma et al. (VA Million Veterans Program)72 and Wiley et al. (the Colorado biobank).

Studies that were used to compare sample sizes of ATLAS’s fine-scale ancestry clusters with other published cohorts with available genetic data (Table S3) included: Belbin et al. (BioMe fine-scale IBD clusters)53, Wu et al. and Li et al. (Ashkenazi Jewish)57,202, Haber et al. and Hovhannisyan et al. (Armenian)55,56, Mehrjoo et al. (Iranian)203, Larena et al. (Filipino)58, Sohail et al. and Ziyatdinov et al. (Mexican)204,205.

IPW analysis

IPW were calculated using logistic regression in R with ATLAS participants as cases and other UCLA Health patients as controls. As described above, we included UCLA Health patients who had at least one encounter within one year of the ATLAS launch date. The following covariates were included: self-reported race, sex, age group, BAS rank, and ADI rank. Individuals of unknown race were excluded. For each ATLAS participant, the probabilities were extracted from the model, and the weights were defined as 1 divided by the predicted probabilities. To assess whether previously unreported ancestry and disease status associations hold after applying IPW, we used the svyglm function (R survey package v4.4.8164) with a quasibinomial model, adjusting for age, sex, BAS rank, and ADI rank, including the calculated weights for ancestry groups with more than 5 cases. To compare comorbidity index values across broad-scale ancestries, considering IPW, linear regression was used, adjusting for age, sex, BAS rank, and ADI rank, including the calculated weights. Adjusted means and CI were obtained using the R emmeans package, v1.10.5132. Patients under 18 years old, with ambiguous sex or with “unclassified” genetic ancestry were excluded.

PheWAS

PheWAS were conducted using the Regenie v4.0 framework66 for all five broad-scale populations (AMR, AFR, EAS, EUR, SAS) and fifteen fine-scale clusters with at least 400 participants (IBD-01 through IBD-15). Related individuals were removed using a 0.05 kinship cutoff in PLINK v2.0a136. Sample sizes for each tested group can be found in Table S4. Association testing was performed separately within each group (i.e., population or cluster) for 1,437 binary traits and 41 quantitative traits, including ICD-derived diagnoses and clinical laboratory measurements (see Retrieving phenotype data). Samples were restricted to those in predefined inclusion lists (--keep), and trait-specific covariates were provided via --covarFile, including age, sex (modeled categorically with –catCovarList sex), BMI, and the top 10 genetic PCs. Quantitative traits were divided into two analysis groups based on missingness patterns, following Regenie’s recommendation that traits with similar levels of missing data are modeled together in Step 1.

Ridge regression (step 1).

For each group, unimputed array genotypes in bed format were filtered using PLINK v2.00a136 to produce a minimal set of variants suitable for estimating genomewide polygenic effects through ridge regression in step 1 of Regenie. We included autosomal variants with call rate ≥99% (--geno 0.01), minor allele frequency ≥1% (--maf 0.01), and Hardy-Weinberg equilibrium p > 1×10−15 (--hwe 1e-15). Additional linkage disequilibrium (LD) pruning was applied using a sliding window of 1000 SNPs, advanced by 100 SNPs, with an r2 threshold of 0.9 (--indep-pairwise 1000 100 0.9), producing an average of 389,209 SNPs per group. For binary traits, step 1 was run with a minimum case count of 50 (--minCaseCount 50). For quantitative traits, step 1 was run with Rank Inverse Normal Transformation (--apply-rint) enabled to stabilize variance across lab values with differing distributions. For all traits, a block size of 1000 (--bsize 1000) was used, leave-one out cross validation (--loocv) was enabled, and sex was marked as a categorical covariate (--catCovarList sex).

Association testing (step 2).

Prediction files from step 1 for each trait and imputed genotypes in BGEN format were used to perform association testing in step 2 of Regenie. For binary traits, a minimum allele count of 20 (--minMAC 20) and a minimum case count of 50 (--minCaseCount 50) were required, and the Firth approximation was performed (--firth --approx) using a p-value threshold of 0.01 (--pThresh 0.01). For quantitative traits, a minimum allele count of 20 (--minMAC 20) was required and Rank Inverse Normal Transformation (--apply-rint) was applied. For all traits, a block size of 500 (--bsize 500) was used and sex was marked as a categorical covariate (--catCovarList sex). Locus pruning was performed separately for each broad- and fine-scale ancestry using PLINK 2.0a to identify unique variant-phenotype associations.

Comparison to EMBL-EBI GWAS Catalog.

To standardize phenotypes for comparison, phecodes for binary traits were mapped to the Experimental Factor Ontology (EFO) using the text2term v4.5.0 Python package. Each phecode was assigned up to three top-matching EFO terms. To determine whether each unique variant–phenotype association was previously unreported, we queried the EMBL-EBI GWAS Catalog (June 27, 2025 freeze) for matching associations. An association from our study was classified as “previously unreported in the EMBL-EBI GWAS Catalog”, only if both of the following searches yielded no results. (1) Associations between the lead variant and any phenotype in the same phecode group as the associated phecode (e.g., for colorectal cancer, we searched for associations with all EFO terms in the broader “neoplasm” category). (2) Associations between the nearest protein-coding gene (within 10kb) and any phenotype in the same category as the associated phecode.

WES sample preparation

Genomic DNA libraries were created by enzymatically shearing high molecular weight genomic DNA to a mean fragment size of 200 base pairs. Multiplexity of exome capture and sequencing was achieved by adding unique asymmetric 10-bp barcodes to the DNA fragments of single samples during library amplification. Equal molar amounts of DNA samples were pooled for exome capture using a slightly modified version probe library of xGen exome research panel from Integrated DNA Technology (IDT). After PCR amplification and quantification of the captured DNA, samples were multiplexed and loaded to Illumina sequencing machines for sequencing to generate 75-base-pair paired-end reads. The samples in this study were sequenced using the Illumina sequencing machines, including NovaSeq 6000 with S2 or S4 flow cells and the NovaSeqX with 25B flow cells.

WES read alignment and variant detection

Sequencing reads in FASTQ format were generated from Illumina image data using the bcl2fastq program (v2.20, Illumina). Whole-exome reads alignment and germline small variant detection were conducted using the Original Quality Functionally Equivalent (OQFE) protocol described in Krasheninina et al., 2020206. Briefly, raw read files (FASTQ) were mapped to the GRCh38 reference obtained from https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/reference/GRCh38_reference_genome/ using BWA-MEM v0.7.17-r1188 in an alt-aware manner139. Duplicate reads were then marked with Picard v2.21.2127. The final CRAM files were compressed with SAMtools v1.2128. Germline variant detection was performed on each CRAM using a Parabricks accelerated version of DeepVariant v0.10.0 with a custom WES model140, resulting in a sample-level gVCF (genomic VCF). Per-sample gVCFs were merged with GLnexus v1.4.3141 into a joint-genotyped multi-sample project-level VCF (pVCF). Variant prediction was restricted to the exome capture region and the 100 base-pairs buffer on each side of the target regions. The WES was of high quality and reached an average coverage of 41.6-fold with a minimum of 24.2-fold, 90% having ≥ 34.7-fold in targeted regions, consistent with a previous publication207. The pVCF was converted to a PLINK file format using PLINK 1.9133 for downstream analyses.

These data underwent extensive quality control to ensure the absence of contamination, duplication, and other technical errors, as well as sufficient read depth to guarantee reliability and accuracy. Genetic duplicates were defined based on the aggregated genotype data of all sequenced samples. Sex was predicted based on the ratio of read coverage on chromosome Y over the whole exome read coverage. VerifyBamID v1.1.3142 was used to estimate sample contamination. Samples were excluded if they showed sex discordance, duplication, cross-individual contamination > 5%, or coverage < 20X in more than 20% off target regions. Overall, 386 samples were excluded for failing QC. Specifically, 249 samples were flagged for gender discordance, 12 had less than 80% of the exome covered at 20X, 80 showed contamination levels above 5%, and 74 were unresolved duplicates. After accounting for 29 samples that appeared in more than one exclusion category, the final number of unique samples excluded due to QC failure was 386. Alignment quality metrics were generated using Picard CollectHsMetrics v2.27.4127 and SAMtools stats v1.15.1 with default settings. MultiQC v1.27.1143 was used to systematically aggregate sample-level quality metrics. RTGtools v 3.12.1129 was used to assess variant QC metrics, including the number of variants for SNPs, small insertions and deletions; genotype counts; heterozygous-to-homozygous ratios for each variant type and transition/transversion (Ti/Tv) ratio. Genotypes were further confirmed using an alternative approach208 with 500 random samples, yielding a high concordance for on-target sites (median concordance of 99.5% for SNPs and indels).

WES summary statistics

For WES-based analyses, we used samples that also had array genotyping data (60,025 out of 61,797 patients with WES), which was necessary to consistently define broad- and fine-scale genetic ancestry. We annotated variants using Ensembl Variant Effect Predictor (VEP) v.112137 for 58,387 samples assigned to the EUR, AFR, SAS, EAS, or AMR ancestry class (i.e., excluding unclassified genetic ancestry samples). Annotations were performed using the GRCh38 cache and the corresponding reference FASTA file. We limited our analysis to autosomal chromosomes to avoid technical artifacts in variant calling caused by the differences in ploidy between males and females, as well as the high-sequence similarity between the X and Y chromosomes in certain regions209. Counts for single-nucleotide variants, indels, multi-allelic, synonymous, missense, and LOF variants were restricted to whole-exome sequencing (WES)-targeted regions. Consistent with a previous UKBB WES study210, we classified LOF variants as those with the following consequences: stop_gained, start_lost, splice_donor, splice_acceptor, stop_lost, and frameshift. To increase reliability, we further restricted LOF variants to those flagged as high confidence by LOFTEE94. For multi-allelic variants, the predicted function for each alternate allele was determined using the --pick-allele option based on the default ordered set of criteria defined by VEP137. To determine the number of variants with MAF < 1% while accounting for ancestry-specific allele frequency differences, we retained variants with MAF < 1% in at least one ancestry group.

For multi-allelic variants, we defined the minor allele as the second most common allele (including the reference) and classified the variant as rare if the cumulative allele frequency of all alternate alleles was < 1%. For rare multi-allelic variants, we incremented a sample’s count only if it carried an alternate allele with MAF < 1%. In multi-allelic cases where the minor allele was the reference or the alternate alleles had AF ≥ 1%, we incremented a sample’s count if it carried any alternate allele. For rare functional variants (e.g., missense), we incremented a sample’s count only if it carried an alternate allele with MAF < 1% corresponding to that functional category.

Replication of known elevated frequencies of rare P/LP variants

We identified variants via a literature search, where variants are known to be enriched within certain populations. For this goal, we selected: 1) Familial Mediterranean Fever (FMF) related variants in the MEFV gene (V726A, M694V, M694I, M680I and E148Q)211 2) the HBB:p.E7V variant, and 3) PCSK9 P/LP variants195.

For FMF and PCSK9, carriers were identified as participants who carry at least one corresponding pathogenic variant within PCSK9 or FMF variant group. For HBB:p.E7V, the frequency was estimated based on this variant alone. We only considered unrelated participants, using a kinship coefficient of 0.05 based on PLINK v2.0a136 king-cutoff. Fisher’s Exact test was used to test the carrier frequency of each cluster against the carrier frequency of the other clusters grouped together. The risk of amyloidosis (a common symptom of FMF) by carrier status was calculated considering carriers of known FMF and patients assigned the amyloidosis phecode (270.33).

Enrichment of ClinPGx and ClinGen P/LP variants

To test differences in frequencies of pharmacogenomic variants, 121 CPIC Level 1A variants were pulled from ClinPGx86,87. An overview of the analyzed variants can be found in Table S4. As variants spanned both common to rare frequencies, we used exome data for consistency. Overall, the variants were highly abundant in ATLAS, with 95.5% of all unrelated individuals carrying at least one. Of these variants, 69 were identified in the UCLA ATLAS exomes. To test enrichment of single alleles across clusters, “carriers” were identified as participants who carry at least one allele of a variant within a gene associated with a pharmacogenomic drug interaction or monogenic condition (unless only a single variant was involved). We then calculated carrier frequency as the number of identified carriers over the total number of participants with WES available and were unrelated using a kinship coefficient of 0.05 based on PLINK v2.0a136 king-cutoff. To identify clusters with a significantly different carrier frequency of variants, we applied a Fisher’s Exact test to the carrier frequency of each cluster against the carrier frequency of the other clusters grouped together. We then used FDR corrections for p-values generated from all Fisher’s exact tests and selected significant clusters (FDR ≤ 0.05) where at least 5 carriers were identified.

To identify clusters where participants were enriched for carriers of monogenic ClinGen variants associated with disease, we used pathogenic variants that were curated by experts within the field and underwent stringent review to be considered pathogenic from ClinGen91. ClinGen variants filtered to identify P/LP variants with autosomal dominant inheritance, autosomal recessive inheritance, or semidominant inheritance, resulting in 3,521 variants associated with 204 conditions overall according to ClinGen. We removed one HNF4A ClinGen variant due to a high allele frequency (MAF>0.01). Enrichment for variants, aggregated per gene, was tested as described above for ClinPGx variants with WES data. PAR was calculated for each cluster, showing significant enrichment of ClinGen P/LP variant frequencies aggregated by gene, focusing on genes associated with monogenic diseases with autosomal dominant or semidominant inheritance. To ensure sufficient power, we included only clusters with at least 30 individuals diagnosed with the disease with WES data available. We used the following formula: PAR = (p(RR-1))/(p(RR-1)+1), where p is the carrier frequency and RR is the ratio of the proportion of carriers with disease to the proportion of noncarriers with disease. For familial hypercholesterolemia, cases were defined as individuals with LDL levels exceeding 190 mg/dL, adjusted for statin use. In case of statin use, the LDL levels were divided by 0.7, as previously done195,212.

To calculate the PAR for breast cancer associated with the three Ashkenazi Jewish founder alleles, we used the formula above. Because individuals in this group are often aware of their genetic status and may pursue preventive breast cancer measures, we excluded individuals who underwent surgery (using the ‘acquired absence of breast’ phecode) but were not diagnosed with breast cancer, to improve the reliability of the estimate.

Differences in total ClinGen allele frequency within ACMG genes across ancestries

The list of all ClinGen91 P/LP variants identified in any ACMG secondary findings (SF) v3.2 genes93 was extracted as described above. Variants were labeled as missense or LOF, based on VEP v.112137 annotations. Missense variants were defined for ‘missense_variant’ variants according to the ‘consequence’ VEP output column. LOF was defined for high-confidence “HC” LOF variants based on LOFTEE94. All variants were rare across broad-scale ancestries. A WES plink file with all participants was filtered to include only the listed variants. Then, this file was broken into ancestry groups (i.e., broad-scale population, or fine-scale cluster) using PLINK v2.0a136. Related participants were removed using a 0.05 kinship cutoff in PLINK v2.0a136. The frequency of each allele in every population group was calculated with the PLINK v2.0a --freq command, and total frequencies were summed up for rare missense and LOF variants separately. To test differences in the allele counts across populations, allele dosages were calculated with PLINK v2.0a136 --recode A option. Dosages were summed to calculate the total alternative (ALT) allele count within each population. The number of reference (REF) alleles was defined as twice the number of individuals in the group minus the number of ALT alleles. Fisher’s Exact tests were applied to test differences between the number of REF and ALT alleles in each population compared to all other participants not assigned to that specific group, considering only groups with over 400 participants. Bonferroni correction was applied for multiple testing correction. Plotting was done using the BPG R package v.7.1.0156.

Variant annotation using a consensus of computational tools

Variants from exome sequencing were annotated using VEP v.112137 with the dbNSFP v4.9a138 and LOFTEE94 plugins installed. The LOFTEE high-confidence “HC” flag was used to select for LOF variants predicted to have deleterious effects.

Missense variants were assigned a 9-point deleteriousness score based on a consensus of nine missense deleteriousness prediction toolkits, similar to methods described in prior biobank-scale rare variant studies. We used the Critical Assessment of Genome Interpretation (CAGI) project95 to prioritize well-performing tools not trained on the same features. We selected five meta-predictors – ClinPred144, MetaRNN145, BayesDel_addAF146, VARITY_R147, REVEL148 – and four stand-alone predictors – AlphaMissense149, MutPred2150, VEST4151, ESM-1b152 – for use in our analysis. We assigned each variant a binary score per tool based on dbNSFP rank scores: 1 if the variant’s score exceeded the threshold score for being more likely a deleterious ClinGen or ClinVar variant than a background variant (Figure S7g–h), and 0 otherwise. Summing these binary scores produced a deleteriousness score ranging from 0 to 9, with predicted damaging missense variants scoring ≥ 5 retained for downstream analysis.

ACMG putative damaging variant distribution across ancestries

A WES PLINK file was filtered to include computationally predicted LOF and predicted damaging missense variants (see above) in ACMG SF v3.2 genes93, providing a more comprehensive evaluation than the one based solely on ClinGen P/LP variants, which included only 17 genes. A WES plink file with all participants was filtered to include only the ACMG putative damaging variants, and the file was split into ancestry groups (broad- and fine-scale ancestries) using PLINK v2.0a136. Related participants were removed using a 0.05 kinship cutoff in PLINK v2.0a136. Allele frequencies were calculated with PLINK v2.0a, and only rare variants (MAF < 1% in all broad-scale ancestries) were kept. Allele dosages per individual were extracted using PLINK v2.0a, and the total rare LOF and predicted damaging missense alleles were counted separately per participant. First, the differences between the total numbers of REF and ALT alleles across ancestries were evaluated with a Fisher’s Exact test as described above (ClinGen ACMG analysis). Second, a Mann-Whitney U test with a Bonferroni correction was applied to test the difference in the distribution of rare LOF and predicted damaging missense counts per individual between any broad- or fine-scale group compared to all others, with groups that included at least 100 participants. Figure S7I presents the Mann-Whitney U test results, with statistically significant differences indicated by an asterisk (*). Visualization was done with BPG156.

ExWAS

ExWAS were conducted using the Regenie v4.0 framework66 for all five broad-scale populations (AMR, AFR, EAS, EUR, SAS) and fifteen fine-scale clusters with at least 400 participants (IBD-01 through IBD-15). Related individuals were removed using a 0.05 kinship cutoff in PLINK v2.0a136. Sample sizes for each tested group can be found in Table S6. Within each group, imputed genotype dosages in BGEN format (--bgen) and prediction scores from PheWAS step 1 (--pred) were used to conduct gene-based association testing for binary and quantitative traits across 17,537 genes. Samples were restricted to those in predefined inclusion lists (--keep), and trait-specific covariates were provided via --covarFile, including age, sex (modeled categorically with --catCovarList sex), BMI, and the top 10 genetic principal components.

Genotype QC (step 2).

Whole-exome genotypes in bed format were filtered using PLINK v2.0a, retaining variants with a call rate ≥ 90% (--geno 0.1), a minor allele count ≥ 1 (--mac 1), and a Hardy-Weinberg equilibrium p-value > 1×10−15 (--hwe 1e-15). In addition, 784,548 variants overlapping low-complexity regions (LCRs) were excluded prior to analysis. Briefly, SNP positions were extracted from the .bim file, intersected using bedtools v2.29.1153 with annotated LCRs from the Genome in a Bottle Consortium213, and filtered from the genotype files, resulting in 12,326,160 exome variants retained for burden testing.

Gene-based association testing (Step 2).

Filtered whole-exome genotypes in BGEN format and prediction files from PheWAS step 1 were used to perform gene-based association testing. Regenie-style annotation, set, and mask files for predicted deleterious LOF and missense variants (see Variant annotation using a consensus of computational tools) were created programmatically using Python v3.11.9 with the polars v1.2.1 package. For binary traits (--bt), a minimum case count of 50 (--minCaseCount 50), a minimum minor allele count of 5 (--minMAC 5), and a block size of 1000 (--bsize 1000) were enforced. Firth logistic regression with saddlepoint approximation (--firth --approx) was applied for variants with p < 0.01 (--pThresh 0.01). For quantitative traits (--qt), rank inverse normal transformation (--apply-rint), a minimum minor allele count of 5 (--minMAC 5), and a block size of 500 (--bsize 500) were enforced. Regenie’s implementation of the RGC gene-based p-value test (–rgc-gene-p) was enabled for all analyses. Additionally, variants were binned by minor allele frequency using 1% bins (--aaf-bins 0.01), and SNP-level membership for each burden mask was recorded (--write-mask-snplist).

Variant-based association testing (step 2).

In addition to gene-based analyses, single variant association testing was performed for selected predicted deleterious LOF and missense variants using filtered whole-exome genotype dosages in BGEN format and prediction scores from PheWAS step 1. For both binary and quantitative traits, tests incorporated the same set of covariates (age, sex modeled categorically, BMI, and the top 10 genetic principal components) and enforced a minimum minor allele count of 5 (--minMAC 5). For binary traits, Firth logistic regression with saddlepoint approximation was applied for variants with p < 0.01, along with a block size of 1000, while quantitative traits were evaluated using rank inverse normal transformed phenotypes with a block size of 500. A Bonferroni-adjusted p < 0.05 was defined as the cutoff for significance.

PheWAS and ExWAS replication in AoU, UKB, and Taiwan Biobank

PheWAS replication in AoU (logistic regression).

Replication analyses were conducted utilizing the Python Hail package v0.2.134165 on AoU Workbench using the Controlled Tier Dataset v8. Phenotypes were obtained by querying OMOP databases for per-patient ICD codes, which were subsequently converted to phecodes v1.2 utilizing the R PheWAS package v0.99.6176. Participants with both phecodes and single-read WGS data were filtered to remove participants flagged from genotype QC or from relatedness QC, resulting in 291,082 total participants. Logistic regression was performed within the given replication genetic ancestry cohort by testing case/control status for a given phenotype and SNP, utilizing age, sex, age^2, sex*age, sex*age^2, and genotype PCs 1-10 as covariates.

PheWAS replication in AoU (All by All).

Replications for variant-trait associations were obtained by querying the Controlled Tier Dataset v8 “All by All” allele count/allele frequency (ACAF) variant result tables using the Python Hail package v0.2.134 on AoU Workbench.

PheWAS replication in Taiwan Biobank.

Replications for variant-trait associations were obtained by querying the summary statistics portal at https://taiwanview.twbiobank.org.tw/pheweb.php7.

ExWAS replication in AoU (All by All).

Replications for gene-trait associations were obtained by querying the Controlled Tier Dataset v8 “All by All” rare variant result tables using the Python Hail package v0.2.134165 on AoU Workbench.

ExWAS replication in UKB.

Replications for gene-trait associations were obtained by querying the AstraZeneca summary statistics portal at http://azphewas.com/69.

PheWAS and ExWAS replication in BioMe

The BioMe Biobank consists of electronic health records and genetic data from approximately 60,000 participants from the Mount Sinai Health System in New York. Participant recruitment was between 2007 and 2023. This study was approved by the Icahn School of Medicine at Mount Sinai’s Institutional Review Board (Institutional Review Board 07–0529). All study participants provided written informed consent.

Genetic data.

BioMe participants were genotyped using the Illumina Infinium Global Diversity Array (GDA; number of participants, N=23,430; number of variants, n=1,833,111) or Infinium Global Screening Array (GSA; N=32,595; n=635,623). Quality control consisted of removing participants with a call rate <95%, a mismatch between self-reported and genetic sex, and high heterozygosity. Duplicated sites and sites with a genotyping rate of <95% were removed. QC was done with PLINK2. Data was imputed with the TOPMED imputation server214 with genome build hg38. Genetic PCs were calculated across participants using PLINK2 for the genotyping sites after LD pruning. Exome sequencing data were generated by the Regeneron Genetic Center215. Quality control consisted of using the Goldilocks Filter210, and variants with quality scores <3 or depth of coverage scores <7 for SNPs, or <5 or depth of coverage <10 for indels were removed. Monomorphic sites were removed. Genetic ancestry was assigned using a random forest classifier trained on principal components derived from the 1000 Genomes data to align with the genetic ancestry assignment in AoU5. Participants were assigned a genetic ancestry using 10 PCs.

Phenotype data.

BioMe laboratory data were processed according to the QualityLab pipeline216. Briefly, quantitative lab values obtained in different clinical contexts (ambulatory, emergency, inpatient, and urgent care), were cleaned. Only labs with at least 100 patients and at least 1000 numeric observations were considered. Only adult (> 18 years) lab values were considered. Per lab, outliers were removed, defined as values as greater or less than 4 standard deviations from the mean of that lab. Non-numeric labs were removed. Labs must have had >70% of the reported units matching. After cleaning, a median value and median age was calculated per person for each lab test in each clinical context and across all contexts. Laboratory phenotypes were manually assigned a match to UKBB phenotypes using available metadata.

Electronic health phenotype data in the form of ICD-10 codes were collapsed into PhecodeX phecodes217. Phecodes were transformed into a binary matrix, where each row was an individual and each column was a phecode. A participant had a 1 if they ever were diagnosed with that phecode, otherwise their value was set to 0. Age was calculated as current age, defined from 01/01/2025.

Association testing.

Association testing was performed using Regenie v4.066. Burden testing was performed using SKATO and the following masks: missense, missense_pLoF, pLoF, and synonymous. PLoF variants were defined using LOFTEE94.

Semaglutide investigation

EHR data curation.

ATLAS GLP1-RAs prescriptions, including medication names, start and end dates, discrete dose, usage instruction, strength and route (oral or subcutaneous), were retrieved and grouped based on the following simple generic names: dulaglutide, semaglutide, liraglutide, exenatide, albiglutide and lixisenatide. For all statistical analyses, only semaglutide users were considered. Usage start dates were defined based on the earliest prescription start date for each patient. In the case of a missing start date, the prescription ordering date was used instead (in most cases, these two fields were identical). Overlapping medication periods were handled such that when a new prescription started, the previous one was considered to have ended. Similarly, missing end dates were determined using the start date of the next prescription when available. In the case of completely overlapping prescriptions with different instructions, the combination of both routes and the weighted average dose was considered. If the discrete dose information was missing, the medication dose was extracted from the instructions’ free text field. If the instructions also omitted the dose information, it was imputed for a given medication type, based on the weighted dose average from the entire cohort on the corresponding type. Initial weight and BMI were defined as the median of all available weight or BMI measurements recorded from in-person visits, within 6 months prior to the first prescription start date. Utilizing the median rather than a single data point helped minimize the likelihood of recording errors in the EHR. The percentage of weight change was obtained for every weight measurement recorded on in-person visits within the period of active prescriptions between 4-60 weeks in total on semaglutide. The medication dose at each time point was defined as the weighted sum of all medication doses (doses multiplied by the number of prescription weeks) by the measurement date. In case both oral and subcutaneous medications were used within a period, a combined route category was defined.

Statistical analysis.

To identify the overall weight loss patterns across time, the FPCA R function from the fdapace package v.0.6.0159,160, which is suited to plot smoothed longitudinal data with repeated measurement, was used. This identified a consistent weight loss pattern up to 60 weeks, and sparse data points beyond ~150 weeks (Figure S8C). Thus, we restricted all analyses to this period.

For all association tests, only participants aged 18 years or older were included. For all analysis parts that involved longitudinal data with repeated measurements, a linear mixed model with the bobyqa optimizer and an increased function evaluation limit (maxfun = 10000) was used (lmer R function; the lmerTest R package v.3.1.3161). In each case, ANOVA was used to identify differences between models to define the best way to model a potential nonlinear relationship between weeks and weight loss, when treating weeks as a fixed effect, based on the Akaike Information Criterion (AIC). In some cases, the best model was achieved using a restricted cubic spline (RCS) with the rcs R function from the rms package (v.7.0.0), applied to weeks. In other cases, a polynomial function of weeks provided a better fit. To plot smoothed longitudinal data with 95% CI, the fitted model values (excluding covariates) and 95% lower and upper confidence bounds, were extracted using the visreg package162 v.2.7.0 and visualized using the BPG package156.

For testing the effect of baseline factors, fixed effects were defined for the medication dose, route, sex, age, initial BMI, first ten genetic PCs (excluding PC6 due to a strong collinearity with PC5 that disrupted the model convergence) and weeks on semaglutide. The model included both random intercepts and slopes for weeks on semaglutide. Bonferroni correction was applied to control for multiple testing.

For testing differences across ancestries, fixed effects were defined for the medication dose, route, sex, age, initial weight, a polynomial function of weeks on semaglutide, and an interaction between a polynomial function of weeks and genetic ancestry categorical class (EUR, AFR, EAS, SAS and AMR). The overall number of weight measurements was 24,145, with a mean of 5 repeated measurements per patient. Ancestry sample sizes were: European (EUR), 3189; African (AFR), 373; admixed American (AMR), 913; East Asian (EAS), 291; South Asian (SAS), 107. The model included both random intercepts and slopes for the polynomial function of weeks on semaglutide. An ANOVA was performed on the model with a Bonferroni correction to assess the global effects of ancestry and ancestry×weeks interaction. When a significant result was found, post hoc comparisons between groups were conducted using the summary lmerTest function (v.3.1.3161), applying a Bonferroni correction to all 12 class or class × time interactions.

To test the relationship between PGS and weight loss, related participants based on their genetic similarity were excluded (defined using PLINK v2.0a136 with the --king-cutoff 0.05). Scaled BMI (PGS000027) and DM2 PGS (PGS000729) in EUR participants were divided into three equal bins each: low, intermediate and high scores. Then, using longitudinal data, for each trait, fixed effects were defined for the categorical PGS bins, medication dose, route, sex, age, initial weight, first ten genetic PCs (excluding PC6 due to a strong collinearity with PC5 that disrupted the model convergence) and a polynomial function of weeks. Random intercepts were defined to account for repeated measures within individuals. Bonferroni correction was applied to control for multiple testing. To test the relationship between PGS and weight loss, relying on a simplified model where the maximum weight loss was considered for each EUR participant, linear regression was applied. The model was adjusted for the medication dose and route, 10 genetic PCs, age, sex, and initial weight. Visualization was made with BPG156, with smoothed data and 95% CI using loess.as R function (fANCOVA package v.0.6.1163).

GWAS.

The analysis was performed to test a relationship between the maximum weight loss on semaglutide and common genetic variants, for each broad-scale ancestry using SAIGE167. For step one, array observed variants were used following PLINK v2.0a127 filtering, with the flags: --maf 0.01 --mind 0.1 --geno 0.1 --hwe 1e-6 . For the second step, array observed and imputed variants were used following plink filtering with --maf 0.01 --geno 0.05 --hwe 1e-6. A quantitative analysis was conducted with the traitType flag. The medication dose and route, the first five genetic PCs, age, sex, and initial weight were used as covariates. METAL168 was used for meta-analysis.

Gene-level tests.

To identify genes genetically associated with weight loss, we used Regenie66 with an additive model for gene-level tests. Only EUR semaglutide users that are not related (relatedness coefficient --king-cutoff 0.05) with WES data were used (n = 2,012), and for each, the maximum weight loss record with the corresponding number of weeks was considered, using information on aggregate dose and route. The list of candidate genes was limited to Bonferroni-significant proteins whose plasma abundance was altered by semaglutide treatment115. We considered variants within these genes with a predicted moderate or high impact on the protein function, according to VEP v.1.2137. For step one, the variant list was limited to observed array SNPs following a PLINK v2.0a136 filtering with: --maf 0.01 --mac 100 --indep-pairwise 1000 100 0.9 --chr 1-22 --snps-only --geno 0.1 --hwe 1e-15. In step 2, we applied PLINK v2.0a filtering to variants across all EUR participants using the parameters --geno 0.05 and --hwe 1e-6. Subsequently, the file was filtered to include only semaglutide users and was used in Step 2. The medication dose and route, the first 10 genetic PCs, age, sex, and initial weight were used as covariates. Bonferroni correction was applied to control for multiple testing. Frequency of variants involved in PTPRU association with weight loss on semaglutide by ancestry can be found in the table below.

Replication.

For replication analyses in AoU the same processes and tools were conducted in EUR individuals. Controlled Tier Dataset v8 was accessed through the Researcher Workbench and imported in PLINK format, including exome variants in chromosome 1 (to extract PTPRU variants) and common ACAF variant datasets (for PGS calculation). Genetic ancestry and PCs derived from pre-defined assignments provided by AoU5. PTPRU VEP137 variant annotations were obtained using Python Hail v0.2.134165 to identify functional consequences. For the Regenie analysis, related individuals were not removed to maximise power, as Regenie is designed to account for relatedness. The combined P-value of ATLAS and AoU was calculated using Fisher’s combined probability test, implemented with the fisher function in the poolr R package v1.2.0166. Multiple testing was addressed using a Bonferroni correction that accounted for the number of gene-level tests in the discovery cohort, along with one replication P-value and one combined P-value.

Frequency of variants involved in PTPRU association with weight loss on semaglutide by ancestry.

Variant AFR AMR EAS EUR SAS
1:29236667:T:C 0 0 0 1.2×10−5 0
1:29258717:G:A 0.00018 0 0 0.00013 0
1:29259276:C:T 0 0 0 1.2×10−5 0
1:29259313:A:G 0 0 0 4.9×10−5 0
1:29259883:C:G 0 0 0 3.7×10−5 0
1:29259912:A:C 0 0 0 1.2×10−5 0
1:29260003:G:C 0 0 0 1.2×10−5 0
1:29260895:A:C 0.00055 0.0025 0 0.0027 0
1:29275496:G:A 0 0 0 7.3×10−5 0
1:29275715:G:T 0.0018 0.011 0.0018 0.0092 0.023
1:29279056:G:A 0 0 0 1.2×10−5 0
1:29279527:G:T 0 6.0×10−5 0 1.2×10−5 0
1:29279541:C:A 0 0 0 1.2×10−5 0
1:29282705:G:A 0.0013 0.0028 0 0.0050 0.0018
1:29282710:C:T 0 0 0 8.5×10−5 0
1:29284754:C:T 0 0 0 9.7×10−5 0
1:29291897:G:A 0 0 0 0.00023 0.00045
1:29291929:G:A 0 0.00030 0 2.4×10−5 0
1:29291942:C:T 0 0.00042 0.00020 0.00052 0
1:29291958:A:G 0.0058 0.00030 0 2.4×10−5 0
1:29291978:G:A 0 0 0 0.00015 0
1:29303853:A:G 0 0.00036 0 0.00024 0
1:29303863:C:T 0 0.00012 0 0.00032 0
1:29303875:G:T 0.094 0.0065 9.9 × 10–5 0.00078 0.00045
1:29304795:G:A 0 0 0 1.2×10−5 0
1:29305397:A:G 0.23 0.38 0.48 0.25 0.38
1:29311488:C:T 0.00018 0 0 4.9×10−5 0.00045
1:29311513:C:A 0 0 0 3.7×10−5 0
1:29311706:G:C 0 0 0 1.2×10−5 0
1:29312579:C:T 0 6.0×10−5 0 9.7×10−5 0
1:29312615:G:A 0 0.0011 0 0.00015 0
1:29315374:C:T 0 5.99×10−5 0 2.4×10−5 0.00045
1:29315448:G:A 0 0 0 1.2×10−5 0
1:29317767:C:T 0 0.00036 0 0.0014 0
1:29317838:G:A 0 0 0 0.00016 0
1:29320709:A:G 0 0 0 6.1×10−5 0
1:29323643:C:T 0 0 0.00030 1.2×10−5 0

QUANTIFICATION AND STATISTICAL ANALYSIS

Statistical analyses were performed using R, PLINK, Regenie, and related tools. Logistic regression, linear regression, Firth-penalized regression, linear mixed models, and Fisher’s exact tests were applied in the relevant analyses, adjusting for appropriate covariates. Multiple-testing correction was applied using the FDR or Bonferroni adjustment, as specified in the relevant analyses. Detailed statistical models and software parameters are provided in the METHOD DETAILS section.

ADDITIONAL RESOURCES

An interactive web portal enables users to explore the PheWAS results, examine fine-scale ancestry associations with clinical phenotypes, and download supplemental data at atlas-phewas.mednet.ucla.edu.

Supplementary Material

1

Figure S1. Whole-exome sequencing quality control, related to the STAR Methods. a. Coverage distribution. The mean and median values were calculated across all samples. Fold enrichment is the degree to which the baited region is enriched compared to the background genomic region. b. Percentages of bases above different coverage thresholds. Every y-axis point shows the percentage of bases above the corresponding x-axis value. The blue area shows 95% of the data, while the surrounding gray lines represent the remaining 5% of the distribution. c. Average base counts by exome capture regions. d. Excluded bases from coverage calculation based on Picard127 e. Insert size distribution. f. Samtools128 summary statistics output. g. Genotype number distribution according to RTGtools129. h. Summary output from RTGtools. Ti/Tv is the ratio of transition (Ti) to transversion (Tv). This included SNPs from off-target regions. i. The distribution of variant counts across all samples after cohort re-genotyping.

2

Figure S2. Phenotype validation and prevalence, related to Figure 1 and the STAR Methods. (A–J) Comparison of vital signs and laboratory test distributions between case and control groups defined by phecodes, with p values indicating significance from Wilcoxon tests. (K and L) Overlap between cancer cases defined from EHR cancer diagnosis data (“Cancer Stage Fact”) and phecode-based case-control groups for breast (K) and prostate cancer (L). (M) Prevalence of frequent phecodes. The red bars indicate the prevalence of the most frequent phecodes in subset of the UCLA ATLAS population who had an encounter between 1–2 years after their initial encounter, at the time of sample collection. The blue bars represent the prevalence of phecodes in this same subset at their encounters 1–2 years after their initial encounter. (N) The change in phecode prevalence over one year post-collection in the subset of the UCLA ATLAS population who had an encounter between 1–2 years after their initial encounter. The triangles represent the number of new patients, and the blue bars represent the percentage of patients. (O and P) The difference in (O) phecode group and (P) phecode prevalence between participants in the UCLA biobank (purple; “UCLA ATLAS”) at their time of collection compared with all other UCLA Health patients with EHR data (green; “Rest of Data Discovery Repository [DDR]”) within one year of the ATLAS launch date. Logistic regression tests were used to compare the two groups adjusted for sex and age (ORs with 95% confidence intervals, shown in the right panels). After Bonferroni multiple testing corrections, ** denotes a significant difference between the two populations.

3

Figure S3. Broad-scale genetic ancestry and health characteristics, related to Figure 1. a. Agreement between self-reported race and genetic ancestry predictions. b. The Elixhauser comorbidity index varies across genetic ancestries after applying inverse probability weighting. ANOVA on a linear regression model tested the overall effect of the categorical ancestry predictor, yielding its P-value. Adjusted means and 95% CI are presented per ancestry group. c-d. The relationship between total and hospital encounters and the comorbidity index. r, Pearson correlation, P, the correlation P-value. e. Variation in laboratory or vital sign measurements across broad-scale ancestries. ANCOVA adjusted for genetic sex and age was performed separately for each tested phenotype.

4

Figure S4. Fine-scale genetic ancestry supplementary results, related to Figure 2. a-b. Quality control of identity-by-descent (IBD) segments. Distribution of IBD segment length per chromosome, before (A) and after (B) removing human leukocyte antigens (HLA), centromere, and IBD depth outliers. After quality control, IBD segments display exponential decay of segment length as expected (most noticeably for chromosomes 6, 15 and 22). The numbers at the top of each plot represent the chromosome number. c. The distribution of genetic ancestry between fine- and broad-scale populations. IBD clusters were sorted by the predominant broad-scale ancestry of participants in each cluster, rather than cluster size as shown in Figure 2a, providing an alternative visualization. D. Cardio-metabolic disease risk for each fine-scale group within the same broad-scale ancestry. Representative cardio-metabolic phecodes were selected, and only populations with at least 100 participants were tested. In cases of small sample sizes, ‘–’ was used instead of numeric values to protect patient privacy. Filled points represent significant results (FDR ≤ 0.05). Firth’s bias-reduced logistic regression adjusted for BMI, sex, and age was used to obtain OR. Among Asian clusters, Filipino individuals were at high risk for all tested medical conditions (essential hypertension: ORIBD-09 =1.6 [1.4, 1.9], FDR = 8.6 × 10−11; type 2 diabetes: ORIBD-09 = 1.5 [1.3, 1.7], FDR = 7.4 × 10−6; coronary atherosclerosis: ORIBD-09 = 1.4 [1.1, 1.7], FDR = 5.3 × 10−3; abdominal aortic aneurysm: ORIBD-09 = 2.8. [1.4, 5.3], FDR = 7.1 × 10−3; hyperlipidemia: ORIBD-09 = 1.2 [1.0, 1.4], FDR = 2.4 × 10−2). The largest Armenian cluster and Jewish and non-Jewish Iranian clusters showed a high risk for type 2 diabetes (Armenian 1: ORIBD-16 = 2.0 [1.5, 2.7], FDR = 2.5 × 10−6; Iranian Jewish: ORIBD-11 = 2.4 [1.9, 2.9], FDR = 6.0 × 10−16; Iranian: ORIBD-17 = 2.2 [1.6, 2.9], FDR = 1.6 × 10−6), hyperlipidemia (Armenian 1: ORIBD-16 = 1.6 [1.2-2.0], FDR = 8.6 × 10−4; Iranian Jewish: ORIBD-11 = 1.5 [1.2, 1.8], FDR = .1.0 × 10−4; Iranian: ORIBD-17 = 1.8 [1.4, 2.3], FDR = 2.9 × 10−5) and coronary atherosclerosis (Armenian 1: ORIBD-16 = 1.9 [1.4, 2.5], FDR = 8.8 × 10−5; Iranian Jewish: ORIBD-11 = 1.8 [1.5, 2.2], FDR = 4.2 × 10−8; Iranian: ORIBD-17 = 1.9 [1.4, 2.5], FDR = 2.9 × 10−5).

5

Figure S5. polygenic score association results, related to Figure 3 and the STAR Methods. a-e Selected associations between polygenic scores (PGS) and diseases in European (EUR) individuals. f-i The odds ratio (OR) and prevalence of cases for the top and bottom PGS deciles across non-EUR ancestries. Logistic regression was used to obtain the OR and p values, which were adjusted using FDR.

6

Figure S6. PhWAS supplementary results, related to Figure 4.

a-b. Distribution of pruned variant-trait associations. The number of genome-wide significant pruned gene-trait associations shared across fine-scale ancestries, using a prioritized gene within 10kb of the pruned variant and matching effect direction. b presents the same associations described in a, but with a more permissive threshold for gene-trait support from less powered ancestral clusters. c. APOE haplotype frequency across fine-scale cohorts. d. Allele frequency for known risk variants across fine-scale cohorts; shading indicates the level of over- (red) or under- (blue) enrichment of a haplotype/allele in a given cohort via Fisher’s exact test. In c-d, Fisher’s exact test was utilized to compare allele frequency. Bold border indicates a significant result (FDR ≤ 0.1). e. Impact of the non-alcoholic cirrhosis risk variant rs738409-G on cirrhosis and clinical sequalae across finescale cohorts. Odds ratios and 95% CI were calculated using logistic regression. f-h. Replicated low-MAF PheWAS associations between rs115750084-G and major depressive disorder (f), rs77742325-G and osteoporosis with no other symptoms (g) and rs202215133-A and migraine (h).

6

Figure S7. Known or putative rare pathogenic variants, related to Figure 5 Figure 5 and the STAR Methods.

a. Frequency differences in Familial Mediterranean Fever (FMF) known risk alleles. b. The risk of carrying the HBB:p.E7V variant. c. The risk of carrying loss-of-function variants in PCSK9. In a-c, the error bars show the 95% Wilson score confidence intervals. d-e. The frequency of BRCA Ashkenazi Jewish founder alleles across broad-scale (d) and fine-scale ancestry groups (e). f. Differences across populations in the total numbers of rare ClinGene P/LP variants in American College of Medical Genetics (ACMG) genes, excluding Ashkenazi Jewish, as a sensitivity analysis to one presented in Figure 5b. ‘Ref’ is the total number of reference alleles, and ‘Alt’ of P/LP ClinGen. Fisher’s exact tests were used to produce odds ratios. Filled points represent significant results at the level of nominal P-value (≤ 0.05). Only fine-scale ancestry clusters with more than 400 participants were included. g-h. Distributions of missense variant pathogenicity rank scores for nine computational tools. Each subplot compares the distribution of all missense variants (purple to green) to known pathogenic variants (red), defined as either: g. ClinGen curated missense variants. h. ClinVar pathogenic/likely pathogenic (P/LP) missense variants. The black dashed line indicates the median rankscore across all missense variants for the given tool. The red dashed line denotes the median rankscore among pathogenic variants in the corresponding dataset. The purple dashed line marks the likelihood-based intersection cutoff derived from the point at which the pathogenic and background distributions cross. i. The number of rare computationally predicted damaging missense and LOF alleles per individual across ancestries. Mann-Whitney U test with a Bonferroni correction was applied to test the difference in the distribution of rare LOF and predicted damaging missense counts per individual between any broad- or fine-scale group compared to all others. Statistically significant differences are indicated by an asterisk (*). Only fine-scale ancestry clusters with more than 100 participants were included. j. Ancestry-specific carrier counts for two known GBA1 rare LOF variants (left) and their impact on Parkinson’s disease risk (right) via logistic regression.

13

Figure S8. GLP1-RAs ATLAS users and semaglutide investigation complementary data, related to Figure 6. a. Age by sex of GLP-1 receptor agonist (GLP1-RAs) users. b. Prescription numbers for GLP1-RAs by simple generic names. c. Weight loss patterns across time. Presented are smoothed longitudinal data using a functional boxplot approach, with median values and pointwise intervals between the 20th and 80th quantiles. Vertical tick marks along the x-axis show the deciles of the data distribution, indicating where most data points are concentrated across weeks. d. The relationship between bins of polygenic scores (PGS) for body mass index (BMI) (left) and type 2 diabetes mellitus (right), and weight loss in response to semaglutide. Shown are Loess smoothed plots with 95% CI based on the maximum weight loss for patients across weeks. p values and effect sizes were obtained from a linear regression model with covariates, and a Bonferroni correction was applied. e. The relationship between type 2 diabetes mellitus PGS and weight loss in response to semaglutide in EUR AoU participants. Scaled PGS were divided into groups and linear mixed model fitted values were plotted with 95% CI, based on longitudinal data with repeated weight measurements. P-values were obtained from a linear mixed-effects model that included covariates. f. GWAS Q-Q plot for weight loss on semaglutide displays no significant findings. g. Gene-level test Q-Q plot for weight-loss on semaglutide (genomic inflation factor: 0.95). h. Differences in carrier numbers of semaglutide efficacy involved alleles in PTPRU across ancestries using Firth’s bias-reduced logistic regression. Only variants that exhibited at least one significant difference between EUR and another ancestry are shown. Filled dots represent significant odds ratios (FDR ≤0.05).

7

Table S1. Summary of lab test results and prescription numbers for the top 50 prescribed medications in ATLAS, related to the STAR Methods.

8

Table S2. Phecodes and broad-scale ancestry associations, related to Figure 1 and the STAR Methods.

9

Table S3. Fine-scale ancestry cluster sample sizes compared with published cohorts and phecode associations, related to Figure 2 and the STAR Methods.

10

Table S4. PheWAS results and pharmacogenomic variants, related to Figure 4.

11

Table S5. Summary of WES variant counts by category and differences across broad-scale ancestries, related to Figure 5 and the STAR Methods.

12

Table S6. ExWAS results, related to Figure 5.

KEY RESOURCES TABLE

REAGENT or RESOURCE SOURCE IDENTIFIER
Deposited data
PheWAS summary statistics This paper atlas-phewas.mednet.ucla.edu
Ancestry disease risk summary statistics This paper atlas-phewas.mednet.ucla.edu
Taiwan BioBank summary statistics Feng et al.7 https://taiwanview.twbiobank.org.tw/pheweb.php
AstraZeneca summary statistics Lei et al.69 http://azphewas.com/
1000 Genomes Project 1000 Genomes Project38 https://www.internationalgenome.org/data
Simons Genome Diversity Project Simons Genome Diversity Project130 https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/vcf_variants/
Human Genome Diversity Project Human Genome Diversity Project131 https://ngs.sanger.ac.uk/production/hgdp/
Software and algorithms
SAMtools v1.2. Danecek et al.128 http://samtools.sourceforge.net/
Emmeans Lenth et al.132 https://rvlenth.github.io/emmeans/
BCFtools Dancek, et al.128 https://samtools.github.io/bcftools/bcftools.html
PLINK Purcell, et al.133 https://www.cog-genomics.org/plink/
SHAPEIT5 Hofmeister, et al.134 https://odelaneau.github.io/shapeit5/
iLASH Shemirani, et al.135 https://github.com/roohy/iLASH
NetworkIt Staudt, et al. https://networkit.github.io/get_started.html
PLINK v2.0a Chang, et al.136 https://www.cog-genomics.org/plink/2.0/
Regeinie v4.0 Mbatchou, et al.66 https://rgcgithub.github.io/regenie/
VEP v112 McLaren, et al.137 https://www.ensembl.org/info/docs/tools/vep/
LOFTEE Karczewski et al. 94 https://github.com/konradjk/loftee
dbNSFP v4.9a Liu et al.138 https://www.dbnsfp.org/
BWA-MEM v0.7.17-r1188 Li et al.139 https://bio-bwa.sourceforge.net/
Picard v2.21.2 NA https://github.com/broadinstitute/picard 127
DeepVariant v0.10.0 Poplin et al.140 https://github.com/google/deepvariant/
VerifyBamID v1.1.3 Jun et al.142 https://genome.sph.umich.edu/wiki/VerifyBamID
GLnexus v1.4.3 Yun et al.141 https://github.com/dnanexus-rnd/GLnexus
MultiQC v1.27.1 Ewels et al.143 https://github.com/MultiQC/MultiQC
RTGtools v3.12.1 Cleary et al.129 https://github.com/RealTimeGenomics/rtg-tools
VCFtools v0.1.16 Danecek et al. https://vcftools.github.io/index.html
pandas v2.2.2 https://pandas.pydata.org/docs/reference/index.html
ClinPred Alirezaie et al.144 https://sites.google.com/site/clinpred/
MetaRNN Li et al.145 http://www.liulab.science/metarnn.html
BayesDel_addAF Feng et al.146 https://fenglab.chpc.utah.edu/BayesDel/BayesDel.html
VARITY_R Wu et al.147 http://varity.varianteffect.org/
REVEL Loannidis et al.148 https://sites.google.com/site/revelgenomics/
AlphaMissense Cheng et al.149 https://alphamissense.hegelab.org/
MutPred2 Pejaver et al.150 https://mutpred.mutdb.org/
VEST4 Carter et al. 151 https://www.cravat.us/CRAVAT/
ESM-1b Rives et al.152 https://github.com/facebookresearch/esm
bedtools v2.29.1 Quinlan et al.153 https://bedtools.readthedocs.io/en/latest/
comorbidity v1.0.7 Gasparini154 https://github.com/ellessenne/comorbidity
bigsnpr Verma et al.155 https://privefl.github.io/bigsnpr/reference/bigsnpr-package.html
Boutros Plotting General v.7.1.0 P’ng et al.156 https://github.com/uclahs-cds/package-BoutrosLab-plotting-general
logistf v1.26.0 Puhr et al.157 https://github.com/georgheinze/logistf
pgsc_calc Lambert et al. 158 https://pgsc-calc.readthedocs.io/en/latest/
fdapace v.0.6.0 Yao et al. & Liu et al159,160 https://github.com/cran/fdapace
lmerTest v.3.1.3 Kuznetsova et al.161 https://github.com/runehaubo/lmerTestR
visreg v.2.7.0 Breheny et al.162 https://pbreheny.github.io/visreg/
fANCOVA v.0.6.1 Wang163 https://www.rdocumentation.org/packages/userfriendlyscience/versions/0.7.2/topics/fanova
svyglm v4.4.8 Lumley164 https://cran.r-project.org/web/packages/survey/survey.pdf
Python Hail v0.2.134 Hail Team165 https://pypi.org/project/hail/
poolr R package v1.2.0 Cinar at al. 166 https://github.com/ozancinar/poolr
SAIGE Zhou et al.167 https://saigegit.github.io/SAIGE-doc/
METAL Willer et al.168 https://github.com/statgen/METAL
  • The UCLA ATLAS links genetic data to electronic health records in a diverse biobank

  • Broad- and fine-scale ancestral mapping reveals disease risk across 92,164 participants

  • Ancestry-specific phenome-wide analyses of common and rare variants identify risk loci

  • Longitudinal electronic health records identify genetic effects on semaglutide response

Acknowledgments

We are deeply grateful to the participants of the UCLA Health Biobank, to the DGSOM and UC Health for their support and development of ATLAS, and to Allen and Charlotte Ginsburg for their support of the UCLA Institute for Precision Health, which oversees ATLAS. WES data for this study were generated as part of the partnership of UCLA Health with Regeneron Genetics Center (RGC). R.H. was supported by EMBO Postdoctoral Fellowship ALTF 1131-2021 and the Prostate Cancer Foundation Young Investigator Award 22YOUN32. M.P.M. was supported by the UCLA-Caltech Medical Scientist Training Program (T32-GM008042) and by the National Institute of Mental Health (F30-MH135712). A.W. was supported by the National Human Genome Research Institute (F31-HG013462). J.F. was supported by the National Institute of Biomedical Imaging and Bioengineering Medical Imaging Informatics Training Grant (T32-EB016640). D.H.G., P.C.B., A.A.T.B. and C.L. were supported by UCLA CTSI. T.S.C. was supported by NIH grants K08AG065519-01A1, R01AG085518-01A1, U54NS123746 and California Department of Public Health, Chronic Disease Control Branch, Alzheimer’s Disease Program, under Contract #22-10079 and #24-10127. N.Zeltser. was supported by the National Human Genome Research Institute (T32-HG002536) and by the National Cancer Institute (F31-CA281168). V.A.A. and N.Zaitlen. are supported by the National Human Genome Research Insitute-R01HG011345.

Declaration of Interests

E.E.K. has received personal fees from Regeneron Pharmaceuticals, 23&Me, Allelica, and Illumina; has received research funding from Allelica; and serves on the advisory boards for Encompass Biosciences, Overtone, and Galateo Bio. P.T.S. is a consultant for 10X Genomics, Illumina, Foresight Diagnostics, Natera, and Twinstrand. P.C.B. sits on the Scientific Advisory Board of Intersect Diagnostics Inc. and previously sat on those of Sage Bionetworks and BioSymetrics Inc. All other authors declare no conflict of interest.

Footnotes

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

Data and code availability
  • The individual-level data reported in this study cannot be deposited in a public repository due to privacy regulations/legal restrictions within the Health System and licensing terms. To request access through collaboration, please contact the lead contact. In addition, PheWAS summary statistics are publicly available at our web portal (atlas-phewas.mednet.ucla.edu.
  • This paper does not report original code.
  • Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request

References

  • 1.Zhou W, Kanai M, Wu K-HH, Rasheed H, Tsuo K, Hirbo JB, Wang Y, Bhattacharya A, Zhao H, Namba S, et al. (2022). Global Biobank Meta-analysis Initiative: Powering genetic discovery across human disease. Cell Genomics 2, 100192. 10.1016/j.xgen.2022.100192. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Chambers DA, Feero WG, and Khoury MJ (2016). Convergence of Implementation Science, Precision Medicine, and the Learning Health Care System: A New Model for Biomedical Research. JAMA 315, 1941. 10.1001/jama.2016.3867. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Bycroft C, Freeman C, Petkova D, Band G, Elliott LT, Sharp K, Motyer A, Vukcevic D, Delaneau O, O’Connell J, et al. (2018). The UK Biobank resource with deep phenotyping and genomic data. Nature 562, 203–209. 10.1038/s41586-018-0579-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Halldorsson BV, Eggertsson HP, Moore KHS, Hauswedell H, Eiriksson O, Ulfarsson MO, Palsson G, Hardarson MT, Oddsson A, Jensson BO, et al. (2022). The sequences of 150,119 genomes in the UK Biobank. Nature 607, 732–740. 10.1038/s41586-022-04965-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.The All of Us Research Program Genomics Investigators, Manuscript Writing Group, Bick AG, Metcalf GA, Mayo KR, Lichtenstein L, Rura S, Carroll RJ, Musick A, Linder JE, et al. (2024). Genomic data in the All of Us Research Program. Nature 627, 340–346. 10.1038/s41586-023-06957-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Kurki MI, Karjalainen J, Palta P, Sipilä TP, Kristiansson K, Donner KM, Reeve MP, Laivuori H, Aavikko M, Kaunisto MA, et al. (2023). FinnGen provides genetic insights from a well-phenotyped isolated population. Nature 613, 508–518. 10.1038/s41586-022-05473-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Feng Y-CA, Chen C-Y, Chen T-T, Kuo P-H, Hsu Y-H, Yang H-I, Chen WJ, Su M-W, Chu H-W, Shen C-Y, et al. (2022). Taiwan Biobank: A rich biomedical research database of the Taiwanese population. Cell Genomics 2, 100197. 10.1016/j.xgen.2022.100197. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Johnson JS, Cote AC, Dobbyn A, Sloofman LG, Xu J, Cotter L, Charney AW, Eating Disorders Working Group of the Psychiatric Genomics Consortium, Birgegård A, Jordan J, et al. (2023). Mapping anorexia nervosa genes to clinical phenotypes. Psychol. Med 53, 2619–2633. 10.1017/S0033291721004554. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Roden D, Pulley J, Basford M, Bernard G, Clayton E, Balser J, and Masys D (2008). Development of a Large-Scale De-Identified DNA Biobank to Enable Personalized Medicine. Clin. Pharmacol. Ther 84, 362–369. 10.1038/clpt.2008.89. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Carey DJ, Fetterolf SN, Davis FD, Faucett WA, Kirchner HL, Mirshahi U, Murray MF, Smelser DT, Gerhard GS, and Ledbetter DH (2016). The Geisinger MyCode community health initiative: an electronic health record–linked biobank for precision medicine research. Genet. Med 18, 906–913. 10.1038/gim.2015.187. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Zawistowski M, Fritsche LG, Pandit A, Vanderwerff B, Patil S, Schmidt EM, VandeHaar P, Willer CJ, Brummett CM, Kheterpal S, et al. (2023). The Michigan Genomics Initiative: A biobank linking genotypes and electronic clinical records in Michigan Medicine patients. Cell Genomics 3, 100257. 10.1016/j.xgen.2023.100257. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.the eMERGE Team, McCarty CA, Chisholm RL, Chute CG, Kullo IJ, Jarvik GP, Larson EB, Li R, Masys DR, Ritchie MD, et al. (2011). The eMERGE Network: A consortium of biorepositories linked to electronic medical records data for conducting genomic studies. BMC Med. Genomics 4, 13. 10.1186/1755-8794-4-13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.COVID-19 Host Genetics Initiative (2023). A second update on mapping the human genetic architecture of COVID-19. Nature 621, E7–E26. 10.1038/s41586-023-06355-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Fatumo S, Chikowore T, Choudhury A, Ayub M, Martin AR, and Kuchenbaecker K (2022). A roadmap to increase diversity in genomic studies. Nat. Med 28, 243–250. 10.1038/s41591-021-01672-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Sirugo G, Williams SM, and Tishkoff SA (2019). The Missing Diversity in Human Genetic Studies. Cell 177, 26–31. 10.1016/j.cell.2019.02.048. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Kullo IJ, Conomos MP, Nelson SC, Adebamowo SN, Choudhury A, Conti D, Fullerton SM, Gogarten SM, Heavner B, Hornsby WE, et al. (2024). The PRIMED Consortium: Reducing disparities in polygenic risk assessment. Am. J. Hum. Genet 111, 2594–2606. 10.1016/j.ajhg.2024.10.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Ding Y, Hou K, Xu Z, Pimplaskar A, Petter E, Boulier K, Privé F, Vilhjálmsson BJ, Olde Loohuis LM, and Pasaniuc B (2023). Polygenic scoring accuracy varies across the genetic ancestry continuum. Nature 618, 774–781. 10.1038/s41586-023-06079-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Tsuo K, Shi Z, Ge T, Mandla R, Hou K, Ding Y, Pasaniuc B, Wang Y, and Martin AR (2024). All of Us diversity and scale improve polygenic prediction contextually with greatest improvements for under-represented populations. BioRxiv Prepr. Serv. Biol, 2024.08.06.606846. 10.1101/2024.08.06.606846. [DOI] [Google Scholar]
  • 19.Naslavsky MS, Suemoto CK, Brito LA, Scliar MO, Ferretti-Rebustini RE, Rodriguez RD, Leite REP, Araujo NM, Borda V, Tarazona-Santos E, et al. (2022). Global and local ancestry modulate APOE association with Alzheimer’s neuropathology and cognitive outcomes in an admixed sample. Mol. Psychiatry 27, 4800–4808. 10.1038/s41380-022-01729-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Cohen J, Pertsemlidis A, Kotowski IK, Graham R, Garcia CK, and Hobbs HH (2005). Low LDL cholesterol in individuals of African descent resulting from frequent nonsense mutations in PCSK9. Nat. Genet 37, 161–165. 10.1038/ng1509. [DOI] [PubMed] [Google Scholar]
  • 21.Jurgens SJ, Wang X, Choi SH, Weng L-C, Koyama S, Pirruccello JP, Nguyen T, Smadbeck P, Jang D, Chaffin M, et al. (2024). Rare coding variant analysis for human diseases across biobanks and ancestries. Nat. Genet 56, 1811–1820. 10.1038/s41588-024-01894-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Manrai AK, Funke BH, Rehm HL, Olesen MS, Maron BA, Szolovits P, Margulies DM, Loscalzo J, and Kohane IS (2016). Genetic Misdiagnoses and the Potential for Health Disparities. N. Engl. J. Med 375, 655–665. 10.1056/NEJMsa1507092. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Venner E, Patterson K, Kalra D, Wheeler MM, Chen Y-J, Kalla SE, Yuan B, Karnes JH, Walker K, Smith JD, et al. (2024). The frequency of pathogenic variation in the All of Us cohort reveals ancestry-driven disparities. Commun. Biol 7, 174. 10.1038/s42003-023-05708-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Graham SE, Clarke SL, Wu K-HH, Kanoni S, Zajac GJM, Ramdas S, Surakka I, Ntalla I, Vedantam S, Winkler TW, et al. (2021). The power of genetic diversity in genome-wide association studies of lipids. Nature 600, 675–679. 10.1038/s41586-021-04064-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Wojcik GL, Graff M, Nishimura KK, Tao R, Haessler J, Gignoux CR, Highland HM, Patel YM, Sorokin EP, Avery CL, et al. (2019). Genetic analyses of diverse populations improves discovery for complex traits. Nature 570, 514–518. 10.1038/s41586-019-1310-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Abul-Husn NS, and Kenny EE (2019). Personalized Medicine and the Power of Electronic Health Records. Cell 177, 58–69. 10.1016/j.cell.2019.02.039. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Hindorff LA, Bonham VL, Brody LC, Ginoza MEC, Hutter CM, Manolio TA, and Green ED (2018). Prioritizing diversity in human genomics research. Nat. Rev. Genet 19, 175–185. 10.1038/nrg.2017.89. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Lajonchere C, Naeim A, Dry S, Wenger N, Elashoff D, Vangala S, Petruse A, Ariannejad M, Magyar C, Johansen L, et al. (2021). An Integrated, Scalable, Electronic Video Consent Process to Power Precision Health Research: Large, Population-Based, Cohort Implementation and Scalability Study. J. Med. Internet Res 23, e31121. 10.2196/31121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Johnson R, Ding Y, Venkateswaran V, Bhattacharya A, Boulier K, Chiu A, Knyazev S, Schwarz T, Freund M, Zhan L, et al. (2022). Leveraging genomic diversity for discovery in an electronic health record linked biobank: the UCLA ATLAS Community Health Initiative. Genome Med. 14, 104. 10.1186/s13073-022-01106-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Johnson R, Ding Y, Bhattacharya A, Knyazev S, Chiu A, Lajonchere C, Geschwind DH, and Pasaniuc B (2023). The UCLA ATLAS Community Health Initiative: Promoting precision health research in a diverse biobank. Cell Genomics 3, 100243. 10.1016/j.xgen.2022.100243. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Feigin VL, Vos T, Nichols E, Owolabi MO, Carroll WM, Dichgans M, Deuschl G, Parmar P, Brainin M, and Murray C (2020). The global burden of neurological disorders: translating evidence into policy. Lancet Neurol. 19, 255–265. 10.1016/S1474-4422(19)30411-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Murray CJL (2024). Findings from the Global Burden of Disease Study 2021. The Lancet 403, 2259–2262. 10.1016/S0140-6736(24)00769-4. [DOI] [PubMed] [Google Scholar]
  • 33.Ferrari AJ, Santomauro DF, Aali A, Abate YH, Abbafati C, Abbastabar H, Abd ElHafeez S, Abdelmasseh M, Abd-Elsalam S, Abdollahi A, et al. (2024). Global incidence, prevalence, years lived with disability (YLDs), disability-adjusted life-years (DALYs), and healthy life expectancy (HALE) for 371 diseases and injuries in 204 countries and territories and 811 subnational locations, 1990–2021: a systematic analysis for the Global Burden of Disease Study 2021. The Lancet 403, 2133–2161. 10.1016/S0140-6736(24)00757-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Sirugo G, Tishkoff SA, and Williams SM (2021). The quagmire of race, genetic ancestry, and health disparities. J. Clin. Invest 131, e150255. 10.1172/JCI150255. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Adigbli G (2020). Race, science and (im)precision medicine. Nat. Med 26, 1675–1676. 10.1038/s41591-020-1115-x. [DOI] [PubMed] [Google Scholar]
  • 36.Krainc T, and Fuentes A (2022). Genetic ancestry in precision medicine is reshaping the race debate. Proc. Natl. Acad. Sci 119, e2203033119. 10.1073/pnas.2203033119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Borrell LN, Elhawary JR, Fuentes-Afflick E, Witonsky J, Bhakta N, Wu AHB, Bibbins-Domingo K, Rodríguez-Santana JR, Lenoir MA, Gavin JR, et al. (2021). Race and Genetic Ancestry in Medicine — A Time for Reckoning with Racism. N. Engl. J. Med 384, 474–480. 10.1056/NEJMms2029562. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.The 1000 Genomes Project Consortium, Corresponding authors, Auton A, Abecasis GR, Steering committee, Altshuler DM, Durbin RM, Abecasis GR, Bentley DR, Chakravarti A, et al. (2015). A global reference for human genetic variation. Nature 526, 68–74. 10.1038/nature15393. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Ong PM, Pech C, Gutierrez NR, and Mays VM (2021). COVID-19 Medical Vulnerability Indicators: A Predictive, Local Data Model for Equity in Public Health Decision Making. Int. J. Environ. Res. Public. Health 18, 4829. 10.3390/ijerph18094829. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Quan H, Li B, Couris CM, Fushimi K, Graham P, Hider P, Januel J-M, and Sundararajan V (2011). Updating and validating the Charlson comorbidity index and score for risk adjustment in hospital discharge abstracts using data from 6 countries. Am. J. Epidemiol 173, 676–682. 10.1093/aje/kwq433. [DOI] [PubMed] [Google Scholar]
  • 41.Quan H, Sundararajan V, Halfon P, Fong A, Burnand B, Luthi J-C, Saunders LD, Beck CA, Feasby TE, and Ghali WA (2005). Coding algorithms for defining comorbidities in ICD-9-CM and ICD-10 administrative data. Med. Care 43, 1130–1139. 10.1097/01.mlr.0000182534.19832.83. [DOI] [PubMed] [Google Scholar]
  • 42.Kind AJH, and Buckingham WR (2018). Making Neighborhood-Disadvantage Metrics Accessible — The Neighborhood Atlas. N. Engl. J. Med 378, 2456–2458. 10.1056/NEJMp1802313. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Banerjee PN, Filippi D, and Allen Hauser W (2009). The descriptive epidemiology of epilepsy—A review. Epilepsy Res. 85, 31–45. 10.1016/j.eplepsyres.2009.03.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Adamu A, Chen R, Li A, and Xue G (2023). Epilepsy in Asian countries. Acta Epileptol. 5, 25. 10.1186/s42494-023-00136-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Minsky S, Vega W, Miskimen T, Gara M, and Escobar J (2003). Diagnostic Patterns in Latino, African American, and European American Psychiatric Patients. Arch. Gen. Psychiatry 60, 637. 10.1001/archpsyc.60.6.637. [DOI] [PubMed] [Google Scholar]
  • 46.Hwang SHJ, Childers ME, Wang PW, Nam JY, Keller KL, Hill SJ, and Ketter TA (2010). Higher prevalence of bipolar I disorder among Asian and Latino compared to Caucasian patients receiving treatment. Asia-Pac. Psychiatry 2, 156–165. 10.1111/j.1758-5872.2010.00080.x. [DOI] [Google Scholar]
  • 47.Mirrakhimov AE, Sooronbaev T, and Mirrakhimov EM (2013). Prevalence of obstructive sleep apnea in Asian adults: a systematic review of the literature. BMC Pulm. Med 13, 10. 10.1186/1471-2466-13-10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Villaneuva ATC, Buchanan PR, Yee BJ, and Grunstein RR (2005). Ethnicity and obstructive sleep apnoea. Sleep Med. Rev 9, 419–436. 10.1016/j.smrv.2005.04.005. [DOI] [PubMed] [Google Scholar]
  • 49.Redline S, Tishler PV, Hans MG, Tosteson TD, Strohl KP, and Spry K (1997). Racial differences in sleep-disordered breathing in African-Americans and Caucasians. Am. J. Respir. Crit. Care Med 155, 186–192. 10.1164/ajrccm.155.1.9001310. [DOI] [PubMed] [Google Scholar]
  • 50.Ong KC, and Clerk AA (1998). Comparison of the severity of sleep-disordered breathing in Asian and Caucasian patients seen at a sleep disorders center. Respir. Med 92, 843– 848. 10.1016/S0954-6111(98)90386-9. [DOI] [PubMed] [Google Scholar]
  • 51.Abdelmoumen I, Jimenez S, Valencia I, Melvin J, Legido A, Diaz-Diaz MM, Griffith C, Massingham LJ, Yelton M, Rodríguez-Hernández J, et al. (2021). Boricua Founder Variant in FRRS1L Causes Epileptic Encephalopathy With Hyperkinetic Movements. J. Child Neurol 36, 93–98. 10.1177/0883073820953001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Caggiano C, Boudaie A, Shemirani R, Mefford J, Petter E, Chiu A, Ercelen D, He R, Tward D, Paul KC, et al. (2023). Disease risk and healthcare utilization among ancestrally diverse groups in the Los Angeles region. Nat. Med 29, 1845–1856. 10.1038/s41591-023-02425-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Belbin GM, Cullina S, Wenric S, Soper ER, Glicksberg BS, Torre D, Moscati A, Wojcik GL, Shemirani R, Beckmann ND, et al. (2021). Toward a fine-scale population health monitoring system. Cell 184, 2068–2083.e11. 10.1016/j.cell.2021.03.034. [DOI] [PubMed] [Google Scholar]
  • 54.Dai CL, Vazifeh MM, Yeang C-H, Tachet R, Wells RS, Vilar MG, Daly MJ, Ratti C, and Martin AR (2020). Population Histories of the United States Revealed through Fine-Scale Migration and Haplotype Analysis. Am. J. Hum. Genet 106, 371–388. 10.1016/j.ajhg.2020.02.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Hovhannisyan A, Delser PM, Hakobyan A, Jones ER, Schraiber JG, Antonosyan M, Margaryan A, Xue Z, Jeon S, Bhak J, et al. (2025). Demographic history and genetic variation of the Armenian population. Am. J. Hum. Genet 112, 11–27. 10.1016/j.ajhg.2024.10.022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Haber M, Mezzavilla M, Xue Y, Comas D, Gasparini P, Zalloua P, and Tyler-Smith C (2016). Genetic evidence for an origin of the Armenians from Bronze Age mixing of Multiple populations. Eur. J. Hum. Genet. EJHG 24, 931–936. 10.1038/ejhg.2015.206. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Li D, Farrell JJ, Mez J, Martin ER, Bush WS, Ruiz A, Boada M, de Rojas I, Mayeux R, Haines JL, et al. (2023). Novel loci for Alzheimer’s disease identified by a genome-wide association study in Ashkenazi Jews. Alzheimers Dement. J. Alzheimers Assoc 19, 5550–5562. 10.1002/alz.13117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Larena M, Sanchez-Quinto F, Sjödin P, McKenna J, Ebeo C, Reyes R, Casel O, Huang J-Y, Hagada KP, Guilay D, et al. (2021). Multiple migrations to the Philippines during the last 50,000 years. Proc. Natl. Acad. Sci 118, e2026132118. 10.1073/pnas.2026132118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Wei W-Q, Bastarache LA, Carroll RJ, Marlo JE, Osterman TJ, Gamazon ER, Cox NJ, Roden DM, and Denny JC (2017). Evaluating phecodes, clinical classification software, and ICD-9-CM codes for phenome-wide association studies in the electronic health record. PloS One 12, e0175508. 10.1371/journal.pone.0175508. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Wu P, Gifford A, Meng X, Li X, Campbell H, Varley T, Zhao J, Carroll R, Bastarache L, Denny JC, et al. (2019). Mapping ICD-10 and ICD-10-CM Codes to Phecodes: Workflow Development and Initial Evaluation. JMIR Med. Inform 7, e14325. 10.2196/14325. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Denny JC, Bastarache L, Ritchie MD, Carroll RJ, Zink R, Mosley JD, Field JR, Pulley JM, Ramirez AH, Bowton E, et al. (2013). Systematic comparison of phenome- wide association study of electronic medical record data and genome-wide association study data. Nat. Biotechnol 31, 1102–1110. 10.1038/nbt.2749. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Duncan L, Shen H, Gelaye B, Meijsen J, Ressler K, Feldman M, Peterson R, and Domingue B (2019). Analysis of polygenic risk score usage and performance in diverse human populations. Nat. Commun 10, 3328. 10.1038/s41467-019-11112-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Martin AR, Kanai M, Kamatani Y, Okada Y, Neale BM, and Daly MJ (2019). Clinical use of current polygenic risk scores may exacerbate health disparities. Nat. Genet 51, 584– 591. 10.1038/s41588-019-0379-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Martin AR, Gignoux CR, Walters RK, Wojcik GL, Neale BM, Gravel S, Daly MJ, Bustamante CD, and Kenny EE (2017). Human Demographic History Impacts Genetic Risk Prediction across Diverse Populations. Am. J. Hum. Genet 100, 635–649. 10.1016/j.ajhg.2017.03.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Curtis D (2018). Polygenic risk score for schizophrenia is more strongly associated with ancestry than with schizophrenia. Psychiatr. Genet 28, 85–89. 10.1097/YPG.0000000000000206. [DOI] [PubMed] [Google Scholar]
  • 66.Mbatchou J, Barnard L, Backman J, Marcketta A, Kosmicki JA, Ziyatdinov A, Benner C, O’Dushlaine C, Barber M, Boutkov B, et al. (2021). Computationally efficient whole-genome regression for quantitative and binary traits. Nat. Genet 53, 1097–1103. 10.1038/s41588-021-00870-7. [DOI] [PubMed] [Google Scholar]
  • 67.Romeo S, Kozlitina J, Xing C, Pertsemlidis A, Cox D, Pennacchio LA, Boerwinkle E, Cohen JC, and Hobbs HH (2008). Genetic variation in PNPLA3 confers susceptibility to nonalcoholic fatty liver disease. Nat. Genet 40, 1461–1465. 10.1038/ng.257. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Cerezo M, Sollis E, Ji Y, Lewis E, Abid A, Bircan KO, Hall P, Hayhurst J, John S, Mosaku A, et al. (2025). The NHGRI-EBI GWAS Catalog: standards for reusability, sustainability and diversity. Nucleic Acids Res. 53, D998–D1005. 10.1093/nar/gkae1070. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Wang Q, Dhindsa RS, Carss K, Harper AR, Nag A, Tachmazidou I, Vitsios D, Deevi SVV, Mackay A, Muthas D, et al. (2021). Rare variant contribution to human disease in 281,104 UK Biobank exomes. Nature 597, 527–532. 10.1038/s41586-021-03855-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Viswanathan L, and Rao SS (2023). Intestinal Disaccharidase Deficiency in Adults: Evaluation and Treatment. Curr. Gastroenterol. Rep 25, 134–139. 10.1007/s11894-023-00870-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Garg A, On KF, Xiao Y, Elkayam E, Cifani P, David Y, and Joshua-Tor L (2025). The molecular basis of Human FN3K mediated phosphorylation of glycated substrates. Nat. Commun 16, 941. 10.1038/s41467-025-56207-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Verma A, Huffman JE, Rodriguez A, Conery M, Liu M, Ho Y-L, Kim Y, Heise DA, Guare L, Panickan VA, et al. (2024). Diversity and scale: Genetic architecture of 2068 traits in the VA Million Veteran Program. Science 385, eadj1182. 10.1126/science.adj1182. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Malloy C, Ahern M, Lin L, and Hoffman DA (2022). Neuronal Roles of the Multifunctional Protein Dipeptidyl Peptidase-like 6 (DPP6). Int. J. Mol. Sci 23, 9184. 10.3390/ijms23169184. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Li C, Hou Y, Ou R, Wei Q, Zhang L, Liu K, Lin J, Chen X, Song W, Zhao B, et al. (2024). GWAS Identifies DPP6 as Risk Gene of Cognitive Decline in Parkinson’s Disease. J. Gerontol. A. Biol. Sci. Med. Sci 79, glae155. 10.1093/gerona/glae155. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Maussion G, Cruceanu C, Rosenfeld JA, Bell SC, Jollant F, Szatkiewicz J, Collins RL, Hanscom C, Kolobova I, de Champfleur NM, et al. (2017). Implication of LRRC4C and DPP6 in neurodevelopmental disorders. Am. J. Med. Genet. A 173, 395–406. 10.1002/ajmg.a.38021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Queiroz-Junior CM, Santos ACPM, Galvão I, Souto GR, Mesquita RA, Sá MA, and Ferreira AJ (2019). The angiotensin converting enzyme 2/angiotensin-(1–7)/Mas Receptor axis as a key player in alveolar bone remodeling. Bone 128, 115041. 10.1016/j.bone.2019.115041. [DOI] [PubMed] [Google Scholar]
  • 77.Sabeh P, Dumas SA, Maios C, Daghar H, Korzeniowski M, Rousseau J, Lines M, Guerin A, Millichap JJ, Landsverk M, et al. (2025). Heterozygous UBR5 variants result in a neurodevelopmental syndrome with developmental delay, autism, and intellectual disability. Am. J. Hum. Genet 112, 75–86. 10.1016/j.ajhg.2024.11.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Zhu Q, Yang J, Shi L, Zhang J, Zhang P, Li J, and Song X (2025). Exploring the role of ubiquitination modifications in migraine headaches. Front. Immunol 16, 1534389. 10.3389/fimmu.2025.1534389. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Yang L, Lewkowich I, Apsley K, Fritz JM, Wills-Karp M, and Weaver TE (2015). Haploinsufficiency for Stard7 is associated with enhanced allergic responses in lung and skin. J. Immunol. Baltim. Md 1950 194, 5635–5643. 10.4049/jimmunol.1500231. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Peng D-F, Hu T-L, Soutto M, Belkhiri A, and El-Rifai W (2014). Loss of glutathione peroxidase 7 promotes TNF-α-induced NF-κB activation in Barrett’s carcinogenesis. Carcinogenesis 35, 1620–1628. 10.1093/carcin/bgu083. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Uhlén M, Fagerberg L, Hallström BM, Lindskog C, Oksvold P, Mardinoglu A, Sivertsson Å, Kampf C, Sjöstedt E, Asplund A, et al. (2015). Proteomics. Tissue-based map of the human proteome. Science 347, 1260419. 10.1126/science.1260419. [DOI] [PubMed] [Google Scholar]
  • 82.Spear ML, Diaz-Papkovich A, Ziv E, Yracheta JM, Gravel S, Torgerson DG, and Hernandez RD (2020). Recent shifts in the genomic ancestry of Mexican Americans may alter the genetic architecture of biomedical traits. eLife 9, e56029. 10.7554/eLife.56029. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Barberena-Jonas C, Medina-Muñoz SG, Cedillo-Castelán V, Sepúlveda-Morales T, Gonzaga-Jáuregui C, ENSA Genomics Consortium, Aguilar-Salinas C, Barberena-Jonas C, Canizales-Quintero S, Cruz-Hervert LP, et al. (2026). Clinical genetic variation across Hispanic populations in the Mexican Biobank. Nat. Med 10.1038/s41591-025-04100-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Backman JD, Li AH, Marcketta A, Sun D, Mbatchou J, Kessler MD, Benner C, Liu D, Locke AE, Balasubramanian S, et al. (2021). Exome sequencing and analysis of 454,787 UK Biobank participants. Nature 599, 628–634. 10.1038/s41586-021-04103-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Gomez F, Hirbo J, and Tishkoff SA (2014). Genetic Variation and Adaptation in Africa: Implications for Human Evolution and Disease. Cold Spring Harb. Perspect. Biol 6, a008524–a008524. 10.1101/cshperspect.a008524. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Relling MV, and Klein TE (2011). CPIC: Clinical Pharmacogenetics Implementation Consortium of the Pharmacogenomics Research Network. Clin. Pharmacol. Ther 89, 464– 467. 10.1038/clpt.2010.279. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Whirl-Carrillo M, Huddart R, Gong L, Sangkuhl K, Thorn CF, Whaley R, and Klein TE (2021). An Evidence-Based Framework for Evaluating Pharmacogenomics Knowledge for Personalized Medicine. Clin. Pharmacol. Ther 110, 563–572. 10.1002/cpt.2350. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Cooper-DeHoff RM, Niemi M, Ramsey LB, Luzum JA, Tarkiainen EK, Straka RJ, Gong L, Tuteja S, Wilke RA, Wadelius M, et al. (2022). The Clinical Pharmacogenetics Implementation Consortium Guideline for SLCO1B1, ABCG2, and CYP2C9 genotypes and Statin-Associated Musculoskeletal Symptoms. Clin. Pharmacol. Ther 111, 1007–1021. 10.1002/cpt.2557. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Johnson JA, Caudle KE, Gong L, Whirl-Carrillo M, Stein CM, Scott SA, Lee MT, Gage BF, Kimmel SE, Perera MA, et al. (2017). Clinical Pharmacogenetics Implementation Consortium (CPIC) Guideline for Pharmacogenetics-Guided Warfarin Dosing: 2017 Update. Clin. Pharmacol. Ther 102, 397–404. 10.1002/cpt.668. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.García-Alfonso P, Saiz-Rodríguez M, Mondéjar R, Salazar J, Páez D, Borobia AM, Safont MJ, García-García I, Colomer R, García-González X, et al. (2022). Consensus of experts from the Spanish Pharmacogenetics and Pharmacogenomics Society and the Spanish Society of Medical Oncology for the genotyping of DPYD in cancer patients who are candidates for treatment with fluoropyrimidines. Clin. Transl. Oncol. Off. Publ. Fed. Span. Oncol. Soc. Natl. Cancer Inst. Mex 24, 483–494. 10.1007/s12094-021-02708-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Rehm HL, Berg JS, Brooks LD, Bustamante CD, Evans JP, Landrum MJ, Ledbetter DH, Maglott DR, Martin CL, Nussbaum RL, et al. (2015). ClinGen — The Clinical Genome Resource. N. Engl. J. Med 372, 2235–2242. 10.1056/NEJMsr1406261. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92.Tanchoco CC, Cruz AJ, Duante CA, and Litonjua AD (2003). Prevalence of metabolic syndrome among Filipino adults aged 20 years and over. Asia Pac. J. Clin. Nutr 12, 271– 276. [PubMed] [Google Scholar]
  • 93.Miller DT, Lee K, Abul-Husn NS, Amendola LM, Brothers K, Chung WK, Gollob MH, Gordon AS, Harrison SM, Hershberger RE, et al. (2023). ACMG SF v3.2 list for reporting of secondary findings in clinical exome and genome sequencing: A policy statement of the American College of Medical Genetics and Genomics (ACMG). Genet. Med. Off. J. Am. Coll. Med. Genet 25, 100866. 10.1016/j.gim.2023.100866. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q, Collins RL, Laricchia KM, Ganna A, Birnbaum DP, et al. (2020). The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 581, 434–443. 10.1038/s41586-020-2308-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Rastogi R, Chung R, Li S, Li C, Lee K, Woo J, Kim D-W, Keum C, Babbi G, Martelli PL, et al. (2025). Critical assessment of missense variant effect predictors on disease-relevant variant data. Hum. Genet 144, 281–293. 10.1007/s00439-025-02732-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96.Erdem M, Ozgul İ, Dioken DN, Gurcuoglu I, Guntekin Ergun S, Cetin-Atalay R, Can T, and Erson-Bensan AE (2021). Identification of an mRNA isoform switch for HNRNPA1 in breast cancers. Sci. Rep 11, 24444. 10.1038/s41598-021-04007-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Cullup T, Kho AL, Dionisi-Vici C, Brandmeier B, Smith F, Urry Z, Simpson MA, Yau S, Bertini E, McClelland V, et al. (2013). Recessive mutations in EPG5 cause Vici syndrome, a multisystem disorder with defective autophagy. Nat. Genet 45, 83–87. 10.1038/ng.2497. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Zhang S, Peng X, Yang S, Li X, Huang M, Wei S, Liu J, He G, Zheng H, Yang L, et al. (2022). The regulation, function, and role of lipophagy, a form of selective autophagy, in metabolic disorders. Cell Death Dis. 13, 132. 10.1038/s41419-022-04593-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 99.Stein CS, Yancey PH, Martins I, Sigmund RD, Stokes JB, and Davidson BL (2010). Osmoregulation of ceroid neuronal lipofuscinosis type 3 in the renal medulla. Am. J. Physiol. Cell Physiol 298, C1388–1400. 10.1152/ajpcell.00272.2009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 100.Huang S, Zhu B, Cheon IS, Goplen NP, Jiang L, Zhang R, Peebles RS, Mack M, Kaplan MH, Limper AH, et al. (2019). PPAR-γ in Macrophages Limits Pulmonary Inflammation and Promotes Host Recovery following Respiratory Viral Infection. J. Virol 93, e00030–19. 10.1128/JVI.00030-19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101.Szot JO, Cuny H, Martin EM, Sheng DZ, Iyer K, Portelli S, Nguyen V, Gereis JM, Alankarage D, Chitayat D, et al. (2024). A metabolic signature for NADSYN1- dependent congenital NAD deficiency disorder. J. Clin. Invest 134, e174824. 10.1172/JCI174824. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 102.Schuurs-Hoeijmakers JHM, Geraghty MT, Kamsteeg E-J, Ben-Salem S, de Bot ST, Nijhof B, van de Vondervoort IIGM, van der Graaf M, Nobau AC, Otte-Höller I, et al. (2012). Mutations in DDHD2, encoding an intracellular phospholipase A(1), cause a recessive form of complex hereditary spastic paraplegia. Am. J. Hum. Genet 91, 1073– 1081. 10.1016/j.ajhg.2012.10.017. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 103.Liao X, Kennel PJ, Liu B, Nash TR, Zhuang RZ, Godier-Furnemont AF, Xue C, Lu R, Colombo PC, Uriel N, et al. (2023). Effect of mechanical unloading on genomewide DNA methylation profile of the failing human heart. JCI Insight 8, e161788. 10.1172/jci.insight.161788. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 104.Lei J, Jiang X, Huang D, Jing Y, Yang S, Geng L, Yan Y, Zheng F, Cheng F, Zhang W, et al. (2024). Human ESC-derived vascular cells promote vascular regeneration in a HIF-1α dependent manner. Protei Cell 15, 36–51. 10.1093/procel/pwad027. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 106.Fernandez-Gonzalez A, Kourembanas S, Wyatt TA, and Mitsialis SA (2009). Mutation of murine adenylate kinase 7 underlies a primary ciliary dyskinesia phenotype. Am. J. Respir. Cell Mol. Biol 40, 305–313. 10.1165/rcmb.2008-0102OC. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107.Toomer KA, Yu M, Fulmer D, Guo L, Moore KS, Moore R, Drayton KD, Glover J, Peterson N, Ramos-Ortiz S, et al. (2019). Primary cilia defects causing mitral valve prolapse. Sci. Transl. Med 11, eaax0290. 10.1126/scitranslmed.aax0290. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108.Deng H, Gong X, Ji G, Li C, and Cheng S (2023). KIF2C promotes clear cell renal cell carcinoma progression via activating JAK2/STAT3 signaling pathway. Mol. Cell. Probes 72, 101938. 10.1016/j.mcp.2023.101938. [DOI] [PubMed] [Google Scholar]
  • 109.Lin F, Hiesberger T, Cordes K, Sinclair AM, Goldstein LSB, Somlo S, and Igarashi P (2003). Kidney-specific inactivation of the KIF3A subunit of kinesin-II inhibits renal ciliogenesis and produces polycystic kidney disease. Proc. Natl. Acad. Sci. U. S. A 100, 5286–5291. 10.1073/pnas.0836980100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 110.den Heijer JM, Cullen VC, Quadri M, Schmitz A, Hilt DC, Lansbury P, Berendse HW, van de Berg WDJ, de Bie RMA, Boertien JM, et al. (2020). A Large-Scale Full GBA1 Gene Screening in Parkinson’s Disease in the Netherlands. Mov. Disord. Off. J. Mov. Disord. Soc 35, 1667–1674. 10.1002/mds.28112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111.Wilding JPH, Batterham RL, Calanna S, Davies M, Van Gaal LF, Lingvay I, McGowan BM, Rosenstock J, Tran MTD, Wadden TA, et al. (2021). Once-Weekly Semaglutide in Adults with Overweight or Obesity. N. Engl. J. Med 384, 989–1002. 10.1056/NEJMoa2032183. [DOI] [PubMed] [Google Scholar]
  • 112.Ryan DH, Lingvay I, Deanfield J, Kahn SE, Barros E, Burguera B, Colhoun HM, Cercato C, Dicker D, Horn DB, et al. (2024). Long-term weight loss effects of semaglutide in obesity without diabetes in the SELECT trial. Nat. Med 30, 2049–2057. 10.1038/s41591-024-02996-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 113.Garvey WT, Batterham RL, Bhatta M, Buscemi S, Christensen LN, Frias JP, Jódar E, Kandler K, Rigas G, Wadden TA, et al. (2022). Two-year effects of semaglutide in adults with overweight or obesity: the STEP 5 trial. Nat. Med 28, 2083–2091. 10.1038/s41591-022-02026-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 114.German J, Cordioli M, Tozzo V, Urbut S, Arumäe K, Smit RAJ, Lee J, Li JH, Janucik A, Ding Y, et al. (2025). Association between plausible genetic factors and weight loss from GLP1-RA and bariatric surgery. Nat. Med 10.1038/s41591-025-03645-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 115.Maretty L, Gill D, Simonsen L, Soh K, Zagkos L, Galanakis M, Sibbesen J, Iglesias MT, Secher A, Valkenborg D, et al. (2025). Proteomic changes upon treatment with semaglutide in individuals with obesity. Nat. Med 10.1038/s41591-024-03355-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 116.Kannel WB (1979). Diabetes and Cardiovascular Disease: The Framingham Study. JAMA 241, 2035. 10.1001/jama.1979.03290450033020. [DOI] [PubMed] [Google Scholar]
  • 117.Tsao CW, and Vasan RS (2015). The Framingham Heart Study: past, present and future. Int. J. Epidemiol 44, 1763–1766. 10.1093/ije/dyv336. [DOI] [PubMed] [Google Scholar]
  • 118.Wiley LK, Shortt JA, Roberts ER, Lowery J, Kudron E, Lin M, Mayer D, Wilson M, Brunetti TM, Chavan S, et al. (2024). Building a vertically integrated genomic learning health system: The biobank at the Colorado Center for Personalized Medicine. Am. J. Hum. Genet 111, 11–23. 10.1016/j.ajhg.2023.12.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 119.Verma A, Damrauer SM, Naseer N, Weaver J, Kripke CM, Guare L, Sirugo G, Kember RL, Drivas TG, Dudek SM, et al. (2022). The Penn Medicine BioBank: Towards a Genomics-Enabled Learning Healthcare System to Accelerate Precision Medicine in a Diverse Population. J. Pers. Med 12, 1974. 10.3390/jpm12121974. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 120.German J, Cordioli M, Tozzo V, Urbut S, Arumäe K, Smit RAJ, Lee J, Li JH, Janucik A, Ding Y, et al. (2024). Association between plausible genetic factors and weight loss from GLP1-RA and bariatric surgery: a multi-ancestry study in 10 960 individuals from 9 biobanks. Preprint at Genetic and Genomic Medicine, 10.1101/2024.09.11.24313458 https://doi.org/10.1101/2024.09.11.24313458. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 121.Palma-Martínez MJ, Posadas-García YS, López-Ángeles BE, Quiroz-López C, Lewis ACF, Bird KA, Lasisi T, Zaidi AA, and Sohail M (2024). The multi-scale complexity of human genetic variation beyond continental groups. Preprint at Genomics, 10.1101/2024.12.11.627824 https://doi.org/10.1101/2024.12.11.627824. [DOI] [Google Scholar]
  • 122.Smith N, Iyer RL, Langer-Gould A, Getahun DT, Strickland D, Jacobsen SJ, Chen W, Derose SF, and Koebnick C (2010). Health plan administrative records versus birth certificate records: quality of race and ethnicity information in children. BMC Health Serv. Res 10, 316. 10.1186/1472-6963-10-316. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 123.Huang QQ, Sallah N, Dunca D, Trivedi B, Hunt KA, Hodgson S, Lambert SA, Arciero E, Wright J, Griffiths C, et al. (2022). Transferability of genetic loci and polygenic scores for cardiometabolic traits in British Pakistani and Bangladeshi individuals. Nat. Commun 13, 4664. 10.1038/s41467-022-32095-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 124.Guo H, Yang J, Huang J, Xu L, Lv Y, Wang Y, Ren J, Feng Y, Zheng Q, and Li L (2025). Comparative efficacy and safety of GLP-1 receptor agonists for weight reduction: A model-based meta-analysis of placebo-controlled trials. Obes. Pillars 13, 100162. 10.1016/j.obpill.2025.100162. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 125.Zhou K, Gheybi K, Soh PXY, and Hayes VM (2025). Evaluating variant pathogenicity prediction tools to establish African inclusive guidelines for germline genetic testing. Commun. Med 5, 157. 10.1038/s43856-025-00883-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 126.Pathak AK, Bora N, Badonyi M, Livesey BJ, SG10K_Health Consortium, Ngeow, J., and Marsh, J.A. (2024). Pervasive ancestry bias in variant effect predictors. Preprint at Bioinformatics, 10.1101/2024.05.20.594987 https://doi.org/10.1101/2024.05.20.594987. [DOI] [Google Scholar]
  • 127.Picard toolkit (2019). (Broad Institute; ). [Google Scholar]
  • 128.Danecek P, Bonfield JK, Liddle J, Marshall J, Ohan V, Pollard MO, Whitwham A, Keane T, McCarthy SA, Davies RM, et al. (2021). Twelve years of SAMtools and BCFtools. GigaScience 10, giab008. 10.1093/gigascience/giab008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 129.Cleary JG, Braithwaite R, Gaastra K, Hilbush BS, Inglis S, Irvine SA, Jackson A, Littin R, Rathod M, Ware D, et al. (2015). Comparing Variant Call Files for Performance Benchmarking of Next-Generation Sequencing Variant Calling Pipelines. Preprint at Bioinformatics, 10.1101/023754 https://doi.org/10.1101/023754. [DOI] [Google Scholar]
  • 130.Mallick S, Li H, Lipson M, Mathieson I, Gymrek M, Racimo F, Zhao M, Chennagiri N, Nordenfelt S, Tandon A, et al. (2016). The Simons Genome Diversity Project: 300 genomes from 142 diverse populations. Nature 538, 201–206. 10.1038/nature18964. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 131.Bergström A, McCarthy SA, Hui R, Almarri MA, Ayub Q, Danecek P, Chen Y, Felkel S, Hallast P, Kamm J, et al. (2020). Insights into human genetic variation and population history from 929 diverse genomes. Science 367, eaay5012. 10.1126/science.aay5012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 132.Lenth Russell V. (2025). emmeans: Estimated Marginal Means, aka Least-Squares Means. Version 1.11.2. [Google Scholar]
  • 133.Purcell S, Neale B, Todd-Brown K, Thomas L, Ferreira MAR, Bender D, Maller J, Sklar P, De Bakker PIW, Daly MJ, et al. (2007). PLINK: A Tool Set for Whole-Genome Association and Population-Based Linkage Analyses. Am. J. Hum. Genet 81, 559–575. 10.1086/519795. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 134.Hofmeister RJ, Ribeiro DM, Rubinacci S, and Delaneau O (2023). Accurate rare variant phasing of whole-genome and whole-exome sequencing data in the UK Biobank. Nat. Genet 55, 1243–1249. 10.1038/s41588-023-01415-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 135.Shemirani R, Belbin GM, Avery CL, Kenny EE, Gignoux CR, and Ambite JL (2021). Rapid detection of identity-by-descent tracts for mega-scale datasets. Nat. Commun 12, 3546. 10.1038/s41467-021-22910-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 136.Chang CC, Chow CC, Tellier LC, Vattikuti S, Purcell SM, and Lee JJ (2015). Second-generation PLINK: rising to the challenge of larger and richer datasets. GigaScience 4, 7. 10.1186/s13742-015-0047-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 137.McLaren W, Gil L, Hunt SE, Riat HS, Ritchie GRS, Thormann A, Flicek P, and Cunningham F (2016). The Ensembl Variant Effect Predictor. Genome Biol. 17, 122. 10.1186/s13059-016-0974-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 138.Liu X, Li C, Mou C, Dong Y, and Tu Y (2020). dbNSFP v4: a comprehensive database of transcript-specific functional predictions and annotations for human nonsynonymous and splice-site SNVs. Genome Med. 12, 103. 10.1186/s13073-020-00803-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 139.Li H, and Durbin R (2009). Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinforma. Oxf. Engl 25, 1754–1760. 10.1093/bioinformatics/btp324. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 140.Poplin R, Chang P-C, Alexander D, Schwartz S, Colthurst T, Ku A, Newburger D, Dijamco J, Nguyen N, Afshar PT, et al. (2018). A universal SNP and small-indel variant caller using deep neural networks. Nat. Biotechnol 36, 983–987. 10.1038/nbt.4235. [DOI] [PubMed] [Google Scholar]
  • 141.Yun T, Li H, Chang P-C, Lin MF, Carroll A, and McLean CY (2021). Accurate, scalable cohort variant calls using DeepVariant and GLnexus. Bioinforma. Oxf. Engl 36, 5582–5589. 10.1093/bioinformatics/btaa1081. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 142.Jun G, Flickinger M, Hetrick KN, Romm JM, Doheny KF, Abecasis GR, Boehnke M, and Kang HM (2012). Detecting and Estimating Contamination of Human DNA Samples in Sequencing and Array-Based Genotype Data. Am. J. Hum. Genet 91, 839–848. 10.1016/j.ajhg.2012.09.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 143.Ewels P, Magnusson M, Lundin S, and Käller M (2016). MultiQC: summarize analysis results for multiple tools and samples in a single report. Bioinforma. Oxf. Engl 32, 3047–3048. 10.1093/bioinformatics/btw354. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 144.Alirezaie N, Kernohan KD, Hartley T, Majewski J, and Hocking TD (2018). ClinPred: Prediction Tool to Identify Disease-Relevant Nonsynonymous Single-Nucleotide Variants. Am. J. Hum. Genet 103, 474–483. 10.1016/j.ajhg.2018.08.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 145.Li C, Zhi D, Wang K, and Liu X (2022). MetaRNN: differentiating rare pathogenic and rare benign missense SNVs and InDels using deep learning. Genome Med 14, 115. 10.1186/s13073-022-01120-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 146.Feng B-J (2017). PERCH: A Unified Framework for Disease Gene Prioritization. Hum. Mutat 38, 243–251. 10.1002/humu.23158. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 147.Wu Y, Li R, Sun S, Weile J, and Roth FP (2021). Improved pathogenicity prediction for rare human missense variants. Am. J. Hum. Genet 108, 1891–1906. 10.1016/j.ajhg.2021.08.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 148.Ioannidis NM, Rothstein JH, Pejaver V, Middha S, McDonnell SK, Baheti S, Musolf A, Li Q, Holzinger E, Karyadi D, et al. (2016). REVEL: An Ensemble Method for Predicting the Pathogenicity of Rare Missense Variants. Am. J. Hum. Genet 99, 877–885. 10.1016/j.ajhg.2016.08.016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 149.Cheng J, Novati G, Pan J, Bycroft C, Žemgulytė A, Applebaum T, Pritzel A, Wong LH, Zielinski M, Sargeant T, et al. (2023). Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science 381, eadg7492. 10.1126/science.adg7492. [DOI] [PubMed] [Google Scholar]
  • 150.Pejaver V, Urresti J, Lugo-Martinez J, Pagel KA, Lin GN, Nam H-J, Mort M, Cooper DN, Sebat J, Iakoucheva LM, et al. (2020). Inferring the molecular and phenotypic impact of amino acid variants with MutPred2. Nat. Commun 11, 5918. 10.1038/s41467-020-19669-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 151.Carter H, Douville C, Stenson PD, Cooper DN, and Karchin R (2013). Identifying Mendelian disease genes with the variant effect scoring tool. BMC Genomics 14 Suppl 3, S3. 10.1186/1471-2164-14-S3-S3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 152.Rives A, Meier J, Sercu T, Goyal S, Lin Z, Liu J, Guo D, Ott M, Zitnick CL, Ma J, et al. (2021). Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc. Natl. Acad. Sci. U. S. A 118, e2016239118. 10.1073/pnas.2016239118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 153.Quinlan AR, and Hall IM (2010). BEDTools: a flexible suite of utilities for comparing genomic features. Bioinforma. Oxf. Engl 26, 841–842. 10.1093/bioinformatics/btq033. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 154.Gasparini A (2018). comorbidity: An R package for computing comorbidity scores. J. Open Source Softw 3, 648. 10.21105/joss.00648. [DOI] [Google Scholar]
  • 155.Privé F, Aschard H, Ziyatdinov A, and Blum MGB (2018). Efficient analysis of large-scale genome-wide data with two R packages: bigstatsr and bigsnpr. Bioinformatics 34, 2781–2787. 10.1093/bioinformatics/bty185. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 156.P’ng C, Green J, Chong LC, Waggott D, Prokopec SD, Shamsi M, Nguyen F, Mak DYF, Lam F, Albuquerque MA, et al. (2019). BPG: Seamless, automated and interactive visualization of scientific data. BMC Bioinformatics 20, 42. 10.1186/s12859-019-2610-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 157.Puhr R, Heinze G, Nold M, Lusa L, and Geroldinger A (2017). Firth’s logistic regression with rare events: accurate effect estimates and predictions? Stat. Med 36, 2302–2317. 10.1002/sim.7273. [DOI] [PubMed] [Google Scholar]
  • 158.Lambert SA, Wingfield B, Gibson JT, Gil L, Ramachandran S, Yvon F, Saverimuttu S, Tinsley E, Lewis E, Ritchie SC, et al. (2024). Enhancing the Polygenic Score Catalog with tools for score calculation and ancestry normalization. Nat. Genet 56, 1989–1994. 10.1038/s41588-024-01937-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 159.Yao F, Müller H-G, and Wang J-L (2005). Functional Data Analysis for Sparse Longitudinal Data. J. Am. Stat. Assoc 100, 577–590. 10.1198/016214504000001745. [DOI] [Google Scholar]
  • 160.Liu B, and Müller H-G (2009). Estimating Derivatives for Samples of Sparsely Observed Functions, With Application to Online Auction Dynamics. J. Am. Stat. Assoc 104, 704–717. 10.1198/jasa.2009.0115. [DOI] [Google Scholar]
  • 161.Kuznetsova A, Brockhoff PB, and Christensen RHB (2017). lmerTest Package: Tests in Linear Mixed Effects Models. J. Stat. Softw 82. 10.18637/jss.v082.i13. [DOI] [Google Scholar]
  • 162.Breheny P, and Burchett W (2017). Visualization of Regression Models Using visreg. R J. 9, 56. 10.32614/RJ-2017-046. [DOI] [Google Scholar]
  • 163.Wang Xiaofeng (2020). fANCOVA: Nonparametric Analysis of Covariance. Version 0.6-1 [Google Scholar]
  • 164.Lumley T (2004). Analysis of Complex Survey Samples. J. Stat. Softw 9. 10.18637/jss.v009.i08. [DOI] [Google Scholar]
  • 165.Hail Team Hail. Version 0.2.134 [Google Scholar]
  • 166.Cinar O, and Viechtbauer W (2022). The poolr Package for Combining Independent and Dependent p Values. J. Stat. Softw 101. 10.18637/jss.v101.i01. [DOI] [Google Scholar]
  • 167.Zhou W, Nielsen JB, Fritsche LG, Dey R, Gabrielsen ME, Wolford BN, LeFaive J, VandeHaar P, Gagliano SA, Gifford A, et al. (2018). Efficiently controlling for casecontrol imbalance and sample relatedness in large-scale genetic association studies. Nat. Genet 50, 1335–1341. 10.1038/s41588-018-0184-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 168.Willer CJ, Li Y, and Abecasis GR (2010). METAL: fast and efficient meta-analysis of genomewide association scans. Bioinformatics 26, 2190–2191. 10.1093/bioinformatics/btq340. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 169.Verma A, Lucas A, Verma SS, Zhang Y, Josyula N, Khan A, Hartzel DN, Lavage DR, Leader J, Ritchie MD, et al. (2018). PheWAS and Beyond: The Landscape of Associations with Medical Diagnoses and Clinical Measures across 38,662 Individuals from Geisinger. Am. J. Hum. Genet 102, 592–608. 10.1016/j.ajhg.2018.02.017. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 170.Shaw DM, Polikowsky HP, Pruett DG, Chen H-H, Petty LE, Viljoen KZ, Beilby JM, Jones RM, Kraft SJ, and Below JE (2021). Phenome risk classification enables phenotypic imputation and gene discovery in developmental stuttering. Am. J. Hum. Genet 108, 2271–2283. 10.1016/j.ajhg.2021.11.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 171.Sherry ST, Ward M, and Sirotkin K (1999). dbSNP-database for single nucleotide polymorphisms and other classes of minor genetic variation. Genome Res. 9, 677–679. [PubMed] [Google Scholar]
  • 172.Danecek P, Auton A, Abecasis G, Albers CA, Banks E, DePristo MA, Handsaker RE, Lunter G, Marth GT, Sherry ST, et al. (2011). The variant call format and VCFtools. Bioinforma. Oxf. Engl 27, 2156–2158. 10.1093/bioinformatics/btr330. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 173.Verma A, and Ritchie MD (2017). Current Scope and Challenges in Phenome-Wide Association Studies. Curr. Epidemiol. Rep 4, 321–329. 10.1007/s40471-017-0127-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 174.Ye Z, Mayer J, Ivacic L, Zhou Z, He M, Schrodi SJ, Page D, Brilliant MH, and Hebbring SJ (2015). Phenome-wide association studies (PheWASs) for functional variants. Eur. J. Hum. Genet. EJHG 23, 523–529. 10.1038/ejhg.2014.123. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 175.Bastarache L (2021). Using Phecodes for Research with the Electronic Health Record: From PheWAS to PheRS. Annu. Rev. Biomed. Data Sci 4, 1–19. 10.1146/annurev-biodatasci-122320-112352. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 176.Carroll RJ, Bastarache L, and Denny JC (2014). R PheWAS: data analysis and plotting tools for phenome-wide association studies in the R environment. Bioinformatics 30, 2375–2376. 10.1093/bioinformatics/btu197. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 177.Ritchie MD, Denny JC, Crawford DC, Ramirez AH, Weiner JB, Pulley JM, Basford MA, Brown-Gentry K, Balser JR, Masys DR, et al. (2010). Robust replication of genotype-phenotype associations across multiple diseases in an electronic medical record. Am. J. Hum. Genet 86, 560–572. 10.1016/j.ajhg.2010.03.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 178.Pacheco Jennifer and Thompson Will (2012). Type 2 Diabetes Mellitus. PheKB. [Google Scholar]
  • 179.American Cancer Society (2024). Cancer Facts & Figures. [Google Scholar]
  • 180.Hartge P, Struewing JP, Wacholder S, Brody LC, and Tucker MA (1999). The prevalence of common BRCA1 and BRCA2 mutations among Ashkenazi Jews. Am. J. Hum. Genet 64, 963–970. 10.1086/302320. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 181.Gillott RG, Willan K, Kain K, Sivananthan UM, and Tayebjee MH (2016). South Asian ethnicity is associated with a lower prevalence of atrial fibrillation despite greater prevalence of established risk factors: a population-based study in Bradford Metropolitan District. Europace, euw010. 10.1093/europace/euw010. [DOI] [PubMed] [Google Scholar]
  • 182.Agbonlahor O, DeJarnett N, Hart JL, Bhatnagar A, McLeish AC, and Walker KL . Racial/Ethnic Discrimination and Cardiometabolic Diseases: A Systematic Review. J. Racial Ethn. Health Disparities 11, 783–807. 10.1007/s40615-023-01561-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 183.Akinyemi RO, Yaria J, Ojagbemi A, Guerchet M, Okubadejo N, Njamnshi AK, Sarfo FS, Akpalu A, Ogbole G, Ayantayo T, et al. (2022). Dementia in Africa: Current evidence, knowledge gaps, and future directions. Alzheimers Dement. 18, 790–809. 10.1002/alz.12432. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 184.Cohen F, Brooks CV, Sun D, Buse DC, Reed ML, Fanning KM, and Lipton RB (2024). Prevalence and burden of migraine in the United States: A systematic review. Headache J. Head Face Pain 64, 516–532. 10.1111/head.14709. [DOI] [PubMed] [Google Scholar]
  • 185.Siddiqi S, Ortiz Z, Simard S, Li J, Lawrence K, Redmond M, Tomlinson JJ, Schlossmacher MG, and Salmaso N (2025). Race and ethnicity matter! Moving Parkinson’s risk research towards diversity and inclusiveness. Npj Park. Dis 11, 45. 10.1038/s41531-025-00891-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 186.Gutiérrez-Rojas L, Porras-Segovia A, Dunne H, Andrade-González N, and Cervilla JA (2020). Prevalence and correlates of major depressive disorder: a systematic review. Rev. Bras. Psiquiatr. Sao Paulo Braz 1999 42, 657–672. 10.1590/1516-4446-2020-0650. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 187.Baxter AJ, Scott KM, Vos T, and Whiteford HA (2013). Global prevalence of anxiety disorders: a systematic review and meta-regression. Psychol. Med 43, 897–910. 10.1017/S003329171200147X. [DOI] [PubMed] [Google Scholar]
  • 188.Prasad P, and Krishnan E (2014). Filipino Gout: A Review. Arthritis Care Res. 66, 337–343. 10.1002/acr.22118. [DOI] [PubMed] [Google Scholar]
  • 189.Feliciano-Astacio BE, Celis K, Ramos J, Rajabli F, Adams LD, Rodriguez A, Rodriguez V, Bussies PL, Sierra C, Manrique P, et al. (2019). The Puerto Rico Alzheimer Disease Initiative (PRADI): A Multisource Ascertainment Approach. Front. Genet 10, 538. 10.3389/fgene.2019.00538. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 190.Kavanagh PL, Fasipe TA, and Wun T (2022). Sickle Cell Disease: A Review. JAMA 328, 57–68. 10.1001/jama.2022.10233. [DOI] [PubMed] [Google Scholar]
  • 191.Hsu L-A, Wu S, Teng M-S, and Ko Y-L (2023). Causal links of α-thalassemia indices and cardiometabolic traits and diabetes: MR study. Life Sci. Alliance 6, e202302204. 10.26508/lsa.202302204. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 192.Naccashian Z, Hattar-Pollara M, Ho CA, and Ayvazian SP (2018). Prevalence and Predictors of Diabetes Mellitus and Hypertension in Armenian Americans in Los Angeles. Diabetes Educ. 44, 130–143. 10.1177/0145721718759981. [DOI] [PubMed] [Google Scholar]
  • 193.Amaryan G, Sarkisian T, Tadevosyan A, and Braegger C (2023). Familial Mediterranean fever in Armenian children with inflammatory bowel disease. Front. Pediatr 11, 1288523. 10.3389/fped.2023.1288523. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 194.Piel FB, Steinberg MH, and Rees DC (2017). Sickle Cell Disease. N. Engl. J. Med 376, 1561–1573. 10.1056/NEJMra1510865. [DOI] [PubMed] [Google Scholar]
  • 195.Goodrich JK, Singer-Berk M, Son R, Sveden A, Wood J, England E, Cole JB, Weisburd B, Watts N, Caulkins L, et al. (2021). Determinants of penetrance and variable expressivity in monogenic metabolic conditions across 77,184 exomes. Nat. Commun 12, 3505. 10.1038/s41467-021-23556-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 196.O’Malley KJ, Cook KF, Price MD, Wildes KR, Hurdle JF, and Ashton CM (2005). Measuring diagnoses: ICD code accuracy. Health Serv. Res 40, 1620–1639. 10.1111/j.1475-6773.2005.00444.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 197.Denny JC, Ritchie MD, Basford MA, Pulley JM, Bastarache L, Brown-Gentry K, Wang D, Masys DR, Roden DM, and Crawford DC (2010). PheWAS: demonstrating the feasibility of a phenome-wide scan to discover gene–disease associations. Bioinformatics 26, 1205–1210. 10.1093/bioinformatics/btq126. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 198.Charlson ME, Pompei P, Ales KL, and MacKenzie CR (1987). A new method of classifying prognostic comorbidity in longitudinal studies: development and validation. J. Chronic Dis 40, 373–383. 10.1016/0021-9681(87)90171-8. [DOI] [PubMed] [Google Scholar]
  • 199.Lambert SA, Wingfield B, Gibson JT, Gil L, Ramachandran S, Yvon F, Saverimuttu S, Tinsley E, Lewis E, Ritchie SC, et al. (2024). The Polygenic Score Catalog: new functionality and tools to enable FAIR research. Preprint, 10.1101/2024.05.29.24307783 https://doi.org/10.1101/2024.05.29.24307783. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 200.Blondel VD, Guillaume J-L, Lambiotte R, and Lefebvre E (2008). Fast unfolding of communities in large networks. J. Stat. Mech. Theory Exp 2008, P10008. 10.1088/1742-5468/2008/10/P10008. [DOI] [Google Scholar]
  • 201.Staudt CL, Sazonovs A, and Meyerhenke H (2016). NetworKit: A tool suite for large-scale complex network analysis. Netw. Sci 4, 508–530. 10.1017/nws.2016.20. [DOI] [Google Scholar]
  • 202.Wu Y, Gettler K, Kars ME, Giri M, Li D, Bayrak CS, Zhang P, Jain A, Maffucci P, Sabic K, et al. (2023). Identifying high-impact variants and genes in exomes of Ashkenazi Jewish inflammatory bowel disease patients. Nat. Commun 14, 2256. 10.1038/s41467-023-37849-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 203.Mehrjoo Z, Fattahi Z, Beheshtian M, Mohseni M, Poustchi H, Ardalani F, Jalalvand K, Arzhangi S, Mohammadi Z, Khoshbakht S, et al. (2019). Distinct genetic variation and heterogeneity of the Iranian population. PLoS Genet. 15, e1008385. 10.1371/journal.pgen.1008385. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 204.Ziyatdinov A, Torres J, Alegre-Díaz J, Backman J, Mbatchou J, Turner M, Gaynor SM, Joseph T, Zou Y, Liu D, et al. (2023). Genotyping, sequencing and analysis of 140,000 adults from Mexico City. Nature 622, 784–793. 10.1038/s41586-023-06595-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 205.Sohail M, Palma-Martínez MJ, Chong AY, Quinto-Cortés CD, Barberena-Jonas C, Medina-Muñoz SG, Ragsdale A, Delgado-Sánchez G, Cruz-Hervert LP, Ferreyra-Reyes L, et al. (2023). Mexican Biobank advances population and medical genomics of diverse ancestries. Nature 622, 775–783. 10.1038/s41586-023-06560-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 206.Krasheninina O, Hwang Y-C, Bai X, Zalcman A, Maxwell E, Reid JG, and Salerno WJ (2020). Open-source mapping and variant calling for large-scale NGS data from original base-quality scores. Preprint at Bioinformatics, 10.1101/2020.12.15.356360 https://doi.org/10.1101/2020.12.15.356360. [DOI] [Google Scholar]
  • 207.Sun KY, Bai X, Chen S, Bao S, Zhang C, Kapoor M, Backman J, Joseph T, Maxwell E, Mitra G, et al. (2024). A deep catalogue of protein-coding variation in 983,578 individuals. Nature 631, 583–592. 10.1038/s41586-024-07556-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 208.Patel Y, Zhu C, Yamaguchi TN, Wang NK, Wiltsie N, Gonzalez AE, Winata HK, Zeltser N, Pan Y, Mootor MFE, et al. (2024). Metapipeline-DNA: A Comprehensive Germline&Somatic Genomics Nextflow Pipeline. Cell Reports Methods, 2026. 10.1016/j.crmeth.2026.101340 https://www.cell.com/cell-reports-methods/fulltext/S2667-2375(26)00040-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 209.Pinto BJ, O’Connor B, Schatz MC, Zarate S, and Wilson MA (2023). Concerning the eXclusion in human genomics: the choice of sex chromosome representation in the human genome drastically affects the number of identified variants. G3 Bethesda Md 13, jkad169. 10.1093/g3journal/jkad169. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 210.Van Hout CV, Tachmazidou I, Backman JD, Hoffman JD, Liu D, Pandey AK, Gonzaga-Jauregui C, Khalid S, Ye B, Banerjee N, et al. (2020). Exome sequencing and characterization of 49,960 individuals in the UK Biobank. Nature 586, 749–756. 10.1038/s41586-020-2853-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 211.Zadeh N, Getzug T, and Grody WW (2011). Diagnosis and management of familial Mediterranean fever: integrating medical genetics in a dedicated interdisciplinary clinic. Genet. Med. Off. J. Am. Coll. Med. Genet 13, 263–269. 10.1097/GIM.0b013e31820e27b1. [DOI] [PubMed] [Google Scholar]
  • 212.Patel AP, Wang M, Fahed AC, Mason-Suares H, Brockman D, Pelletier R, Amr S, Machini K, Hawley M, Witkowski L, et al. (2020). Association of Rare Pathogenic DNA Variants for Familial Hypercholesterolemia, Hereditary Breast and Ovarian Cancer Syndrome, and Lynch Syndrome With Disease Risk in Adults According to Family History. JAMA Netw. Open 3, e203959. 10.1001/jamanetworkopen.2020.3959. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 213.Olson ND, Wagner J, McDaniel J, Stephens SH, Westreich ST, Prasanna AG, Johanson E, Boja E, Maier EJ, Serang O, et al. (2022). PrecisionFDA Truth Challenge V2: Calling variants from short and long reads in difficult-to-map regions. Cell Genomics 2, 100129, S2666-979X(22)00058-1. 10.1016/j.xgen.2022.100129. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 214.Das S, Forer L, Schönherr S et al. Next-generation genotype imputation service and methods. Nat Genet 48, 1284–1287 (2016). 10.1038/ng.3656. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 215.Forrest IS, Chaudhary K, Vy HMT, Petrazzini BO, Bafna S, Jordan DM, Rocheleau G, Loos RJF, Nadkarni GN, Cho JH, et al. (2022). Population-Based Penetrance of Deleterious Clinical Variants. JAMA 327, 350. 10.1001/jama.2021.23686. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 216.Dennis JK, Sealock JM, Straub P, Lee YH, Hucks D, Actkins K, Faucon A, Feng Y-CA, Ge T, Goleva SB, et al. (2021). Clinical laboratory test-wide association scan of polygenic scores identifies biomarkers of complex disease. Genome Med. 13, 6. 10.1186/s13073-020-00820-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 217.Shuey MM, Stead WW, Aka I, Barnado AL, Bastarache JA, Brokamp E, Campbell M, Carroll RJ, Goldstein JA, Lewis A, et al. (2023). Next-generation phenotyping: Introducing phecodeX for enhanced discovery research in medical phenomics. Bioinformatics 39, btad655. 10.1093/bioinformatics/btad655. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1

Figure S1. Whole-exome sequencing quality control, related to the STAR Methods. a. Coverage distribution. The mean and median values were calculated across all samples. Fold enrichment is the degree to which the baited region is enriched compared to the background genomic region. b. Percentages of bases above different coverage thresholds. Every y-axis point shows the percentage of bases above the corresponding x-axis value. The blue area shows 95% of the data, while the surrounding gray lines represent the remaining 5% of the distribution. c. Average base counts by exome capture regions. d. Excluded bases from coverage calculation based on Picard127 e. Insert size distribution. f. Samtools128 summary statistics output. g. Genotype number distribution according to RTGtools129. h. Summary output from RTGtools. Ti/Tv is the ratio of transition (Ti) to transversion (Tv). This included SNPs from off-target regions. i. The distribution of variant counts across all samples after cohort re-genotyping.

2

Figure S2. Phenotype validation and prevalence, related to Figure 1 and the STAR Methods. (A–J) Comparison of vital signs and laboratory test distributions between case and control groups defined by phecodes, with p values indicating significance from Wilcoxon tests. (K and L) Overlap between cancer cases defined from EHR cancer diagnosis data (“Cancer Stage Fact”) and phecode-based case-control groups for breast (K) and prostate cancer (L). (M) Prevalence of frequent phecodes. The red bars indicate the prevalence of the most frequent phecodes in subset of the UCLA ATLAS population who had an encounter between 1–2 years after their initial encounter, at the time of sample collection. The blue bars represent the prevalence of phecodes in this same subset at their encounters 1–2 years after their initial encounter. (N) The change in phecode prevalence over one year post-collection in the subset of the UCLA ATLAS population who had an encounter between 1–2 years after their initial encounter. The triangles represent the number of new patients, and the blue bars represent the percentage of patients. (O and P) The difference in (O) phecode group and (P) phecode prevalence between participants in the UCLA biobank (purple; “UCLA ATLAS”) at their time of collection compared with all other UCLA Health patients with EHR data (green; “Rest of Data Discovery Repository [DDR]”) within one year of the ATLAS launch date. Logistic regression tests were used to compare the two groups adjusted for sex and age (ORs with 95% confidence intervals, shown in the right panels). After Bonferroni multiple testing corrections, ** denotes a significant difference between the two populations.

3

Figure S3. Broad-scale genetic ancestry and health characteristics, related to Figure 1. a. Agreement between self-reported race and genetic ancestry predictions. b. The Elixhauser comorbidity index varies across genetic ancestries after applying inverse probability weighting. ANOVA on a linear regression model tested the overall effect of the categorical ancestry predictor, yielding its P-value. Adjusted means and 95% CI are presented per ancestry group. c-d. The relationship between total and hospital encounters and the comorbidity index. r, Pearson correlation, P, the correlation P-value. e. Variation in laboratory or vital sign measurements across broad-scale ancestries. ANCOVA adjusted for genetic sex and age was performed separately for each tested phenotype.

4

Figure S4. Fine-scale genetic ancestry supplementary results, related to Figure 2. a-b. Quality control of identity-by-descent (IBD) segments. Distribution of IBD segment length per chromosome, before (A) and after (B) removing human leukocyte antigens (HLA), centromere, and IBD depth outliers. After quality control, IBD segments display exponential decay of segment length as expected (most noticeably for chromosomes 6, 15 and 22). The numbers at the top of each plot represent the chromosome number. c. The distribution of genetic ancestry between fine- and broad-scale populations. IBD clusters were sorted by the predominant broad-scale ancestry of participants in each cluster, rather than cluster size as shown in Figure 2a, providing an alternative visualization. D. Cardio-metabolic disease risk for each fine-scale group within the same broad-scale ancestry. Representative cardio-metabolic phecodes were selected, and only populations with at least 100 participants were tested. In cases of small sample sizes, ‘–’ was used instead of numeric values to protect patient privacy. Filled points represent significant results (FDR ≤ 0.05). Firth’s bias-reduced logistic regression adjusted for BMI, sex, and age was used to obtain OR. Among Asian clusters, Filipino individuals were at high risk for all tested medical conditions (essential hypertension: ORIBD-09 =1.6 [1.4, 1.9], FDR = 8.6 × 10−11; type 2 diabetes: ORIBD-09 = 1.5 [1.3, 1.7], FDR = 7.4 × 10−6; coronary atherosclerosis: ORIBD-09 = 1.4 [1.1, 1.7], FDR = 5.3 × 10−3; abdominal aortic aneurysm: ORIBD-09 = 2.8. [1.4, 5.3], FDR = 7.1 × 10−3; hyperlipidemia: ORIBD-09 = 1.2 [1.0, 1.4], FDR = 2.4 × 10−2). The largest Armenian cluster and Jewish and non-Jewish Iranian clusters showed a high risk for type 2 diabetes (Armenian 1: ORIBD-16 = 2.0 [1.5, 2.7], FDR = 2.5 × 10−6; Iranian Jewish: ORIBD-11 = 2.4 [1.9, 2.9], FDR = 6.0 × 10−16; Iranian: ORIBD-17 = 2.2 [1.6, 2.9], FDR = 1.6 × 10−6), hyperlipidemia (Armenian 1: ORIBD-16 = 1.6 [1.2-2.0], FDR = 8.6 × 10−4; Iranian Jewish: ORIBD-11 = 1.5 [1.2, 1.8], FDR = .1.0 × 10−4; Iranian: ORIBD-17 = 1.8 [1.4, 2.3], FDR = 2.9 × 10−5) and coronary atherosclerosis (Armenian 1: ORIBD-16 = 1.9 [1.4, 2.5], FDR = 8.8 × 10−5; Iranian Jewish: ORIBD-11 = 1.8 [1.5, 2.2], FDR = 4.2 × 10−8; Iranian: ORIBD-17 = 1.9 [1.4, 2.5], FDR = 2.9 × 10−5).

5

Figure S5. polygenic score association results, related to Figure 3 and the STAR Methods. a-e Selected associations between polygenic scores (PGS) and diseases in European (EUR) individuals. f-i The odds ratio (OR) and prevalence of cases for the top and bottom PGS deciles across non-EUR ancestries. Logistic regression was used to obtain the OR and p values, which were adjusted using FDR.

6

Figure S6. PhWAS supplementary results, related to Figure 4.

a-b. Distribution of pruned variant-trait associations. The number of genome-wide significant pruned gene-trait associations shared across fine-scale ancestries, using a prioritized gene within 10kb of the pruned variant and matching effect direction. b presents the same associations described in a, but with a more permissive threshold for gene-trait support from less powered ancestral clusters. c. APOE haplotype frequency across fine-scale cohorts. d. Allele frequency for known risk variants across fine-scale cohorts; shading indicates the level of over- (red) or under- (blue) enrichment of a haplotype/allele in a given cohort via Fisher’s exact test. In c-d, Fisher’s exact test was utilized to compare allele frequency. Bold border indicates a significant result (FDR ≤ 0.1). e. Impact of the non-alcoholic cirrhosis risk variant rs738409-G on cirrhosis and clinical sequalae across finescale cohorts. Odds ratios and 95% CI were calculated using logistic regression. f-h. Replicated low-MAF PheWAS associations between rs115750084-G and major depressive disorder (f), rs77742325-G and osteoporosis with no other symptoms (g) and rs202215133-A and migraine (h).

6

Figure S7. Known or putative rare pathogenic variants, related to Figure 5 Figure 5 and the STAR Methods.

a. Frequency differences in Familial Mediterranean Fever (FMF) known risk alleles. b. The risk of carrying the HBB:p.E7V variant. c. The risk of carrying loss-of-function variants in PCSK9. In a-c, the error bars show the 95% Wilson score confidence intervals. d-e. The frequency of BRCA Ashkenazi Jewish founder alleles across broad-scale (d) and fine-scale ancestry groups (e). f. Differences across populations in the total numbers of rare ClinGene P/LP variants in American College of Medical Genetics (ACMG) genes, excluding Ashkenazi Jewish, as a sensitivity analysis to one presented in Figure 5b. ‘Ref’ is the total number of reference alleles, and ‘Alt’ of P/LP ClinGen. Fisher’s exact tests were used to produce odds ratios. Filled points represent significant results at the level of nominal P-value (≤ 0.05). Only fine-scale ancestry clusters with more than 400 participants were included. g-h. Distributions of missense variant pathogenicity rank scores for nine computational tools. Each subplot compares the distribution of all missense variants (purple to green) to known pathogenic variants (red), defined as either: g. ClinGen curated missense variants. h. ClinVar pathogenic/likely pathogenic (P/LP) missense variants. The black dashed line indicates the median rankscore across all missense variants for the given tool. The red dashed line denotes the median rankscore among pathogenic variants in the corresponding dataset. The purple dashed line marks the likelihood-based intersection cutoff derived from the point at which the pathogenic and background distributions cross. i. The number of rare computationally predicted damaging missense and LOF alleles per individual across ancestries. Mann-Whitney U test with a Bonferroni correction was applied to test the difference in the distribution of rare LOF and predicted damaging missense counts per individual between any broad- or fine-scale group compared to all others. Statistically significant differences are indicated by an asterisk (*). Only fine-scale ancestry clusters with more than 100 participants were included. j. Ancestry-specific carrier counts for two known GBA1 rare LOF variants (left) and their impact on Parkinson’s disease risk (right) via logistic regression.

13

Figure S8. GLP1-RAs ATLAS users and semaglutide investigation complementary data, related to Figure 6. a. Age by sex of GLP-1 receptor agonist (GLP1-RAs) users. b. Prescription numbers for GLP1-RAs by simple generic names. c. Weight loss patterns across time. Presented are smoothed longitudinal data using a functional boxplot approach, with median values and pointwise intervals between the 20th and 80th quantiles. Vertical tick marks along the x-axis show the deciles of the data distribution, indicating where most data points are concentrated across weeks. d. The relationship between bins of polygenic scores (PGS) for body mass index (BMI) (left) and type 2 diabetes mellitus (right), and weight loss in response to semaglutide. Shown are Loess smoothed plots with 95% CI based on the maximum weight loss for patients across weeks. p values and effect sizes were obtained from a linear regression model with covariates, and a Bonferroni correction was applied. e. The relationship between type 2 diabetes mellitus PGS and weight loss in response to semaglutide in EUR AoU participants. Scaled PGS were divided into groups and linear mixed model fitted values were plotted with 95% CI, based on longitudinal data with repeated weight measurements. P-values were obtained from a linear mixed-effects model that included covariates. f. GWAS Q-Q plot for weight loss on semaglutide displays no significant findings. g. Gene-level test Q-Q plot for weight-loss on semaglutide (genomic inflation factor: 0.95). h. Differences in carrier numbers of semaglutide efficacy involved alleles in PTPRU across ancestries using Firth’s bias-reduced logistic regression. Only variants that exhibited at least one significant difference between EUR and another ancestry are shown. Filled dots represent significant odds ratios (FDR ≤0.05).

7

Table S1. Summary of lab test results and prescription numbers for the top 50 prescribed medications in ATLAS, related to the STAR Methods.

8

Table S2. Phecodes and broad-scale ancestry associations, related to Figure 1 and the STAR Methods.

9

Table S3. Fine-scale ancestry cluster sample sizes compared with published cohorts and phecode associations, related to Figure 2 and the STAR Methods.

10

Table S4. PheWAS results and pharmacogenomic variants, related to Figure 4.

11

Table S5. Summary of WES variant counts by category and differences across broad-scale ancestries, related to Figure 5 and the STAR Methods.

12

Table S6. ExWAS results, related to Figure 5.

RESOURCES