Summary
Precision medicine initiatives across the globe have led to a revolution of repositories linking large-scale genomic data with electronic health records, enabling genomic analyses across the entire phenome. Many of these initiatives focus solely on research insights, leading to limited direct benefit to patients. We describe the biobank at the Colorado Center for Personalized Medicine (CCPM Biobank) that was jointly developed by the University of Colorado Anschutz Medical Campus and UCHealth to serve as a unique, dual-purpose research and clinical resource accelerating personalized medicine. This living resource currently has more than 200,000 participants with ongoing recruitment. We highlight the clinical, laboratory, regulatory, and HIPAA-compliant informatics infrastructure along with our stakeholder engagement, consent, recontact, and participant engagement strategies. We characterize aspects of genetic and geographic diversity unique to the Rocky Mountain region, the primary catchment area for CCPM Biobank participants. We leverage linked health and demographic information of the CCPM Biobank participant population to demonstrate the utility of the CCPM Biobank to replicate complex trait associations in the first 33,674 genotyped individuals across multiple disease domains. Finally, we describe our current efforts toward return of clinical genetic test results, including high-impact pathogenic variants and pharmacogenetic information, and our broader goals as the CCPM Biobank continues to grow. Bringing clinical and research interests together fosters unique clinical and translational questions that can be addressed from the large EHR-linked CCPM Biobank resource within a HIPAA- and CLIA-certified environment.
Keywords: biobanking, electronic health records, precision medicine, pharmacogenomics, learning health system
The CCPM Biobank is a dual-purpose research and clinical resource to accelerate genetic discovery and translation into clinical practice. We have enrolled >200,000 participants and returned genetic findings to thousands. We are uniquely positioned among EHR-linked biobanks, due to our size, integrated learning health system design, and unique catchment area.
Introduction
Population-scale biobanks increasingly provide opportunities to understand patterns of health and disease within and across populations, discover the factors that drive these patterns, and enable advances in healthcare.1,2,3,4 The biobank at the Colorado Center for Personalized Medicine (CCPM Biobank) was jointly developed by the University of Colorado Anschutz Medical Campus and UCHealth to serve as a unique, dual-purpose research and clinical resource accelerating personalized medicine. The CCPM Biobank has broad objectives to collect, store, and generate data on biological samples within a clinically certified environment. The goals of these activities are to support scientific discovery, collaboration with partners, and return of actionable clinical results. As a resource comprising electronic health records (EHRs), genotype data, and other integrated data sources (e.g., geocoded data and survey data), the CCPM Biobank is available to the University of Colorado (CU) community for discovery research and pragmatic clinical trials across a range of both acute and chronic conditions. From its inception, the CCPM Biobank was designed as a translational medicine initiative that delivers biobank-fueled innovations back to participants and their healthcare providers to inform healthcare decisions.
In the context of the ongoing revolution of personalized medicine initiatives worldwide, the CCPM Biobank fills a unique niche with diverse data from the Rocky Mountain region. UCHealth serves both a highly urban population in the Denver, Colorado Springs, and Fort Collins metropolitan areas as well as numerous rural populations in the intermountain west and western great plains regions of the United States. UCHealth is composed of 679 clinics and 13 hospitals, including a quaternary care facility and academic medical headquarters at the Anschutz Medical Campus in Aurora, Colorado. UCHealth has served more than 6 million people in the last 10 years including approximately 2 million active patients seen within the past 2 years. UCHealth members enrolled in the CCPM Biobank receive care for a broad range of conditions and come from diverse and often unique populations with respect to medical conditions, socioeconomic status, geographic context, and cultural norms. In particular, our catchment includes many individuals living in rural and/or high-altitude settings that are rare in other biobanks. Their participation enables unprecedented insights into interactions between genetic and environmental factors influencing health in these understudied settings.
Here we describe the overall structure of the CCPM Biobank and present initial findings supporting its simultaneous utility for research and clinical care transformation with more than 200,000 enrolled participants and 33,674 genotyped participants as of March 2022. We characterize the current base of consented individuals and our first cohort of genotyped participants. We demonstrate the utility for research via association studies for ten common cardiovascular, immune, metabolic, and anthropometric polygenic traits. Finally, we describe the opportunities to translate research insights directly into clinical applications, including return of actionable clinical genetic test results, as we realize the Anschutz Medical Campus and UCHealth vision for personalized medicine for all.
Subjects and methods
All procedures and analyses described here were performed in accordance with the ethical standards of Colorado Multiple Institutional Review Board and in accordance with the informed consent of CCPM participants.
CCPM planning and development
In 2014 the major partners at the Anschutz Medical Campus—University of Colorado Anschutz Medical Campus, UCHealth, Children’s Hospital Colorado, and CU Medicine—came together to create the CCPM. At its inception, CCPM launched a campus-wide survey of needs and interests in personalized medicine present in both the university and health system for both clinical implementation and research. CCPM members, in discussions with the Colorado Multiple Institutional Review Board (COMIRB) with input from UCHealth Patient and Family Advisory Councils, began developing an innovative consent model that supported the dual goals of the CCPM Biobank and held focus groups to develop consenting materials that balanced scientific accuracy and approachability.5
Participant recruitment and engagement
The CCPM Biobank began enrollment in September 2015 with a self-consent, paper-based process for adult participants (age ≥18 years, who could consent for themselves in English) receiving medical care at UCHealth University of Colorado Hospital. Children were not enrolled. The two-page consent form authorized collection of blood samples left over after clinical testing at UCHealth and provided broad consent for research and participant recontact, leaving open the opportunities for return of clinical genetic test results. Supporting information for consent included clinician education, patient information brochures, a frequently asked questions (FAQs) sheet, and a telephone helpline. This initial enrollment process was piloted within the Cardiac and Vascular Center at the UCHealth University of Colorado Hospital. In February 2016, the consent form was updated to allow for an additional tube of blood to be collected, dedicated to the Biobank. In July 2016, the consent process was expanded to include three additional clinics at the UCHealth University of Colorado Hospital and to allow for consent via the UCHealth My Health Connection patient portal.
In October 2016, enrollment expanded again to additional UCHealth clinics in the Denver metropolitan area and we added an informational video to augment the online consent process.6 In August 2018, we transitioned entirely to an electronic self-consent model using My Health Connection, UCHealth’s online patient portal, and enrollment was opened to patients across the entire UCHealth system. To enable the return of clinical genetic test results to participants and their healthcare providers, in the Spring of 2018, we launched a pilot of return of results by recontacting CCPM Biobank participants with high-impact pathogenic variant findings. The consent form was also revised at this time to include additional information on clinical and research aims, sharing of data, Colorado legal statutes, and the Genetic Information Non-discrimination Act (GINA). In October 2018, we began a large secondary consent initiative to authorize actual return of clinically actionable test results. In November 2019, we launched a revised consent form to simultaneously enroll in research and authorize clinical return of results (referred to as the “unified” consent). Many versions of the consenting forms/approaches have been offered in both English and Spanish.
To keep participants informed and involved in CCPM activities, we have implemented several engagement efforts including sending monthly on-boarding messages to welcome newly enrolled CCPM Biobank participants and quarterly newsletters to all participants. Newsletters update participants on new partnerships, recent research activities within CCPM, and opportunities to become involved in CCPM Biobank activities, such as focus groups and new research studies performed in CCPM. Our website (www.cobiobank.org) is used to maintain participant engagement and to support educational efforts for both patients and providers around return of clinical genetic test results such as those related to pharmacogenetics, disease risk, and carrier status. Participant engagement approaches are in continual development due to the dynamic nature of the scope of the Biobank for research discovery and clinical care, and our continuously evolving capability with respect to return of clinical genetic test results.
Sample acquisition and processing
Samples from early participants consisted of blood samples left over after clinical testing at UCHealth clinical laboratories. The process of identifying consented participants and acquiring leftover specimens were manual processes and therefore resource and time intensive. To improve sample acquisition workflows, in conjunction with the February 2016 consent update, we launched an automated sample collection process whereby participant consent recorded in the EHR triggers a blood collection order to collect a dedicated CCPM Biobank sample at the participant’s next clinical blood draw. This improved our collection of biospecimens, but there remained a significant gap between participants enrolled and samples collected, in part because not all active UCHealth patients would ever receive an order for a blood draw (e.g., individuals being seen by the ophthalmology department). This was exacerbated by the onset of the COVID-19 pandemic. In March 2021 we piloted the collection of saliva samples at limited UCHealth locations during COVID-19 vaccination clinics for individuals who consented to participate in the Biobank.
The CCPM Biobank Laboratory is accredited by the College of American Pathologists (CAP) and certified by the Clinical Laboratory Improvement Amendments (CLIA) for high-complexity testing. Samples are accessioned into a clinically compliant Laboratory Information System (LIS) (SLIMS, Agilent, Inc.), which links patient demographics to samples and their derivatives. Samples collected in clinical labs, including blood (4mL, EDTA) and saliva (2mL, Genefi, IsoHelix), are transported to the CCPM Biobank Laboratory. Genomic DNA is extracted into 2D barcoded tubes in a 96-well plate format on a FlexSTAR+ instrument (Autogen Inc.) utilizing Flexigene chemistry (Qiagen). Nucleic acid yield is determined by fluorescence (QuantiFluor ds DNA system, Promega) and UV absorbance, and these values, along with characteristics such as volume and storage location, are linked to each sample in the LIS. DNA samples are stored at 4°C prior to genotyping and at −20°C for long-term storage.
Genotyping was performed on one of two different customized versions of Illumina’s Infinium Expanded Multi-Ethnic Genotyping Array (MEGA-EX) according to manufacturer’s instructions. We chose this platform as the primary genotyping array given our extensive experience in development of genotyping platforms in The Consortium on Asthma among African-ancestry Populations in the Americas (CAAPA)7,8 and The Population Architecture using Genetics and Epidemiology (PAGE)9,10,11 studies and because it was designed with marker representation from numerous global populations to cover both clinical and scientific interests.12 Customized content on the MEGA-EX was composed of probes for >40,000 SNPs requested by University of Colorado Anschutz Medical Campus research community, with prior extensive QC on call rates and to ensure sample consistency (see supplemental methods). Among the ∼2.1 million genetic variants on MEGA-EX are several thousand with potential clinical relevance, a subset of which have either been analytically and clinically validated, or had a confirmatory assay developed, for later return to participants and their providers.
Research applications
In addition to genomic data, the CCPM Biobank contains rich phenotypic information through linkage to EHR. UCHealth has maintained a single Epic-based EHR system across all of its sites since 2011 and currently has records for more than 8 million unique individuals. Health Data Compass (HDC) (healthdatacompass.org), our institutional research data warehouse, extracts individuals’ data from Epic including demographics, diagnosis codes, notes, flow sheets, lab values, admissions/discharges, and medications. Data are harmonized to the Observational Medical Outcomes Partnership (OMOP) common data model13 and linked to external data sources such as immunization records and vital statistics including death records from the Colorado Department of Health and Environment as well as Colorado All Payers Claims Database (APCD) (see web resources). These integrated data are stored in a HIPAA-compliant environment in Google Cloud. Approved researchers can securely access approved analytic datasets via Google BigQuery and secure cloud-based workspaces, allowing for easy access using structured query language (SQL) or with application programming interfaces (APIs) in R and Python. Data can be made available to researchers in identified or de-identified form and can be refreshed at regular intervals to support longitudinal research.
Oversight of proposed research using the CCPM Biobank resources is governed by the Access to Biobank Committee (ABC), composed of regulatory specialists, bioethicists, epidemiologists, informaticians, and statistical geneticists. In addition to regulatory approvals required by Health Data Compass (i.e., COMIRB approval for identified protocols, approval by UCHealth Privacy and Security Officers, etc.), the CCPM Biobank requires additional approval to ensure research requests respect availability of resources and are consistent with both the letter and spirit of the CCPM Biobank consent. The application process includes a description of the proposed research project, the analytical team, workflow, resource storage and security, participant inclusion and exclusion criteria, as well as a description of the project in lay language to allow participants to learn about active research projects through the CCPM Biobank newsletter.
Clinical applications
As of this submission, three types of clinical genetic test results are being returned to CCPM Biobank participants and their providers when the participant has signed an appropriate consent: (1) high-impact pathogenic variants, (2) risk for hereditary hemochromatosis (HFE), and (3) pharmacogenetic results. We have implemented separate processes for returning these three types of results.
Our decision to return high-impact pathogenic variants was taken with consideration of the policy from the American College of Medical Genetics and Genomics (ACMG)14,15 for reporting of secondary findings in clinical exome and genome sequencing, with a concentration on those genes for which actionable interventions exist to alter risk and/or disease outcome. Variants from MEGA genotyping in the current list of secondary genes recommended for return by ACMG14,16,17 that are interpreted to be pathogenic or likely pathogenic according to the ACMG and Association for Molecular Pathology (AMP) guidelines for variant interpretation18 are confirmed in the CCPM Biobank Laboratory with a validated Sanger sequencing method. Due to the potential impact on participants and their families, a certified genetic counselor returns all high-impact pathogenic variant findings to participants. Although all participants consented to receive clinically relevant findings as part of the secondary or unified consent process, when recontacting participants we first re-confirm their desire to receive clinically relevant results. Once consent is confirmed, results are returned via telephone by a certified genetic counselor, and the clinical laboratory report and additional documentation of the result (e.g., notes, direct messaging to the participant’s healthcare provider(s), adding a diagnosis to the participant’s problem list) is placed in the EHR, as appropriate. The laboratory report is available to the participant through their online patient EHR portal, and participants are referred for clinical follow-up with specialists with the participant’s consent.
For those participants identified to have an increased risk for hereditary hemochromatosis and who have signed a secondary or unified consent, an order for genetic HFE testing is completed and the result placed directly into the participant’s EHR without contact by a certified genetic counselor. This approach was chosen due to the relatively low penetrance of hereditary hemochromatosis, the relatively high prevalence of homozygosity for the variant of interest (c.845C>G [p.Cys282Tyr]) in the general population, and the ability of primary care providers to provide follow-up care and confirmatory assays, including blood ferritin testing.
A summary of the CCPM Biobank clinical pharmacogenomics implementation process has been previously published.19 In brief, the MEGA-EX array contains a number of known pathogenic and functional variants selected from ClinVar (www.ncbi.nlm.nih.gov/clinvar/),20 Online Mendelian Inheritance of Man (OMIM, omim.org),21 PharmGKB (pharmgkb.org),22,23 Clinical Pharmacogenetics Implementation Consortium (CPIC) (cpicpgx.org), and other databases of known variants that could be relevant to patient care. To streamline the pharmacogenomic implementation process, variants with clinical potential are validated on the array using complementary data generation approaches to ensure high fidelity and technical validity of results generated using the MEGA-EX platform directly rather than relying on confirmatory testing as in the case of high-impact pathogenic variants. A custom, automated pipeline using BC Platforms (https://www.bcplatforms.com/) software provides the translation layer from a standard text format into a version that can be ingested by the UCHealth EHR system and filed as structured data elements. Structured results are then used to trigger drug-gene-specific clinical decision support (CDS) tools (e.g., interruptive alerts, passive warnings), which provide clinicians with pertinent pharmacogenetic results and clinical recommendations at the point of prescribing.
Due to well-known limits in genomics knowledge among clinicians,24,25,26 local health professions’ groups and representative stakeholders were engaged throughout the development of our return of results process. Educational materials were created with involvement of relevant practitioner groups, including physicians, advanced practice providers, and pharmacists. These materials undergo at least annual assessment and revision to promote the use of genetic test results generated by the CCPM Biobank in clinical care.
CCPM biobank characterization and genomic replication
We identified all participants enrolled in the CCPM Biobank as of February 24, 2022 inclusive of those genotyped to calculate descriptive statistics on population demographics (gender, race/ethnicity, and age). Racial and gender identities with representation of less than 1% in the cohort were removed for privacy reasons. Median observation period was calculated using the entry dates of the first and last diagnosis codes in each patient’s record, excluding medical history diagnoses which are associated with prior dates of presentation, not the date entered in the record. The frequency and coverage (i.e., the percent of individuals with one or more entries) of different data types (encounters, diagnoses, procedures, flowsheets, medications, laboratory tests) were calculated for the entire CCPM Biobank population from January 1, 2011 (the time of EHR implementation) to February 24, 2022. We also analyzed the prevalence of medical conditions in the CCPM Biobank using more than 1,200 phecode-based phenotypes27 derived from ICD9-CM and ICD10-CM codes using the PheWAS R package.28,29 Finally, maps of participant location were created using 3-digit zip code (as required by HIPAA de-identification regulations) and zip code tabulation area (ZCTA) from the United States Census Bureau (see web resources). CCPM Biobank participant data were aggregated over the resulting regions and censored to remove 17 ZCTAs with a population of 20,000 or fewer persons based on census data from the most recent 3-digit count in 2000 (see web resources). All analyses were conducted using R v.3.6.0 (see web resources) and a variety of packages for data processing, graphics, and reporting (see Magrittr, Glue, Here, rgdal, shades, albersusa, bigquery, gtsummary, maps, and tidylog in the web resources).30,31,32,33,34
We performed association tests at ten well-established loci in the GWAS catalog using imputed genotypes in the first set of 33,674 patients with MEGA genotype data (“Research Pilot”). We conducted extensive quality control on genotype batch, inter-plate concordance, and SNP and sample filtering on sites directly genotyped on MEGA (see supplemental methods), then imputed additional genotypes using the TOPMed Imputation Server35,36 in two batches, keeping the intersection of SNPs in both batches with r2 > 0.7 (see supplemental methods). Each association test used genotype, sex, age, the first five kinship-adjusted principal components by GENESIS37,38,39,40 as fixed effects, along with the genomic background correction (for highly polygenic traits and relatedness) present in REGENIE.41 We assessed proper calibration of association tests via the genomic control inflation factor lambda42 and standard qq plots. We then compared the odds ratios of the significant associations found within the CCPM Research Pilot study to previously reported findings in the GWAS catalog43 (accession numbers available in supplemental information) for each phenotype. To ensure consistent stranding, we ignored variants with ambiguous or unreported effect alleles. Confidence intervals were reported using a normal approximation to the effect size estimate to calculate the standard error.
Results
Enrollment and participant summary
As of March 2022, 200,673 participants have enrolled in the CCPM Biobank and 103,600 have provided a blood (n = 100,542) or saliva (n = 3,058) sample (Figure 1). Enrollment increased significantly starting in August 2018 when we moved to an electronic consent process that became available to patients across the entire UCHealth system that represents approximately 2 million active patients. Since that time, an average of more than 3,000 participants enroll in the CCPM Biobank every month and an average of more than 1,700 samples are collected. The ratio of consent to samples (∼2:1) has remained fairly consistent over the past several years and reflects the delay in sample collection inherent in our process to collect an additional tube of blood at a future clinical blood collection. About 60% of all participants have signed a secondary or unified consent to opt in to receiving clinical genetic test results from the CCPM Biobank (see participant recruitment and engagement in Subjects and methods). We have maintained participant engagement via on-going outreach informing participants of new research and industry partners with whom we share data, facilitating recontact for research opportunities and return of results. We also provide education to our participants around genomics and personalized medicine, and to showcase studies for which they have contributed. E-mail open rates for our newsletter and on-boarding messages hover around 30% and 60%, respectively. These efforts have contributed to high retention rates, with a median time of enrollment of 33.5 months; 726 participants have withdrawn from the study, and 6,398 have died.
Figure 1.
Cumulative enrollment and sample collection of CCPM Biobank participants over time
Enrollment shown in blue and sample collection in red.
CCPM Biobank participants live in all 50 states as well as the District of Columbia (Figure 2A), although the majority (∼67%) reside along the Front Range region of Colorado, which includes the Denver metropolitan area, Fort Collins, and Colorado Springs (Figure 2B). Participants are predominantly female (59%), white (85%), and non-Hispanic (88%), reflecting the demographics of the underlying UCHealth patient population. The age of participants is mostly uniformly distributed for age groups from 30 to 69 and is somewhat lower for younger adults (18–29) and those aged 70 and older (Table 1).
Figure 2.
Density of recruitment of CCPM Biobank participants
Density across (A) the USA and (B) Colorado shown, at the level of the first three digits of zip codes.
Table 1.
CCPM biobank participants and research pilot demographics
|
CCPM Biobanka |
Research pilotab |
|
|---|---|---|
| (n = 196,559) | (n = 25,125) | |
| Sex | ||
| Female | 116,812 (59%) | 16,129 (64%) |
| Male | 79,747 (41%) | 8,996 (36%) |
| Age in years | 50 (36, 65) | 54 (40, 67) |
| Age group | ||
| 18–29 | 23,391 (12%) | 1,093 (4.4%) |
| 30–49 | 76,074 (39%) | 9,854 (39%) |
| 50–59 | 32,096 (16%) | 4,500 (18%) |
| 60–69 | 33,892 (17%) | 5,024 (20%) |
| 70–79 | 24,009 (12%) | 3,374 (13%) |
| 80+ | 7,097 (3.6%) | 1,280 (5.1%) |
| EHR race | ||
| African American | 7,466 (3.9%) | 1,160 (4.7%) |
| Asian | 4,080 (2.1%) | 593 (2.4%) |
| White | 163,207 (85%) | 20,758 (84%) |
| Multiple races | 5,995 (3.1%) | 712 (2.9%) |
| Otherc | 11,169 (5.8%) | 1,368 (5.6%) |
| Unknown | 4,642 (0%) | 534 (0%) |
| EHR ethnicity | ||
| Hispanic | 18,137 (9.2%) | 2,181 (8.7%) |
| Non-Hispanic | 172,728 (88%) | 22,304 (89%) |
| Unknown | 5,694 (2.9%) | 640 (2.5%) |
| APCDdavailability | 72,270 (37%) | 10,373 (41%) |
| Length of record in years | 5.2 (2.4, 7.9) | 7.2 (4.8, 9.4) |
Statistics presented: n (%); median (IQR)
The research pilot consisted of two genotype batches (see Supplemental Methods). Counts presented in this table represent participants genotyped in the first batch
Native Hawaiian, Other Pacific Islander, Native American, Alaska Native, Other
All Payers Claims Database
Rich clinical history for each CCPM Biobank participant is available in the form of more than 600 million data points in the EHR. Across a median observation period of 5.2 years, more than 99% of participants have information for at least one clinical encounter, diagnosis codes, and procedure codes (Table 2). We used ICD-9-CM and ICD-10-CM diagnosis codes in the EHR to derive more than 1,200 phecodes covering a wide range of complex diseases for each participant. Nearly 50% of participants have a phecode for an endocrine/metabolic disorder, largely driven by disorders in lipid metabolism (Figure 3). Nearly 25% of participants have the phecode for hypertension, which is the single most frequent disorder noted in our population (Figure 3).
Table 2.
Clinical data availability for CCPM biobank participants
| Event type | Total event count | Participant count (% cohort represented) |
|---|---|---|
| Encounter entry | 40,052,867 | 196,370 (99.9%) |
| Diagnosis code | 40,576,473 | 194,554 (99.0%) |
| Procedure code | 125,637,476 | 196,236 (99.8%) |
| Flowsheet entry | 502,664,992 | 195,313 (99.4%) |
| Medication order | 18,583,715 | 181,989 (92.6%) |
| Laboratory value | 85,632,533 | 183,352 (93.3%) |
Figure 3.
CCPM Biobank participant proportional phecode use by domain and specific phecode as derived from the EHR across participants in the entire CCPM Biobank
Domain shown at left and phecode at right.
Genotyping and imputation
We chose the first 34,435 CCPM Biobank participants with verified consents to be genotyped on our custom MEGA12 platform, which contains ∼2 million markers selected for representation across diverse populations. Extensive quality control metrics were used to filter markers and samples (see supplemental methods). The genotyping success rate exceeded 98% of samples, and sex concordance between inferred genetic sex and sex or gender as reported in the EHR was >99.8%; of the 75 sex discrepancies, only 8 were unable to be resolved by EHR review for sex chromosome abnormality or transgender status. After QC was complete, 33,674 participants genotyped at 1,696,932 sites remained. We leveraged these sites to perform whole-genome imputation using the TOPMed reference panel to yield more than 98 million loci. After filtering to loci with confident imputation quality, ∼50M sites remained (see supplemental methods).
Population structure and relatedness
Although family relationships often are not documented in EHR data, precise family structures can be inferred using dense genotype information. To infer familial relationships, we used our PONDEROSA algorithm,44 a method robust to arbitrary levels of endogamy, and KING-robust.40 We identified 761 first-degree relative pairs (comprising both parent-offspring and sibling pairs), 300 second-degree relative pairs, and 465 third-degree relative pairs (Figure 4A). From these relatives, we reconstructed numerous pedigrees, the largest of which is composed of 5 individuals and spans 3 generations.
Figure 4.
Ancestry and population structure estimates in the CCPM Biobank
(A) Counts of closely related pairs of individuals as determined by PONDEROSA.
(B) Principal components analysis of CCPM Biobank participants (gray points) overlaid with individuals from a global reference panel (colored points).
(C) Admixture estimates for CCPM Biobank participants from the five largest EHR race/ethnicity categories. 1000 Genomes and Human Genome Diversity Project (HGDP) data are provided for reference.
Race and ethnicity information encoded in the EHR in the form of census categories is available for each Biobank participant, but ambiguities in racial and ethnic categories or mis-classification of race/ethnicity during clinical visits may result in classifications that are inappropriate in genetic analyses (e.g., when controlling for population structure). We used genome-wide SNP data with different dimension reduction methods (see supplemental methods) to infer genetic ancestry cluster memberships for each CCPM Biobank participant. Although the majority (∼80%) of CCPM Biobank participants are of primarily European descent, participants represent a wide spectrum of diverse ancestries (Figures 4B and 4C), including substantial populations with genetic ancestry from Africa, South Asia, East Asia, and the Americas. The full spectrum of ancestry is further highlighted with global ancestry proportions (Figure 4C).
GWAS catalog replication findings
To evaluate the utility of the CCPM Biobank as a research tool, we conducted genetic association tests on 10 traits with previously well-established causal loci including Alzheimer disease, rheumatoid arthritis, psoriasis, multiple sclerosis, hypothyroidism, type 1 and type 2 diabetes, obesity, breast cancer, and asthma (see pilot analyses in Subjects and methods and supplemental methods), and we compared the effect sizes of the most significant SNP for each trait with overlapping associations previously reported in the GWAS catalog.43 The effect sizes of overlapping SNPs are largely consistent with those previously reported (Figure 5), with the most significant SNPs having effect sizes in the same direction and general magnitude as reported. We note that although most of these associations would be significant after correction for multiple testing (as in a genome-wide association study, for example), two SNPs associated with asthma and breast cancer would not. We recognize that our power to replicate previously validated SNPs is limited in some instances by low case counts, especially when considering SNPs with modest effect sizes, as is consistent with other non-ascertained biobank cohorts.
Figure 5.
Replication of known associations in the CCPM Biobank across a range of traits, comparing CCPM Biobank findings with REGENIE to those found in the GWAS catalog. Error bars represent the 95% confidence interval for the odds ratio
For all, the risk-increasing allele is compared, and with multiple reporters, the largest dataset in the GWAS catalog was used for reference. OMIM identifiers for nearest genes are as follows: APOE (MIM: 107741), SMAD3 (MIM: 603109), TERT (MIM: 187270), HCG22 (MIM: 613918), PTCSC2 (MIM: N/A), HLA-DRA (MIM: 142860), FTO (MIM: 610966), HCP5 (MIM: 604676), HLA-DRB1 (MIM: 142857), HLA-DQB1 (MIM: 604305), and TCF7L2 (MIM: 602228).
Research applications
As of December 2022, ABC has completed 79 service requests including 36 study feasibility assessments, 7 polygenic risk score analyses, and 22 data delivery requests to CCPM researchers and external scientists. Additionally, the CCPM Biobank administered surveys in 2020–2021 to ∼180,000 CCPM Biobank participants to gain information about their experiences during the COVID-19 pandemic including testing, symptoms, healthcare utilization, and impact on health and well-being. More than 25,000 participants responded to the surveys (14% response). These data are linked with clinical data from the EHR and are available for additional research.45
Clinical applications
Nearly 10,000 of the 33,674 genotyped participants have signed a secondary or universal consent to receive clinical genetic test results.
As of May 1, 2022, we have completed a pilot study for the return of HIPV findings to 24 participants (unpublished data). Six participants have variants in genes that confer increased hereditary breast and ovarian cancer risk, 17 participants have variants in genes that confer increased risk for cardiac-related conditions (e.g., cardiomyopathy, familial hypercholesterolemia), and one participant has a variant associated with increased risk for malignant hyperthermia. Counts of these conditions and their affected genes are found in Table S1. Feedback from our first 24 participants who have received these results has been positive (unpublished data). Most of the participants who were referred for clinical follow-up have either had or scheduled appointments with specialty providers and/or their primary care provider, and the majority have shared their results with family members. We continue to return these types of results as more participants become eligible to receive clinical genetic test results, and we expect in the future that 1%–2% of participants will have HIPV findings.
We performed an additional pilot with 36 participants who were found to be homozygous for the c.845G>A (p.Cys282Tyr) variant in HFE that is associated with an increased risk for hereditary hemochromatosis. For these results, a clinical test report was returned directly to the participants’ EHR without contact by a certified genetic counselor. Of note, 12 of the 36 participants had a prior diagnosis of hereditary hemochromatosis, and 10 of those 12 had prior genetic testing for HFE (unpublished data).
As of December 2022, we preemptively returned CYP2C19 (MIM: 124020), SLCO1B1 (MIM: 604843), and/or DPYD (MIM: 612779) clinical pharmacogenetic results to 16,436 CCPM Biobank participants and their providers, amounting to 47,867 unique genomic indicators in Epic. CDS tools are in operation for 20 medications: ten medications affected by CYP2C19 variants (i.e., clopidogrel, voriconazole, citalopram, escitalopram, proton pump inhibitors [dexlansoprazole, lansoprazole, omeprazole, pantoprazole], brivaracetam, and clobazam), seven statin medications affected by SLCO1B1 variation (i.e., atorvastatin, fluvastatin, lovastatin, pitavastatin, pravastatin, rosuvastatin, simvastatin), and three fluoropyrimidines affected by DPYD variation (i.e., systemic 5-fluorouracil, capecitabine, topical 5-fluorouracil). As of December 2022, 1,456 drug-gene interaction CDS alerts have been triggered for 1,249 unique CCPM Biobank participants. The most common drug-gene interaction alerts are for proton pump inhibitors (54.3%), followed by statins (30.2%), es/citalopram (14.1%), clopidogrel (1.3%), and systemic 5-fluorouracil (0.1%). An in-depth analysis of the relationship between drug-gene interaction alerts and subsequent clinical actions is underway. We expect nearly 100% of biobank participants to have at least one actionable pharmacogenomic variant in the future.
Discussion
The dual research and clinical mission of the CCPM Biobank provides unique opportunities for researchers, clinicians, and participants to be involved in and contribute to translational research that informs healthcare. The CCPM Biobank resource presently contains clinical data from the EHR for more than 200,000 participants and dense array genotyping for nearly 34,000 participants, and this resource continues to grow steadily. The convenience of our self-consent model via the Epic patient portal (My Health Connection) allows us to reach patients across the entire UCHealth system and the ability to tailor educational messages in both English and Spanish. Participants provide broad consent to use collected samples and associated EHR data in wide-ranging health related research, including various -omics domains. They also consent to future contact for other studies, which serves as an invaluable resource for new study recruitment. We also screen for clinically important genetic findings including high-impact pathogenic variants, hereditary hemochromatosis, and pharmacogenomic results. Returning these results for use in the clinical setting provides an important pathway to improve health and responsibly advance the frontiers of personalized medicine.
Research efforts in CCPM are supported by rich clinical data from the EHR, data from outside sources such as the Colorado All Payers Claims Database (APCD) (see web resources) and Colorado Department of Public Health and Environment (CDPHE), and the ability to include geocoded environmental information, which are all integrated into the robust cloud-based health data warehouse. These resources, in addition to the use of standard data formats, such as the OMOP common data model13 used by the All of Us Research Program46 and other large initiatives, enable the CCPM Biobank to participate in collaborative science and contribute to other large-scale research efforts, such as the recently assembled Global Biobank Meta-Analysis47 and COVID-19 Host Genetics48 Initiatives. The CCPM consent allows external partnerships, leading us to develop industry relationships supporting advanced computation (we were one of the first health data warehouses in the cloud) and sequencing data acquisition.
Through clinical return of results, the CCPM Biobank operates as a genomic learning health system. Nearly 60% of our participants have provided consent for return of clinical genetic test results, and this percentage will increase over time as more participants sign the unified consent for both research and return of results. We follow a continuous improvement approach in our clinical return of results implementation. For example, our pharmacogenetic program is grounded in the Practical, Robust Implementation and Sustainability Model (PRISM).19,49,50 We monitor the frequency and type of CDS alerts triggered in clinical practice and associated clinical actions (e.g., medication changes) on a weekly basis through automated reports and manual chart reviews. This systematic operational evaluation allows for nearly real-time assessment of CDS effectiveness and an opportunity to provide additional interventions, resulting in the initiation of weekly pharmacogenetic clinical rounds where clinical pharmacists and a physician review high-risk pharmacogenetic cases. In these situations, the team determines whether an in-basket message (i.e., asynchronous CDS) to the provider providing additional pharmacogenetic information is warranted, with the ultimate goal of enhancing the quality of patient care. As advancements are made in gene therapies and other genomic technologies, variants with limited to no clinical evidence today may be returned at a future time when clinical utility has been demonstrated.
The wide support among CCPM participants to receive genetic testing presents opportunities to further understand the potential impacts of population genomic screening. Screening of clinically actionable variants before onset of symptoms can improve diagnosis and treatment outcomes.51,52 However, there can be unintended consequences including increased out-of-pocket healthcare costs to receive recommended interventions, such as imaging, monitoring, and blood tests. Increased screening for rare conditions can also lead to negative impacts for those who receive false positive results. While our approach minimizes technical false positives through confirmatory, clinically validated laboratory-developed tests, expressivity may be more variable and penetrance lower in the general population than estimates derived from kindreds segregating disease. Further, we note that there are limits to our understanding and capture of pathogenic variants in diverse populations. CCPM attempts to partially mitigate this through the use of the MEGA array which was designed for use in diverse populations. Future efforts that leverage non-biased variant ascertainment (e.g., next-generation sequencing) will enable CCPM to better fulfill its clinical mission for individuals of all ancestries.
The genetic diversity in large-scale biobanks and the CCPM Biobank is of increasing importance to fully understand population health disparities and improve personalized medicine applications within the Rocky Mountain region and elsewhere. In the CCPM Biobank, as is common elsewhere (e.g., Belbin et al.53), EHR Race/Ethnicity imperfectly recapitulates genetic ancestry and may be sub-optimal in many genetic analyses. However, genetic ancestry analyses reveal a diverse spectrum of ancestries present in the CCPM Biobank that reflect much of worldwide diversity. In comparison to other biobanks,54,55,56 the CCPM Biobank has a substantial proportion of non-European ancestry populations, including ∼10% Hispanic individuals, ∼5% African American individuals, and other historically underrepresented populations in genomics and in population-based biobanks.53,54,55,56 Future research will involve finer-scale characterization of the ancestry of participants, particularly focusing on the unique population demographics of the Rocky Mountain region.
CCPM Biobank participants also come from diverse environments. With the intersection of geocoded data, we plan to monitor the health of the rural participants in the CCPM Biobank, which make up a substantial proportion of biobank participants and UCHealth patients. Rural health remains highly underrepresented in biomedical research,57 particularly in institutional biobanks concentrated in major metropolitan areas. The UCHealth catchment also contains areas with moderate and high altitude (including participants living >10,000 feet above sea level), which provides a unique context that would be rare to observe in most health systems in the United States.
Our approach has some limitations. First, our electronic consent model has limitations in outreach to patient populations who may be less likely to interact with UCHealth via the web portal. While UCHealth has seen overall growth of adoption of the patient portal since 2020 due to the increase in telemedicine, our utilization of an electronic consent does have the potential to increase the digital divide. To address inequities in enrollment, including patients who require consent materials in other languages, we have started to work with community-based research groups to develop tailored materials and enhance our reach to under-represented populations. Second, like most biobanks, there are limitations to phenotypes derived solely from billing codes and other discrete EHR fields. We have developed a system for implementing computational phenotyping algorithms and chart review to reduce phenotypic misclassification. With our access to surveys with our participants, we hope to be able to access relevant covariate information as well as phenotypic outcomes that may not be present, or in limited scope, in traditional EHRs (e.g., sleep disorders, intensity and duration of symptoms during illness).
Overall, the CCPM Biobank is uniquely positioned in the EHR-linked biobanking landscape, due to our size, integrated learning health system design, and unique catchment area. Many similar initiatives have focused predominantly on research,4 or only recently implemented clinical return of results focused on HIPVs with limited pharmacogenomics implementation due to the cost/scale required to perform clinical validation of these common variants.58,59 Our Biobank Laboratory is both CLIA certified and CAP accredited, allowing for direct return of clinically actionable results without the need for orthogonal confirmatory testing or to collect a new clinical sample. Other integrated programs have much smaller enrollment targets (J. Wagner et al., 2022, ASHG, conference) or lack advanced EHR integrations to enable real-time clinical decision support.60,61
Our goal over the next few years is to continue to increase the number of participants with both clinical and genomic data in the CCPM Biobank into the hundreds of thousands, and to optimize return of clinically actionable results to eligible participants. We also strive to expand the reach and utility of our data assets internally and externally by collaborating with colleagues involved in global efforts such as the GBMI (gbmi.org) as well as disease domain-specific efforts such as the PAGE Study (pagestudy.org). Through CCPM-initiated efforts and collaborations with national and international partners facilitated by CCPM resources, we can discover novel insights to inform targeted personalized strategies to improve health outcomes.
Data and code availability
There are restrictions to the availability of the CCPM genetic and EHR datasets due to the sensitive nature of these datasets and HIPAA compliance. Summary data for the GWAS presented here are available at the Colorado Biobank Portal (https://hdc-sandbox-bioengine.uw.r.appspot.com/) (registration and user approval required).
Acknowledgments
We are deeply indebted to the participants who have made the CCPM Biobank a reality, and financial support from UCHealth, Children's Hospital Colorado (CHCO), CU Medicine, the Department and School of Medicine, and the Skaggs School of Pharmacy and Pharmaceutical Sciences. We especially acknowledge Elizabeth Concordia, Jena Hausmann, Donald Elliman, and John Reilly, David A. Schwartz, David Ross, Dan Theodorescu, Ann Thor, and the CCPM Governance Committee for their vision in launching a cross-institutional enterprise focused on delivering personalized medicine at the point of clinical care. The authors credit oversight and guidance from the CCPM Governance Committee, Advisory Committee, the entire Health Data Compass team, and past and present members of the Pharmacogenomics Implementation Committee Colorado (PICColo). Support for title page creation and format was provided by AuthorArranger, a tool developed at the National Cancer Institute.
M.L. and C.R.G. were partially supported by the National Institutes of Health (R01HL151152, R01HG011345, and U01HG011715). K.C.B. was partially supported by the Department Endowed Chair of Medicine, National Institutes of Health (R01AI132476, R01HL104608, R25HL146166, 5UM1AI109565-07, 5U19 AI117673-05), and Regeneron Pharmaceuticals, Inc. (Project #4841-4450-1168.2). L.K.W. was partially supported by the National Institutes of Health (K01LM013088). C.S.G. was partially supported by the National Institutes of Health (R01 HG010067). This project was partially supported by NIH/NCATS Colorado CTSA Grant Number UL1 TR002535 and UM1 TR004399. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
Declaration of interests
K.C.B. owns stock in Tempus and Galatea Bio and is an employee of Oxford Nanopore Technologies. C.R.G. owns stock in 23andMe, Inc.
Published: January 4, 2024
Footnotes
Supplemental information can be found online at https://doi.org/10.1016/j.ajhg.2023.12.001.
Contributor Information
Kathleen C. Barnes, Email: kathleen.barnes@cuanschutz.edu.
Christopher R. Gignoux, Email: chris.gignoux@cuanschutz.edu.
Web resources
albersusa, https://github.com/hrbrmstr/albersusa
APCD (accessed February 12, 2023), https://www.civhc.org/get-data/co-apcd-info/
bigrquery, https://bigrquery.r-dbi.org/
gtsummary, https://CRAN.R-project.org/package=gtsummary
Here, https://here.r-lib.org/
Magrittr, https://magrittr.tidyverse.org/
Methods for De-identification of PHI (accessed August 3, 2021), https://www.hhs.gov/hipaa/for-professionals/privacy/special-topics/de-identification/index.html
OMIM, https://www.omim.org/
TIGER/Line Shapefiles: Zip Code Tabulation Areas (accessed February 12, 2023), https://www.census.gov/cgi-bin/geo/shapefiles/index.php?year=2019&layergroup=ZIP+Code+Tabulation+Areas
Supplemental information
References
- 1.Bycroft C., Freeman C., Petkova D., Band G., Elliott L.T., Sharp K., Motyer A., Vukcevic D., Delaneau O., O'Connell J., et al. The UK Biobank resource with deep phenotyping and genomic data. Nature. 2018;562:203–209. doi: 10.1038/s41586-018-0579-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Carey D.J., Fetterolf S.N., Davis F.D., Faucett W.A., Kirchner H.L., Mirshahi U., Murray M.F., Smelser D.T., Gerhard G.S., Ledbetter D.H. The Geisinger MyCode community health initiative: an electronic health record-linked biobank for precision medicine research. Genet. Med. 2016;18:906–913. doi: 10.1038/gim.2015.187. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Dewey F.E., Murray M.F., Overton J.D., Habegger L., Leader J.B., Fetterolf S.N., O'Dushlaine C., Van Hout C.V., Staples J., Gonzaga-Jauregui C., et al. Distribution and clinical impact of functional variants in 50,726 whole-exome sequences from the DiscovEHR study. Science. 2016;354 doi: 10.1126/science.aaf6814. [DOI] [PubMed] [Google Scholar]
- 4.Roden D.M., Pulley J.M., Basford M.A., Bernard G.R., Clayton E.W., Balser J.R., Masys D.R. Development of a Large-Scale De-Identified DNA Biobank to Enable Personalized Medicine. Clin. Pharmacol. Ther. 2008;84:362–369. doi: 10.1038/clpt.2008.89. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Coors M.E., Westfall N., Zittleman L., Taylor M., Westfall J.M. Translating Biobank Science into Patient-Centered Language. Biopreserv. Biobank. 2018;16:59–63. doi: 10.1089/bio.2017.0089. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.UCHealth . UCHealth. Youtube; 2021. Biobank at the Colorado Center for Personalized Medicine (CCPM)https://www.youtube.com/watch?v=8ij-qNLYUFU [Google Scholar]
- 7.Johnston H.R., Hu Y.-J., Gao J., O’Connor T.D., Abecasis G.R., Wojcik G.L., Gignoux C.R., Gourraud P.A., Lizee A., Hansen M., et al. Identifying tagging SNPs for African specific genetic variation from the African Diaspora Genome. Sci. Rep. 2017;7 doi: 10.1038/srep46398. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Mathias R.A., Taub M.A., Gignoux C.R., Fu W., Musharoff S., O’Connor T.D., Vergara C., Torgerson D.G., Pino-Yanes M., Shringarpure S.S., et al. A continuum of admixture in the Western Hemisphere revealed by the African Diaspora genome. Nat. Commun. 2016;7 doi: 10.1038/ncomms12522. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Wojcik G.L., Graff M., Nishimura K.K., Tao R., Haessler J., Gignoux C.R., Highland H.M., Patel Y.M., Sorokin E.P., Avery C.L., et al. Genetic analyses of diverse populations improves discovery for complex traits. Nature. 2019;570:514–518. doi: 10.1038/s41586-019-1310-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Wojcik G.L., Fuchsberger C., Taliun D., Welch R., Martin A.R., Shringarpure S., Carlson C.S., Abecasis G., Kang H.M., Boehnke M., et al. Imputation-Aware Tag SNP Selection To Improve Power for Large-Scale, Multi-ethnic Association Studies. G3. 2018;8:3255–3267. doi: 10.1534/g3.118.200502. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Bien S.A., Wojcik G.L., Hodonsky C.J., Gignoux C.R., Cheng I., Matise T.C., Peters U., Kenny E.E., North K.E. The Future of Genomic Studies Must Be Globally Representative: Perspectives from PAGE. Annu. Rev. Genomics Hum. Genet. 2019;20:181–200. doi: 10.1146/annurev-genom-091416-035517. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Bien S.A., Wojcik G.L., Zubair N., Gignoux C.R., Martin A.R., Kocarnik J.M., Martin L.W., Buyske S., Haessler J., Walker R.W., et al. Strategies for Enriching Variant Coverage in Candidate Disease Loci on a Multiethnic Genotyping Array. PLoS One. 2016;11 doi: 10.1371/journal.pone.0167758. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Overhage J.M., Ryan P.B., Reich C.G., Hartzema A.G., Stang P.E. Validation of a common data model for active safety surveillance research. J. Am. Med. Inform. Assoc. 2012;19:54–60. doi: 10.1136/amiajnl-2011-000376. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Miller D.T., Lee K., Chung W.K., Gordon A.S., Herman G.E., Klein T.E., Stewart D.R., Amendola L.M., Adelman K., Bale S.J., et al. ACMG SF v3.0 list for reporting of secondary findings in clinical exome and genome sequencing: a policy statement of the American College of Medical Genetics and Genomics (ACMG) Genet. Med. 2021;23:1381–1390. doi: 10.1038/s41436-021-01172-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Miller D.T., Lee K., Gordon A.S., Amendola L.M., Adelman K., Bale S.J., Chung W.K., Gollob M.H., Harrison S.M., Herman G.E., et al. Recommendations for reporting of secondary findings in clinical exome and genome sequencing, 2021 update: a policy statement of the American College of Medical Genetics and Genomics (ACMG) Genet. Med. 2021;23:1391–1398. doi: 10.1038/s41436-021-01171-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Kalia S.S., Adelman K., Bale S.J., Chung W.K., Eng C., Evans J.P., Herman G.E., Hufnagel S.B., Klein T.E., Korf B.R., et al. Recommendations for reporting of secondary findings in clinical exome and genome sequencing, 2016 update (ACMG SF v2.0): a policy statement of the American College of Medical Genetics and Genomics. Genet. Med. 2017;19:249–255. doi: 10.1038/gim.2016.190. [DOI] [PubMed] [Google Scholar]
- 17.Miller D.T., Lee K., Abul-Husn N.S., Amendola L.M., Brothers K., Chung W.K., Gollob M.H., Gordon A.S., Harrison S.M., Hershberger R.E., et al. ACMG SF v3.1 list for reporting of secondary findings in clinical exome and genome sequencing: A policy statement of the American College of Medical Genetics and Genomics (ACMG) Genet. Med. 2022;24:1407–1414. doi: 10.1016/j.gim.2022.04.006. [DOI] [PubMed] [Google Scholar]
- 18.Richards S., Aziz N., Bale S., Bick D., Das S., Gastier-Foster J., Grody W.W., Hegde M., Lyon E., Spector E., et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genet. Med. 2015;17:405–424. doi: 10.1038/gim.2015.30. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Aquilante C.L., Kao D.P., Trinkley K.E., Lin C.-T., Crooks K.R., Hearst E.C., Hess S.J., Kudron E.L., Lee Y.M., Liko I., et al. Clinical implementation of pharmacogenomics via a health system-wide research biobank: the University of Colorado experience. Pharmacogenomics. 2020;21:375–386. doi: 10.2217/pgs-2020-0007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Landrum M.J., Lee J.M., Benson M., Brown G.R., Chao C., Chitipiralla S., Gu B., Hart J., Hoffman D., Jang W., et al. ClinVar: improving access to variant interpretations and supporting evidence. Nucleic Acids Res. 2018;46:D1062–D1067. doi: 10.1093/nar/gkx1153. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Amberger J.S., Bocchini C.A., Schiettecatte F., Scott A.F., Hamosh A. Omim.org: Online Mendelian Inheritance in Man (OMIM®), an online catalog of human genes and genetic disorders. Nucleic Acids Res. 2015;43:D789–D798. doi: 10.1093/nar/gku1205. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Whirl-Carrillo M., McDonagh E.M., Hebert J.M., Gong L., Sangkuhl K., Thorn C.F., Altman R.B., Klein T.E. Pharmacogenomics knowledge for personalized medicine. Clin. Pharmacol. Ther. 2012;92:414–417. doi: 10.1038/clpt.2012.96. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Whirl-Carrillo M., Huddart R., Gong L., Sangkuhl K., Thorn C.F., Whaley R., Klein T.E. An Evidence-Based Framework for Evaluating Pharmacogenomics Knowledge for Personalized Medicine. Clin. Pharmacol. Ther. 2021;110:563–572. doi: 10.1002/cpt.2350. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Hamilton J.G., Abdiwahab E., Edwards H.M., Fang M.-L., Jdayani A., Breslau E.S. Primary care providers’ cancer genetic testing-related knowledge, attitudes, and communication behaviors: A systematic review and research agenda. J. Gen. Intern. Med. 2017;32:315–324. doi: 10.1007/s11606-016-3943-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Owusu Obeng A., Fei K., Levy K.D., Elsey A.R., Pollin T.I., Ramirez A.H., Weitzel K.W., Horowitz C.R. Physician-Reported Benefits and Barriers to Clinical Implementation of Genomic Medicine: A Multi-Site IGNITE-Network Survey. J. Pers. Med. 2018;8 doi: 10.3390/jpm8030024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.White S., Jacobs C., Phillips J. Mainstreaming genetics and genomics: a systematic review of the barriers and facilitators for nurses and physicians in secondary and tertiary care. Genet. Med. 2020;22:1149–1155. doi: 10.1038/s41436-020-0785-6. [DOI] [PubMed] [Google Scholar]
- 27.Wei W.-Q., Bastarache L.A., Carroll R.J., Marlo J.E., Osterman T.J., Gamazon E.R., Cox N.J., Roden D.M., Denny J.C. Evaluating phecodes, clinical classification software, and ICD-9-CM codes for phenome-wide association studies in the electronic health record. PLoS One. 2017;12 doi: 10.1371/journal.pone.0175508. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Wu P., Gifford A., Meng X., Li X., Campbell H., Varley T., Zhao J., Carroll R., Bastarache L., Denny J.C., et al. Mapping ICD-10 and ICD-10-CM Codes to Phecodes: Workflow Development and Initial Evaluation. JMIR Med. Inform. 2019;7 doi: 10.2196/14325. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Carroll R.J., Bastarache L., Denny J.C. PheWAS: data analysis and plotting tools for phenome-wide association studies in the R environment. Bioinformatics. 2014;30:2375–2376. doi: 10.1093/bioinformatics/btu197. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Wickham H., Averick M., Bryan J., Chang W., McGowan L., François R., Grolemund G., Hayes A., Henry L., Hester J., et al. Welcome to the Tidyverse. JOSS. 2019;4:1686. [Google Scholar]
- 31.Pebesma E. Simple Features for R: Standardized Support for Spatial Vector Data. The R Journal. 2018;10:439. [Google Scholar]
- 32.Kahle D., Wickham H. ggmap: Spatial Visualization with ggplot2. The R Journal. 2013;5:144–161. [Google Scholar]
- 33.Grolemund G., Wickham H. Dates and Times Made Easy with lubridate. J. Stat. Softw. 2011;40:1–25. [Google Scholar]
- 34.Ooms J. The jsonlite Package: A Practical and Consistent Mapping Between JSON Data and R Objects.Preprint at arXiv:1403 2805 [Stat CO] 2014.10.48550/arXiv.1403.2805
- 35.Taliun D., Harris D.N., Kessler M.D., Carlson J., Szpiech Z.A., Torres R., Taliun S.A.G., Corvelo A., Gogarten S.M., Kang H.M., et al. Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program. Nature. 2021;590:290–299. doi: 10.1038/s41586-021-03205-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Das S., Forer L., Schönherr S., Sidore C., Locke A.E., Kwong A., Vrieze S.I., Chew E.Y., Levy S., McGue M., et al. Next-generation genotype imputation service and methods. Nat. Genet. 2016;48:1284–1287. doi: 10.1038/ng.3656. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Conomos M.P., Miller M.B., Thornton T.A. Robust inference of population structure for ancestry prediction and correction of stratification in the presence of relatedness. Genet. Epidemiol. 2015;39:276–293. doi: 10.1002/gepi.21896. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Conomos M.P., Reiner A.P., Weir B.S., Thornton T.A. Model-free Estimation of Recent Genetic Relatedness. Am. J. Hum. Genet. 2016;98:127–148. doi: 10.1016/j.ajhg.2015.11.022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Gogarten S.M., Bhangale T., Conomos M.P., Laurie C.A., McHugh C.P., Painter I., Zheng X., Crosslin D.R., Levine D., Lumley T., et al. GWASTools: an R/Bioconductor package for quality control and analysis of genome-wide association studies. Bioinformatics. 2012;28:3329–3331. doi: 10.1093/bioinformatics/bts610. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Manichaikul A., Mychaleckyj J.C., Rich S.S., Daly K., Sale M., Chen W.-M. Robust relationship inference in genome-wide association studies. Bioinformatics. 2010;26:2867–2873. doi: 10.1093/bioinformatics/btq559. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Mbatchou J., Barnard L., Backman J., Marcketta A., Kosmicki J.A., Ziyatdinov A., Benner C., O'Dushlaine C., Barber M., Boutkov B., et al. Computationally efficient whole-genome regression for quantitative and binary traits. Nat. Genet. 2021;53:1097–1103. doi: 10.1038/s41588-021-00870-7. [DOI] [PubMed] [Google Scholar]
- 42.Devlin B., Roeder K. Genomic control for association studies. Biometrics. 1999;55:997–1004. doi: 10.1111/j.0006-341x.1999.00997.x. [DOI] [PubMed] [Google Scholar]
- 43.Buniello A., MacArthur J.A.L., Cerezo M., Harris L.W., Hayhurst J., Malangone C., McMahon A., Morales J., Mountjoy E., Sollis E., et al. The NHGRI-EBI GWAS Catalog of published genome-wide association studies, targeted arrays and summary statistics 2019. Nucleic Acids Res. 2019;47:D1005–D1012. doi: 10.1093/nar/gky1120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Williams C.M., Scelza B.A., Daya M., Lange E.M., Gignoux C.R., Henn B.M. A rapid, accurate approach to inferring pedigrees in endogamous populations. bioRxiv. 2020 Preprint at. [Google Scholar]
- 45.Johnson R.K., Marker K.M., Mayer D., Shortt J., Kao D., Barnes K.C., Lowery J.T., Gignoux C.R. COVID-19 surveillance in the Biobank at the Colorado Center for Personalized Medicine: Observational study. JMIR Public Health Surveill. 2022;8 doi: 10.2196/37327. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Klann J.G., Joss M.A.H., Embree K., Murphy S.N. Data model harmonization for the All Of Us Research Program: Transforming i2b2 data into the OMOP common data model. PLoS One. 2019;14 doi: 10.1371/journal.pone.0212463. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Zhou W., Kanai M., Wu K.-H.H., Rasheed H., Tsuo K., Hirbo J.B., Wang Y., Bhattacharya A., Zhao H., Namba S., et al. Global Biobank Meta-analysis Initiative: Powering genetic discovery across human disease. Cell Genom. 2022;2 doi: 10.1016/j.xgen.2022.100192. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Niemi M.E.K., Karjalainen J., Liao R.G., Neale B.M., Daly M., Ganna A., Pathak G.A., Andrews S.J., Kanai M., Veerapen K., et al. Mapping the human genetic architecture of COVID-19. Nature. 2021;600:472–477. doi: 10.1038/s41586-021-03767-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Feldstein A.C., Glasgow R.E. A practical, robust implementation and sustainability model (PRISM) for integrating research findings into practice. Jt. Comm. J. Qual. Patient Saf. 2008;34:228–243. doi: 10.1016/s1553-7250(08)34030-6. [DOI] [PubMed] [Google Scholar]
- 50.Bates D.W., Kuperman G.J., Wang S., Gandhi T., Kittler A., Volk L., Spurr C., Khorasani R., Tanasijevic M., Middleton B. Ten commandments for effective clinical decision support: making the practice of evidence-based medicine a reality. J. Am. Med. Inform. Assoc. 2003;10:523–530. doi: 10.1197/jamia.M1370. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Zhang L., Bao Y., Riaz M., Tiller J., Liew D., Zhuang X., Amor D.J., Huq A., Petelin L., Nelson M., et al. Population genomic screening of all young adults in a health-care system: a cost-effectiveness analysis. Genet. Med. 2019;21:1958–1968. doi: 10.1038/s41436-019-0457-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Adams M.C., Evans J.P., Henderson G.E., Berg J.S. The promise and peril of genomic screening in the general population. Genet. Med. 2016;18:593–599. doi: 10.1038/gim.2015.136. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Belbin G.M., Cullina S., Wenric S., Soper E.R., Glicksberg B.S., Torre D., Moscati A., Wojcik G.L., Shemirani R., Beckmann N.D., et al. Toward a fine-scale population health monitoring system. Cell. 2021;184:2068–2083.e11. doi: 10.1016/j.cell.2021.03.034. [DOI] [PubMed] [Google Scholar]
- 54.Abul-Husn N.S., Kenny E.E. Personalized Medicine and the Power of Electronic Health Records. Cell. 2019;177:58–69. doi: 10.1016/j.cell.2019.02.039. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Staples J., Maxwell E.K., Gosalia N., Gonzaga-Jauregui C., Snyder C., Hawes A., Penn J., Ulloa R., Bai X., Lopez A.E., et al. Profiling and Leveraging Relatedness in a Precision Medicine Cohort of 92,455 Exomes. Am. J. Hum. Genet. 2018;102:874–889. doi: 10.1016/j.ajhg.2018.03.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Ritchie M.D., Denny J.C., Crawford D.C., Ramirez A.H., Weiner J.B., Pulley J.M., Basford M.A., Brown-Gentry K., Balser J.R., Masys D.R., et al. Robust Replication of Genotype-Phenotype Associations across Multiple Diseases in an Electronic Medical Record. Am. J. Hum. Genet. 2010;87:310. doi: 10.1016/j.ajhg.2010.03.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Committee on Ways and Means, US House of Representatives . 2020. Left Out: Barriers to Health Equity for Rural and Underserved Communities.https://democrats-waysandmeans.house.gov/sites/evo-subsites/democrats-waysandmeans.house.gov/files/documents/WMD%20Health%20Equity%20Report_07.2020_FINAL.pdf [Google Scholar]
- 58.Abul-Husn N.S., Soper E.R., Braganza G.T., Rodriguez J.E., Zeid N., Cullina S., Bobo D., Moscati A., Merkelson A., Loos R.J.F., et al. Implementing genomic screening in diverse populations. Genome Med. 2021;13:17. doi: 10.1186/s13073-021-00832-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Uber R., Wright E. Pharmacogenomics implementation and multidisciplinary genomics collaboration: Real-world experience from Geisinger. Am. J. Health Syst. Pharm. 2022;79:1038–1041. doi: 10.1093/ajhp/zxac065. [DOI] [PubMed] [Google Scholar]
- 60.May T., Cannon A., Moss I.P., Nakano-Okuno M., Hardy S., Miskell E.L., Kelley W.V., Curry W., East K.M., Acemgil A., et al. Recruiting diversity where it exists: The Alabama Genomic Health Initiative. J. Genet. Couns. 2020;29:471–478. doi: 10.1002/jgc4.1258. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Limdi N.A., Absher D., Asif I., Bateman L., Barsh G., Bowling K.M., Cooper G.M., Davis B.H., East K.M., Finnila C.R., et al. 338 The Alabama Genomic Health Initiative: Integrating Genomic Medicine into Primary Care. J. Clin. Transl. Sci. 2023;7:100–101. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
There are restrictions to the availability of the CCPM genetic and EHR datasets due to the sensitive nature of these datasets and HIPAA compliance. Summary data for the GWAS presented here are available at the Colorado Biobank Portal (https://hdc-sandbox-bioengine.uw.r.appspot.com/) (registration and user approval required).





