Skip to main content
AACR Open Access logoLink to AACR Open Access
. 2024 Dec 3;34(2):234–245. doi: 10.1158/1055-9965.EPI-24-1247

Evaluation of Multiple Breast Cancer Polygenic Risk Score Panels in Women of Latin American Heritage

Xiaosong Huang 1,2,#, Paul C Lott 2,#, Donglei Hu 3, Valentina A Zavala 1,2, Zoeb N Jamal 2, Tatiana Vidaurre 4, Sandro Casavilca-Zambrano 4, Jeannie Navarro Vásquez 4, Carlos A Castañeda 4, Guillermo Valencia 4, Zaida Morante 4, Mónica Calderón 4, Julio E Abugattas 4, Hugo A Fuentes 4, Ruddy Liendo-Picoaga 4, Jose M Cotrina 4, Silvia P Neciosup 4, Patricia Rioja Viera 4, Luis A Salinas 4, Marco Galvez-Nino 4, Scott Huntsman 3, Sixto E Sanchez 5, Michelle A Williams 6, Bizu Gelaye 6,7, Ana P Estrada-Florez 2,8, Guadalupe Polanco-Echeverry 2, Magdalena Echeverry 8, Alejandro Velez 9,10, Jenny A Carmona-Valencia 11, Mabel E Bohorquez-Lozano 8, Javier Torres 12, Miguel Cruz 13, Weang-Kee Ho 14,15, Soo Hwang Teo 15,16, Mei Chee Tai 15, Esther M John 17,18,19, Christopher A Haiman 20,21, David V Conti 20,21, Fei Chen 20,21, Gabriela Torres-Mejía 22, Lawrence H Kushi 23, Susan L Neuhausen 24, Elad Ziv 3,‡, Luis G Carvajal-Carmona 2,25,26,‡; for the COLUMBUS Consortium, Laura Fejerman 1,2,25,*,‡
PMCID: PMC11799839  PMID: 39625644

Abstract

Background:

A substantial portion of the genetic predisposition for breast cancer is explained by multiple common genetic variants of relatively small effect. A subset of these variants, which have been identified mostly in individuals of European (EUR) and Asian ancestries, have been combined to construct a polygenic risk score (PRS) to predict breast cancer risk, but the prediction accuracy of existing PRSs in Hispanic/Latinx individuals (H/L) remain relatively low. We assessed the performance of several existing PRS panels with and without addition of H/L-specific variants among self-reported H/L women.

Methods:

PRS performance was evaluated using multivariable logistic regression and the area under the ROC curve.

Results:

Both EUR and Asian PRSs performed worse in H/L samples compared with original reports. The best EUR PRS performed better than the best Asian PRS in pooled H/L samples. EUR PRSs had decreased performance with increasing Indigenous American (IA) ancestry, while Asian PRSs had increased performance with increasing IA ancestry. The addition of two H/L SNPs increased performance for all PRSs, most notably in the samples with high IA ancestry, and did not impact the performance of PRSs in individuals with lower IA ancestry.

Conclusions:

A single PRS that incorporates risk variants relevant to the multiple ancestral components of individuals from Latin America, instead of a set of ancestry-specific panels, could be used in clinical practice.

Impact:

The results highlight the importance of population-specific discovery and suggest a straightforward approach to integrate ancestry-specific variants into PRSs for clinical application.

Introduction

Breast cancer is the most common cancer in women in the Unites States and worldwide (1). The risk of developing breast cancer is determined by genetic and environmental factors, with approximately 30% of the total variability attributed to the former (2). Whereas genetic factors include both variants in high-penetrance genes and common variants, a substantial portion of the genetic susceptibility for breast cancer is explained by multiple common genetic variants of relatively small effect (3, 4). Through genome-wide association studies (GWAS) conducted mostly in European (EUR) populations, more than 300 common genetic variants associated with breast cancer have been discovered (5–12). Even though the effect of each variant is relatively small, multiple variants can be combined into a polygenic risk score (PRS) that can be used to stratify women into different levels of risk of developing breast cancer (3, 4, 13, 14).

Most PRSs for breast cancer have been developed specifically in women of EUR ancestry (6, 15). PRSs developed by the Breast Cancer Association Consortium (BCAC) include 313-variant and 3,820-variant models and showed good performance in EUR studies, with an area under the ROC curve of 0.63 to 0.64 (4). Because the GWAS and subsequent development of PRSs has been done mostly using data from individuals of EUR ancestry, the performance of these models in non-EUR populations has been comparably less robust (16, 17).

Recently, there have been efforts to develop breast cancer PRSs for non-EUR populations. PRSs developed for Asian women have had limited success (13, 18–21). The best performing Asian PRSs include a 46-variant and 2,985-variant model with AUC values of 0.60 and 0.61, respectively (13). In the same study, linear combinations of the Asian and EUR PRSs had markedly higher performance in Asian samples compared with the Asian PRSs alone but modestly higher performance compared with the EUR PRSs alone (13). A recent study developed a joint PRS by combining the 313-variant EUR PRS and an African (AFR)-specific PRS with modest performance in individuals of AFR ancestry (AUC = 0.577; ref. 22).

Hispanic/Latino (H/L) communities are comprised of genetically admixed individuals who have varying degrees of Indigenous American (IA), EUR, AFR, and Asian ancestries (23). There have been limited efforts in developing H/L-specific PRSs for breast cancer. We previously developed and tested a 180-variant PRS that comprised 178 variants discovered in an EUR GWAS and 2 variants (rs140068132 and rs3778609) discovered in a H/L GWAS (11, 14). This PRS performed well (AUC = 0.63) in H/L studies when compared with EURs, but the analysis did not test the extent to which adding the H/L variants improved performance in H/L individuals.

A unified PRS applicable to individuals from diverse ancestral backgrounds termed multiple-ancestry PRS (MA-PRS) has been previously published (24). The MA-PRS combined ancestry-specific PRSs to a single PRS based on genetic ancestral composition of three reference ancestries [AFR, East Asian (EAS), and EUR]. This unified model had modest or no performance improvement in non-EUR populations. There was no analysis of performance in stratified populations based on different countries or different proportions of genetic ancestry.

In this study, we assessed the latest PRS panels developed in populations of EUR and Asian ancestries in combined analyses of multiple H/L studies from different countries and with different proportions of genetic ancestries. We evaluated the performance of panels with or without H/L-specific variants, and we evaluated the performance of different PRS panels in H/L women of different genetic ancestries.

Materials and Methods

Study participants

We included 18,444 women (5,697 breast cancer cases and 12,747 controls) self-identified as H/L from seven studies as described below. All studies obtained local Institutional Review Board approval and written informed consent from participants.

The Peruvian Genetics and Genomics of Breast Cancer Study (PEGEN-BC) is a hospital-based cohort study. As of August 2022, we have recruited 2,089 participants from the Instituto Nacional de Enfermedades Neoplasicas (INEN) in Lima, Peru. Women were invited to participate if they had a diagnosis of invasive breast cancer in 2010 or later and were between 21 and 79 years of age when diagnosed. A blood sample was drawn by a certified phlebotomist at the INEN central laboratory. The present report includes analyses with a subset of 1,811 unrelated patients with available genome-wide genotype data. This study was approved by the institutional review boards of the INEN and the University of California, Davis. All individuals provided written informed consent to participate. The Pregnancy Outcomes, Maternal and Infant Cohort Study (PrOMIS) study recruited 3,347 women ages 18 or older who attended prenatal care clinics at the Instituto Nacional Materno Perinatal in Lima, Peru (25). The 3,339 unrelated PrOMIS participants with available genome-wide genotype data were included in the present study as convenience controls to be compared with the PEGEN-BC study participants in association analyses. The study was approved by the institutional review boards of the Instituto Nacional Materno Perinatal and the Office of Human Research Administration, Harvard T.H. Chan School of Public Health (Boston, MA). All participants provided written informed consent (26, 27).

We combined samples from the San Francisco Bay Area Breast Cancer Study (SFBCS) and the Northern California Breast Cancer Family Registry (NC-BCFR) study to create the SFBCS/NC-BCFR dataset in our analysis because these two studies recruited from the same geographic area and within similar time frames. The SFBCS is a population-based case–control study for breast cancer (28). Self-identified H/L women newly diagnosed with breast from 1995 to 2002 cancer were identified through the Greater Bay Area Cancer Registry. Controls matched on race/ethnicity and 5-year age group were identified using random-digit dialing. The NC-BCFR is a population-based family study and one of six sites of the Breast Cancer Family Registry (29). The Greater Bay Area Cancer Registry was used to identify self-identified H/L women newly diagnosed with breast cancer from 1995 to 2009. Controls matched on race/ethnicity and 5-year age group were identified through random-digit dialing. After data cleaning and removing relatives, the SFBCS/NC-BCFR dataset has 942 cases and 699 controls.

The Cancer de Mama study is a population-based case–control study of breast cancer in Mexico City, Monterrey, and Veracruz, Mexico. It includes 709 cases diagnosed at hospitals in the study region from 2005 to 2007, ages 35 to 69 years, and 702 controls matched by frequency in 5-year age groups (11, 30).

The Multiethnic Cohort (MEC) study is a prospective cohort study in Los Angeles County, California, and Hawaii. For the current analysis, we included a total of 520 self-identified H/L women ages 45 years or above with breast cancer as cases and 2,491 unaffected women matched on sex, age, and ethnicity as controls (31).

We analyzed 225 H/L breast cancer cases and 3,576 H/L controls from the Genetic Epidemiology Research on Adult Health and Aging cohort (RRID: SCR_010472) which is a genotyped subset of the Kaiser Permanente Northern California Research Project on Genes, Environment, and Health (RPGEH) study. Kaiser-RPGEH is a multiracial/multiethnic biobank of consenting members of the Kaiser Permanente Health Plan (32).

The Post-Columbian Study of Environmental and Heritable Causes of Breast Cancer (COLUMBUS) is a case–control study of breast cancer with individuals recruited throughout large cancer hospitals in central and southern Colombia and Mexico City. We divided COLUMBUS participants into two studies, namely, COLUMBUS-Colombia and COLUMBUS-Mexico, based on recruiting regions. In the COLUMBUS-Colombia study, we included 955 breast cancer cases diagnosed from 2011 to 2017, ages 21 to 88 years, and 1,053 controls who did not have first-degree relatives with breast cancer and matched controls with cases by recruiting hospital, education level, socioeconomic status, and local origin data. In the COLUMBUS-Mexico study, we included 535 cases diagnosed with breast cancer from 2010 to 2018 at the Social Security Hospital in Mexico City, ages 24 to 89 years, and 887 controls matched by recruiting hospital (11).

Genotyping, imputation, and genetic ancestry

Genotyping and imputation methods for studies included in this analysis were described elsewhere (12, 14, 27) and below. Genetic ancestry estimates for our subjects were available from previous studies (14, 26) and described below.

We performed genotyping using the following arrays: Affymetrix (RRID: SCR_007817) 6.0 for the SFBCS/NC-BCFR samples; Affymetrix LAT for the Kaiser Permanente RPGEH samples; Illumina 660 W-Quad for the MEC cases and Illumina 2.5M for the MEC controls; Illumina OncoArray for the Cancer de Mama samples; Affymetrix Axiom Genome-Wide LAT and Custom for the COLUMBUS controls, Axiom UK Biobank for the COLUMBUS Colombian cases, and Affymetrix Axiom Precision Medicine Research Array for the COLUMBUS Mexican cases; Affymetrix Axiom Precision Medicine Research Array for the PEGEN-BC cases, and Illumina Multi-Ethnic Global Array for the PEGEN-BC controls.

The genotyping quality control and imputation for SFBCS/NC-BCFR, Kaiser-RPGEH, and MEC datasets were described in a previous publication (12).

The genotyping quality control and imputation for the PEGEN-BC dataset were described in (27). We assessed the 1000 Genomes (RRID: SCR_006828) and the TOPMed (RRID: SCR_015677) reference panels for imputation and found that using 1000 Genomes for imputation yielded a higher proportion of overlapping polymorphisms with available PRS panels, so we report in this study the results from the dataset imputed with the 1000 Genomes reference panel on the Michigan Imputation Server (RRID: SCR_023554; ref. 33).

The genotyping quality control for COLUMBUS-Colombia and COLUMBUS-Mexico studies were described previously (14). The genotype imputation was performed using the Michigan Imputation Server (using 1000 Genomes as reference panel) or TOPMed imputation server (using TOPMed as reference panel; refs. 33–36). We retrieved genotype data for polymorphisms in the PRS panels from imputation with 1000 Genomes if available. We retrieved genotype data for any remaining polymorphisms in the PRS panels from imputation with the TOPMed panel if available.

Genetic ancestry were estimated from genome-wide genotype data of study participants merged with data from four reference populations from the 1000 Genomes project (34): admixed Americans (Peru, Colombia, Mexico, and Puerto Rico), EURs (Americans with northern and western EUR ancestries and those from Italy, Spain, Finland, England, and Scotland), EASs (China, Japan, and Vietnam), and AFR populations (Nigeria, Kenya, Gambia, and Sierra Leone). Individual continental, global genetic ancestry was estimated using ADMIXTURE (RRID: SCR_001263; ref. 37; unsupervised, k = 4).

Assessed PRSs

We evaluated five PRSs previously developed and tested in women with EUR and Asian ancestries (Supplementary Table S1). PRS313 is a 313-variant PRS developed by the BCAC in EUR populations using P value and linkage disequilibrium (LD) hard thresholding followed by a stepwise forward selection (4). We found 295 variants from this set in our dataset, so it was renamed PRS295. PRS178 included 178 variants that reached genome-wide statistical significance in GWASs in EUR and Asian populations (14). We found 164 variants from this set in our dataset, so it was renamed PRS164. PRS3820 is a 3820-variant PRS developed by the BCAC in EUR populations using Least Absolute Shrinkage and Selection Operator (LASSO) penalized regression (4). We found 3,503 variants from this set in our dataset, so it was renamed PRS3503. PRS46 is a 46-variant PRS developed in Asian populations using a clumping and thresholding method (13). We found 43 variants from this set in our dataset, so it was renamed PRS43. PRS2985 is a 2,985-variant PRS developed in Asian populations using LASSO penalized regression method (13). We found 2,750 variants from this set in our dataset, so it was renamed PRS2750. We constructed additional five modified PRSs by supplementing original models with two variants (rs140068132 and rs3778609) discovered in GWASs conducted in H/L individuals, and the modified models were named PRS297, PRS166, PRS3505, PRS45, and PRS2752 correspondingly. We included all SNPs in our analysis regardless of imputation quality because we found there were no substantive differences in the associations with breast cancer based on PRS panels (PRS3822 as an example) constructed with different imputation r2 cutoffs (Supplementary Table S2).

Statistical analyses

For each person i, a PRS was constructed by the following equation:

PRSi = β1x1 + β2x2 + … + βkxk + … + βmxm

In which xk is the risk allele dosage at variant k, βk is the corresponding weight, and m is the total number of risk variants. For EUR and Asian PRSs, we used weights from the original publications (3, 13) for several reasons. First, there has not been another large and independent H/L GWAS study, and many samples used in this PRS evaluation study were used in an original H/L GWAS; to avoid overfitting, we decided to not use the effect sizes in the H/L GWAS as weights. Second, the original weights should be more accurate in the absence of heterogeneity by ancestry because they are derived from studies with larger sample sizes. Third, it has been shown that weights from other populations work well in H/L samples (14). For the two H/L variants supplemented to original PRSs, we used weights from a previous publication (14). As there could be overfitting for these particular SNPs, we also conducted analyses (included in the supplementary materials section) excluding samples that were part of the original GWAS.

We first tested the adjusted association of PRS with breast cancer risk in all samples and samples stratified by country. We adjusted for study, and IA and AFR genetic ancestries were assessed in our analysis with the following procedure. A linear regression of the study and genetic ancestry on PRS (as dependent variable) was performed using controls within each country. The resultant linear model was then used to derive predicted PRSs for cases and controls within each country. Then we derived residuals as the difference between the actual PRS and the predicted PRS for each sample. Residuals were then normalized to the mean and SD of controls. We used the standardized residuals as main predictors in univariate logistic regression with breast cancer as the outcome. We estimated ORs per unit SD and the area under the ROC curves to compare with results from previous studies. We tested calibration using the Hosmer–Lemeshow test across deciles of the adjusted PRSs, with the 40th to 50th and 50th to 60th deciles combined as the reference group.

To evaluate the ancestry-specific performance of PRS, we divided the pooled dataset into ranges of IA ancestry: 0 to 0.4, 0.4 to 0.6, and 0.6 to 1. Ranges were chosen to include large number of observations in each group while representing clearly differentiated ranges of low, intermediate, and high IA ancestry (Supplementary Table S3). We performed logistic regression using the same adjustment procedure as above but only adjusting for study. We tested the difference of coefficients (log odds ratio) using z-test and the difference of AUCs using the DeLong test (38).

All tests for statistical significance used a two-sided alpha of 0.05. We developed the script to calculate PRSs and statistical analysis using R programming language (version 4.3.1, RRID: SCR_001905) and a specific package for AUC analysis (pROC, v1.18.4, RRID: SCR_024286; ref. 39).

Data availability

The datasets generated and/or analyzed during the current study are available from the corresponding author upon request.

Results

Study participant characteristics

Our pooled samples for analysis included 18,444 H/L women from seven studies recruited from four countries (Colombia, Mexico, Peru, and United States) for a total of 5,697 cases and 12,747 controls (Table 1). The mean age at diagnosis for cases ranged from 50 to 66 years across studies, and for controls, the mean age at enrollment ranged from 28 to 64 years. Genetic ancestry was predominantly EUR and IA and varied within and between studies (Supplementary Fig. S1). For instance, the PEGEN-BC study had the highest average IA ancestry (75.6% in cases and 83% in controls), whereas the Kaiser-PRGEH study had the lowest average IA ancestry (29.9% in cases and 31.1% in controls).

Table 1.

Participant characteristics by study and case–control status.

COLUMBUS (Colombia) COLUMBUS (Mexico) Cancer de Mama (Mexico) PEGEN-BC (Peru) Kaiser (United States) MEC (United States) SFBCS/NC-BCFR (United States)
Characteristic Controls Cases Controls Cases Controls Cases Controls Cases Controls Cases Controls Cases Controls Cases
# of individuals 1,053 955 887 535 702 709 3,339 1,811 3,576 225 2,491 520 699 942
Age at diagnosis (cases) or interview (controls) in years, mean (SD) 64 (10) 52.5 (10.9) 35 (12) 56.5 (12.8) 52.2 (9.1) 52.4 (9.8) 28.1 (6.3) 49.8 (10.9) 54.7 (13.1) 62.6 (10.6) 66.4 (8.1) 66 (8) 49.3 (14.1) 50.6 (11.2)
IA ancestry, mean % (SD) 41.6 (13.1) 41.2 (11.3) 64.9 (17) 64.4 (17.4) 67.7 (18.4) 62.9 (18.5) 83 (18.5) 75.6 (16.8) 31.1 (18) 29.9 (17.4) 40.9 (13.9) 39.5 (13.3) 43.1 (16.7) 38.9 (17.5)
EUR ancestry, mean % (SD) 50.4 (12.4) 50.9 (10.8) 29.1 (15.3) 29.6 (15.7) 29.4 (17) 34.2 (17.8) 13.1 (9.8) 18 (12.5) 63.3 (19.2) 65.3 (18.4) 55.1 (14.6) 56.2 (13.9) 52 (16.9) 56 (17.9)
AFR ancestry, mean % (SD) 7.2 (6.3) 7.1 (4.6) 5 (1.9) 5 (1.8) 2.9 (3) 3 (2.8) 2.7 (4.9) 4.7 (7.6) 5.6 (8.3) 4.8 (4.4) 4 (2.3) 4.3 (3.5) 5 (7.3) 5.1 (7.1)
EAS ancestry, mean % (SD) 0.6 (0.9) 0.6 (0.8) 0.7 (1.1) 0.8 (1.2) NA NA 1.3 (2.1) 1.9 (5.7) NA NA NA NA NA NA

Evaluating EUR PRSs in H/L populations

We evaluated the performance of three previously published PRSs developed using EUR data in the pooled H/L samples. Their development and derivation are described elsewhere (4, 14) and briefly in the methods section.

We evaluated the association between individual variants of EUR PRSs and breast cancer risk in H/L samples using multivariable logistic regression models adjusted for genetic ancestry and study. Of the 3,597 variants in EUR PRSs, 497 were nominally statistically significant (P < 0.05), with 340 also directionally consistent. Fourteen variants remained genome-wide statistically significant (P value ≤ 5 × 10−8). Of the 3,597 variants, 1,936 had associations that were directionally consistent with those reported in EUR populations (Supplementary Table S4).

The performance of these three PRSs were similar to each other in EUR ancestry populations (the reported AUC is between 0.63 and 0.64; refs. 4, 14), with PRS3820 slightly better than the other two. We first evaluated the performance of the three EUR panels in our total pooled H/L samples from all participating studies adjusting for study and global ancestry. Our results showed that the performance of all three panels diminished in the H/L combined set (AUC between 0.606 and 0.611) compared with the performance in the EUR studies (Fig. 1A, E; Supplementary Table S5). For instance, PRS3503 had an AUC of 0.611 [95% confidence interval (CI), 0.602–0.62] in the H/L studies compared with an AUC of 0.636 in the EUR studies. Among the three PRSs, PRS3503 showed slightly greater OR and AUC values than the other two panels. We use PRS3503 to represent the best EUR PRS in most of our subsequent analyses. Women in the top 10% of the PRS3503 had a 2.07-fold (95% CI, 1.84–2.32) elevated risk compared with women at average score (PRS in the 40–60th percentile; Fig. 2A). The calibration curve and Hosmer–Lemeshow test suggested good fit for all three EUR panels (Fig. 3A–C; Supplementary Table S6). Next, we evaluated the performance of EUR PRSs in samples from different countries and found that the performance of EUR PRSs varied by country in the H/L samples (Fig. 1A; Supplementary Table S5). For example, PRS3503 performed better in Colombia (AUC = 0.626, 95% CI, 0.601–0.65) than in Peru (AUC = 0.598, 95% CI, 0.581–0.614), although the AUC difference was not statistically significant (DeLong test P value = 0.0596). The average EUR ancestry of the Colombia samples (0.51) is higher than the average EUR ancestry of the Peruvian samples (0.15), and the average IA ancestry of the Colombia samples (0.41) is lower than the average IA ancestry of the Peruvian samples (0.80), suggesting that the difference in genetic ancestry among populations might affect the performance of the PRSs.

Figure 1.

Figure 1.

Association between PRSs from three EUR SNP panels (and modified panels with addition of two H/L SNPs) and breast cancer risk in H/L samples stratified by countries or IA genetic ancestry. Area under the ROC curve (AUC) or OR were calculated according to Statistical analyses in Materials and Methods. A, AUC of EUR PRS panels tested in H/L samples stratified by countries, (B) AUC of EUR PRS panels modified with addition of two H/L SNPs tested in H/L samples stratified by countries, (C) AUC of EUR PRS panels tested in H/L samples stratified by IA ancestry ranges, (D) AUC of EUR PRS panels modified with addition of two H/L SNPs tested in H/L samples stratified by IA ancestry ranges, (E) OR of EUR PRS panels tested in H/L samples stratified by countries, (F) OR of EUR PRS panels modified with addition of two H/L SNPs tested in H/L samples stratified by countries, (G) OR of EUR PRS panels tested in H/L samples stratified by IA ancestry ranges, and (H) OR of EUR PRS panels modified with addition of two H/L SNPs tested in H/L samples stratified by IA ancestry ranges. Error bars represent the 95% CI.

Figure 2.

Figure 2.

Association between PRSs from tested PRS panels with overall breast cancer risk in H/L samples relative to the middle quantile (40%–60%). A, the three EUR PRS panels, (B) the three PRS panels modified from EUR panels (with addition of two H/L SNPs), (C) the two Asian PRS panels, and (D) the two PRS panels modified from Asian panels (with addition of two H/L SNPs). Error bars represent the 95% CI.

Figure 3.

Figure 3.

Calibration of PRSs from tested PRS panels. The graph depicts the predicted vs. observed proportions of cases within each decile of the log-normalized PRS. Calibration plots for (A, B, and C) the three EUR PRS panels and the three PRS panels modified from the EUR panels (with addition of two H/L SNPs), and plots for (D and E) the two Asian PRS panels and the two PRS panels modified from Asian panels (with addition of two H/L SNPs). Error bars represent the 95% CI.

To explore whether IA ancestry affects the performance of the EUR PRSs in H/L samples, we evaluated the performance of three EUR PRSs in H/L populations stratified by the proportion of global IA ancestry. We chose three IA ancestry ranges to stratify our samples: 0 to 0.4, 0.4 to 0.6, and 0.6 to 1. The performance of all three EUR PRSs in the high IA range (0.6–1) was significantly worse than the performance in the low IA range (0–0.4). There was no significant difference of performance in modest versus high IA ranges (Fig. 1C; Supplementary Tables S7 and S8). For instance, PRS3503 had an AUC of 0.631 (95% CI, 0.615–0.648) in the low IA range, whereas the AUC was 0.601 (95% CI, 0.585–0.617) in the intermediate IA range and 0.602 (95% CI, 0.588–0.616) in the high IA range. The DeLong test of AUC difference between low and high IA ranges was statistically significant (P value = 0.0071). We saw a similar result in performance measured by the OR per SD; PRS3503 had an OR per SD of 1.62 (95% CI, 1.52–1.72) in the low IA range, whereas the OR per SD was 1.45 (95% CI, 1.37–1.52) in the high IA range. The OR difference between the low and high IA ranges was statistically significant (P value = 0.0053; Fig. 1G; Supplementary Table S8).

Supplementing EUR PRSs with H/L-specific risk variants

We had previously discovered three independent breast cancer risk variants in the 6q25 region (rs140068132, rs3778609, and rs851980) specific to H/L populations (11, 12). Two of them (rs140068132 and rs3778609) are novel and not in LD with any variant in EUR or Asian PRS panels whereas rs851980 is in strong LD with several variants in EUR and Asian PRS panels (Supplementary Table S9). We added the two variants (rs140068312 and rs3778609) to the EUR PRS to assess how much they improved risk prediction. The new panels were labeled PRS166, PRS297, and PRS3505. We evaluated the performance of these three new panels the same way as we did for the three original ones. We saw a significant improvement of all three new panels when evaluated in all pooled H/L samples (Fig. 1B and F; Supplemental Table S5). For instance, PRS3505 had an AUC of 0.626 (95% CI, 0.617–0.635) and an OR per SD of 1.59 (95% CI, 1.54–1.64) which was significantly higher than the AUC and OR of PRS3503 when evaluated in our total pooled samples (AUC difference P value = 2.5e–14 and OR difference P value = 0.01). Women in the top 10% of the PRS3505 panel had a 2.41-fold (95% CI, 2.15–2.70) elevated risk compared with women at average score (PRS in the 40th–60th percentile; Fig. 2B). The calibration curve and Hosmer–Lemeshow test suggested good fit for all three modified models (Fig. 3A–C; Supplementary Table S6). When evaluated in samples within different countries, we saw improvements in all countries, but the improvements were statistically significant only in the Peruvian and U.S. studies. When evaluated in H/L populations stratified by different ranges of global IA ancestry, we saw performance improvement in all strata for all three models, but the most profound improvement was in the high IA range (0.6–1; Fig. 1D and H; Supplementary Table S7). We also tested adding rs851980 in addition to the two variants on the performance of PRSs and found that it did not add further improvement (Supplementary Fig. S2A and S2B).

Evaluating Asian PRSs in H/L populations

We evaluated the performance of two previously published Asian PRSs in our pooled H/L data. Their development and derivation are described elsewhere (13) and briefly in the methods section.

We evaluated the association between individual variants of Asian PRSs and breast cancer risk in H/L samples using multivariable logistic regression models adjusted for genetic ancestry and study. Of the 2,750 variants in the Asian PRS, 1,762 had associations that were directionally consistent with those reported in Asian populations (Supplementary Table S4). Of these, 581 variants were nominally statistically significant (P value < 0.05), with 476 also directionally consistent. Forty-nine variants remained genome-wide statistically significant (P value less than 5 × 10−8).

The performance of these two PRSs in Asian studies was similar, with PRS2750 being slightly better than PRS43 (the reported AUC is 0.60 for PRS43 and 0.61 for PRS2750; ref. 13) . We first evaluated the performance of the two Asian PRSs in our total pooled H/L data from all participating studies adjusting for study and global ancestry. They performed worse in H/L (AUC is between 0.580 and 0.594, with an OR per SD between 1.33 and 1.39) than in Asian individuals (Fig. 4A and E; Supplementary Table S5). For PRS2750, the OR per 1 SD was 1.39 (95% CI, 1.35–1.44) and the AUC was 0.594 (95% CI, 0.585–0.603). PRS2750 showed slightly better performance than PRS43. We used PRS2750 to represent the best Asian PRSs in subsequent analyses. Women in the top 10% of the PRS2750 had a 1.76-fold (95% CI, 1.57–1.97) elevated risk compared with women at average score (PRS in the 40th–60th percentile; Fig. 2C). The calibration curve and Hosmer–Lemeshow test suggested good fit for both Asian panels (Fig. 3D and E; Supplementary Table S6). Next, we evaluated the performance of the Asian PRSs in samples from different countries. The performance of the Asian PRSs varied by countries in the H/L samples (Fig. 4A and E; Supplementary Table S5). For example, PRS2750 performed better in Peru (AUC = 0.603, 95% CI, 0.587–0.619) than in Mexico (AUC = 0.587, 95% CI, 0.566–0.609).

Figure 4.

Figure 4.

Association between PRSs from two Asian SNP panels (and modified panels with addition of two H/L SNPs) and breast cancer risk in H/L samples stratified by countries or IA genetic ancestry. Area under the ROC curve (AUC) or OR were calculated according to Statistical analyses in Materials and Methods. A, AUC of Asian PRS panels tested in H/L samples stratified by countries, (B) AUC of Asian PRS panels modified with addition of two H/L SNPs tested in H/L samples stratified by countries, (C) AUC of Asian PRS panels tested in H/L samples stratified by IA ancestry ranges, (D) AUC of Asian PRS panels modified with addition of two H/L SNPs tested in H/L samples stratified by IA ancestry ranges, (E) OR of Asian PRS panels tested in H/L samples stratified by countries, (F) OR of Asian PRS panels modified with addition of two H/L SNPs tested in H/L samples stratified by countries, (G) OR of Asian PRS panels tested in H/L samples stratified by IA ancestry ranges, and (H) OR of Asian PRS panels modified with addition of two H/L SNPs tested in H/L samples stratified by IA ancestry ranges. Error bars represent the 95% CI.

To explore whether IA ancestry affected the performance of the Asian PRSs, we evaluated the performance of the two Asian PRSs in H/L samples stratified by the different ranges of global IA ancestry. The performance of PRS2750 in the high IA range (0.6–1) was significantly better than the performance in the low IA range (0–0.4; Fig. 4C and G; Supplementary Tables S7 and S8). Specifically, PRS2750 had an OR per SD of 1.36 (95% CI, 1.28–1.45) in the low IA range, whereas the OR per SD was 1.50 (95% CI, 1.42–1.57) in the high IA range, and the P value for the OR difference was 0.018. Similarly, PRS2750 had an AUC of 0.59 (95% CI, 0.573–0.607) in the low IA range, whereas the AUC was 0.611 (95% CI, 0.598–0.625) in the high IA range. The DeLong test of AUC difference between the low and high IA ranges was suggestive (P value = 0.0519). The best Asian PRS, PRS2750, had a better AUC than the best EUR PRS PRS3503 (AUC difference P value = 0.048) in the highIA range (0.6–1; Figs 1C and 4C).

Supplementing Asian PRSs with H/L-specific risk variants

We added the two H/L-specific risk variants (rs140068312 and rs3778609) to the two Asian PRSs (PRS43 and PRS2750), which resulted in two new panels: PRS45 and PRS2752, respectively. We evaluated the performance of the two new models the same way as above for the two original models. We saw significant improvements of the new Asian models compared with the original PRS43 and PRS2750 when evaluated in the combined H/L dataset (Fig. 4B and F; Supplementary Table S5). For instance, PRS2752 had an AUC of 0.606 (95% CI, 0.597–0.615) and an OR per SD = 1.47 (95% CI, 1.42–1.52), which is significantly higher than the original PRS2750 when evaluated in our complete pooled samples (AUC difference P value = 1.03e–06 and OR difference P value = 0.017). Women in the top 10% of the PRS2752 had a 1.89-fold (95% CI = 1.69–2.12) elevated risk compared with women at average score (PRS in the 40–60th percentile; Fig. 2D). The calibration curve and Hosmer–Lemeshow test suggested good fit for both modified models (Fig. 3D and E; Supplementary Table S6). When evaluated in samples from different countries, we saw improvements in all countries for both models, but the improvements were statistically significant only in Peru and the United States. When evaluated in H/L samples stratified by different ranges of global IA ancestry, we saw performance improvements in all ranges for both models (Fig. 4D and H; Supplementary Table S7).

Performance of a combined EUR/Asian PRS

We also evaluated the performance of PRSs combining the best Asian PRS (PRS2750) with the three EUR PRSs with or without the addition of H/L-specific variants. The combined PRSs had markedly improved performance compared with the Asian PRS2750 alone (Supplementary Fig. S3A and S3B). The combined PRSs had modest improvement in performance compared with the EUR PRSs when evaluated in pooled samples, samples stratified by countries, or samples in low or modest IA ranges whereas the improvement was more significant in the high IA range (Supplementary Fig. S3A and S3B). Addition of H/L-specific variants had similar effects to the combined PRSs as to the EUR PRSs alone (Supplementary Fig. S3C and S3D). With the addition of H/L-specific variants, combining the Asian PRS2750 with the EUR models did not provide significant further improvement on performance in the high IA range (Supplementary Fig. S3D).

Discussion

There is increasing interest in using PRS for personalized risk stratification for breast cancer prevention and early detection (4). However, it is crucial that women representing diverse ancestries are included in PRS studies to reduce health disparities. Our study provides important information about the utility of PRSs for breast cancer risk prediction in H/L women. We tested the performance of different models previously developed in EUR and Asian studies and compared them with models that also included two risk-associated variants discovered in H/L studies. Our key findings were that (i) the best performing EUR or Asian PRSs both perform worse in the H/L studies; (ii) the EUR PRSs performed better in H/L samples with relatively low IA ancestry proportion than in H/L samples with relatively high IA ancestry proportion, whereas the Asian PRSs performed better in H/L samples with relatively high IA ancestry proportion than in H/L samples with relatively low IA ancestry proportion, which is consistent with what would be expected based on the distribution of genomic variation established by early human migration; (iii) the best EUR PRS performed better than the best Asian PRS in pooled H/L samples and particularly in H/L samples with relatively low IA ancestry proportion, whereas the best Asian PRS performed slightly better than the best EUR PRS in H/L samples with relatively high IA ancestry proportion, although the difference was not statistically significant; (iv) a combination of the EUR PRS with the best Asian PRS performed similarly to the EUR PRS in pooled H/L samples and H/L samples with low or modest IA ancestry whereas similarly to the Asian PRS in H/L samples with high IA ancestry; and (v) addition of the two H/L variants improved performance across all models (performance improvement was mostly due to the addition of SNP rs140068132, see Supplementary Figs. S4A–S4F and S5A–S5F), whereas the improvement was more significant for EUR models and in women of relatively high IA ancestry proportions.

Only a few studies have reported PRS development and/or evaluation in H/L samples. An early study using only 147 H/L cases and 3,201 controls evaluated a 71-SNP (PRS with the SNPs identified in an EUR GWAS (16) and reported a modest performance (OR per SD = 1.39, AUC = 0.59). This study had not tested the effect of ancestry on PRS performance. A more recent study conducted by Shieh and colleagues (14) developed a 180-SNP which included SNPs mostly discovered in EUR studies but also the two SNPs discovered in H/L-specific GWAS. This larger-panel PRS performed better than the smaller 71-SNP PRS in H/L samples; however, the study did not evaluate how the performance of PRS varied with inclusion of more EUR SNPs and/or the two H/L-specific SNPs and/or the PRS developed in EAS populations. Although the study by Shieh and colleagues tested the association of ancestry with PRS performance and found no significant association, it is possible that the relatively limited range of IA ancestry variation in the data set had limited power to detect it. Another more recent study had examined the performance of previously developed PRSs (including the two PRSs mentioned above) in a cohort including a modest number of H/L samples (40). This study replicated the positive predictive performance of previously developed models and confirmed that PRS180 performed better than the 71-SNP PRS in H/L samples. No test of the effect of ancestry on PRS performance was conducted in that study.

Genetic ancestry has been speculated and shown to affect the performance of PRSs for various diseases (41, 42). The current study includes a relatively large proportion of H/L participants with very high IA ancestry proportion compared with previous studies. Our study included 2,394 cases and 4,667 controls with IA ancestry greater than 0.6, whereas the previous 180-SNP PRS study had only 1,666 cases and 1,404 controls with IA ancestry greater than 0.55 (14).

Our study is the largest to date in evaluating breast cancer PRS in H/L samples with relatively high IA ancestry contribution. The increased sample size and diversity allowed us to detect the variability of PRS performance by ancestry which was not detected in smaller studies. Although previous studies have developed and evaluated PRSs with the addition of H/L-specific variants (14, 24), they did not include information about performance in strata based on IA ancestry proportion or national origin (i.e., Peruvians, Colombians, Mexicans, and U.S. H/L).

Previous studies used approaches that combined existing ancestry-specific PRSs into a joint model weighted by ancestry proportions (22, 24, 43) to improve the PRS performance in admixed populations. In the current study, we integrated the ancestry-specific variants directly into existing PRSs, confirming that adding IA ancestry–specific variants to existing models does not decrease the PRS performance in samples with relatively low IA ancestry. Our results suggest that it is more important to discover ancestry-specific variants and add them to existing models than to refine existing models by ancestry proportions of individuals. This approach has an important practical advantage of making the implementation of PRS in clinical settings more straightforward without worrying about estimating the genetic ancestry portions of individuals.

Our study has several limitations. First, our dataset does not include all the SNPs in the original models due to genotyping/imputation differences. Although this limitation could potentially bias our analysis, we showed that the PRSs with slightly fewer SNPs were able to reproduce the performance of original PRSs using PRS166 (AUC = 0.62) versus PRS180 (AUC = 0.63) as an example. Second, our study samples do not represent all H/L populations, as they do not include, for example, H/L individuals of Caribbean origin who have higher AFR ancestry proportions (23, 44, 45). Performance of breast cancer PRS has been lower in AFR ancestry populations (46). Therefore, improving PRS in H/L populations from the Caribbean will likely benefit from improving PRSs for AFR ancestry populations. A recent GWAS of breast cancer in women of AFR ancestry identified new susceptibility variants which improved risk prediction (47). Including some of these variants in future PRS panels might improve performance in H/L individuals of African ancestry. Further studies including studies with Caribbean representation are needed to make the analysis more generalizable and develop more robust PRSs applicable to all H/L individuals. Third, the current analysis included samples from two studies (SFBCS/NC-BCFR and MEC) used in the discovery H/L GWASs. This might have introduced a bias in the performance of the PRSs with the two H/L SNPs. Because the two discovery studies included only U.S. H/L participants, the bias only affects the U.S. H/L group in the AUC estimates stratified by country and the IA ancestry range 0 to 0.4 estimate. We performed the same analysis on a dataset excluding the two discovery studies and confirmed the robustness of the estimates (Supplementary Fig. S6A–S6D). Fourth, there are substantial differences in the age of cases and controls in some of our studies (such as COLUMBUS, PEGEN-BC, and Kaiser) which might have biased PRS performance. We performed the same analysis on a subset of the datasets with matched age between cases and controls (1:1) and found no substantial difference in performance from our main analysis (Supplementary Fig. S7A–S7H). Changes in estimates are expected due to differences in sample size in the matched analysis. Finally, our study included individuals from community-based and/or familial breast cancer clinics and likely included a number of individuals carrying pathogenic variants in moderate- or high-penetrance genes. The association of PRS with breast cancer risk in mutation carriers might have different effect sizes compared with the general population (48, 49).

In summary, our results suggest that PRSs developed in EUR or Asian studies have predictive value in H/L individuals and that their performance can be further enhanced by the addition of H/L-specific variants. We also demonstrate that the addition of H/L-specific variants has the most impact in improving PRS-based prediction in the subset of individuals with the highest IA ancestry proportions without hurting prediction in the lower IA ancestry proportion subsets. Overall, this study supports the continuation of efforts to enhance diversity for breast cancer risk variant discovery as to improve equity in prediction accuracy.

Supplementary Material

Supplemental Table S1

Supplemental Table S1 shows description of the existing and modified PRS panels evaluated in our studies.

Supplemental Figure S1

Supplemental Figure S1 shows global ancestry distribution of individuals by countries across seven studies included in this analysis.

Supplemental Table S2

Supplemental Table S2 shows performance of PRS panels constructed from PRS3820 European panel plus 2 H/L SNPs using different imputation r2 cutoffs.

Supplemental Figure S2

Supplemental Figure S2 shows performance of PRS panels with addition of 2 or 3 H/L SNPs

Supplemental Table S3

Supplemental Table S3 shows sample sizes in the Indigenous American ancestry ranges selected for our analysis.

Supplemental Figure S3

Supplemental Figure S3 shows performance of combinations of European and Asian PRS panels in H/L samples.

Supplemental Table S4

Supplemental Table S4 shows a list of polymorphisms in the PRS panels with association analyses summary statistics from 7 pooled Hispanic/Latino datasets.

Supplemental Figure S4

Supplemental Figure S4 shows effect of adding only one H/L SNP to European PRS panels on performance in H/L samples.

Supplemental Table S5

Supplemental Table S5 shows Comparisons of Areas under the receiver operating characteristics curve (AUC) and odd ratios (OR) per standard deviation of the PRS panels.

Supplemental Figure S5

Supplemental Figure S5 shows effect of adding only one H/L SNP to Asian PRS panels on performance in H/L samples.

Supplemental Table S6

Supplemental Table S6 shows two-sided Hosmer-Lemeshow P values for calibration of the PRS panels tested in our datasets.

Supplemental Figure S6

Supplemental Figure S6 shows performance of PRS panels in H/L samples excluding samples in discovery GWAS.

Supplemental Table S7

Supplemental Table S7 shows comparisons of performance for the PRS panel pairs in different ranges of Indigenous American (IA) ancestry.

Supplemental Figure S7

Supplemental Figure S7 shows performance of PRS panels in H/L samples with matched age in cases and controls.

Supplemental Table S8

Supplemental Table S8 shows comparisons of performance of the PRS panel for different pair of ranges of Indigenous American (IA) ancestry.

Supplemental Table S9

Supplemental Table S9 shows variants in European or Asian PRS panels with high linkage disequilibrium (LD) with rs851980

Acknowledgments

We first want to thank the participants of all the studies that were part of the present analyses. We are also grateful for the financial support from the Placer Breast Cancer Endowed Chair, University of California Davis (L. Fejerman), The NCI (R01-CA286650, R01-CA273313, and R01-CA204797 to L. Fejerman), the California Precision Medicine Initiative Grant, Office of the California Governor (OPR18111 to E. Ziv, S.L. Neuhausen, and L.G. Carvajal-Carmona), the Universidad del Tolima, Colombia (projects 10112 and 520115); MINCIENCIAS, Colombia program “Becas Doctorales Nacionales” (grant 528-2011 to C. Ramirez and 647-2015 to A.P. Estrada-Florez); MINCIENCIAS, Colombia program “Formación de Capital Humano de Alto Nivel para el Departamento de Tolima- 2016” (grant 755-2016, J. Benavides); L’OREAL-UNESCO-ICETEX-COLCIENCIAS, Colombia (project 3900917/2017, A.P. Estrada-Florez); Coordinacion Nacional de Investigación en Salud, Instituto Mexicano del Seguro Social, México (grant FIS/IMSS/PROT/PRIO/13/027, J. Torres) and Consejo Nacional de Ciencia y Tecnología, México (Fronteras de la Ciencia grant 773, J. Torres); 2021 AACR Fellowship to Further Diversity, Equity, and Inclusion in Cancer Disparities Research, Grant Number 21-40-69-ESTR, A.P. Estrada-Florez; The GSK Oncology Ethnic Research Initiative (M. Echeverry de Polanco; L.G. Carvajal-Carmona), The Auburn Community Cancer Endowed Chair in Basic Research, U.S.(L.G. Carvajal-Carmona); the Heart, BrEast, and BrAin HeaLth Equity Research Program (P.C. Lott; L.G. Carvajal-Carmona; L. Fejerman), The residual class settlement funds in the matter of Krueger v. Wyeth, Inc. [396 F.Supp.3d 931 (S.D. Cal. 2012), L.G. Carvajal-Carmona], Towards Health Equity Oncology Grant, Gilead Sciences (grant no. 19425, L.G. Carvajal-Carmona), and the NCI/NIH (grant 5D43CA260869-01, L.G. Carvajal-Carmona and L. Fejerman). The COLUMBUS Consortium (in alphabetical order) includes the following: Jennyfer Benavides (Universidad del Tolima, Ibagué, Colombia), Mabel Bohorquez (Universidad del Tolima, Ibagué, Colombia), Fernando Bolaños (Hospital Hernando Moncaleano Perdomo, Neiva, Colombia), Luis G Carvajal-Carmona (University of California Comprehensive Cancer Center, Sacramento), Jenny Carmona (Dinámica IPS, Medellín, Colombia), Ángel Criollo (Universidad del Tolima, Ibagué, Colombia), Magdalena Echeverry (Universidad del Tolima, Ibagué, Colombia), Ana Estrada (Universidad del Tolima, Ibagué, Colombia),Gilbert Mateus (Hospital Federico Lleras Acosta, Ibagué, Colombia), Raúl Murillo (Pontificia Universidad Javeriana, Bogotá, Colombia), Justo Ramirez (Hospital Hernando Moncaleano Perdomo, Neiva, Colombia), Yesid Sánchez (Universidad del Tolima, Ibagué, Colombia), Carolina Sanabria (Instituto Nacional de Cancerología, Bogotá, Colombia), Martha Lucia Serrano (Instituto Nacional de Cancerología, Bogotá, Colombia), John Jairo Suarez (Universidad del Tolima, Ibagué, Colombia), and Alejandro Vélez (Dinámica IPS, Medellín, Colombia, Hospital Pablo Tobón Uribe, Medellín, Colombia).

Footnotes

Note: Supplementary data for this article are available at Cancer Epidemiology, Biomarkers & Prevention Online (http://cebp.aacrjournals.org/).

Authors’ Disclosures

E.M. John reports grants from the NIH during the conduct of the study. L.H. Kushi reports grants from the NIH during the conduct of the study. L. Fejerman reports grants from the NIH and Precision Medicine Initiative Office of the Governor during the conduct of the study and grants from Gilead Corporation outside the submitted work. No disclosures were reported by the other authors.

Authors’ Contributions

X. Huang: Conceptualization, data curation, software, formal analysis, validation, investigation, visualization, methodology, writing–original draft, writing–review and editing. P.C. Lott: Conceptualization, resources, data curation, formal analysis, methodology, writing–review and editing. D. Hu: Data curation, formal analysis, writing–review and editing. V.A. Zavala: Data curation, validation, writing–review and editing. Z.N. Jamal: Data curation, writing–review and editing. T. Vidaurre: Data curation, writing–review and editing. S. Casavilca–Zambrano: Data curation, writing–review and editing. J. Navarro Vásquez: Data curation, writing–review and editing. C.A. Castañeda: Data curation, writing–review and editing. G. Valencia: Data curation, writing–review and editing. Z. Morante: Data curation, writing–review and editing. M. Calderón: Data curation, writing–review and editing. J.E. Abugattas: Data curation, writing–review and editing. H.A. Fuentes: Data curation, writing–review and editing. R. Liendo-Picoaga: Data curation, writing–review and editing. J.M. Cotrina: Data curation, writing–review and editing. S.P. Neciosup: Data curation, writing–review and editing. P. Rioja Viera: Data curation, writing–review and editing. L.A. Salinas: Data curation, writing–review and editing. M. Galvez-Nino: Data curation, writing–review and editing. S. Huntsman: Data curation, writing–review and editing. S.E. Sanchez: Data curation, writing–review and editing. M.A. Williams: Data curation, writing–review and editing. B. Gelaye: Data curation, writing–review and editing. A.P. Estrada-Florez: Data curation, writing–review and editing. G. Polanco-Echeverry: Data curation, writing–review and editing. M. Echeverry: Data curation, writing–review and editing. A. Velez: Data curation, writing–review and editing. J.A. Carmona-Valencia: Data curation, writing–review and editing. M.E. Bohorquez-Lozano: Data curation, writing–review and editing. J. Torres: Data curation, writing–review and editing. M. Cruz: Data curation, writing–review and editing. W.-K. Ho: Data curation, writing–review and editing. S.H. Teo: Data curation, writing–review and editing. M.C. Tai: Data curation, writing–review and editing. E.M. John: Data curation, funding acquisition, writing–review and editing. C.A. Haiman: Data curation, funding acquisition, writing–review and editing. D.V. Conti: Data curation, writing–review and editing. F. Chen: Data curation, writing–review and editing. G. Torres-Mejía: Data curation, writing–review and editing. L.H. Kushi: Data curation, writing–review and editing. S.L. Neuhausen: Data curation, funding acquisition, writing–review and editing. E. Ziv: Data curation, funding acquisition, writing–review and editing. L.G. Carvajal-Carmona: Data curation, funding acquisition, writing–review and editing. L. Fejerman: Conceptualization, resources, formal analysis, supervision, funding acquisition, investigation, methodology, writing–original draft, project administration, writing–review and editing.

References

  • 1. Sedeta E, Sung H, Laversanne M, Bray F, Jemal A. Recent mortality patterns and time trends for the major cancers in 47 countries worldwide. Cancer Epidemiol Biomarkers Prev 2023;32:894–905. [DOI] [PubMed] [Google Scholar]
  • 2. Mucci LA, Hjelmborg JB, Harris JR, Czene K, Havelick DJ, Scheike T, et al. Familial risk and heritability of cancer among twins in Nordic countries. JAMA 2016;315:68–76. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Mavaddat N, Pharoah PDP, Michailidou K, Tyrer J, Brook MN, Bolla MK, et al. Prediction of breast cancer risk based on profiling with common genetic variants. J Natl Cancer Inst 2015;107:djv036. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Mavaddat N, Michailidou K, Dennis J, Lush M, Fachal L, Lee A, et al. Polygenic risk scores for prediction of breast cancer and breast cancer subtypes. Am J Hum Genet 2019;104:21–34. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Adedokun B, Du Z, Gao G, Ahearn TU, Lunetta KL, Zirpoli G, et al. Cross-ancestry GWAS meta-analysis identifies six breast cancer loci in African and European ancestry women. Nat Commun 2021;12:4198. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Zhang H, Ahearn TU, Lecarpentier J, Barnes D, Beesley J, Qi G, et al. Genome-wide association study identifies 32 novel breast cancer susceptibility loci from overall and subtype-specific analyses. Nat Genet 2020;52:572–81. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Lilyquist J, Ruddy KJ, Vachon CM, Couch FJ. Common genetic variation and breast cancer risk-past, present, and future. Cancer Epidemiol Biomarkers Prev 2018;27:380–94. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Michailidou K, Lindström S, Dennis J, Beesley J, Hui S, Kar S, et al. Association analysis identifies 65 new breast cancer risk loci. Nature 2017;551:92–4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Michailidou K, Beesley J, Lindstrom S, Canisius S, Dennis J, Lush MJ, et al. Genome-wide association analysis of more than 120,000 individuals identifies 15 new susceptibility loci for breast cancer. Nat Genet 2015;47:373–80. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Michailidou K, Hall P, Gonzalez-Neira A, Ghoussaini M, Dennis J, Milne RL, et al. Large-scale genotyping identifies 41 new loci associated with breast cancer risk. Nat Genet 2013;45:353–61, 361e1-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Fejerman L, Ahmadiyeh N, Hu D, Huntsman S, Beckman KB, Caswell JL, et al. Genome-wide association study of breast cancer in Latinas identifies novel protective variants on 6q25. Nat Commun 2014;5:5260. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Hoffman J, Fejerman L, Hu D, Huntsman S, Li M, John EM, et al. Identification of novel common breast cancer risk variants at the 6q25 locus among Latinas. Breast Cancer Res 2019;21:3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Ho W-K, Tai M-C, Dennis J, Shu X, Li J, Ho PJ, et al. Polygenic risk scores for prediction of breast cancer risk in Asian populations. Genet Med 2022;24:586–600. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Shieh Y, Fejerman L, Lott PC, Marker K, Sawyer SD, Hu D, et al. A polygenic risk score for breast cancer in US Latinas and Latin American women. J Natl Cancer Inst 2020;112:590–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Ferreira MA, Gamazon ER, Al-Ejeh F, Aittomäki K, Andrulis IL, Anton-Culver H, et al. Genome-wide association and transcriptome studies identify target genes and risk loci for breast cancer. Nat Commun 2019;10:1741. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Allman R, Dite GS, Hopper JL, Gordon O, Starlard-Davenport A, Chlebowski R, et al. SNPs and breast cancer risk prediction for African American and Hispanic women. Breast Cancer Res Treat 2015;154:583–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Wang S, Qian F, Zheng Y, Ogundiran T, Ojengbede O, Zheng W, et al. Genetic variants demonstrating flip-flop phenomenon and breast cancer risk prediction among women of African ancestry. Breast Cancer Res Treat 2018;168:703–12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Wen W, Shu X-O, Guo X, Cai Q, Long J, Bolla MK, et al. Prediction of breast cancer risk based on common genetic variants in women of East Asian ancestry. Breast Cancer Res 2016;18:124. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Chan CHT, Munusamy P, Loke SY, Koh GL, Yang AZY, Law HY, et al. Evaluation of three polygenic risk score models for the prediction of breast cancer risk in Singapore Chinese. Oncotarget 2018;9:12796–804. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Hsieh Y-C, Tu S-H, Su C-T, Cho E-C, Wu C-H, Hsieh M-C, et al. A polygenic risk score for breast cancer risk in a Taiwanese population. Breast Cancer Res Treat 2017;163:131–8. [DOI] [PubMed] [Google Scholar]
  • 21. Ho W-K, Tan M-M, Mavaddat N, Tai M-C, Mariapun S, Li J, et al. European polygenic risk score for prediction of breast cancer shows similar performance in Asian women. Nat Commun 2020;11:3833. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Gao G, Zhao F, Ahearn TU, Lunetta KL, Troester MA, Du Z, et al. Polygenic risk scores for prediction of breast cancer risk in women of African ancestry: a cross-ancestry approach. Hum Mol Genet 2022;31:3133–43. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Bryc K, Velez C, Karafet T, Moreno-Estrada A, Reynolds A, Auton A, et al. Colloquium paper: genome-wide patterns of population structure and admixture among Hispanic/Latino populations. Proc Natl Acad Sci U S A 2010;107(Suppl 2):8954–61. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Hughes E, Wagner S, Pruss D, Bernhisel R, Probst B, Abkevich V, et al. Development and validation of a breast cancer polygenic risk score on the basis of genetic ancestry composition. JCO Precis Oncol 2022;6:e2200084. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Gelaye B, Zhong Q-Y, Basu A, Levey EJ, Rondon MB, Sanchez S, et al. Trauma and traumatic stress in a sample of pregnant women. Psychiatry Res 2017;257:506–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Zavala VA, Casavilca-Zambrano S, Navarro-Vásquez J, Tamayo LI, Castañeda CA, Valencia G, et al. Breast cancer subtype and clinical characteristics in women from Peru. Front Oncol 2023;13:938042. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Zavala VA, Casavilca-Zambrano S, Navarro-Vásquez J, Castañeda CA, Valencia G, Morante Z, et al. Association between ancestry-specific 6q25 variants and breast cancer subtypes in Peruvian women. Cancer Epidemiol Biomarkers Prev 2022;31:1602–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. John EM, Phipps AI, Davis A, Koo J. Migration history, acculturation, and breast cancer risk in Hispanic women. Cancer Epidemiol Biomarkers Prev 2005;14:2905–13. [DOI] [PubMed] [Google Scholar]
  • 29. John EM, Sangaramoorthy M, Koo J, Whittemore AS, West DW. Enrollment and biospecimen collection in a multiethnic family cohort: the Northern California site of the Breast Cancer Family Registry. Cancer Causes Control 2019;30:395–408. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Angeles-Llerenas A, Ortega-Olvera C, Pérez-Rodríguez E, Esparza-Cano JP, Lazcano-Ponce E, Romieu I, et al. Moderate physical activity and breast cancer risk: the effect of menopausal status. Cancer Causes Control 2010;21:577–86. [DOI] [PubMed] [Google Scholar]
  • 31. Kolonel LN, Henderson BE, Hankin JH, Nomura AM, Wilkens LR, Pike MC, et al. A multiethnic cohort in Hawaii and Los Angeles: baseline characteristics. Am J Epidemiol 2000;151:346–57. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Kvale MN, Hesselson S, Hoffmann TJ, Cao Y, Chan D, Connell S, et al. Genotyping informatics and quality control for 100,000 subjects in the Genetic Epidemiology Research on Adult health and aging (GERA) cohort. Genetics 2015;200:1051–60. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Das S, Forer L, Schönherr S, Sidore C, Locke AE, Kwong A, et al. Next-generation genotype imputation service and methods. Nat Genet 2016;48:1284–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. 1000 Genomes Project Consortium; Auton A, Brooks LD, Durbin RM, Garrison EP, Kang HM, Korbel JO, et al. A global reference for human genetic variation. Nature 2015;526:68–74. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Fuchsberger C, Abecasis GR, Hinds DA. minimac2: faster genotype imputation. Bioinformatics 2015;31:782–4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Taliun D, Harris DN, Kessler MD, Carlson J, Szpiech ZA, Torres R, et al. Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program. Nature 2021;590:290–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. Alexander DH, Novembre J, Lange K. Fast model-based estimation of ancestry in unrelated individuals. Genome Res 2009;19:1655–64. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. DeLong ER, DeLong DM, Clarke-Pearson DL. Comparing the areas under two or more correlated receiver operating characteristic curves: a nonparametric approach. Biometrics 1988;44:837–45. [PubMed] [Google Scholar]
  • 39. Robin X, Turck N, Hainard A, Tiberti N, Lisacek F, Sanchez J-C, et al. pROC: an open-source package for R and S+ to analyze and compare ROC curves. BMC Bioinformatics 2011;12:77. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Liu C, Zeinomar N, Chung WK, Kiryluk K, Gharavi AG, Hripcsak G, et al. Generalizability of polygenic risk scores for breast cancer among women with European, African, and Latinx ancestry. JAMA Netw Open 2021;4:e2119084. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Ding Y, Hou K, Xu Z, Pimplaskar A, Petter E, Boulier K, et al. Polygenic scoring accuracy varies across the genetic ancestry continuum. Nature 2023;618:774–81. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Wang Y, Kanai M, Tan T, Kamariza M, Tsuo K, Yuan K, et al. Polygenic prediction across populations is influenced by ancestry, genetic architecture, and methodology. Cell Genomics 2023;3:100408. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Tshiaba PT, Ratman DK, Sun JM, Tunstall TS, Levy B, Shah PS, et al. Integration of a cross-ancestry polygenic model with clinical risk factors improves breast cancer risk stratification. JCO Precis Oncol 2023;7:e2200447. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Bertoni B, Budowle B, Sans M, Barton SA, Chakraborty R. Admixture in Hispanics: distribution of ancestral population contributions in the Continental United States. Hum Biol 2003;75:1–11. [DOI] [PubMed] [Google Scholar]
  • 45. Ziv E, John EM, Choudhry S, Kho J, Lorizio W, Perez-Stable EJ, et al. Genetic ancestry and risk factors for breast cancer among Latinas in the San Francisco Bay Area. Cancer Epidemiol Biomarkers Prev 2006;15:1878–85. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Du Z, Gao G, Adedokun B, Ahearn T, Lunetta KL, Zirpoli G, et al. Evaluating polygenic risk scores for breast cancer in women of African ancestry. J Natl Cancer Inst 2021;113:1168–76. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. Jia G, Ping J, Guo X, Yang Y, Tao R, Li B, et al. Genome-wide association analyses of breast cancer in women of African ancestry identify new susceptibility loci and improve risk prediction. Nat Genet 2024;56:819–26. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Kuchenbaecker KB, McGuffog L, Barrowdale D, Lee A, Soucy P, Dennis J, et al. Evaluation of polygenic risk scores for breast and ovarian cancer risk prediction in BRCA1 and BRCA2 mutation carriers. J Natl Cancer Inst 2017;109:djw302. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49. Barnes DR, Rookus MA, McGuffog L, Leslie G, Mooij TM, Dennis J, et al. Polygenic risk scores and breast and epithelial ovarian cancer risks for carriers of BRCA1 and BRCA2 pathogenic variants. Genet Med 2020;22:1653–66.trun ‐1 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplemental Table S1

Supplemental Table S1 shows description of the existing and modified PRS panels evaluated in our studies.

Supplemental Figure S1

Supplemental Figure S1 shows global ancestry distribution of individuals by countries across seven studies included in this analysis.

Supplemental Table S2

Supplemental Table S2 shows performance of PRS panels constructed from PRS3820 European panel plus 2 H/L SNPs using different imputation r2 cutoffs.

Supplemental Figure S2

Supplemental Figure S2 shows performance of PRS panels with addition of 2 or 3 H/L SNPs

Supplemental Table S3

Supplemental Table S3 shows sample sizes in the Indigenous American ancestry ranges selected for our analysis.

Supplemental Figure S3

Supplemental Figure S3 shows performance of combinations of European and Asian PRS panels in H/L samples.

Supplemental Table S4

Supplemental Table S4 shows a list of polymorphisms in the PRS panels with association analyses summary statistics from 7 pooled Hispanic/Latino datasets.

Supplemental Figure S4

Supplemental Figure S4 shows effect of adding only one H/L SNP to European PRS panels on performance in H/L samples.

Supplemental Table S5

Supplemental Table S5 shows Comparisons of Areas under the receiver operating characteristics curve (AUC) and odd ratios (OR) per standard deviation of the PRS panels.

Supplemental Figure S5

Supplemental Figure S5 shows effect of adding only one H/L SNP to Asian PRS panels on performance in H/L samples.

Supplemental Table S6

Supplemental Table S6 shows two-sided Hosmer-Lemeshow P values for calibration of the PRS panels tested in our datasets.

Supplemental Figure S6

Supplemental Figure S6 shows performance of PRS panels in H/L samples excluding samples in discovery GWAS.

Supplemental Table S7

Supplemental Table S7 shows comparisons of performance for the PRS panel pairs in different ranges of Indigenous American (IA) ancestry.

Supplemental Figure S7

Supplemental Figure S7 shows performance of PRS panels in H/L samples with matched age in cases and controls.

Supplemental Table S8

Supplemental Table S8 shows comparisons of performance of the PRS panel for different pair of ranges of Indigenous American (IA) ancestry.

Supplemental Table S9

Supplemental Table S9 shows variants in European or Asian PRS panels with high linkage disequilibrium (LD) with rs851980

Data Availability Statement

The datasets generated and/or analyzed during the current study are available from the corresponding author upon request.


Articles from Cancer Epidemiology, Biomarkers & Prevention are provided here courtesy of American Association for Cancer Research

RESOURCES