Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 Aug 12.
Published in final edited form as: Cancer Epidemiol Biomarkers Prev. 2025 Sep 2;34(9):1593–1599. doi: 10.1158/1055-9965.EPI-25-0371

An analytic pipeline to obtain reliable genetic ancestry estimates from tumor-derived RNA sequencing data

Courtney E Johnson 1, Ximing Ran 2, Julia Wrobel 2, Natalie R Davidson 3, Casey S Greene 4, Michael P Epstein 5, Jeffrey R Marks 6, Lauren C Peres 7, Jennifer A Doherty 8, Joellen M Schildkraut 1
PMCID: PMC12340684  NIHMSID: NIHMS2096950  PMID: 40622249

Abstract

Background:

Germline genetics may influence tumor molecular characteristics and ultimately cancer survival. Studies of tumor characteristics, including our epithelial ovarian cancer (EOC) studies of Black women in the United States, may have RNASeq data from archival tumor tissue but lack germline DNA for at least some individuals. Incomplete germline DNA measurements impede analyses of important measures like global genetic ancestry, often used in downstream analyses, by reducing sample sizes.

Methods:

The study population consists of 184 women who participated in two population-based studies of EOC with both germline and formalin-fixed paraffin-embedded (FFPE) tumor samples and an additional 58 women diagnosed with EOC from the same two studies with only FFPE tumor tissue. We used tumor RNASeq data to calculate proportions of African, European, and Asian genetic ancestry using a pipeline built on the packages SeqKit, HISAT2, SAMtools, BCFtools, plink, and ADMIXTURE. Women from the 1000 Genomes Project were used as the reference populations, and germline genetic ancestry estimates from blood or saliva were used as the baseline comparison. We evaluated multiple quality control strategies to improve genetic ancestry estimation.

Results:

Correlations between tumor RNASeq-derived estimates of genetic ancestry from our pipeline and germline-derived African and European genetic ancestry ranged between 0.76-0.94.

Conclusions:

RNASeq data from archival FFPE tumor tissue can be confidently and efficiently used to approximate global genetic ancestry in an admixed population when germline DNA is unavailable.

Impact:

This approach supports analyses of genetic ancestry and cancer when germline samples are not available.

Keywords: genetic ancestry, RNASeq, methods

INTRODUCTION

There is evidence across many cancer types that tumor features associated with differential outcomes vary by genetic ancestry1-9, and that within self-reported race, variation in genetic ancestry is associated with tumor features that have an impact on clinical outcomes3,4,9. Lee et. al. identified differential survival, gene expression, and methylation by genetic ancestry across multiple cancer types3. In non-small cell lung cancer, KRAS mutations were associated with European genetic ancestry9. Among Black women with breast cancer, an increased proportion of African genetic ancestry was related to an increased association with estrogen receptor-negative (ER-) or triple-negative breast cancer (TNBC), two subtypes typified by poorer overall survival2. Martini et. al. found further evidence in a cohort of Black women with breast cancer that genetic ancestry is related to both the immunologic landscape and molecular signatures in breast cancer, specifically TNBC4. In an analysis of breast cancer conducted by Telonis et. al., women with less African genetic ancestry had better 10-year overall survival, and genetic ancestry was associated with distinct subtype-specific gene expression and methylation profiles8. Within breast cancer subtypes, Telonis et. al. further identified the top 10 genes and top 50 methylation probes able to predict genetic ancestry via supervised principal components8.

Thus, the incorporation of genetic ancestry is invaluable for genomic studies of cancer. However, a common practical issue in some studies is that only formalin-fixed paraffin embedded (FFPE) tumor tissues are available, and germline DNA is lacking. For example, in the multicenter study of epithelial ovarian cancer (EOC), the African American Cancer Epidemiology Study (AACES)10, the availability of germline DNA is incomplete or absent. Genetic ancestry is thus harder to infer study-wide, reducing sample size and limiting power for analyses. Leveraging information from available tissue sequencing data to accurately characterize genetic ancestry in the absence of germline DNA would alleviate this limitation.

Here, we propose a computational pipeline that allows for the extraction of necessary data for genetic ancestry estimation from tumor RNASeq. The pipeline uses several existing software packages, SeqKit11, HISAT212, SAMtools13, BCFtools13, PLINK v1.914, and ADMIXTURE15. We illustrate the utility of this pipeline using admixed women from studies of EOC, an uncommon, fatal, and heterogenous disease where sample size issues related to biospecimens are especially exacerbated16.

MATERIALS AND METHODS

Study Population

Black women diagnosed with EOC enrolled in either AACES10 or the North Carolina Ovarian Cancer Study (NCOCS)17 and who consented to the release of tumor tissue were included in the analysis. Written informed consent was obtained for NCOCS participants, whereas AACES participants provided verbal consent for interviews and signed medical record and pathology release forms to allow for access to tumor tissue. The NCOCS and AACES studies were conducted in accordance with the US Common Rule and approved by the Duke Medical Center Institutional Review Board (IRB) and the IRBs of participating enrollment sites. If they consented to the release of blood or saliva, they were included in the overlapping informative subset, and if they did not, they were included in the exploratory subset. NCOCS enrolled participants in North Carolina between 1999-2005, and AACES enrolled participants from 11 metropolitan areas or states across the United States between 2010-201510,17. A total of 184 women provided both samples, and 58 women provided only tumor tissue.

FFPE tumor tissue samples were obtained from the institutions where the patients were treated. RNA was extracted and sequenced to create paired-end FASTQ files for each participant as previously described18 (database of Genotypes and Phenotypes, dbGaP, accession number phs003632.v1.p1). As in Davidson & Barnard, et. al., we refer to the study populations of AACES and Black NCOCS participants collectively as SchildkrautB18. SNPs from DNA extracted from blood or saliva were genotyped via the OncoArray and phased via SHAPEIT v219, and this process in the germline data has also been previously described in detail20,21 (dbGaP accession number phs001882.v1.p1).

Genetic Ancestry and Similarity

In this pipeline, our goal is to calculate a measure approximating genetic ancestry. However, we cannot directly measure “genetic ancestry”, as this would involve collecting genetic information from their ancestors and comparing the participant’s genome to these genomes. Instead, using publicly available data resources, such as the 1000 Genomes Project,22 we can calculate a proxy of genetic ancestry. The 1000 Genomes Project collects genetic information on populations from ethnic groups around the world, and they are categorized into groups referred to as “superpopulations”, including an African group. When we compare participant genetic information to these superpopulations, we are determining a measure of genetic similarity to the individuals in these reference populations, and considering these reference populations to be proxy populations for our participants’ ancestors who originated from these continents23. In this way, our measure of genetic similarity is a proxy for genetic ancestry, and while we refer to genetic ancestry for the remainder of this paper for simplicity, we are truly describing genetic similarity.

Tumor Genetic Ancestry Estimation Pipeline

Broadly, our method to calculate genetic ancestry from tumor tissue gene expression data consists of aligning RNASeq reads to the human genome, calling single nucleotide polymorphisms (SNPs), and comparing the called SNPs to a reference panel with a standard method of global genetic ancestry inference (Figure 1A).

Figure 1.

Figure 1.

A: Tumor RNASeq-derived global genetic ancestry pipeline. B: Distribution of germline DNA-derived global genetic ancestry among Black women diagnosed with epithelial ovarian cancer in the United States, N=184.

First, we converted the paired-end FASTQ files to one FASTA file using SeqKit11. Next, we aligned the FASTA files to the human genome with HISAT2 v2.2.112 using a reference index for GRCh38. This program outputs a SAM file, which was sorted via SAMtools13 in preparation for further analysis. Upon sorting of the SAM file, variant calling was performed via the BCFtools13 mpileup command, producing a VCF. From here, PLINK v1.914 was used to convert the VCF into BED/BIM/FAM format. A total of 1,010 women from the 1000 Genomes Project22 were selected from African, European, South Asian, and East Asian superpopulations. Each superpopulation contained 229-263 women. Women from the American superpopulations were excluded because they were already admixed.

Because the number of SNPs called per participant varies, the linkage to the reference panel and global genetic ancestry calculation is run for one participant at a time to maximize the information each participant’s data provides. PLINK v1.914 was used to merge the sample and the reference panel. We iteratively applied multiple methods of pruning SNPs for quality control (QC) to identify the method with the best performance, described below. After the optimal QC filters were identified, ADMIXTURE15 was run on the clean, merged data. By including the reference populations from the 1000 Genomes Project, ADMIXTURE first trains the model on these individual’s genomes, assuming each of these individuals to be entirely from the population they represent, with a genetic ancestry coefficient equaling 1. Then, with this information, ADMIXTURE iterates through different values of genetic ancestry coefficients ranging from 0 to 1 for the admixed individuals in our sample dataset until the estimates converge. ADMIXTURE outputs these genetic ancestry coefficients, and with the framework previously described, we consider them to be a proxy for global genetic ancestry proportions. Superpopulation genetic ancestry was summarized as European, African, and Asian (collapsing the combined proportion of South Asian and East Asian due to lower proportions).

Quality Control Filters

Standard QC

QC thresholds using PLINK were iterated through standard values to identify the ideal QC setting for tumor RNASeq-derived SNP data for genetic ancestry calculation. Minimum allele frequency (MAF) and Hardy-Weinberg Equilibrium (HWE) p-values were allowed to vary through 0.05, 0.01, and 0.001.

FFPE Artifacts

To assess the impact of formalin-fixation on genetic ancestry correlations, a sensitivity analysis removing C→T and G→A mutations, features which are far more abundant in FFPE tissue24, was done. As samples degrade over time25,26, generally producing less coverage within the sequencing, we inspected the distribution of SNPs, number of reads, and percentage of low-quality reads called from the tumor RNASeq data by time between tumor collection and extraction. We used the median time between tumor collection and extraction (9 years) to dichotomize the data into tumor samples 9 years and older vs tumor samples collected within 9 years. Additionally, the correlation between genetic ancestry estimates was compared separately for these two groups.

Intronic SNP QC

We examined the distribution of the SNPs called from the tumor RNASeq data across exons and introns and discovered a large proportion of intronic SNPs. Although this has been reported in the literature27, we retained only the most likely and highest quality intronic SNPs during the alignment step: multi-mapping was disabled, and the value for alignment score validity was increased. This was done by modifying the options of -k and --score-min in HISAT2, respectively.

Minimum SNP Threshold

While there is a suggested minimum number of SNPs to use ADMIXTURE to calculate genetic ancestry (10,000 per the developer’s recommendations15), there was wide variation in the number of SNPs called from RNASeq data across individuals, and we thought it relevant to determine if a higher minimum is required when beginning from noisier. A range of minimum SNP thresholds were considered to identify the influence of the number of SNPs on genetic ancestry calculation. Correlations at each possible threshold, from the minimum number of SNPs called to the maximum, were calculated and plotted.

Using Different Alignment Software

To inspect the variability in downstream SNP calling and genetic ancestry estimation, a second alignment method, called STAR28 was run instead of HISAT2, and the rest of the pipeline steps remained the same.

Germline Genetic Ancestry Estimation

Germline-derived global genetic ancestry was also calculated via ADMIXTURE15 with the same reference panel of 1,010 women from the 1000 Genomes project.

Correlation of Genetic Ancestry Estimates

After both germline DNA- and tumor RNASeq-derived global genetic ancestry were calculated, genetic ancestry was summarized into three groups: African, European, and Asian. The proportions of each genetic ancestry in the tumor and in the germline were plotted against each other, and the Pearson’s correlation coefficients between the values were calculated. These coefficients were calculated at every QC setting investigated: for each MAF & HWE combination, at each level of minimum SNP threshold, when assessing FFPE artifacts, by time between sample collection and sequencing, and for each alignment software.

Data Availability Statement:

The germline DNA data set is currently available at the database of Genotypes and Phenotypes (dbGaP) under accession number phs001882.v1.p1 (OncoArray – FOCI data). The RNASeq data set is currently available at dbGaP under accession number phs003632.v1.p1. The rest of the data generated in this study are available upon reasonable request from the corresponding author.

RESULTS

Distribution of Germline DNA-Derived Genetic Ancestry

In SchildkrautB, the highest proportion of global genetic ancestry was African genetic ancestry, with a median of 0.84 [Min: 0.56, Max: 0.98], followed by European genetic ancestry, with a median of 0.14 [Min: <0.01, Max: 0.43] and genetic Asian ancestry with a median of 0.01, [Min: <0.01, Max: 0.12; Figure 1B].

Variant Calling from FFPE tumor RNAseq

Among the 184 women with germline and tumor data, the median number of SNPs called from the RNASeq data was 31,379 SNPs, with a wide range (Min: 1,469, Max: 173,937). More recently collected tumors (<9 years) had more SNPs called (Median: 35,094, Min: 1,760, Max: 120,429) than those collected 9+ years prior (Median: 27,908, Min: 1,469, Max: 173,937, Mann-Whitney U test p=0.003; Figure 2A). While not statistically significant, total reads and percentage of low-quality reads appeared to be partially related to tissue age. All samples with fewer than 25,000,000 reads were among those collected 9+ years prior to extraction, as were samples with more than 1% of reads found to be low quality (Figure 2B-C).

Figure 2.

Figure 2.

Density distribution of number of SNPs called (A), total reads (B), and percentage of low-quality reads (C) by age of the tumor sample.

Genetic ancestry correlation

The correlation between African, European, and Asian genetic ancestry calculated from germline DNA tumor RNASeq from women in SchildkrautB ranged from 0.75-0.83, 0.70-0.80, and 0.15-0.44, respectively (Table 1). Correlations between African subpopulations derived from tumor RNASeq and germline DNA were low and unstable, ranging from −0.07-0.34 for Esan genetic ancestry, −0.02-0.16 for Gambian genetic ancestry, −0.14-0.14 for Luhya genetic ancestry, 0.11-0.20 for Mende genetic ancestry, and −0.11-0.10 for Yoruba genetic ancestry (Supplementary Figure S1 & Supplementary Table S1).

Table 1.

Pearson correlation coefficient (r) and 95% confidence interval (CI) between germline DNA- and tumor RNASeq-derived genetic ancestry in SchildkrautB at different quality control (QC) metrics, N=184.

MAF HWE African Genetic
Ancestry
European Genetic
Ancestry
Asian Genetic
Ancestry
r 95% CI r 95% CI r 95% CI
0.05 0.05 0.81 (0.76, 0.86) 0.70 (0.62, 0.77) 0.15 (0.01, 0.29)
0.05 0.01 0.82 (0.77, 0.87) 0.74 (0.66, 0.80) 0.24 (0.10, 0.37)
0.05 0.001 0.83 (0.78, 0.87) 0.76 (0.69, 0.82) 0.29 (0.15, 0.42)
0.01 0.05 0.79 (0.73, 0.84) 0.78 (0.71, 0.83) 0.30 (0.16, 0.43)
0.01 0.01 0.79 (0.73, 0.84) 0.79 (0.73, 0.84) 0.39 (0.26, 0.51)
0.01 0.001 0.80 (0.74, 0.85) 0.80 (0.74, 0.84) 0.34 (0.21, 0.47)
0.001 0.05 0.78 (0.71, 0.83) 0.78 (0.71, 0.83) 0.26 (0.12, 0.39)
0.001 0.01 0.77 (0.70, 0.82) 0.77 (0.71, 0.82) 0.44 (0.31, 0.55)
0.001 0.001 0.75 (0.67, 0.80) 0.75 (0.68, 0.80) 0.33 (0.19, 0.45)
Removing C→T & G→A
0.05 0.001 0.81 (0.76, 0.86) 0.76 (0.69, 0.82) 0.25 (0.11, 0.38)

Optimal QC Filters

After iterating through different values for MAF and HWE, MAF ≥ 0.05 and HWE p-value ≥ 0.001 provided the highest Pearson correlation coefficients for African and European genetic ancestry estimates (Table 1). At this optimal pruning level, the correlations between germline- and tumor-derived genetic African, European, and Asian genetic ancestry proportions were 0.83, 0.76, and 0.29, respectively. A scatter plot of the correlation of genetic ancestry estimates at optimal QC metrics is shown in Figure 3A. Generally, African and European genetic ancestry were slightly underestimated, and Asian genetic ancestry was overestimated. After removing C→T and G→A conversions, which are known to be inflated among FFPE tissue, the correlation was not impacted. After iterating through minimum SNP thresholds, the estimates for African and European genetic ancestry were most improved with a 32,995 SNP minimum to 0.94 and 0.86, respectively (Figure 3B). 100 women (54%) of the women in SchildkrautB did not meet the 32,995 SNP threshold. (Further information on the annotation of the SNPs overall, as well as among women who met this optimal SNP threshold compared with women who did not, can be found in Supplementary Table S2.)

Figure 3.

Figure 3.

Scatter plots of correlation between germline DNA- and tumor RNASeq-derived genetic ancestry at optimal QC metrics; A. MAF=0.05, HWE=0.001, B. Correlation of genetic ancestries at different tumor RNASeq-derived SNP thresholds. The optimal threshold was >=32,995 SNPs (n=84). C. Correlation between germline DNA-derived and tumor RNASeq-derived estimates of genetic ancestry by tissue age at extraction. AFR=African genetic ancestry, ASI=Asian genetic ancestry, EUR=European genetic ancestry.

The correlation between Asian genetic ancestry estimates was optimized to 0.44 using MAF ≥ 0.001 and HWE ≥ 0.01 (Table 1). In this setting, the correlation between African and European genetic ancestry estimates was similar, at 0.77 for both estimates.

When comparing estimates by tissue age, tumors collected within the last 9 years had higher correlations for African and European genetic ancestries, at 0.91 and 0.89, respectively (Figure 3C). However, correlations among the older samples were still relatively high at 0.78 for African genetic ancestry and 0.71 for European genetic ancestry.

Using Different Alignment Software

When using STAR instead of HISAT2 as the alignment software, the correlations between germline and tumor estimates were comparable (African Pearson’s r=0.87, European Pearson’s r=0.79, Asian Pearson’s r=0.27).

Comparison Between Women with and without Germline DNA Available

When we calculated genetic ancestry from tumor RNASeq among women without corresponding germline DNA available, we compared the overall distribution of the estimates to those of women with germline DNA available. We plotted the distributions of each subset’s tumor RNASeq derived genetic ancestry and found the estimates to be consistent, regardless of which analytical group they belonged to (Figure 4).

Figure 4.

Figure 4.

Violin plots of the distribution of tumor-derived genetic ancestry among Black women diagnosed with epithelial ovarian cancer in the informative subset with germline DNA available (n=184) and the exploratory subset without germline DNA available (n=58).

DISCUSSION

We developed a pipeline of existing software that provides for the accurate estimation of genetic ancestry using RNASeq data derived from archival FFPE tumor tissue. While genetic ancestry estimation has been conducted using gene expression data for specific subtypes within breast cancer8, our pipeline does not rely on a set of specific differential profiles or zero in on certain regions or genes, and instead leverages as much data as possible. Such genetic ancestry estimation that can stand on its own can be used in tissue repositories that do not have corresponding germline DNA samples, such as pathology repositories and studies on rare diseases. Our method is robust to alignment and genetic ancestry software packages, QC settings, tissue age, and FFPE artifacts. While our pipeline has some fine-tuning options, foregoing some of these options still results in accurate genetic ancestry estimates.

To optimize the estimation of genetic ancestry proportions, using continental groups as the smallest unit for prediction, we excluded rare variants (MAF < 0.05) and potential genotyping errors (HWE < 0.001). Extreme p-values for HWE tend to reflect genotyping errors, so this more stringent threshold may have improved the correlation coefficient by removing problematic SNPs14. The default for MAF thresholds is typically 0.01, so relaxing the MAF to 0.05 is likely balancing the exclusion of rare SNPs while preserving enough SNPs to allow for accurate prediction. Further, setting a minimum SNP threshold (approx. 33,000) allowed for even greater confidence in estimation, although only just under half of the participants met this metric. Using more tumor samples from more recent diagnoses improved the estimates but was not required for solid estimates of global genetic ancestry. The largest outliers were removed when setting a minimum SNP threshold, intuitively showing that more SNPs generate more reliable estimates. Reassuringly, this method is robust to the inflated levels of C→T and G→A mutations caused by FFPE, and the information gleaned from these SNPs is still useful when analyzing genetic ancestry. Additionally, the findings of proportions of genetic ancestry among Black women in the United States with ovarian cancer are in line with prior research characterizing US admixed populations29,30. Where there are slight variations, accounting for measurement error can overcome this and lead to reasonable effect estimates.

Limitations to this method include that Asian genetic ancestry, a smaller but non-zero contribution to this population (Figure 1B), is not consistently estimated across germline DNA and tumor RNASeq data. In our sample, there is a low proportion of Asian genetic ancestry when estimated from germline DNA, but a higher proportion estimated from RNASeq data. This could be due to noise and more chance for error within the tumor RNASeq data.

If the research question is to assess the impacts of a genetic population that is known to be a small proportion in a study sample, careful consideration of quality control should be made. Additionally, disentangling continental genetic ancestries into finer detail is unstable and unreliable. We attempted to calculate African subpopulation genetic ancestry with tumor data, but the correlation with germline-derived subpopulation genetic ancestry was low and inconsistent. Despite these caveats, combining tumor RNASeq-derived estimates can be used to increase power in downstream analysis.

Archival tumor tissue can confidently and efficiently approximate global genetic ancestry in an admixed population when germline DNA is unavailable. This can be incorporated into various downstream analyses focused on examining genetic ancestry in cancer. We believe this pipeline will address an existing need to distinguish the contribution of genetic ancestry from other factors in many analyses of cancers and to increase sample size and representation of admixed groups in molecular epidemiologic studies, especially for uncommon cancers where the only sequencing data available is tumor-derived RNASeq. Many large consortia have sparse germline data available, so the approach we describe here will allow for more investigation of genetic ancestry within studies that may not have been previously able to do so.

Supplementary Material

1
2
3

Acknowledgments:

This work was supported by NCI of NIH (R01 CA237318 to J.M. Schildkraut, M.P. Epstein, J.R. Marks, C.E. Johnson, & L.C. Peres, R01 CA200854 to J.A. Doherty, J.M. Schildkraut, J.R. Marks, C.S. Greene, N.R. Davidson, & C.E. Johnson, R01 CA142081 to J.M. Schildkraut, R01 CA076016 to J.M. Schildkraut, R01 CA188943 to J.M. Schildkraut, R01 CA237170 to C.S. Greene & J.A. Doherty, K99/R00CA218681 to L.C. Peres and National Cancer Institute GAME-ON Post-GWAS Initiative U19-CA148112 to J.M. Schildkraut). Research reported in this publication utilized the High-Throughput Genomics and Cancer Bioinformatics Shared Resource at Huntsman Cancer Institute at The University of Utah and was supported by NCI of the NIH under award number P30CA042014. We would also like to thank the AACES investigators Anthony J. Alberg, Elisa V. Bandera, Jill Barnholtz-Sloan, Melissa Bondy, Michele L. Cote, Ellen Funkhouser, Edward Peters, Ann G. Schwartz, Paul Terry, and Patricia G. Moorman. This study would not have been possible without the efforts of the North Carolina Central Tumor Registry and all of the staff of the NCOCS. We also thank Rex C. Bentley for the review of the pathology in the NCOCS and Rex C. Bentley and Ann M. Mills for the review of the pathology in the AACES.

Footnotes

Conflict of interest:

The authors declare no potential conflicts of interest.

References

  • 1.Arora K, Tran TN, Kemel Y, Mehine M, Liu YL, Nandakumar S, et al. Genetic Ancestry Correlates with Somatic Differences in a Real-World Clinical Cancer Sequencing Cohort. Cancer Discov. 2022;12(11):2552. doi: 10.1158/2159-8290.CD-22-0312 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Iyer HS, Zeinomar N, Omilian AR, Perlstein M, Davis MB, Omene CO, et al. Neighborhood Disadvantage, African Genetic Ancestry, Cancer Subtype, and Mortality Among Breast Cancer Survivors. JAMA Netw Open. 2023;6(8):e2331295. doi: 10.1001/jamanetworkopen.2023.31295 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Lee KK, Rishishwar L, Ban D, Nagar SD, Mariño-Ramírez L, McDonald JF, et al. Association of Genetic Ancestry and Molecular Signatures with Cancer Survival Disparities: A Pan-Cancer Analysis. Cancer Res. 2022;82(7):1222. doi: 10.1158/0008-5472.CAN-21-2105 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Martini R, Delpe P, Chu TR, Arora K, Lord B, Verma A, et al. African Ancestry-Associated Gene Expression Profiles in Triple-Negative Breast Cancer Underlie Altered Tumor Biology and Clinical Outcome in Women of African Descent. Cancer Discov. 2022;12(11):2530–2551. doi: 10.1158/2159-8290.CD-22-0138 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.McHugh J, Saunders EJ, Dadaev T, McGrowder E, Bancroft E, Kote-Jarai Z, et al. Prostate cancer risk in men of differing genetic ancestry and approaches to disease screening and management in these groups. Br J Cancer. 2021;126(10):1366. doi: 10.1038/s41416-021-01669-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Myer PA, Lee JK, Madison RW, Pradhan K, Newberg JY, Isasi CR, et al. The Genomics of Colorectal Cancer in Populations with African and European Ancestry. Cancer Discov. 2022;12(5):1282–1293. doi: 10.1158/2159-8290.CD-21-0813 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Tamayo LI, Day-Friedland E, Zavala VA, Marker KM, Fejerman L. Genetic Ancestry and Breast Cancer Subtypes in Hispanic/Latina Women. In: Ramirez AG, Trapido EJ, eds. Advancing the Science of Cancer in Latinos: Building Collaboration for Action. Springer; 2023: 79–88. Accessed November 14, 2024. http://www.ncbi.nlm.nih.gov/books/NBK595792/ [Google Scholar]
  • 8.Telonis AG, Rodriguez DA, Spanheimer PM, Figueroa ME, Goel N. Genetic Ancestry-specific Molecular and Survival Differences in Admixed Patients With Breast Cancer. Ann Surg. 2024;279(5):866–873. doi: 10.1097/SLA.0000000000006135 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Wang X, Hou K, Ricciuti B, Alessi JV, Li X, Pecci F, et al. Additional impact of genetic ancestry over race/ethnicity to prevalence of KRAS mutations and allele-specific subtypes in non-small cell lung cancer. HGG Adv. 2024;5(3):100320. doi: 10.1016/j.xhgg.2024.100320 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Schildkraut JM, Johnson C, Dempsey LF, Qin B, Terry P, Akonde M, et al. Survival of epithelial ovarian cancer in Black women: a society to cell approach in the African American cancer epidemiology study (AACES). Cancer Causes Control CCC. 2023;34(3):251–265. doi: 10.1007/s10552-022-01660-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Shen W, Sipos B, Zhao L. SeqKit2: A Swiss army knife for sequence and alignment processing. iMeta. 2024;3(3):e191. doi: 10.1002/imt2.191 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Kim D, Paggi JM, Park C, Bennett C, Salzberg SL. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat Biotechnol. 2019;37(8):907–915. doi: 10.1038/s41587-019-0201-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Danecek P, Bonfield JK, Liddle J, Marshall J, Ohan V, Pollard MO, et al. Twelve years of SAMtools and BCFtools. GigaScience. 2021;10(2):giab008. doi: 10.1093/gigascience/giab008 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Chang CC, Chow CC, Tellier LC, Vattikuti S, Purcell SM, Lee JJ. Second-generation PLINK: rising to the challenge of larger and richer datasets. GigaScience. 2015;4:7. doi: 10.1186/s13742-015-0047-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Alexander DH, Novembre J, Lange K. Fast model-based estimation of ancestry in unrelated individuals. Genome Res. 2009;19(9):1655–1664. doi: 10.1101/gr.094052.109 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Schildkraut JM, Alberg AJ, Bandera EV, Barnholtz-Sloan J, Bondy M, Cote ML, et al. A multi-center population-based case–control study of ovarian cancer in African-American women: the African American Cancer Epidemiology Study (AACES). BMC Cancer. 2014;14:688. doi: 10.1186/1471-2407-14-688 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Moorman PG, Schildkraut JM, Calingaert B, Halabi S, Vine MF, Berchuck A. Ovulation and ovarian cancer: a comparison of two methods for calculating lifetime ovulatory cycles (United States). Cancer Causes Control. 2002;13(9):807–811. doi: 10.1023/A:1020678100977 [DOI] [PubMed] [Google Scholar]
  • 18.Davidson NR, Barnard ME, Hippen AA, Campbell A, Johnson CE, Way GP, et al. Molecular subtypes of high-grade serous ovarian cancer across racial groups and gene expression platforms. Cancer Epidemiol Biomark Prev. 2024;33(8):1114–1125. doi: 10.1158/1055-9965.EPI-24-0113 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.O’Connell J, Gurdasani D, Delaneau O, Pirastu N, Ulivi S, Cocca M, et al. A general approach for haplotype phasing across the full spectrum of relatedness. PLoS Genet. 2014;10(4):e1004234. doi: 10.1371/journal.pgen.1004234 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Manichaikul A, Peres LC, Wang XQ, Barnard ME, Chyn D, Sheng X, et al. Identification of novel epithelial ovarian cancer loci in women of African ancestry. Int J Cancer. 2020;146(11):2987–2998. doi: 10.1002/ijc.32653 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Phelan CM, Kuchenbaecker KB, Tyrer JP, Kar SP, Lawrenson K, Winham SJ, et al. Identification of 12 new susceptibility loci for different histotypes of epithelial ovarian cancer. Nat Genet. 2017;49(5):680–691. doi: 10.1038/ng.3826 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Auton A, Abecasis GR, Altshuler DM, Durbin RM, Abecasis GR, Bentley DR, et al. A global reference for human genetic variation. Nature. 2015;526(7571):68–74. doi: 10.1038/nature15393 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.National Academies of Sciences, Engineering, and Medicine; Division of Behavioral and Social Sciences and Education; Health and Medicine Division; Committee on Population; Board on Health Sciences Policy; Committee on the Use of Race, Ethnicity, and Ancestry as Population Descriptors in Genomics Research. Using Population Descriptors in Genetics and Genomics Research: A New Framework for an Evolving Field. National Academies Press; (US: ); 2023: 60-64, 127, 132–134. Accessed April 18, 2025. http://www.ncbi.nlm.nih.gov/books/NBK589855/ [PubMed] [Google Scholar]
  • 24.Guo Q, Lakatos E, Bakir IA, Curtius K, Graham TA, Mustonen V. The mutational signatures of formalin fixation on the human genome. Nat Commun. 2022;13:4487. doi: 10.1038/s41467-022-32041-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Ura H, Niida Y. Comparison of RNA-Sequencing Methods for Degraded RNA. Int J Mol Sci. 2024;25(11):6143. doi: 10.3390/ijms25116143 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Schuierer S, Carbone W, Knehr J, Petitjean V, Fernandez A, Sultan M, et al. A comprehensive assessment of RNA-seq protocols for degraded and low-quantity samples. BMC Genomics. 2017;18(1):442. doi: 10.1186/s12864-017-3827-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Eghbalnia HR, Wilfinger WW, Mackey K, Chomczynski P. Coordinated analysis of exon and intron data reveals novel differential gene expression changes. Sci Rep. 2020;10(1):15669. doi: 10.1038/s41598-020-72482-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Dobin A, Davis CA, Schlesinger F, Drenkow J, Zaleski C, Jha S, et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics. 2013;29(1):15–21. doi: 10.1093/bioinformatics/bts635 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Zakharia F, Basu A, Absher D, Assimes TL, Go AS, Hlatky MA, et al. Characterizing the admixed African ancestry of African Americans. Genome Biol. 2009;10(12):R141. doi: 10.1186/gb-2009-10-12-r141 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Bryc K, Durand EY, Macpherson JM, Reich D, Mountain JL. The Genetic Ancestry of African Americans, Latinos, and European Americans across the United States. Am J Hum Genet. 2015;96(1):37–53. doi: 10.1016/j.ajhg.2014.11.010 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1
2
3

Data Availability Statement

The germline DNA data set is currently available at the database of Genotypes and Phenotypes (dbGaP) under accession number phs001882.v1.p1 (OncoArray – FOCI data). The RNASeq data set is currently available at dbGaP under accession number phs003632.v1.p1. The rest of the data generated in this study are available upon reasonable request from the corresponding author.

RESOURCES