Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Mar 1.
Published in final edited form as: Pancreas. 2025 Mar 1;54(3):e171–e178. doi: 10.1097/MPA.0000000000002408

Somatic Genomic Profiling of Pancreatic Ductal Adenocarcinomas from a Diverse Cohort of Patients

Andrea N Riner 1, Enrique I Velazquez-Villarreal 13, Seeta Rajpara 2, Jing Qian 2, Yuxin Jin 2, Donna Loza 2, Ashwin Akki 3, Kelly M Herremans 1, Rohit Raj 4, Terence M Williams 5, Nipun Merchant 6, Thomas J George 7, Steven J Hughes 1, Mariana C Stern 8, Renee Reams 9, Ken Redda 9, Diana J Wilkie 10, Folakemi T Odedina 11, Srikar Chamala 12, Bo Han 2, Edward Agyare 9, David W Craig 13, John D Carpten 13, Jose G Trevino 14
PMCID: PMC12203564  NIHMSID: NIHMS2091834  PMID: 39999309

Abstract

Objectives:

Black/African American (B/AA) pancreatic ductal adenocarcinoma (PDAC) patients have worse clinical outcomes than White patients and are underrepresented in genomic databases. We aimed to expand our understanding of the PDAC somatic landscape from a diverse cohort.

Methods:

FFPE specimens from 24 surgically resected PDAC cases were collected, with self-reported race/ethnicity. Whole exome sequencing was performed on malignant and benign tissue. Bioinformatics analysis included deduction of genetic ancestry and somatic mutational analysis, with comparisons to public datasets.

Results:

Out of 24 cases, 17 identified as B/AA race; genetic ancestry analysis confirmed proportions of Sub-Saharan African ancestry greater than 47%. The most commonly mutated genes included KRAS, TP53, SMAD4, and CDKN2A. Comparison of mutations in our cohort versus publicly available, predominantly White datasets showed higher mutation frequencies of ATM, RREB1, BRCA1/2, KDM6A, ARID1A, BRAF and MYC (p < 0.04). When cohorts were combined and analyzed by race, no mutation frequencies differences were observed, including KRAS.

Conclusion:

Genomic analysis of PDAC tumors from B/AA and White patients demonstrate similarities in mutation frequencies. Larger studies are needed to further understand molecular characterizations across continental subpopulations. This study provides further rationale for equitable representation of diverse patients in genomic databases and clinical trials.

Keywords: somatic mutations, pancreatic ductal adenocarcinoma, disparities

Introduction

Pancreatic ductal adenocarcinoma (PDAC) remains a lethal disease that is plagued by racial and ethnic disparities, reflected in lower one-year relative survival in 2017 for Black/African American (B/AA) patients compared to White (W) patients across all stages of disease, despite relatively similar rates of disease presentation (B/AA: 58.1% localized, 51.7% regional, 15.3% distant; W: 63.9% localized, 58.5% regional, 21.8% distant), across all stages of disease.[1] Although health equity is heavily impacted by social determinants of health and access to care, disparities may also be influenced by molecular differences in tumor biology. Somatic and germline variants, and alternative RNA splicing frequencies differ across diverse populations, which potentially alters cancer development, pharmacokinetics, therapeutic resistance, response, and toxicity to targeted therapies.[24] In addition to tumor characteristics, pharmacokinetics may also be affected by drug metabolizing enzymes that vary by genetic ancestry.[57]

Despite the observed disparities for PDAC among B/AA individuals, there is a paucity of data on genomic drivers and potential targetable alterations in PDAC among racially/ethnically diverse populations because these patients have historically been underrepresented in cancer research studies and clinical trials. To date, key whole exome sequencing (WES) studies of PDAC have not been inclusive of B/AA or Hispanic/Latino patients, including large national datasets such as The Cancer Genome Atlas (TCGA)[8]. Prior et al. collated mutation frequencies from Catalogue of Somatic Mutations in Cancer (COSMIC), cBioPortal, International Cancer Genome Consortium (ICGC) and TCGA to estimate Ras mutation frequencies and found that 88% of PDAC cases harbor KRAS mutations, but racial/ethnic subgroup data were not presented.[9] To our knowledge, differences in KRAS mutation frequencies among PDAC patients with diverse ancestral heritage has yet to be studied. Interestingly, a systematic review of KRAS mutation rates in colorectal cancer, another cancer with prevalent KRAS mutations, demonstrated increased frequency of KRAS mutations among B/AA compared to White patients.[10] Among KRAS mutations across all cancers, approximately 70% are G12D, G12V, G12C, G13D or Q61R mutations.[11] In PDAC, the most common KRAS mutations are G12D (~40%−50%) and G12V (~30%−34%).[11, 12] Although G12C mutations are less common, targeted inhibitors of G12C-mutated KRAS have been developed, such as sotorasib, which showed promise in the CodeBreaK100 trial.[13] Taking KRAS as an example, specific mutations and mutation frequencies among PDAC patients from diverse ancestral heritages has not yet been explored despite this being the most ubiquitous somatic mutation in PDAC. Such lack of biologic underpinnings of the disease could contribute to even larger disparities in precision oncology among these patients.[14]

Lack of diversity in genomic databases and clinical investigations leaves providers with incomplete data on molecular drivers of cancer, safety, and efficacy of cancer therapeutics, and therefore potentially limits our ability to provide optimized personalized care and make strides toward eliminating PDAC disparities. Although targetable molecular alterations remain relatively uncommon in PDAC,[15] precision medicine has potential to improve survival and is the future of cancer care, particularly relative to clinical activity for targeting oncogenic KRAS.[16] To address this gap of knowledge, we sought to conduct exome sequencing to identify mutations implicated in key cancer genes in PDAC tumors from B/AA and Hispanic patients with the aim of deepening our understanding of potential differential drivers of PDAC that could contribute to disparities.

Methods

Sample Preparation

Retrospectively collected de-identified formalin-fixed paraffin-embedded (FFPE) pancreatic ductal adenocarcinoma (PDAC) pancreatectomy specimens were obtained. This study was approved by the Institutional Review Boards of the University of Florida (IRB201600873), The Ohio State University and University of Miami (IRB 2014CC0077 and 20060858, respectively). Clinical data were obtained from retrospective chart review. For each case, serial 5-micron sections from FFPE blocks were placed on unstained slides. One section from each case was used for hematoxylin and eosin (H&E) staining using standard approaches.[17] H&E sections from each case were used to mark regions of PDAC, uninvolved normal, and where identified, pancreatitis by an expert histopathologist. These H&E slides were used as maps for scrape microdissection from 5–10 serial sections from unstained slides. Thus, for each case, we collected pancreatic ductal adenocarcinoma, uninvolved normal, and where identified, pancreatitis tissue.

For this study we utilized histologically confirmed archival FFPE PDAC tissue samples surgically resected from newly diagnosed patients. Hematoxylin and eosin stained sections were histologically characterized and marked for regions of tumor and normal/uninvolved pancreas tissue regions. Unstained serial sections were used to collect tumor and normal/uninvolved pancreas tissue by scrape microdissection. DNA was then extracted for WES of matched tumor and normal/uninvolved from each case for somatic mutational analysis.

DNA Isolation

Genomic DNA and RNA was extracted from FFPE tissue using the Covaris truXTRAC© FFPE total NA Column kit (Woburn, MA) based on the manufacturer’s recommendations. Briefly, tissue samples were placed into Covaris Adaptive Focused Acoustics (AFA) and suspended in Lysis buffer followed by Proteinase K digestion. Tissue mixtures were then equally split for separate DNA and RNA extraction. Tissue mixtures were then ultrasonically emulsed using the Covaris E220 System. DNA or RNA sample emulsions were then separately chemically de-crosslinked and pipetted onto appropriate spin columns. DNA and RNA were then collected using provided elution buffers. DNA and RNA were quality assessed and quantified using both the (Thermo Fisher Scientific, Inc. Waltham, MA) and the Genomic DNA Screen Tape Assay utilizing the Agilent 4200 TapeStation System (Santa Clara, CA). DNA and RNA samples were stored at −80°C.

Whole exome sequencing

For WES we utilized a custom expanded exome bait set (Agilent Technologies, Inc.). Briefly, components of the expanded exome included the following probe groups: original baits from SureSelect Human All Exon V6, (Agilent Technologies, Inc.) and custom baits for select genomic regions.[18] Genomic DNA in the amount of 50–200 nanograms (ng) from pancreatic tumor tissue and paired uninvolved (normal) pancreatic tissue from each case was sheared in 50 microliters (μl) of TE low EDTA buffer employing the Covaris E220 system (Covaris, Inc., Woburn, MA) to target fragment sizes of 150 – 200 bp. Fragmented DNA was then converted to an adapter-ligated whole genome library using the Kapa Hyper Prep Library Prep kit (Kapa Biosciences, Inc., Wilmington, MA) according to the manufacturer’s protocol. SureSelect XT Adaptor Oligo Mix was utilized in the ligation step (Agilent Technologies, Inc.). Individual tumor adapter-ligated libraries were enriched into the exome capture reaction, and for germline each adapter-ligated library was pooled before proceeding to capture using Agilent’s SureSelect Human All Exon V6 + custom probes capture library kit. Samples that had successful libraries created were then sequenced on Illumina MiSeq technology for quality control to assess the ability of the libraries to be sequenced. Subsequently, each library was pooled and sequenced on Illumina’s NovaSeq 6000 (Illumina, San Diego, CA) using 300 cycle kit. Raw FASTQs were generated using the industry standard BCL2FASTQ v1.8.4. Mean target coverage was 105X for tumor samples and 55X for uninvolved control samples.

Sequencing data analytics

All sequencing reads were converted to industry standard FASTQ files using the Bcl Conversion and Demultiplexing tool (Illumina, Inc). Sequencing reads were aligned to the GRCh38 reference genome using the MEM module of BWA v0.7.17[19] and SAMTOOLS v1.9[19] to produce BAM files. After alignment, the base quality scores were recalibrated and joint indel realignment was performed on the BAM files using GATK v4.0.10.1.[20] Duplicate read pairs were marked using PICARD v2.18.22.[21] Final BAM files were then used to identify germline and somatic events. Germline SNP and INDELS were identified using GATK haplotype caller in the constitutional sample.

We used GATK HaplotypeCaller v4.0.10.1 to generate germline VCF files for ancestry analysis. Known population genotype data were downloaded from the 1000 Genomes Project phase 3 and grouped by super population for ancestry analysis.[22] VCF subsetting and merging were performed using VCFtools v0.1.17 and SnpSift v4.3t.[23, 24] VCF genotype allele coding (0/1) was converted to numeric (012) using VCFtools and PLINK v1.90b6.7.[25] All ancestry analyses were performed on autosomal chromosomes. Global population admixture was estimated using STRUCTURE v2.3.4.[26] We subset and merged all VCF files by a list of 1766 ancestry-informative markers (AIMs). The result was converted to Structure format by PLINK. To run STRUCTURE, we set up population numbers k=5, NUMREPS=2000, and BURNIN=50000. Principal component analysis (PCA) was performed on the same dataset by R function prcomp. Local Ancestry analysis was performed using the tool LAMP-LD v1.3.[27] Five super population ancestral haplotype files were prepared using an expanded AIM list of 20803 SNPs (single nucleotide polymorphisms). To run LAMP-LD, we set the window length to 50 SNPs, and the number of states was set to 20. All ancestry analysis results were visualized using R v3.6.0 ggplot2 v3.4.1[28] package. For initial analysis, genotypes for a total of 1,766 AIMs were extracted and used for estimating global genetic ancestry based upon five ancestral super continental populations including African (AFR), American Indian (AMR), European (EUR), East Asian (EAS), and South Asian (SAS) using STRUCTURE (Figure 1). We further used LAMP-LD to estimate local ancestry for 20,803 SNPs extracted from WES data across the genome for deducing chromosome level admixture for each individual relative to the same five super populations (Figure 2).

Figure 1:

Figure 1:

Ancestry and admixture Patterns. This figure illustrates the frequencies of admixture in each tumor sample (P02 - P62) from a cohort of 24 PDAC patients. The admixed genomes are categorized into five super populations: AFR (blue), AMR (red), EAS (yellow), EUR (green), and SAS (purple).

Figure 2:

Figure 2:

Chromosome-Level Admixture Patterns. This figure presents the local ancestry estimation of each tumor sample (P02 - P62) at the chromosomal level within a cohort of 24 PDAC patients. The admixed genomes of each chromosome are grouped into five super populations: AFR (blue), AMR (red), EAS (yellow), EUR (green), and SAS (purple).

Somatic variant callers identified single nucleotide variants (SNVs) and small insertions and deletions (indels). Somatic mutation analysis was performed by identifying overlapping variant calls generated by Strelka[29] and Mutect.[30] Due to the diffuse nature of these tumor samples, we retained somatic mutations that had a minor allele frequency of 4%, and manually reviewed variants in the KRAS region using IGV.[31] We compared these mutations against germline results to confirm and verify the somatic mutations. After filtering and manually reviewing somatic variants, VCFs were then annotated with Ensembl’s VEP[32] tool and converted to the MAF file format for visualization purposes. MAFtools[33] was utilized in the generation of oncoplots, annotated with demographic information associated with each sample.

Results

Sample description and whole exome sequencing (WES)

Mean age at diagnosis for the 24 newly diagnosed patients was 65 years (Table 1). Self-reported race was available for 19 of 24 patients, with 17 patients self-reporting as Black (B/AA) and 2 self-reporting as White. Among Black patients, 7 self-reported as Non-Hispanic, 6 identified as Hispanic, and 11 did not indicate ethnicity (Table 1). A total of 6 patients identified as Hispanic, with two of them also identifying as White and the others not reporting race. In addition to self-reported race and ethnicity, we also deduced genetic ancestral proportions and admixture for each patient using germline Ancestry Informative Markers (AIMs) from WES data. Among our cohort, 70.8% self-reported as B/AA, all with higher proportions of Sub-Saharan African Ancestry than the other Ancestry components. However, as expected, our cases have varying degrees of genetic admixture, including those who self-report as B/AA. WES data summary quality metrics and statistics are provided in Supplementary Table 1.

Table 1.

Patient Demographics, Clinical Characteristics and Genomic Ancestry Proportions. This table provides details on patient demographics (race and ethnicity were self-reported) and clinical characteristics, along with the proportions of admixed genomes classified into five super populations: AFR (African), AMR (Amerindian), EAS (East Asian), EUR (European), and SAS (South Asian). The admixed genomic ancestry proportions offer insights into the continental ancestral origins of the patients’ genomes.

Sample ID Sex Race Ethnicity AFR AMR EAS EUR SAS Age at resection Tumor stage Histology (1=well, 2=mod, 3=poor, 4=undiff) Survival from date of surgery (months)
P02 M Black Non-Hispanic 0.9 0 0 0.1 0 49 3 2 8.9
P06 F Black Non-Hispanic 1 0 0 0 0 69 2 1-to-2 37.8
P07 F White Hispanic 0.2 0.1 0 0.7 0 54 3 3 10.2
P11 M NA Non-Hispanic 0.8 0.1 0 0 0 62 3 2 79.5
P14 M Black Non-Hispanic 0.9 0 0 0.1 0 68 3 2 16.5
P20 F Black Non-Hispanic 0.7 0 0 0.3 0 72 3 2 26.3
P22 M Black Non-Hispanic 0.8 0 0 0.2 0 69 2 2 16.2
P25 F Black Non-Hispanic 0.9 0 0 0.1 0 74 2 2 14.8
P26 M White Hispanic 0.2 0.1 0 0.6 0.1 58 3 2 11.6
P31 NA NA Hispanic 0.3 0.2 0 0.6 0 NA NA NA NA
P32 NA NA Hispanic 0.2 0.3 0 0.3 0.2 NA NA NA NA
P34 NA NA Hispanic 0.2 0 0 0.7 0 NA NA NA NA
P35 NA NA Hispanic 0.2 0.1 0.3 0.4 0 NA NA NA NA
P46 M Black Unknown 0.7 0 0 0.3 0 63 3 2-to-3 25
P47 F Black Unknown 0.5 0 0 0.3 0.2 72 3 2-to-3 24
P48 M Black Unknown 0.7 0 0 0.3 0 62 3 2-to-3 67
P49 M Black Unknown 0.6 0.1 0 0.3 0 65 1 2 35
P50 M Black Unknown 0.5 0 0 0.3 0.2 65 3 2 13
P51 M Black Unknown 0.9 0 0 0.1 0 67 3 2-to-3 15
P55 F Black Unknown 0.8 0 0 0.1 0 63 3 2 12
P59 M Black Unknown 0.7 0 0 0.3 0 54 2 1 33
P60 M Black Unknown 1 0 0 0 0 68 3 2 35
P61 F Black Unknown 0.9 0.1 0 0 0 69 2 2 16
P62 F Black Unknown 0.9 0 0 0.1 0 78 2 2 8

Somatic mutation frequencies in PDAC

Somatic mutations were detected in 24 PDAC cases using WES data of tumor and matched uninvolved normal DNA. We were able to calculate the variant allele frequency of somatic mutations for each tumor, and on average the variant allele frequency was 9.8, suggesting that the tumor cellularity across cases was <20%. This relatively low tumor cellularity is indicative of the diffuse nature of PDAC. An annotated oncoplot of mutated genes and frequencies in our PDAC cases is illustrated in Figure 3. The most commonly mutated gene in our cohort was KRAS, which was mutated in all cases (100%). Other genes known to be commonly mutated in PDAC were identified in our cohort, including TP53 (75%), SMAD4 (21%) and CDKN2A (21%).

Figure 3:

Figure 3:

Oncoplot of PDAC Identifies Driver Genes and Somatic Mutations. Top Panel: Total Mutational Burden (TMB) calculated for each sample. Middle Panel: Driver genes identified based on the type of mutation; light blue: Truncating/Frameshift Deletion; dark blue: Missense Mutation; light green: Nonsense Mutation; and dark green: Multi-hit. Mutation frequencies are shown for each gene. Bottom Panel: Race/Ethnicity (self-reported); green: Black Non-Hispanic; blue: White Hispanic; purple: White Non-Hispanic; red: Not Provided.

We compared the frequencies of commonly mutated cancer genes from our PDAC cohort to two publicly available PDAC genomic datasets including 160 (4% B/AA) cases from The Cancer Genome Atlas Firehouse Legacy (TCGA-FL) cohort[34] and 140 (1% B/AA) cases from the National Cancer Institute Clinical Proteomic Tumor Analysis Consortium (CPTAC) cohort[35] (Table 2). In all three datasets KRAS was the most commonly mutated gene. Mutation frequencies for KRAS, TP53, CDKN2A, and SMAD4, were similar across all groups (Table 2). Genes with highly statistically significant differences were ATM and RREB1, which demonstrated higher frequencies in our cohort (Table 2). Other genes with notable variant differences include ARID1A, BRCA1, BRCA2, KDM6A, BRAF, NF1 and MYC (Table 2).

Table 2.

Mutation Frequencies across PDAC Studies. This table includes gene names, mutation frequencies from the current study (USC), and frequencies from external studies (TCGA and CPTAC). The included p-values demonstrate the significance levels when comparing frequencies with our study.

Gene Our Study Our Study TCGA Firehose Legacy P-values CPTAC P-values
Mutated Samples (N=24) (N=160) (USC vs. TCGA) (N=140) (USC vs. CPTAC)
KRAS 24 100.00% 91.70% 0.55 96.40% 0.8
TP53 18 75.00% 68.40% 0.58 75.00% 1
CDKN2A 5 20.80% 16.50% 0.48 20.70% 0.99
SMAD4 5 20.80% 25.60% 0.48 17.10% 0.55
ARID1A 4 16.70% 6.00% 0.02 5.70% 0.02
ATM 4 16.70% 3.00% 0.002 1.40% 0.0003
BRCA1 3 12.50% 1.50% 0.003 0.70% 0.001
BRCA2 3 12.50% 1.50% 0.003 0.70% 0.001
KDM6A 3 12.50% 3.00% 0.016 1.40% 0.003
RNF43 3 12.50% 6.00% 0.13 6.40% 0.16
RREB1 3 12.50% 6.00% 0.13 0.00% 0.0004
BRAF 2 8.30% 1.50% 0.03 0.70% 0.01
NF1 2 8.30% 2.30% 0.07 0.70% 0.01
TGFBR2 2 8.30% 5.30% 0.42 3.60% 0.17
GNAS 1 4.20% 3.80% 0.89 5.00% 0.8
PALB 1 4.20% 0.80% 0.13 0.70% 0.11
PRSS1 0 0.00% 3.00% 0.08 0.70% 0.4
MYC 1 4.20% 0.00% 0.04 0.70% 0.11

To further determine if there are possible differences in the frequency of commonly mutated genes in PDAC we explored somatic mutations in PDAC cases from the AACR Project GENIE dataset[36] (Table 3), which also includes self-identified race information. We combined cases from our cohort, TCGA-FL, CPTAC, and AACR Genie, while stratifying by self-reported race. This dataset included a total of 273 (6%) cases who self-identified as B/AA and 4226 (94%) cases who self-identified as White (Table 3). Based upon this large cohort, we did not identify any statistically significant differences in the frequency of commonly mutated genes by self-identified race stratification (Table 3).

Table 3.

Mutation frequencies in in Black/African American (B/AA) and White cases across our study, TCGA-FL, CPTAC and AACR GENIE. It lists gene names and their respective mutation frequencies as percentages for each racial group. The last column presents p-values, indicating the statistical significance of the observed differences in mutation frequencies between B/AA and White patients.

Gene Black Patients (N=280) White Patients (N=4612) p-value
KRAS 89.30% 87.60% 0.89
TP53 81.80% 70.50% 0.36
BRCA1 1.80% 1.90% 0.96
BRCA2 3.90% 4.40% 0.86
BRAF 1.40% 1.80% 0.82
KDM6A 4.30% 3.80% 0.86
ATM 3.90% 5.10% 0.69
ARID1A 5.70% 8.40% 0.47
NF1 0.70% 1.50% 0.59

Patterns of KRAS mutations in PDAC

KRAS mutations were present in 100% of the cases that we interrogated in our study, which is similar to rates presented by other studies as shown in Table 2. However, we also wanted to assess the types of KRAS mutations that we identified in our cohort and to see if they are consistent with other larger publicly available PDAC datasets. We detected KRAS G12D, G12V, G12R, T20M, and Q61H mutations in our study (Table 4). Our study cohort had significantly higher frequency of G12V mutations (p=0.001) and T20M mutations (p=0.041), yet lower frequency of G12R mutations (p=0.051). As our PDAC study cohort is enriched for B/AA cases, we determined if there was a difference in the distribution of KRAS mutations by race. Our KRAS allelic mutation case distribution was compared to predominantly non-B/AA cases from TCGA-FL and CPTAC, stratified by available self-reported race information. We did not detect any significant differences in the distribution of KRAS mutation subtypes by race (Figure 4).

Table 4.

Distribution of KRAS Mutation Types in PDAC Studies. This table shows the distribution of mutation codon changes in our PDAC cohort, TCGA Firehouse cohort, and CPTAC cohort.

KRAS Mutation Codon Change Our Study (N=24) TCGA Firehouse PDAC (N=140) CPTAC PDAC (N=136) P-value
G12D 33.33% 43.57% 44.85% 0.172
G12V 45.83% 28.57% 29.41% 0.001
G12R 8.33% 19.29% 16.91% 0.051
Q61H 8.33% 4.29% 4.41% 0.705
Q61R 0.00% 1.43% 1.47% 0.621
G12A 0.00% 0.71% 0.00% 1.000
G12C 0.00% 0.71% 0.74% 0.877
G13C 0.00% 0.71% 0.00% 1.000
G12S 0.00% 0.71% 0.74% 0.877
G13D 0.00% 0.00% 0.74% 0.254
L23V 0.00% 0.00% 0.74% 0.254
T20M 4.17% 0.00% 0.00% 0.041

Figure 4.

Figure 4.

Frequency of KRAS Mutation Types by Race in Combined Cohorts. The figure presents the distribution of KRAS mutation types in Black/African American (B/AA) patients (N=289) and White patients (N=4311) obtained by merging data from our PDAC cohort, TCGA-FL, CPTAC, and AACR Genie. KRAS mutations are categorized by types, with the following color codes: red: G12D; blue: G12V; purple: G12R; orange: Q61H; light blue: Q61R; pink: G12C; brown: Q61L; gray: G12L; black: Q61K; yellow: G12F; dark red: G12S; and dark pink: A146V.

Discussion

To date, the vast majority of our current understanding of PDAC biology has been derived from data collected from largely European descent cohorts, even given observed emerging disparities in disease incidence and outcomes seen among African American patients.[1, 37] These limitations in representation extend to clinical trials, where trial participants have not been representative of the overall PDAC patient population with regard to race and ethnicity.[38] Despite improvements in reporting of race or ethnicity of trial participants over time, it remains imperfect; as an example, the CodeBreaK100 trial in PDAC patients does not report race or ethnicity of participants.[13, 38] While the larger CodeBreak 100 trial in non-small cell lung cancer (NSCLC) included 1.2% B/AA and 2.9% Hispanic or Latino patients in the sotorasib arm versus 0 B/AA and 5.2% Hispanic or Latino patients in the control arm, disproportionately underrepresenting non-White patients with NSCLC.[39] Using a diverse set of cases, albeit relatively small, we were able to begin to assess potential differences in the molecular genomic make up of PDAC across populations through comparison of our results to those presented through publicly available scientific reports.

Our analysis of somatic mutation profiles revealed a relatively similar general pattern of commonly mutated genes in PDAC from B/AA as has been presented in other studies consisting of largely European descent patients. This includes similar mutation frequencies for commonly mutated genes such as KRAS, TP53, CDKN2A and SMAD4. This also included our observing similar KRAS mutation type frequencies in PDAC across racial groups. Understandably, the most significant limitation of our study is the relatively small size of our study. However, in comparison to two of the largest published studies, including TCGA and CPTAC, our PDAC dataset includes a proportionally larger number of self-reported B/AA patients. TCGA-FL contains largely (86%) self-identified Non-Hispanic White cases, while there is a significant limitation of self-identified race information in CPTAC with race data being unavailable for 76% of cases. While AACR Project Genie also has a relatively large number of cases but the percentage of B/AA cases is similarly low compared to White cases (5.7%). Where possible, we attempted to compare our genomic results with that of these other datasets, given their own limitations of relatively small proportional minority representation.

A limitation was in the context of race stratification. While stratifying cohorts based on race or ethnicity can provide crucial insights into health disparities, it’s important to acknowledge that sometimes the availability of public data, particularly based on mutation frequencies, makes it technically infeasible to stratify the cohorts along these lines. This limitation underscores the challenges researchers face in comparing cohorts with differing demographic compositions. However, this very limitation underscores the importance of generating cohorts similar to those in our study. In our research, we were able to molecularly describe a population and stratify it by various variables, as illustrated in Figure 3. Such efforts are crucial for better understanding the nuanced relationships between genetic factors and demographics.

One strategy to better determine possible mutation frequencies differences was to merge public data with our cohort, even though it was small in comparison, to increase our sample size (Table 3). However, no differences were detected at this level of resolution. We believe that by increasing our cohort’s sample size, we could better determine any possible differences that did not occur in this case after the merging process. This assumption arises because we are potentially dealing with genetic admixture within this population, which has been previously reported.

While our analysis considered genetically investigated individuals of Black/African American (B/AA) descent, along with data from publicly available databases categorized by self-reported race/ethnicity, it is important to note the potential for discrepancies between self-reported race/ethnicity and genetic ancestry. However, recent studies have demonstrated a significant correlation between these two methods of obtaining race and ethnicity.

Another limitation of any PDAC study is the diffuse histologic nature of these cancers, which results in reduced tumor cellularity, making mutation detection more difficult without higher sequencing coverage. To account for this, we included predominantly cases that came from surgical resected specimens that allow pathologists to identify areas of high tumor cell concentration, rather than relatively scant biopsy specimens. We have also provided information on the variant allele frequencies in our sequencing data (Supplementary Table 2).

Ancestry analysis offers a nuanced understanding of genetic backgrounds. Within our cohort, comprising individuals predominantly self-identified as B/AA, it was revealed that 70.8% exhibited higher proportions of Sub-Saharan African Ancestry compared to other ancestry components. This observation underscores the complexity of genetic heritage within this population and emphasizes the necessity of a comprehensive ancestry analysis. By elucidating ancestral proportions and admixture patterns, we gain insights beyond conventional racial categories, allowing for more tailored approaches in healthcare and research. Understanding the genetic diversity within populations affected by pancreatic cancer is pivotal for developing targeted interventions and personalized treatment strategies that account for individual genetic ancestry variations. Thus, integrating ancestry analysis, particularly through AIMs derived from WES data, serves as a crucial step towards unraveling the intricate interplay between genetics, race, and cancer susceptibility.

However, this may increase the false negative rate leading to a potential under estimation of overall mutation frequencies. Limiting our samples to those collected from surgery also potentially introduces a selection bias given that more biologically aggressive tumors may be missed since patients would be inoperable at diagnosis. Despite these limitations, our study makes an important contribution by including an ethnic and racial diverse set of samples. Our findings will need to be validated and confirmed in a larger independent study of PDAC tumors from B/AA and Hispanic individuals. Furthermore, although we did not identify differences in somatic DNA alterations, there is a need to also assess other molecular measurements including transcriptional and epigenetic, which could also shed potential light on differences in PDAC incidence and outcomes across racial groups.

Conclusion

In this comprehensive genomic analysis of PDAC tumors from racially and ethnically diverse patients, similarities in mutation frequencies were identified between B/AA and White patients. Although these data suggest that molecular differences in tumor biology may have less contribution to disparities than hypothesized, larger studies with more diverse patent cohorts are needed to further our understanding of the molecular characterization of PDAC across continental subpopulations. In the era of precision medicine and disparate outcomes, this study provides further rationale for equitable representation of all patients in genomic databases and clinical trials.

Supplementary Material

Suppl Table 1

Supplementary Table 1. WES data summary quality metrics and statics.

Suppl Table 2

Supplementary Table 2. Variant allele frequencies from WES data.

Funding:

Research reported in this publication was supported by the National Cancer Institute of the National Institutes of Health under Award Numbers (U54 CA233444, U54 CA233396, and U54 CA233465) as part of the Florida-California Cancer Research, Education & Engagement (CaRE2) Health Equity Center. Authors are also supported by the National Human Genome Research Institute (T32 HG008958 to ANR), National Cancer Institute (R01 CA242003, U54 CA233444, U54 CA233444-03S1, P30CA247796) of the National Institutes of Health and the Joseph and Ann Matella Fund for Pancreatic Cancer Research (JGT). The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. The final peer-reviewed manuscript is subject to the National Institutes of Health Public Access Policy. Drs. Riner and Herremans are also supported by the Collaborative Alliance for Pancreatic Education and Research.

References

  • 1.National Cancer Institute. SEER*Explorer: An interactive website for SEER cancer statistics (2023). [Accessed on July 28, 2023. Available at: https://seer.cancer.gov/explorer/].
  • 2.Huang SM, Temple R. Is this the drug or dose for you? Impact and consideration of ethnic factors in global drug development, regulatory review, and clinical practice. Clin Pharmacol Ther. 2008. Sep;84(3):287–94. doi: 10.1038/clpt.2008.144. [DOI] [PubMed] [Google Scholar]
  • 3.Ramamoorthy A, Pacanowski MA, Bull J, et al. Racial/ethnic differences in drug disposition and response: review of recently approved drugs. Clin Pharmacol Ther. 2015. Mar;97(3):263–73. doi: 10.1002/cpt.61. Epub 2015 Jan 20. [DOI] [PubMed] [Google Scholar]
  • 4.Wang BD, Ceniccola K, Hwang S, et al. Alternative splicing promotes tumour aggressiveness and drug resistance in African American prostate cancer. Nat Commun. 2017. Jun 30;8:15921. doi: 10.1038/ncomms15921. Erratum in: Nat Commun. 2017 Sep 27;8:16161. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Zanger UM, Schwab M. Cytochrome P450 enzymes in drug metabolism: regulation of gene expression, enzyme activities, and impact of genetic variation. Pharmacol Ther. 2013. Apr;138(1):103–41. doi: 10.1016/j.pharmthera.2012.12.007. Epub 2013 Jan 16. [DOI] [PubMed] [Google Scholar]
  • 6.Rajman I, Knapp L, Morgan T, et al. African Genetic Diversity: Implications for Cytochrome P450-mediated Drug Metabolism and Drug Development. EBioMedicine. 2017. Mar;17:67–74. doi: 10.1016/j.ebiom.2017.02.017. Epub 2017 Feb 20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Li Y, Steppi A, Zhou Y, et al. Tumoral expression of drug and xenobiotic metabolizing enzymes in breast cancer patients of different ethnicities with implications to personalized medicine. Sci Rep. 2017. Jul 6;7(1):4747. doi: 10.1038/s41598-017-04250-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Spratt DE, Chan T, Waldron L, et al. Racial/Ethnic Disparities in Genomic Sequencing. JAMA Oncol. 2016. Aug 1;2(8):1070–4. doi: 10.1001/jamaoncol.2016.1854. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Prior IA, Hood FE, Hartley JL. The Frequency of Ras Mutations in Cancer. Cancer Res. 2020. Jul 15;80(14):2969–2974. doi: 10.1158/0008-5472.CAN-19-3682. Epub 2020 Mar 24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Staudacher JJ, Yazici C, Bul V, et al. Increased Frequency of KRAS Mutations in African Americans Compared with Caucasians in Sporadic Colorectal Cancer. Clin Transl Gastroenterol. 2017. Oct 19;8(10):e124. doi: 10.1038/ctg.2017.48. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Timar J, Kashofer K. Molecular epidemiology and diagnostics of KRAS mutations in human cancer. Cancer Metastasis Rev. 2020. Dec;39(4):1029–1038. doi: 10.1007/s10555-020-09915-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Waters AM, Der CJ. KRAS: The Critical Driver and Therapeutic Target for Pancreatic Cancer. Cold Spring Harb Perspect Med. 2018. Sep 4;8(9):a031435. doi: 10.1101/cshperspect.a031435. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Strickler JH, Satake H, George TJ, et al. Sotorasib in KRAS p.G12C-Mutated Advanced Pancreatic Cancer. N Engl J Med. 2023. Jan 5;388(1):33–43. doi: 10.1056/NEJMoa2208470. Epub 2022 Dec 21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Riner AN, Girma S, Vudatha V, et al. Eligibility Criteria Perpetuate Disparities in Enrollment and Participation of Black Patients in Pancreatic Cancer Clinical Trials. J Clin Oncol. 2022. Jul 10;40(20):2193–2202. doi: 10.1200/JCO.21.02492. Epub 2022 Mar 22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Flaherty KT, Gray RJ, Chen AP, et al. ; NCI-MATCH team. Molecular Landscape and Actionable Alterations in a Genomically Guided Cancer Clinical Trial: National Cancer Institute Molecular Analysis for Therapy Choice (NCI-MATCH). J Clin Oncol. 2020. Nov 20;38(33):3883–3894. doi: 10.1200/JCO.19.03010. Epub 2020 Oct 13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Pishvaian MJ, Blais EM, Brody JR, et al. Overall survival in patients with pancreatic cancer receiving matched therapies following molecular profiling: a retrospective analysis of the Know Your Tumor registry trial. Lancet Oncol. 2020. Apr;21(4):508–518. doi: 10.1016/S1470-2045(20)30074-7. Epub 2020 Mar 2. Erratum in: Lancet Oncol. 2020 Apr;21(4):e182. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Ramesh PS, Madegowda V, Kumar S, et al. DNA extraction from archived hematoxylin and eosin-stained tissue slides for downstream molecular analysis. World J Methodol. 2019. Nov 14;9(3):32–43. doi: 10.5662/wjm.v9.i3.32. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Manojlovic Z, Christofferson A, Liang WS, et al. Comprehensive molecular profiling of 718 Multiple Myelomas reveals significant differences in mutation frequencies between African and European descent cases. PLoS Genet. 2017. Nov 22;13(11):e1007087. doi: 10.1371/journal.pgen.1007087. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Li H, Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics. 2009. Jul 15;25(14):1754–60. doi: 10.1093/bioinformatics/btp324. Epub 2009 May 18. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.McKenna A, Hanna M, Banks E, et al. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res. 2010. Sep;20(9):1297–303. doi: 10.1101/gr.107524.110. Epub 2010 Jul 19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Broad Institute. Picard toolkit (2023). [Accessed on July 28, 2023. Available at: https://broadinstitute.github.io/picard/].
  • 22.1000 Genomes Project Consortium; Auton A, Brooks LD, Durbin RM, Garrison EP, Kang HM, Korbel JO, Marchini JL, McCarthy S, McVean GA, Abecasis GR. A global reference for human genetic variation. Nature. 2015. Oct 1;526(7571):68–74. doi: 10.1038/nature15393. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Danecek P, Auton A, Abecasis G, et al. ; 1000 Genomes Project Analysis Group. The variant call format and VCFtools. Bioinformatics. 2011. Aug 1;27(15):2156–8. doi: 10.1093/bioinformatics/btr330. Epub 2011 Jun 7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Cingolani P, Patel VM, Coon M, et al. Using Drosophila melanogaster as a Model for Genotoxic Chemical Mutational Studies with a New Program, SnpSift. Front Genet. 2012. Mar 15;3:35. doi: 10.3389/fgene.2012.00035. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Chang CC, Chow CC, Tellier LC, et al. Second-generation PLINK: rising to the challenge of larger and richer datasets. Gigascience. 2015. Feb 25;4:7. doi: 10.1186/s13742-015-0047-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Falush D, Stephens M, Pritchard JK. Inference of population structure using multilocus genotype data: linked loci and correlated allele frequencies. Genetics. 2003. Aug;164(4):1567–87. doi: 10.1093/genetics/164.4.1567. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Baran Y, Pasaniuc B, Sankararaman S, et al. Fast and accurate inference of local ancestry in Latino populations. Bioinformatics. 2012. May 15;28(10):1359–67. doi: 10.1093/bioinformatics/bts144. Epub 2012 Apr 11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Wickham Hadley. Ggplot2: Elegant Graphics for Data Analysis. 2nd ed. Springer International Publishing, 2016. 10.1007/978-3-319-24277-4. [DOI] [Google Scholar]
  • 29.Kim S, Scheffler K, Halpern AL, et al. Strelka2: fast and accurate calling of germline and somatic variants. Nat Methods. 2018. Aug;15(8):591–594. doi: 10.1038/s41592-018-0051-x. Epub 2018 Jul 16. [DOI] [PubMed] [Google Scholar]
  • 30.Benjamin D, Sato T, Cibulskis K, et al. Calling Somatic SNVs and Indels with Mutect2. bioRxiv 861054; doi: 10.1101/861054. [DOI] [Google Scholar]
  • 31.Robinson JT, Thorvaldsdóttir H, Wenger AM, et al. Variant Review with the Integrative Genomics Viewer. Cancer Res. 2017. Nov 1;77(21):e31–e34. doi: 10.1158/0008-5472.CAN-17-0337. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.McLaren W, Gil L, Hunt SE, et al. The Ensembl Variant Effect Predictor. Genome Biol. 2016. Jun 6;17(1):122. doi: 10.1186/s13059-016-0974-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Mayakonda A, Lin DC, Assenov Y, et al. Maftools: efficient and comprehensive analysis of somatic variants in cancer. Genome Res. 2018. Nov;28(11):1747–1756. doi: 10.1101/gr.239244.118. Epub 2018 Oct 19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.National Cancer Institute. TCGA Research Network. [Accessed on July 28, 2023. Available at: https://www.cancer.gov/ccg/research/genome-sequencing/tcga].
  • 35.Edwards NJ, Oberti M, Thangudu RR, et al. The CPTAC Data Portal: A Resource for Cancer Proteomics Research. A Resource for Cancer Proteomics Research. J Proteome Res. 2015. Apr 15. [DOI] [PubMed] [Google Scholar]
  • 36.The AACR Project GENIE Consortium. AACR Project GENIE: Powering Precision Medicine Through An International Consortium, Cancer Discov. 2017. Aug;7(8):818–831. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Giaquinto AN, Miller KD, Tossas KY, Winn RA, et al. Cancer statistics for African American/Black People 2022. CA Cancer J Clin. 2022. May;72(3):202–229. doi: 10.3322/caac.21718. Epub 2022 Feb 10. [DOI] [PubMed] [Google Scholar]
  • 38.Herremans KM, Riner AN, Winn RA, et al. Diversity and Inclusion in Pancreatic Cancer Clinical Trials. Gastroenterology. 2021. Dec;161(6):1741–1746.e3. doi: 10.1053/j.gastro.2021.06.079. Epub 2021 Aug 17. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.de Langen AJ, Johnson ML, Mazieres J, et al. ; CodeBreaK 200 Investigators. Sotorasib versus docetaxel for previously treated non-small-cell lung cancer with KRASG12C mutation: a randomised, open-label, phase 3 trial. Lancet. 2023. Mar 4;401(10378):733–746. doi: 10.1016/S0140-6736(23)00221-0. Epub 2023 Feb 7. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Suppl Table 1

Supplementary Table 1. WES data summary quality metrics and statics.

Suppl Table 2

Supplementary Table 2. Variant allele frequencies from WES data.

RESOURCES