Summary
Identifying risk protein targets and their therapeutic drugs is crucial for effective cancer prevention. Here, we conduct integrative and fine-mapping analyses of large genome-wide association studies data for breast, colorectal, lung, ovarian, pancreatic, and prostate cancers and characterize 710 lead variants independently associated with cancer risk. Through mapping protein quantitative trait loci (pQTLs) for these variants using plasma proteomics data from over 75,000 participants, we identify 365 proteins associated with cancer risk. Subsequent colocalization analysis identifies 101 proteins, including 74 not reported in previous studies. We further characterize 36 potential druggable proteins for cancers or other disease indications. Analyzing >3.5 million electronic health records, we conducted analyses of emulated trials for 11 drugs across 290 comparisons and identified three drugs significantly associated with reduced colorectal cancer risk: caffeine vs. paroxetine (hazard ratio [HR], 0.51; 95% confidence interval [CI], 0.41–0.64), haloperidol vs. prochlorperazine (HR, 0.47; 95% CI, 0.33–0.68), and trazodone hydrochloride vs. paroxetine (HR, 0.49; 95% CI, 0.38–0.63). Conversely, caffeine was associated with increased cancer risk in comparison with finasteride (colorectal cancer) and fluoxetine (breast cancer). Meta-analysis identified six drugs significantly associated with cancer risk, including acetazolamide, which was associated with reduced colorectal cancer risk (HR, 0.79; 95% CI, 0.72–0.87). This study identifies previously unreported protein biomarkers and candidate drug targets across six major cancer types and highlights several approved drugs with potential chemopreventive effects.
Keywords: human cancers, risk variants, GWAS, pQTL, proteomics, EHR, therapeutic drug, caffeine, acetazolamide, drug trial emulation
This study established an analytical framework that integrates large-scale cancer GWASs, proteomic data, and EHRs to identify druggable risk proteins and potential therapeutic candidates for cancer prevention. The authors reported 36 potentially druggable proteins and demonstrated that several approved drugs (e.g., acetazolamide) may confer preventive effects against colorectal cancer.
Introduction
Human genetic research has not only advanced our understanding of disease mechanisms but has also significantly contributed to drug discovery and development. Drugs supported by genetic evidence exhibit enhanced therapeutic validity compared to those lacking such support, highlighting the importance of incorporating genetic evidence in drug development initiatives.1,2 Common risk variants implicated in diseases can dysregulate nearby gene or protein expression, which can mimic the effects of therapeutic drugs on the targetable proteins. These proteins could serve as potential targets for therapeutic intervention.3 Thus, concerted efforts for cancer prevention based on proteins influenced by common polymorphisms that modulate cancer risk are urgently needed.4 To date, genome-wide association studies (GWASs) have identified several hundred common genetic risk loci for each of three prevalent cancer types—breast, colorectal, and prostate5,6,7,8—and several dozen risk loci have been identified for other cancers, such as cancers of the lung, pancreas, and ovarian.9,10,11,12,13 Previous research, including our work, has identified hundreds of putative cancer-susceptible genes potentially regulated by these risk variants, using methods such as expression quantitative trait locus (eQTL) analysis8,9,10,11,12,14,15,16,17,18,19,20 and transcriptome-wide association studies (TWASs).7,19,21,22,23,24,25,26,27,28,29 However, most dysregulated gene expressions have not been thoroughly investigated at the protein level.
To deepen the understanding of causal mechanisms and enhance drug discovery endeavors, it is imperative to explore data from transcriptomic to proteomic studies. Proteins, the ultimate products of mRNA translation, play critical roles in cellular activities and represent promising therapeutic targets, as evidenced by successful drug targeting of enzymes, transporters, ion channels, and receptors.30 Recent studies include protein QTL (pQTL) mapping and Mendelian randomization (MR) analysis by integrating cancer GWASs and blood proteomics data to identify potential risk proteins. However, only a few dozen cancer risk proteins have been reported, with a false discovery rate (FDR) < 0.05.31,32,33,34,35,36 Most reported proteins have not been directly linked to the GWAS-identified risk variants in common cancer types. Furthermore, research is lacking in integrating multiple population-scale proteomic studies, such as the recent emerging UK Biobank Pharma Proteomics Project (UKB-PPP),37 which offers an unprecedented opportunity to establish extensive pQTL databases, accelerating therapeutic drug discovery for therapeutic prevention and intervention in human cancers.
Traditional drug discovery faces numerous challenges, including escalating costs, lengthy timelines, and high failure rates.38 Drug repurposing presents a promising strategy by identifying new applications for existing drugs, leveraging their well-documented characteristics.39 With the widespread adoption of modern electronic health record (EHR) systems, vast amounts of real-world patient data are available to augment pre-clinical outcomes and facilitate drug repurposing screening. Recently, drug repurposing using EHRs has successfully discovered repurposing hypotheses for preventing Alzheimer disease,40 reducing cancer mortality,41,42 and treating COVID-1943,44 and coronary artery disease.45 However, for therapeutic drugs that have been used long term to treat disease indications with evidence of affecting the expression of cancer risk proteins, their potential association with the risk of human cancers remains largely unclear. Some of these drugs may be linked to an increased cancer risk due to long-neglected side effects.
In this work, we integrate large GWAS data for breast, colorectal, lung, ovarian, pancreatic, and prostate cancers and population-scale proteomics data from over 75,000 participants combined from the Atherosclerosis Risk in Communities (ARIC) study,46 deCODE genetics,47 and the UKB-PPP to identify risk proteins associated with each cancer. We further characterized therapeutic drugs based on druggable risk proteins targeted by approved drugs or undergoing clinical trials for cancer treatment or other indications. We further evaluate the effect of cancer risk for those drugs approved for the indications, using over 3.5 million EHRs from the database at Vanderbilt University Medical Center (VUMC). Findings from this study offer additional insights into therapeutic drugs targeting risk proteins for cancer prevention and intervention.
Material and methods
This study is covered by IRB#250258 approval for a population-based study using de-identified omics and electronic health records.
Characterization of lead variants in six types of cancer
We performed a comprehensive analysis to characterize the lead genetic variants associated with breast, colorectal, lung, ovarian, pancreatic, and prostate cancers (Figure 1A). For breast cancer, we included strong independent association signals at p < 1 × 10−6 from a fine-mapping study48 and risk variants from another GWAS.6 We combined the reported lead variants from these two studies after removing those variants in linkage disequilibrium (LD) (r2 < 0.1 in European populations) (Figure S1). Next, we further included additional lead variants from SuSiE fine-mapping analysis on GWASs (N = 247,173),49 with fine-mapping windows of 500 kilobases (kb) and allowed a maximum of five causal variants. LD reference was based on the British ancestry UKB samples (N = 337,000).50 We identified a credible set of causal variants with a 95% posterior inclusion probability (95% PIP) for each independent risk signal and a lead variant was represented by the variant with the minimum p. We included additional lead variants from our SuSiE analysis with LD r2 < 0.1 in European populations with the above set of lead variants for those with independent risk-associated signals at GWAS p < 5 × 10−8 and located in GWAS loci with independent risk-associated signals at p < 1 × 10−6 in European populations (Figure S1).
Figure 1.
Overview of the analytical framework
(A) An illustration depicting the identification of proteins associated with the risk of the six major cancers: breast, lung, colorectal, ovarian, pancreatic, and prostate. Population-based proteomics data (for pQTLs) and GWAS data resources (for identifying lead variants) utilized in this study are shown on the left. Meta-analyses of cis-pQTLs from ARIC and deCODE, conducted through the SOMAscan platform, were combined with pQTL results from the UKB-PPP to identify potential risk proteins, as depicted in the middle images. Colocalization analyses between GWAS summary statistics and cis-pQTLs were performed to identify cancer risk proteins with high confidence, as illustrated on the right.
(B) The proteins with evidence of colocalization annotated based on drug-protein information from four databases: DrugBank, ChEMBL, TTD, and Open Targets.
(C) The framework for evaluating the effects of drugs approved for indications on cancer risk. The inverse probability of treatment weighting (IPTW) framework was utilized to construct emulations of treated-control drug trials based on millions of patients’ electronic health records stored at VUMC SD (left). In these emulations, the Cox proportional hazard model was conducted for each trial to assess the hazard ratio (HR) of cancer risk between the treated focal drug and the control drug (right).
For colorectal cancer, we analyzed lead variants from our recent fine-mapping study51 based on the GWAS data from 254,791 participants in both European and East Asian populations. We characterized lead variants with independent risk-associated signals at a minimal p < 1 × 10−6 in European populations, from the analysis based on GWAS from trans-ancestry and European populations, respectively (Figure S1). For prostate cancer, we first identified lead variants with independent risk-associated signals at GWAS p < 5 × 10−8 from our SuSiE fine-mapping analysis on GWAS summary statistics (N = 140,306). We also included additional GWAS-identified risk variants with p < 1 × 10−6 in European populations from the previous trans-ancestry GWAS8 and r2 < 0.1 with any lead variants from the above set in the fine-mapping analysis (Figure S1). Similarly, we used the above strategy to characterize lead variants from fine-mapping analysis for ovarian cancer (N = 63,347) and pancreatic cancer (N = 21,536). Specifically, we utilized an LD reference based on the British ancestry UKB samples (N = 337,000)50 for prostate, ovary, and pancreatic fine-mapping analyses. We next included additional risk variants that are missed in the above set of lead variants from previous GWASs for ovarian11 and pancreatic10 cancer. For lung cancer, we included risk variants at p < 1 × 10−6 in European populations from the trans-ancestry GWAS (N = 85,716).9 Further details on lead variant identification are summarized in Figure S1, and details on GWAS resources can be found in Table S1 and the supplemental material and methods.
Identification of putative target proteins for lead variants
To identify potential cancer risk proteins, we mapped GWAS lead variants to cis-pQTL (±500-kb region of a gene) results from three previous studies among European populations: UKB-PPP,37 ARIC,44 and deCODE45 (Figure 1A). To minimize technical variation for pQTL analyses, stringent quality control (QC) procedures have been conducted in previous projects, including the removal of low-quality samples and proteins. Protein expression levels were normalized across samples, and the resulting values were further inverse-rank normalized. Association analyses were then conducted between protein expression and each genetic variant, adjusting for relevant covariates. We harmonized pQTL results across datasets by aligning reference and effect alleles. Specifically, when the reference and effect alleles were switched in one or more studies, we inverted the sign of the corresponding pQTL beta values. To increase the power of pQTLs, we combined cis-pQTLs from the ARIC and the deCODE (both assayed through the SOMAscan platform) via a fixed-effect meta-analysis using META.52 Cis-pQTLs from the UKB-PPP (assayed through the Olink platform) were independently analyzed. In a few cases where the lead variant did not overlap with any cis-pQTLs, we substituted it with the correlated variant exhibiting the strongest association signal. The putative cancer risk protein was defined based on pQTL significance at a Bonferroni threshold of p < 0.05 (nominal p = 2.3 × 10−5, corresponding to 2,164 variant-protein tests for UKB-PPP; nominal p = 3.8 × 10−5, corresponding to 1,322 variant-protein tests for ARIC+deCODE). We additionally evaluate the pQTL results for the lead variants using data from individuals of African (n = 971) and East Asian (n = 262) ancestry from the UKB-PPP. The putative cancer risk protein was additionally supported based on pQTL significance at a nominal p < 0.05.
Colocalization analyses between pQTL and GWAS signals
To identify cancer risk proteins, we conducted colocalization analysis using two approaches: Bayesian method coloc53 and summary-data-based MR (SMR).54 For the SMR approach, a HEIDI test was performed on significant SMR results to determine if the colocalized signals could be explained by one single causal variant or by linkage. For each protein, SNPs with p < 0.5 from GWASs, a minor-allele frequency (MAF) of >0.01, and within 50 kb of the lead variant were included. To estimate the posterior probability (PP) of colocalization, we utilized the default priors and coloc.abf function. In our study, we particularly focused on the assumption that one genetic variant is simultaneously associated with both traits, which was quantified by PP.H4. We reported proteins with PP.H4 of >0.5 and >0.8 to support protein discoveries from the colocalization analysis. Additionally, we also performed SMR+HEIDI analysis for significant cis-pQTLs with default parameter settings. Specifically, significant SMR+HEIDI results were defined as a tested locus with a Bonferroni-adjusted SMR p < 0.05 (nominal p = 8.5 × 10−5, corresponding to 589 tests for UKB-PPP; nominal p = 2.8 × 10−4, corresponding to 174 tests for ARIC+deCODE) and a HEIDI p ≥ 0.05 (no obvious evidence of heterogeneity of estimated effects or linkage). In addition, we also performed colocalization analyses using European GWASs and cis-pQTLs from African American and East Asian populations (UKB-PPP) to provide further support for the identified risk proteins.
Functional genomic analyses
For our identified cancer risk proteins, we examined their xQTLs, including eQTLs, splicing QTLs (sQTLs), and alternative polyadenylation QTLs (apaQTLs) using the resource from the GTEx (v.8). We collected eQTLs and sQTLs from six normal tissues and whole blood from GTEx studies, and we collected apaQTLs from Li’s work.55 A nominal p < 0.05 for at least one xQTL in either tissue or blood samples was considered supportive of the pQTL results. We also applied a Bonferroni threshold of p < 0.05 for each cancer type to support the pQTL findings (e.g., nominal p = 2.2 × 10−3, corresponding to 23 tests for breast cancer) from the analyses of xQTL, Clinical Proteomic Tumor Analysis Consortium (CPTAC), and The Cancer Genome Atlas (TCGA) below.
We identified putative regulatory variants in strong LD (r2 > 0.8 in the European population) for lead variants with significant colocalization between GWAS and cis-pQTL signals. Using the HaploReg tool,56 we annotated these variants with a variety of epigenetic annotations, including regulatory chromatin states based on DNase and histone chromatin immunoprecipitation sequencing (ChIP-seq) from the Roadmap Epigenomics Project, histone marks for promoters and enhancers, binding sites of transcription factors, and gene annotation from GENCODE and RefSeq. We denoted variants as “proximal” if they overlapped with these functional annotations near the closest target gene. We analyzed a variety of chromatin-chromatin interaction data, from a 4D genome,57 FANTOM5,58 EnhancerAtlas,59 and a super-enhancer.60 We examined the overlap between putative regulatory variants and enhancer elements in corresponding cell lines or tissues of these six cancer types. We further determined enhancer-promoter loops after combining these data with ChIP-seq data of the histone modification H3K27ac (an active enhancer mark). We focused on interacting loops in which a fragment overlapped an H3K27ac peak (enhancer-like elements). In contrast, the other fragment overlapped the promoter of a gene (defined as a region 2 kb upstream and 100 bp downstream of the transcript start site). We denoted variants as “distal” if they overlapped with these chromatin-chromatin variants.
For our identified cancer risk proteins, we assessed the statistical significance of their differential protein expression between tumors and normal tissues in breast, colorectal, lung, ovarian, and pancreatic cancer samples using data from CPTAC, accessed through the UALCAN website.61,62 Similarly, we analyzed their differential gene expression between tumor and normal tissue using data from TCGA, also through the UALCAN website.
Identification of focal drugs and formulation of drug-treated patient groups
We assessed the therapeutic relevance of cancer risk proteins by mapping them to known drug-target pairs using four major databases: DrugBank,63 ChEMBL,64 the Therapeutic Target Database (TTD),65 and Open Targets.66 This approach identified multiple drugs targeting cancer risk proteins, which we refer to as focal drugs (Figure 1B). To evaluate their potential impact on cancer development, we conducted comparative analyses between focal drugs and corresponding control drugs. To minimize potential confounding factors associated with a drug prescription, we selected control drugs that belong to the same second-level Anatomical Therapeutic Chemical classification category (ATC-L2) as the focal drug. For each focal drug, we constructed emulated clinical trials, each comprising a treated group (patients prescribed the focal drug) and a control group (patients prescribed a control drug). One focal drug may have multiple trials, depending on the number of potential control drugs belonging to the same ATC-L2 category. After defining focal-control drug pairs, we selected patients from the Synthetic Derivative (SD), a de-identified EHR database of 3.5 million individuals at VUMC.67 Using strict eligibility criteria (supplemental material and methods), we ensured each trial included at least 500 patients in both the treated and control groups.40
Calculation of overall hazard ratio for focal drug on cancer development risk
To rigorously evaluate the impact of focal drugs on cancer development, we applied the inverse probability of treatment weighting (IPTW) framework68 to reduce confounding and ensure balanced baseline characteristics between the treated and control groups (supplemental material and methods). Within each eligible trial emulation, we used the weighted Cox proportional hazards model to estimate the hazard ratio (HR) for cancer risk, comparing patients exposed to the focal drug vs. those exposed to the control drug over a 10-year follow-up period (Figure 1C). For each focal drug, we also performed a random-effects meta-analysis across all balanced trials to estimate its overall association with cancer risk (supplemental material and methods). The meta-analysis was condcuted for emulation trials tested and control drugs that belong to the same ATC-L3 category. Finally, to assess the robustness of our findings, we conducted sensitivity analyses using a 5-year follow-up window and evaluated the effects of drug-cancer association based on tested and control drugs belonging to the same ATC-L2 or ATC-L3 category. We applied a Bonferroni threshold of p < 0.05 within all trial sets (accounting for all 290 comparisons across cancers) to report significant drug-cancer associations. Similarly, a Bonferroni threshold of p < 0.05 (accounting for the 11 treated drugs evaluated) was used for meta-analyses. This structure defines two distinct families of statistical testing: one at the trial level and another at the meta-analysis. This hierarchical family design reflects the way hypotheses were organized and allows type I error control at both stages. We selected the Bonferroni correction to provide a strong control of false positives given the large number of comparisons and the interdependence among tests.
Results
Characterizing lead variants for breast, ovarian, prostate, colorectal, lung, and pancreas cancers
For breast cancer, we included 196 lead variants with independent association signals at loci from a previous fine-mapping study48 and an additional 32 genetic variants from another GWAS.6 Our additional fine-mapping analysis using SuSiE49 on the Breast Cancer Association Consortium GWAS (N = 247,173) resulted in five previously unreported lead variants. After integrating the previous results with additional fine-mapping efforts, we identified 227 lead variants independently associated with cancer risk at each locus. We applied similar strategies to identify lead variants for colorectal51 cancer and other cancer types, leveraging both previously published GWASs and our fine-mapping analyses (material and methods). In total, we identified 710 lead variants, including 227 for breast cancer, 213 for colorectal cancer, 213 for prostate cancer, 26 for lung cancer, 13 for ovarian cancer, and 18 for pancreatic cancer (Table S2).
Identifying cancer risk proteins from pQTL mapping and colocalization analyses
We mapped the 710 lead variants to cis-pQTLs to identify cancer risk proteins. At a Bonferroni-corrected p < 0.05, we identified a total of 459 pQTL association signals (corresponding to 365 proteins after combining proteins unique to each cancer) for 222 lead variants across six cancer types, including 74 for breast, 127 for colorectal, 37 for lung, 5 for ovarian, 9 for pancreatic, and 113 for prostate cancers (Figure 2; Table S3). Notably, 312 of the identified proteins (85.4% of 365) among these cancer types have not been reported in previous proteomics-based MR studies32,33,34,35,36,69,70 (Table S4). Furthermore, through analysis of the identified proteins commonly observed in multiple cancers, we found that 60 proteins were commonly observed in at least two of these six cancers (Figure S2). In particular, we observed that several well-known cancer-related proteins, such as major histocompatibility complex, class I, A (HLA-A) and major histocompatibility complex, class I, E (HLA-E), were linked to lead variants located in the major histocompatibility complex (MHC) in breast, colorectal, and lung cancers, highlighting the potential role of these proteins in cancer pleiotropy and shared cancer risk mechanisms (Figure 2).
Figure 2.
Genome-wide distribution of lead variants and putative risk proteins among six types of cancer
Proteins identified for each cancer are represented by different colors. Each circle represents a single lead variant-protein pair. For each lead variant, only proteins with the smallest pQTLs are presented in the figure; other proteins are denoted by an asterisk (∗) and listed in Table S3. A dashed box highlights several well-known cancer-related proteins, such as HLA-A and HLA-E, which are linked to lead variants located in the major histocompatibility complex (MHC).
A further colocalization analysis identified 101 proteins linked to 88 GWAS lead variants after combining proteins unique to each cancer that showed strong evidence supported by either colocalization or SMR+HEIDI analysis, including 51 high-confidence cancer risk proteins at PP.H4 > 0.8 (material and methods; Table S5). Specifically, we identified 23 proteins for breast (17 lead variants), 38 proteins for colorectal (32 lead variants), 7 proteins for lung (5 lead variants), 2 for ovarian (1 lead variant), 2 for pancreatic (2 lead variants), and 29 for prostate (31 lead variants) cancers (Figures 3A and 3B; Table S5). Of these, 74 proteins (73.2% of 101) have not been previously linked to cancer risk (Table S6). Of the 101 risk proteins identified, we found 15 proteins mapped to the HLA region,71 including 13 previously unreported candidates and 6 druggable targets. This substantial number of proteins in the HLA region underscores the critical role of immune regulation in cancer etiology, with MHC class I polypeptide-related sequence B (MICB), natural cytotoxicity triggering receptor 3 (NCR3), and HLA-E involved in natural killer (NK) and T cell activation,72,73 and advanced glycosylation end-product specific receptor (AGER) and tripartite motif containing 40 (TRIM40), which modulate nuclear factor κB (NF-κB)-mediated inflammatory responses.74,75 Furthermore, seven proteins were identified as cancer-driver proteins based on previous studies76,77 and included in the Cancer Gene Census (CGC),78 including aldehyde dehydrogenase 2 family (ALDH2), HLA-A, and SUB1 homolog, transcriptional regulator (SUB1) for breast cancer; ALDH2 and HLA-A for colorectal cancer; 5'-nucleotidase, cytosolic II (NT5C2) for lung cancer; and NT5C2, ring finger protein 43 (RNF43), TYRO3 protein tyrosine kinase (TYRO3), and ubiquitin specific peptidase 28 (USP28) for prostate cancer.
Figure 3.
Identification of 101 cancer risk proteins through pQTL and colocalization analyses
(A) Number of proteins showing evidence of colocalizations between pQTLs and GWAS association signals for six cancer types.
(B) Percentage of proteins showing evidence of colocalizations between pQTLs and GWAS summary statistics for six cancer types.
(C) A plot illustrating the high consistency of pQTL p values for 22 cancer risk proteins between the ARIC+deCODE and the UKB-PPP (proteins commonly assayed from the SOMAscan and Olink platforms).
Of note, 71 proteins were only assayed by either the SOMAscan (n = 32) or Olink (n = 39) platform. For the remaining 22 significant proteins commonly assayed, all showed a pQTL significance signal with a minimal nominal p < 1 × 10−5 in both ARIC+deCODE and UKB-PPP (r = 0.66, p = 2 × 10−4; Figure 3C), supporting the robustness of our integrative approach for identifying cancer risk proteins. Additionally, we analyzed pQTL data from individuals of African (n = 971) and East Asian (n = 262) ancestry from the UKB-PPP.37 At a nominal p < 0.05, we identified 50 pQTLs (corresponding to 42 proteins) in the African ancestry population and 26 pQTLs (corresponding to 24 proteins) in the East Asian ancestry population (Table S3). Colocalization analysis additionally revealed nine proteins in the African ancestry population and an additional protein from the East Asian ancestry population with a PP.H4 > 0.5 (Table S5; material and methods).
Cancer risk proteins supported by functional genomics analyses
Of the identified 101 proteins among the six cancers, we examined whether they are supported by functional genomics analyses. Specifically, we first evaluated multi-omics QTL analysis (xQTL, including eQTLs, alternative splicing [sQTLs], and alternative polyadenylation [apaQTLs]) results in their respective target tissues and whole-blood samples (material and methods). We found 63 proteins that were supported by at least one xQTL at a nominal p < 0.05, including 12 for breast (52% of 23), 22 for colorectal (57% of 38), 5 for lung (71% of 7), 2 for ovarian, 2 for pancreatic, and 20 for prostate (68% of 29) cancers (Table S6). At a Bonferroni threshold of p < 0.05, we found 36 proteins supported by at least one xQTL, including 6 from breast cancer; 14 from colorectal cancer; 2 each in lung, ovarian, and pancreatic cancers; and 10 from prostate cancer (Table S6).
Second, we used functional genomic data generated in their cancer-related tissues/cells (i.e., promoters and enhancers) to characterize putative functional variants that are in strong LD (r2 > 0.8 in the European population) with the lead variants (material and methods). Our results showed that 17 genes were likely regulated by the closest putative regulatory variants with either promoter and/or enhancer activities (Table S7). We further investigated the potential distal regulatory effects of putative functional variants on these genes by analyzing chromatin interaction data (material and methods). We found that 39 genes were regulated distally by putative functional variants through long-term promoter-enhancer interactions (Table S8). Lastly, we examined differential protein expression between normal and tumor tissues available for breast, colon, lung, and pancreatic cancers using data from CPTAC. We showed evidence of the 18 identified proteins with consistent association directions being supported by significantly differential expression at a nominal p < 0.05, including 3 for breast cancer, 10 for colorectal cancer, 4 for lung cancer, and 1 for pancreatic cancer (Table S9). At a Bonferroni threshold of p < 0.05, 11 proteins (3 from breast, 7 from colorectal, 4 from lung, and 1 from pancreatic cancers) showed significant differential protein expression and consistent direction with pQTLs from CPTAC (Table S9). Similarly, we showed evidence of the 40 identified proteins supported by significantly differential mRNA expression using data from TCGA Program, including 6 for breast cancer, 15 for colorectal cancer, 3 for lung cancer, 1 for pancreatic cancer, and 15 for prostate cancer (Table S9). At a Bonferroni threshold of p < 0.5, 35 proteins (5 from breast, 13 from colorectal, 3 from lung, 1 from pancreatic, and 13 from prostate cancers) shared significant differential gene expression and consistent direction with pQTLs from TCGA (Table S9).
Furthermore, we conducted gene-set enrichment analyses using Enrichr79 and provided additional evidence that these 101 risk proteins were significantly overrepresented in breast cancer (FDR = 9.4 × 10−5), colorectal carcinoma (FDR = 1.1 × 10−5), squamous cell carcinoma (FDR = 4.2 × 10−5), ovarian neoplasm (FDR = 1.5 × 10−4), pancreatic carcinoma (FDR = 1.1 × 10−7), and prostate neoplasms (FDR = 8.8 × 10−6) (Table S10). Additional analyses based on risk proteins identified for each cancer type were consistent with these observations (Table S10). Taken together, the results of our analysis provided additional evidence that most of the identified proteins are partially or wholly supported by functional genomics analyses.
Identifying druggable proteins and their corresponding drugs
Using data from DrugBank,63 ChEMBL,64 the TTD,65 and Open Targets,66 we comprehensively annotated our proteins as therapeutic targets of approved or clinical-stage drugs. Of the 101 proteins among the six cancers, we identified 36 druggable proteins potentially targeted by 404 approved drugs or undergoing clinical trials for cancer treatment or other indications (Figure 4; Table S11). Specifically, we found 19 proteins targeted by 133 drugs either approved or under clinical trials to treat cancers (Figure 5; Table S12). Our results also provide evidence that the remaining druggable proteins are targeted by 197 drugs used for treating indications other than cancer (Figure S3).
Figure 4.
A circular plot showing 36 druggable proteins potentially targeted by 404 approved drugs or undergoing clinical trials for cancer treatment or other indications
Presented from inner to outer layers are cancer types, proteins, and drugs. Each drug-protein association is annotated by DrugBank, ChEMBL, TTD, and OpenTargets, with lines in different colors representing each database. Associations where proteins are annotated by two databases are linked to drugs with thick lines. The eight protein-drug pairs are highlighted for strong binding affinity, with reported KD, Ki, or IC50 values below 50 nM.
Figure 5.
A circular plot showing 19 druggable proteins potentially targeted by 133 approved drugs or undergoing clinical trials for cancer treatment
Presented from inner to outer layers are cancer types, proteins, drugs, and cancers. Drugs approved and undergoing clinical trials for cancer treatment are highlighted in green and gray, respectively. Approved drug indications are formatted in bold, while indications under clinical trial are in regular font.
Evaluating associations of drugs approved for indications with cancer risk
We next evaluated the effect on cancer risk of therapeutic drugs that have been used long term to treat non-cancer indications based on real-world EHRs from the VUMC SD database. Following the stringent criteria described in material and methods, we formed 290 balanced trials from 11 treated drugs, associated with seven proteins from three cancers (Figure 6A), each with more than 10 eligible control drugs. Specifically, we identified three drugs for breast cancer (caffeine targeting ALDH2 and haloperidol and trazodone hydrochloride, with both targeting HLA-A), several drugs for colorectal cancer (caffeine and theophylline, with both targeting ALDH2; acetazolamide and captopril, with both targeting transferrin (TF); haloperidol and trazodone hydrochloride, with both targeting HLA-A; and vilazodone, targeting advanced glycosylation end-product specific receptor (AGER), and one drug for prostate cancer (sirolimus, targeting TYRO3) (Figure 6B). All patients analyzed had substantial medical records during the 10-year follow-up period, with a median follow-up of ∼7 years (85 months) and an interquartile range (IQR) of 6–8.5 years (72–102 months). Additionally, across the 290 trials, a median of 99.3% of patients were either cancer free at the end of the follow-up period (median: 58%) or were lost to follow-up, while the remaining 0.7% were diagnosed with cancer (Table S14).
Figure 6.
Drugs approved for treated indications showing significant effects on cancer risk
(A) A table showing cancer risk alleles of lead variants, risk proteins, and drug names approved for indications. Positive associations are indicated by upward arrows, while negative associations are indicated by downward arrows. Odds ratio (OR) refers to the exponential transformation of logistic regression coefficient, and beta refers to the linear regression coefficient.
(B) An illustration of six drugs, showing significant associations with cancer risk in meta-analysis and their potential targeted risk proteins.
(C) Kaplan-Meier plots depict the statistically significant difference in the probability of being cancer free for patients in the treated group (taking a focal drug, shown in green) and those in the control group (primarily selected by the smallest p value across eligible treated-control trials for each significant drug in the meta-analyses, shown in purple; for all detailed results, see Table S11). The shaded area represents the 95% confidence interval. The hazard ratio and p for the focal drug compared to the control drug, determined through weighted Cox proportional hazards models, are presented in the top right corner of each panel. Numbers at risk (patients remaining cancer free) are displayed below the x axis for both the treated and control drug groups. The IPTW-weighted event counts at 10 years were as follows: breast cancer, caffeine (721.97) vs. fluoxetine (162.66) and haloperidol (55.05) vs. prochlorperazine (103.57); colorectal cancer: acetazolamide (13.13) vs. alteplase (28.69), captopril (16.04) vs. ramipril (61.51), and vilazodone (12.13) vs. duloxetine (142.33); and prostate cancer: sirolimus (24.39) vs. alteplase (12.63).
For breast cancer, we applied the weighted Cox proportional hazards model to calculate the HR for each trial within the 10-year follow-up period for the three analyzed drugs (caffeine, haloperidol, and trazodone hydrochloride) across 73 eligible treated-control drug trials. Particularly, at a Bonferroni-corrected p < 0.05 (p < 0.05/290), we found that caffeine was associated with an increased risk when compared to fluoxetine (HR = 1.43, 95% CI = [1.19, 1.71]) (Table S14). Furthermore, we conducted a random-effects meta-analysis for each of these three drugs. At a Bonferroni-corrected p < 0.05 (p < 0.05/11), we found that caffeine was associated with an increased risk (HR = 1.15, 95% CI = [1.09, 1.20]), while haloperidol was associated with a decreased risk (HR = 0.86, 95% CI = [0.78, 0.95]) (Table S15; Figure 6C).
For colorectal cancer, we analyzed seven drugs (acetazolamide, caffeine, captopril, haloperidol, theophylline, trazodone hydrochloride, and vilazodone) across 185 eligible treated-control trials. At a Bonferroni-corrected p < 0.05, we found that caffeine was associated with a decreased risk when compared to paroxetine (HR = 0.51, 95% CI = [0.41, 0.64]) and magnesium sulfate (HR = 0.36, 95% CI = [0.27, 0.48]). Conversely, caffeine, when compared to the control drug finasteride (HR = 2.32, 95% CI = [1.67, 3.23]), was associated with an increased risk. We also identified that haloperidol and trazodone hydrochloride were associated with an increased risk when compared to prochlorperazine (HR = 0.47, 95% CI = [0.33, 0.68]) and paroxetine (HR = 0.49, 95% CI = [0.38, 0.63]), respectively (Table S14). Further meta-analyses showed that acetazolamide (HR = 0.79, 95% CI = [0.72, 0.87]) was associated with a decreased risk, while captopril (HR = 2.15, 95% CI = [1.81, 2.57]) and vilazodone (HR = 1.73, 95% CI = [1.45, 2.05]) were associated with an increased risk (Table S15; Figure 6C).
For prostate cancer, we analyzed one drug (sirolimus) across 32 eligible treated-control trials. At a Bonferroni-corrected p < 0.05, none of the trials comparing sirolimus to its control drugs showed a significant association with cancer risk. However, our meta-analysis revealed that sirolimus was significantly associated with an increased risk of prostate cancer (HR = 1.37, 95% CI = [1.2, 1.55]) (Table S15; Figure 6C).
In addition, we conducted several additional analyses to evaluate the robustness of our findings. Firstly, we conducted the heterogeneity analysis among the meta-analysis results for each of the six significant drugs and found no significant heterogeneity for them (Figures S4A–S4F). Secondly, we conducted a leave-one-out analysis for these six drugs, in which each control drug from the ATC-L2 class was excluded in each test. Our results showed that HRs remained consistent across all leave-one-out tests, demonstrating that no single control drug disproportionately influenced the observed associations (Figure S5). Thirdly, we conducted a weighted Cox proportional hazards model using control drugs from ATC-L3, a narrow pharmacological or therapeutic subgroup. We showed that three drugs (colorectal cancer: captopril and vilazodone and prostate cancer: sirolimus) are consistently associated with cancer risk after Bonferroni corrections (Table S15). The remaining three did not reach the significance threshold but showed a consistent direction of association with cancer risk (e.g., HR), likely due to the limited number of controls in the analysis (Figure S6). Thirdly, we conducted analyses using the weighted Cox proportional hazards model within the 5-year follow-up period for the significant treated-control trials identified above, demonstrating consistent associations across all six trials, at a nominal p < 0.05 (Table S16). Similarly, with meta-analyses based on the results from the 5-year follow-up period, we showed five of the identified six drugs (breast cancer: caffeine; colorectal cancer: acetazolamide, captopril, and vilazodone; and prostate cancer: sirolimus) showed consistent associations at a nominal p < 0.05 (Table S17). Lastly, we additionally evaluated the reversal potential of candidate compounds by calculating the negative correlation between cancer-associated gene expression profiles (based on TWAS-derived Z scores) and compound-induced transcriptomic changes from LINCS perturbation experiments (supplemental material and methods). Our analysis provided additional evidence of four compounds exhibiting notable reversal effects (Pearson correlation P ranging from 0.02 to 0.09): haloperidol and trazodone for breast cancer, acetazolamide for colorectal cancer, and sirolimus for prostate cancer (Table S18).
Discussion
In this study, we conducted a comprehensive investigation of cancer risk proteins by integrating lead variants and pQTLs for six common cancer types using large-scale GWAS and population-based proteomics data. Through pQTL mapping and subsequent colocalization analysis, we identified 101 risk proteins across the six cancer types, with over three-quarters of them not previously linked to cancer susceptibility. Moreover, most of the proteins we identified are supported by functional genomics analyses. Our findings not only significantly expand the pool of known cancer risk proteins but also offer additional insights into the biology and susceptibility of common cancers.
Through analysis of drug-protein interaction databases, we identified 36 druggable proteins potentially targeted by 404 therapeutic drugs. Among these, 30 drugs have already received approval for cancer treatment, while 73 are currently undergoing clinical trials for cancer treatment. These findings offer genetic evidence supporting the effectiveness of certain drugs and suggest potential opportunities for repurposing them to treat additional cancers that share common risk proteins. However, it is crucial to acknowledge that while the cancer risk proteins identified in our study hold promise as therapeutic targets for cancer treatment, drugs may also have adverse effects, potentially exacerbating cancer development through these targets (i.e., depending on their inhibitory or promotive effects).80 Additionally, our analysis characterized 197 drugs used for indications other than cancer, which may influence cancer risk due to their interactions with cancer risk proteins. Overall, our findings have the potential to accelerate therapeutic drug discovery for the prevention and intervention of human cancers.
Our results showed that acetazolamide exhibits a notable effect in preventing colorectal cancer development, through targeting the colorectal cancer risk protein transferrin (TF). Previous literature suggested that the TF protein may function similarly to an inhibitor of converting enzyme (e.g., ICA protein), as ICA belongs to the same TF protein superfamily.81,82 Thus, our genetic discovery provides strong evidence supporting acetazolamide as a promising drug for colorectal cancer prevention, potentially influencing a TF-regulated CA etiological pathway. In line with our findings, prior studies demonstrated acetazolamide’s role in inhibiting the cell viability, migration, and colony formation ability of colorectal cancer cells,83 as well as its ability to suppress the development of intestinal polyps in Min mice.84 Interestingly, we also identified another drug, captopril, which potentially targets the cancer risk protein TF and is associated with an increased colorectal cancer risk. Captopril may influence cancer risk by inhibiting angiotensin-converting enzyme (ACE), thereby affecting downstream TF protein levels through the regulation of its receptors.85,86 Consistent with our findings, a previous study also highlighted that ACE inhibition may play a risk-promoting role in colorectal cancer.87
Our results indicated that caffeine, a drug prioritized by the risk protein ALDH2, may have a protective role in colorectal cancer when compared to several control drugs. These findings align with previous studies suggesting a potential protective effect of caffeine in colorectal cancer.88,89 Specifically, caffeine demonstrated a significant protective role in colorectal cancer risk when compared to paroxetine (p = 5.66 × 10−9) and magnesium sulfate (p = 9.10× 10−12). This is supported by previous studies highlighting associations between ALDH2 variants and caffeine consumption90,91 and another study demonstrating that caffeine inhibits ALDH2 activity, resulting in reduced nucleophilicity and partial alterations in the enzyme’s secondary structure.92 Additionally, ALDH2 has been shown to enhance colorectal cancer stemness by activating β-catenin signaling.93,94 Conversely, our results also revealed that caffeine exhibits a risk-promoting role in colorectal cancer when compared to finasteride (p = 4.72 × 10−7). This suggested that finasteride may have a stronger protective effect against colorectal cancer risk than caffeine, as previous studies have shown that finasteride may play a role in the causal pathways of colorectal neoplasms95 and demonstrated effectiveness in the treatment of prostate neoplasms.96,97,98
Our meta-analysis results showed that haloperidol, which targets the known cancer risk immune protein HLA-A, is associated with a reduced breast cancer risk. This contrasts with previous studies suggesting that haloperidol increases breast cancer risk.80,99,100 The meta-analysis result appears to be primarily driven by the most significant individual trial (haloperidol vs. prochlorperazine, HR = 0.69, p = 4.72 × 10−2), which may overshadow other signals showing the risk-promoting role played by haloperidol. Interestingly, our meta-analysis results showed that haloperidol plays a risk-promoting role in colorectal cancer (HR = 1.25, CI = [1.01, 1.54], nominal p = 4.0 × 10−2). Although these findings are in line with the observation that HLA-A is associated with decreased risk in breast cancer and increased risk in colorectal cancer, replication of these findings in an independent dataset is necessary for future studies.
We also uncovered an additional drug, vilazodone, that is associated with an increased colorectal cancer risk by targeting AGER, a protein known to promote chronic inflammation by activating the NF-κB pathway.101 Our study also found that sirolimus, a mechanistic target of rapamycin (mTOR) inhibitor, exhibits a risk-promoting role in prostate cancer, consistent with previous research.102,103 While our findings highlight potential etiological pathways involving these drugs, it is essential to evaluate the effects of the reported candidate drugs through both in vitro and in vivo assays in future research.
To strengthen the statistical power, we conducted a meta-analysis of the pQTL results from the ARIC and deCODE genetics projects, both of which used the same SOMAscan technology for protein assays (covering >4,500 proteins). We noticed that several proteins, such as ABO, cathepsin S (CTSS), paired immunoglobulin-like type 2 receptor alpha (PILRA), CD177 molecule (CD177), MICB, plasminogen (PLG), and asparaginase and isoaspartyl peptidase 1 (ASRGL1), showed significant heterogeneity, while their corresponding pQTLs achieved genome-wide significance in at least one individual dataset, providing strong support for their inclusion as significant risk proteins (Table S3). We examined the concordance of cis-pQTLs for 1,729 proteins measured on both the Olink (UKB-PPP) and SOMAscan (ARIC and deCODE) platforms. We observed a notable concordance for pQTLs with consistent association directions between the UKB-PPP and SOMAscan platforms (correlation r = 0.52; Figure S7). It should also be noted that our identified drug-cancer associations are observational from data rather than causal inferences, as these results are based on VUMC EHR data that may carry residual confounding and bias, even with the application of IPTW to balance covariates. While matching based on similar drug usage profiles (e.g., context, dosage, and frequency) helps reduce indication bias, this approach is constrained by missing data, which may reduce the effective sample size and affect the generalizability of results. Furthermore, the estimated HRs may be influenced by inherent selection biases arising from loss to follow-up, censoring, and competing risks. These issues underscore the need for future work by employing time-dependent models and competing risk frameworks to better capture treatment effects and improve robustness. Further validation of results in an independent EHR dataset would also be beneficial.
Data and code availability
Table S1 provides the download information for the summary statistics of GWAS data for the six common cancers, including breast,6 ovarian,11 prostate,104 colorectal,105 lung,13 and pancreatic.29 Metadata and pQTL summary statistics from UKB-PPP can be downloaded from Synapse with project SynID, syn51364943: https://www.synapse.org/Synapse:syn51364943, and pQTLs from ARIC46 and deCODE genetics47 can be accessed through previous publications. Functional genomic data include TCGA and CPTAC differential expression results accessible through https://ualcan.path.uab.edu/index.html; 4DGenome: https://4dgenome.research.chop.edu/; Depmap: https://depmap.org/portal/; FANTOM5: http://fantom.gsc.riken.jp/5/. HaploReg v.3: https://pubs.broadinstitute.org/mammals/haploreg/; and GTEx v.8: https://gtexportal.org/home/downloads/adult-gtex/qtl. GENCODE (v.26.GRCh38): https://www.gencodegenes.org/human/release_26.html. The National Cancer Institute can be accessed through https://www.cancer.gov/about-cancer/treatment/drugs; CGC can be accessed via the COSMIC website: https://cancer.sanger.ac.uk/census. Drugs and compound data can be downloaded from four drug database websites: ChEMBL: https://www.ebi.ac.uk/chembl/, TTD: https://db.idrblab.net/ttd/, Open Targets: https://www.opentargets.org/, and DrugBank: https://go.drugbank.com/. The EHR data, containing de-identified clinical information, can be accessed through the VUMC SD database: https://victr.vumc.org/data-resource-sd/.
The developed pipeline and main source R codes that are used in this work are available from the GitHub website of X.G.’s lab (https://github.com/XingyiGuo/PQTL_EHR/).
Acknowledgments
This work was supported by the US National Institutes of Health grants 1R37CA227130-01A1 and R01CA269589-01A1 to X.G. and R01CA297582 to X.G. and Z.Y. The data analyses were conducted using the Advanced Computing Center for Research and Education (ACCRE) at Vanderbilt University. This work was supported by the New Frontiers in Research Fund (NFRFE-2018-00748) and an NSERC Discovery Grant (RGPIN-2024-04679) to Q. Long. The computational infrastructure was partly supported by a Canada Foundation for Innovation JELF grant (36605) to Q. Long.
Declaration of interests
The authors declare no competing interests.
Published: December 2, 2025
Footnotes
Supplemental information can be found online at https://doi.org/10.1016/j.ajhg.2025.11.008.
Contributor Information
Zhijun Yin, Email: zhijun.yin.1@vumc.org.
Xingyi Guo, Email: xingyi.guo@vumc.org.
Supplemental information
References
- 1.Nelson M.R., Tipney H., Painter J.L., Shen J., Nicoletti P., Shen Y., Floratos A., Sham P.C., Li M.J., Wang J., et al. The support of human genetic evidence for approved drug indications. Nat. Genet. 2015;47:856–860. doi: 10.1038/ng.3314. [DOI] [PubMed] [Google Scholar]
- 2.Diogo D., Tian C., Franklin C.S., Alanne-Kinnunen M., March M., Spencer C.C.A., Vangjeli C., Weale M.E., Mattsson H., Kilpeläinen E., et al. Phenome-wide association studies across large population cohorts support drug target validation. Nat. Commun. 2018;9:4285. doi: 10.1038/s41467-018-06540-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Finan C., Gaulton A., Kruger F.A., Lumbers R.T., Shah T., Engmann J., Galver L., Kelley R., Karlsson A., Santos R., et al. The druggable genome and support for target identification and validation in drug development. Sci. Transl. Med. 2017;9 doi: 10.1126/scitranslmed.aag1166. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Peters U., Tomlinson I. Utilizing Human Genetics to Develop Chemoprevention for Cancer-Too Good an Opportunity to be Missed. Cancer Prev. Res. 2024;17:7–12. doi: 10.1158/1940-6207.CAPR-22-0523. [DOI] [PubMed] [Google Scholar]
- 5.Jia G., Ping J., Shu X., Yang Y., Cai Q., Kweon S.S., Choi J.Y., Kubo M., Park S.K., Bolla M.K., et al. Genome- and transcriptome-wide association studies of 386,000 Asian and European-ancestry women provide new insights into breast cancer genetics. Am. J. Hum. Genet. 2022;109:2185–2195. doi: 10.1016/j.ajhg.2022.10.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Zhang H., Ahearn T.U., Lecarpentier J., Barnes D., Beesley J., Qi G., Jiang X., O'Mara T.A., Zhao N., Bolla M.K., et al. Genome-wide association study identifies 32 novel breast cancer susceptibility loci from overall and subtype-specific analyses. Nat. Genet. 2020;52:572–581. doi: 10.1038/s41588-020-0609-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Fernandez-Rozadilla C., Timofeeva M., Chen Z., Law P., Thomas M., Schmit S., Díez-Obrero V., Hsu L., Fernandez-Tajes J., Palles C., et al. Deciphering colorectal cancer genetics through multi-omic analysis of 100,204 cases and 154,587 controls of European and east Asian ancestries. Nat. Genet. 2023;55:89–99. doi: 10.1038/s41588-022-01222-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Conti D.V., Darst B.F., Moss L.C., Saunders E.J., Sheng X., Chou A., Schumacher F.R., Olama A.A.A., Benlloch S., Dadaev T., et al. Trans-ancestry genome-wide association meta-analysis of prostate cancer identifies new susceptibility loci and informs genetic risk prediction. Nat. Genet. 2021;53:65–75. doi: 10.1038/s41588-020-00748-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Byun J., Han Y., Li Y., Xia J., Long E., Choi J., Xiao X., Zhu M., Zhou W., Sun R., et al. Cross-ancestry genome-wide meta-analysis of 61,047 cases and 947,237 controls identifies new susceptibility loci contributing to lung cancer. Nat. Genet. 2022;54:1167–1177. doi: 10.1038/s41588-022-01115-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Klein A.P., Wolpin B.M., Risch H.A., Stolzenberg-Solomon R.Z., Mocci E., Zhang M., Canzian F., Childs E.J., Hoskins J.W., Jermusyk A., et al. Genome-wide meta-analysis identifies five new susceptibility loci for pancreatic cancer. Nat. Commun. 2018;9:556. doi: 10.1038/s41467-018-02942-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Phelan C.M., Kuchenbaecker K.B., Tyrer J.P., Kar S.P., Lawrenson K., Winham S.J., Dennis J., Pirie A., Riggan M.J., Chornokur G., et al. Identification of 12 new susceptibility loci for different histotypes of epithelial ovarian cancer. Nat. Genet. 2017;49:680–691. doi: 10.1038/ng.3826. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Lawrenson K., Song F., Hazelett D.J., Kar S.P., Tyrer J., Phelan C.M., Corona R.I., Rodríguez-Malavé N.I., Seo J.H., Adler E., et al. Genome-wide association studies identify susceptibility loci for epithelial ovarian cancer in east Asian women. Gynecol. Oncol. 2019;153:343–355. doi: 10.1016/j.ygyno.2019.02.023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.McKay J.D., Hung R.J., Han Y., Zong X., Carreras-Torres R., Christiani D.C., Caporaso N.E., Johansson M., Xiao X., Li Y., et al. Large-scale association analysis identifies new lung cancer susceptibility loci and heterogeneity in genetic susceptibility across histological subtypes. Nat. Genet. 2017;49:1126–1132. doi: 10.1038/ng.3892. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Wen W., Chen Z., Bao J., Long Q., Shu X.O., Zheng W., Guo X. Genetic variations of DNA bindings of FOXA1 and co-factors in breast cancer susceptibility. Nat. Commun. 2021;12:5318. doi: 10.1038/s41467-021-25670-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Moreno V., Alonso M.H., Closa A., Vallés X., Diez-Villanueva A., Valle L., Castellví-Bel S., Sanz-Pamplona R., Lopez-Doriga A., Cordero D., Solé X. Colon-specific eQTL analysis to inform on functional SNPs. Br. J. Cancer. 2018;119:971–977. doi: 10.1038/s41416-018-0018-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Chen Z., Wen W., Beeghly-Fadiel A., Shu X.O., Díez-Obrero V., Long J., Bao J., Wang J., Liu Q., Cai Q., et al. Identifying Putative Susceptibility Genes and Evaluating Their Associations with Somatic Mutations in Human Cancers. Am. J. Hum. Genet. 2019;105:477–492. doi: 10.1016/j.ajhg.2019.07.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Guo X., Lin W., Bao J., Cai Q., Pan X., Bai M., Yuan Y., Shi J., Sun Y., Han M.R., et al. A Comprehensive cis-eQTL Analysis Revealed Target Genes in Breast Cancer Susceptibility Loci Identified in Genome-wide Association Studies. Am. J. Hum. Genet. 2018;102:890–903. doi: 10.1016/j.ajhg.2018.03.016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.He J., Wen W., Beeghly A., Chen Z., Cao C., Shu X.O., Zheng W., Long Q., Guo X. Integrating transcription factor occupancy with transcriptome-wide association analysis identifies susceptibility genes in human cancers. Nat. Commun. 2022;13:7118. doi: 10.1038/s41467-022-34888-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Guo X., Lin W., Wen W., Huyghe J., Bien S., Cai Q., Harrison T., Chen Z., Qu C., Bao J., et al. Identifying Novel Susceptibility Genes for Colorectal Cancer Risk From a Transcriptome-Wide Association Study of 125,478 Subjects. Gastroenterology. 2021;160:1164–1178.e6. doi: 10.1053/j.gastro.2020.08.062. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Yuan Y., Bao J., Chen Z., Villanueva A.D., Wen W., Wang F., Zhao D., Fu X., Cai Q., Long J., et al. Multi-omics analysis to identify susceptibility genes for colorectal cancer. Hum. Mol. Genet. 2021;30:321–330. doi: 10.1093/hmg/ddab021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Bien S.A., Su Y.R., Conti D.V., Harrison T.A., Qu C., Guo X., Lu Y., Albanes D., Auer P.L., Banbury B.L., et al. Genetic variant predictors of gene expression provide new insight into risk of colorectal cancer. Hum. Genet. 2019;138:307–326. doi: 10.1007/s00439-019-01989-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Wu L., Shi W., Long J., Guo X., Michailidou K., Beesley J., Bolla M.K., Shu X.O., Lu Y., Cai Q., et al. A transcriptome-wide association study of 229,000 women identifies new candidate susceptibility genes for breast cancer. Nat. Genet. 2018;50:968–978. doi: 10.1038/s41588-018-0132-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Gao G., Fiorica P.N., McClellan J., Barbeira A.N., Li J.L., Olopade O.I., Im H.K., Huo D. A joint transcriptome-wide association study across multiple tissues identifies candidate breast cancer susceptibility genes. Am. J. Hum. Genet. 2023;110:950–962. doi: 10.1016/j.ajhg.2023.04.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Mancuso N., Gayther S., Gusev A., Zheng W., Penney K.L., Kote-Jarai Z., Eeles R., Freedman M., Haiman C., Pasaniuc B., PRACTICAL consortium Large-scale transcriptome-wide association study identifies new prostate cancer risk regions. Nat. Commun. 2018;9:4079. doi: 10.1038/s41467-018-06302-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Liu D., Zhu J., Zhou D., Nikas E.G., Mitanis N.T., Sun Y., Wu C., Mancuso N., Cox N.J., Wang L., et al. A transcriptome-wide association study identifies novel candidate susceptibility genes for prostate cancer risk. Int. J. Cancer. 2022;150:80–90. doi: 10.1002/ijc.33808. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Bosse Y., Li Z., Xia J., Manem V., Carreras-Torres R., Gabriel A., Gaudreault N., Albanes D., Aldrich M.C., Andrew A., et al. Transcriptome-wide association study reveals candidate causal genes for lung cancer. Int. J. Cancer. 2020;146:1862–1878. doi: 10.1002/ijc.32771. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Lu Y., Beeghly-Fadiel A., Wu L., Guo X., Li B., Schildkraut J.M., Im H.K., Chen Y.A., Permuth J.B., Reid B.M., et al. A Transcriptome-Wide Association Study Among 97,898 Women to Identify Candidate Susceptibility Genes for Epithelial Ovarian Cancer Risk. Cancer Res. 2018;78:5419–5430. doi: 10.1158/0008-5472.CAN-18-0951. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Gusev A., Lawrenson K., Lin X., Lyra P.C., Jr., Kar S., Vavra K.C., Segato F., Fonseca M.A.S., Lee J.M., Pejovic T., et al. A transcriptome-wide association study of high-grade serous epithelial ovarian cancer identifies new susceptibility genes and splice variants. Nat. Genet. 2019;51:815–823. doi: 10.1038/s41588-019-0395-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Zhong J., Jermusyk A., Wu L., Hoskins J.W., Collins I., Mocci E., Zhang M., Song L., Chung C.C., Zhang T., et al. A Transcriptome-Wide Association Study Identifies Novel Candidate Susceptibility Genes for Pancreatic Cancer. J. Natl. Cancer Inst. 2020;112:1003–1012. doi: 10.1093/jnci/djz246. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Zheng C.J., Han L.Y., Yap C.W., Ji Z.L., Cao Z.W., Chen Y.Z. Therapeutic targets: Progress of their exploration and investigation of their characteristics (vol 58, pg 259, 2006) Pharmacol. Rev. 2006;58:682. doi: 10.1124/pr.58.2.4. [DOI] [PubMed] [Google Scholar]
- 31.Zhu J., Shu X., Guo X., Liu D., Bao J., Milne R.L., Giles G.G., Wu C., Du M., White E., et al. Associations between Genetically Predicted Blood Protein Biomarkers and Pancreatic Cancer Risk. Cancer Epidemiol. Biomarkers Prev. 2020;29:1501–1508. doi: 10.1158/1055-9965.EPI-20-0091. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Shu X., Bao J., Wu L., Long J., Shu X.O., Guo X., Yang Y., Michailidou K., Bolla M.K., Wang Q., et al. Evaluation of associations between genetically predicted circulating protein biomarkers and breast cancer risk. Int. J. Cancer. 2020;146:2130–2138. doi: 10.1002/ijc.32542. [DOI] [PubMed] [Google Scholar]
- 33.Wu L., Shu X., Bao J., Guo X., Kote-Jarai Z., Haiman C.A., Eeles R.A., Zheng W., PRACTICAL CRUK BPC3 CAPS PEGASUS Consortia Analysis of Over 140,000 European Descendants Identifies Genetically Predicted Blood Protein Biomarkers Associated with Prostate Cancer Risk. Cancer Res. 2019;79:4592–4598. doi: 10.1158/0008-5472.CAN-18-3997. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Gregga I., Pharoah P.D., Gayther S.A., Manichaikul A., Im H.K., Kar S.P., Schildkraut J.M., Wheeler H.E. Predicted proteome association studies of breast, prostate, ovarian, and endometrial cancers implicate plasma protein regulation in cancer susceptibility. Cancer Epidemiol. Biomarkers Prev. 2023;32:1198–1207. doi: 10.1158/1055-9965.EPI-23-0309. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Jia G., Yang Y., Ping J., Xu S., Liu L., Guo X., Tao R., Long J., Zheng W. Identification of target proteins for breast cancer genetic risk loci and blood risk biomarkers in a large study by integrating genomic and proteomic data. Int. J. Cancer. 2023;152:2314–2320. doi: 10.1002/ijc.34472. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Considine D.P.C., Jia G., Shu X., Schildkraut J.M., Pharoah P.D.P., Zheng W., Kar S.P., Ovarian Cancer Association Consortium Genetically predicted circulating protein biomarkers and ovarian cancer risk. Gynecol. Oncol. 2021;160:506–513. doi: 10.1016/j.ygyno.2020.11.016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Sun B.B., Chiou J., Traylor M., Benner C., Hsu Y.H., Richardson T.G., Surendran P., Mahajan A., Robins C., Vasquez-Grinnell S.G., et al. Plasma proteomic associations with genetics and health in the UK Biobank. Nature. 2023;622:329–338. doi: 10.1038/s41586-023-06592-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Tautermann C.S. Current and Future Challenges in Modern Drug Discovery. Methods Mol. Biol. 2020;2114:1–17. doi: 10.1007/978-1-0716-0282-9_1. [DOI] [PubMed] [Google Scholar]
- 39.Pushpakom S., Iorio F., Eyers P.A., Escott K.J., Hopper S., Wells A., Doig A., Guilliams T., Latimer J., McNamee C., et al. Drug repurposing: progress, challenges and recommendations. Nat. Rev. Drug Discov. 2019;18:41–58. doi: 10.1038/nrd.2018.168. [DOI] [PubMed] [Google Scholar]
- 40.Zang C., Zhang H., Xu J., Zhang H., Fouladvand S., Havaldar S., Cheng F., Chen K., Chen Y., Glicksberg B.S., et al. High-throughput target trial emulation for Alzheimer's disease drug repurposing with real-world data. Nat. Commun. 2023;14:8180. doi: 10.1038/s41467-023-43929-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Wu Y., Warner J.L., Wang L., Jiang M., Xu J., Chen Q., Nian H., Dai Q., Du X., Yang P., et al. Discovery of Noncancer Drug Effects on Survival in Electronic Health Records of Patients With Cancer: A New Paradigm for Drug Repurposing. JCO Clin. Cancer Inform. 2019;3:1–9. doi: 10.1200/CCI.19.00001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Xu H., Aldrich M.C., Chen Q., Liu H., Peterson N.B., Dai Q., Levy M., Shah A., Han X., Ruan X., et al. Validating drug repurposing signals using electronic health records: a case study of metformin associated with reduced cancer mortality. J. Am. Med. Inform. Assoc. 2015;22:179–191. doi: 10.1136/amiajnl-2014-002649. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Bejan C.A., Cahill K.N., Staso P.J., Choi L., Peterson J.F., Phillips E.J. DrugWAS: Drug-wide Association Studies for COVID-19 Drug Repurposing. Clin. Pharmacol. Ther. 2021;110:1537–1546. doi: 10.1002/cpt.2376. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Reznikov L.R., Norris M.H., Vashisht R., Bluhm A.P., Li D., Liao Y.S.J., Brown A., Butte A.J., Ostrov D.A. Identification of antiviral antihistamines for COVID-19 repurposing. Biochem. Biophys. Res. Commun. 2021;538:173–179. doi: 10.1016/j.bbrc.2020.11.095. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Liu R., Wei L., Zhang P. A deep learning framework for drug repurposing via emulating clinical trials on real-world patient data. Nat. Mach. Intell. 2021;3:68–75. doi: 10.1038/s42256-020-00276-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Zhang J., Dutta D., Köttgen A., Tin A., Schlosser P., Grams M.E., Harvey B., CKDGen Consortium. Yu B., Boerwinkle E., et al. Plasma proteome analyses in individuals of European and African ancestry identify cis-pQTLs and models for proteome-wide association studies. Nat. Genet. 2022;54:593–602. doi: 10.1038/s41588-022-01051-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Ferkingstad E., Sulem P., Atlason B.A., Sveinbjornsson G., Magnusson M.I., Styrmisdottir E.L., Gunnarsdottir K., Helgason A., Oddsson A., Halldorsson B.V., et al. Large-scale integration of the plasma proteome with genetics and disease. Nat. Genet. 2021;53:1712–1721. doi: 10.1038/s41588-021-00978-w. [DOI] [PubMed] [Google Scholar]
- 48.Fachal L., Aschard H., Beesley J., Barnes D.R., Allen J., Kar S., Pooley K.A., Dennis J., Michailidou K., Turman C., et al. Fine-mapping of 150 breast cancer risk regions identifies 191 likely target genes. Nat. Genet. 2020;52:56–73. doi: 10.1038/s41588-019-0537-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Wang G., Sarkar A., Carbonetto P., Stephens M. A simple new approach to variable selection in regression, with application to genetic fine mapping. J. R. Stat. Soc. Series B Stat. Methodol. 2020;82:1273–1300. doi: 10.1111/rssb.12388. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Weissbrod O., Hormozdiari F., Benner C., Cui R., Ulirsch J., Gazal S., Schoech A.P., van de Geijn B., Reshef Y., Márquez-Luna C., et al. Functionally informed fine-mapping and polygenic localization of complex trait heritability. Nat. Genet. 2020;52:1355–1363. doi: 10.1038/s41588-020-00735-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Chen Z., Guo X., Tao R., Huyghe J.R., Law P.J., Fernandez-Rozadilla C., Ping J., Jia G., Long J., Li C., et al. Fine-mapping analysis including over 254,000 East Asian and European descendants identifies 136 putative colorectal cancer susceptibility genes. Nat. Commun. 2024;15:3557. doi: 10.1038/s41467-024-47399-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Schwarzer G., Carpenter J.R., Rücker G. Springer; 2015. Meta-analysis with R. [Google Scholar]
- 53.Giambartolomei C., Vukcevic D., Schadt E.E., Franke L., Hingorani A.D., Wallace C., Plagnol V. Bayesian test for colocalisation between pairs of genetic association studies using summary statistics. PLoS Genet. 2014;10 doi: 10.1371/journal.pgen.1004383. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Zhu Z., Zhang F., Hu H., Bakshi A., Robinson M.R., Powell J.E., Montgomery G.W., Goddard M.E., Wray N.R., Visscher P.M., Yang J. Integration of summary data from GWAS and eQTL studies predicts complex trait gene targets. Nat. Genet. 2016;48:481–487. doi: 10.1038/ng.3538. [DOI] [PubMed] [Google Scholar]
- 55.Li L., Huang K.L., Gao Y., Cui Y., Wang G., Elrod N.D., Li Y., Chen Y.E., Ji P., Peng F., et al. An atlas of alternative polyadenylation quantitative trait loci contributing to complex trait and disease heritability. Nat. Genet. 2021;53:994–1005. doi: 10.1038/s41588-021-00864-5. [DOI] [PubMed] [Google Scholar]
- 56.Ward L.D., Kellis M. HaploReg: a resource for exploring chromatin states, conservation, and regulatory motif alterations within sets of genetically linked variants. Nucleic Acids Res. 2012;40:D930–D934. doi: 10.1093/nar/gkr917. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.He B., Chen C., Teng L., Tan K. Global view of enhancer-promoter interactome in human cells. Proc. Natl. Acad. Sci. USA. 2014;111:E2191–E2199. doi: 10.1073/pnas.1320308111. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Lizio M., Harshbarger J., Shimoji H., Severin J., Kasukawa T., Sahin S., Abugessaisa I., Fukuda S., Hori F., Ishikawa-Kato S., et al. Gateways to the FANTOM5 promoter level mammalian expression atlas. Genome Biol. 2015;16:22. doi: 10.1186/s13059-014-0560-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Gao T., Qian J. EnhancerAtlas 2.0: an updated resource with enhancer annotation in 586 tissue/cell types across nine species. Nucleic Acids Res. 2020;48:D58–D64. doi: 10.1093/nar/gkz980. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Wang Y., Song C., Zhao J., Zhang Y., Zhao X., Feng C., Zhang G., Zhu J., Wang F., Qian F., et al. SEdb 2.0: a comprehensive super-enhancer database of human and mouse. Nucleic Acids Res. 2023;51:D280–D290. doi: 10.1093/nar/gkac968. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Chandrashekar D.S., Bashel B., Balasubramanya S.A.H., Creighton C.J., Ponce-Rodriguez I., Chakravarthi B.V.S.K., Varambally S. UALCAN: A Portal for Facilitating Tumor Subgroup Gene Expression and Survival Analyses. Neoplasia. 2017;19:649–658. doi: 10.1016/j.neo.2017.05.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Chandrashekar D.S., Karthikeyan S.K., Korla P.K., Patel H., Shovon A.R., Athar M., Netto G.J., Qin Z.S., Kumar S., Manne U., et al. UALCAN: An update to the integrated cancer data analysis platform. Neoplasia. 2022;25:18–27. doi: 10.1016/j.neo.2022.01.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Knox C., Wilson M., Klinger C.M., Franklin M., Oler E., Wilson A., Pon A., Cox J., Chin N.E.L., Strawbridge S.A., et al. DrugBank 6.0: the DrugBank Knowledgebase for 2024. Nucleic Acids Res. 2024;52:D1265–D1275. doi: 10.1093/nar/gkad976. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Zdrazil B., Felix E., Hunter F., Manners E.J., Blackshaw J., Corbett S., de Veij M., Ioannidis H., Lopez D.M., Mosquera J.F., et al. The ChEMBL Database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods. Nucleic Acids Res. 2024;52:D1180–D1192. doi: 10.1093/nar/gkad1004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Zhou Y., Zhang Y., Zhao D., Yu X., Shen X., Zhou Y., Wang S., Qiu Y., Chen Y., Zhu F. TTD: Therapeutic Target Database describing target druggability information. Nucleic Acids Res. 2024;52:D1465–D1477. doi: 10.1093/nar/gkad751. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Buniello A., Suveges D., Cruz-Castillo C., Llinares M.B., Cornu H., Lopez I., Tsukanov K., Roldán-Romero J.M., Mehta C., Fumis L., et al. Open Targets Platform: facilitating therapeutic hypotheses building in drug discovery. Nucleic Acids Res. 2025;53:D1467–D1475. doi: 10.1093/nar/gkae1128. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Roden D.M., Pulley J.M., Basford M.A., Bernard G.R., Clayton E.W., Balser J.R., Masys D.R. Development of a large-scale de-identified DNA biobank to enable personalized medicine. Clin. Pharmacol. Ther. 2008;84:362–369. doi: 10.1038/clpt.2008.89. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Austin P.C., Stuart E.A. Moving towards best practice when using inverse probability of treatment weighting (IPTW) using the propensity score to estimate causal treatment effects in observational studies. Stat. Med. 2015;34:3661–3679. doi: 10.1002/sim.6607. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Shu X., Zhou Q., Sun X., Flesaker M., Guo X., Long J., Robson M.E., Shu X.O., Zheng W., Bernstein J.L. Associations between circulating proteins and risk of breast cancer by intrinsic subtypes: a Mendelian randomisation analysis. Br. J. Cancer. 2022;127:1507–1514. doi: 10.1038/s41416-022-01923-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Sun J., Zhao J., Jiang F., Wang L., Xiao Q., Han F., Chen J., Yuan S., Wei J., Larsson S.C., et al. Identification of novel protein biomarkers and drug targets for colorectal cancer by integrating human plasma proteome with genome. Genome Med. 2023;15 doi: 10.1186/s13073-023-01229-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Shiina T., Hosomichi K., Inoko H., Kulski J.K. The HLA genomic loci map: expression, interaction, diversity and disease. J. Hum. Genet. 2009;54:15–39. doi: 10.1038/jhg.2008.5. [DOI] [PubMed] [Google Scholar]
- 72.Mora-Bitria L., Asquith B. Germline natural killer cell receptors modulating the T cell response. Front. Immunol. 2024;15 doi: 10.3389/fimmu.2024.1477991. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Dendrou C.A., Petersen J., Rossjohn J., Fugger L. HLA variation and disease. Nat. Rev. Immunol. 2018;18:325–339. doi: 10.1038/nri.2017.143. [DOI] [PubMed] [Google Scholar]
- 74.Jia X., Zhao C., Zhao W. Emerging Roles of MHC Class I Region-Encoded E3 Ubiquitin Ligases in Innate Immunity. Front. Immunol. 2021;12 doi: 10.3389/fimmu.2021.687102. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Yang Q., Tang Y., Tang C., Cong H., Wang X., Shen X., Ju S. Diminished LINC00173 expression induced miR-182-5p accumulation promotes cell proliferation, migration and apoptosis inhibition via AGER/NF-kappaB pathway in non-small-cell lung cancer. Am. J. Transl. Res. 2019;11:4248–4262. [PMC free article] [PubMed] [Google Scholar]
- 76.Bailey M.H., Tokheim C., Porta-Pardo E., Sengupta S., Bertrand D., Weerasinghe A., Colaprico A., Wendl M.C., Kim J., Reardon B., et al. Comprehensive Characterization of Cancer Driver Genes and Mutations. Cell. 2018;174:1034–1035. doi: 10.1016/j.cell.2018.07.034. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Dietlein F., Weghorn D., Taylor-Weiner A., Richters A., Reardon B., Liu D., Lander E.S., Van Allen E.M., Sunyaev S.R. Identification of cancer driver genes based on nucleotide context. Nat. Genet. 2020;52:208–218. doi: 10.1038/s41588-019-0572-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Sondka Z., Bamford S., Cole C.G., Ward S.A., Dunham I., Forbes S.A. The COSMIC Cancer Gene Census: describing genetic dysfunction across all human cancers. Nat. Rev. Cancer. 2018;18:696–705. doi: 10.1038/s41568-018-0060-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Chen E.Y., Tan C.M., Kou Y., Duan Q., Wang Z., Meirelles G.V., Clark N.R., Ma'ayan A. Enrichr: interactive and collaborative HTML5 gene list enrichment analysis tool. BMC Bioinf. 2013;14:128. doi: 10.1186/1471-2105-14-128. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80.Taipale H., Solmi M., Lähteenvuo M., Tanskanen A., Correll C.U., Tiihonen J. Antipsychotic use and risk of breast cancer in women with schizophrenia: a nationwide nested case-control study in Finland. Lancet Psychiatry. 2021;8:883–891. doi: 10.1016/S2215-0366(21)00241-8. [DOI] [PubMed] [Google Scholar]
- 81.Durdagi S., Vullo D., Pan P., Kähkönen N., Määttä J.A., Hytönen V.P., Scozzafava A., Parkkila S., Supuran C.T. Protein-protein interactions: inhibition of mammalian carbonic anhydrases I-XV by the murine inhibitor of carbonic anhydrase and other members of the transferrin family. J. Med. Chem. 2012;55:5529–5535. doi: 10.1021/jm3004587. [DOI] [PubMed] [Google Scholar]
- 82.Eckenroth B.E., Mason A.B., McDevitt M.E., Lambert L.A., Everse S.J. The structure and evolution of the murine inhibitor of carbonic anhydrase: a member of the transferrin superfamily. Protein Sci. 2010;19:1616–1626. doi: 10.1002/pro.439. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Karakus F., Eyol E., Yilmaz K., Unuvar S. Inhibition of cell proliferation, migration and colony formation of LS174T Cells by carbonic anhydrase inhibitor. Afr. Health Sci. 2018;18:1303–1310. doi: 10.4314/ahs.v18i4.51. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Noma N., Fujii G., Miyamoto S., Komiya M., Nakanishi R., Shimura M., Tanuma S.I., Mutoh M. Impact of Acetazolamide, a Carbonic Anhydrase Inhibitor, on the Development of Intestinal Polyps in Min Mice. Int. J. Mol. Sci. 2017;18 doi: 10.3390/ijms18040851. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85.Shen Y., Li X., Dong D., Zhang B., Xue Y., Shang P. Transferrin receptor 1 in cancer: a new sight for cancer therapy. Am. J. Cancer Res. 2018;8:916–931. [PMC free article] [PubMed] [Google Scholar]
- 86.Chen Y.J., Qian Z.M., Sheng Y., Zheng J., Liu Y. Angiotensin II down-regulates transferrin receptor 1 and ferroportin 1 expression in Neuro-2a cells via activation of type-1 receptor. Neurosci. Lett. 2020;716 doi: 10.1016/j.neulet.2019.134684. [DOI] [PubMed] [Google Scholar]
- 87.Yarmolinsky J., Díez-Obrero V., Richardson T.G., Pigeyre M., Sjaarda J., Paré G., Walker V.M., Vincent E.E., Tan V.Y., Obón-Santacana M., et al. Genetically proxied therapeutic inhibition of antihypertensive drug targets and risk of common cancers: A mendelian randomization analysis. PLoS Med. 2022;19 doi: 10.1371/journal.pmed.1003897. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Schmit S.L., Rennert H.S., Rennert G., Gruber S.B. Coffee consumption and the risk of colorectal cancer. Cancer Epidemiol. Biomarkers Prev. 2016;25:634–639. doi: 10.1158/1055-9965.EPI-15-0924. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89.Sartini M., Bragazzi N.L., Spagnolo A.M., Schinca E., Ottria G., Dupont C., Cristina M.L. Coffee Consumption and Risk of Colorectal Cancer: A Systematic Review and Meta-Analysis of Prospective Studies. Nutrients. 2019;11 doi: 10.3390/nu11030694. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.Low J.J.L., Tan B.J.W., Yi L.X., Zhou Z.D., Tan E.K. Genetic susceptibility to caffeine intake and metabolism: a systematic review. J. Transl. Med. 2024;22:961. doi: 10.1186/s12967-024-05737-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91.Chen C.H., Kraemer B.R., Mochly-Rosen D. ALDH2 variance in disease and populations. Dis. Model. Mech. 2022;15 doi: 10.1242/dmm.049601. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92.Laskar A.A., Alam M.F., Ahmad M., Younus H. Kinetic and biophysical investigation of the inhibitory effect of caffeine on human salivary aldehyde dehydrogenase: Implications in oral health and chemotherapy. J. Mol. Struct. 2018;1157:61–68. [Google Scholar]
- 93.Wei P.L., Prince G.M.S.H., Batzorig U., Huang C.Y., Chang Y.J. ALDH2 promotes cancer stemness and metastasis in colorectal cancer through activating beta-catenin signaling. J. Cell. Biochem. 2023;124:907–920. doi: 10.1002/jcb.30418. [DOI] [PubMed] [Google Scholar]
- 94.Li R., Zhao Z., Sun M., Luo J., Xiao Y. ALDH2 gene polymorphism in different types of cancers and its clinical significance. Life Sci. 2016;147:59–66. doi: 10.1016/j.lfs.2016.01.028. [DOI] [PubMed] [Google Scholar]
- 95.Keum N., Cao Y., Lee D.H., Park S.M., Rosner B., Fuchs C.S., Wu K., Giovannucci E.L. Male pattern baldness and risk of colorectal neoplasia. Br. J. Cancer. 2016;114:110–117. doi: 10.1038/bjc.2015.438. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 96.Thompson I.M., Jr., Goodman P.J., Tangen C.M., Parnes H.L., Minasian L.M., Godley P.A., Lucia M.S., Ford L.G. Long-term survival of participants in the prostate cancer prevention trial. N. Engl. J. Med. 2013;369:603–610. doi: 10.1056/NEJMoa1215932. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 97.Wang L., Lei Y., Gao Y., Cui D., Tang Q., Li R., Wang D., Chen Y., Zhang B., Wang H. Association of finasteride with prostate cancer: A systematic review and meta-analysis. Medicine (Baltim.) 2020;99 doi: 10.1097/MD.0000000000019486. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 98.Thompson I.M., Lucia M.S., Redman M.W., Darke A., La Rosa F.G., Parnes H.L., Lippman S.M., Coltman C.A. Finasteride decreases the risk of prostatic intraepithelial neoplasia. J. Urol. 2007;178:107–110. doi: 10.1016/j.juro.2007.03.012. ; discussion 110. [DOI] [PubMed] [Google Scholar]
- 99.Rahman T., Sahrmann J.M., Olsen M.A., Nickel K.B., Miller J.P., Ma C., Grucza R.A. Risk of Breast Cancer With Prolactin Elevating Antipsychotic Drugs: An Observational Study of US Women (Ages 18-64 Years) J. Clin. Psychopharmacol. 2022;42:7–16. doi: 10.1097/JCP.0000000000001513. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 100.Hicks B.M., Busby J., Mills K., O'Neil F.A., McIntosh S.A., Zhang S.D., Liberante F.G., Cardwell C.R. Post-diagnostic antipsychotic use and cancer mortality: a population based cohort study. BMC Cancer. 2020;20:804. doi: 10.1186/s12885-020-07320-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 101.Zhou M., Zhang Y., Shi L., Li L., Zhang D., Gong Z., Wu Q. Activation and modulation of the AGEs-RAGE axis: Implications for inflammatory pathologies and therapeutic interventions - A review. Pharmacol. Res. 2024;206 doi: 10.1016/j.phrs.2024.107282. [DOI] [PubMed] [Google Scholar]
- 102.Chamie K., Ghosh P.M., Koppie T.M., Romero V., Troppmann C., deVere White R.W. The effect of sirolimus on prostate-specific antigen (PSA) levels in male renal transplant recipients without prostate cancer. Am. J. Transplant. 2008;8:2668–2673. doi: 10.1111/j.1600-6143.2008.02430.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 103.Yanik E.L., Siddiqui K., Engels E.A. Sirolimus effects on cancer incidence after kidney transplantation: a meta-analysis. Cancer Med. 2015;4:1448–1459. doi: 10.1002/cam4.487. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 104.Schumacher F.R., Al Olama A.A., Berndt S.I., Benlloch S., Ahmed M., Saunders E.J., Dadaev T., Leongamornlert D., Anokian E., Cieza-Borrella C., et al. Association analyses of more than 140,000 men identify 63 new prostate cancer susceptibility loci. Nat. Genet. 2018;50:928–936. doi: 10.1038/s41588-018-0142-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 105.Huyghe J.R., Bien S.A., Harrison T.A., Kang H.M., Chen S., Schmit S.L., Conti D.V., Qu C., Jeon J., Edlund C.K., et al. Discovery of common and rare genetic risk variants for colorectal cancer. Nat. Genet. 2019;51:76–87. doi: 10.1038/s41588-018-0286-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Table S1 provides the download information for the summary statistics of GWAS data for the six common cancers, including breast,6 ovarian,11 prostate,104 colorectal,105 lung,13 and pancreatic.29 Metadata and pQTL summary statistics from UKB-PPP can be downloaded from Synapse with project SynID, syn51364943: https://www.synapse.org/Synapse:syn51364943, and pQTLs from ARIC46 and deCODE genetics47 can be accessed through previous publications. Functional genomic data include TCGA and CPTAC differential expression results accessible through https://ualcan.path.uab.edu/index.html; 4DGenome: https://4dgenome.research.chop.edu/; Depmap: https://depmap.org/portal/; FANTOM5: http://fantom.gsc.riken.jp/5/. HaploReg v.3: https://pubs.broadinstitute.org/mammals/haploreg/; and GTEx v.8: https://gtexportal.org/home/downloads/adult-gtex/qtl. GENCODE (v.26.GRCh38): https://www.gencodegenes.org/human/release_26.html. The National Cancer Institute can be accessed through https://www.cancer.gov/about-cancer/treatment/drugs; CGC can be accessed via the COSMIC website: https://cancer.sanger.ac.uk/census. Drugs and compound data can be downloaded from four drug database websites: ChEMBL: https://www.ebi.ac.uk/chembl/, TTD: https://db.idrblab.net/ttd/, Open Targets: https://www.opentargets.org/, and DrugBank: https://go.drugbank.com/. The EHR data, containing de-identified clinical information, can be accessed through the VUMC SD database: https://victr.vumc.org/data-resource-sd/.
The developed pipeline and main source R codes that are used in this work are available from the GitHub website of X.G.’s lab (https://github.com/XingyiGuo/PQTL_EHR/).






