Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2019 Aug 1.
Published in final edited form as: Int J Cancer. 2018 Mar 14;143(3):527–534. doi: 10.1002/ijc.31341

Prospective study of blood metabolites associated with colorectal cancer risk

Xiang Shu 1, Yong-Bing Xiang 2, Nathaniel Rothman 3, Danxia Yu 1, Hong-Lan Li 2, Gong Yang 1, Hui Cai 1, Xiao Ma 2, Qing Lan 3, Yu-Tang Gao 2, Wei Jia 4,5, Xiao-Ou Shu 1, Wei Zheng 1
PMCID: PMC6019169  NIHMSID: NIHMS950098  PMID: 29479691

Abstract

Few prospective studies, and none in Asians, have systematically evaluated the relationship between blood metabolites and colorectal cancer risk. We conducted a nested case-control study to search for risk-associated metabolite biomarkers for colorectal cancer in an Asian population using blood samples collected prior to cancer diagnosis. Conditional logistic regression was performed to assess associations of metabolites with cancer risk. In the current study, we included 250 incident cases with colorectal cancer and individually matched controls nested within two prospective Shanghai cohorts. We found 35 metabolites associated with risk of colorectal cancer after adjusting for multiple comparisons. Among them, 12 metabolites were glycerophospholipids including nine associated with reduced risk of colorectal cancer and three with increased risk [odds ratios (ORs) per standard deviation (SD) increase of transformed metabolites: 0.31 to 1.98; p values: 0.002 to 1.25×10−10]. The other 23 metabolites associated with colorectal cancer risk included nine lipids other than glycerophospholipid, seven aromatic compounds, five organic acids, and four other organic compounds. After mutual adjustment, nine metabolites remained statistically significant for colorectal cancer. Together, these independently associated metabolites can separate cancer cases from controls with an area under the curve of 0.76 for colorectal cancer. We have identified that dysregulation of glycerophospholipids may contribute to risk of colorectal cancer.

Keywords: Metabolomics, biomarkers, colorectal cancer, nested case-control study

Introduction

Colorectal cancer is the most commonly diagnosed malignancy of the digestive system with approximately 1.3 million new patients diagnosed annually worldwide.1 Multiple risk factors of colorectal cancer, including obesity, type 2 diabetes, cigarette smoking, and heavy alcohol consumption,24 are related to metabolic perturbation.5 It has been recognized that alteration of energy metabolism and metabolic reprogramming are critical for tumor initiation and progression.6,7 A systematic evaluation of metabolomic profiles prior to cancer diagnosis can help facilitate the discovery of biomarkers for early detection of cancer and improve the understanding of cancer etiology and biology.

Most metabolomics studies conducted to date for colorectal cancer are retrospective,8,9 and thus results from these studies could be affected by prevalent biases due to the presence of cancer or changes of lifestyles after cancer diagnosis. A prospective study conducted in Caucasians investigated blood metabolites in relation to risk of colorectal cancer and identified a suggestive inverse association of leucyl-leucine.10 However, no statistically significant association was found after adjusting for multiple comparisons.10 Herein, we report results from a metabolomics study conducted in Chinese to systematically search novel blood biomarkers for colorectal cancer risk using samples collected prior to cancer diagnosis.

Material and Methods

Study population

This case-control study was nested within two population-based prospective cohorts, the Shanghai Women’s Health Study (SWHS) and the Shanghai Men’s Health Study (SMHS). Detailed methodology for these two cohort studies was described in previous publications.11,12 Briefly, 74,941 women aged 40–70 years were recruited for the SWHS from 1996 to 2000, and 61,480 men aged 40–74 years were recruited for the SMHS from 2002 to 2006 from urban Shanghai, China. In-person interviews were conducted at baseline to collect information on socio-demographics, disease history and medication use, family history of cancer, dietary intakes, cigarette smoking, alcohol consumption and other lifestyle factors. Participants’ height, weight, and waist and hip circumferences were measured following a standard protocol. At the time of the baseline survey, a 10 ml blood sample was collected and placed in an EDTA tube. Blood samples were kept at 4°C during transportation and processed within 6 hours. Plasma was aliquoted to 2 ml vials for long-term storage at −75°C until needed for analysis. All participants are followed by a combination of annual linkage with the population-based Shanghai Tumor Registry and Vital Statistics Registry and in-person surveys taking place every 2–4 years. Incident cancer cases are verified via reviews of medical charts. Institutional review boards of all involved institutes approved the studies, and written informed consent was obtained from all participants prior to interview.

For this study we selected cases and controls from cohort members who provided blood samples at baseline, accounting for approximately 75% of each cohort. In this study, we included 250 incident colorectal cancers (125 women; 125 men) randomly selected from the 1,855 incident cases diagnosed before 2013. A control was randomly selected from the risk set of cancer-free cohort members for each index case using the incidence density sampling method and matched to the case on age (± 2 years), sex, date of sample collection (± 30 days), time of sample collection (morning or afternoon), time interval after last meal (± 2 hours), recent antibiotic use (yes or no), and menopausal status (for women).

Global metabolic profiling of plasma samples by GC-TOFMS

Metabolites in the plasma were derivatized and subsequently analyzed by gas chromatography time-of-flight mass spectrometry (GC-TOFMS) based on the previously published method with minor modifications.13 Briefly, an aliquot of 100 µL plasma sample was spiked with 10 µL of L-2-chlorophenylalanine in water (0.3 mg/mL) and 10 µL of heptadecanoic acid in methanol (1 mg/mL). The mixture was extracted with 300 µL of methanol/chloroform (3:1). After vortexed for 30 s and stored at −20 °C for 10 min, samples were centrifuged at 16,100 g for 10 min. An aliquot of the 300 µL supernatant was transferred to a glass sampling vial to vacuum-dry at room temperature. Then, 80-µL of methoxyamine (15 mg/mL in pyridine) was added to the vial and kept at 30 °C for 90 min, followed by 80 µL BSTFA (1%TMCS) at 70 °C for 120 min for derivatization to take place. The derivatized samples were then analyzed by an Agilent 6890N GC coupled with a Pegasus HT TOFMS (Leco Corporation, St Joseph, USA) in the splitless mode. A DB-5 ms capillary column (30 m × 250 µm I.D., 0.25-µm film thickness; Agilent J&W Scientific, Folsom, CA, USA) was used for separation. The measurements were made with electron impact ionization (70 eV) in the full scan mode (m/z 30–600).

Pooled quality control (QC) samples were prepared by mixing aliquots of the study samples such that the pooled samples broadly represent the biological average of the whole sample set. The QC samples for this project were prepared with the test samples and injected after every 10 test samples throughout the analytical run. Solvent blank samples were also added for detection of potential contamination. All samples were prepared in a random order and the lab team was blinded to the case-control status.

Global metabolic profiling of plasma samples by UPLC-QTOFMS13

Briefly, an aliquot of 100 µL of plasma sample was extracted with a 400 µL mixture of methanol and acetonitrile (5:3, v/v, containing 2-chlorophenylalanine as the internal standard). After being vortexed and centrifuged for 20 minutes at 16,100 g, the supernatant was used for ultra-performance liquid chromatography quadrupole-time-of-flight mass spectrometry (UPLC-QTOFMS) analysis. A Waters ACQUITY UPLC system equipped with a binary solvent delivery manager and a sample manager (Waters Corporation, Milford, MA) is used with chromatographic separations performed on a 2.1 × 100 mm 1.7 µm ACQUITY BEH C18 chromatography column. Mass spectrometry is performed using a Micromass Q-TOF Premier mass spectrometer equipped with an electrospray interface (Waters Corporation, Milford, MA).

Metabolomics data analysis

Metabolomics data analysis followed the previous publications.13 The raw data from UPLC-QTOFMS were analyzed with the MarkerLynx Applications Manager version 4.1 (Waters, Manchester, U.K.). The raw data from GC-TOFMS was processed with ChromaTOF software (v4.22, Leco Co., CA, USA). A total of 131 features were detected by GC-TOFMS and 900 features by UPLC-QTOFMS (205 from positive mode and 695 from negative mode) (Figure 1). Compound annotation for GC-TOFMS was performed by comparing the mass fragments with NIST 08 Standard mass spectral database with a similarity of more than 70 % and then verified by available reference standards in the in-house standard library (containing 1200 endogenous metabolites). For UPLC-QTOFMS data, compound annotation was carried out with the aid of available reference standards or the web-based resources such as the Human Metabolome Database (HMDB, www.hmdb.ca). A total of 618 metabolites were identified, of which 167 were annotated by in-house library.

Figure 1.

Figure 1

Flowchart of data acquisition and processing steps employed for current study

To eliminate analytical variation and minor sensitivity shifts throughout the study, the three data sets (UPLC-QTOFMS POS, NEG, and GC-TOFMS) were normalized by their pooled QCs respectively (Equation 1).

Peak intensity=raw intensity×(i=1n peak intensity of QCi)/npeak intensity of ((QCi+QC(i+1))/2) Equation 1

The normalized data from two platforms were combined and repetitive and ectogenic compounds were removed. Missing values were replaced by minimum of valid data from corresponding compounds.

Statistical analysis

Natural logarithm transformation and standardization according to the mean and standard deviation of all samples were used to normalize the distribution of metabolite data. Levels of a specific metabolite below its detection limit were assigned the lowest detectable value in data analysis. Univariate analyses were performed for each covariate using paired t-test (continuous), McNemar’s test (binary), or conditional logistic regression model (categorical). For the multiple regression analyses, variables adjusted for included age at interview (years), ever smoking (Yes/No), pack-years smoking, BMI (kg/m2), leisure-time physical activity (Yes/No), alcohol drinking (Yes/No), time interval between last meal and sample collection (hours), intakes of red meat, vegetables and fruits, fiber, and fat (all measured by g/d), and total calorie intake (kcal/d), at baseline. We used the Benjamini-Hochberg false discovery rate (BH-FDR) <0.05 (based on 618 tests) as the significance level to reduce type I errors due to multiple comparisons. We started the association analyses by estimating ORs associated with per 1 SD increase of log-transformed and standardized metabolites. To assess potential influence of outliers on study results, we also performed categorical data analyses based on tertile distribution of controls for all metabolites that showed a statistically significant association in the analysis using continuous variables.

To understand the intra-relation of metabolites statistically significantly associated with risk of colorectal cancer, we calculated pairwise Spearman’s correlation coefficients. We also performed pairwise partial correlation analyses by including all statistically significant metabolites to control effects of other metabolites. To identify metabolites that are independently associated with colorectal cancer risk, we include all metabolites from the same class (lipids, aromatic compounds, organic acids and derivatives, and other compounds) that are statistically significantly associated with the risk in a single regression model. Metabolites remained statistically significant in each of these classes were then included in a single model to search for metabolites independently associated with cancer risk. We further estimated the area under the curve (AUC) for separating cancer cases and controls based on levels of these metabolites. In a sensitivity analysis, we excluded subjects who reported a diagnosis of type 2 diabetes and who had cancer diagnosed within first two years. Analyses stratified by time interval between blood collection and cancer diagnosis (<4 years, ≥4 years) and by cancer site (colon or rectum) were also performed. All statistical analyses were conducted using Stata 14 (StataCorp, College Station, TX) and R (version 3.31).

Results

After excluding samples with poor quality, 245 matched pairs were remained for further analysis. Cases and controls were, in general, well matched on all intended matching covariates. Male cancer cases were 0.3 year younger than their matched controls, and the difference was statistically significant (p<0.05). Cancer cases had a slightly shorter interval between last meal and sample collection than their controls (cases: 5.0 hours, controls: 6.6 hours, p<0.001). No statistically significant case-control differences were observed for other characteristics of participants such as body mass index, waist-hip ratio, leisure- time physical activity, alcohol drinking, and dietary intakes (Table 1). In total we found 35metabolites that showed a statistically significant association with colorectal cancer at the FDR-p of <0.05 (Table 2). Similar associations were obtained from univariable models (data not shown). The statistically significantly associated metabolites were mainly glycerophospholipids and other lipids (five associated with increased risk and 14 with decreased risk). The most significant glycerophospholipid and lipid other than glycerophospholipid are PE(20:0/18:2) (OR=0.40, p=1.25×10−10) and MG(0:0/18:3/0:0) (OR=4.66, p=5.75×10−11), respectively. Majority of identified glycerophospholipids were inversely associated with colorectal cancer risk (nine with decreased risk and three with increased risk). In addition, seven aromatic compounds, five organic acids, and four other organic compounds were also associated with colorectal cancer risk (10 with increased risk and six with decreased risk). Excluding subjects with prevalent type 2 diabetes did not change the results appreciably (Supplementary Table 1). Categorical analyses by tertile distributions of these metabolites showed a similar association as those derived from the continuous variable analyses, indicating that the results of these analyses were robust (Supplementary Table 2).

Table 1.

Participant characteristics at baseline by sex and case-control status in the Shanghai Women Health’s Study and Shanghai Men Health’s Study

Variables Female (N=122 pairs) Male (N=123 pairs)


Case Control p Case Control p
Age (years)a 56.9±8.4 57.0±8.4 0.664 56.2±6.8 56.5±6.6 0.025
Body mass index (kg/m2)b
  <25 82 (67.2) 67 (54.9) 79 (64.2) 83 (67.5)
  25–29.9 35 (28.7) 45 (36.9) 39 (31.7) 39 (31.7)
  ≥30 5 (4.1) 10 (8.2) 0.1 5 (4.1) 1 (0.8) 0.338
Waist-hip-ratioa 0.82±0.05 0.82±0.06 0.323 0.91±0.06 0.89±0.06 0.075
Regular leisure-time physical activity engagementb 56 (45.9) 58 (47.5) 0.773 45 (36.6) 47 (38.2) 0.793
Smoking statusb
  Never 121 (99.2) 117 (95.9) 45 (36.6) 38 (30.9)
  Ever 1 (0.82) 5 (4.1) 0.219 78 (63.4) 85 (69.1) 0.317
Pack-years of smokingb
  Non-smoker 121 (99.2) 117 (95.9) 45 (36.6) 38 (30.9)
  <30 pack-years 0 (0) 4 (3.3) 47 (38.2) 61 (49.6)
  ≥30 pack-years 1 (0.8) 1 (0.8) 1 31 (25.2) 24 (19.5) 0.2
Ever regularly drank alcohol beverageb 1 (0.8) 4 (3.3) 0.375 40 (32.5) 43 (35.0) 0.691
Prior type 2 diabetes diagnosisb, c 9 (7.44) 11 (9.2) 0.607 12 (9.9) 5 (4.1) 0.057
Red meat intaked 48.3±45.8 50.1±37.3 0.934 65.0±38.5 60.4±52.4 0.529
Vegetable and fruit intaked 562.1±289.8 541.2±251.2 0.429 562.1± 289.8 494.6±225.9 0.611
Fiber intaked 11.0±3.8 11.1±3.5 0.95 11.8±4.1 11.9±5.0 0.565
Fat intaked 28.5±13.0 28.8±13.0 0.719 36.7±13.1 34.1±16.9 0.223
Time intervala, e 4.1±3.5 4.1±3.5 0.593 5.0±4.3 6.6±5.0 <0.001
a

Presented in Mean ± SD, p value obtained from paired t-test.

b

Presented in N (%), p value obtained from McNemar's test (binary) or conditional logistic regression (categorical).

c

Type 2 diabetes prior to cancer diagnosis.

d

The unit of dietary intake is grams/day and presented in Mean ± SD. Further adjusted for age at interview (years) and total calorie intake (kcal) in conditional logistic regression.

e

Time interval between last meal and sample collection (hours).

Table 2.

Associations of metabolites with colorectal cancer risk in multivariable conditional logistic regression with BH-FDR<0.05

Metabolite, per standard deviation increase ORa 95% CI p Classb
  MG(0:0/18:3/0:0) 4.66 2.94–7.38 5.75×10−11 Lipids
  PE(20:0/18:2) 0.40 0.30–0.53 1.25×10−10 Lipids
  Tetracosanoic acid 0.44 0.33–0.59 2.36×10−8 Lipids
  PE(22:6/p-18:1) 0.37 0.25–0.53 8.68×10−8 Lipids
  PE(p-16:0/20:4) 0.31 0.20–0.48 2.20×10−7 Lipids
  PC(22:6/18:0) 0.55 0.43–0.70 8.30×10−7 Lipids
  5,6:8,9-diepoxyergost-22-ene-3,7beta-diol 0.49 0.37–0.65 1.30×10−6 Lipids
  PE(22:6/16:0) 0.41 0.28–0.59 2.81×10−6 Lipids
  PE(p-18:0/18:2) 0.47 0.34–0.64 3.21×10−6 Lipids
  PS(18:0/18:0) 0.53 0.39–0.73 7.56×10−5 Lipids
  PE(o-18:1/20:4) 1.98 1.41–2.78 7.81×10−5 Lipids
  2,3-epoxymenaquinone 1.71 1.30–2.25 1.18×10−4 Lipids
  Butenylcarnitine 0.54 0.39–0.74 1.56×10−4 Lipids
  PE(18:3/18:0) 0.61 0.47–0.80 2.53×10−4 Lipids
  PC(18:3/16:0) 1.50 1.20–1.88 3.73×10−4 Lipids
  PE(p-16:0/18:1) 0.68 0.54–0.86 0.002 Lipids
  PC(16:0/16:0) 1.38 1.12–1.69 0.002 Lipids
  Ethyl 4-(methylthio)butyrate 0.62 0.45–0.85 0.003 Lipids
  13'-carboxy-alpha-tocopherol 0.70 0.55–0.88 0.003 Lipids
  N-undecylbenzenesulfonic acid 5.25 3.07–8.98 1.36×10−9 Aromatic cpd
  2-dodecylbenzenesulfonic acid 5.39 3.07–9.46 4.36×10−9 Aromatic cpd
  4-hydroxy-5-(dihydroxyphenyl)-valeric acid-o-methyl-o-sulphate 0.36 0.23–0.57 1.08×10−5 Aromatic cpd
  Coumarin 1.74 1.35–2.25 2.28×10−5 Aromatic cpd
  3-aminobenzoic acid 1.74 1.31–2.30 1.09×10−4 Aromatic cpd
  2-methyl-4-phenyl-2-butyl 2-methylpropanoate 0.56 0.40–0.78 7.21×10−4 Aromatic cpd
  Benzoic acid 1.55 1.19–2.00 9.64×10−4 Aromatic cpd
  Selenocystine 1.99 1.48–2.68 6.41×10−6 Organic acids
  Isoputreanine 0.31 0.18–0.53 1.85×10−5 Organic acids
  Glutamyl-glutamate 0.47 0.31–0.69 1.37×10−4 Organic acids
  Homocarnosine 1.53 1.19–1.98 0.001 Organic acids
  2-keto-glutaramic acid 1.71 1.21–2.43 0.002 Organic acids
  Picolinic acid 4.03 2.40–6.76 1.22×10−7 Others
  4-amino-1-piperidinecarboxylic acid 0.50 0.36–0.71 1.13×10−4 Others
  2-acetyl-5-methylpyridine 0.67 0.53–0.85 7.27×10−4 Others
  1-deoxy-d-xylulose 5-phosphate 1.47 1.16–1.85 0.001 Others
a

ORs, 95% CI and p values were derived from multivariable conditional logistic regression. Model included age at interview (years), smoking status (never/ever), pack-years smoking, BMI (kg/m2), regular leisure-time physical activity engagement (Yes/No), ever regularly drank alcohol (Yes/No), time interval between last meal and sample collection (hours), dietary red meat intake (g/d), vegetable and fruit intake (g/d), fiber intake (g/d), fat intake (g/d), and total calorie intake (kcal) at baseline as covariates. The ORs were estimated with respect to 1 SD increase of natural logarithm transformed standardized metabolite levels.

b

Lipids: lipids and lipid-like molecules; aromatic cpd: aromatic compounds; organic acids: organic acids and derivatives; others: other organic compounds.

Many metabolites are correlated (Supplementary Figure 1–A). For example, a close correlation (r >0.7) was found between N-undecylbenzenesulfonic acid and 2-dodecylbenzenesulfonic acid, isoputreanine and 4-amino-1-piperidinecarboxylic acid, and glutamyl-glutamate and isoputreanine. Partial correlation analysis showed a diminished correlation for most pairs (Supplementary Figure 1–B). After mutual adjustment, nine metabolites were found to be independently associated with colorectal cancer (Table 3). Risk prediction models built using these metabolites showed an AUC of 0.76.

Table 3.

Metabolites independently associated with colorectal cancer risk after mutual adjustment

Metabolite, per standard deviation increase ORa 95% CIa pa
  Picolinic acid 5.11 2.33–11.20 4.74×10−5
  PE(20:0/18:2) 0.45 0.29–0.70 3.24×10−4
  5,6:8,9-diepoxyergost-22-ene-3,7beta-diol 0.47 0.30–0.73 7.37×10−4
  2-methyl-4-phenyl-2-butyl 2-methylpropanoate 0.38 0.22–0.68 8.89×10−4
  Selenocystine 2.10 1.28–3.44 0.003
  2,3-epoxymenaquinone 1.87 1.22–2.84 0.004
  PC(22:6/18:0) 0.57 0.39–0.84 0.005
  Ethyl 4-(methylthio)butyrate 0.53 0.33–0.87 0.013
  PE(p-16:0/20:4) 0.47 0.26–0.85 0.013
a

ORs, 95% CI and p values were derived from multivariable conditional logistic regression including all metabolites listed in the table for each cancer. Same covariates were adjusted for as listed in Table 2.

AUC =0.76 area under the curve was derived from unconditional logistic regression including all metabolites listed in the table for each cancer.

We conducted further stratified analyses for these independently associated metabolites by time interval between blood collection and cancer diagnosis (Table 4). Subjects were divided into two groups, <4 years and ≥4 years, with roughly similar sample size in each group. None of the metabolites associated with colorectal cancer risk exhibited a notable heterogeneity by baseline-diagnosis time interval (Phet>0.05). Excluding subjects who had cancer diagnosed within first two years did not materially change the associations for the identified metabolites (Supplementary Table 3). The stratified analysis within colorectal cancer by cancer site (colon/rectum, Supplementary Table 4) suggested stronger associations with rectal cancer risk for PE(p-18:0/18:2) and 4-amino-1-piperidinecarboxylic acid. None of the heterogeneity tests was statistically significant when multiple comparisons were considered.

Table 4.

Associations of selected metabolites with colorectal cancer risk by time interval of cancer diagnosis

Metabolites, per standard deviation increase < 4 yearsb ≥ 4 yearsb Phetc


ORa 95% CI p ORa 95% CI p
  Picolinic acid 3.58 1.94–6.61 4.33×10−5 5.79 1.56–21.49 0.009 0.517
  PE(20:0/18:2) 0.41 0.29–0.59 6.48×10−7 0.19 0.08–0.47 3.17×10−4 0.113
  5,6:8,9-diepoxyergost-22-ene-3,7beta-diol 0.51 0.35–0.74 4.22×10−4 0.38 0.19–0.77 0.007 0.481
  2-methyl-4-phenyl-2-butyl 2-methylpropanoate 0.51 0.32–0.81 0.005 0.69 0.36–1.33 0.273 0.457
  Selenocystine 1.76 1.20–2.57 0.004 2.66 1.32–5.35 0.006 0.31
  2,3-epoxymenaquinone 1.88 1.31–2.70 6.04×10−4 1.26 0.73–2.17 0.398 0.232
  PC(22:6/18:0) 0.43 0.29–0.62 9.29×10−6 0.56 0.33–0.94 0.029 0.415
  Ethyl 4-(methylthio)butyrate 0.48 0.29–0.79 0.004 0.90 0.52–1.58 0.717 0.099
  PE(p-16:0/20:4) 0.26 0.14–0.47 1.05×10−5 0.28 0.11–0.68 0.005 0.909
a

ORs, 95% CI and p values were derived from multivariable conditional logistic regression. Same covariates were adjusted for as listed in Table 2.

b

158 and 87 case-control pairs were included in stratified analysis when the time interval between baseline and cancer diagnosis was < 4 years and ≥ 4 years, respectively

c

Phet was derived from Cochran's Q test of heterogeneity.

DISCUSSION

To our knowledge, this is the first prospective metabolomics study conducted in an Asian population to investigate metabolomic biomarkers for colorectal cancer risk. A previous prospective metabolomics study conducted in Caucasians reported an inverse association between leucyl-leucine and colorectal cancer risk for both sexes, although the association was not statistically significant after adjusting for multiple comparisons.10 Leucyl-leucine is a dipeptide, which is resulted from incomplete breakdown of proteins. However, prior knowledge regarding the link between this metabolite and cancer development has been limited. We cannot evaluate this association as the metabolite was not measured in our study.

Multiple glycerophospholipids, specifically phosphatidylcholines and phosphatidylethanolamines, were inversely associated with risk of colorectal cancer in our study. The associations did not statistically significantly vary by time intervals between blood draw and cancer diagnosis, suggesting that these associations were less likely to be affected by potential bias due to reverse causation. Phosphatidylcholines can be converted from cytidine diphosphocholines, which is originally derived from choline.14 Prospective studies have reported that a high level of circulating choline is associated with an increased risk of colorectal cancer.15,16 Phosphatidylcholines can also be synthesized in the liver via the phosphatidylethanolamine methylation pathways.14,17 Phosphatidylethanolamines are derived from cytidine diphosphate ethanolamines, which are downstream derivatives of ethanolamine. Phosphatidylcholines and phosphatidylethanolamines are key components of cell membranes, responsible for maintaining structural integrity. Furthermore, phosphatidylethanolamines are enriched in the mitochondria inner membrane. They regulate key mitochondrial functions such as energy production.18 Phosphatidylcholines and phosphatidylethanolamines also play critical roles in lipoprotein secretion, lipid droplet formation, de novo lipogenesis, and influence the insulin sensitivity of muscles. Therefore, dysregulation of phosphatidylcholines and phosphatidylethanolamines is likely to have a significant impact on lipid profiles, energy balance, and insulin signaling,18 all of which support key functions in cells. Another important property of phosphatidylcholines, linking it to tumorigenesis, is their potential anti-inflammatory effects.19 Chronic inflammation, for example inflammatory bowel disease, is an established risk factor for colorectal cancer.20 With a few exceptions, most of the phosphatidylcholines and phosphatidylethanolamines were found in our study to be inversely associated with risk of colorectal cancer, consistent with their potential protective roles in the etiology of these cancers either via maintaining cell normal functions or anti-inflammatory functions.

Several aromatic compounds were found to be associated with colorectal cancer risk, although most of the observed associations do not appear to be independent from other lipid metabolites. In particular, a positive association was found with coumarin. Coumarin exists in many plants that have a sweet scent. It was initially used as a food additive but later was banned because of its hepatotoxicity in animals. It was also used as a medicine to treat edemas caused by venous and lymphatic drainage disorders. Severe hepatotoxicity adverse events were reported that led to its market withdrawal.21 Currently, no convincing evidence has been found for potential carcinogenic effect of coumarin in humans. Another risk-associated aromatic compound, 3-aminobenzoic acid, which is used as an intermediate for dyes and pesticides, has been labeled as a toxin or pollutant (from HMDB).

We observed increased risk of colorectal cancer in subjects with high levels of picolinic acid. Picolinic acid, a metabolite of tryptophan catabolism, is produced under inflammatory conditions.22 It may mediate immunosuppression by inhibiting proliferation and metabolic activity of CD4+ T cells.23 On the other hand, picolinic acid has also been reported to have an anti-tumor effect, with a reduced tumor cell proliferation found in an animal study administered with picolinic acid.24 More research on this metabolite is needed.

The observed increased risk of colorectal cancer in subjects with high levels of selenocystine, an organic selenium compound, is in contrast to previous laboratory findings that selenocystine can induce apoptosis of cancer cells via reactive oxygen species generation25,26. Nonetheless, reactive oxygen species itself is a double-edged sword in cancer etiology27 raising the possibility that long-term exposure to high selenocystine levels may be cancer-promoting, which needs to be further investigated.

With a prospective study design, the concern of reverse causation is mitigated, compared to retrospective case-control studies. No statistically significant difference was found for the identified associations when stratified by time interval between blood draw and cancer diagnosis; thus, our findings are less likely to be explained by the presence of occult disease. One limitation of our study is that we can only assess metabolites captured by the metabolomics platform used in this study, and thus other metabolites of interest could not be included in the current study. Although we have adjusted for multiple comparisons using the FDR method, we could not rule out entirely that some of the positive association may be due to chance alone, and thus replicating findings from this study would be helpful. The generalisability of our findings remains unclear until similar studies are conducted in other ethnic groups. With an AUC of 0.76, the models including metabolite markers identified in this study showed a moderate discriminatory accuracy in identifying high-risk individuals for colorectal cancer. However, these models need to be validated in future studies as over-fitting might be a potential issue. Some of the findings from this study could not be explained satisfactorily given our current limited knowledge on cancer biology. Some of the cancer-associated metabolites identified in this study may be associated with cancer risk via their correlation with other metabolites not evaluated in this study. Therefore, future studies are needed to understand the mechanisms for the associations identified in this study.

In summary, we have identified 35 metabolites associated with colorectal cancer risk in a prospective metabolomics study. Dysregulated glycerophospholipids are most abundantly related to colorectal cancer in this study. Independent studies are warranted to replicate our findings.

Supplementary Material

Supp FigS1
Supp TableS1
Supp TableS2
Supp TableS3
Supp TableS4

Novelty and Impact.

Metabolomics has been widely applied to search for novel biomarkers for risk of cancer and other diseases. However, few metabolomics studies have been conducted for colorectal cancer risk. In this large metabolomics study using samples collected prior to cancer diagnosis conducted in Chinese, we identified 35 metabolites were associated with risk of colorectal cancer after adjusting for multiple comparisons. The majority of risk-associated metabolites were lipids with glycerophospholipids being the most abundant. Our findings provide novel evidence that dysregulation in glycerophospholipids may be important in the development of colorectal cancer. The metabolite biomarkers found in this study may be useful in identifying high risk individuals for screening and chemoprevention.

Acknowledgments

The authors thank Guochong (Damon) Jia for his help with literature reviews. The authors also thank study participants and research staff of the SWHS and SMHS for their contributions and commitment to this project.

Funding

This work was supported in part by National Institutes of Health/ National Cancer Institute grants (U.S.A), (UM1 CA182910, UM1 CA173640), and the National Key Basic Research Program "973 project" grant (China), No. 2015CB554000.

Abbreviations

OR

Odds ratio

CI

Confidence interval

SD

Standard deviation

BH-FDR

The Benjamini-Hochberg false discovery rate

SWHS

The Shanghai Women Health Study

SMHS

The Shanghai Men Health Study

GC-TOFMS

Gas chromatography time-of-flight mass spectrometry

UPLC-QTOFMS

Ultra-performance liquid chromatography quadrupole-time-of-flight mass spectrometry

QC

Quality control

BMI

Body mass index

AUC

The area under the curve

Footnotes

Competing Interests: None declared

Author Contributions

Study concept and design: Xiao-Ou Shu, Wei Zheng

Acquisition of data: Yong-Bing Xiang, Hong-Lan Li, Gong Yang, Yu-Tang Gao, Wei Jia, Xiao-Ou Shu, Wei Zheng

Statistical analysis: Xiang Shu, Danxia Yu, Hui Cai

Interpretation of data: Xiang Shu, Xiao-Ou Shu, Wei Zheng

Drafting of the manuscript: Xiang Shu, Wei Zheng

Critical revision of the manuscript for important intellectual content: Nathaniel Rothman, Danxia Yu, Qing Lan, Yu-Tang Gao, Wei Jia, Xiao-Ou Shu

Obtained funding: Xiao-Ou Shu, Wei Zheng

Study supervision: Wei Zheng

References

  • 1.Ferlay J, Soerjomataram I, Ervik M, Dikshit R, Eser S, Mathers C, Rebelo M, Parkin D, Forman D, Bray F. GLOBOCAN 2012 v1.0, Cancer Incidence and Mortality Worldwide: IARC CancerBase No. 11 [Internet] Lyon, France: International Agency for Research on Cancer; 2013. [accessed on 24/05/2017]. Available from: http://globocan.iarc.fr. [Google Scholar]
  • 2.Center MM, Jemal A, Smith RA, Ward E. Worldwide variations in colorectal cancer. CA Cancer J Clin. 2009;59:366–78. doi: 10.3322/caac.20038. [DOI] [PubMed] [Google Scholar]
  • 3.Haggar FA, Boushey RP. Colorectal cancer epidemiology: incidence, mortality, survival, and risk factors. Clin Colon Rectal Surg. 2009;22:191–7. doi: 10.1055/s-0029-1242458. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Guraya SY. Association of type 2 diabetes mellitus and the risk of colorectal cancer: A meta-analysis and systematic review. World J Gastroenterol. 2015;21:6026–31. doi: 10.3748/wjg.v21.i19.6026. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Zhu S, St-Onge M-P, Heshka S, Heymsfield SB. Lifestyle behaviors associated with lower risk of having the metabolic syndrome. Metabolism. 2004;53:1503–11. doi: 10.1016/j.metabol.2004.04.017. [DOI] [PubMed] [Google Scholar]
  • 6.Dang CV. Links between metabolism and cancer. Genes Dev. 2012;26:877–90. doi: 10.1101/gad.189365.112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Ward PS, Thompson CB. Metabolic reprogramming: a cancer hallmark even warburg did not anticipate. Cancer Cell. 2012;21:297–308. doi: 10.1016/j.ccr.2012.02.014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Zhang F, Zhang Y, Zhao W, Deng K, Wang Z, Yang C, Ma L, Openkova MS, Hou Y, Li K. Metabolomics for biomarker discovery in the diagnosis, prognosis, survival and recurrence of colorectal cancer: a systematic review. Oncotarget. 2017;8:35460–72. doi: 10.18632/oncotarget.16727. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Farshidfar F, Weljie AM, Kopciuk KA, Hilsden R, McGregor SE, Buie WD, MacLean A, Vogel HJ, Bathe OF. A validated metabolomic signature for colorectal cancer: exploration of the clinical value of metabolomics. Br J Cancer. 2016;115:848–57. doi: 10.1038/bjc.2016.243. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Cross AJ, Moore SC, Boca S, Huang W-Y, Xiong X, Stolzenberg-Solomon R, Sinha R, Sampson JN. A prospective study of serum metabolites and colorectal cancer risk. Cancer. 2014;120:3049–57. doi: 10.1002/cncr.28799. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Zheng W, Chow W-H, Yang G, Jin F, Rothman N, Blair A, Li H-L, Wen W, Ji B-T, Li Q, Shu X-O, Gao Y-T. The Shanghai Women’s Health Study: rationale, study design, and baseline characteristics. Am J Epidemiol. 2005;162:1123–31. doi: 10.1093/aje/kwi322. [DOI] [PubMed] [Google Scholar]
  • 12.Shu X-O, Li H, Yang G, Gao J, Cai H, Takata Y, Zheng W, Xiang Y-B. Cohort Profile: The Shanghai Men’s Health Study. Int J Epidemiol. 2015;44:810–8. doi: 10.1093/ije/dyv013. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Qiu Y, Cai G, Su M, Chen T, Zheng X, Xu Y, Ni Y, Zhao A, Xu LX, Cai S, Jia W. Serum metabolite profiling of human colorectal cancer using GC-TOFMS and UPLC-QTOFMS. J Proteome Res. 2009;8:4844–50. doi: 10.1021/pr9004162. [DOI] [PubMed] [Google Scholar]
  • 14.Ueland PM. Choline and betaine in health and disease. J Inherit Metab Dis. 2011;34:3–15. doi: 10.1007/s10545-010-9088-4. [DOI] [PubMed] [Google Scholar]
  • 15.Bae S, Ulrich CM, Neuhouser ML, Malysheva O, Bailey LB, Xiao L, Brown EC, Cushing-Haugen KL, Zheng Y, Cheng T-YD, Miller JW, Green R, Lane DS, Beresford SA, Caudill MA. Plasma choline metabolites and colorectal cancer risk in the Women’s Health Initiative Observational Study. Cancer Res. 2014;74:7442–52. doi: 10.1158/0008-5472.CAN-14-1835. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Guertin KA, Li XS, Graubard BI, Albanes D, Weinstein SJ, Goedert JJ, Wang Z, Hazen SL, Sinha R. Serum Trimethylamine N-oxide, Carnitine, Choline, and Betaine in Relation to Colorectal Cancer Risk in the Alpha Tocopherol, Beta Carotene Cancer Prevention Study. Cancer Epidemiol Biomark Prev Publ Am Assoc Cancer Res Cosponsored Am Soc Prev Oncol. 2017;26:945–52. doi: 10.1158/1055-9965.EPI-16-0948. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.David L, Nelson, Michael M. Cox. Principles of Biochemistry. (5) 2008:827–829. [Google Scholar]
  • 18.van der Veen JN, Kennelly JP, Wan S, Vance JE, Vance DE, Jacobs RL. The critical role of phosphatidylcholine and phosphatidylethanolamine metabolism in health and disease. Biochim Biophys Acta. 2017;1859:1558–72. doi: 10.1016/j.bbamem.2017.04.006. [DOI] [PubMed] [Google Scholar]
  • 19.Treede I, Braun A, Sparla R, Kühnel M, Giese T, Turner JR, Anes E, Kulaksiz H, Füllekrug J, Stremmel W, Griffiths G, Ehehalt R. Anti-inflammatory effects of phosphatidylcholine. J Biol Chem. 2007;282:27155–64. doi: 10.1074/jbc.M704408200. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Beaugerie L, Itzkowitz SH. Cancers complicating inflammatory bowel disease. N Engl J Med. 2015;372:1441–52. doi: 10.1056/NEJMra1403718. [DOI] [PubMed] [Google Scholar]
  • 21.Abraham K, Wöhrlin F, Lindtner O, Heinemeyer G, Lampen A. Toxicology and risk assessment of coumarin: focus on human data. Mol Nutr Food Res. 2010;54:228–39. doi: 10.1002/mnfr.200900281. [DOI] [PubMed] [Google Scholar]
  • 22.Bosco MC, Rapisarda A, Reffo G, Massazza S, Pastorino S, Varesio L. Macrophage activating properties of the tryptophan catabolite picolinic acid. Adv Exp Med Biol. 2003;527:55–65. doi: 10.1007/978-1-4615-0135-0_6. [DOI] [PubMed] [Google Scholar]
  • 23.Prodinger J, Loacker LJ, Schmidt RLJ, Ratzinger F, Greiner G, Witzeneder N, Hoermann G, Jutz S, Pickl WF, Steinberger P, Marculescu R, Schmetterer KG. The tryptophan metabolite picolinic acid suppresses proliferation and metabolic activity of CD4+ T cells and inhibits c-Myc activation. J Leukoc Biol. 2016;99:583–94. doi: 10.1189/jlb.3A0315-135R. [DOI] [PubMed] [Google Scholar]
  • 24.Leuthauser SW, Oberley LW, Oberley TD. Antitumor activity of picolinic acid in CBA/J mice. J Natl Cancer Inst. 1982;68:123–6. [PubMed] [Google Scholar]
  • 25.Chen T, Wong Y-S. Selenocystine induces reactive oxygen species-mediated apoptosis in human cancer cells. Biomed Pharmacother Biomedecine Pharmacother. 2009;63:105–13. doi: 10.1016/j.biopha.2008.03.009. [DOI] [PubMed] [Google Scholar]
  • 26.Misra S, Boylan M, Selvam A, Spallholz JE, Björnstedt M. Redox-active selenium compounds--from toxicity and cell death to cancer treatment. Nutrients. 2015;7:3536–56. doi: 10.3390/nu7053536. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Pan J-S, Hong M-Z, Ren J-L. Reactive oxygen species: a double-edged sword in oncogenesis. World J Gastroenterol. 2009;15:1702–7. doi: 10.3748/wjg.15.1702. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supp FigS1
Supp TableS1
Supp TableS2
Supp TableS3
Supp TableS4

RESOURCES