Abstract
Context
Type 2 diabetes (T2D) is a major global concern, with Asia at its epicenter in recent years. Proteins, products of gene transcription, serve as dynamic biomarkers for pinpointing perturbed pathways in disease development. Previous T2D proteomic association studies primarily focused on European populations.
Objective
The aim of this study was to investigate the relationship between plasma proteins and the incidence of T2D in Asian individuals.
Methods
We examined the association of 4775 plasma proteins with incident T2D in a Singapore multi-ethnic cohort of 1659 Asian individuals (539 cases and 1120 controls) using logistic regression. We used 2-sample mendelian randomization and colocalization analysis to evaluate the causal relationship between proteins and T2D.
Results
Our analysis revealed 522 proteins that were associated with incident T2D after adjusting for age, sex, and ethnicity, and 17 proteins that remained statistically significantly associated after adjusting for other T2D risk factors such as fasting glucose, waist circumference, and triglycerides. Among the 522 proteins associated with incident T2D, the change in 205 plasma proteins, observed in parallel with the development of T2D at baseline and 6-year follow-up, were further associated with incident T2D. The associated proteins showed enrichment in neuron generation, glycosaminoglycan binding, and insulin-like growth factor binding. Two-sample mendelian randomization analysis suggested 3 plasma proteins, GSTA1, INHBC, and FGL1, play causal roles in the development of T2D, with colocalization evidence supporting GSTA1 and INHBC.
Conclusion
Our findings reveal plasma protein profiles linked to the onset of T2D in Asian populations, offering insights into the biological mechanisms of T2D development.
Keywords: asian populations, proteomics, incident type 2 diabetes, longitudinal study, mendelian randomization, colocalization
Type 2 diabetes (T2D) and its complications substantially contribute to the burden of mortality and disability globally (1, 2). Lifestyle factors, genetic predisposition, and other risk factors jointly play a role in the pathogenesis of T2D, resulting in multiple dysfunctional pathways involving many organ systems (3, 4). More than 650 T2D-associated genetic signals have been identified in transancestry genetic association studies, but translating these genetic associations to better understand T2D pathophysiology has not been straightforward (5-9). Most associated variants are in noncoding regions, and the molecular mechanisms linking variants to T2D development are unclear (6). In addition, genetic variants fail to capture the effects of lifestyle exposures and their interactions with genetic factors leading to T2D.
Proteins, as products of gene transcription, function as dynamic biomarkers reflecting genetic regulation and environmental changes, which can aid in pinpointing the perturbed proteins and pathways in T2D pathogenesis (10). Large-scale population proteomic studies on plasma proteins and complex diseases, including T2D, have become possible in recent years as high-throughput proteomic profiling technologies improve (11-16). Mendelian randomization (MR) uses genetic variants to assess causality and has been used to integrate the genome and proteome to reveal causal proteins and potential T2D therapeutic targets (12, 13, 16, 17). Many proteins, including ACY1, MXRA8, and ADIPOQ, are associated with incident T2D after accounting for factors such as body mass index (BMI) and fasting plasma glucose (12, 13), and are highly consistent across European and African American populations (15). Compared to T2D-proteomics studies conducted in European populations, similar studies in Asian populations are limited and tended to have smaller sample sizes (18, 19). A recent study from the China Kadoorie Biobank (CKB) that included 1896 Chinese individuals identified 33 proteins associated with incident T2D, and 3 of these proteins had a potential causal link to T2D (18). As the pathophysiology of T2D may vary across populations (20, 21), the transferability of these findings and previously unreported proteins to Asian populations can highlight similar and unique pathways associated with the onset of T2D to better understand this disorder.
Using data from 1659 participants (539 incident cases and 1120 controls) of the Singapore Multi-Ethnic Cohort Phase 1 (MEC1), we examined the association of 4775 plasma proteins with incident T2D. In a subset of participants, we further assessed the change in plasma proteins with the development of T2D using repeated measurements collected at 2 time points (baseline and follow-up). Finally, we evaluated the causality of the relationship between these proteins and T2D using MR and colocalization analyses to identify potential therapeutic targets.
Materials and Methods
Study Cohort
Participants in this study were sampled from the Singapore MEC1, a population-based cohort study that aims to discover how lifestyle factors, genetic factors, and their interactions affect the development of chronic health conditions in Singapore (https://blog.nus.edu.sg/sphs/population-studies/multi-ethnic-cohort-phase-1-mec1/) (22). Participants were recruited between 2004 and 2010 for a baseline assessment and were invited to a follow-up visit between 2011 and 2016 (mean follow-up duration of 6.3 ± 1.6 years). At both baseline and follow-up, participants completed a standardized interviewer-administered questionnaire that included questions on sociodemographic and lifestyle factors, personal and family medical history, and medication use. Participants underwent a physical examination during which anthropometric and blood pressure measurements were taken by trained staff (supplemental note (23)) (22). Fasting blood samples were drawn for biomarker measurements and stored at −80 °C the same day. Written consent was obtained from all participants, and this study was approved by the National University of Singapore Institutional Review Board (reference codes B-16-158 and N-18-059).
Using MEC1, we carried out a nested case-control study for incident T2D. Incident T2D cases were ascertained through linkage with records on the national health care database, self-reported history of T2D at follow-up (physician diagnosis or use of diabetes medication), or having fasting glucose of 7 mmol/L or greater, or glycated hemoglobin A1c of 6.5% or greater, or random blood glucose of 11 mmol/L or greater at follow-up following the American Diabetes Association criteria for diabetes (24). Controls were selected by incident density sampling matching (1 case to 2 controls) by age (±5 years), sex, ethnicity, and date of blood collection (±2 years). The final sample for the analysis of baseline proteomics and incident T2D association included 539 cases and 1120 controls (Supplementary Fig. S1 and supplemental note (23)). For protein change analysis, we included only newly diagnosed cases at follow-up visit using fasting glucose or glycated hemoglobin A1c, resulting in 216 cases and 458 controls at baseline and follow-up (see Supplementary Fig. S1 and supplemental note (23)).
Protein Profiling
The relative concentrations of proteins were measured in plasma samples from MEC1 participants using the aptamer-based SomaScan assay v4 platform (SomaLogic Inc) (25) (https://mohanlab.bme.uh.edu/wp-content/uploads/2017/02/SSM-002-Rev-4-SOMAscan-Technical-White-Paper.pdf). SomaLogic employed an in-house library (Naptamers = 5272) of modified aptamers to target specific proteins. The protein concentrations were quantified using a DNA microarray and expressed in terms of relative fluorescent units (RFU).
The routine quality control processes used by SomaLogic consist of several steps with different control samples as the reference (see supplemental note (23)). After normalization, 82 samples were excluded with scale factors beyond the range of acceptance criteria (ie, 0.4-2.5) (see Supplementary Fig. S1 (23)). In addition, we applied principal component analysis on log2-transformed protein levels at the sample and aptamer level to identify potential outliers. No additional samples or aptamers were excluded due to excessive principal component values and/or interquartile ranges values (Supplementary Figs. S2 and S3 (23)). The median coefficient of variation for aptamers with both calibrator controls and quality controls were below 15% for all plates (Supplementary Fig. S4 (23)). Based on the aptamer annotation, we excluded 294 aptamers, such as nonhuman proteins, nonprotein aptamers, and aptamers without Entrez gene symbols or UniProt IDs. Finally, 4978 aptamers targeting 4775 unique proteins or protein isoforms remained for subsequent analysis.
Statistical Analysis
We applied 2 transformations to protein RFU levels at each time point: log2 transformation for interpretability of the effect sizes and rank-based inverse normal transformation to calculate P values to assess statistical significance. We reported the effect sizes obtained from log2-transformed protein level associations and their corresponding P values resulting from rank-based inverse normal transformation, unless otherwise specified. The log2-transformed protein levels and phenotypic trait values at different time points were winsorized separately within a range of ±5 SD to reduce the influence of outliers. Insulin and triglyceride (TG) levels were log-transformed due to skewed distributions. Logistic regression model was used to assess the association of baseline protein levels with incident T2D with adjustment for age, sex, and ethnicity (model 1) to maximize samples that no longer had matched case/control due to sample loss (see supplemental note and Supplementary Fig. S5 (23)). A Bonferroni threshold of P equal to 1.05 × 10−5 (0.05/4775 proteins) was considered statistically significant. If more than one aptamer was targeted for the same protein, only the aptamer with the smallest P value in model 1 was retained for subsequent analysis. We sequentially included fasting glucose (model 2), waist circumference (model 3), and TGs (model 4). The final model was derived using stepwise logistic regression analysis using the Bayesian information criterion (BIC), which imposes a higher penalty on each parameter and tends to select simpler models. Variables in the full model (model 5) included age, sex, ethnicity, fasting glucose, waist circumference, TG, systolic blood pressure (SBP), and family history of diabetes. Based on the BIC, other T2D-associated variables, such as BMI, hip circumference, diastolic blood pressure, high-density lipoprotein (HDL), low-density lipoprotein (LDL), C-reactive protein, blood pressure-lowering medication, and a history of hypertension, were not selected for model 5 (see Supplementary Table S1 (23)).
For each of the plasma proteins associated with incident T2D in our study, we further examined the associations between their longitudinal changes over time with T2D. We first computed the log2-transformed fold change in protein levels between two time points, referred to as the “protein difference”. Similar to the baseline association, we used this protein difference for effect size interpretation, and both the protein differences and the baseline protein levels were winsorized within a range of ±5 SD. Additionally, we applied a rank-based inverse normal transformation of these protein differences to assess the statistical significance. Logistic regression models were used to investigate the relationship between these protein differences and incident T2D. The basic model adjusted for age at baseline, sex, ethnicity, and corresponding baseline protein level. We sequentially included baseline fasting glucose, family history of diabetes, and changes in waist circumference, TG, and SBP between baseline and follow-up into the models. All analyses were performed in R version 4.2.2.
Protein Signatures for Risk Factors
We applied elastic net with logistic regression models to derive protein signatures altered by risk factors and their association with incident T2D. Using 522 proteins associated with T2D, we implemented a 100-times repeated 10-fold cross-validation approach to identify protein signatures of key risk factors, as well as to evaluate the association of these protein signatures with incident T2D. Elastic net regression was used to optimize the predictive model by selecting the most informative protein predictors with the R package “glmnet” (version 4.1-8) (26). This elastic net analysis was performed in all 1659 participants. For each participant, protein signature scores were calculated as the weighted sum of the abundances of the selected proteins multiplied by the β coefficients for these proteins derived from the elastic net model based on all participants. The association between risk factors or their corresponding protein signatures and incident T2D were adjusted for age, sex, and ethnicity using R version 4.2.2.
Functional Enrichment Analysis
Enrichment analyses of baseline proteins and changes in proteins associated with incident T2D were performed to identify biological pathways. We applied the g:Profiler web server (https://biit.cs.ut.ee/gprofiler/gost) to perform functional enrichment for each protein category with a statistical significance threshold of Benjamini-Hochberg false discovery rate (FDR)-corrected P value less than .05 (27). We also estimated tissue specificity using TissueEnrich with Human Protein Atlas data (28). Both enrichment analyses used the SomaScan v4 panel protein (N = 4775) as the background list (see supplemental note (23)).
Whole-Genome Data
Whole-genome data were available for a subset of MEC1 participants from the National Precision Medicine Programme Phase I (https://npm.a-star.edu.sg/) (29). Samples were whole-genome sequenced to an average of 15× coverage. Read alignment was performed with BWA-MEM (version 0.7.17). Variant discovery and genotyping were performed with GATK (version 4.0.6.0). Site-level filtering included only retaining VQSR-PASS and non-STAR allele variants. At the sample level, samples with call rates less than 95%, BAM cross-contamination rates greater than 2%, or BAM error rates greater than 1.5% were excluded. At the genotype call level, genotypes with depth (DP) less than 5 in any ancestry group or genotype quality (GQ) less than 20 or allele balance (AB) greater than 0.8 (heterozygotes calls) were set to missing. Genetic variants were filtered to exclude those with robust, unified test for Hardy-Weinberg equilibrium (RUTH) P value less than .01, a variant call rate less than 90%, being monomorphic, or having a minor allele count less than 5 prior to phasing with Eagle version 2.4 (30, 31). After quality control, the data set included 16 014 113 genetic variants in a sample set of 1753 individuals with both genetic and proteomic data.
Mendelian Randomization
To evaluate the potential causal role of the proteins associated with T2D, we performed 2-sample MR analyses using transancestry protein quantitative trait locus (pQTL) data from 1753 individuals both with genetic and proteomic information, along with summary statistics from independent multiancestry genome-wide association studies (GWAS) for T2D (N = 1 339 889) (see supplemental note (23)) (9). Among the 522 proteins associated with T2D, we identified genetic instruments for 139 (26.6%) proteins by examining independent genetic variants within a 1-Mb local region around each protein-encoding gene (see supplemental note (23)). The inverse variance-weighted method was used to estimate the combined effect size of multiple instrumental variables, and the Wald ratio was used when only one variant instrument was available. FDR-corrected P values less than .05 were considered statistically significant. We used the Cochran Q statistic to assess instrument heterogeneity for proteins with more than 1 genetic instrument and the MR-Egger regression intercept to detect horizontal pleiotropy for proteins with more than 2 genetic instruments. A nonzero intercept in MR-Egger regression indicates the presence of pleiotropy in the variant (32). R package “TwoSampleMR” was used for MR analyzes (33).
Colocalization Analysis
To assess the colocalization of the pQTLs and T2D association, we estimated the posterior probability of a shared genetic variant both for protein and T2D within a Bayesian framework at each locus (34). Using the same independent trans-ancestry T2D GWAS summary statistics used for MR analysis (9), colocalization evidence was evaluated using posterior probabilities (PP) for hypothesis 4 (H4), indicating the joint presence of associations for both traits driven by the same causal variant. Associations with PP H4 greater than 0.5 were defined as “likely to colocalize.” while PP H4 greater than 0.8 indicated a “highly likely” colocalization. The colocalization analysis was performed using the “coloc” package in R.
Results
Study Sample Characteristics
Fig. 1 summarizes the study design and the number of samples included for each analysis. For the association of baseline proteins and T2D incidence, we included 539 incident T2D cases (218 Chinese, 165 Malay, 156 Indian) and 1120 controls (409 Chinese, 348 Malay, 363 Indian). For the longitudinal analysis, we excluded self-reported T2D cases and included only 216 newly diagnosed, treatment-naive cases (88 Chinese, 64 Malay, 64 Indian) identified at the follow-up visit and 458 controls (200 Chinese, 139 Malay, 119 Indian) who had plasma proteins measured at both time points (see Supplementary Fig. S1 (23)). The baseline and follow-up characteristics of incident T2D cases and controls, and the longitudinal phenotype traits changes, are presented in Supplementary Table S1 (23). At baseline, participants who later developed T2D had a higher BMI, waist circumference, systolic blood pressure, and TGs, and lower HDL-cholesterol than the controls (all P < 1 × 10−12). However, in the longitudinal analysis, changes in these risk factors were not associated with T2D (all P > .05).
Figure 1.
Summary of the study design, methods, and primary results.
Incident Type 2 Diabetes Association With Proteins at Baseline
We identified 522 proteins significantly associated with incident T2D after adjusting for age, sex, and ethnicity at baseline (model 1; PBonferroni < 1.05 × 10−5). The strongest associations observed at P less than 1.05 × 10−5 included MXRA8 (odds ratio [OR] 0.05; 95% CI, 0.03-0.08; P = 6.79 × 10−36), ADIPOQ (OR 0.25; 95% CI, 0.20-0.31; P = 1.30 × 10−34), SLITRK3 (OR 0.30; 95% CI, 0.25-0.37; P = 1.80 × 10−33), and ADH4 (OR 2.12; 95% CI, 1.87-2.41; P = 1.53 × 10−29) (Fig. 2A and Supplementary Table S2 (23)). Among the 522 associated proteins identified in our study, 314 (60.64%) proteins have not been previously reported to be associated with T2D risk (Supplementary Fig. S6 (23)). We compared our findings to 6 published studies (AGES–Reykjavik Study, Atherosclerosis Risk in Communities [ARIC], Jackson Heart Study [JHS], Framingham Heart Study and Malmö Diet and Cancer Study [FHS-MDCS], CKB, and the Guangzhou Nutrition and Health Study [GNHS] on proteomics and T2D incidence in Chinese, European, and African populations (12, 13, 15, 16, 18, 19)). Across the 6 published studies, 550 proteins were reported to be statistically significant at a Bonferroni threshold (∼1 × 10−5) or had an FDR q value less than 0.05 in at least one study. Of these, 488 (88.73%) were present in our data and 390 (79.92% of 488) showed directionally consistent associations with MEC1 (Supplementary Fig. S6 and Supplementary Table S3 (23)). Additionally, 16 proteins were statistically significant across all 5 (MEC1, AGES-Reykjavik, ARIC, JHS, and FHS-MDCS) studies in their basic models (see Supplementary Table S3 (23)).
Figure 2.
Plasma protein associations with incident type 2 diabetes (T2D) using logistic regression. A, Volcano plot of -log10(P) of plasma protein associations with incident T2D with each point representing proteins with P less than 1.05 × 10−5. Proteins with P less than 1 × 10−20 are indicated in the plot. B, Forest plot of odds ratios, CIs, and P values for the 17 proteins associated with incident T2D in model 5 from MEC1 (Ncases = 539 and Ncontrols = 1120). Protein levels were inverse-normal transformed. Model 1 includes age, sex, and ethnicity. Model 5 includes age, sex, ethnicity, fasting glucose, waist circumference, triglycerides, systolic blood pressure, and family history of diabetes. Odds ratio on the x-axis is presented on the log scale. C, Box plots depicting baseline Insulin-like growth factor-binding protein 7 (IGFBP7) protein levels in T2D cases and controls across diverse ethnicity groups (Chinese, Malay, and Indian). Y-axis represents log2-transformed IGFBP7 protein levels at baseline. D, Forest plot of odds ratios, CIs, and P values for the association between IGFBP7 and incident T2D risk in Chinese, Malay, and Indian separately. MEC1, Singapore Multi-Ethnic Cohort Phase 1.
We sequentially adjusted for additional T2D risk factors through a stepwise model selection process. After adjusting for fasting glucose, 344 proteins (65.90% of 522 in model 1) remained statistically significant. After further adjustment for waist circumference (105 proteins; 20.11% of 522) and both waist circumference and TGs (22 proteins; 4.21% of 522), a markedly smaller number of proteins remained statistically significant, demonstrating the substantial influence of these risk factors on the association between proteins and incident T2D (Supplementary Fig. S7; Supplementary Tables S2, S4, and S5 (23)). Specifically for waist circumference adjustment in model 3, 45.79% (N = 239) of the signals were attenuated and no longer statistically significant (see Supplementary Table S5 (23)). Lastly, in model 5 with additional variables SBP and family history of diabetes, 17 (3.26% of 522) proteins remained statistically significantly associated with T2D risk at the Bonferroni-corrected threshold (Table 1 and Fig. 2B). Additionally, associations adjusted for BMI instead of waist circumference in models 3, 4, and 5 showed high consistency with waist circumference–adjusted models with Pearson correlation coefficients exceeding 0.99 (Supplementary Table S6 and Supplementary Fig. S8 (23)).
Table 1.
List of 17 plasma proteins associated with incident type 2 diabetes in model 5 adjusted for age, sex, ethnicity, fasting glucose, waist circumference, triglycerides, systolic blood pressure, and family history of diabetes at baseline in the Singapore Multi-Ethnic Cohort Phase 1 study
| Target full name | Entrez gene symbol | UniProt ID | Baseline (N = 1659) | Longitudinal (N = 674) | ||||
|---|---|---|---|---|---|---|---|---|
| Model 1 | Model 5 | Protein change model | ||||||
| OR (95% CI) | P | OR (95% CI) | P | OR (95% CI) | P | |||
| SLIT and NTRK-like protein 3 | SLITRK3 | O94933 | 0.30 (0.25-0.37) | 1.80E-33 | 0.48 (0.39-0.61) | 1.08E-10 | 0.19 (0.11-0.33) | 9.44E-10 |
| Matrix-remodeling-associated protein 8 | MXRA8 | Q9BRK3 | 0.05 (0.03-0.08) | 6.79E-36 | 0.20 (0.11-0.35) | 1.79E-08 | 0.03 (0.01-0.08) | 4.44E-13 |
| Growth arrest-specific protein 7 | GAS7 | O60861 | 0.21 (0.12-0.37) | 2.90E-12 | 0.27 (0.15-0.48) | 3.89E-08 | 0.22 (0.07-0.64) | 5.52E-04 |
| ADAMTS-like protein 2 | ADAMTSL2 | Q86TH1 | 5.39 (3.82-7.61) | 2.14E-22 | 2.87 (1.96-4.22) | 8.18E-08 | 11.92 (5.98-23.77) | 1.01E-14 |
| Adiponectin | ADIPOQ | Q15848 | 0.25 (0.20-0.31) | 1.30E-34 | 0.51 (0.39-0.66) | 1.27E-07 | 0.36 (0.20-0.63) | 4.54E-04 |
| Advanced glycosylation end product-specific receptor, soluble | AGER | Q15109 | 0.59 (0.52-0.67) | 6.35E-16 | 0.67 (0.58-0.78) | 1.50E-07 | 0.69 (0.47-1.00) | 4.21E-02 |
| Apolipoprotein C-I | APOC1 | P02654 | 0.10 (0.05-0.20) | 1.06E-11 | 0.13 (0.06-0.28) | 1.87E-07 | 0.06 (0.02-0.18) | 7.84E-07 |
| Lutropin-choriogonadotropic hormone receptor | LHCGR | P22888 | 0.40 (0.26-0.63) | 3.07E-08 | 0.42 (0.26-0.68) | 1.60E-06 | 0.24 (0.07-0.80) | 5.31E-03 |
| Protein S100-A7 | S100A7 | P31151 | 0.33 (0.21-0.52) | 1.54E-10 | 0.42 (0.27-0.67) | 1.77E-06 | 3.14 (1.47-6.74) | 5.48E-02 |
| Aminoacylase-1 | ACY1 | Q03154 | 2.20 (1.91-2.54) | 2.70E-29 | 1.44 (1.23-1.69) | 2.05E-06 | 3.41 (2.53-4.59) | 4.48E-14 |
| Stanniocalcin-1 | STC1 | P52823 | 3.65 (2.60-5.14) | 1.59E-13 | 2.49 (1.70-3.65) | 2.29E-06 | 1.26 (0.66-2.43) | 3.73E-01 |
| T-cell surface glycoprotein CD5 | CD5 | P06127 | 0.34 (0.21-0.54) | 5.03E-11 | 0.46 (0.29-0.75) | 2.45E-06 | 0.58 (0.22-1.58) | 3.28E-02 |
| Cell adhesion molecule 2 | CADM2 | Q8N3J6 | 0.14 (0.09-0.20) | 1.26E-26 | 0.37 (0.24-0.57) | 2.59E-06 | 0.26 (0.13-0.52) | 1.13E-04 |
| Glycerol-3-phosphate dehydrogenase [NAD(+)], cytoplasmic | GPD1 | P21695 | 2.73 (2.27-3.27) | 7.86E-29 | 1.61 (1.31-1.98) | 3.68E-06 | 4.67 (3.11-7.01) | 1.05E-12 |
| Tumor necrosis factor ligand superfamily member 6, soluble form | FASLG | P48023 | 0.67 (0.54-0.83) | 9.77E-10 | 0.75 (0.60-0.93) | 4.16E-06 | 0.31 (0.17-0.54) | 2.37E-07 |
| Myelin-oligodendrocyte glycoprotein | MOG | Q16653 | 0.30 (0.18-0.49) | 2.61E-11 | 0.44 (0.26-0.74) | 5.75E-06 | 0.59 (0.24-1.48) | 9.03E-02 |
| Vesicular, overexpressed in cancer, prosurvival protein 1 | VOPP1 | Q96AW1 | 0.16 (0.11-0.23) | 2.67E-27 | 0.45 (0.30-0.66) | 6.59E-06 | 0.09 (0.04-0.20) | 2.91E-10 |
Model 1: the basic protein-T2D model adjusted for age, sex, and ethnicity at baseline. Model 5: the full protein-T2D model adjusted for age, sex, ethnicity, and T2D risk factors, including fasting glucose, waist circumference, triglycerides, systolic blood pressure, and family history of diabetes, identified 17 significantly associated proteins at baseline (Bonferroni-corrected P = 1.05 × 10−5). Protein change model: the association of protein changes with incident T2D adjusted for age, sex, ethnicity, and baseline protein level. Protein levels were log2 transformed for interpretability of the ORs and rank-based inverse normal transformed for P values.
Abbreviations: OR, odds ratio; T2D, type 2 diabetes.
To assess possible differential effects on T2D across different Asian ethnicity groups, we examined the interaction between 4775 proteins and ethnicity (Chinese, Malay, and Indian) in association with incident T2D. We found a statistically significant interaction between insulin-like growth factor-binding protein 7 (IGFBP7) and ethnicity (Pinteraction = 7.35 × 10−6) in the model adjusted for age, sex, and ethnicity (see Supplementary Table S2 (23)). Further analyses stratified by ethnicity suggested that the interaction was mainly driven by the Indian participants (Indian: OR 23.84; 95% CI, 7.63-74.53; P = 2.55 × 10−9; Malay: OR 5.55; 95% CI, 2.00-15.37; P = 5.35 × 10−4; Chinese: OR 0.94; 95% CI, 0.41-2.16; P = 9.81 × 10−1) (Fig. 2C and 2D). We note the wide CI as a result of smaller sample sizes for the strata. The lead variant (rs1543178, P = 1.48× 10−6) associated with IGFBP7 in the pQTL meta-analysis exhibited variation in minor allele frequency across these ethnic groups: 0.01 in Chinese, 0.05 in Malay, and 0.12 in Indian (Supplementary Fig. S9 (23)).
Association Between Protein Signatures for Risk Factors and Incident Type 2 Diabetes
We used elastic net regression to perform feature selection on 522 incident T2D-associated proteins to derive protein signatures for risk factors, including BMI, waist circumference, hip circumference, SBP, diastolic blood pressure, insulin, C-reactive protein, TG, total cholesterol, HDL, and LDL (Supplementary Table S7 (23)). Among the 1659 MEC1 participants, all of the risk factors and corresponding protein signatures were significantly associated with incident T2D after adjusting for age, sex, and ethnicity, except for the total cholesterol and total cholesterol-protein signature (Supplementary Table S8 (23)). Protein signatures generally showed stronger associations with incident T2D compared to their corresponding clinical measurements, except for total cholesterol, TGs, and LDL.
Change in Plasma Proteins and Type 2 Diabetes
We further assessed the longitudinal change of the 522 incident T2D-associated protein biomarkers in relation to T2D onset. In 216 treatment-naive cases based on glycemic measures at follow-up and 458 controls, changes in the levels of 205 (39.27%) proteins were statistically significantly associated with T2D adjusted for age, sex, ethnicity, and baseline protein level (P < 9.58 × 10−5 = .05/522) (Supplementary Fig. S10A (23)). A consistent direction of association was observed in 96.10% (197/205) of these proteins with a Pearson correlation coefficient of 0.76 (Supplementary Fig. S10B (23)). Adjusting for baseline fasting glucose in the longitudinal analysis reduced the number of statistically significant associations from 205 to 180 (see Supplementary Table S5 (23)). Among these 180 proteins, 125 (69.44%) were associated with differences in fasting glucose measured at 2 time points (see Supplementary Table S4 (23)). Finally, 178 proteins remained associated in the full model with additional adjustment for baseline fasting glucose, family history of diabetes at baseline, and changes in waist circumference, TG, and SBP (Supplementary Tables S5 and S9 (23)).
Functional Enrichment Analysis
The 522 proteins associated with incident T2D were enriched for pathways, including insulin-like growth factor binding, glycosaminoglycan binding, and hormone metabolic processes (Supplementary Fig. S11A and Supplementary Table S10 (23)). While the functional enrichment analysis suggested enrichment in the liver, consistent with previous research in Western populations (Supplementary Fig. S11B (23)) (12), we note that proteins included on the SomaScan panel may be enriched in the liver (see supplemental note (23)). Similar enriched pathways were also observed for the subset of 205 proteins that changed significantly between baseline and follow-up (Supplementary Figs. S11C and S11D and S12 (23)).
Evaluating Causal Relationships Between Proteins and Type 2 Diabetes
Using 2-sample MR analysis to evaluate the causal relationship between proteins and T2D, only 26.63% (N = 139) of the 522 incident T2D-associated proteins had valid genetic instruments for proteins and summary statistics available from Mahajan et al (9). (see supplemental note (23)). In the 2-sample MR analysis, 3 proteins (GSTA1, INHBC, and FGL1) had evidence (P FDR < .05) of a causal effect on T2D (see Table 2 and Supplementary Table S11 (23)). Directions of effect were consistent between causal and observational estimates and both incident and longitudinal analysis (see Table 2). GSTA1 has previously been reported to be causally associated with T2D, while INHBC was identified as a consequence of T2D (12, 13). Colocalization analysis between pQTL instruments and external T2D GWAS showed evidence of colocalization for GSTA1 (PP H4 = 83.70%) and INHBC (PP H4 = 71.20%) but not FGL1 (PP H4 = 1.10%) (see Supplementary Table S11 and Supplementary Fig. S13 (23)).
Table 2.
Proteins with evidence of causal associations with type 2 diabetes in mendelian randomization analysis
| Protein name | Protein | MR analysis | Colocalization analysis | Association with incident T2Da | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Baseline protein (model 1) | Protein change | |||||||||||
| Method | IV (N) | β | SE | P | FDR.P | PP H4 | β | P | β | P | ||
| Glutathione S-transferase A1 | GSTA1 | Wald ratio | 1 | 0.05 | 0.01 | 4.04 × 10−5 | 6.95 × 10−3 | 83.70% | 0.92 | 1.44 × 10−22 | 1.08 | 6.00 × 10−8 |
| Inhibin β C chain | INHBC | Wald ratio | 1 | 0.13 | 0.03 | 2.97 × 10−4 | 2.55 × 10−2 | 71.20% | 1.51 | 1.08 × 10−27 | 2.6 | 1.50 × 10−10 |
| Fibrinogen-like protein 1 | FGL1 | Wald ratio | 1 | −0.06 | 0.02 | 6.17 × 10−4 | 3.54 × 10−2 | 1.10% | −0.39 | 2.35 × 10−6 | −0.39 | 4.50 × 10−2 |
Proteins with statistically significant levels of FDR-corrected P less than .05 are displayed here, while the remaining proteins are listed in Supplementary Table S11.
Abbreviations: FDR.P, false discovery rate–corrected P value; IV (N), number of instrument variants; MR, mendelian randomization; PP H4: posterior probability of hypothesis 4; T2D, type 2 diabetes.
aBoth incident T2D association models were adjusted for age, sex, and ethnicity.
Discussion
We measured 4775 proteins in 1659 participants (539 incident cases and 1120 controls) of Chinese, Indian, and Malay ethnicity with the SomaScan assay. About 80% of previously reported associated proteins (N = 488) showed a consistent direction of effect in our study, representing a validated set of diabetes-associated proteins that were replicated across geographically and diverse populations (12, 13, 15, 16). A test of heterogeneity identified a single protein (IGFBP7) for which there is some evidence of interaction with ethnicity for the association with T2D. However, we recognize that our study had limited power to detect interactions. These findings suggest that many of the biological processes in the pathophysiology of T2D (those captured by measurements of plasma proteins) are shared across ancestry groups despite apparently heterogeneous T2D phenotype manifestations.
We identified 522 proteins associated with incident T2D at the Bonferroni-corrected threshold, including 316 that have not been previously reported to be associated with T2D. Many of these previously unreported associations may have arisen due to the larger number of proteins assayed in our study than earlier versions of the SomaScan platform. We also reported associated proteins with basic age, sex, and ethnicity adjustment. In some previous studies, models adjusted for other physiologic factors known to affect plasma protein levels that may be altered early in the natural history of disease, particularly in T2D with a prolonged natural history, were reported (12, 13, 15, 18). It is commonly accepted that obesity and insulin resistance are early events in the pathogenesis of T2D, with β-cell dysfunction leading to hyperglycemia at a later stage (3). Our data showed that (i) excess adiposity and elevated TG (which are markers for insulin resistance) were already observed at baseline, and (ii) adjustment for adiposity and TG attenuated protein-T2D associations reflecting correlations of proteins with these risk factors. The differences in association results between the fully adjusted model 5 and model 1 adjusted for age, sex, and ethnicity primarily stemmed from adjustments for waist circumference (239 proteins, 45.8%), fasting glucose (178 proteins, 34.1%), and TG (85 proteins, 16.3%) out of 522 proteins in model 1. Adiposity, as a key risk factor for T2D, promotes the elevated release of nonesterified fatty acids and glycerol from adipocyte and influences glucose regulation (35). Fasting glucose levels at baseline may rise mildly in cases, but glucose toxicity develops, leading to β-cell dysfunction and early dysregulation of glucose metabolism (35). TGs, the most abundant lipid in adipose tissue, independently predict higher T2D risk as visceral TG storage produces insulin-resistant adipocytes, increasing nonesterified fatty acids and glycerol release that exacerbate insulin resistance in muscle and liver (36). The 3 proteins identified by MR analysis were also associated with baseline fasting glucose and TGs, and their association significance was attenuated after adjustment for these factors. This highlights the limitations of observational analysis to distinguish between confounding and mediation by other biological risk factors and the value of MR and colocalization analyses to provide evidence on causality.
We derived protein signatures for key T2D risk factors to enhance the predictive power for incident T2D. Our findings demonstrated that protein signatures exhibited stronger associations with incident T2D than their corresponding clinical measurements, suggesting that protein signatures not only reflect the clinical phenotype but also capture additional composite biological variations that may not be directly measured by traditional risk factors. However, these findings lack external validation, and further studies are needed to confirm the generalizability and utility of these protein signatures in other populations. Despite this limitation, integrating protein biomarkers provides a more precise assessment of T2D risk, which could offer potential for personalized prevention strategies and targeted interventions.
The proteins associated with incident T2D and those exhibiting associations across baseline and follow-up longitudinal change with incident T2D were involved in metabolism and the insulin response system, and enriched for liver-specific gene expression, which was consistent with previous studies in Western populations using the SomaScan panel (12, 13). After adjustment for other known risk factors for T2D, 17 proteins remained associated with incident T2D, of which 7 proteins (STC1, GAS7, APOC1, MOG, CD5, S100A7, and LHCGR) have been previously unreported. While these protein biomarkers will likely provide limited predictive gain over more conventional biomarkers such as obesity and fasting plasma glucose, omics biomarkers associated with T2D have the potential to provide insight into the biological pathways involved in the pathophysiology linking known risk factors with T2D development and potentially to the complications of T2D. For example, STC1 receptors are found in mouse pancreatic β calls and that ligand colocalizes with insulin (37), and APOC1 helps prevent insulin resistance by reducing fat storage, even when plasma TG levels are high in experimental studies in mice (38). Insulin-like growth factor-binding protein 7 (IGFBP7) is the only protein that showed an interaction with ethnicity in its association with T2D. IGFBP7 plays a key role in protein synthesis, cell growth, and cell survival (39). This protein's strong binding affinity for insulin suggests its potential role in insulin resistance in T2D (40-42). Further research is needed to evaluate whether these previously unreported protein biomarkers for T2D risk may result from genetic or lifestyle differences between different populations.
MR identified 3 proteins (GSTA1, INHBC, and FGL1) that may have a causal role in T2D development. GSTA1 and INHBC both exhibited consistent directions of effect in causal and observational estimates, further substantiated by colocalization evidence within our data set. Limited by a small sample size, none of the 3 proteins had more than 2 independent genetic instruments that would allow us to test for pleiotropy. We replicated the finding of the causal role of GSTA1 in T2D from the US ARIC study (13). GSTA1 (glutathione S-transferase A1), expressed in liver and kidney tissues, functions as an indicator of hepatic function and plays a pivotal role in the body's antioxidant system in mice (43). INHBC (inhibin β C chain, a protein of the transforming growth factor β family) is secreted by the liver and regulates the secretion of gonadal hormones and insulin (44, 45). However, MR analysis in the AGES-Reykjavik study suggested that altered concentrations of INHBC were a consequence and not a cause of T2D, and that it is associated with both incident and prevalent T2D. We note that INHBC showed statistically significant changes between baseline and follow-up in our data. No causal effects on T2D have previously been reported for FGL1 (fibrinogen-like protein 1), but it was found to promote adipogenesis in animal and cell models (46). None of these 3 proteins remained statistically significant after adjustment for other risk factors. These protein associations with T2D risk were primarily accounted for by specific adjustments: FGL1 by fasting glucose adjustment, and GSTA1 and INHBC by TG adjustment. These findings imply that glucose metabolism, lipid metabolism, and abdominal obesity may mediate the causal link between these proteins and T2D. To bolster the robustness and reliability of our findings, additional validation with larger sample sizes is needed.
Our study has several limitations. First, as Asian studies with data on proteomics and T2D incidence were currently limited in size, we did not validate our findings in independent Asian populations. We compared our results with 2 previous proteomic studies conducted in Chinese populations and demonstrated consistent directional effect sizes with MEC1 results. Among the 5 T2D-associated proteins in GNHS and 24 in CKB detected in both the previous studies and MEC1 protein detection panels (mass spectrometry) or Olink verse SomaScan v4), only 2 (40.0%) in GNHS and 14 (58.3%) in CKB showed consistent directional effect size with MEC1 results. Prioritizing larger Asian population sets in future proteomics research will improve the robustness and transferability of findings. Second, the prolonged natural history of T2D can affect protein levels even in prospective studies, as highlighted by the poorer metabolic risk profile prior to T2D diagnosis. Third, the presence of different statistical significance levels among different aptamers targeting the same protein in our analysis suggests differences in the capture specificity of the assay with a median correlation of about 0.4 observed for overlapping proteins between the SomaLogic and Olink platforms as previously reported (47-49). Therefore, future research should explore combining results across panels. Finally, while we observed an enrichment for liver-specific transcripts from the 522 incident-T2D associated proteins, other organ systems involved in T2D such as the β cell may not secrete significant amounts of protein to be detected using plasma proteomics. This underscores a potential limitation on plasma proteomics panels for understanding disease pathophysiology, emphasizing the need for caution in interpretation and consideration of complementary approaches.
In conclusion, we have identified multiple plasma proteins, including 316 previously unreported associations, associated with incident T2D in a population of diverse Asian ancestry. Our data suggest that many of the pathophysiologic pathways leading to T2D are likely shared between populations of different ancestry. Plasma proteomics is particularly well suited to detect changes in proteins secreted by the liver and for T2D, particularly those associated with obesity, insulin resistance, and glucose metabolism disorder (even relatively mild changes that occur well before the onset of T2D). Our findings on these plasma proteins provide evidence for causality and implicate novel biological pathways associated with the identified protein risk factors, which could be targeted by future preventive interventions.
Acknowledgments
We thank all participants, the study team, and the investigators for their research contributions.
Abbreviations
- ARIC
Atherosclerosis Risk in Communities
- BIC
Bayesian information criterion
- BMI
body mass index
- CKB
China Kadoorie Biobank
- FDR
false discovery rate
- FHS-MDCS
Framingham Heart Study and Malmö Diet and Cancer Study
- FGL1
fibrinogen-like protein 1
- GNHS
Guangzhou Nutrition and Health Study
- GSTA1
glutathione S-transferase A1
- GWAS
genome-wide association study
- H4
hypothesis 4
- HDL
high-density lipoprotein
- IGFBP7
insulin-like growth factor-binding protein 7
- INHBC
inhibin β C chain
- JHS
Jackson Heart Study
- LDL
low-density lipoprotein
- MEC1
Singapore Multi-Ethnic Cohort Phase 1
- MR
mendelian randomization
- PP
posterior probabilities
- pQTL
protein quantitative trait locus
- RFU
relative fluorescent units
- SBP
systolic blood pressure
- TG
triglycerides
- T2D
type 2 diabetes
Contributor Information
Yujian Liang, Saw Swee Hock School of Public Health, National University of Singapore and National University Health System, Singapore 117549, Singapore.
Charlie G Y Lim, Saw Swee Hock School of Public Health, National University of Singapore and National University Health System, Singapore 117549, Singapore.
Scott C Ritchie, Cambridge Baker Systems Genomics Initiative, Department of Public Health and Primary Care, University of Cambridge, Cambridge CB2 0SR, UK; British Heart Foundation Centre of Research Excellence, University of Cambridge, Cambridge CB2 0BB, UK; Victor Phillip Dahdaleh Heart and Lung Research Institute, University of Cambridge, Cambridge CB2 0BB, UK; Health Data Research UK Cambridge, Wellcome Genome Campus and University of Cambridge, Cambridge CB10 1SA, UK; British Heart Foundation Cardiovascular Epidemiology Unit, Department of Public Health and Primary Care, University of Cambridge, Cambridge CB2 0SR, UK.
Nicolas Bertin, Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore 138672, Singapore.
Jin-Fang Chai, Saw Swee Hock School of Public Health, National University of Singapore and National University Health System, Singapore 117549, Singapore.
Jiali Yao, Saw Swee Hock School of Public Health, National University of Singapore and National University Health System, Singapore 117549, Singapore.
Yun Li, Department of Genetics, University of North Carolina at Chapel Hill, Chapel Hill, NC 27599, USA; Department of Biostatistics, University of North Carolina at Chapel Hill, Chapel Hill, NC 27599, USA.
E Shyong Tai, Saw Swee Hock School of Public Health, National University of Singapore and National University Health System, Singapore 117549, Singapore; Department of Medicine, Yong Loo Lin School of Medicine, National University of Singapore and National University Health System, Singapore 117597, Singapore.
Rob M van Dam, Saw Swee Hock School of Public Health, National University of Singapore and National University Health System, Singapore 117549, Singapore; Departments of Exercise and Nutrition Sciences and Epidemiology, Milken Institute School of Public Health, The George Washington University, Washington, DC 20052, USA.
Xueling Sim, Saw Swee Hock School of Public Health, National University of Singapore and National University Health System, Singapore 117549, Singapore.
Funding
The MEC1_T1 and MEC1_T2 studies are supported by individual research and clinical scientist award programs from the National Medical Research Council (NMRC) and the Biomedical Research Council (BMRC) of Singapore, and infrastructure funding from the Singapore Ministry of Health (Population Health Metric and Analytics [PHMA]), National University of Singapore and National University Health System, Singapore. This study made use of whole-genome data generated on MEC1 as part of the Singapore National Precision Medicine program funded by the Industry Alignment Fund (Pre-Positioning) (IAF-PP: H17/01/a0/007). S.C.R was supported by core funding from the: British Heart Foundation (RG/18/13/33946), the Munz Chair of Cardiovascular Prediction and Prevention and the NIHR Cambridge Biomedical Research Centre (NIHR203312), the Cambridge BHF Centre of Research Excellence (RE/18/1/34212), and the BHF Chair Award (CH/12/2/29428) and by Health Data Research UK, which is funded by the UK Medical Research Council, Engineering and Physical Sciences Research Council, Economic and Social Research Council, Department of Health and Social Care, Chief Scientist Office of the Scottish Government Health and Social Care Directorates, Health and Social Care Research and Development Division (Welsh Government), Public Health Agency (Northern Ireland), British Heart Foundation and the Wellcome Trust. The views expressed are those of the authors and not necessarily those of the NIHR or the Department of Health and Social Care.
Author Contributions
X.S., R.M.vD., and E.S.T. designed the study. Y.L., C.G.Y.L., S.C.R., N.B., J.C., J.Y., Yun.L., E.S.T., R.M.vD., and X.S. contributed to the analysis and interpretation of data. Y.L., X.S., E.S.T., and R.M.vD. wrote the draft of the manuscript. All authors contributed to critical revision of the manuscript.
Disclosures
The authors have nothing to disclose. The National University of Singapore has signed a collaboration agreement with SomaLogic to conduct SomaScan of MEC1 stored samples at no charge in exchange for the rights to analyze linked MEC1 phenotype data.
Data Availability
Summary statistics for all measured proteins are provided in the supplemental tables. Data from the Singapore Multi-Ethnic Cohort Phase 1 study can be requested by researchers for scientific purposes through an application process at the listed website (https://blog.nus.edu.sg/sphs/data-and-samples-request/). Data will be shared through an institutional data-sharing agreement.
References
- 1. Saeedi P, Petersohn I, Salpea P, et al. Global and regional diabetes prevalence estimates for 2019 and projections for 2030 and 2045: results from the International Diabetes Federation Diabetes Atlas, 9(th) edition. Diabetes Res Clin Pract. 2019;157:107843. [DOI] [PubMed] [Google Scholar]
- 2. Zheng Y, Ley SH, Hu FB. Global aetiology and epidemiology of type 2 diabetes mellitus and its complications. Nat Rev Endocrinol. 2018;14(2):88‐98. [DOI] [PubMed] [Google Scholar]
- 3. Chatterjee S, Khunti K, Davies MJ. Type 2 diabetes. Lancet. 2017;389(10085):2239‐2251. [DOI] [PubMed] [Google Scholar]
- 4. Pearson ER. Type 2 diabetes: a multifaceted disease. Diabetologia. 2019;62(7):1107‐1112. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. McCarthy MI. Painting a new picture of personalised medicine for diabetes. Diabetologia. 2017;60(5):793‐799. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Vujkovic M, Keaton JM, Lynch JA, et al. Discovery of 318 new risk loci for type 2 diabetes and related vascular outcomes among 1.4 million participants in a multi-ancestry meta-analysis. Nat Genet. 2020;52(7):680‐691. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Mahajan A, Taliun D, Thurner M, et al. Fine-mapping type 2 diabetes loci to single-variant resolution using high-density imputation and islet-specific epigenome maps. Nat Genet. 2018;50(11):1505‐1513. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Spracklen CN, Horikoshi M, Kim YJ, et al. Identification of type 2 diabetes loci in 433,540 east Asian individuals. Nature. 2020;582(7811):240‐245. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Mahajan A, Spracklen CN, Zhang W, et al. Multi-ancestry genetic study of type 2 diabetes highlights the power of diverse populations for discovery and translation. Nat Genet. 2022;54(5):560‐572. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Liu Y, Buil A, Collins BC, et al. Quantitative variability of 342 plasma proteins in a human twin population. Mol Syst Biol. 2015;11(1):786. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Suhre K, McCarthy MI, Schwenk JM. Genetics meets proteomics: perspectives for large population-based studies. Nat Rev Genet. 2021;22(1):19‐37. [DOI] [PubMed] [Google Scholar]
- 12. Gudmundsdottir V, Zaghlool SB, Emilsson V, et al. Circulating protein signatures and causal candidates for type 2 diabetes. Diabetes. 2020;69(8):1843‐1853. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Rooney MR, Chen J, Echouffo-Tcheugui JB, et al. Proteomic predictors of incident diabetes: results from the atherosclerosis risk in communities (ARIC) study. Diabetes Care. 2023;46(4):733‐741. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Zaghlool SB, Halama A, Stephan N, et al. Metabolic and proteomic signatures of type 2 diabetes subtypes in an Arab population. Nat Commun. 2022;13(1):7121. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Chen ZZ, Gao Y, Keyes MJ, et al. Protein markers of diabetes discovered in an African American cohort. Diabetes. 2023;72(4):532‐543. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Ngo D, Benson MD, Long JZ, et al. Proteomic profiling reveals biomarkers and pathways in type 2 diabetes risk. JCI Insight. 2021;6(5):e144392. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Davey Smith G, Hemani G. Mendelian randomization: genetic anchors for causal inference in epidemiological studies. Hum Mol Genet. 2014;23(R1):R89‐R98. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Yao P, Iona A, Pozarickij A, et al. Proteomic analyses in diverse populations improved risk prediction and identified new drug targets for type 2 diabetes. Diabetes Care. 2024;47(6):1012‐1019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Gou W, Yue L, Tang XY, et al. Circulating proteome and progression of type 2 diabetes. J Clin Endocrinol Metab. 2022;107(6):1616‐1625. [DOI] [PubMed] [Google Scholar]
- 20. Rhee EJ. Diabetes in asians. Endocrinol Metab (Seoul). 2015;30(3):263‐269. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Ma RC, Chan JC. Type 2 diabetes in east asians: similarities and differences with populations in Europe and the United States. Ann N Y Acad Sci. 2013;1281(1):64‐91. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Tan KHX, Tan LWL, Sim X, et al. Cohort profile: the Singapore multi-ethnic cohort (MEC) study. Int J Epidemiol. 2018;47(3):699‐699j. [DOI] [PubMed] [Google Scholar]
- 23. Liang Y, Lim CGY, Ritchie SC, et al. Data from: Circulating proteomic profiles are associated with the onset of type 2 diabetes in a multi-ethnic Asian population—a longitudinal study. figshare. 10.6084/m9.figshare.26787955.v2. Deposited 13 December 2024. [DOI]
- 24. Addendum. 2. Classification and diagnosis of diabetes: standards of medical care in diabetes-2021. Diabetes Care. 2021;44(Suppl 1):S15‐S33. [DOI] [PubMed] [Google Scholar]
- 25. Gold L, Ayers D, Bertino J, et al. Aptamer-based multiplexed proteomic technology for biomarker discovery. PLoS One. 2010;5(12):e15004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Friedman J, Hastie T, Tibshirani R. Regularization paths for generalized linear models via coordinate descent. J Stat Softw. 2010;33(1):1‐22. [PMC free article] [PubMed] [Google Scholar]
- 27. Raudvere U, Kolberg L, Kuzmin I, et al. G:Profiler: a web server for functional enrichment analysis and conversions of gene lists (2019 update). Nucleic Acids Res. 2019;47(W1):W191‐W198. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Jain A, Tuteja G. TissueEnrich: tissue-specific gene enrichment analysis. Bioinformatics. 2019;35(11):1966‐1967. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Wong E, Bertin N, Hebrard M, et al. The Singapore national precision medicine strategy. Nat Genet. 2023;55(2):178‐186. [DOI] [PubMed] [Google Scholar]
- 30. Kwong AM, Blackwell TW, LeFaive J, et al. Robust, flexible, and scalable tests for Hardy-Weinberg equilibrium across diverse ancestries. Genetics. 2021;218(1):iyab044. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Loh PR, Danecek P, Palamara PF, et al. Reference-based phasing using the haplotype reference consortium panel. Nat Genet. 2016;48(11):1443‐1448. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Burgess S, Thompson SG. Interpreting findings from Mendelian randomization using the MR-Egger method. Eur J Epidemiol. 2017;32(5):377‐389. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Hemani G, Zheng J, Elsworth B, et al. The MR-Base platform supports systematic causal inference across the human phenome. Elife. 2018:7:e34408. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Giambartolomei C, Vukcevic D, Schadt EE, et al. Bayesian test for colocalisation between pairs of genetic association studies using summary statistics. PLoS Genet. 2014;10(5):e1004383. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. Stumvoll M, Goldstein BJ, van Haeften TW. Type 2 diabetes: principles of pathogenesis and therapy. Lancet. 2005;365(9467):1333‐1346. [DOI] [PubMed] [Google Scholar]
- 36. Zhao JZY, Wei F, Song J, et al. Triglyceride is an independent predictor of type 2 diabetes among middle-aged and older adults: a prospective study with 8-year follow-ups in two cohorts. J Transl Med. 2019;7(1):403. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Zaidi D, Turner JK, Durst MA, Wagner GF. Stanniocalcin-1 co-localizes with insulin in the pancreatic islets. ISRN Endocrinol. 2012;2012:834359. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Jong MC, Voshol PJ, Muurling M, et al. Protection from obesity and insulin resistance in mice overexpressing human apolipoprotein C1. Diabetes. 2001;50(12):2779‐2785. [DOI] [PubMed] [Google Scholar]
- 39. Watanabe J, Takiyama Y, Honjyo J, et al. Role of IGFBP7 in diabetic nephropathy: TGF-β1 induces IGFBP7 via smad2/4 in human renal proximal tubular epithelial cells. PLoS One. 2016;11(3):e0150897. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40. Liu Y, Wu M, Ling J, et al. Serum IGFBP7 levels associate with insulin resistance and the risk of metabolic syndrome in a Chinese population. Sci Rep. 2015;5(1):10227. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. Gu HF, Gu T, Hilding A, et al. Evaluation of IGFBP-7 DNA methylation changes and serum protein variation in Swedish subjects with and without type 2 diabetes. Clin Epigenetics. 2013;5(1):20. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Evdokimova V, Tognon CE, Benatar T, et al. IGFBP7 binds to the IGF-1 receptor and blocks its activation by insulin-like growth factors. Sci Signal. 2012;5(255):ra92. [DOI] [PubMed] [Google Scholar]
- 43. Oniki K, Umemoto Y, Nagata R, et al. Glutathione S-transferase A1 polymorphism as a risk factor for smoking-related type 2 diabetes among Japanese. Toxicol Lett. 2008;178(3):143‐145. [DOI] [PubMed] [Google Scholar]
- 44. Namwanje M, Brown CW. Activins and inhibins: roles in development, physiology, and disease. Cold Spring Harb Perspect Biol. 2016;8(7):a021881. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45. Zanetti D, Stell L, Gustafsson S, et al. Plasma proteomic signatures of a direct measure of insulin sensitivity in two population cohorts. Diabetologia. 2023;66(9):1643‐1654. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46. Qian W, Zhao M, Wang R, Li H. Fibrinogen-like protein 1 (FGL1): the next immune checkpoint target. J Hematol Oncol. 2021;14(1):147. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47. Pietzner M, Wheeler E, Carrasco-Zanini J, et al. Synergistic insights into human health from aptamer- and antibody-based proteomic profiling. Nat Commun. 2021;12(1):6822. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48. Lundberg M, Eriksson A, Tran B, Assarsson E, Fredriksson S. Homogeneous antibody-based proximity extension assays provide sensitive and specific detection of low-abundant proteins in human blood. Nucleic Acids Res. 2011;39(15):e102. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49. Eldjarn GH, Ferkingstad E, Lund SH, et al. Large-scale plasma proteomics comparisons through genetics and disease associations. Nature. 2023;622(7982):348‐358. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Citations
- Liang Y, Lim CGY, Ritchie SC, et al. Data from: Circulating proteomic profiles are associated with the onset of type 2 diabetes in a multi-ethnic Asian population—a longitudinal study. figshare. 10.6084/m9.figshare.26787955.v2. Deposited 13 December 2024. [DOI]
Data Availability Statement
Summary statistics for all measured proteins are provided in the supplemental tables. Data from the Singapore Multi-Ethnic Cohort Phase 1 study can be requested by researchers for scientific purposes through an application process at the listed website (https://blog.nus.edu.sg/sphs/data-and-samples-request/). Data will be shared through an institutional data-sharing agreement.


