Skip to main content
Nature Communications logoLink to Nature Communications
. 2026 May 9;17:6271. doi: 10.1038/s41467-026-73017-z

Multi-omics integration predicts the incidence of 17 diseases in the UK Biobank

Jiawen Du 1,#, Muqing Zhou 2,#, Hanling Wang 3, Jianqiao Wang 1, Laura M Raffield 2, Ruihai Zhou 4, Yun Li 1,2,, Can Chen 1,5,6,7,, Quan Sun 8,9,10,
PMCID: PMC13376913  PMID: 42106324

Abstract

Multi-omics technologies, such as metabolomics and proteomics, offer deep molecular perspectives that could enhance risk prediction, but large-scale studies integrating both are scarce. Here we show the predictive values of these two omics across 17 incident diseases in 23,776 UK Biobank participants with complete baseline for 159 NMR-based metabolites and 2,923 Olink affinity-based proteins. We found that adding omics data significantly improved risk prediction for all 17 diseases compared to clinical predictors alone. Proteomics-only models generally outperformed metabolomics-only models for 16 of the 17 diseases, and integrating both omics added little prediction power over proteomics-only models. Furthermore, we identified key omics features, including both well-established (e.g., KLK3/PSA for prostate cancer) and potential novel ones (e.g., PRG3 for skin cancer). We further connected diseases with medication and socioeconomic factors through key proteins, highlighting the clinical utility of omics data for enhancing individual risk prediction, providing molecular insights into disease mechanisms, and potentially guiding future therapeutic development.

Subject terms: Predictive markers, Diseases, Epidemiology, Risk factors


This study utilizing data from the UK biobank suggests that plasma proteomics profiles significantly improve prediction of the incidence of 17 diseases, for example some types of cancer and cardiometabolic diseases, over standard clinical predictors and metabolomics profiles.

Introduction

Effective risk stratification is fundamental to the prevention, early detection, and management of various diseases1. Clinical risk assessment of disease outcome has primarily relied on established risk factors2, including demographic attributes, routine laboratory measurements, and sometimes lifestyle behaviors. These traditional predictors, while valuable, may not capture the underlying complex disease mechanisms. Recent advances in high-throughput omics technologies have significantly enhanced our ability to characterize biological processes at the molecular level, thereby holding the potential for improving risk assessment, disease prediction, and personalized interventions35.

Metabolomics has emerged as a robust and cost-effective method to quantify metabolites circulating in the blood. Previous studies have demonstrated the potential clinical utility of nuclear magnetic resonance (NMR)-based metabolomics, establishing their additive values beyond conventional predictors in predicting various clinical endpoints69. Similarly, proteomics, measuring proteins in various sample types including biofluids and tissues, has also been utilized to understand disease mechanisms and improve clinical risk predictions1013. The UK Biobank has played a key role in exploring such disease-omics relationships at a large scale9,10. However, studies to comprehensively investigate the prediction power integrating these two omics remain scarce. Leveraging these complementary omics profiles holds significant promise to enhance disease prediction in clinical practice, potentially leading to more accurate diagnoses, refined prognostic prediction, and ultimately more personalized treatment strategies14.

To fill in the gap, in this work we leverage 159 metabolites (NMR based Nightingale platform) and 2923 proteins (Affinity-based Olink 3k proteomics) measured in 23,776 individuals from the UK Biobank (UKB) to systematically evaluate whether incorporating metabolomics and proteomics data enhances prediction of 17 incident disease endpoints compared to traditional clinical predictors alone. We further investigate top contributing omics features and identify enriched demographic, medical and socioeconomic protein correlates for each disease, leveraging external results15. Our study not only develops better risk prediction models, but also provides insights into a wide range of disease risk factors.

Results

Study overview

The UKB metabolomics data9 was collected from the targeted Nightingale NMR based platform and the proteomics profiles13 was collected using Olink on blood plasma samples. We included 23,776 individuals as our study participants who had complete records of 159 metabolites and 2,923 proteins available (Supplementary Data 1-2), after quality control (QC; Methods). The mean age at enrollment was 56.9 years, ranging from 39 to 70, with 10,870 (45.7%) males.

One goal of this study is to investigate the value of metabolomics and proteomics in predicting incident diseases compared to clinical predictors only. Thus, we investigated 17 clinical endpoints, encompassing cancers, cardiovascular disorders, neurological disorders, systemic diseases, respiratory disorders, ophthalmic conditions, and musculoskeletal conditions, which were classified into three categories—cardiometabolic, cancers, and others (Supplementary Data 3; Methods). Observed incidence for the 17 investigated diseases varied in our cohort, ranging from 3.0% for dementia to 14.6% for cataracts (Table 1). For baseline clinical predictors, we considered three nested sets following a previous study7: (1) age+sex; (2) ASCVD, a set of cardiovascular-related predictors, and (3) PANEL, a more comprehensive set of clinical predictors including demographic, lifestyle, and selected laboratory measurements (Fig. 1a, Supplementary Data 4). Missing covariates were imputed within each recruitment center (Methods).

Table 1.

Characteristics of the study cohort

Variable Participants, No. (%) (n = 23,776)
Age, yr
Mean (SD) 56.9 (8.2)
Median (range) 58.0 (50.0- 64.0)
Sex
Female 12,906 (54.3%)
Male 10,870 (45.7%)
Population
European 22423 (94.3%)
African 630 (2.6%)
South Asian 419 (1.8%)
East Asian 108 (0.5%)
Other 196 (0.8%)
Body mass index
Mean (SD) 27.5 (4.8)
Median (range) 26.8 (24.2–29.9)
Missing 114 (0.5%)
Smoking status
Current smoker 2461 (10.4%)
Former smoker 8248 (34.7%)
Never smoke 12,965 (54.5%)
Missing 102 (0.4%)
MACE 3351 (14.09%)
Heart Failure 1186 (4.99%)
CHD 3134 (13.18%)
PAD 1220 (5.13%)
Atrial Fibrillation 2113 (8.89%)
T2 Diabetes 2330 (9.80%)
Prostate Cancer 737 (3.10%)
Skin Cancer 1294 (5.44%)
Breast Cancer 870 (3.66%)
Dementia 723 (3.04%)
Cataracts 3475 (14.61%)
Glaucoma 748 (3.14%)
COPD 1659 (6.98%)
Asthma 2441 (10.27%)
Renal Disease 2946 (12.39%)
Liver Disease 1628 (6.85%)
Fractures 2279 (9.59%)

MACE Major Adverse Cardiac Event, CHD Coronary Heart Disease, PAD Peripheral Artery Disease, T2 Diabetes Type 2 Diabetes, COPD Chronic Obstructive Pulmonary Disease.

Fig. 1. Study overview.

Fig. 1

a Overview of the three baseline predictor sets. Three nested baseline clinical predictor sets were considered in this study: (1) age+sex; (2) ASCVD, a set of commonly used cardiovascular-related predictors, and (3) PANEL, a more comprehensive set of clinical predictors including demographic, lifestyle, and selected laboratory measurements. Details of (2) and (3) are included in Supplementary Data 3. b Overview of the study design. This study included 23,776 UK Biobank individuals with complete records of 159 metabolites and 2923 proteins. Within this cohort, we first adjusted for one baseline clinical predictor set (X) and obtained martingale residuals using the Cox proportional hazards (PH) models for each endpoint (step 1). We performed five repeated five-fold stratified cross validation (CV). In each five-fold CV, we split the entire cohort into five folds separately for each trait to ensure an equal number of cases in each fold, and trained the model on four folds to generate predictions in the remaining fold using mixOmics. Models were evaluated by adding the predicted residuals back to the CoxPH model using Harrell’s C-index (step 2). We repeated this five-fold splitting process five times using different random seeds to assess variations in model performance (step 3). We also identified key omics features (i.e. metabolites and proteins) for each endpoint by estimating the feature importance scores (step 4). This figure was created in BioRender. Du, J. (2026) https://BioRender.com/ye0ab87.

To control for covariates in omics models, we fitted Cox proportional hazard (PH) models for each outcome adjusting for the recruitment center and covariates. We calculated martingale residuals (Fig. 1b, step 1), which represented the survival outcome after accounting for baseline predictor effects and were subsequently used as labels (prediction outcomes) for downstream modeling. Then we used mixOmics16 to perform multivariate model training with omics data for outcome (martingale residuals) prediction. We performed five repeated five-fold stratified cross-validation (CV; Methods). In each five-fold CV, we split the entire cohort into five folds, training the model on four folds and generating predictions in the remaining fold (Fig. 1b, step 2). To ensure the robustness of our results, we repeated this five-fold splitting process by five times using different random seeds (Fig. 1b, step 3). We evaluated the prediction performance of three omics models, namely metabolomics-only, proteomics-only, and combined-omics, by fitting the CoxPH models again with the predicted residuals added to the model (i.e. endpoint ~ baseline clinical predictor set + recruitment center + predicted residual). Finally, Harrell’s C-index was calculated and used as the evaluation metric (Fig. 1b, step 3). We also assessed important omics features contributing most to omics prediction models (Fig. 1b, step 4; Methods), separately for each endpoint and each baseline adjustment.

Enhanced prediction accuracy including omics profiles

We first evaluated whether integrating metabolomics and/or proteomics offered improved prediction power over baseline models based solely on one of the three clinical predictor sets. Across all the 17 traits, adding omics data (metabolomics, proteomics, or their combination) to any baseline consistently yielded statistically significant improvement in predictive performance (Fig. 2, Supplementary Data 5-6), with mean C-index increment of 0.063, 0.038 and 0.030 compared to the three baseline models, respectively (p-values < 1E-6). For example, compared to the baseline PANEL model, adding proteomics alone resulted in a C-index increment from 0.66 to 0.75 for prostate cancer (p-value < 1E-6), indicating the value of omics data beyond clinical risk factors.

Fig. 2. Absolute C-index for all 17 investigated endpoints.

Fig. 2

Each dot represents the mean C-index in the five repeated five-fold CV with error bars indicating one standard deviation (SD). For baseline models, C-index was calculated from CoxPH models only including baseline predictors. For omics models, C-index was calculated by adding the predicted martingale residuals back to the baseline CoxPH model. Source data are provided as a Source Data file.

Comparison between proteomics, metabolomics, and combined models

Since single-omics models are well-established in previous studies7,10,11, we first evaluated if integrating proteomics and metabolomics together would enhance prediction performance. In general, the combined models achieved 0.034 and 0.014 average C-index increment compared to baseline predictors and metabolomics-only models respectively, under PANEL adjustment, while showing slight inferiority (average C-index decrement = 0.002) compared to proteomics-only model. We also observed that the performance varied across traits and baseline predictors. In peripheral artery disease (PAD), chronic obstructive pulmonary disease (COPD), renal disease, and fractures, the combined models demonstrated consistently significant improvements across all three baseline predictor sets (Fig. 2). For a broader range of conditions, the combined models significantly outperform single-omics models for 6 additional traits (major adverse cardiac events [MACE], T2 diabetes, heart failure, coronary heart disease [CHD], liver disease, and glaucoma) in at least one baseline predictor set (Supplementary Fig. 1, Supplementary Data 6). For example, in CHD, the combined-omics models showed the best performance with age+sex (C-index difference ranging from 0.008 to 0.050, p-value < 1E-6) and ASCVD clinical predictors (C-index difference ranging from 0.001 to 0.014, p-value < 1E-6) but demonstrated comparable performance to proteomics-only models with the comprehensive PANEL clinical predictors set (C-index difference = 2.5E-4, p-value = 0.43).

Additionally, we observed a trend showing less pronounced advantages of combined-omics models when including more comprehensive clinical predictors in PANEL models. Specifically, proteomics-only models significantly outperformed or performed comparably (the second best with no statistically significant difference from the best) with combined-omics models for seven traits with age+sex, nine traits with ASCVD and ten traits with PANEL as baseline (Supplementary Fig. 1, Supplementary Data 6), although the absolute difference in C-index is minimal. For example, for the 7 out of 10 traits where proteomics significantly outperformed combined-omics models with PANEL predictors, combined-omics models showed an average decrease in C-index of 0.007 compared to proteomics-only models. We note that the PANEL predictor set included lipids measurements, which consist of a substantial proportion of metabolomics (Supplementary Fig. 2). Thus, the benefits of including a more comprehensive metabolite profile in the combined-omics models might be surpassed by the costs of higher degrees of freedom, resulting in less predictive performance for combined-omics models compared to proteomics-only models.

Consistent with the results that proteomics-only models performed similarly to combined models, proteomics profiles also showed superior predictive capability than metabolomics. Specifically, with PANEL baseline predictors, for 16 out of 17 traits (except for glaucoma), proteomics-only models achieved significantly higher C-index values than metabolomics-only models (mean C-index difference 0.019, Fig. 2). As an example, asthma consistently showed the greatest advantages of proteomics-only models over metabolomics-only models, with C-index differences of 0.081, 0.078 and 0.060 (p-values < 1E-6) across age+sex, ASCVD and PANEL clinical predictors, respectively. Note that the number of measured proteins (n = 2,923) is substantially larger than metabolites (n = 159), which likely explains the larger contribution of proteomics models to risk predictions. Caution is needed to generalize this conclusion which may be specific to the diseases under study, and to the specific proteomics and metabolomics platforms used.

On the other hand, we only observed one trait, glaucoma, for which including metabolomics alone (C-index = 0.716) showed significant while small improvements compared to either proteomics-only (C-index = 0.712, p-value < 1E-6) or combined-omics (C-index = 0.712, p-value < 1E-6) models with PANEL baseline predictors.

Known disease associations confirmed by top contributing omics features

Focusing on the combined-omics models, we then sought to identify key omics features contributing most to the predictions, especially those that remained significant even after accounting for the extensive PANEL predictors (Supplementary Data 7, Supplementary Fig. 3-22). Since feature importance scores were weighted by block weights (Methods), we observed in the combined models that proteomics was 25 times, 41 times, and 97 times more important than metabolomics under adjustment for age+sex, ASCVD and PANEL, respectively, indicating the dominant role of proteomics in model prediction as well as feature selection. We first highlight some notable examples with established metabolite/protein-disease relationships. For example, even with huge difference in block weight, glucose (a metabolite) stood out as the 3rd and 4th top feature for T2 diabetes under adjustment of age+sex and ASCVD respectively. Note that it disappeared in PANEL adjustment as the PANEL baseline included glucose (Supplementary Data 7). In addition, albumin and total fatty acids were also identified as key metabolites for T2 diabetes when only adjusting for age+sex. Except for T2 diabetes, the key omics features are all dominated by proteins for all the other traits. For example, we identified kallikrein-related peptidase 3/prostate-specific antigen (KLK3/PSA), a well-established prognostic marker, for prostate cancer17, which showed the strongest association among all the protein-disease pairs (z-score = 11.79, Supplementary Fig. 7). Additionally, we confirmed a key determinant protein, Crystallin Beta B2 (CRYBB2), for cataracts18,19 (z-score = 12.48, Supplementary Fig. 10); Hepatitis A Virus Cellular Receptor 1 (HAVCR1, z-score = 2.96) for kidney disease20 (Supplementary Fig. 16); as well as Apolipoprotein E (APOE, z-score = −5.92)21, Glial Fibrillary Acidic Protein (GFAP, z-score = 5.14)22, and VGF (z-score = 3.88)23 for dementia (Supplementary Fig. 19).

We also observed a consistent pattern of top-ranking proteins across cardiovascular traits (Fig. 3). Even after adjusting for PANEL baseline, several well-established biomarkers were found prominent across atrial fibrillation (AFib), heart failure, CHD and MACE. For example, brain natriuretic peptide (BNP) and N-terminal pro-brain natriuretic peptide (NT-proBNP), widely used as indicators for clinical diagnosis of cardiac dysfunction2427, showed strong positive associations with all the four cardiovascular conditions (Fig. 3, Supplementary Data 7). Moreover, a significant negative association with Protein S (PROS1) was also consistently observed in AFib and CHD, aligning with previous studies linking PROS1 deficiency to increased risks of CHD28,29.

Fig. 3. Key omics features for cardiometabolic diseases adjusting for PANEL baseline predictors.

Fig. 3

This heatmap shows identified key omics features associated with cardiometabolic disease (Major Adverse Cardiovascular Events [MACE], heart failure, Coronary Heart Disease [CHD], Peripheral Artery Disease [PAD], Atrial Fibrillation [AFib], and type 2 diabetes), after adjusting for the PANEL baseline clinical predictors. We selected the top 5 important features (which are all proteins here) for each disease and displayed the union of them. Color indicates the direction of association, with intensity showing the association strength. The normalized importance scores are displayed inside each grid. Source data are provided as a Source Data file.

At a more granular level, within these four cardiometabolic traits, we observed sub-patterns. MACE and CHD tend to have more similar proteomic signatures, sharing a significant positive association with lipoprotein(a)30, SERPIND131 and LTBP232. This sub-pattern was further confirmed when calculating correlations of feature importance scores for all omics features across traits (Fig. 4), where we observed that cardiovascular traits formed a cluster with high correlations among each other, particularly for MACE and CHD (correlation = 0.587). Likewise, heart failure and AFib shared ANGPT2 and BGLAP which were less significant in MACE or CHD. PAD showed the most distinct patterns among these conditions. While it shared the strong positive associations with NT-proBNP, PAD was also uniquely characterized by strong associations with BOC and CDON. Although these are not yet well-established biomarkers for PAD, emerging evidence suggests a plausible role for these proteins in the adaptive response to chronic limb ischemia, specifically reflecting the reactivation of the Hedgehog signaling pathway for angiogenesis33,34, highlighting potential avenues for further investigation. These observations were largely consistent using age+sex and ASCVD baseline sets (Supplementary Figs. 3-4, Supplementary Data 7).

Fig. 4. Correlations of feature importance scores for all proteins across 17 investigated endpoints.

Fig. 4

This correlation heatmap shows Pearson correlations of feature importance scores for all involved metabolites and proteins across each endpoint. There are some obvious clusters across traits. For example, cardiovascular traits (i.e., Major Adverse Cardiovascular Events [MACE], heart failure, Coronary Heart Disease [CHD], Peripheral Artery Disease [PAD], Atrial Fibrillation [AFib]) formed a cluster showing high correlations, with the strongest relationship observed between MACE and CHD (correlation = 0.587). Other notable sub-patterns include a significant correlation between the two ophthalmologic conditions, cataracts and glaucoma (correlation = 0.314), and between the two metabolic traits, liver disease and renal disease (correlation = 0.360). Source data are provided as a Source Data file.

Besides, some other well-established protein-disease associations were also revealed. For the respiratory diseases including COPD and asthma, well-known shared biomarkers were observed, with Secretoglobin Family 1 A Member 1 (SCGB1A1)35,36 identified as proteins associated with protective effects and peptidoglycan recognition protein 4 (PRR4)3739 with increased risk (Supplementary Data 7). Metabolic traits (i.e. renal disease and liver disease) also exhibited a shared key protein pattern that remained evident even after adjustment for PANEL predictors (Supplementary Fig. 16). Specifically, both of them demonstrated a significant association with fibroblast activation protein (FAP) and ITGAM, suggesting a convergent mechanism of metabolic organ damage, where ITGAM drives the sustained recruitment of inflammatory immune cells40,41, while FAP mediates the subsequent transition to pathological fibrosis and tissue remodeling4244.

Novel protein-disease relationships from top omics features

Besides well-established protein evidence, our analysis also uncovered several novel relationships in models adjusting for PANEL predictors. For example, we identified two related novel proteins for skin cancer, Proteoglycan 3 (PRG3) and gamma-glutamyltransferase 5 (GGT5). Further Mendelian Randomization (MR) analysis (Methods) revealed that there is a potential causal relationship between PRG3 protein level and skin cancer (inverse-variance weighted45 [IVW] OR = 0.91, 95% CI: 0.89–0.93, p-value = 1.04E-16), with increased protein expression leading to reduced cancer risk. Sensitivity analysis using MR-Egger46 to correct for potential horizontal pleiotropy also supported this relationship (MR-Egger intercept p-value = 0.07; MR-Egger OR = 0.94 (0.90-0.99), p-value = 0.019). We further validated this potential causal relationship using SomaScan pQTLs47 (Methods), where the weighted median estimator48 showed significant result (OR = 0.89, 95% CI: 0.81-0.98, p-value = 0.019) while the IVW estimator was close to significance with the same direction of effect (OR = 0.94, 95% CI: 0.87–1.02, p-value = 0.12).

Shared pathways between top metabolites and proteins

To investigate the unique and shared mechanisms captured by proteomics and metabolomics, we used the top features from the two single-omics models with PANEL adjusted (Supplementary Data 8-9) to conduct KEGG49 pathway analysis, separately for metabolites and proteins. We found a fundamental divergence of disease signatures between the two omics. Top metabolites mainly captured a shared systemic background of dysregulated lipid homeostasis, resulting in pathways exhibited substantial overlap across traits. For example, HIF-1 signaling and biosynthesis of unsaturated fatty acids are two consistent top pathways (Supplementary Data 10). This pattern is likely driven by the high correlation structure of the Nightingale metabolomics profiles (Supplementary Fig. 2). On the other hand, proteomics captured more specific mechanisms in the structural and signaling machinery of disease progression (Supplementary Data 11). For example, cholesterol metabolism is the strongest pathway related to MACE (p-value = 2.0E-4, FDR = 1.4E-3), integrin signaling and cell adhesion molecule (CAM) interaction are significantly related to renal disease (p-value = 1.3E-3, 1.4E-3, FDR = 0.04, 0.04), complement and coagulation cascades are highlighted for liver disease (p-value = 9.5E-4, FDR = 0.02), and cytokine-cytokine receptor interaction showed significance for COPD (p-value = 1.6E-3, FDR = 0.03). These results demonstrate that metabolomics provide an overall landscape of systemic energetic stress, while proteomics reflect the distinct pathophysiological mechanisms and tissue-specific remodeling unique to the etiology of individual diseases.

In addition to the unique patterns captured by each omics, we also found multiple shared pathways between metabolomics and proteomics: insulin resistance for T2 diabetes (p-value = 3.1E-3 and 0.037; FDR = 0.03 and 0.07 for metabolites and proteins, respectively), African trypanosomiasis for prostate cancer (p-value = 0.021 and 0.033; FDR = 0.03 and 0.07), thyroid hormone synthesis for breast cancer (p-value = 7.2E-4 and 0.026; FDR = 0.02 and 0.06) and asthma (p-value = 7.2E-4 and 7.1E-3; FDR = 0.04 and 0.11), AGE-RAGE signaling pathway in diabetic complications for COPD (p-value = 2.6E-3 and 3.2E-3; FDR = 0.03 and 0.04), and efferocytosis for renal disease (p-value = 1.1E-4 and 0.022; FDR = 6.5E-3 and 0.13). These shared signatures highlight biological hubs where metabolic and protein signaling are likely coupled. For instance, the identification of insulin resistance in both omics for T2 diabetes serves as a robust cross-validation. Similarly, the concordance on efferocytosis in renal disease points to a coordinated multi-molecular response to tissue injury, involving both the protein receptors required for phagocytosis and the lipid mediators generated during cell clearance. Further, the appearance of AGE-RAGE signaling in COPD underscores a unified axis of oxidative stress and inflammation visible across the two biological layers. Collectively, these shared pathways represent mechanistic drivers where functional metabolic shifts are in parallel with the dysregulation of cellular signaling machinery.

Enrichment of top proteins in medication usage

Leveraging our identified top contributing proteins (after PANEL adjusted), we queried results from an external study of UKB proteomics15 and found significant enrichment for proteins associated with particular medications for 7 of the 17 investigated traits, resulting in 18 distinct drug-disease associations where the drugs were shared across traits (Supplementary Data 12). We confirmed some expected drug-disease associations, including Bisoprolol (for all cardiovascular traits), the widely prescribed anticoagulant Warfarin (for heart failure, CHD, AFib), and the loop diuretic Furosemide for treating cardiovascular diseases (MACE, heart failure, and AFib) (Supplementary Data 12). Additionally, top proteins of COPD and asthma showed strong enrichment for common respiratory medications, including Salbutamol (OR = 18.12, 47.96; FDR = 9.08E-7, 2.33E-16) and the combination of Salmeterol and Fluticasone Propionate (OR = 17.21, 30.15; FDR = 1.25E-3, 9.80E-7), while asthma was also uniquely associated with Prednisolone (OR = 9.67; FDR = 3.50E-7). We noted that there are few current studies exploring the medication correlates of Olink-measured proteins in true external data (i.e. not in UKB itself). As such resources become available, it would be interesting to further contextualize the putative drivers of our top proteins contributing to disease prediction.

We also uncovered new evidence supporting some existing but less well-established associations. For example, we found associations between Warfarin and top proteins for fracture risk with ASCVD baseline model (OR = 7.84, FDR = 0.014), adding evidence to early studies suggesting the association between long-term Warfarin usage and increased osteoporotic fracture50,51. Future studies are needed to better understand this relationship.

Enrichment of top proteins in socioeconomics, lifestyle, and demographic factors

In addition to the enriched medications, we also identified several other factors associated with various diseases, especially for cardiometabolic traits which had the most associations with socioeconomic and lifestyle factors (including alcohol assumption, leisure/social activities, and smoking status) (Supplementary Data 13-14), consistent with previous findings that socioeconomic factors are a major determinant in cardiovascular disease risk52. Besides, a strong enrichment in average total household income and Townsend deprivation index were observed for COPD-associated proteins, which remained highly significant even after considering top proteins in the PANEL-adjusted model (OR = 23.52, 12.94 FDR = 2.52E-9, 6.09E-5, Supplementary Data 13). This molecular-level finding is consistent with epidemiological data, which indicates a higher prevalence of COPD among adults with lower income levels53,54. To further investigate the causality of this relationship and to provide a better explanation, we conducted an MR analysis (Methods), which revealed a significant negative association between total household income and COPD risk (OR = 0.51, 95% CI: 0.34-0.77, p-value = 0.001). This finding substantiates the causal impact of socioeconomic status on COPD, bridging observational epidemiology with genetic evidence. We also confirmed expected relationships between diseases and basic demographic factors, including a strong association between incident COPD-associated proteins and smoking status (current, ever or never smoker) (OR = 19.66, FDR = 7.32E-15, Supplementary Data 14), reinforcing the central role of smoking in COPD pathophysiology even after accounting for extensive clinical risk factors including smoking status, possibly due to the misreported smoking status or smoking intensity information (for example pack-years, which is not included in our PANEL predictor set) that is captured by top proteins.

Discussion

In this study, we evaluated the incremental predictive performance of metabolomics and proteomics profiles, individually and in combination, upon three traditional clinical predictor sets for 17 incident diseases in 23,776 UKB participants. Compared to previous studies using metabolomics7 or top 5 proteins for risk prediction11, we achieved further enhanced predictive power for most diseases by including comprehensive proteomics and demonstrated that proteomics generally contributed more than metabolomics under these specific platforms. We also found that including both omics may not necessarily achieve better prediction under the same training sample sizes, suggesting potential model over-fitting. To the best of our knowledge, our study is currently the largest to systematically evaluate contributions of both metabolomics and proteomics profiles in incident disease prediction.

Our results were based on a multi-omics integration method, mixOmics16, but we also performed sensitivity analyses using a suite of traditional and deep learning approaches, including penalized Cox regression, Elastic Net, Random Survival Forest and MOGONET55 (Methods). We observed that the penalized Cox model exhibited severe overfitting when adjusting PANEL covariates, resulting in performance deterioration below the baseline model (Supplementary Fig. 23). Further, Random Survival Forests were computationally intractable given our dataset of over 3,000 predictors, with model fitting time exceeding several weeks per fold. Elastic net and MOGONET produced similar overall performance, but both failed to distinguish different omics integration schemes as effectively as mixOmics (Supplementary Figs. 24-25), where all the omics models showed similar performance. Results from these comparative analyses highlight the specific advantages of mixOmics in our high-dimensional multi-omics setting, justifying our choice of this method.

We identified top prioritized features from single-omics models and conducted pathway analysis separately for metabolites and proteins. Our results showed that metabolites reflect a shared systemic background of lipids, whereas proteins capture circulating signatures of disease-specific structural remodeling. We also identified multiple shared pathways, indicating biological hubs where functional metabolic shifts and protein signaling are mechanistically coupled to drive disease progression. In the combined models with both omics, top contributing omics features were dominated by proteins, consistent with the observation that proteomics-only models performed similarly to combined models. Our identified top proteins reinforced discoveries from previous studies focusing on proteomics-only associations and predictions10,11. We also uncovered some novel protein-disease associations that were to our knowledge not reported before, including a potential causal relationship of PRG3 for skin cancer. Beyond the underlying important proteins, we further identified protein-associated medication, socioeconomic and demographic factors for each disease. Our results revealed known drug-disease treatment effects, for example, the association between Bisoprolol and cardiovascular diseases (Supplementary Data 12). We also identified some less documented drug-disease relationships, suggesting potential novel repurposing drug candidates or unexpected side effects, which warrant future closer investigations.

Despite providing important insights, our study also has some limitations, most of which are due to data availability issues and point to potential future directions. First, our definition of prevalent cases relied exclusively on ICD-10 codes from Hospital Episode Statistics (HES) data and did not incorporate primary care records. While HES data provides robust phenotyping for severe morbidity requiring hospitalization, it likely underestimates the overall disease burden by excluding milder cases managed solely in primary care settings. Consequently, our study cohorts may be biased toward individuals with higher disease severity. We noted that the primary care data from UKB was made available in the middle of our study, thus we were not able to incorporate this data into our analysis. Future investigations could combine hospital and primary care datasets to capture the full spectrum of disease phenotypes and further validate these multi-omics associations across broader patient populations. Second, we only included 159 metabolites measured on the NMR platform and 2923 proteins on the Olink platform. Future studies are warranted to investigate whether our conclusions can be generalized to omics data from other platforms, for example untargeted metabolomics platforms, and to other cohorts and biobanks with different recruitment strategies. Third, both metabolomic and proteomic data were derived from plasma samples. Although in large-scale risk prediction, plasma represents a practical, minimally invasive medium that captures the aggregated physiological state, they may not fully reflect tissue-specific pathophysiology (e.g., brain processes relevant to dementia or lung-specific processes relevant to COPD). Our analysis relies on the assumption that plasma biomarkers serve as effective proxies for prediction because local tissue pathology often induces systemic responses or results in the leakage of tissue-specific markers into circulation. Fourth, our omics profiles were collected at baseline recruitment, precluding capturing the information on disease progression. Another interesting question is to leverage longitudinal measurements, or omics trajectories, to predict disease incidence, which is not currently available. Fifth, we only included two omics types in this study. Future studies may integrate other omics, for example, genomics through polygenic risk scores or epigenomic markers. Sixth, we imputed the missing covariates assuming randomness within each recruitment center, which may not be true in reality. We might also have missed some cases using hospitalization records only, without incorporation of death or primary care records. Seventh, feature importance scores in our model represent the statistical contribution to prediction power, which can be influenced by the high correlation structure inherent in omics data. The top-ranked omics features identified in our study should be interpreted as key predictive markers or representatives of correlated biological modules, rather than isolated causal drivers. Future studies utilizing orthogonal validation or causal inference methods (e.g., MR) are needed to disentangle these correlation structures and/or to establish specific causality. Lastly, given the observation that proteomics models showed robust performance, we largely focused on top proteins for various diseases, without regard to whether top proteins were disease-shared or specific. A previous study discussed alternative strategies using disease-specific proteins (i.e., unique to each disease) to separate out proteins that have broad influences on a wide spectrum of traits11, which deserves further investigation of a large number of diseases at the same time. Future studies could also explore the interactive effects between metabolomics and proteomics, which is beyond the scope of this study.

In summary, we leveraged metabolomics and proteomics profiles in 23,776 UKB participants to investigate their predictive ability for 17 common incidence diseases. Our models improved risk prediction even accounting for the comprehensive PANEL clinical predictors, suggesting the values of omics profiles. Additionally, our pathway analysis from top contributing omics features highlighted that metabolomics and proteomics capture distinct aspects of disease pathology. Further, the identified key proteins and their enriched factors provide insights into potential novel disease mechanisms, repurposing drugs or unknown side effects, and a broad category of risk factors, holding promising clinical utility.

Methods

This research was conducted using the UKB resource (application number 25953), subject to a data transfer agreement. The UKB has obtained ethics approval from the Northwest Multi-Center Research Ethics Committee (approval number 11/NW/0382) and obtained written informed consent from all participants prior to the study.

Omics data and covariates imputation

The UKB metabolomics data, collected from the targeted Nightingale NMR based platform, were measured in 275,241 participants9. For metabolomics data, we performed QC for any metabolites flagged as QC+ in the released data from UKB by assessing the percentage below the lower limit of detection. We dropped the metabolites with such a percentage > 10%. For the remaining metabolites, we imputed missing data with the lowest observed value.

Proteomics profiling was conducted using Olink on blood plasma samples from 54,219 individuals13. For proteomics data, we dropped proteins with missing rate greater than 20%. For the remaining proteins, we scaled the values within each batch and imputed missing values to half of the minimum observed value. Only randomized baseline participants were selected, and proteome values were subsequently inverse-normalized across batches.

To handle missing covariate data, we assumed they were occurring at random within each recruitment center, then employed Multivariate Imputation by Chained Equations (MICE)56 with random forest. To preserve the unique data characteristics of each recruitment location, this imputation was performed separately within each assessment center. We implemented the MICE algorithm in Python using the scikit-learn library, with a Random Forest estimator to capture complex, non-linear relationships. For binary covariates, imputed values were rounded to maintain their original dichotomous scale.

Survival endpoints definition

The 17 investigated clinical endpoints were defined by ICD10 codes following a previous publication7. Participants diagnosed with the specific condition prior to the baseline assessment were excluded. We confirmed that each endpoint has an incidence rate exceeding 3% in the study cohort. Analysis for highly sex-differentiated endpoints (i.e., breast cancer and prostate cancer) was restricted to the respective populations. T2 diabetes status was excluded from baseline predictor sets for T2 diabetes endpoint.

Model fitting, evaluation, and feature importance assessment

We utilized the R package mixOmics v6.30.016 (the “block.spls” function) to perform a multivariate model training with omics data for outcome (martingale residuals) prediction. We specified a fixed number of latent components to be 10, following the established convention in genomic studies of using 10 principal components to avoid the potential data leakage associated with extensive data-dependent tuning. We then maintained the default settings for the remaining parameters, and performed five repeated five-fold stratified CV. In each five-fold CV, we split the entire cohort into five folds separately for each trait, ensuring an equal number of cases in each fold to prevent data imbalance, and trained the model on four folds to generate predictions in the remaining fold. (Fig. 1b, step 2). We then repeated this five-fold splitting process by five times using different random seeds (Fig. 1b, step 3). We bootstrapped one million times to obtain the p-values for the comparisons between model performance to account for the dependent structure of C-index across the five observations.

To identify the most influential omics features, we employed the “block.spls” function based on Partial Least Squares (PLS) in the mixOmics16 package to the combined-omics model. In this model, the relationship between the predictors (i.e. omics features) and outcome is captured through a set of latent components, with loadings Ljk measuring the contribution of omics feature j to latent component k. Within each omics block, we have j=1PLjk2=1, where P is the total number of features. Furthermore, the Average Variance Extracted (AVE) for latent component k (AVEk) quantifies its overall explanatory power. Therefore, we defined the importance score of each feature as a weighted average of its loadings across all latent components, using the corresponding AVE values as weights. Mathematically, the importance of feature j within each omics block can be expressed as k=1KLjk*AVEkk=1KAVEk where K is the total number of latent components. To make the feature importance comparable across blocks, we multiplied the numerator by P to put them into the same scale, and then multiplied by the block weight (SSomic=cory,Tomic2), where T is the embedding matrix. The direction of feature importance scores was determined by the direction of correlation between original feature values and the outcome. Therefore, the final feature importance score for feature j from the combined model can be expressed as Impj=signcory,X*SSomic*k=1KP*Ljk*AVEkk=1KAVEk. For single-omics models, we remove the term SSomic when selecting top features. We chose a fixed random seed 42 in sample split and treated the first fold as the test set with the remaining for training to derive feature importance scores (Fig. 1b, step d). To select the top features, we normalized the importance scores to z-scores, and the key omics features were identified by having absolute normalized importance scores greater than 1.96.

Mendelian randomization analysis

To investigate the potential causal relationship between PRG3 protein and the risk of skin cancer, we conducted a Mendelian randomization (MR) analysis with PRG3 protein level as exposure and skin cancer as outcome. GWAS summary statistics for PRG3 protein was collected from UKB13, and instrument variants were collected after performing linkage disequilibrium (LD) clumping with European TOP-LD57 reference panel using r2 > 0.001 as threshold. Summary statistics for the skin cancer were obtained from the FinnGen consortium58 using “Other malignant neoplasms of skin (=non-melanoma skin cancer), excluding all cancers (controls excluding all cancers)” category. After harmonizing the datasets on effect alleles and removing the SNPs that are significant in both exposure and outcome GWAS, the causal effect was primarily estimated using the random-effects inverse-variance weighted (IVW) method45. To assess the robustness of our results against potential violations of MR assumptions, particularly horizontal pleiotropy, we performed several sensitivity analyses using MR-Egger regression46 and weighted median method48. For external validation, we used SomaScan pQTL from deCODE47 as the instrument variables for PRG3 protein, and conducted MR analysis using the above-mentioned FinnGen GWAS for skin cancer as the outcome.

To investigate the causal relationship between total household income and COPD, we used GWAS summary statistics for household income from UKB participants59 and performed the same LD-clumping as mentioned earlier to obtain instrument variables. Summary statistics for COPD was also obtained from FinnGen. Analysis was conducted with the R package MendelianRandomization v0.10.0 in R v4.4.2.

KEGG pathway analysis

To investigate the biological mechanisms associated with the top features, we performed pathway enrichment analysis using the Kyoto Encyclopedia of Genes and Genomes (KEGG) database49. Top features were selected from corresponding single-omics models with PANEL adjusted. Proteins were mapped to NCBI Entrez IDs, while metabolites were mapped to their corresponding KEGG Compound IDs, with aggregate measures mapped to representative proxy compounds (i.e., linoleic acid for PUFA and Omega-6; DHA for Omega-3; cholesterol for cholesterol-related; triglyceride for triglyceride-related; palmitic acid for fatty acids, saturated fatty acids and MUFA) to capture pathway connectivity. Enrichment analysis was conducted using clusterProfiler60 v4.14.6 R package. Significance was defined as p-value < 0.05.

Factors enriched with top contributing proteins

To further explore the potential known drivers of variance in the identified disease-associated proteins, we queried enrichment results from an external study15 linking proteins with multiple categories of factors, including basic demographics, medication usage, and socioeconomic factors. For each disease, we selected key proteins with an absolute normalized importance score (z-score) > 1.96 in our models, corresponding to a two-tailed p-value < 0.05, from models with the comprehensive PANEL baseline adjustment. For our enrichment queries, significance was defined as false discovery rate (FDR) < 0.05 and the corresponding enrichment odds ratio (OR) > 1.

Sensitivity analysis

Besides mixOmics16, we also applied penalized Cox model, Elastic Net, and a deep learning method, MOGONET55 for sensitivity analysis. Both the penalized Cox model and Elastic Net were implemented using the glmnet v4.1.10 R package61. For penalized Cox model, we applied a Lasso penalty for features in both omics to enforce sparsity. For Elastic Net, we used α=0.5 to balance Lasso and Ridge regularization. The hyperparameter tuning for the regularization parameter (λ) in both models was performed using the Bayesian Information Criterion (BIC) to prevent overfitting. The models were evaluated using a 5-fold cross-validation scheme across five different random seeds to ensure robustness. For MOGONET, as its original architecture was designed exclusively for classification, we implemented specific structural modifications to enable regression on continuous labels (martingale residuals). Specifically, we only modified the following three places in the original model’s architecture: we replaced the classification-specific Cross-Entropy loss with the Mean Squared Error (MSE) loss. The terminal Softmax activation was also removed to allow for raw, continuous output values. Furthermore, the View Correlation Discovery Network (VCDN) module was modified to fuse view representations via concatenation rather than the original tensor cross-product. These modifications effectively convert MOGONET from a classifier into a multi-view regression model. After this modification, we applied this model under the exact same study setting. Note that this approach represents a direct translation of a classification framework rather than a specialized regression architecture, and that the modification is not for improvement but only for feasibility. Thus, the resulting predictive performance remains limited.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

Supplementary Figs. (1.8MB, pdf)
Supplementary Data 1-14 (820.6KB, xlsx)
Reporting Summary (2MB, pdf)
41467_2026_73017_MOESM4_ESM.pdf (115.2KB, pdf)

Description of Additional Supplementary Files

Source data

Source Data (237.9KB, xlsx)

Acknowledgements

This research has been conducted using the UKB Resource under Application Number 25953. We thank the UKB participants and research team for enabling this study. Y.L. discloses support for the research of this work from Funder R01AR083790, Funder U01HG011720, and Funder R01HL146500. All the other authors declare no relevant funding.

Author contributions

J.D., Y.L., C.C. and Q.S. designed the project. J.D., M.Z., H.W. and J.W. implemented models, conducted data analysis. L.M.R. and R.Z. advised in interpretation and discussion of the results. J.D. and Q.S. wrote the manuscript. All authors reviewed and approved the manuscript.

Peer review

Peer review information

Nature Communications thanks the anonymous reviewers for their contribution to the peer review of this work. [A peer review file is available.]

Data availability

The UKB data are available for health-related research that is in the public interest with application and approval required. For details on how to apply to use UKB data, please see https://www.ukbiobank.ac.uk/use-our-data/apply-for-access/Source data are provided with this paper.

Code availability

Scripts that are developed and used in this study have been made publicly available on https://github.com/Dolce-99/ukb-multiomics62.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Jiawen Du, Muqing Zhou.

These authors jointly supervised this work: Yun Li, Can Chen, Quan Sun.

Contributor Information

Yun Li, Email: yun_li@med.unc.edu.

Can Chen, Email: canc@unc.edu.

Quan Sun, Email: sunq@chop.edu.

Supplementary information

The online version contains supplementary material available at 10.1038/s41467-026-73017-z.

References

  • 1.WHO CVD, R. isk & Chart, W. orking Group. World Health Organization cardiovascular disease risk charts: revised models to estimate risk in 21 global regions. Lancet Glob. Health7, e1332–e1345 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Goff, D. C. et al. 2013 ACC/AHA guideline on the assessment of cardiovascular risk: a report of the American College of Cardiology/American Heart Association Task Force on Practice Guidelines. Circulation129, S49–S73 (2014). [DOI] [PubMed] [Google Scholar]
  • 3.Chen, C. et al. Applications of multi-omics analysis in human diseases. MedComm4, e315 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Garg, M. et al. Disease prediction with multi-omics and biomarkers empowers case-control genetic discoveries in the UK Biobank. Nat. Genet56, 1821–1831 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Collins, F. S. & Varmus, H. A new initiative on precision medicine. N. Engl. J. Med.372, 793–795 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Galal, A., Talal, M. & Moustafa, A. Applications of machine learning in metabolomics: disease modeling and classification. Front. Genet.13, 1017340 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Buergel, T. et al. Metabolomic profiles predict individual multidisease outcomes. Nat. Med.28, 2309–2320 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Würtz, P. et al. Metabolite profiling and cardiovascular event risk: a prospective study of 3 population-based cohorts. Circulation131, 774–785 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Julkunen, H. et al. Atlas of plasma NMR biomarkers for health and disease in 118,461 individuals from the UK Biobank. Nat. Commun.14, 604 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Deng, Y.-T. et al. Atlas of the plasma proteome in health and disease in 53,026 adults. Cell188, 253–271.e7 (2025). [DOI] [PubMed] [Google Scholar]
  • 11.Carrasco-Zanini, J. et al. Proteomic signatures improve risk prediction for common and rare diseases. Nat. Med.30, 2489–2498 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Suhre, K. et al. Connecting genetic risk to disease end points through the human blood plasma proteome. Nat. Commun.8, 14357 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Sun, B. B. et al. Plasma proteomic associations with genetics and health in the UK Biobank. Nature622, 329–338 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Hasin, Y., Seldin, M. & Lusis, A. Multi-omics approaches to disease. Genome Biol.18, 83 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Pietzner, M. et al. Machine learning-guided deconvolution of plasma protein levels. Mol Syst Biol.21, 1822–1844 (2025). [DOI] [PMC free article] [PubMed]
  • 16.Rohart, F., Gautier, B., Singh, A. & Lê Cao, K.-A. mixOmics: An R package for’omics feature selection and multiple data integration. PLoS Comput. Biol.13, e1005752 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Stamey, T. A. et al. Prostate-specific antigen as a serum marker for adenocarcinoma of the prostate. N. Engl. J. Med.317, 909–916 (1987). [DOI] [PubMed] [Google Scholar]
  • 18.Yao, K. et al. Characterization of a novel mutation in the CRYBB2 gene associated with autosomal dominant congenital posterior subcapsular cataract in a Chinese family. Mol. Vis.17, 144–152 (2011). [PMC free article] [PubMed] [Google Scholar]
  • 19.Zhou, Y. et al. A novel CRYBB2 stopgain mutation causing congenital autosomal dominant cataract in a Chinese family. J. Ophthalmol.2016, 4353957 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Han, W. K., Bailly, V., Abichandani, R., Thadhani, R. & Bonventre, J. V. Kidney injury molecule-1 (KIM-1): a novel biomarker for human renal proximal tubule injury. Kidney Int.62, 237–244 (2002). [DOI] [PubMed] [Google Scholar]
  • 21.Raulin, A.-C. et al. ApoE in Alzheimer’s disease: pathophysiology and therapeutic strategies. Mol. Neurodegener.17, 72 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Kim, K. Y., Shin, K. Y. & Chang, K.-A. GFAP as a potential biomarker for Alzheimer’s disease: a systematic review and meta-analysis. Cells12, 1309 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Quinn, J. P., Kandigian, S. E., Trombetta, B. A., Arnold, S. E. & Carlyle, B. C. VGF as a biomarker and therapeutic target in neurodegenerative and psychiatric diseases. Brain Commun.3, fcab261 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Hijazi, Z., Oldgren, J., Siegbahn, A., Granger, C. B. & Wallentin, L. Biomarkers in atrial fibrillation: a clinical review. Eur. Heart J.34, 1475–1480 (2013). [DOI] [PubMed] [Google Scholar]
  • 25.Maalouf, R. & Bailey, S. A review on B-type natriuretic peptide monitoring: assays and biosensors. Heart Fail Rev.21, 567–578 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Cao, Z., Jia, Y. & Zhu, B. BNP and NT-proBNP as diagnostic biomarkers for cardiac dysfunction in both clinical and forensic medicine. Int J. Mol. Sci.20, 1820 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Yuan, S. et al. Cross-population GWAS and proteomics improve risk prediction and reveal mechanisms in atrial fibrillation. Nat. Commun.16, 6426 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.ten Kate, M. K. & van der Meer, J. Protein S deficiency: a clinical perspective. Haemophilia14, 1222–1228 (2008). [DOI] [PubMed] [Google Scholar]
  • 29.Ken-Dror, G., Cooper, J. A., Humphries, S. E., Drenos, F. & Ireland, H. A. Free protein S level as a risk factor for coronary heart disease and stroke in a prospective cohort study of healthy United Kingdom men. Am. J. Epidemiol.174, 958–968 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.O’Donoghue, M. L. et al. Lipoprotein(a), PCSK9 Inhibition, and cardiovascular risk. Circulation139, 1483–1492 (2019). [DOI] [PubMed] [Google Scholar]
  • 31.Rau, J. C., Beaulieu, L. M., Huntington, J. A. & Church, F. C. Serpins in thrombosis, hemostasis and fibrinolysis. J. Thrombosis Haemost.5, 102–115 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Ho, F. K. et al. A proteomics-based approach for prediction of different cardiovascular diseases and dementia. Circulation151, 277–287 (2025). [DOI] [PubMed] [Google Scholar]
  • 33.Lavine, K. J., Kovacs, A. & Ornitz, D. M. Hedgehog signaling is critical for maintenance of the adult coronary vasculature in mice. J. Clin. Invest.118, 2404–2414 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Thomas, N. A., Koudijs, M., van Eeden, F. J. M., Joyner, A. L. & Yelon, D. Hedgehog signaling plays a cell-autonomous role in maximizing cardiac developmental potential. Development135, 3789–3799 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Zhu, L. et al. The club cell marker SCGB1A1 downstream of FOXA2 is reduced in asthma. Am. J. Respir. Cell Mol. Biol.60, 695–704 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Chen, X. et al. SCGB1A1 as a key regulator of splenic immune dysfunction in COPD: insights from a murine model. Int J. Chron. Obstruct Pulmon Dis.20, 497–509 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Zhong, X. et al. Machine learning-based screening of asthma biomarkers and related immune infiltration. Front. Allergy6, 1506608 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Gong, R., Wang, Z., Tan, G. & Huang, Y. Bioinformatics analysis revealed underlying molecular mechanisms associated with asthma severity and identified GABAergic related pathway as a potential therapy for Th2-high endotype asthma. Heliyon10, e28401 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Ho, S.-C. et al. Proteoglycan 4 is a diagnostic biomarker for COPD. COPD 1999 (2015) 10.2147/COPD.S90926. [DOI] [PMC free article] [PubMed]
  • 40.Nakashima, H. et al. Activation of CD11b+ Kupffer cells/macrophages as a common cause for exacerbation of TNF/Fas-ligand-dependent hepatitis in hypercholesterolemic mice. PLoS ONE8, e49339 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Khan, S. Q., Khan, I. & Gupta, V. CD11b activity modulates pathogenesis of lupus nephritis. Front. Med.5, 52 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Basalova, N., Alexandrushkina, N., Grigorieva, O., Kulebyakina, M. & Efimenko, A. Fibroblasta (FAPα) in fibrosis: beyond a perspective marker for activated stromal cells? Biomolecules13, 1718 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Park, J. E. et al. Fibroblast activation protein, a dual specificity serine protease expressed in reactive human tumor stromal fibroblasts. J. Biol. Chem.274, 36505–36512 (1999). [DOI] [PubMed] [Google Scholar]
  • 44.Wang, X. M., Yu, D. M. T., McCaughan, G. W. & Gorrell, M. D. Fibroblast activation protein increases apoptosis, cell adhesion, and migration by the LX-2 human stellate cell line. Hepatology42, 935–945 (2005). [DOI] [PubMed] [Google Scholar]
  • 45.Bowden, J., Davey Smith, G. & Burgess, S. Mendelian randomization with invalid instruments: effect estimation and bias detection through Egger regression. Int. J. Epidemiol.44, 512–525 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Minelli, C. et al. The use of two-sample methods for Mendelian randomization analyses on single large datasets. Int. J. Epidemiol.50, 1651–1659 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Ferkingstad, E. et al. Large-scale integration of the plasma proteome with genetics and disease. Nat. Genet53, 1712–1721 (2021). [DOI] [PubMed] [Google Scholar]
  • 48.Bowden, J., Davey Smith, G., Haycock, P. C. & Burgess, S. Consistent estimation in mendelian randomization with some invalid instruments using a weighted median estimator. Genet Epidemiol.40, 304–314 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Kanehisa, M. & Goto, S. KEGG: Kyoto encyclopedia of genes and genomes. Nucleic Acids Res.28, 27–30 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Gage, B. F., Birman-Deych, E., Radford, M. J., Nilasena, D. S. & Binder, E. F. Risk of osteoporotic fracture in elderly patients taking warfarin: results from the National Registry of Atrial Fibrillation 2. Arch. Intern Med. 166, 241–246 (2006). [DOI] [PubMed] [Google Scholar]
  • 51.Peng, S. et al. Effect of fracture risk in inhaled corticosteroids in patients with chronic obstructive pulmonary disease: a systematic review and meta-analysis. BMC Pulm. Med.23, 304 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Argentieri, M. A. et al. Integrating the environmental and genetic architectures of aging and mortality. Nat. Med.31, 1016–1025 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Weeks, J. & Elgaddal, N. Chronic Obstructive Pulmonary Disease Among Adults Aged 18 Years and Older: United States, 2023. https://stacks.cdc.gov/view/cdc/174596, 10.15620/cdc/174596 (2025). [DOI] [PMC free article] [PubMed]
  • 54.Shohaimi, S. et al. Area deprivation predicts lung function independently of education and social class. Eur. Respir. J.24, 157–161 (2004). [DOI] [PubMed] [Google Scholar]
  • 55.Wang, T. et al. MOGONET integrates multi-omics data using graph convolutional networks allowing patient classification and biomarker identification. Nat. Commun.12, 3445 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Buuren, S. V. & Groothuis-Oudshoorn, K. Mice: multivariate imputation by chained equations in R. J. Stat. Soft. 45, 207–214 (2011).
  • 57.Huang, L. et al. TOP-LD: a tool to explore linkage disequilibrium with TOPMed whole-genome sequence data. Am. J. Hum. Genet.109, 1175–1181 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Kurki, M. I. et al. FinnGen provides genetic insights from a well-phenotyped isolated population. Nature613, 508–518 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Hill, W. D. et al. Genome-wide analysis identifies molecular systems and 149 genetic loci associated with income. Nat. Commun.10, 5741 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Yu, G., Wang, L.-G., Han, Y. & He, Q.-Y. clusterProfiler: an R package for comparing biological themes among gene clusters. OMICS A J. Integr. Biol.16, 284–287 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Friedman, J. et al. GLMNET: lasso and elastic-net regularized generalized linear models. 4.1-10 10.32614/CRAN.package.glmnet (2008).
  • 62.Du, J. Multi-omics integration predicts the incidence of 17 diseases in the UK Biobank. 10.5281/zenodo.19422825. [DOI] [PMC free article] [PubMed]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Figs. (1.8MB, pdf)
Supplementary Data 1-14 (820.6KB, xlsx)
Reporting Summary (2MB, pdf)
41467_2026_73017_MOESM4_ESM.pdf (115.2KB, pdf)

Description of Additional Supplementary Files

Source Data (237.9KB, xlsx)

Data Availability Statement

The UKB data are available for health-related research that is in the public interest with application and approval required. For details on how to apply to use UKB data, please see https://www.ukbiobank.ac.uk/use-our-data/apply-for-access/Source data are provided with this paper.

Scripts that are developed and used in this study have been made publicly available on https://github.com/Dolce-99/ukb-multiomics62.


Articles from Nature Communications are provided here courtesy of Nature Publishing Group

RESOURCES