Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Aug 29;40(17):e72201. doi: 10.1096/fj.202600647R

Explainable Plasma Proteomics–Based Machine Learning for Osteoporosis Diagnosis, Prognosis, and Protein Biomarker Discovery in the UK Biobank

Wenxiang Zhang 1,2,3, Wenjing Zhang 1,2,3, Hanwen Cheng 3,4, Weijie Gong 5, Yuhui Kou 3,4,✉, Baoguo Jiang 1,3,4,✉
PMCID: PMC13525639  PMID: 42667149

ABSTRACT

Osteoporosis (OP) is often underdiagnosed, highlighting the need for tools that can both detect existing disease and predict future risk; large‐scale plasma proteomics combined with explainable machine learning enables integrated diagnostic and prognostic modeling while prioritizing clinically relevant protein markers. This study aims to develop and validate an explainable plasma proteomics machine‐learning framework for osteoporosis diagnosis, future risk prediction, and biomarker discovery. We further tested whether a combined marker panel could distinguish normal, prevalent OP, and future incident OP states from baseline samples. Using UK Biobank plasma proteomic data, we established SPX‐OP, which separately models prevalent OP and incident OP based on Extreme Gradient Boosting (XGBoost) and SHapley Additive exPlanations (SHAP), and then evaluates whether the union of diagnostic and prognostic markers supports integrated baseline stratification. In the experiments, both the diagnostic and prognostic XGBoost models showed robust discrimination for osteoporosis status and future risk, respectively. SHAP‐derived protein markers, including FSHB, ADIPOQ, SOST, COL9A1, and CHAD, were linked to osteoporosis and enriched in bone‐related pathways involving bone development and remodeling, extracellular matrix organization, and inflammatory processes. Using only these SHAP‐selected protein markers, the XGBoost model outperformed the full‐proteome models and provided robust, simultaneous diagnostic and prognostic prediction of osteoporosis. In summary, this work transforms high‐dimensional proteomic data into interpretable marker sets, paving the way for improved risk stratification and further validation of plasma protein biomarkers in osteoporosis.

Keywords: clinically relevant protein markers, diagnosis and prognosis of osteoporosis, plasma proteomic


SPX‐OP is an explainable machine‐learning framework integrating large‐scale UK Biobank plasma proteomics for osteoporosis diagnosis, prognosis, and marker discovery. Olink proteomic data undergo quality control and differential analysis, followed by XGBoost modeling with five‐fold cross‐validation. SHAP interpretation identifies diagnostic and prognostic protein markers, enabling biologically informed risk stratification and integrated classification of normal, prevalent osteoporosis, and incident osteoporosis.

graphic file with name FSB2-40-e72201-g005.webp

1. Introduction

Osteoporosis is a chronic skeletal disorder characterized by reduced bone strength and an increased risk of fragility fractures, particularly in older adults [1, 2]. Osteoporotic fractures are associated with substantial morbidity, excess mortality, loss of independence, and significant healthcare and societal costs, yet the disease itself is often clinically silent until a fracture occurs [3, 4, 5]. As a result, many individuals are first diagnosed only after sustaining a low‐trauma fracture, when irreversible bone loss has already accumulated, and opportunities for early intervention have been missed [6]. These challenges highlight a central clinical need: Approaches that can both identify individuals who already have osteoporosis and, at the same time, reliably stratify those at high future risk.

Current clinical diagnosis of osteoporosis mainly relies on dual‐energy X‐ray absorptiometry (DXA) to measure bone mineral density (BMD) [7]. Clinical risk assessment tools such as FRAX further combine BMD with age and other risk factors to estimate 10‐year fracture probability and support decisions on pharmacologic prevention [8]. However, many fragility fractures occur in people whose BMD is only osteopenic or even within the normal range, indicating that BMD‐based criteria alone do not fully capture skeletal fragility and can miss high‐risk patients, contributing to persistent underdiagnosis and undertreatment of osteoporosis [9, 10]. These limitations have stimulated growing interest in complementary imaging indices and biomarkers that refine risk stratification beyond BMD and better align diagnostic tools with actual fracture risk in routine practice.

Proteins are key functional effectors of cellular processes and pathways [11, 12, 13], and proteins have emerged as informative biomarkers that capture both local tissue pathology and systemic pathophysiological states across a wide range of diseases [14, 15]. Recent large‐scale plasma proteomic studies have demonstrated that high‐throughput profiling of thousands of proteins can reveal disease‐associated signatures, elucidate underlying biological mechanisms, and improve risk prediction in cardiovascular, neurodegenerative, and malignant conditions [11, 16, 17]. Building on this progress, the UK Biobank Pharma Proteomics Project has generated standardized Olink Explore measurements for more than 50 000 participants, providing one of the largest population‐scale resources for plasma proteomics and an attractive platform for systematic discovery of protein markers of osteoporosis [18, 19, 20, 21, 22]. Machine‐learning methods are well suited to analyzing high‐dimensional plasma proteomic data because they can capture non‐linear relationships and complex feature interactions [23, 24, 25, 26, 27, 28]. However, such models are often regarded as “black boxes”, which limits biological interpretability and can hinder clinical translation. Shapley additive explanations (SHAP) [12, 29] provide a principled, game‐theoretic framework to decompose model predictions into additive feature contributions, thereby quantifying the influence of each protein on the predicted risk at both the individual and cohort level. Although interpretable machine‐learning strategies such as XGBoost combined with SHAP have been applied to multiple diseases, most prior studies have focused on either single‐task risk prediction or disease‐specific feature prioritization. In contrast, our study was designed around three osteoporosis‐oriented objectives: (i) separate modeling of current disease status and future disease risk using the same population‐scale plasma proteomic resource; (ii) derivation of compact diagnostic and prognostic marker panels from full‐proteome models; and (iii) evaluation of whether the union of these markers can support integrated baseline stratification across normal, prevalent OP, and future incident OP states. Thus, the main contribution of SPX‐OP lies not in introducing a new generic algorithm, but in establishing an explainable proteomics workflow tailored to osteoporosis diagnosis, prognosis, and marker discovery.

In this study, we used UK Biobank plasma proteomic and clinical data to develop SPX‐OP for osteoporosis diagnosis, future risk prediction, and marker discovery. We first built separate models for prevalent and incident OP, and then used SHAP to identify markers from protein features and clinical features. We further examined whether combining these diagnostic and prognostic markers could support baseline stratification of individuals into normal, prevalent OP, and future incident OP groups. Overall, our aim was to move from high‐dimensional proteomic data toward smaller, more interpretable marker sets with potential translational value for osteoporosis risk assessment.

2. Materials and Methods

2.1. The Framework of SPX‐OP

SPX‐OP is an explainable plasma proteomics–based machine‐learning framework for integrated osteoporosis diagnosis, prognosis and marker discovery (Figure 1). It links large‐scale UK Biobank plasma proteomic data to task‐specific prediction models and interpretable protein markers, with the overall aim of moving from high‐dimensional measurements to clinically meaningful, biologically grounded markers and risk scores.

FIGURE 1.

FIGURE 1

Framework of SPX‐OP.

First, we obtained Olink Explore plasma proteomic profiles and osteoporosis outcomes from UK Biobank, applied quality control, and restricted analyses to individuals aged > 50 years. This focus on midlife and older adults reflects the typical onset of primary osteoporosis and reduces heterogeneity from very low‐risk younger participants. We then constructed age‐stratified diagnostic, prognostic, and integrated datasets, using 5‐year age bands and controlled OP: Normal ratios within each band. This design mitigates severe class imbalance, aligns age distributions between cases and controls, and limits the influence of potentially noisy excess controls.

Second, to identify proteins associated with current disease and future risk, we performed two differential expression analyses: Normal vs. prevalent OP (diagnostic comparison) and Normal vs. incident OP (prognostic comparison). This step reduces dimensionality and enriches for proteins showing clear case–control contrasts, thereby providing focused feature sets for downstream models and an initial pool of candidate markers.

Finally, we trained XGBoost classifiers on these proteomic features for three prediction settings: Binary diagnosis (prevalent OP vs. Normal), binary prognosis (incident OP vs. Normal), and integrated three‐class classification (Normal, prevalent OP, incident OP). Besides, we applied SHAP to the fitted XGBoost models to quantify the contribution of each protein to model outputs. Proteins were ranked by mean absolute SHAP value to define diagnostic and prognostic marker sets, which were then used to retrain reduced “marker‐only” models and to build the integrated three‐class classifier. Gene Ontology enrichment analysis was applied to the SHAP‐selected markers to place them in a bone‐related biological context.

2.2. Ethics Statement

This research has been conducted using the UK Biobank Resource under application number (773 436). The UK Biobank study received ethical approval from the National Health Service North‐West Center Research Ethics Committee (reference: NW/0382). All participants provided written informed consent at recruitment, including consent for long‐term follow‐up and linkage to health‐related records. The present study involved secondary analysis of de‐identified UK Biobank data only; therefore, no additional recruitment or direct contact with participants was required.

2.3. Plasma Proteomic Measurement

Plasma proteomic profiles were obtained from the UK Biobank Research Analysis Platform (UKB‐RAP, https://ukbiobank.dnanexus.com) [30, 31]. At the baseline assessment (2006–2010), venous blood was collected into EDTA vacutainers, processed according to UK Biobank standard operating procedures, and plasma aliquots were stored at −80°C [11, 32]. Proteins were quantified at Olink Analytical Services using the Olink Explore Proximity Extension Assay in combination with next‐generation sequencing across four panels (cardiometabolic, inflammation, neurology, and oncology) [33]. The concentrations of 2 923 unique proteins were quantified, and their full names are listed in Table S1. Protein abundances were provided as Normalized Protein eXpression (NPX) values. We quantified missingness in the proteomic matrix before imputation, including the overall proportion of missing values and the distributions of missingness across proteins and samples (Figure 2). Because directly discarding partially missing NPX values would reduce sample size and potentially bias estimates, we addressed missingness using multivariate imputation. To reduce the potential impact of excessive missingness, proteins with a missing rate greater than 20% were excluded before imputation. For the remaining proteins, we applied a k‐nearest neighbors–based imputer implemented in Python (KNNImputer, scikit‐learn), which replaces missing entries by a distance‐weighted average of the k nearest neighbors in the joint proteomic space. KNN‐based imputation was chosen because it preserves local similarity structure across samples and has been widely used and methodologically evaluated in proteomics and other omics datasets [34, 35].

FIGURE 2.

FIGURE 2

Overview of missingness in the plasma proteomics matrix. (A) Histogram showing the distribution of missing‐value percentages across proteins. (B) Histogram showing the distribution of missing‐value percentages across samples. (C) Bar plot of the 20 proteins with the highest missingness rates. (D) Bar plot of the 20 samples with the highest missingness rates.

2.4. Data Collection and Processing

We analyzed plasma proteomic data from the UK Biobank (UKB) with a focus on osteoporosis (OP) [31, 36]. The source cohort comprised three outcome groups with marked class imbalance: Participants without OP (Normal, n = 41 312), with prevalent OP at baseline (n = 1 099), and with incident OP during follow‐up (n = 1 750). OP outcomes were ascertained using UK Biobank first‐occurrence records mapped to the International Classification of Diseases, 10th Revision (ICD‐10) codes M80, M81, and M82 [37]. Specifically, prevalent OP was defined as a first recorded OP diagnosis on or before the baseline assessment date, whereas incident OP was defined as a first recorded OP diagnosis after baseline. Participants without any recorded OP diagnosis (M80/M81/M82) during the available follow‐up were classified as Normal. For incident OP cases, follow‐up time was defined as the interval between baseline assessment and the first recorded OP diagnosis. For participants without recorded OP, follow‐up time was defined as the interval between baseline assessment and the last available follow‐up record. As shown in Figure 3, Normal participants vastly outnumber OP cases, a pattern that can bias machine‐learning models toward the majority class, inflate overall accuracy while masking poor sensitivity for OP, and distort feature selection. Therefore, the raw cohort was not used directly for modeling; instead, we further processed the data to rebalance class proportions and reduce confounding.

FIGURE 3.

FIGURE 3

Age distribution of UK Biobank participants across osteoporosis phenotypes. (A) Frequency histogram of baseline age for participants without osteoporosis. (B) Frequency histogram of baseline age for participants who developed incident osteoporosis during follow‐up. (C) Frequency histogram of baseline age for participants with prevalent osteoporosis at baseline. Bars represent the number of individuals within each 5‐year age band (counts shown above bars).

Because primary, age‐related OP typically occurs after midlife, all analyses were restricted to individuals aged > 50 years [38]. This focuses on postmenopausal and age‐related OP, and avoids dilution of signal by very low‐risk younger participants. Consistently, Figure 3 shows that only 6.3% of incident OP and 5.7% of prevalent OP cases were < 50 years, compared with 23.9% of non‐OP participants, further supporting this age restriction. Given that age is one of the strongest determinants of OP [35], we further controlled for age by constructing age‐stratified datasets using 5‐year age bands (e.g., 50–55, 55–60, 60–65, etc.). We chose this granularity because 5‐year age groups are commonly used in epidemiologic age‐standardization and standardized health reporting, and have also been widely adopted in age‐related osteoporosis analyses [39, 40]. Thus, 5‐year age bands provided a practical balance between reducing age‐related confounding and preserving adequate sample size. Within each band, we applied random downsampling of non‐OP controls to achieve an approximate 1:5 OP: Normal ratio. This ratio was selected as a pragmatic compromise between reducing extreme class imbalance and retaining sufficient control samples for stable model training.

Based on this procedure, we derived three task‐specific datasets tailored to the diagnostic, prognostic, and integrated settings, as shown in Tables 1 and 2. For diagnostic modeling (prevalent OP vs. Normal), the final dataset comprised 5 180 Normal participants and 1 036 participants with prevalent OP, enabling evaluation of models for current disease detection. For prognostic modeling (incident OP vs. Normal), the dataset included 8 200 Normal participants and 1 640 individuals who developed OP during follow‐up, allowing assessment of prediction of future risk. Following the strategy used in previous interpretable machine‐learning and plasma proteomics studies, we extracted participants who developed osteoporosis only after more than 10 years of follow‐up as a non‐overlapping independent temporal validation cohort within UK Biobank. These individuals were not used for model development, SHAP‐based marker selection, or hyperparameter tuning.

TABLE 1.

Summary of diagnostic and prognostic plasma proteomics datasets for osteoporosis.

Normal OP
Participants (Diagnosis) 5180 1036
Participants (Prognosis) 8200 1640

TABLE 2.

Summary of the integrated plasma proteomics dataset distinguishing normal, prevalent osteoporosis, and incident osteoporosis cases.

Normal OP
Prevalent OP (baseline) Incident OP (follow‐up)
Participants 13 380 1036 1640

For the integrated diagnostic and prognostic task, we combined these groups to obtain 13 380 Normal participants, 1 036 with prevalent OP, and 1 640 with incident OP. Separating these task‐specific datasets allows us to disentangle diagnostic from prognostic performance, while the integrated dataset supports evaluation of models designed to handle both tasks simultaneously.

In addition to plasma proteomic measurements and osteoporosis outcomes, baseline clinical and lifestyle covariates were extracted from UK Biobank to characterize potential confounding factors and to construct clinical and combined clinical–proteomic models. These covariates included sex, body mass index (BMI), smoking status, alcohol drinker status, and other covariates. The full list of clinical covariates is provided in Table S2. These variables were included to account for demographic, anthropometric, metabolic, inflammatory, and lifestyle factors known to be associated with osteoporosis risk.

2.5. Differential Expression Analysis

Differential expression analysis was performed to identify plasma proteins associated with osteoporosis status and future osteoporosis. After quality control of the proteomic data (as described in Section 2.3), we compared protein levels between (i) Normal and prevalent OP participants in the diagnostic dataset and (ii) Normal and incident OP participants in the prognostic dataset (Table 1). For each comparison, we used the limma package in R to fit a separate linear model for every protein, with group (OP vs. Normal) as the primary predictor. Proteins with p < 0.05 and |log2FC| > 1 were retained as an initial candidate feature set for downstream machine‐learning modeling. The differential protein analysis was used as an exploratory feature‐screening step, whereas subsequent XGBoost modeling, SHAP interpretation, cross‐validation, and feature‐number sensitivity analyses were used to further prioritize predictive protein markers.

2.6. Machine‐Learning Classification Models for Osteoporosis Diagnosis and Prognosis

We used gradient boosting decision trees (XGBoost) as the core classifier. XGBoost has been widely applied to UK Biobank–scale data for disease risk prediction and phenotypic classification, and offers good performance, scalability, and handling of non‐linear effects and feature interactions [23, 25, 41]. For each task, models were trained on the age‐stratified datasets described above. We considered three prediction settings: (i) diagnostic binary classification of prevalent OP vs. Normal; (ii) prognostic binary classification of incident OP vs. Normal; and (iii) integrated three‐class classification of Normal, prevalent OP, and incident OP.

We adopted a nested cross‐validation strategy. In the outer loop, 5‐fold cross‐validation was used to obtain unbiased estimates of model performance. In the inner loop, hyperparameters were tuned using GridSearchCV in Python. The learning rate (learning_rate) and number of trees (n_estimators) were optimized over the grids learning_rate ∈ {0.01, 0.03, 0.05, 0.10, 0.15, 0.20, 0.25, 0.30} and n_estimators ∈ {50, 100, 200, 300, 500, 700, 900, 1 000}. Optimal parameter combinations in each outer fold were selected based on the highest cross‐validated average precision score in the inner cross‐validation loop. The best model from each outer fold was then refitted on the corresponding training split and evaluated on the held‐out test split for the diagnostic, prognostic, and three‐class tasks. Performance was summarized using receiver operating characteristic (ROC) curves, and macro‐averaged area under the ROC curve (macro AUC) as reported in the Results section.

For each prediction task, we compared three feature settings: Clinical covariates alone, proteomic features alone, and combined clinical–proteomic features. The clinical covariates consisted of sex, BMI, smoking status, and other variables (Table S2). Categorical variables were encoded using one‐hot encoding before model training, whereas continuous variables such as BMI were used as numerical predictors. This design allowed us to evaluate whether plasma proteomic profiles improved osteoporosis classification beyond conventional clinical and lifestyle risk factors.

2.7. Model Interpretation

To interpret the machine‐learning models and derive compact marker sets, we used SHAP (Shapley additive explanations) applied to the fitted XGBoost classifiers [29]. SHAP provides additive feature attributions for individual predictions and has been widely used to quantify the contribution of each variable in tree‐based models. For each trained XGBoost model (diagnostic, prognostic, and integrated), we computed SHAP values using the TreeExplainer implementation in the Python shap package. For every protein, we then calculated the mean absolute SHAP value across all individuals and used this as a global measure of feature importance. Proteins were ranked in descending order of mean absolute SHAP value to identify those contributing most strongly to the model output.

3. Results

3.1. Osteoporosis Diagnostic Classification and Marker Discovery

To characterize plasma proteomic signatures of established osteoporosis and derive an interpretable diagnostic marker set, we focused on prevalent OP vs. Normal participants in the diagnostic dataset (Table 1 and Figure 4). This analysis was designed to address three questions: (i) which plasma proteins are differentially expressed in prevalent OP; (ii) whether these proteins can accurately classify OP at the individual level; and (iii) which subset of proteins drives model predictions and what biological pathways they represent.

FIGURE 4.

FIGURE 4

Diagnostic differential proteomics, XGBoost classification, and SHAP‐based marker discovery for prevalent osteoporosis. (A) Volcano plot of differentially expressed plasma proteins between Normal and prevalent OP (p < 0.05, |log2 Fold Change| > 1). (B) Comparison of AUC values for XGBoost models based on different feature sets in osteoporosis diagnostic classification. (C) Sensitivity analysis of SHAP‐selected feature numbers. (D) SHAP beeswarm summary plot showing both feature importance and direction of effects in the diagnostic XGBoost model. (E–G) Gene Ontology enrichment analysis (biological process, molecular function, and cell component) of the SHAP‐selected protein markers, highlighting bone‐related pathways.

We began with differential expression analysis of plasma proteins between Normal and prevalent OP. We identified a set of significantly upregulated and downregulated proteins using the R package limma (Figure 4A and Table S3). Using p < 0.05 as the initial screening criterion, we identified candidate differentially expressed proteins for downstream model construction. The complete differential expression results, including log2FC, nominal p values, and Benjamini–Hochberg adjusted p values, are provided in Table S3. This step served to reduce dimensionality, highlight proteins with strong case–control contrasts, and provide an initial list of candidate diagnostic markers. Next, we compared XGBoost models built using clinical features alone, proteomic features alone, and combined clinical–proteomic features. As shown in Figure 4B and Table S4, the combined model achieved the highest AUC, whereas the clinical‐only model showed lower discrimination than models incorporating proteomic information.

Next, to obtain a more compact and interpretable diagnostic marker set, we applied SHAP to the full XGBoost model and ranked proteins by their mean absolute SHAP values. To determine the optimal number of SHAP‐selected features and avoid an arbitrary cutoff, we performed a sensitivity analysis by constructing XGBoost models using different numbers of top‐ranked SHAP features, ranging from the top 10 to the top 200 features at intervals of 10. As shown in Figure 4C, model performance gradually improved as more SHAP‐ranked features were included and reached the highest AUC when the top 170 features were used. Therefore, the top 170 SHAP‐ranked features were selected as the optimal feature subset. The SHAP beeswarm summary plot of the top‐ranked features is shown in Figure 4D, illustrating both feature importance and the direction of feature effects. Several proteins previously implicated in bone metabolism. For example, Mann et al. demonstrated that the COL1A1 Sp1 polymorphism is a functional variant associated with lower BMD and impaired bone material properties, thereby predisposing carriers to osteoporotic fractures by altering both bone quantity and quality [42]. Chondroadherin (CHAD) acts as a novel regulator of bone metabolism that is downregulated in osteoporotic bone and, via its integrin‐binding cyclic peptide, can attenuate ovariectomy‐induced bone loss by inhibiting preosteoclast motility and bone resorption, highlighting its translational potential for osteoporosis therapy [43].

Finally, to place these markers in a biological context, we performed Gene Ontology enrichment analysis on the SHAP‐selected proteins. Enriched terms in the biological process, molecular function categories, and cell component were strongly related to bone biology, including regulation of bone resorption, regulation of bone remodeling, and bone‐associated processes (Figure 4E–G and Table S5). Together, these results indicate that XGBoost can exploit plasma proteomic differences between Normal and prevalent OP for accurate diagnosis, while SHAP‐guided selection yields a compact, biologically coherent marker set for further modeling.

3.2. Osteoporosis Prognosis Classification and Marker Discovery

To extend our analysis from diagnosis to prognosis, we next examined whether baseline plasma proteomics could predict incident OP and yield a SHAP‐selected panel of plasma proteins predictive of future osteoporosis. Using the same overall pipeline as in the diagnostic analysis, we compared participants who developed OP during follow‐up (incident OP) with age‐matched Normal participants in the prognostic dataset (Table 1 and Figure 5).

FIGURE 5.

FIGURE 5

Prognostic differential proteomics, XGBoost classification, and SHAP‐based marker discovery for incident osteoporosis. (A) Volcano plot of differentially expressed plasma proteins between Normal and incident OP (p < 0.05, |log2 Fold Change| > 1). (B) Comparison of AUC values for XGBoost models based on different feature sets in osteoporosis prognostic classification. (C) Sensitivity analysis of SHAP‐selected feature numbers for incident OP prediction. (D) Independent validation performance of the optimized Top180 prognostic XGBoost model. (E) SHAP beeswarm summary plot showing both feature importance and direction of effects. (F, G) Gene Ontology enrichment analysis (biological process and molecular function) of the SHAP‐selected protein markers, highlighting bone‐related pathways.

Baseline differential expression analysis between Normal and incident OP identified a set of significantly upregulated and downregulated proteins (Figure 5A and Table S6). Using p < 0.05 as the initial screening criteria, we identified candidate differentially expressed proteins for downstream model construction. The complete differential expression results, including log2FC, nominal p values, and Benjamini–Hochberg adjusted p values, are provided in Table S6. For the prognostic classification task, we compared XG Boost models based on clinical features alone, proteomic features alone, and combined clinical–proteomic features. As shown in Figure 5B and Table S7, models incorporating proteomic features achieved higher AUC values than the clinical‐only model. The combined clinical–proteomic model showed performance comparable to the proteomic‐only model, suggesting that plasma proteomic profiles captured the major prognostic information for predicting incident OP. For the prognostic classification task, we further evaluated how the number of SHAP‐selected features influenced XGBoost performance. Models were trained using the top 10 to top 200 SHAP‐ranked features, with an interval of 10 features. The AUC values were compared across all feature subsets. As shown in Figure 5C, the AUC increased with the inclusion of more SHAP‐ranked features and reached its maximum at the top 180 features. These results support the use of the top 180 SHAP‐ranked features as the optimal prognostic marker subset. Accordingly, the Top180 feature set was selected to establish the optimized prognostic model. This trained Top180 model was then applied to an independent validation cohort as shown in Figure 5D. When applied to the independent validation cohort, the trained Top180 XGBoost model achieved an AUC of 0.714, supporting the robustness and generalizability of the SHAP‐selected prognostic model.

Applying SHAP to the full prognostic XGBoost model, we ranked proteins by mean absolute SHAP values and obtained the top 10 features (Figure 5E), several of which overlap with diagnostic markers and have reported roles in bone metabolism. As an illustrative example, sclerostin (SOST) emerged as a shared marker for prevalent and incident OP (Figures 4D and 5E). This is consistent with its established role in bone metabolism and with current therapies: Romosozumab, a humanized monoclonal antibody targeting SOST, is approved for the treatment of postmenopausal osteoporosis [44, 45]. The convergence between our data‐driven marker selection and an existing drug target highlights the biological and translational relevance of the identified plasma proteomic markers [46, 47, 48].

Gene Ontology enrichment of the SHAP‐selected proteins again highlighted bone‐related biological processes and molecular functions, including bone development and remodeling, extracellular matrix organization (Figure 5F,G and Table S8). Together with the diagnostic results, these findings show that baseline plasma proteomic profiles can support both accurate prognostic risk classification for incident OP and SHAP‐guided identification of a concise, biologically coherent marker panel.

3.3. Exploratory Three‐Class Stratification of Prevalent and Future Osteoporosis States

To further examine whether diagnostic and prognostic proteomic signals could be jointly captured in a single baseline sample, we performed an exploratory three‐class analysis distinguishing Normal, prevalent OP, and incident OP. We emphasize that this integrated model was intended as a complementary stratification analysis built upon the separate diagnostic and prognostic tasks, rather than as a replacement for them. Relative to the baseline blood draw, these three labels represent mutually exclusive states: No recorded OP, recorded prevalent OP, and future incident OP. We therefore used this setting to test whether the union of SHAP‐selected diagnostic and prognostic markers could provide a compact feature representation for baseline stratification across these clinically relevant states (Table 2 and Figure 6).

FIGURE 6.

FIGURE 6

Multi‐class XGBoost Model for joint osteoporosis diagnosis and prognosis. (A) ROC curves of multiclass XGBoost models using all features. (B) ROC curves of multiclass XGBoost models using SHAP‐selected diagnosis and prognosis features.

We constructed an integrated three‐class XGBoost model to distinguish Normal, prevalent OP, and incident OP participants. Two feature settings were compared: One using all available clinical and proteomic features (Figure 6A), and another using a reduced marker set formed by the union of diagnosis‐related SHAP Top170 features and prognosis‐related SHAP Top180 features (Figure 6B). Model performance was evaluated using the same outer cross‐validation splits, and macro‐averaged one‐vs.‐rest AUC was used as the primary metric for three‐class discrimination. Interestingly, the reduced marker set outperformed the all‐feature model despite containing far fewer variables. Therefore, the reduced marker set not only improved model interpretability and compactness but also enhanced the discriminative performance of the integrated three‐class XGBoost model.

Overall, our findings suggest that a single baseline proteomic panel may support stratification into no recorded OP, prevalent OP, and elevated future OP risk groups. This integrated model is best viewed as a complementary proof‐of‐concept analysis rather than a standalone clinical decision tool.

4. Discussion

In this study, we examined whether baseline protein profiles and clinical features could support osteoporosis diagnosis, future risk prediction, and marker discovery. Separate models were developed for prevalent and incident OP, after which SHAP was used to identify smaller diagnostic and prognostic protein and clinical marker sets. We further assessed whether these markers could be combined for integrated stratification.

Our work extends current approaches to osteoporosis risk assessment, which primarily rely on DXA‐derived BMD and clinical risk tools [7, 8]. These tools are valuable but have well‐recognized limitations, including limited sensitivity for early or subclinical disease and incomplete capture of bone quality and systemic influences on bone turnover. In contrast, plasma proteomics integrates signals from bone and extra‐skeletal tissues, including endocrine, inflammatory, and metabolic pathways [11, 16, 17]. By leveraging UK Biobank's large‐scale, standardized Olink proteomic resource, SPX‐OP demonstrates that plasma protein signatures can be used not only to distinguish existing osteoporosis but also to predict future disease and to unify diagnostic and prognostic assessment within a single modeling framework [18, 19, 20].

The SHAP‐prioritized protein markers are biologically and clinically plausible. Proteins such as FSHB, ADIPOQ, SOST, COL9A1, and CHAD have been implicated in bone metabolism, matrix organization, and endocrine regulation of skeletal homeostasis [42, 43, 46, 47, 48]. FSHB is biologically relevant to osteoporosis because FSH has been shown to stimulate TNFα production from bone marrow immune cells, thereby enhancing osteoclast and osteoblast formation and contributing to hypogonadal bone loss [49]. ADIPOQ encodes adiponectin, an adipose‐derived hormone that has been reported to regulate bone mass. This suggests that ADIPOQ may reflect the interaction between fat metabolism and bone metabolism in osteoporosis [50]. Many of the SHAP‐selected proteins were enriched in Gene Ontology categories related to bone development, bone remodeling, extracellular matrix organization, and inflammatory processes. The identification of sclerostin (SOST) as a shared diagnostic and prognostic marker is particularly notable, given its central role in Wnt signaling and the availability of Romosozumab, a monoclonal antibody targeting SOST for the treatment of postmenopausal osteoporosis [44, 45]. The convergence between our data‐driven marker selection and known bone biology, including an established drug target, supports the mechanistic relevance of the SPX‐OP marker panels and suggests that other identified proteins may point to new pathways or therapeutic candidates.

From a clinical perspective, the integrated three‐class SPX‐OP model should be interpreted cautiously. Rather than replacing the separate diagnostic and prognostic models, it serves as a complementary proof‐of‐concept analysis showing that a single baseline plasma proteomic profile may support stratification of individuals into three groups: No recorded OP, prevalent OP, and elevated future OP risk. The marker‐based multiclass model, built from the union of diagnostic and prognostic SHAP‐selected proteins, achieved higher macro AUC than the full‐proteome model while using far fewer features, indicating that a focused panel captures most of the useful signal. SPX‐OP should be viewed as a potential adjunctive risk‐stratification tool rather than a replacement for DXA. If externally validated, it may help identify individuals who warrant confirmatory DXA testing or closer follow‐up within existing screening pathways.

Methodologically, this study has several strengths. We used a large, deeply phenotyped cohort with standardized plasma proteomic measurements and clearly defined prevalent and incident OP outcomes. Restricting analyses to individuals aged > 50 years and constructing age‐stratified datasets with controlled OP: Normal ratios helped to reduce confounding by age and mitigate extreme class imbalance. Nested cross‐validation with systematic hyperparameter tuning provided robust performance estimates for diagnostic, prognostic, and integrated models. The use of SHAP enabled transparent, protein‐level attributions, facilitating protein marker selection, biological interpretation, and construction of reduced “marker‐only” models that are more amenable to clinical translation than opaque, high‐dimensional predictors.

Several limitations should be acknowledged. First, OP outcomes were defined using UK Biobank first‐occurrence records linked to ICD‐10 codes rather than universal DXA‐confirmed diagnoses. Thus, prevalent and incident OP labels reflect the timing of first recorded clinical diagnosis rather than the true biological onset of disease, and some undiagnosed individuals may have been misclassified as controls. Second, the identified protein markers were derived from large‐scale UK Biobank plasma proteomics and were not validated in independent clinical samples or functional model systems. Therefore, these markers should be regarded as statistically prioritized and biologically plausible candidates rather than definitively validated clinical biomarkers or causal regulators of osteoporosis. We did not perform a formal cost‐effectiveness analysis or implementation study, and therefore the economic and workflow implications of deploying SPX‐OP relative to DXA‐based screening remain uncertain. Future work should evaluate external validity, workflow feasibility, and health‐economic performance in clinically annotated screening cohorts.

Importantly, our goal was not simply to apply XGBoost and SHAP to osteoporosis, but to determine whether full‐proteome models could be distilled into smaller and biologically meaningful marker panels for both diagnosis and future risk prediction.

Author Contributions

Yuhui Kou, Wenxiang Zhang, and Wenjing Zhang conceived the concept. Wenxiang Zhang designed the methodology and performed experiments. Wenxiang Zhang and Wenjing Zhang wrote the manuscript. Yuhui Kou, Baoguo Jiang, and Hanwen Cheng contributed to the revision of the manuscript. Yuhui Kou, Weijie Gong, and Baoguo Jiang supervised this work.

Funding

This work was supported by the National Key R&D Program of China (MOST), grant no. 2023YFC2509900, Shenzhen Medical Research Fund, grant no. D2401015, the National Key R&D Program Priority Special Projects of China, grant no. 2024YFC3016605, Guangdong Medical Research Fund, grant no. B2025088, Beijing Municipal Natural Science Foundation‐Fengtai Joint Fund, grant no. 2024FTQY037, and Shenzhen Clinical Research Center for Trauma Treatment, grant no. 20230731111952004.

Conflicts of Interest

The authors declare no conflicts of interest.

Supporting information

Tables S1–S8: fsb272201‐sup‐0001‐TableS1‐S8.xlsx. Table S1: Full names of the 2923 unique proteins in this study. Table S2: The clinical covariate information. Table S3: Differential expression analysis of plasma proteins between the normal and prevalent OP groups. Table S4: Performance metrics of the clinical‐only, proteomics‐only, and combined models for osteoporosis diagnostic classification. Table S5: Gene Ontology enrichment analysis (biological process, molecular function, and cell component) for prevalent osteoporosis. Table S6: Differential expression analysis of plasma proteins between the normal and incident osteoporosis. Table S7: Performance metrics of the clinical‐only, proteomics‐only, and combined models for osteoporosis prognosis classification. Table S8: The results of Gene Ontology enrichment analysis (biological process, molecular function, and cell component).

FSB2-40-e72201-s001.xlsx (1.1MB, xlsx)

Acknowledgments

We are grateful to all the lab members in the Jiang laboratory for their assistance in this study. We sincerely thank the reviewers and journal staff for their constructive comments and support.

Contributor Information

Yuhui Kou, Email: yuhuikou@bjmu.edu.cn.

Baoguo Jiang, Email: jiangbaoguo@vip.sina.com.

Data Availability Statement

The plasma proteomic and osteoporosis data used in this study were obtained from the UK biobank database (https://www.ukbiobank.ac.uk/enable‐your‐research/apply‐for‐access). In accordance with UK biobank policies, these individual‐level data cannot be shared by the authors. Researchers can apply for access directly from UK Biobank through its established data access procedures. The SPX‐OP code is openly available on GitHub (https://github.com/zwx94/SPX‐OP.git).

References

  • 1. Rachner T. D., Khosla S., and Hofbauer L. C., “Osteoporosis: Now and the Future,” Lancet 377, no. 9773 (2011): 1276–1287. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Seletkov D., Starck S., Mueller T. T., et al., “AI‐Driven Preclinical Disease Risk Assessment Using Imaging in UK Biobank,” npj Digital Medicine 8, no. 1 (2025): 480. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Eastell R., O'Neill T. W., Hofbauer L. C., et al., “Postmenopausal osteoporosis,” Nature Reviews Disease Primers 2 (2016): 16069. [DOI] [PubMed] [Google Scholar]
  • 4. Li Q., Sun T., Su J., Liu S., Zhang W., and Kou Y., “Multifunctional Bone Repair Platform: Cytokines Scavenging and Sequential Drug Release for Osteoporotic Bone Defects,” Materials & Design 260 (2025): 115127. [Google Scholar]
  • 5. Zhao P., Ying Z., Yuan C., et al., “Shared Genetic Architecture Highlights the Bidirectional Association Between Major Depressive Disorder and Fracture Risk,” General Psychiatry 37, no. 3 (2024): e101418. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Ito K., “Cost‐Effectiveness of Screening for Osteoporosis in Older Men With a History of Falls,” JAMA Network Open 3, no. 12 (2020): e2027584. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Khan A. A., Slart R., Ali D. S., et al., “Osteoporotic Fractures: Diagnosis, Evaluation, and Significance From the International Working Group on DXA Best Practices,” Mayo Clinic Proceedings 99, no. 7 (2024): 1127–1141. [DOI] [PubMed] [Google Scholar]
  • 8. Force USPST , Nicholson W. K., Silverstein M., et al., “Screening for Osteoporosis to Prevent Fractures: US Preventive Services Task Force Recommendation Statement,” Journal of the American Medical Association 333, no. 6 (2025): 498–508. [DOI] [PubMed] [Google Scholar]
  • 9. Malgo F., Hamdy N. A., Papapoulos S. E., and Appelman‐Dijkstra N. M., “Bone Material Strength as Measured by Microindentation in Vivo Is Decreased in Patients With Fragility Fractures Independently of Bone Mineral Density,” Journal of Clinical Endocrinology and Metabolism 100, no. 5 (2015): 2039–2045. [DOI] [PubMed] [Google Scholar]
  • 10. Chia R., Moaddel R., Kwan J. Y., et al., “A Plasma Proteomics‐Based Candidate Biomarker Panel Predictive of Amyotrophic Lateral Sclerosis,” Nature Medicine 31, no. 10 (2025): 3440–3450. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Deng Y. T., You J., He Y., et al., “Atlas of the Plasma Proteome in Health and Disease in 53,026 Adults,” Cell 188, no. 1 (2025): 71e7–71e253. [DOI] [PubMed] [Google Scholar]
  • 12. Chen S., Yan K., Li X., and Liu B., “Protein Language Pragmatic Analysis and Progressive Transfer Learning for Profiling Peptide‐Protein Interactions,” IEEE Transactions on Neural Networks and Learning Systems 36, no. 8 (2025): 15385–15399. [DOI] [PubMed] [Google Scholar]
  • 13. Zhang J., Zeng H., Chen J., and Zhu Z., “INAB: Identify Nucleic Acid Binding Domain via Cross‐Modal Protein Language Models and Multiscale Computation,” Briefings in Bioinformatics 26, no. 5 (2025): bbaf509. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Wang F., Ding Y., Lei X., Liao B., and Wu F. X., “Human Protein Complex‐Based Drug Signatures for Personalized Cancer Medicine,” IEEE Journal of Biomedical and Health Informatics 25, no. 11 (2021): 4079–4088. [DOI] [PubMed] [Google Scholar]
  • 15. Pan Y., Lei X., and Zhang Y., “Association Predictions of Genomics, Proteinomics, Transcriptomics, Microbiome, Metabolomics, Pathomics, Radiomics, Drug, Symptoms, Environment Factor, and Disease Networks: A Comprehensive Approach,” Medicinal Research Reviews 42, no. 1 (2022): 441–461. [DOI] [PubMed] [Google Scholar]
  • 16. Eldjarn G. H., Ferkingstad E., Lund S. H., et al., “Large‐Scale Plasma Proteomics Comparisons Through Genetics and Disease Associations,” Nature 622, no. 7982 (2023): 348–358. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Afshar S., Dammer E. B., Bian S., et al., “Plasma Proteomic Associations With Alzheimer's Disease Endophenotypes,” Nature Aging 5, no. 10 (2025): 2104–2124. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Carrasco‐Zanini J., Pietzner M., Davitte J., et al., “Proteomic Signatures Improve Risk Prediction for Common and Rare Diseases,” Nature Medicine 30, no. 9 (2024): 2489–2498. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. You J., Guo Y., Zhang Y., et al., “Plasma Proteomic Profiles Predict Individual Future Health Risk,” Nature Communications 14, no. 1 (2023): 7817. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Goeminne L. J. E., Vladimirova A., Eames A., et al., “Plasma Protein‐Based Organ‐Specific Aging and Mortality Models Unveil Diseases as Accelerated Aging of Organismal Systems,” Cell Metabolism 37, no. 1 (2025): 22.e6–22.e205. [DOI] [PubMed] [Google Scholar]
  • 21. Chen Y., Long T., Wang M., et al., “Prospective Cohort Study Integrating Plasma Proteomics and Machine Learning for Early Risk Prediction of Prostate Cancer,” International Journal of Surgery 111, no. 9 (2025): 6123–6134. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Liu W., Chen W., Dong D., Zhang G., and Xing N., “Pre‐Diagnostic Plasma Proteomics Profile Uncovers New Biomarkers and Mechanistic Insights for Incident Kidney Cancer,” International Journal of Surgery 112, no. 1 (2025): 873–886. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Nielsen R. L., Monfeuga T., Kitchen R. R., et al., “Data‐Driven Identification of Predictive Risk Biomarkers for Subgroups of Osteoarthritis Using Interpretable Machine Learning,” Nature Communications 15, no. 1 (2024): 2817. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Chen R., Duffy A., Petrazzini B. O., et al., “Expanding Drug Targets for 112 Chronic Diseases Using a Machine Learning‐Assisted Genetic Priority Score,” Nature Communications 15, no. 1 (2024): 8891. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Pan Z., Zhang R., Shen S., et al., “OWL: An Optimized and Independently Validated Machine Learning Prediction Model for Lung Cancer Screening Based on the UK Biobank, PLCO, and NLST Populations,” eBioMedicine 88 (2023): 104443. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Jiang Y., Zhao B., Wang X., et al., “UKB‐MDRMF: A Multi‐Disease Risk and Multimorbidity Framework Based on UK Biobank Data,” Nature Communications 16, no. 1 (2025): 3767. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Julkunen H. and Rousu J., “Comprehensive Interaction Modeling With Machine Learning Improves Prediction of Disease Risk in the UK Biobank,” Nature Communications 16, no. 1 (2025): 6620. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Li M., He X., Gong W., et al., “Plasma Proteomics Profiles Predict the Risk of Future Aortic Aneurysm and Aortic Dissection,” International Journal of Surgery 111, no. 10 (2025): 6894–6904. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Lundberg S. M. and Lee S.‐I., “A Unified Approach to Interpreting Model Predictions,” in Proceedings of the 31st International Conference on Neural Information Processing Systems (Curran Associates, Inc, 2017), 4768–4777. [Google Scholar]
  • 30. Julkunen H., Cichonska A., Tiainen M., et al., “Atlas of Plasma NMR Biomarkers for Health and Disease in 118,461 Individuals From the UK Biobank,” Nature Communications 14, no. 1 (2023): 604. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Sudlow C., Gallacher J., Allen N., et al., “UK Biobank: An Open Access Resource for Identifying the Causes of a Wide Range of Complex Diseases of Middle and Old Age,” PLoS Medicine 12, no. 3 (2015): e1001779. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Sun B. B., Chiou J., Traylor M., et al., “Plasma Proteomic Associations With Genetics and Health in the UK Biobank,” Nature 622, no. 7982 (2023): 329–338. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Wik L., Nordberg N., Broberg J., et al., “Proximity Extension Assay in Combination With Next‐Generation Sequencing for High‐Throughput Proteome‐Wide Analysis,” Molecular & Cellular Proteomics 20 (2021): 100168. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Robbins J. M., Peterson B., Schranner D., et al., “Human Plasma Proteomic Profiles Indicative of Cardiorespiratory Fitness,” Nature Metabolism 3, no. 6 (2021): 786–797. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Zheng Y., Li J., Li Y., et al., “Plasma Proteomic Profiles Reveal Proteins and Three Characteristic Patterns Associated With Osteoporosis: A Prospective Cohort Study,” Journal of Advanced Research 75 (2025): 491–503. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Bycroft C., Freeman C., Petkova D., et al., “The UK Biobank Resource With Deep Phenotyping and Genomic Data,” Nature 562, no. 7726 (2018): 203–209. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. Di D., Zhou H., Cui Z., et al., “Early‐Life Tobacco Smoke Elevating Later‐Life Osteoporosis Risk: Mediated by Telomere Length and Interplayed With Genetic Predisposition,” Journal of Advanced Research 68 (2025): 331–340. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Morin S. N., Feldman S., Funnell L., et al., “Clinical Practice Guideline for Management of Osteoporosis and Fracture Prevention in Canada: 2023 Update,” CMAJ 195, no. 39 (2023): E1333–E1348. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Diaz T., Strong K. L., Cao B., et al., “A Call for Standardised Age‐Disaggregated Health Data,” Lancet Healthy Longevity 2, no. 7 (2021): e436–e443. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Boschitsch E. P., Durchschlag E., and Dimai H. P., “Age‐Related Prevalence of Osteoporosis and Fragility Fractures: Real‐World Data From an Austrian Menopause and Osteoporosis Clinic,” Climacteric 20, no. 2 (2017): 157–163. [DOI] [PubMed] [Google Scholar]
  • 41. Garg M., Karpinski M., Matelska D., et al., “Disease Prediction With Multi‐Omics and Biomarkers Empowers Case‐Control Genetic Discoveries in the UK Biobank,” Nature Genetics 56, no. 9 (2024): 1821–1831. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Mann V., Hobson E. E., Li B., et al., “A COL1A1 Sp1 Binding Site Polymorphism Predisposes to Osteoporotic Fracture by Affecting Bone Density and Quality,” Journal of Clinical Investigation 107, no. 7 (2001): 899–907. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Capulli M., Olstad O. K., Onnerfjord P., et al., “The C‐Terminal Domain of Chondroadherin: A New Regulator of Osteoclast Motility Counteracting Bone Loss,” Journal of Bone and Mineral Research: the Official Journal of the American Society for Bone and Mineral Research 29, no. 8 (2014): 1833–1846. [DOI] [PubMed] [Google Scholar]
  • 44. Knox C., Wilson M., Klinger C. M., et al., “DrugBank 6.0: The DrugBank Knowledgebase for 2024,” Nucleic Acids Research 52, no. D1 (2024): D1265–D1275. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. Meng S., Liu Q., Dai R., et al., “Development of a Novel Macroscopic Regulation and Microscopic Intervention Mode Nanosystem for Osteoporosis Treatment,” Materials Today Bio 32 (2025): 101829. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Reid I. R., “Short‐Term and Long‐Term Effects of Osteoporosis Therapies,” Nature Reviews Endocrinology 11, no. 7 (2015): 418–428. [DOI] [PubMed] [Google Scholar]
  • 47. Wu D., Li L., Wen Z., and Wang G., “Romosozumab in Osteoporosis: Yesterday, Today and Tomorrow,” Journal of Translational Medicine 21, no. 1 (2023): 668. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Ferrari S. L., “Osteoporosis: Romosozumab to Rebuild the Foundations of Bone Strength,” Nature Reviews Rheumatology 14, no. 3 (2018): 128. [DOI] [PubMed] [Google Scholar]
  • 49. Iqbal J., Sun L., Kumar T. R., Blair H. C., and Zaidi M., “Follicle‐Stimulating Hormone Stimulates TNF Production From Immune Cells to Enhance Osteoblast and Osteoclast Formation,” Proceedings of the National Academy of Sciences of the United States of America 103, no. 40 (2006): 14925–14930. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50. Kajimura D., Lee H. W., Riley K. J., et al., “Adiponectin Regulates Bone Mass via Opposite Central and Peripheral Mechanisms Through FoxO1,” Cell Metabolism 17, no. 6 (2013): 901–915. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Tables S1–S8: fsb272201‐sup‐0001‐TableS1‐S8.xlsx. Table S1: Full names of the 2923 unique proteins in this study. Table S2: The clinical covariate information. Table S3: Differential expression analysis of plasma proteins between the normal and prevalent OP groups. Table S4: Performance metrics of the clinical‐only, proteomics‐only, and combined models for osteoporosis diagnostic classification. Table S5: Gene Ontology enrichment analysis (biological process, molecular function, and cell component) for prevalent osteoporosis. Table S6: Differential expression analysis of plasma proteins between the normal and incident osteoporosis. Table S7: Performance metrics of the clinical‐only, proteomics‐only, and combined models for osteoporosis prognosis classification. Table S8: The results of Gene Ontology enrichment analysis (biological process, molecular function, and cell component).

FSB2-40-e72201-s001.xlsx (1.1MB, xlsx)

Data Availability Statement

The plasma proteomic and osteoporosis data used in this study were obtained from the UK biobank database (https://www.ukbiobank.ac.uk/enable‐your‐research/apply‐for‐access). In accordance with UK biobank policies, these individual‐level data cannot be shared by the authors. Researchers can apply for access directly from UK Biobank through its established data access procedures. The SPX‐OP code is openly available on GitHub (https://github.com/zwx94/SPX‐OP.git).


Articles from The FASEB Journal are provided here courtesy of Wiley

RESOURCES