Skip to main content
BMC Medical Informatics and Decision Making logoLink to BMC Medical Informatics and Decision Making
. 2026 Jul 8;26:379. doi: 10.1186/s12911-026-03632-9

A generalizable computational framework for integrating heterogeneous biomarkers into interpretable scalar risk representations

Fernanda Oliveira Duarte 1,✉, Mauro Masili 2,✉, Luciana Camillo 1, Jaqueline Bianchi 1, Krissia Franco de Godoy 1, Bruna Dias de Lima Fragelli 1, Joice Margareth de Almeida Rodolpho 1, Juliana de Almeida Prado 3, Carlos Speglich 4, Fernanda de Freitas Aníbal 1,✉
PMCID: PMC13628972  PMID: 42421067

Abstract

A central challenge in biomedical informatics remains the integration of heterogeneous biomarker data into unified, interpretable, and clinically meaningful representations. In this study, we present a computational framework for the construction of unified scalar representations derived from multiple blood-based biomarkers, with the objective of capturing latent structures associated with mental health status. As a proof-of-principle, we instantiate this framework through the Mental Disorder Risk Index (MDRI), a scalar and dimensionless representation that integrates multiple biomarkers through an optimization-based and severity-sensitive aggregation strategy. The proposed framework includes linear and nonlinear formulations, with parameter estimation performed through a combination of deterministic and meta-heuristic optimization strategies. Model performance is evaluated using established metrics, mainly the Matthews correlation coefficient. The framework is applied to a dataset composed of eight biomarkers and respective mental health scores. The results indicate that stable scalar representations can be obtained within the normalized biomarker space, with robustness observed in several optimization trajectories, confirming the stability of the underlying latent structure of the biomarkers. Sensitivity analysis reveals heterogeneous contributions of the biomarkers, indicating that the proposed index captures non-trivial multivariate relationships. These findings establish the computational viability of the construction of unified biomarker-based indices and provide a structured basis for the development of interpretable, scalable, and generalizable representations of complex biomedical data.

Keywords: Computational framework, Biomedical informatics, Biomarker integration, Risk representation, Interpretability, Multivariate modeling

Introduction

Mental health disorders are characterized by high biological and phenotypic heterogeneity, representing significant challenges for the development of quantitative and generalizable diagnostic tools. Despite the growing interest in biomarker-based approaches, the integration of multivariate biological data into unified, interpretable, and clinically meaningful representations remains an open problem. From a systems perspective, this challenge can be framed as the mapping of high-dimensional biomarker data into lower-dimensional representations capable of preserving relevant latent structural relationships.

Biomedical research increasingly relies on the integration of heterogeneous data sources and computational approaches, particularly machine learning, to model complex relationships between biological variables and support clinical decision-making [1, 2]. In this context, blood-based biomarkers provide valuable quantitative information [3]. However, their computational synthesis remains challenging when multiple markers with different units, scales, and interactions need to be analyzed in an integrated manner [4]. Current analytical approaches typically oscillate between isolated biomarker thresholds, which oversimplify biological complexity, and black-box predictive models, which often lack the interpretability, reproducibility, and transferability required for clinical adoption [5, 6]. As a result, findings remain inconsistent, and there is no consensus on how to systematically integrate these biomarkers into a unified and interpretable risk metric.

Existing approaches for biomarker integration in mental health generally rely on statistical dimensionality reduction techniques or machine learning models. Although these methods can capture multivariate relationships, there is still a critical need for formal computational frameworks capable of constructing stable, interpretable, and generalizable scalar representations across different contexts. In this scenario, there is a critical need for formal computational structures capable of transforming multivariate biomarker profiles into stable, interpretable, and reusable risk representations. Such structures must explicitly address fundamental challenges in biomedical computation, including numerical conditioning, feature heterogeneity, and uncertainty propagation, particularly in clinical domains where objective and standardized risk metrics have not yet been established [4, 7].

Mental health provides a particularly relevant domain to address this challenge. Psychiatric disorders are among the leading causes of disability worldwide and contribute substantially to the global burden of disease, representing a significant proportion of years lived with disability [8]. Their etiology involves complex interactions among biological, psychological, and environmental factors [8]. Due to the limited understanding of these mechanisms, mental disorders are characterized by a heterogeneous range of signs and symptoms [9]. Thus, blood-based biomarkers related to mental health act as indicators of physiological, biochemical, and molecular alterations triggered by different stressors, whether physical, psychological, or environmental, that activate complex cascades of physiological responses aimed at maintaining homeostasis and facilitating adaptation [10]. Indeed, mental disorders are often associated with low-grade inflammation and dysregulation of stress hormones. Among the main blood biomarkers are CRP, hsCRP, IL-6, IL-1β, TNF-α, IL-8, cortisol, and vitamin D [11–18].

In medical science, established health indices such as the Homeostasis Model Assessment for Insulin Resistance (HOMA-IR), the Fibrosis-4 Index (FIB-4), and the Atherogenic Index of Plasma (AIP) are widely used to simplify complex physiological states into actionable metrics [19–21], often benefiting from well-established pathophysiological pathways, validated thresholds, and large homogeneous cohorts. However, mental health disorders present additional methodological challenges, including high biological variability, complex interactions among biomarkers, symptom heterogeneity, and the absence of validated composite risk indices. This gap contributes to underdiagnosis and is intensified by symptom overlap across different disorders [22, 23], with negative impacts on patient clinical outcomes [24]. Recent studies have demonstrated the utility of composite inflammation-based indices, biomarker panels, and machine learning models for integrated mental health assessment [25–28]. However, these approaches still require a unified architecture that simultaneously addresses numerical stability, data heterogeneity, and model-independent interpretability, highlighting the need for systematic approaches capable of integrating biomarkers into standardized and reproducible risk representations. For these reasons, mental health was deliberately selected as the primary domain for the proof-of-principle application of the proposed computational framework. The Mental Disorder Risk Index (MDRI) is presented as an instantiation of this framework, employing an optimization-driven and severity-sensitive aggregation strategy capable of transforming heterogeneous biomarkers into a dimensionless scalar representation, expanding the flexibility of conventional linear or threshold-based composite approaches [29].

In this context, it is worth emphasizing that the main contributions of this study include the formulation of a generalizable computational framework to integrate heterogeneous blood-based biomarkers into unified scalar representations; the implementation of linear and nonlinear model formulations combined with deterministic, meta-heuristic, and machine learning–based parameter estimation strategies; the proof-of-principle instantiation of the framework through the Mental Disorder Risk Index (MDRI); and the evaluation of the structural properties of the framework, including stability, sensitivity, and interpretability, within a multivariate biomarker context.

In light of these contributions, this study investigates the feasibility of constructing a unified scalar index derived from multiple blood-based biomarkers, aiming to capture latent patterns associated with mental health status. The proposed framework incorporates both linear and nonlinear formulations, with parameter estimation performed through a combination of deterministic and metaheuristic algorithms. Emphasis is placed on internal consistency and the structural properties of the resulting index, evaluated through sensitivity analysis and model-dependent interpretability techniques applied to a dataset composed of eight biomarkers and their respective mental health scores.

Methods

Subjects

This study utilized a dataset comprising 158 patients with mental disorders and 144 healthy control volunteers, each of whom underwent comprehensive blood tests to quantify the levels of the biomarkers cortisol, CRP, hsCRP, vitamin D, IL-6, IL-1Inline graphic, TNF-Inline graphic, and IL-8.

The study sample included healthy individuals recruited via a public announcement at the Federal University of São Carlos (UFSCar) and patients diagnosed with mental disorders from two psychiatric institutions and one intermediate healthcare unit between outpatient and hospital care located in three cities in São Paulo State, Brazil. Blood collection was performed on a single occasion by qualified healthcare professionals at the aforementioned institutions. The inclusion criteria for the patient group were age over 18 years, the absence of autoimmune diseases or cancer, and a confirmed mental disorder diagnosis by a psychiatrist, according to the International Classification of Diseases (ICD). The control group comprised healthy adults aged 18 years or older with no history of autoimmune disease, cancer, mental illness, substance abuse, or chronic medication use who were recruited from diverse populations across multiple institutions, geographic regions, and data collection periods via identical protocols.

To characterize the sample and evaluate participants’ mental health, a questionnaire collecting sociodemographic and clinical data was administered. To measure levels of depression, anxiety, and stress, all participants completed the Depression, Anxiety, and Stress Scale (DASS-21). The DASS-21, a widely used self-report instrument, evaluates the frequency and intensity of symptoms related to these three mental disorders over the preceding weeks. The DASS-21 consists of 21 items, with seven items for each of the evaluated disorders. The participants indicate the degree to which each symptom applies to them via a 4-point Likert scale. A final score is then attributed to the participant, which is categorized as ‘Subclinical’, Mild’, ‘Moderate’, or ‘Severe’ [30].

The DASS-21 was selected for its proven reliability and validity in simultaneously assessing symptoms of depression, anxiety, and stress across different populations. Unlike many other scales that focus on a single construct, which require multiple tools for multidimensional analysis, the DASS-21 provides a comprehensive evaluation, enhancing the understanding of cooccurring symptoms [30].

Post-traumatic stress disorder (PTSD) was assessed via a subset of items from the Post-Traumatic Stress Disorder Checklist (PCL-5), which was validated in a Brazilian sample by Osório et al. [31]. This subset focused on the PTSD core symptoms, with a score of 6 or higher used as the cutoff for probable PTSD. The self-reported responses reflected symptoms experienced in the past month.

Experimental procedure

Peripheral venous blood samples were collected from each participant in a single session between 7:00 and 10:00 AM; vacuum tubes with EDTA were used for whole blood analysis, and dry tubes with separator gel were used for serum extraction. Immediately after collection, whole blood samples were analyzed via semiautomated portable immunofluorescence analyzers: the QuickSTAR (In Vitro) for CRP, hsCRP, and cortisol quantification and the ECO (SD Biosensor) for vitamin D quantification. Serum was isolated by refrigerated centrifugation at 4 °C at 4000 rpm for 10 min and then stored at − 80 °C until analysis. Cytokine concentrations (TNF-Inline graphic, IL-6, IL-8, and IL-1Inline graphic) were measured via chemiluminescence with an Immulite® 1000 analyzer (Siemens Healthcare Diagnostics).

Conceptual framework

This section outlines the computational framework underlying the transformation of multivariate biomarker data into a unified and interpretable risk representation for mental health risk assessment.

The mathematical formulation adopted in this study is intentionally constructed as a flexible parametric representation rather than a strictly mechanistic model. The objective is not to impose a priori biological constraints, but to investigate whether a low-dimensional scalar mapping can capture structured relationships within a heterogeneous biomarker space. Accordingly, the proposed formulations should be interpreted as phenomenological representations, whose validity is assessed in terms of internal consistency, numerical behavior, and sensitivity to parameter variations, rather than strict physiological interpretability.

To facilitate understanding of the proposed methodology, a schematic overview of the computational framework is presented in Fig. 1, illustrating the main components, data flow, and integration of model formulation, optimization strategies, and analysis procedures leading to the construction of the MDRI.

Fig. 1.

Fig. 1

Conceptual overview of the proposed generalizable computational framework for constructing the Mental Disorder Risk Index (MDRI). Blood-based biomarkers and mental health outcomes are integrated through a unified parametric formulation. Multiple parameter estimation strategies, including deterministic optimization, metaheuristic optimization, and machine learning approaches, are employed employed to generate independent MDRI variants. Model interpretability and robustness are assessed through permutation feature importance (PFI), SHAP analysis, and Sobol′ sensitivity analysis. The resulting indices are combined through a weighted-voting ensemble strategy and subsequently validated, yielding a unified scalar representation of mental health risk

Computational resources

The dataset was stored as Excel-compatible .csv files. The main computational codes were developed in Python (version 3.12.1) within the JupyterLab environment (version 4.1.8). Some auxiliary code was written in Fortran and compiled with Intel compiler (version 2022.1.0). The codes were executed on an Intel Core i7 12th-generation CPU with 32 GB of RAM, typically requiring 4–5 h to analyze each mental health disorder.

Data cleaning and preprocessing

The dataset was cleaned by removing patients with missing bloodwork results or mental health assessments. Outliers were excluded by employing a combination of the jackknife distance and the Mahalanobis distance methods [32], with a 95% upper confidence limit as the threshold. This process excluded seven outliers and 18 subjects with missing data from the patient group, resulting in  n=133 patients. In the control group, 10 outliers and 12 subjects with missing data were removed, resulting in a sample size of  n=122 subjects.

The mental health scores for depression, anxiety, and stress were binarized into two schemes:

  • Scheme 1 × 3: Scores categorized as ‘Subclinical’ were labeled ‘NO’, whereas ‘Mild’, ‘Moderate’, or ‘Severe’ were labeled ‘YES’. This binarization strategy distinguished participants with virtually no mental disorders from those experiencing any level of mental health condition.

  • Scheme 2 × 2: Scores categorized as ‘Subclinical’ or ‘Mild’ were labeled ‘NO’, whereas ‘Moderate’ or ‘Severe’ were labeled ‘YES’. This scheme aimed to identify participants with higher severity levels of mental disorders.

PTSD scores were inherently binary within the dataset. The cleaned dataset was then divided into two portions via stratified random splitting, with 80% of the patients used for training the models and the remaining 20% reserved for testing.

To prevent any feature from overwhelming the results, we employed  z-score standardization and min–max normalization. These methods address different scaling goals.  Z‑score standardization re‑expresses each biomarker in units of its own standard deviation, centering at zero and giving all biomarkers comparable variance regardless of their original units or dispersion. Applying min–max normalization to the standardized features maps them into a fixed interval Inline graphic, which is often preferred for algorithms that assume bounded inputs or when using regularization and optimization schemes. This second step also enables interpretability and comparison of feature importances across biomarkers because all final variables lie on the same numeric scale. This transformation ensures that subsequent algebraic operations, including nonlinear transformations, are applied to commensurable quantities. Consequently, the resulting index is dimensionless by construction. Therefore, the index should be interpreted as a relative measure within the normalized feature space rather than a physically grounded quantity.

The blood test data in the training sample were first standardized via  z-score normalization to reduce variability, as expressed by:

graphic file with name d33e481.gif 1

in which Inline graphic and Inline graphic represent the mean and standard deviation of the variableInline graphic, respectively. The standardized data were then normalized to ensure that each biomarker’s range fell within the Inline graphic interval, according to:

graphic file with name d33e504.gif 2

In Eq. (1), the biomarker variables Inline graphic are Inline graphic cortisol (nmol/L), Inline graphic CRP (mg/L), Inline graphic hsCRP (mg/L), Inline graphic vitamin D (ng/mL), Inline graphic IL-6 (pg/mL), Inline graphic IL-1Inline graphic (pg/mL), Inline graphic TNF-Inline graphic (pg/mL), and Inline graphic IL-8 (pg/mL). The normalized variables Inline graphic are dimensionless.

Notably, the remaining 20% of the data, reserved for testing the models, were standardized and normalized via the statistics (Inline graphic, Inline graphic, Inline graphic, Inline graphic) derived from the training sample. This procedure prevented information leakage from the test set to the training set (train-test contamination), which would occur if standardization and normalization were performed on the entire dataset before splitting. This train-test contamination unintentionally influences the model during training, leading to overly optimistic performance estimates and poor generalizability [33]. Hereafter, the normalized biomarker variables Inline graphic in Eq. (2) are redefined as Inline graphic for brevity.

Exploratory data analysis and feature selection

We conducted descriptive statistical analysis of the demographic and clinical features of the participants to characterize the datasets.

Relationships among the biomarkers were assessed via Hoeffding’s D test of independence, which allowed for the identification of potential correlations. Hoeffding’s D test is a nonparametric measure of association that can detect whether there are significant relationships between biomarkers and identify potential patterns or general dependencies between variables, including nonlinear relationships. It is particularly useful when variables might be related in ways that cannot be captured by linear correlation measures, such as Pearson’s correlation coefficient.

Theoretical foundations

This section outlines a multistep formulation for developing a unified metric to analyze biomarkers potentially associated with mental health disorders comprehensively. Subsequent subsections detail regression analyses using machine learning, metaheuristic, and deterministic models.

The primary rationale for employing multiple models is empirical: given the heterogeneity of biomarker-outcome relationships, the relatively small sample size, and the exploratory nature of this investigation, no single model structure can be assumed optimal a priori. Consistent findings across methodologically distinct approaches (metaheuristic, deterministic, and machine learning) provide stronger evidence of structural stability than any single method in isolation, a consideration aligned with the broader principle that no algorithm universally outperforms all others across all possible problems, known as “No Free Lunch” (NFL) theorem [34]. Thus, we employed a set of methodologically distinct approaches to assess the internal coherence of the framework across different optimization dynamics.

Machine learning modeling

We employed machine learning techniques to explore the complex relationships between biomarkers and target variables. We evaluated the predictive performance of 12 classification algorithms, including AdaBoost, CN2 rule induction, decision tree, gradient boosting,  k-nearest neighbors, logistic regression, naïve Bayes, neural network, random forest, stacking, stochastic gradient descent, and support vector machines. This approach allowed us to identify the most effective models for the dataset. Prior to evaluation, we heuristically fine-tuned the hyperparameters for each algorithm to optimize their performance.

Algorithm performance was evaluated via the Matthews correlation coefficient (MCC) [35]. The MCC metric is more suitable than other metrics for assessing binary classification models because of its robustness with imbalanced datasets and its comprehensive consideration of all four confusion matrix categories: true positives, true negatives, false positives, and false negatives. Investigations have shown that the MCC outperforms the receiver operating characteristic area under the curve (ROC AUC) for evaluating binary classification performance [36, 37]. The MCC can be interpreted as a specialized variant of the Pearson correlation coefficient adapted to binary classification problems. Although not directly related, the MCC shares a similar range of values, spanning from − 1 to 1. A score close to 0 suggests that the model performs no better than random chance. In this context, scores above 0.3 are considered moderate, and those above 0.5 are typically considered strong. The highest MCC score from the 12 algorithms was used to determine the weight of each biomarker, as described in the subsequent sections.

The mental disorder risk index (MDRI)

A key challenge in constructing a general biomarker-based risk index is defining a computationally stable and interpretable weighting structure to biomarkers. Among several possible strategies for this task, one approach relies on expert opinion or on the literature concerning mental health and biomarkers. However, a lack of consensus can hinder this approach because its clinical relevance can be subjective and may vary among experts.

Statistical methods such as principal component analysis or factor analysis are often used to identify the most important components and reduce dimensionality. Nevertheless, these techniques are constrained by assumptions about data structure, including the sufficiency of second-order statistics (variance-covariance) and the linearity of relationships among variables, which may be restrictive for heterogeneous biomedical data.

An alternative is a data-driven approach to identify the biomarkers most strongly associated with mental health outcomes. Methods such as regression analysis or machine learning algorithms can be used to identify the most important biomarkers and their corresponding weights. These algorithms can handle large datasets and detect correlations between biomarkers and mental health outcomes. However, they may overfit to training data, reducing generalizability to new observations.

To overcome the weaknesses of previous methods and develop a comprehensive index, henceforth referred to as the Mental Disorder Risk Index (MDRI), we adopted a hybrid multistep approach to determine the optimal combination of biomarkers. We developed a reliable weighting system for each biomarker in the MDRI. The optimal weights may differ depending on the clinical context, as biomarkers significant in one population may not be equally important in another. To address this variability and capture the complex relationships between biomarkers and mental health outcomes, we propose a general MDRI mathematical formulation that allows for flexible weighting and potential adjustments, given by:

graphic file with name d33e652.gif 3

where Inline graphic are the biomarker variables as defined in Eq. (2), with Inline graphic in the present work. The coefficients Inline graphic denote the respective weights, and Inline graphic represent exponents accounting for nonlinear relationships between the biomarkers and the outcomes. The independent term Inline graphic allows for a baseline offset in the index.

The inclusion of nonlinear terms through exponentiation introduces variable sensitivity scaling per biomarker, allowing the model to weight each variable’s contribution non-uniformly across its normalized range. Rather than assuming linear additivity, the formulation allows for monotonic but non-uniform transformations of individual variables, enabling the model to approximate a broader class of functional relationships. This choice is not meant to imply a specific biological mechanism, but to provide a parsimonious means of capturing potential nonlinear effects within a low-dimensional parametric structure.

Estimation of unknown parameters

Given the flexibility of the proposed formulations, parameter identifiability is not guaranteed in a strict statistical sense. Multiple parameter configurations may yield comparable objective values, particularly in the presence of correlated or weakly informative variables. For this reason, the estimated parameters should not be interpreted as uniquely determined or directly associated with specific biological effects. Instead, they are regarded as elements of an equivalence class of solutions that produce similar aggregate behavior at the level of the index.

In regression analysis, we aim to identify the function that most accurately represents the given data by minimizing estimation error. To achieve this, common error estimates can be employed, including the mean absolute error (MAE) and mean squared error (MSE), among others. Each metric varies in strengths and weaknesses, depending on dataset characteristics, such as outliers or bias tendencies toward underestimation or overestimation. To avoid favoring any particular error measurement, we adopted an ensemble error scoring approach to optimize the fitting procedure. Specifically, we selected a combination of the metrics median absolute deviation (MedAD), smooth MAE (SMAE) [38], and a novel loss function introduced in the present work, named Power Root Mean Log Square Error (PRMLSE). The respective formulas are defined as follows:

graphic file with name d33e693.gif 4
graphic file with name d33e697.gif 5

and

graphic file with name d33e704.gif 6

in which  y denotes the binarized mental health label, assigned under the relevant scheme, Inline graphic denotes the scalar MDRI output produced by Eq. (3) for the same observation, and Inline graphic is a small regularization constant introduced to prevent the logarithm from being evaluated at zero.

We implemented an ensemble approach minimizing the arithmetic mean of these three metrics. This ensemble method is designed to not prioritize any individual metric and to balance their contributions, reducing individual biases to some extent, as follows:

graphic file with name d33e730.gif 7

The objective function Eq. (7) is defined as a composite of multiple error metrics, each capturing distinct aspects of the discrepancy between predicted and observed values.

The PRMLSE metric holds a scale-symmetric sensitivity profile, growing linearly (analogous to MAE) for residuals exceeding unity, and hyperbolically for residuals below unity. Within the ensemble objective of Eq. (7), this structure introduces a complementary optimization signal that counterbalances the tendency of MedAD and SMAE to drive all residuals toward zero: the joint minimization reflects a multi-criteria balance rather than a single-objective accuracy measure. Accordingly, the PRMLSE component should not be interpreted as a standalone accuracy metric but as an auxiliary structural term whose influence is modulated by the two direct error measures in the ensemble. This formulation is not intended to replace established loss functions, but to provide a complementary penalty component within a multi-criteria evaluation framework tailored to the present exploratory setting.

The use of multiple components reflects the absence of a single metric capable of fully characterizing model performance in this context. By aggregating complementary measures, the formulation aims to balance sensitivity to central tendency, dispersion, and relative deviations, although without imposing a formal probabilistic interpretation.

The optimal method for determining the weighting coefficients (Inline graphic) and exponents (Inline graphic) in Eq. (3) is not known a priori. Therefore, various alternatives should be implemented to establish one that best fits the dataset. Hence, we implemented and evaluated five variations of the general MDRI formula given by Eq. (3) to determine the best fit for each mental health outcome. These function trials were denoted as Inline graphic, in which Inline graphic and Inline graphic represent depression, anxiety, stress, and PTSD, respectively.

We employed Eq. (3) to investigate both linear (by setting Inline graphic) and nonlinear relationships between the biomarkers (Inline graphic) and the four target variables ( C). Specifically, Inline graphic are linear versions of Eq. (3), and Inline graphic are nonlinear versions.

Both linear ( k=1) and nonlinear (Inline graphic) analyses apply 20 metaheuristic algorithms, including ant colony-based optimization (ACOR), atom search optimization (ASO), biogeography-based optimization (BBO), the circle search algorithm (CircleSA), the culture algorithm (CA), differential evolution (JADE), electromagnetic field optimization (EFO), elephant herd optimization (EHO), the gain sharing knowledge-based algorithm (GSKA), Gaussian simulated annealing (GaussianSA), dual annealing (DA), harmony search (HS), hill climbing (HC), the improved queuing search algorithm (QSA), the Levy-flight Jaya algorithm (LJA), the memetic algorithm (MA), moth-flame optimization (MFO), particle swarm optimization (PSO), adaptive inertial weighted particle swarm optimization (AIW-PSO), and enhanced warning optimization (ETWO). The selection of these metaheuristic algorithms was partially based on Kůdela’s work [39]. Each method yields a fitness value computed via Eq. (7), and the coefficients (Inline graphic) and exponents (Inline graphic) from the function with the lowest fitness values are selected for Inline graphic in Inline graphic.

The use of multiple metaheuristic algorithms is intended to explore the structure of the optimization landscape rather than to identify a single globally optimal solution, serving as a robustness check to ensure that the resulting scalar mapping is not an artifact of a specific optimization trajectory, but a stable feature of the biomarker-index manifold.

Given the potential non-convexity of the objective function, different algorithms may converge to distinct local minima, providing insight into the variability of the parameter estimation process. Accordingly, the diversity of optimization strategies is employed as an exploratory framework to assess the stability of the formulation across different search dynamics.

The lowest quartile of the 20 fitness values was subsequently derived from both linear ( k=1) and nonlinear ( k=2) models. The medians of the corresponding coefficients and, where applicable, exponents were calculated, providing initial estimates for further deterministic curve-fitting regression to compute Inline graphic for Inline graphic. We perform deterministic regressions via five methods: DIviding RECTangles (DIRECT), simple homology global optimization (SHGO), Bayesian-Gaussian (BG), hybrid adaptive Lipschitzian optimization (HALO), and Powell’s hybrid (PH).

Finally, we applied the previous 12 algorithms from the machine learning approach, selecting the algorithm that yielded the highest MCC score in the classification of mental health-related biomarkers. Afterwards, we exploited the permutation feature importance (PFI) technique to quantify the impact of each biomarker on the overall MCC score for the selected algorithm. A biomarker was considered “important” if randomly permuting its values across data points significantly reduced the model’s predictive performance. Conversely, a biomarker is considered “unimportant” if such permutation has minimal or no impact on performance. By measuring the decay in the model’s predictive performance after a biomarker’s values are randomly shuffled, the PFI reveals the extent to which the model depends on the biomarker. The PFI is quantified as the difference between the model’s performance and its performance after permutation. The importance values were used as weights in an additional linear formulation of Eq. (3) for  k=5, completing the five function trials. Given that the weights were predetermined in this instance, we set Inline graphic to avoid an additional degree of freedom.

Notably, the present investigation involves an ill-conditioned mathematical problem, where small changes in the input variables can lead to substantial changes in the results. The analysis of condition numbers reveals that certain nonlinear formulations exhibit substantial ill-conditioning, indicating a high sensitivity of the solution to perturbations in the input data or parameter values. Rather than being treated solely as a numerical limitation, this behavior provides insight into the structural properties of the model, suggesting that increased functional flexibility is accompanied by reduced numerical stability. In contrast, linear formulations display significantly lower condition numbers, reflecting a trade-off between representational capacity and stability. This observation informs the interpretation of the results and motivates a cautious assessment of models with high sensitivity.

Sensitivity analyses

Once weights (Inline graphic) and exponents (Inline graphic, where applicable) are assigned to the MDRI [Eq. (3)], a sensitivity analysis aids in assessing model reliability under different conditions and identifying critical biomarkers for control or monitoring. This can help in identifying the contribution of each individual independent variable (Inline graphic) to the overall variability in the output of the mathematical model. Moreover, it can reveal potential weaknesses or biases in the model.

Local sensitivity analysis focuses on the impact of small changes in input variables around specific data points. In contrast, global sensitivity analysis inspects the overall impact of input variables over a broad range of values. In this study, we conducted global sensitivity analysis via first-order Sobol′ sensitivity indices [40, 41], which quantify the individual contribution of each input variable to output variance without interaction terms.

Global sensitivity analysis is employed to quantify the contribution of individual biomarkers to the variability of the index. This analysis plays a central role in the present framework, as it provides a model-agnostic characterization of variable influence that is independent of specific parameter realizations. In the absence of strong assumptions regarding model structure, such sensitivity measures offer a principled means of assessing the internal coherence of the formulation and identifying dominant factors within the normalized feature space.

Index validation

To further validate the predictive accuracy of Inline graphic we computed mental health scores with optimized parameters obtained from both linear and nonlinear analyses and compared them with actual scores.

Validation was performed under a strict training–holdout protocol. The patient dataset was randomly partitioned once into an 80% training portion and a 20% test portion, stratified by severity label to preserve class proportions. The healthy control group was never used for training and was treated as a fully independent evaluation set. All model-selection steps were confined to the training portion only:

  • Hyperparameter tuning for the 12 machine learning algorithms,

  • Selection of the best metaheuristic and deterministic solutions by minimizing Eq. (7),

  • Computation of permutation feature importance for the ML-based index,

  • Derivation of the percentile thresholds (25th, 50th, 75th) used to binarize Inline graphic outputs.

No information from the 20% test subset or from the control group was used at any stage of fitting, tuning, or selection.

For evaluation, each finalized index was applied without retraining to the 20% patient test subset and the control group. Predictions were binarized using the training-derived thresholds. For each mental health condition and for both binarization schemes (1 × 3: any severity vs. none; 2 × 2: moderate-to-severe vs. none/mild), the MCC was computed from the four counts (TP, TN, FP, FN) using the standard definition [35].

The same procedure was applied to the control group: control subjects received a predicted class from each fixed index, and MCC was calculated against their known health status. This yields a direct measure of specificity under distribution shift, without any parameter adjustment on controls [35].

Composite index formulation

The five individual indices exhibit distinct response patterns across severity levels: some demonstrate stronger performance at low severity, others at high severity, and some at intermediate severity. To exploit this complementarity, a composite index was constructed by weighting and combining the outputs of all five indices into an aggregated scalar representation.

A dynamic weighting strategy adapted the contribution of each index on the basis of its class-conditional performance on the training set, such that indices with stronger class-specific discriminative contribution receive proportionally higher weight, and indices performing at or below chance level (i.e., MCC Inline graphic 0) were assigned a weight of zero to avoid contributing to the final prediction.

Ensemble methods, such as voting systems, combine multiple models with the objective of aggregating complementary classification behaviors into a more comprehensive representation. A voting system pools multiple models by having them “vote” on the final classification. This method is particularly useful when each index has unique strengths, capitalizing on their collective power. There are two voting systems:

  • Majority Voting: Each index predicts a category for a new input, and the final prediction is the category that receives the most “votes” across the indices.

  • Weighted Voting: Weights are assigned to each index based on its performance in predicting specific categories. That is, the “vote” is multiplied by the weight of that index. The category with the highest weighted vote is selected.

In this work, we employed the weighted voting system to ensure that indices with stronger class-conditional performance contribute more to the final aggregated score, with indices performing at or below chance level contributing zero weight. The weighted vote for each category Inline graphic was computed for each subject (observation) as:

graphic file with name d33e987.gif 8

in which the weights Inline graphic are derived from the classification performance of index Inline graphic on the training set. Specifically, for each index  k and each condition  C, the weight Inline graphic was computed as follows: the full training set was partitioned into the subset of observations for which index Inline graphic predicted class  j, and the subset for which it predicted the complementary class. Within the first subset, observations correctly predicted as class  j constitute true positives for that class, and those incorrectly assigned constitute false positives; the remaining observations contribute as false negatives. A class-specific MCC was then computed from the resulting counts, and this value was assigned as Inline graphic. This procedure yields a separate weight for each class Inline graphic and each index  k, capturing the class-specific discriminative contribution of that index on the training data. The Inline graphic operator ensures that indices whose class-specific contribution is negative (i.e., performing worse than chance for that class) do not invert the vote, contributing zero weight instead. The Inline graphic is an indicator function (analogous to the Kronecker delta symbol) that equals 1 for index Inline graphic predicting class  j and 0 otherwise. The composite prediction, named Inline graphic, was the class with the highest weighted score across all classes, i.e.,

graphic file with name d33e1059.gif 9

The Inline graphic function identifies the index of the maximum element in sequence  z, i.e., Inline graphic, where Inline graphic.

The performance of the composite index was subsequently compared with that of the individual indices to characterize the behavior of the aggregated representation relative to its components within the same evaluation framework.

Composite weights Inline graphic in Eq. (8) are not fitted on test data. For each index  k and each class  C, Inline graphic was set equal to the class-specific MCC obtained on the training portion only, truncated at zero by the Inline graphic operator. Indices performing at or below chance on the training data for a given class therefore contribute zero weight for that class. The composite prediction is thus a training-derived, weighted vote, and its subsequent MCC values on the test subset and on controls reflect internal consistency rather than independent validation.

Results and discussion

Mental disorders are complex, multifactorial health conditions influenced by biological, psychological, and environmental factors. Beyond their etiological complexity, these disorders are phenotypically heterogeneous and exhibit significant symptomatic overlap, posing substantial diagnostic challenges.

The present results are interpreted within the proof-of-principle scope, primarily from a computational and structural perspective. In this context, the primary contribution of this study lies in the formalization of an integrative risk-modeling architecture designed to process high-dimensional, heterogeneous biological signals within a unified parametric structure. Rather than targeting a specific disorder or predefined biomarker set, the proposed MDRI framework provides a methodological strategy for synthesizing heterogeneous biomarker profiles into a scalar representation, whose clinical relevance remains to be established through future validation on larger, independent cohorts [8, 9]. The framework is designed to operate on non-Gaussian, heterogeneous biomarker data by combining deterministic and metaheuristic optimization, nonlinear modeling, and ensemble aggregation within a single parametric structure, addressing the methodological challenges inherent to this class of biomedical data.

Study population and biomarker profiles

The characteristics of the study population, including demographic data and clinical profiles of the biomarkers, are presented in Table 1.

Table 1.

Clinical and demographic characteristics of the study sample, categorized by mental health patient group and healthy controls

Characteristic Control group (Inline graphic) Patient Group (Inline graphic)
Age 34.34 Inline graphic 11.07 [18–65] 41.65 Inline graphic 11.48 [19–68]
Gender Inline graphic (%)
Male 40 (32.79%) 99 (74.44%)
Female 82 (67.21%) 34 (25.56%)
Racial Identification Inline graphic (%)
White 77 (63.11%) 67 (50.38%)
Black 43 (35.25%) 64 (48.12%)
Indigenous 0 (0%) 1 (0.75%)
Other 2 (1.64%) 1 (0.75%)
Biomarker Mean ± SD [min.–max.]
Cortisol 424.93 Inline graphic 211.6 [9.61–1000.00] 551.94 Inline graphic 225.99 [68.18–1000.00]
CRP 4.1 Inline graphic 4.99 [0.5–24.38] 8.51 Inline graphic 8.16 [0.50–37.70]
HsCRP 4.1 Inline graphic 5.01 [0.5–20.16] 4.80 Inline graphic 6.28 [0.10–37.26]
Vitamin D 34.93 Inline graphic 14.96 [8.7–79.2] 33.61 Inline graphic 16.04 [9.90–95.90]
IL-6 2.2 Inline graphic 0.44 [2–4.27] 4.77 Inline graphic 3.25 [2.00–19.50]
IL-1Inline graphic 6.2 Inline graphic 2.53 [5–16.3] 8.20 Inline graphic 7.01 [5.00–39.40]
TNF-Inline graphic 10.18 Inline graphic 2.94 [4–22.5] 9.96 Inline graphic 4.42 [4.00–29.50]
IL-8 11.13 Inline graphic 5.34 [5–31.1] 9.37 Inline graphic 5.26 [5.00–31.90]

*Values are presented as the mean Inline graphic standard deviation [minimum–maximum] where applicable

Table 1 reveals differences between the control and patient groups. The patient group was older on average and exhibited an imbalance in sex distribution, with a predominance of male participants, whereas the control group was predominantly female. These demographic disparities may serve as confounding factors, particularly for biomarkers sensitive to age and sex.

Regarding biomarker profile, the patient group exhibited higher mean values for inflammatory markers such as CRP and IL-6, suggesting a potential association between systemic inflammation and mental health conditions within the studied population. However, variability across biomarkers indicates a heterogeneous biological profile, reinforcing the need for integrative approaches capable of capturing multivariate relationships.

Biomarker relationships and independence structure

The heterogeneous biomarker patterns observed in Table 1 indicate that mental health conditions are not characterized by uniform biological alterations. Instead, distinct biomarkers exhibit variable behavior across individuals, suggesting the need for approaches capable of capturing multivariate relationships among biomarkers.

This methodological flexibility is particularly relevant in the current context, in which factors such as chronic stress and low vitamin D levels have been associated with alterations in neuroendocrine and immune regulation, including prolonged exposure to elevated cortisol levels and reduced modulatory effects of vitamin D on immunity [12, 42, 43]. Within the present framework, biomarker directionality is not assumed a priori; instead, the weighting and sign of each variable’s contribution are determined empirically through optimization, accommodating both elevated and attenuated signals.

Scientific evidence has shown that inflammation may be involved in the pathophysiology of psychiatric disorders. Dysregulation of markers such as CRP, IL-6, IL-1β, and TNF-α has been associated with inflammatory and oxidative stress mechanisms in conditions such as depression, schizophrenia, and bipolar disorder [11, 44–47]. In the present dataset, the patient group exhibited higher mean concentrations of cortisol, CRP, hsCRP, IL-6, and IL-1Inline graphic compared to controls, whereas TNF-Inline graphic and IL-8 levels were marginally lower and vitamin D levels were comparable. These divergences are consistent with the biological heterogeneity of psychiatric conditions, in which the direction of cytokine dysregulation may vary across disorder subtypes and clinical stages.

The multi-biomarker strategy employed in this study to construct the MDRI is aligned with the context outlined by Schmidt et al. [48], who highlight the limitations of single-biomarker approaches and advocate for multiplex biomarker panels to improve diagnostic sensitivity and specificity. Our methodology reflects this approach by combining machine learning, metaheuristic optimization, and regression modeling to weight and aggregate biomarkers, including cortisol, hsCRP, IL-6, and vitamin D. This convergence supports the translational potential of composite biomarker panels by addressing the limitations of single-marker strategies for precision diagnostics in mental health. In this context, the MDRI represents a preliminary instantiation of a unified quantitative model that integrates systemic physiological information, whose ability to capture biologically relevant complexity in mental health still needs to be evaluated in larger and more rigorously controlled studies.

To investigate the nonlinear relationships among the eight biomarkers used in this study, Hoeffding’s D sample independence test was applied, generating a matrix as shown in Fig. 2.

Fig. 2.

Fig. 2

Hoeffding’s D sample independence test matrix visualizing nonlinear correlations between biomarkers. The color intensity represents the correlation magnitude

Each cell represents a statistical measure of the dependence between two biomarkers, with color intensity indicating the magnitude of the correlation. All biomarkers, except CRP and hsCRP, exhibited negligible mutual dependencies, indicating minimal or nonexistent correlations. The weak correlation between CRP and hsCRP is expected, as hsCRP corresponds to a high-sensitivity measurement of the same protein assessed by the standard CRP test. The absence of correlation among the remaining biomarkers suggests that each variable provides complementary information about the physiological state of the patients. Consequently, all biomarkers were retained in the formulation of the MDRI.

Machine learning performance and model variability

To evaluate the predictive performance of different modeling strategies, multiple machine learning algorithms were applied to clarify mental health condition based on biomarker data. In the data processing, mental health scores for anxiety, depression, and stress were binarized into 1 × 3 and 2 × 2 schemes, described in detail in the Methods section. PTSD scores are inherently binary. Although both binarization schemes were considered throughout the analysis, representative results are presented here for clarity.

Thus, to increase the confidence in the results, multiple-input regression analyses were performed via metaheuristic, deterministic, and machine learning methods, yielding five indices. For example, to obtain Inline graphic (where superscript  C represents one of four mental health conditions—depression, anxiety, stress, or PTSD—and subscript  5 indicates the machine learning approach), 12 machine learning classification algorithms were employed for each of the four mental health conditions. The five indices were validated using the remaining 20% of the dataset and the entirety of the control group. Each index produces a numerical value within a determined range, which is stratified into four proportional subdivisions based on the reference scores established for each evaluated mental disorder. These subdivisions correspond to the following clinical severity categories: absent, mild, moderate, and severe. The value generated by each index is classified into one of these categories, predicting the manifestation or no manifestation of the psychiatric condition.

As a representative example of the predictive efficacy of each algorithm, Table 2 lists the MCC values for each mental health disorder under the 1 × 3 scheme. Higher MCC values indicate better classification performance of the respective mental health disorder. The highest MCC value for each disorder is highlighted in bold. This table emphasizes the importance of testing multiple algorithms, each with distinct strengths and characteristics, since the most effective algorithm for a given dataset cannot be known a priori.

Table 2.

Comparative performance of 12 machine learning algorithms in the 1 × 3 scheme, evaluated on the held-out test set (20% of the patient dataset), in predicting each mental health disorder. The best algorithm performances are highlighted in bold

Model MCC
Anxiety Depression Stress PTSD
AdaBoost 0.18 0.01 0.14 –0.14
CN2 Rule Induction 0.38 0.22 0.49 0.21
Decision Tree 0.22 0.21 0.40 0.25
Gradient Boosting 0.35 0.40 0.40 0.13
k-Nearest Neighbors 0.24 0.55 0.46 0.21
Logistic Regression 0.38 0.36 0.37 0.15
Naïve Bayes 0.15 0.25 0.19 –0.01
Neural Network 0.44 0.44 0.28 0.25
Random Forest 0.28 0.47 0.49 0.27
Stacking 0.21 0.39 0.28 0.10
Stochastic Gradient Descent 0.36 0.42 0.40 0.15
Support Vector Machines 0.27 0.37 0.40 0.34

The results in Table 2 demonstrate variability in model performance across different mental health conditions. No single algorithm consistently outperformed the others across all disorders, highlighting the absence of a universally optimal model for this dataset. For example, the CN2 rule induction algorithm achieved the highest highest MCC for stress (0.49), whereas k-nearest neighbors performed best for depression (0.55), and support vector machines showed the strongest performance for PTSD (0.34).

This variability reflects differences in how each algorithm captures underlying patterns in the biomarker space and reinforces the importance of evaluating multiple modeling approaches. These findings support the use of a multi-model framework, in which complementary strengths of different algorithms can be leveraged to improve robustness and generalizability. Within this context, the MDRI integrates information derived from distinct modeling strategies, contributing to the construction of a more stable and interpretable scalar representation.

Biomarker importance and interpretability analyses

The analyses of biomarkers importance revealed consistent patterns across complementary approaches. Figure 3a and b provide complementary perspectives on biomarker relevance within the model.

Fig. 3.

Fig. 3

Biomarker importance. Panel (a) shows the decrease in model performance (MCC) for the CN2 rule induction model for predicting stress in the 1 × 3 scheme. Panel (b) shows the impact of each biomarker on the model output. Points represent the impact of a single feature value on the model prediction. Red points indicate high impact, and blue points indicate low impact

Figure 3a depicts PFI, in which biomarkers values were randomly permuted across 100 cycles of permutations, and the corresponding decrease in MCC was computed. The final importance value of biomarkers (blue bars in Fig. 3a) is the average of these 100 values, with black lines (Fig. 3a) indicating the standard deviation. Biomarkers are ranked vertically from most to least relevant. The horizontal axis represents the decrease in the MCC when a biomarker is randomized. Biomarkers with greater decreases in MCCs are more important for model predictions, i.e., changes in the values of these biomarkers cause greater decreases in the respective MCCs, impacting model predictability. These importance values were used as weights in construction of the stress disorder index Inline graphic. Based on this analysis, TNF-Inline graphic and IL-6 have emerged as the most influential biomarkers in the model, whereas IL-8 shows negligible contribution. These importance values were subsequently used as weights in the construction of the stress-related MDRI variant (Inline graphic). Figure 3b shows the SHapley Additive exPlanations (SHAP) analyses for the CN2 rule induction model, offering a complementary interpretation of biomarker contributions. SHAP values, derived from a game theory framework, quantify the contribution of each biomarker to individual predictions, allowing both the magnitude and direction of influence to be assessed. In the SHAP summary plot, each point represents the contribution of a specific biomarker value on the model’s prediction, with color indicating the relative magnitude of the biomarker (low to high values). Red points indicate high impact, and blue points indicate low impact. The distance of the point from the central zero line denotes the magnitude of the impact. Biomarkers are ordered by the mean absolute SHAP value, reflecting their overall importance to the model.

The joint interpretation of Fig. 3a and b reveals consistent patterns of biomarker influence. For example, vitamin D demonstrates substantial importance in Fig. 3a, and the SHAP analysis indicates that lower vitamin D levels are associated with higher model outputs, whereas moderate to higher concentrations are associated with reduced outputs. Similarly, TNF-Inline graphic exhibits a strong positive association with model predictions, with higher concentrations contributing to increased model outputs.

A comparative analysis between the PFI and SHAP approaches shows that the most and least influential biomarkers are largely consistent across both methods, although their rankings differ. This agreement between independent interpretability techniques increases confidence in the robustness of the observed feature importance patterns.

Sensitivity analyses and structural consistency

Figure 4 presents the first-order Sobol′ sensitivity indices for the eight blood biomarkers in Inline graphic, which are calculated for depression (a), anxiety (b), stress (c) under the 1 × 3 scheme, and PTSD (d). The Sobol′ index quantifies the proportion of variance in Inline graphic attributable to each biomarker, with higher values indicating greater influence on the model output. Overall, the results show that the relative importance of biomarkers varies depending on the specific mental health condition, indicating the presence of disorder-specific sensitivity patterns. TNF-Inline graphic consistently shows high importance across depression, stress, and PTSD, whereas hsCRP and IL-6 show more condition-dependent relevance, with hsCRP being particularly prominent in anxiety and IL-6 in stress.

Fig. 4.

Fig. 4

First-order Sobol′ sensitivity indices for depression (a), anxiety (b), stress (c), and PTSD (d), quantifying the fractional contribution of each individual biomarker to the total variance of the Inline graphic output

The results also reveal distinct disorder-specific signatures. For example, hsCRP shows a unique contribution in in anxiety (Fig. 4), whereas vitamin D demonstrates relevance in depression and stress. Cortisol also presents notable importance in depression and PTSD predictions. In contrast, CRP and hsCRP show minimal influence across most condition, except for hsCRP in anxiety. Similarly, IL-1Inline graphic and IL-8 contribute minimally to the index outputs in most cases, with limited influence observed outside PTSD. A comparison between the stress-related results in Fig. 3 and the Sobol′ indices in Fig. 4c reveals strong consistency across analyses approaches. In particular, TNF-Inline graphic and IL-6 are consistently identified as key contributors across PFI, SHAP, and Sobol′ analyses, whereas IL-8 and IL-1Inline graphic show negligible contributions. This agreement is noteworthy given that these methods are conceptually distinct: PFI and SHAP are model-based interpretability techniques derived from machine learning, whereas the Sobol′ index is a variance-based global sensitivity analysis method. The alignment of these independent approaches reinforces the structural coherence of the biomarker importance patters observed in the dataset and supports the robustness of the MDRI formulation. Importantly, these sensitivity and importance analyses (PFI, SHAP, and Sobol′) reflect model-derived associations within the normalized biomarker space, and should be interpreted as indicative of patterns captured by the model, rather than as direct causal biological relationships.

MDRI variants and distributional behavior

Twenty-five independent indices (linear and nonlinear) were generated for each of the 20 metaheuristic and five deterministic methods. Additionally, a linear index was constructed using PFI values from 12 machine learning algorithms. The best-performing index for the metaheuristic and deterministic sets was selected based on the lowest fitness value [Eq. (7)], while the best machine learning index was chosen based on the their highest MCC. This process yields five final indices for each mental health condition.

Figure 5 presents a representative violin plot, which is a type of density plot, illustrating the distribution of PTSD risk indices. The plot depicts the density and dispersion of values for each index, with the width of each violin reflecting the concentration of observations and embedded points reflecting individual data values. The yellow markers indicate the median values for each distribution. The distribution revels distinct behavioral pattern across the indices. Inline graphic (blue) has the highest median values, with a concentration toward the upper range (closer to 1), suggesting a tendency to classify cases as high risk and indicating increased sensitivity to PTSD risk. In contrast, Inline graphic (orange) is centered around lower values (Inline graphic0.4), indicating a tendency to assign lower average risk for PTSD. The Inline graphic (green) shows a more balanced distribution, with peak density in the intermediate range (0.5–0.6), suggesting more sensitivity. The Inline graphic (red) presents a vertically compressed distribution, indicating low variability in predictions and reduced sensitivity to variations in PTSD risk. The Inline graphic (purple) demonstrates a broader distribution with a pronounced central peak, suggesting higher variability and potentially reduced specificity. The composite index Inline graphic integrates the outputs of the five individual indices and exhibits a multipeaked distribution, reflecting its aggregated nature. This pattern indicates that the composite index captures a broader spectrum of risk profiles, combining predictive behaviors into a unified representation (depicted as a thick-edged line in the plot).

Fig. 5.

Fig. 5

Violin plot of kernel density estimates (KDEs) comparing the score distributions of five PTSD indices (Inline graphic) and the composite index (Inline graphic). The width of each violin reflects the density of observations across index values, with embedded points representing individual samples and yellow markers indicating median values. Differences in distribution shape, central tendency, and dispersion highlight variability in predictive behavior among individual indices, while the composite index exhibits a broader and multipeaked structure, reflecting its integrative aggregation of diverse risk patterns

Note that the five indices exhibit negligible mutual statistical dependence, as evidenced by near-zero Hoeffding’s D values in Fig. 6, supporting the rationale for their aggregation into a composite index.

Fig. 6.

Fig. 6

Hoeffding’s D sample independence test matrices visualizing nonlinear relationships between the five anxiety indices in the 1×3 scheme, evaluated on the 20% patient dataset (a) and the control group (b). Panels (c) and (d) show the 2×2 scheme evaluated in both groups. Color intensity represents the magnitude of dependence

Independence of indices and rationale for aggregation

Figure 6 illustrates the correlation matrices between the indices for anxiety under the 1 × 3 (a) and 2 × 2 (b) schemes, evaluated on 20% of the dataset (c) and the control group (d), using Hoeffding’s D sample independence test. The correlation coefficients are notably close to 0, indicating negligible interindex correlation.

The near-zero Hoeffding’s D values among the indices are consistent with statistical independence, suggesting that each index captures distinct aspects of the biomarker space. This supports the use of independent validation procedures for each index and provides a rationale for their subsequent integration into composite formulation.

Comparative performance of individual and composite indices

Table 3 summarizes the performance of the individual indices (Inline graphic and the composite index Inline graphic in detecting depression, stress, anxiety, and PTSD, evaluated using MCC values. The indices were evaluated on two groups: a held-out subset of the patient dataset (20% retained for testing) and the control group, under both binarization schemes (1 × 3 and 2 × 2). Overall, the individual indices exhibit variable and generally low performance across conditions, with MCC values fluctuating across datasets and schemes. No single index consistently achieved high performance across all conditions, indicating that individual formulations are limited in their ability to capture the full spectrum of mental health severity. However, individual indices performed better in specific conditions or datasets, indicating that each index captures different patterns in the biomarker data.

Table 3.

Performance of individual (Inline graphic) and composite (Inline graphic) indices in predicting mental health risk, evaluate using MCC values under the 1 × 3 and 2 × 2 schemes in both test (20% patient subset) and control groups

Method Index Scheme 1 × 3 Scheme 2 × 2
Test (Dataset) Test (Control) Test (Dataset) Test (Control)
MCC for Depression
Metaheuristic Inline graphic 0.071 –0.036 0.053 –0.093
Inline graphic 0.189 –0.046 0.316 –0.206
Curve fitting Inline graphic 0.189 –0.115 0.135 –0.135
Inline graphic 0.225 –0.187 0.271 0.000
ML Inline graphic –0.337 0.114 –0.105 0.011
Composite Inline graphic 0.756 0.725 0.848 0.957
MCC for Stress
Metaheuristic Inline graphic 0.195 0.036 –0.250 –0.026
Inline graphic –0.107 –0.015 0.139 0.086
Curve fitting Inline graphic 0.218 –0.032 –0.222 –0.057
Inline graphic 0.015 0.036 0.127 –0.008
ML Inline graphic –0.104 –0.168 0.030 0.044
Composite Inline graphic 0.646 0.623 0.602 0.654
MCC for Anxiety
Metaheuristic Inline graphic 0.091 0.053 0.018 0.045
Inline graphic 0.301 0.043 –0.255 –0.063
Curve fitting Inline graphic 0.208 0.050 0.317 0.079
Inline graphic 0.138 –0.181 0.108 –0.030
ML Inline graphic –0.467 0.034 –0.269 0.102
Composite Inline graphic 0.846 0.847 0.892 0.964
MCC for PTSD
Metaheuristic Inline graphic 0.012 0.010 0.012 0.043
Inline graphic 0.125 0.038 –0.279 0.070
Curve fitting Inline graphic –0.071 –0.028 –0.071 –0.028
Inline graphic –0.006 0.094 –0.145 0.094
ML Inline graphic 0.104 –0.018 –0.006 0.035
Composite Inline graphic 1.000 0.961 1.000 0.980

The subscripts Inline graphic refer to metaheuristic optimization for linear and nonlinear models (shaded rows) from Eq. (3), respectively. The subscripts Inline graphic indicate deterministic curve fitting for the linear and nonlinear models, respectively. The subscript  k=5 denotes the ML-based model

Comparing linear (Inline graphic) and nonlinear (Inline graphic models did not reveal a clear advantage for nonlinear models. Although nonlinear models achieved slightly higher MCC values in 17 out of 32 comparisons (53%), linear models exhibited greater numerical stability, characterized by substantially lower condition numbers. This indicates reduced sensitivity to input perturbations and suggests that, within the present proof-of-principle framework, linear formulations provide a more stable representation without compromising predictive performance.

The machine learning-based index ( k=5) was superior to the metaheuristic and deterministic indices (Inline graphic) in only two instances both corresponding to depression in the control group under both schemes. In most cases, its MCC performance was comparable or slightly superior to some individual models, but without statistically significant. These results indicate that, although machine learning contributes to model construction, it does not independently resolve the variability observed across individual indices.

In contrast, the composite index (Inline graphic) produces numerically higher MCC values across all conditions, datasets, and binarization schemes. All composite MCC values exceeded 0.6, with most above 0.8 and some approaching or reaching 1.0. This marked improvement reflects the integration of complementary predictive behaviors across individual indices, a phenomenon consistent with ensemble theory where combining weakly correlated predictors yields improved aggregate performance. As demonstrated in Fig. 6, the individual indices exhibit near-zero mutual correlations, supporting the validity of their aggregation. The weighted voting scheme amplifies concordant predictions while suppressing inconsistent contributions, resulting in enhanced overall performance [49]. However, it is important to emphasize that the composite MCC values reflect the internal consistency within the dataset rather than generalizable predictive performance. The weighting coefficients Inline graphicwere derived from the same training data used to construct the indices, and no resampling-based validation was performed. Therefore, the reported values should be interpreted descriptively and within the proof-of-principle scope of this study. In particular, the MCC values, including the PTSD result of 1.000, obtained from a relatively small sample (approximately 24–27 observations), should not be interpreted as evidence of perfect classification. No systematic differences were observed between the 1 × 3 and 2 × 2 binarization schemes or between the two test groups. Additionally, class distributions were imbalanced, as expected in severity-stratified mental health data, with positive class prevalence ranging from approximately 30% to 55% under the 1 × 3 scheme and from 15% to 35% under the 2 × 2 scheme. Given this imbalance, MCC was used as the primary evaluation metric, as incorporates all elements of the confusion matrix and provides a balance assessment of classification performance. This choice is consistent with previous studies demonstrating that MCC outperforms metrics such as ROC AUC, accuracy, and F1 score in imbalanced biomedical classification settings [36, 37]. Full confusion matrices were not reported, as individual cell counts may be misleading under class imbalance whereas MCC provides a single, balanced summary that directly reflects the contingency table.

Interpretation within the proof-of-principle scope

Confidence intervals were not computed, as the present study is explicitly framed as a proof-of-principle investigation with a restricted sample size. The 20% test subset comprises approximately 24 to 27 observations per condition, and the control group is of comparable size. Under these conditions, bootstrap or analytic interval estimates would be inherently unstable and could overstate inferential certainty. Therefore, the results are interpreted descriptively, focusing on internal consistency and structural patterns rather than formal statistical inference.

Clinical and methodological implications

The development of a unified biomarker-based index is particularly relevant given the methodological challenges associated with diagnosing mental disorders, which largely rely on subjective assessments. It is estimated that over 50% of potential depression cases remain undiagnosed, and individuals who initiate treatment often encounter unstructured and inadequate access to appropriate care [23, 50]. Unlike many medical conditions diagnosed through precise measurements, mental health diagnoses are primarily based on clinical interviews, in which professionals evaluate self-reported symptoms such as mood changes, sleep disturbances, and cognitive difficulties [51]. These symptoms are often nonspecific and may arise a wide range of factors, including psychological stressors, physiological conditions, medical treatments, or underlying health issues. This variability limits the precision and reproducibility of current diagnoses approach and highlighting the complexity of accurately identifying mental disorders, as well as the need for complementary and more objective assessment strategies [52]. In this context, blood-based biomarker analysis remains underutilized in routine clinical practice. The proposed framework contributes to this gap by offering a structured computational approach for integrating heterogeneous biomarker data into unified representations. While not intended as a diagnostic tool stage, such approaches may support future developments in risk stratification and contribute to more objective and scalable mental health assessment framework.

Conclusion

This study presented a computational framework for the formulation of a unified scalar index derived from multiple blood-based biomarkers, with the objective of examining the feasibility of constructing low-dimensional representations of complex biomedical data. The proposed approach integrates linear and nonlinear formulations, multiple optimization strategies, and model-dependent evaluation procedures within a single analytical structure.

The results indicate that it is possible to obtain coherent scalar representations that capture structured relationships within the normalized biomarker space. At the same time, the analysis reveals important trade-offs between model flexibility and numerical stability, as well as variability in parameter estimates arising from the non-convex nature of the optimization problem. These findings underscore the fact that multiple parametrizations can produce similar aggregate behavior, highlighting the non-uniqueness of the resulting index.

The primary limitations of this study are the restricted sample size, absence of cross-validation, potential data-dependent dependencies in the model aggregation procedure, and a demographic imbalance between the patient and control groups that introduces a potential confound for biomarkers sensitive to sex and age. The analysis is based on a single dataset, which may restrict the generalizability of the findings. And the absence of external validation prevents definitive conclusions regarding the robustness of the proposed framework across independent populations. In addition, this study is positioned as a proof-of-principle investigation and does not aim to establish clinical applicability. Potential confounding factors, including demographic and biological variability, were not fully controlled and may influence the observed relationships. These factors limit the scope of the present conclusions to the framework’s internal behavior within the current dataset. Within these bounds, the framework provides a testable basis for exploring composite biomarker indices and characterizes the interplay between model structure, optimization dynamics, and sensitivity properties.

Future work should focus on validating the proposed formulations using larger and independent datasets, incorporating statistically grounded evaluation protocols, and refining the mathematical structure to improve stability and interpretability. Within this context, the proposed framework stablishes a formal and generalizable basis for the construction of interpretable biomarker-based risk representations.

Future directions

In this study, the results establish a computational basis for new research on integrative biomarker modeling. Future work should focus on evaluating the proposed framework using larger and independent datasets, aiming to evaluate its generalization and robustness across diverse biological and clinical conditions. In this context, the Mental Disorder Risk Index (MDRI) may be explored to examine the behavior of the framework across different populations and settings.

The incorporation of additional biomarkers and multimodal data sources, such as genetic markers, imaging data, and longitudinal clinical records, holds significant promise for enhancing the framework’s capabilities. Imaging-derived biomarkers offer rich structural and functional information, but are highly dependent on reliable medical image segmentation. Recent advances in conventional, deep learning-based, hybrid, and graph-theoretical medical image segmentation methods provide a robust foundation for such integration [53, 54].

In addition, the use of demographically matched cohort designs and covariate-adjusted models will be essential to isolate the contribution of mental health status from confounding biological and demographic factors. Finally, the application of formal cross-validation procedures and independent external cohorts will be necessary to assess the generalization of the framework beyond the present proof-of-principle context. From a translational perspective, future studies may evaluate the application of the framework in different clinical contexts, including its potential integration with existing diagnostic and decision-support systems. In this sense, the framework is positioned as an auxiliary computational tool for biomarker integration, risk stratification, and longitudinal analysis in clinical research.

Acknowledgements

Not applicable.

Author contributions

F. O. D. and M. M.: Conceptualization, methodology, investigation, data collection, analysis, writing of the main manuscript text, and review and editing. L. C., J. B., K. F. G., B. D. L. F., and J. M. A. R.: Methodology, data collection, writing – review. J. A. P.: Writing – review. C. S. PETROBRAS Technical representative, resources, review. F. F. A.: Supervision, resources, funding acquisition, methodology, writing – review and editing. All the authors have read and approved the final manuscript.

Funding

This work was supported by PETROBRAS, Brazil (grant number 2022/00220-0). B. D. L. F. was funded by the São Paulo Research Foundation (FAPESP), Brazil (process numbers 2013/07296-2 and 2022/06210-6).

Data availability

The datasets generated and/or analyzed during the current study are not publicly available due to privacy restrictions but are available from the corresponding author upon reasonable request.

Declarations

Ethics approval and consent to participate

This study was approved by the Human Research Ethics Committee (CEP) of the Federal University of São Carlos (UFSCar) under registration number CAAE: 54381621.0.0000.5504, in accordance with ethical guidelines for research involving human participants. All volunteers provided written informed consent.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Fernanda Oliveira Duarte, Email: fefa.duarte74@gmail.com.

Mauro Masili, Email: mmasili@gmail.com.

Fernanda de Freitas Aníbal, Email: ffanibal@ufscar.br.

References

  • 1.Verhoeven JE, Wolkowitz OM, Barr Satz I, Conklin Q, Lamers F, Lavebratt C, et al. The researcher’s guide to selecting biomarkers in mental health studies. BioEssays. 2024;46(3):e2300246. 10.1002/bies.202300246. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Zhao Q, Zhang C, Zhang W, Zhang S, Liu Q, Guo Y. Applications and challenges of biomarker-based predictive models in proactive health management. Front Public Health. 2025;13:1633487. 10.3389/fpubh.2025.1633487. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat Mach Intell. 2019;1(5):206–15. 10.1038/s42256-019-0048-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Ciobanu-Caraus O, Aicher A, Kernbach JM, Regli L, Serra C, Staartjes VE. A critical moment in machine learning in medicine: on reproducible and interpretable learning. Acta Neurochir. 2024;166:14. 10.1007/s00701-024-05892-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.McDermott JE, Wang J, Mitchell H, Webb-Robertson BJ, Hafen R, Ramey J, et al. Challenges in biomarker discovery: combining expert insights with statistical analysis of complex omics data. Expert Opin Med Diagn. 2013;7(1):37–51. 10.1517/17530059.2012.718329. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Emmerich CH, Gamboa LM, Hofmann MCJ, Bonin-Andresen M, Arbach O, Schendel P, et al. Improving target assessment in biomedical research: the GOT-IT recommendations. Nat Rev Drug Discov. 2021;20:64–81. 10.1038/s41573-020-0087-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Climente-González H, Oh M, Chajewska U, Hosseini R, Mukherjee S, Gan W, et al. Interpretable machine learning leverages proteomics to improve cardiovascular disease risk prediction and biomarker identification. Commun Med. 2025;5:170. 10.1038/s43856-025-00872-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Wu Y, Wang L, Tao M, Cao H, Yuan H, Ye M, et al. Changing trends in the global burden of mental disorders from 1990 to 2019 and predicted levels in 25 years. Epidemiol Psychiatr Sci. 2023;32:e63. 10.1017/S2045796023000756. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Levchenko A, Nurgaliev T, Kanapin A, Samsonova A, Gainetdinov RR. Current challenges and possible future developments in personalized psychiatry with an emphasis on psychotic disorders. Heliyon. 2020;6(5):e03990. 10.1016/j.heliyon.2020.e03990. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Djuric Z, Bird CE, Furumoto-Dawson A, Rauscher GH, Ruffin MT 4th, Stowe RP, et al. Biomarkers of psychological stress in health disparities research. Open Biomark J. 2008;1:7–19. 10.2174/1875318300801010007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Janelidze S, Mattei D, Westrin Å, Träskman-Bendz L, Brundin L. Cytokine levels in the blood may distinguish suicide attempters from depressed patients. Brain Behav Immun. 2011;25(2):335–9. 10.1016/j.bbi.2010.10.010. [DOI] [PubMed] [Google Scholar]
  • 12.Berk M, Williams LJ, Jacka FN, O’Neil A, Pasco JA, Moylan S, et al. So depression is an inflammatory disease, but where does the inflammation come from? BMC Med. 2013;11:200. 10.1186/1741-7015-11-200. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Serafini G, Pompili M, Elena Seretti M, Stefani H, Palermo M, Coryell W, et al. The role of inflammatory cytokines in suicidal behavior: a systematic review. Eur Neuropsychopharmacol. 2013;23(12):1672–86. 10.1016/j.euroneuro.2013.06.002. [DOI] [PubMed] [Google Scholar]
  • 14.Zhang C, Wu Z, Zhao G, Wang F, Fang Y. Identification of IL6 as a susceptibility gene for major depressive disorder. Sci Rep. 2016;6:31264. 10.1038/srep31264. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Köhler CA, Freitas TH, Maes M, de Andrade NQ, Liu CS, Fernandes BS, et al. Peripheral cytokine and chemokine alterations in depression: a meta-analysis of 82 studies. Acta Psychiatr Scand. 2017;135(5):373–87. 10.1111/acps.12698. [DOI] [PubMed] [Google Scholar]
  • 16.Uint L, Bastos GM, Thurow HS, Borges JB, Hirata TDC, França JID, et al. Increased levels of plasma IL-1b and BDNF can predict resistant depression patients. Rev Assoc Med Bras. 2019;65(3):361–9. 10.1590/1806-9282.65.3.361. [DOI] [PubMed] [Google Scholar]
  • 17.Osimo EF, Pillinger T, Rodriguez IM, Khandaker GM, Pariante CM, Howes OD. Inflammatory markers in depression: a meta-analysis of mean differences and variability in 5,166 patients and 5,083 controls. Brain Behav Immun. 2020;87:901–9. 10.1016/j.bbi.2020.02.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Göver T, Slezak M. Targeting glucocorticoid receptor signaling pathway for treatment of stress-related brain disorders. Pharmacol Rep. 2024;76:1333–45. 10.1007/s43440-024-00654-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Sterling RK, Lissen E, Clumeck N, Sola R, Correa MC, Montaner J, et al. Development of a simple noninvasive index to predict significant fibrosis in patients with HIV/HCV coinfection. Hepatology. 2006;43(6):1317–25. 10.1002/hep.21178. [DOI] [PubMed] [Google Scholar]
  • 20.Petersen MC, Shulman GI. Mechanisms of insulin action and insulin resistance. Physiol Rev. 2018;98(4):2133–223. 10.1152/physrev.00063.2017. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Fernández-Macías JC, Ochoa-Martínez AC, Varela-Silva JA, Pérez-Maldonado IN. Atherogenic index of plasma: novel predictive biomarker for cardiovascular illnesses. Arch Med Res. 2019;50(5):285–94. 10.1016/j.arcmed.2019.08.009. [DOI] [PubMed] [Google Scholar]
  • 22.Kapur S, Phillips AG, Insel TR. Why has it taken so long for biological psychiatry to develop clinical tests and what to do about it? Mol Psychiatry. 2012;17(12):1174–9. 10.1038/mp.2012.105. [DOI] [PubMed] [Google Scholar]
  • 23.Fekadu A, Demissie M, Birhane R, Medhin G, Bitew T, Hailemariam M, et al. Under detection of depression in primary care settings in low and middle-income countries: a systematic review and meta-analysis. Syst Rev. 2022;11(1):21. 10.1186/s13643-022-01893-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Faisal-Cury A, Ziebold C, Rodrigues DMO, Matijasevich A. Depression underdiagnosis: prevalence and associated factors. A population-based study. J Psychiatr Res. 2022;151:157–65. 10.1016/j.jpsychires.2022.04.025. [DOI] [PubMed] [Google Scholar]
  • 25.Huang SM, Wu FH, Ma KJ, Wang JY. Individual and integrated indexes of inflammation predicting the risks of mental disorders - statistical analysis and artificial neural network. BMC Psychiatry. 2025;25:226. 10.1186/s12888-025-06652-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Huang J, Hou X, Li M, Xue Y, An J, Wen S, et al. A preliminary composite of blood-based biomarkers to distinguish major depressive disorder and bipolar disorder in adolescents and adults. BMC Psychiatry. 2023;23:755. 10.1186/s12888-023-05204-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Benacek J, Lawal N, Ong T, Tomasik J, Martin-Key N, Funnell E, et al. Identification of predictors of mood disorder misdiagnosis and subsequent help-seeking behavior in individuals with depressive symptoms: gradient-boosted tree machine learning approach. JMIR Ment Health. 2024;11:e50738. 10.2196/50738. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Oliveira GS, Rodrigues FHS, Pontes JGM, Tasic L. Artificial intelligence-based methods and omics for mental illness diagnosis: a review. Bioengineering. 2025;12(10):1039. 10.3390/bioengineering12101039. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Zhu T, Liu X, Wang J, Kou R, Hu Y, Yuan M, et al. Explainable machine-learning algorithms to differentiate bipolar disorder from major depressive disorder using self-reported symptoms, vital signs, and blood-based markers. Comput Methods Programs Biomed. 2023;240:107723. 10.1016/j.cmpb.2023.107723. [DOI] [PubMed] [Google Scholar]
  • 30.Vignola RCB, Tucci AM. Adaptation and validation of the depression, anxiety and stress scale (DASS) to Brazilian Portuguese. J Affect Disord. 2014;155(1):104–9. 10.1016/j.jad.2013.10.031. [DOI] [PubMed] [Google Scholar]
  • 31.Osório FL, Silva TDA, Santos RG, Chagas MHN, Chagas NMS, Sanches RF, et al. Posttraumatic stress disorder checklist for DSM-5 (PCL-5): transcultural adaptation of the Brazilian version. Arch Clin Psychiatry. 2017;44(1):10–9. 10.1590/0101-60830000000107. [DOI] [Google Scholar]
  • 32.Mahalanobis PC. Reprint of: On the generalised distance in statistics (1936). Sankhya A. 2018;80(S1):1–7. 10.1007/s13171-019-00164-5. [DOI] [Google Scholar]
  • 33.Bernett J, Blumenthal DB, Grimm DG, Haselbeck F, Joeres R, Kalinina OV, et al. Guiding questions to avoid data leakage in biological machine learning applications. Nat Methods. 2024;21(8):1444–53. 10.1038/s41592-024-02362-y. [DOI] [PubMed] [Google Scholar]
  • 34.Wolpert DH, Macready WG. No free lunch theorems for optimization. IEEE Trans Evol Comput. 1997;1(1):67–82. 10.1109/4235.585893. [DOI] [Google Scholar]
  • 35.Matthews BW. Comparison of the predicted and observed secondary structure of T4 phage lysozyme. Biochim Biophys Acta. 1975;405(2):442–51. 10.1016/0005-2795(75)90109-9. [DOI] [PubMed] [Google Scholar]
  • 36.Chicco D, Jurman G. The Matthews correlation coefficient (MCC) should replace the ROC AUC as the standard metric for assessing binary classification. Biodata Min. 2023;16(1):4. 10.1186/s13040-023-00322-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Chicco D, Jurman G. A statistical comparison between Matthews correlation coefficient (MCC), prevalence threshold, and Fowlkes–Mallows index. J Biomed Inf. 2023;144:104426. 10.1016/j.jbi.2023.104426. [DOI] [PubMed] [Google Scholar]
  • 38.Noel MM, Banerjee A, Oswal Y, Amali DG, Muthiah-Nakarajan V. Alternate loss functions for classification and robust regression can improve the accuracy of artificial neural networks. arXiv Prepr. 2024. 10.48550/arXiv.2303.09935. arXiv:2303.09935. [DOI] [Google Scholar]
  • 39.Kůdela J. The evolutionary computation methods no one should use. arXiv Prepr. 2023. 10.48550/arXiv.2301.01984. arXiv:2301.01984. [DOI] [Google Scholar]
  • 40.Sobol′ IM. Sensitivity estimates for nonlinear mathematical models. Math Model Comput Exp. 1993;1(4):407–14. [Google Scholar]
  • 41.Sobol′ IM. Global sensitivity indices for nonlinear mathematical models and their Monte Carlo estimates. Math Comput Simul. 2001;55(1–3):271–80. 10.1016/S0378-4754(00)00270-6. [DOI] [Google Scholar]
  • 42.Sic A, Cvetkovic K, Manchanda E, Knezevic NN. Neurobiological implications of chronic stress and metabolic dysregulation in inflammatory bowel diseases. Diseases. 2024;12(9):220. 10.3390/diseases12090220. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Battersby AJ, Kampmann B, Burl S. Vitamin D in early childhood and the effect on immunity to Mycobacterium tuberculosis. Clin Dev Immunol. 2012;2012:430972. 10.1155/2012/430972. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Howren MB, Lamkin DM, Suls J. Associations of depression with C-reactive protein, IL-1, and IL-6: a meta-analysis. Psychosom Med. 2009;71(2):171–86. 10.1097/PSY.0b013e3181907c1b. [DOI] [PubMed] [Google Scholar]
  • 45.Dowlati Y, Herrmann N, Swardfager W, Liu H, Sham L, Reim EK, et al. A meta-analysis of cytokines in major depression. Biol Psychiatry. 2010;67(5):446–57. 10.1016/j.biopsych.2009.09.033. [DOI] [PubMed] [Google Scholar]
  • 46.Haapakoski R, Mathieu J, Ebmeier KP, Alenius H, Kivimäki M. Cumulative meta-analysis of interleukins 6 and 1, tumour necrosis factor α and C-reactive protein in patients with major depressive disorder. Brain Behav Immun. 2015;49:206–15. 10.1016/j.bbi.2015.06.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Goldsmith DR, Rapaport MH, Miller BJ. A meta-analysis of blood cytokine network alterations in psychiatric patients: comparisons between schizophrenia, bipolar disorder and depression. Mol Psychiatry. 2016;21(12):1696–709. 10.1038/mp.2016.3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Schmidt H, Shelton R, Duman R. Functional biomarkers of depression: diagnosis, treatment, and pathophysiology. Neuropsychopharmacology. 2011;36:2375–94. 10.1038/npp.2011.151. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Yuan C, Liu Z, Li X, Zhou X, Wang D, Fan Y, et al. A dynamic weighted ensemble learning framework for cardiovascular risk prediction in type 2 diabetes: a comparative study with SHAP-based interpretability. Sci Rep. 2025;15:45029. 10.1038/s41598-025-28786-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Mitchell AJ, Vaze A, Rao S. Clinical diagnosis of depression in primary care: a meta-analysis. Lancet. 2009;374(9690):609–19. 10.1016/S0140-6736(09)60879-5. [DOI] [PubMed] [Google Scholar]
  • 51.American Psychiatric Association. Diagnostic and Statistical Manual of Mental Disorders. 5th ed. Arlington, VA: American Psychiatric Publishing; 2013. 10.1176/appi.books.9780890425596. [DOI] [Google Scholar]
  • 52.Fried EI, Flake JK, Robinaugh DJ. Revisiting the theoretical and methodological foundations of depression measurement. Nat Rev Psychol. 2022;1(6):358–68. 10.1038/s44159-022-00050-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Xu Y, Quan R, Xu W, Huang Y, Chen X, Liu F. Advances in Medical Image Segmentation: A Comprehensive Review of Traditional, Deep Learning and Hybrid Approaches. Bioengineering. 2024;11(10):1034. 10.3390/bioengineering11101034. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Xu Y, Liu F, Xu W, Quan R. Overview of Graph Theoretical Approaches in Medical Image Segmentation. In: Zhou K, editor. Computational and Experimental Simulations in Engineering. ICCES 2024. Mechanisms and Machine Science. Volume 175. Cham: Springer; 2025. pp. 819–35. 10.1007/978-3-031-81673-4_61. [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The datasets generated and/or analyzed during the current study are not publicly available due to privacy restrictions but are available from the corresponding author upon reasonable request.


Articles from BMC Medical Informatics and Decision Making are provided here courtesy of BMC

RESOURCES