Skip to main content
STAR Protocols logoLink to STAR Protocols
. 2024 Jul 31;5(3):103213. doi: 10.1016/j.xpro.2024.103213

Integrative protocol for quantifying cholesterol-related sterols in human serum samples and building decision support systems

Eva Kočar 1,6,7, Sonja Katz 2,3,6,7, Žiga Pušnik 4, Cene Skubic 1, Tadeja Režen 1, Vitor AP Martins dos Santos 3,5, Miha Mraz 4, Miha Moškon 4,, Damjana Rozman 1,8,∗∗
PMCID: PMC11342266  PMID: 39088327

Summary

The growing interest in clinical diagnostics has recently focused on metabolic biomarkers. Here, we present a protocol for sample preparation, extraction of cholesterol-related sterols, and quantification of 10 sterols in human blood serum samples using targeted liquid chromatography-tandem mass spectrometry (LC-MS/MS). We also describe steps of machine learning techniques to develop novel decision-making systems that offer potential benefits in disease monitoring and surveillance by measuring metabolic pathways.

For complete details on the use and execution of this protocol, please refer to Kočar et al.1 and Skubic et al.2

Subject areas: Metabolism, Molecular Biology, Computer sciences

Graphical abstract

graphic file with name fx1.jpg

Highlights

  • Instructions for the extraction of sterols from human serum samples

  • Steps for detecting and quantifying 10 non-polar sterols using LC-MS/MS

  • Guidelines for machine learning in developing novel decision support systems


Publisher’s note: Undertaking any experimental protocol requires adherence to local institutional guidelines for laboratory safety and ethics.


The growing interest in clinical diagnostics has recently focused on metabolic biomarkers. Here, we present a protocol for sample preparation, extraction of cholesterol-related sterols, and quantification of 10 sterols in human blood serum samples using targeted liquid chromatography-tandem mass spectrometry (LC-MS/MS). We also describe steps of machine learning techniques to develop novel decision-making systems that offer potential benefits in disease monitoring and surveillance by measuring metabolic pathways.

Before you begin

This protocol delineates the procedure for the extraction of cholesterol-related sterols (sterols from the post-squalene part of cholesterol synthesis; see Figure 1) from human blood serum samples, followed by targeted lipidomics – the quantification of cholesterol related sterol intermediates by LC-MS/MS. As the use of metabolic biomarkers is lately attracting growing interest in clinical diagnostics, we also describe the machine learning techniques for building new decision-making systems for disease monitoring and surveillance. This gives us a unique opportunity to search for potential benefits of measuring metabolic pathways. Moreover, our approach and protocol can also be used for readily available clinical parameters (e.g., clinical parameters collected upon patient admission to the hospital), as was recently published by our group (please refer to Kočar et al.1). Briefly, serum samples were collected from COVID-19 patients at the time of their hospitalization. Sterol intermediates (referred also as sterols throughout the manuscript) were extracted and identified/quantified by LC-MS/MS. Based on the measurements we developed prediction models for disease severity using machine learning techniques, applying sterol concentrations and simple clinical parameters (both measured upon admission), and their combination. This protocol is divided into 4 parts: (i) sample preparation, (ii) extraction of sterols from serum samples, (iii) LC-MS/MS analysis, and (iv) machine learning techniques for the development of novel decision-making systems. The timing of the individual steps is determined based on experimental experience and is also influenced by the quantity of samples.

Figure 1.

Figure 1

Post-squalene part of cholesterol biosynthesis

Cholesterol synthesis is divided into pre-squalene (not shown) and post-squalene parts. A more detailed description of the pre-squalene pathway is given elsewhere.3 Lanosterol is the first molecule of the post-squalene part of cholesterol synthesis. From this point on, the synthesis branches either to a Bloch or a Kandutsch-Russel (K-R) pathway. They are enzymatically identical except for the first and the last step. In the Bloch pathway, CYP51A1 catalyzes lanosterol conversion to FF-MAS, while in the K-R pathway DHCR24 catalyzes it to 24,25-dihydrolanosterol. The last step of the Bloch pathway is defined by a conversion of desmosterol to cholesterol (by DHCR24), while the K-R pathway finishes with the conversion of 7-dehydrocholesterol to cholesterol by DHCR7. As all sterol intermediates from lanosterol to desmosterol contain C24 double bond, DHCR24 can theoretically metabolize any of them, therefore, both branches intertwine.3,4,5 The enzymes catalyzing each step of the synthesis are shown in gray.

We believe that our approach demonstrates how combining different research areas can create novel decision-support tools capable of improving clinical standards and shows great potential for different applications in the clinic (e.g., new biomarkers for disease severity, outcome, and long-term effects).

Institutional permissions

The present protocol uses human serum samples. All participants enrolled in our study provided written informed consent and the study was approved by the Medical Ethics Committee of the Republic of Slovenia (No. 0120–211/2020/7 and No. 0120–33/2022/3). Prior to replicating the present protocol ensure to acquire permissions from relevant institutions and patients included in your study.

Sample collection

Inline graphicTiming: 30–45 min

This section describes the steps of blood sample collection and processing for further serum isolation (Figure 2).

  • 1.
    Collect blood samples into vacutainer tubes.
    • a.
      Use Vacutainer Serum tubes to collect serum (e.g., BD Vacutainer Serum tubes, Cat. No.: 367815). They are available in different sizes and can be recognized by a red cap.
  • 2.

    Centrifuge samples at 1800 × g at 22°C for 10 min.

  • 3.

    Transfer serum to a new cryovial or microcentrifuge tube.

  • 4.

    Make aliquots if necessary (at least 250 μL for one extraction).

  • 5.

    Store at −80°C until further use (long-term storage) or proceed with sterol extraction (see step-by-step method details, sterol extraction).

Note: Step 1 is performed at the clinic and is not included in 30–45 min timing.

Inline graphicCRITICAL: With this method you are dealing with biological samples and hazardous solvents (toluene, methanol, cyclohexane, 1-propanol, formic acid) and must therefore proceed very carefully and observe the precautions according to the safety data sheets (obtained from supplier). Work under a fume hood and wear personal protective equipment.

Figure 2.

Figure 2

Schematic overview of sample preparation process

All organic solvents used in this protocol are HPLC/MS grade (Please refer to key resources table).

Preparation of hydrolysis solution

Inline graphicTiming: 10 min

This section describes the steps for preparing the hydrolysis solution used for sterol extraction (see step-by-step method details, sterol extraction).

  • 6.

    Dissolve 4 g of NaOH in 10 mL of Milli-Q water.

  • 7.
    Add 90 mL of HPLC grade ethanol.
    • a.
      Final concentation of NaOH is 1M (Milli-Q water:ethanol, 1:9, v/v).
  • 8.

    Store at +4°C until further use.

Note: After adding ethanol (step 7), the color of the solution becomes milky. When the hydrolysis solution turns a yellow/orange/brown color (during storage at +4°C), prepare a fresh stock.

Preparation of internal standard

Inline graphicTiming: 45 min

This section describes the steps for preparing internal standard for further sterol extraction (see step-by-step method details, sterol extraction).

Note: A stock solution of the internal standard (lathosterol-D7, Avanti Polar Lipids, Cat. No.: 700056P) should be prepared at 1 mg/mL in toluene and stored at −20°C. It is recommended to prepare smaller aliquots to avoid freeze/thaw cycles.

  • 9.
    Bring the stock solution of internal standard to 22 ± 2°C for at least 20 min before use.
    • a.
      Warming the stock solution is necessary to improve solubility and thus ensure precise and reliable results.
  • 10.
    Dilute the stock solution (lathosterol-D7, 1 mg/mL) in HPLC grade methanol to obtain a final concentration of lathosterol-D7 of 10 ng/μL.
    • a.
      Use glass vial with screw cap.
    • b.
      Add 1980 μL of HPLC grade methanol to 2 mL glass vial with screw cap.
    • c.
      Vortex the stock solution of lathosterol-D7.
    • d.
      Add 20 μL of stock solution to the vial with HPLC grade methanol.
  • 11.

    Vortex for 30 s.

  • 12.

    Make aliquots of 500 μL.

  • 13.
    Overlay samples with N2 to prevent oxidation.
    • a.
      Purge the samples with nitrogen by carefully pouring a stream of nitrogen over the vial containing the sample to displace any oxygen present.
    • b.
      Seal the vial.
  • 14.

    Store at −20°C until further use.

Preparation of calibration standards

Inline graphicTiming: 1.5 h

This section describes the steps for preparing standard solution consisting of the same concentration of each sterol standard used for generating calibration curve and further sterol quantification.

Calibration standard mix No.1: lanosterol, 24,25-dihydrolanosterol, T-MAS, dihydro-T-MAS, zymosterol, zymostenol, 24-dehydrolathosterol, lathosterol, desmosterol.

Calibration standard No.2: cholesterol.

See key resources table for specific information regarding each sterol standard.

  • 15.
    Dissolve each sterol in toluene to a suitable stock concentration.
    • a.
      Be careful to not over dilute.
    • b.
      See Table 1 for an example.
  • 16.

    Prepare a mix consisting of each sterol with a final concentration of 40 μg/mL for each sterol (in HPLC grade methanol).

  • 17.

    Vortex vigorously.

  • 18.
    Prepare two serial dilutions of standard solutions with known concentrations in HPLC grade methanol in a given concentration range (to cover the expected concentration range in your samples).
    • a.
      e.g., sterol mix: 0, 1.2, 2.4, 4.9, 9.8, 19.5, 39.1, 78.1, 156.3, 312.5, 625, 1250, 2500, 5000 ng/μL.
    • b.
      e.g., cholesterol: 0, 9.8, 19.5, 39.1, 78.1, 156.3, 312.5, 625, 1250, 2500 ng/μL.
    • c.
      For accuracy, use fresh mix each time. While 5 μL injections don’t require large standards volume, preparing 200–500 μL can ease pipetting. Use glass vials for this purpose (refer to the key resources table).
  • 19.

    Vortex vigorously.

  • 20.
    Overlay samples with N2.
    • a.
      Purge the samples with nitrogen by carefully pouring a stream of nitrogen over the vial containing the sample to displace any oxygen present.
    • b.
      Seal the vial.
  • 21.

    Store at −20°C until further use.

Inline graphicCRITICAL: To ensure accuracy and prevent degradation, it is ideal to dissolve powdered lipid standards directly in their supplied glass container, typically amber glass, shielding them from light. This precaution is particularly important for lipid standards that are sensitive to UV and visible light. To prevent the organic solvent from dripping, aspirate and expel it three times before adding the organic solvent to the vial. Also, make sure that you pipette quickly but carefully to avoid evaporation and to ensure the correct final concentration of internal standard. Be careful to not over dilute sterol stock samples.

Table 1.

Stock concentrations of sterols

Sterol Stock concentration [mg/mL]
lanosterol 1
24,25-dihydrolanosterol 1
T-MAS 1
dihydro-T-MAS 1
zymosterol 1
zymostenol 1
24-dehydrolathosterol 1
lathosterol 2.5
desmosterol 2.5
cholesterol 10

Dissolve sterols in toluene to a final stock concentration, except cholesterol in methanol.

Note: Since the sterols are supplied as 1–10 mg powder, the stocks are diluted in toluene to prevent evaporation. Cholesterol is purchased as a 5–100 g powder and can therefore be diluted directly in methanol and you can prepare it fresh several times during your work.

Preparation of mobile phase for LC-MS/MS analysis

Inline graphicTiming: 15 min

This section describes the steps for preparing the mobile phase for LC-MS/MS analysis.

  • 22.
    Based on the number of your samples, calculate the appropriate volume of mobile phase plus 10%–20% as a safety margin.
    • a.
      Flow rate is 200 μL/min for sterols and 300 μL/min for cholesterol.
    • b.
      Method is 36 min long per sample.
  • 23.

    The composition of mobile phase is shown in Table 2.

Note: Use HPLC/MS grade solvents.

Table 2.

Mobile phase composition for LC-MS/MS analysis

Solvent Volume [%]
methanol 80
1-propanol 10
formic acid 0.05
Milli-Q water 9.95

Key resources table

REAGENT or RESOURCE SOURCE IDENTIFIER
Biological samples

Human serum from patients with COVID-19 This study N/A

Chemicals, peptides, and recombinant proteins

Lanosterol Avanti Polar Lipids (Alabaster, AL, USA) Cat# 700063
24,25-Dihydrolanosterol Avanti Polar Lipids (Alabaster, AL, USA) Cat# 700067
T-MAS Avanti Polar Lipids (Alabaster, AL, USA) Cat# 700073
dihydro-T-MAS Avanti Polar Lipids (Alabaster, AL, USA) Cat# 700173
Zymosterol Avanti Polar Lipids (Alabaster, AL, USA) Cat# 700068
Zymostenol Avanti Polar Lipids (Alabaster, AL, USA) Cat# 700118
24-Dehydrolathosterol Avanti Polar Lipids (Alabaster, AL, USA) Cat# 700114
Lathosterol Avanti Polar Lipids (Alabaster, AL, USA) Cat# 700069
Lathosterol-D7 Avanti Polar Lipids (Alabaster, AL, USA) Cat# 700056
Desmosterol Avanti Polar Lipids (Alabaster, AL, USA) Cat# 700060
Cholesterol Merck (Darmstadt, Germany) Cat# C3045
Ethanol Various manufacturers For example: Honeywell, Riedel-de Haën, Cat# 34852
Methanol Various manufacturers For example: Honeywell, Riedel-de Haën, Cat# 34860
Cyclohexane Various manufacturers For example: Sigma-Aldrich, Cat# 34855
Formic acid Various manufacturers For example: Honeywell, Fluka, Cat# 94318
1-Propanol Various manufacturers For example: Honeywell, Riedel-de Haën, Cat# 34871
Toluene Various manufacturers For example: Merck, Cat# 1.00849
Water Various manufacturers For example: Merck, Cat# 1.15333
Sodium hydroxide Various manufacturers For example: Honeywell, Fluka, Cat# 30620

Software and algorithms

scikit-learn https://scikit-learn.org/stable/
Code for model development and validation: available upon request Kočar et al.1 and this paper https://github.com/sonjakatz/covid_sterols_ML
https://doi.org/10.5281/zenodo.12167403

Other

Shaking water bath Various manufacturers For example: BioLink Laboratories, SB-12L
Vortex Various manufacturers For example: IKA, MS 3 basic
Centrifuge Eppendorf, Germany 5810R
Concentrator Eppendorf, Germany 5301
HPLC Shimadzu, Japan Shimadzu Nexera XR HPLC
Mass spectrometer AB Sciex LLC, USA SCIEX Triple Quad 3500
Luna 3 μm PFP(2) 100A column Shimadzu, Japan Cat# 00D-4447-B0 and Cat# 00F-4447-B0
BD Vacutainer serum tubes BD, USA Cat# 367815 (6 mL)
Cryo tube Various manufacturers For example: TPP, Cat# 89020
Microcentrifuge tubes Various manufacturers For example: Sarstedt, Cat# 72.690.001
Glass Pasteur pipettes Various manufacturers For example: Brand, Cat# 7477 20
15 mL glass conical centrifuge tube with screw cap (#1) Various manufacturers For example: centrifuge tube – Merck, Corning PYREX Disposable Glass Conical Centrifuge Tubes, Cat# CLS9950215; screw cap – Merck, Corning disposable phenolic cap, Cat# CLS9999915
15 mL glass conical centrifuge tube without screw cap Various manufacturers For example: BRAND, Cat# BR779012
Screw cap (#2) Various manufacturers For example: Carl ROTH, Screw caps LABOCAP for Ø 15/16 mm
2 mL clear HPLC vials Various manufacturers For example: Thermo Scientific, Cat# C4000-1
0.2 mL integ. insert HPLC vials Various manufacturers For example: Macherey-Nagel, Cat# 702007
Screw caps for HPLC vials Various manufacturers For example: Macherey-Nagel, Cat# 702287

Step-by-step method details

Sterol extraction

Inline graphicTiming: 3.5–4 h

This section describes the steps for extracting sterols from serum samples.

  • 1.

    Add 250 μL of serum sample to a 15 mL glass conical centrifuge tube vial with a screw cap (See key resources table, Other, screw cap #1).

  • 2.
    Add 20 μL of internal standard (10 ng/μL lathosterol-D7).
    • a.
      Internal standard should be at 22 ± 2°C for at least 20 min before use.
    • b.
      Vortex internal standard before adding to the serum.
    • c.
      Overlay internal standard solution with N2 before sealing it and store at −20°C.
  • 3.

    Add 1 mL of hydrolysis solution and vortex the sample vigorously for 10 s.

  • 4.

    Incubate samples at 190 rpm at 65°C for 1 h in shaking water bath.

  • 5.

    Add 0.5 mL Milli-Q water.

  • 6.

    Add 3 mL of cyclohexane.

  • 7.

    Vortex the sample vigorously for 20 s.

  • 8.

    Centrifuge the sample at 1300 × g at 22°C for 10 min.

  • 9.

    Transfer upper organic phase to a new 15 mL glass conical centrifuge tube without screw cap using bore (Pasteur) glass pipette.

  • 10.

    Add 3 mL of cyclohexane to the leftover of the original sample.

  • 11.

    For pooling the second extract, repeat steps 7–9.

  • 12.
    After combining both extracts, evaporate the organic solvent using a vacuum centrifuge/concentrator at 45°C.
    • a.
      Usually takes up to 35 min.
  • 13.
    Dissolve lipid film in 100 μL of HPLC grade methanol.
    • a.
      Add 100 μL of HPLC grade methanol to the sample.
    • b.
      Cover the glass conical centrifuge tube with a screw cap to avoid evaporation of the methanol (See key resources table, Other, screw cap #2).
    • c.
      Vortex the sample.
  • 14.

    Transfer the sample immediately to HPLC glass vial to avoid evaporation of the methanol.

  • 15.

    Overlay the sample with N2 before sealing it.

  • 16.

    You can proceed with LC-MS/MS analysis or store the sample at −20°C until further use.

Inline graphicCRITICAL: After evaporation of the organic solvent (step 12), allow the glass centrifuge tubes to stand at room temperature for 1–2 min before adding methanol and dissolving the lipid film (to avoid evaporation of the methanol at a higher temperature). When pipetting organic solvents, especially cyclohexane, only use chemically resistant tips to ensure safety and accuracy.

Note: If you do not have a shaking bath available (needed in step 4), you can also incubate the sample in a water bath and vortex it every few min (vigorously for 20 s) before placing the sample back in the water bath. When sample volume is less than 250 μL use vial with inserts (step 14). Examples of different conical 15 mL glass centrifuge tubes and different screw caps can be found in the key resources table.

LC-MS/MS analysis

Inline graphicTiming: 36 min/sample

This section introduces the steps for chromatographic separation and MS detection of cholesterol-related sterols. MS experimental settings are shown in Tables 3 and 4. This method is a refined version of our previously published approach.2 Representative chromatogram of sterol intermediates isolated from serum samples is shown in Figure 3.

  • 17.
    Chromatographic separation.
    • a.
      Use two pentafluorophenyl columns Phenomenex Luna 3 mm (Phenomenex, USA) with combined length of 250 mm.
    • b.
      Set the temperature of the autosampler to 15°C.
    • c.
      Set the oven temperature to 40°C.
    • d.
      Use mobile phase with a methanol/1-propanol/formic acid/water (v/v/v/v, 80:10:0.05:9.95) composition.
    • e.
      Set an isocratic flow to 200 μL/min for sterol intermediates, except for cholesterol at 300 μL/min.
    • f.
      Set the injection volume of the standard or sample to 5 μL, except for cholesterol to 1 μL.

Inline graphicCRITICAL: Concentration of cholesterol is significantly higher compared to other sterols, therefore, cholesterol has to be measured separately in a new independent run with different injection volume and flow rate. If necessary, dilute the samples.

  • 18.
    Detection.
    • a.
      Detailed information about MS experimental conditions and parameters are listed in Table 3.
    • b.
      Calibration: Acquire MS data for each calibration standard at different concentrations and generate a calibration curve (response vs. concentration). See Preparation of sterol calibration standards.

Inline graphicCRITICAL: If you have stored your samples at −20°C, bring them to 37°C for at least 15–20 min and then vortex the samples before placing them in the sample tray and starting the method. Warming the samples prior to LC-MS/MS analysis is crucial for optimizing results by increasing solubility, decreasing viscosity and improving ionization. This ensures reliable, sensitive and accurate analysis while preventing clogging of the device and incomplete ionization, facilitating sample handling and enabling the detection of lipid molecules even at low concentrations.

Table 3.

MS experimental settings

Parameter Condition – sterol intermediates Condition – cholesterol
Ionization source APCI
Ionization mode positive
Curtain Gas (CUR) 20 30
Collision Gas (CAD) 4 6
IonSpray Voltage (IS) 5500
Temperature (TEM) 350
Ion Source Gas 1 (GS1) 25
Ion Source Gas 2 (GS2) 0
Collision energy (CE) 26

Table consists of Source/Gas and Compound parameters for sterol intermediates and cholesterol.

Table 4.

Detailed parameters for each compound

Sterol Molecular weight (g/mol) Q1 Mass (Da) Q3 Mass (Da) DP (volts) CE (volts) CXP (volts)
lanosterol 426.72 409 191 80 22 13
24,25-dihydrolanosterol 428.73 411 191 100 24 16
T-MAS 412.69 395 243 100 30 16
dihydro-T-MAS 414.71 397 177 110 30 14
zymosterol 384.64 367 215 85 26 14
zymostenol 386.65 369 215 80 30 13
24-dehydrolathosterol 384.64 367 215 85 26 14
lathosterol 386.65 369 215 80 30 13
lathosterol-D7∗ 393.70 376 215 100 27 16
desmosterol 384.64 367 215 85 26 14
cholesterol 386.65 369.3 215.1 80 30 13

CE, collision energy; CXP, collision cell exit potential; DP, declustering potential; Q1 Mass, mass-to-charge ratio (m/z) of ions selected by the first quadrupole; Q3 Mass, mass-to-charge ratio (m/z) of ions selected by the third quadrupole; ∗, internal standard.

Figure 3.

Figure 3

Representative chromatogram of sterol standards

1 – zymosterol, 2–24-dehydrolathosterol, 3 – desmosterol, 4 – zymostenol, 5 – lathosterol, 6 – T-MAS, 7 – lanosterol, 8 – dihydro-T-MAS, 9–24,25-dihydrolanosterol. Cholesterol is not shown, as it is measured in a separated run.

Note: In our case chromatographic separation was performed on a Shimadzu Nexera XR HPLC (Shimadzu, Japan) and detection on a SCIEX Triple Quad 3500 mass spectrometer (AB Sciex LLC, USA). Data was evaluated with Analyst software 1.6.3 (AB Sciex LLC, USA).

Setting up the computational environment

> git clone git@github.com:sonjakatz/covid_sterols_ML.git .

  • 20.

    Install the environment.

> conda env create -f environment.yml

> source activate env_covid

if(!require(devtools)){

install.packages("devtools") # If not already installed

}

devtools::install_github('Sbuttery/protocols')

Note: The time estimates provided derive from our computations conducted via a virtual machine which is hosted on a physical Supermicro H11Dsi server, running a MicroGB of available memorysoft Server 2019 hypervisor with 2 AMD EPYC 7F72 24-core processors and 256 GB RAM. Respective computational virtual machine has 16 virtual processors with 2 threads per processor (32 threads), dynamically allocated memory with maximum amount of 64 GB and 2 TB of disk space.

Data preprocessing

Inline graphicTiming: 1 h

This section explains the steps for data preprocessing. Flowchart depicting fitting classification models on the training set is shown in Figure 4.

Figure 4.

Figure 4

Flowchart depicting fitting classification models on the training set

Blue segments show system initialization and program execution, purple segments represent iterations through selected classification models and testing all scenarios, while the process of cross-validation is shown in green segments. Figure is adapted from Kočar et al.1

All preprocessing steps and classification algorithms used in this study were implemented using the scikit-learn Python library (version 1.1.3).6 The source code and fitted models are accessible via a public GitHub repository.

  • 21.

    Load the required packages.

import os

import pandas as pd

  • 22.

    Load the data.

> data = pd.read_csv("data.csv", index_col=0)

  • 23.

    Select your target variable for predictions (e.g., disease severity).

> endpoint = "disease_severity"

  • 24.
    Determine the number of classes/categories.
    • a.
      e.g., categories according to increasing severity – mild (class: 0), moderate (class: 1), severe (class: 2), critical (class: 3).
    • b.
      In the case of small number of members in a specific class, try to combine groups in order to achieve similar number of samples in each class (e.g., Class 1: mild, moderate; Class 2: severe, critical).

Note: This counteracts model inaccuracies accountable to class imbalances during training, reducing the risk of under-fitting of the classifier.

data[endpoint].replace(1,0,inplace=True)

data[endpoint].replace(2,0,inplace=True)

data[endpoint].replace([3,4],1,inplace=True)

  • 25.

    Discard samples/patients with missing target variable.

> data = data[data[endpoint].notna()]

  • 26.
    Discard irrelevant or potentially biasing variables from data preprocessing.
    • a.
      e.g., timestamps (e.g., date).

idx = data.columns[data.columns.str.contains("date")]

data_cleaned = data.drop(idx, axis=1)

  • 27.

    Remove features with a rate of missingness of more than 15%.

cutoff = 0.15

cols2keep = data.columns[data.isna().sum(axis=0) / data.shape[0] <= cutoff]

data_cleaned = data.loc[:,cols2keep]

  • 28.

    Save preprocessed data.

> data_cleaned.to_csv("data_cleaned.csv")

Feature selection

Inline graphicTiming: ∼6 h; dependent on computational resources

Before training the models, select relevant features for predicting your target variable to avoid overfitting of machine learning models (feature selection). How to achieve this is explained in the following section.

  • 29.

    Load the required packages.

import os

import pandas

from preprocessing import run_iterativeBoruta, imputation_scaling

import json

  • 30.

    Load the data.

data = pd.read_csv("data_cleaned.csv", index_col=0)

endpoint = "disease_severity"

X = data.drop(endpoint, axis=1)

y = data[endpoint].ravel()

  • 31.

    Perform data imputation for features with a degree of data missingness of ≤ 15%.

    We suggest to use the following methods for different data types of features.
    • a.
      Continuous features: use IterativeImputer from scikit-learn6 and scale them through min-max normalization. This scales and translates each feature individually to lie between zero and one, ultimately scaling the maximum absolute value of each feature to unit size.
    • b.
      Binary features: impute using K-Nearest Neighbors (kNN) method. For each sample’s missing value, kNN imputation fills missingness by averaging the nearest neighbors' values based on the Euclidean distance.
    • c.
      Categorical features: impute by the most frequently occurring value and encode to numerical representation by an ordinal encoder. Categorical variables commonly include clinical metadata such as information on types and name of prescribed medication, description of symptoms (e.g., type of cough - dry, productive, mixed with blood), presence of infections with different viral, bacterial, and fungal strains (e.g., Streptococcus pneumoniae).
      Note: Imputation is required, as the machine learning models utilized cannot handle missing data points.
      num_columns = X.select_dtypes(include=["float64"]).columns
      bin_columns = X.select_dtypes(include=["int64"]).columns
      cat_columns = X.select_dtypes(include=["object"]).columns
      preprocessor = imputation_scaling(num_columns, bin_columns, cat_columns, X)
      columnOrderAfterPreprocessing = [ele[5:] for ele in preprocessor.get_feature_names_out()]
      X_preproc = preprocessor.fit_transform(X)
  • 32.
    Conduct feature selection using iterative variation of the unsupervised feature selection method called Boruta.7
    • a.
      Carry out 100 iterations of Boruta with random initialization cord the selected features of each iteration.
      Note: We suggest to use the python implementation of Boruta due to its flexibility in selecting a threshold for comparison between shadow and real features. Although the number of conducted iterations may be varied we suggest a minimum of 100 iterations for statistical relevance.8
      n_iter = 100
      perc = 100
      dict_iterBoruta = run_iterativeBoruta(
       X=X_preproc, y=y,
       cols=columnOrderAfterPreprocessing, perc=perc,
       n_iter=n_iter)
      with open(f"iterativeBoruta_{perc}perc.json", "w") as f:
      json.dump(dict_iterBoruta, f, indent=4)
    • b.
      Only keep features recorded in at least 50% of iterations for further analysis.
      perc=100
      with open(f"iterativeBoruta_{perc}perc.json", "r") as f: dict_iterBoruta = json.load(f)
      thresh = 0.5
      with open(f"iterativeBoruta.txt", "w") as f:
       for key, val in dict_iterBoruta.items():
       if val > thresh:
       f.write(key+"∖n")

Classification models

Inline graphicTiming: ∼12–24 h; dependent on computational resources

Inline graphicCRITICAL: Prior to training, choose a few suitable predictive models (e.g. for classification problems: AdaBoost,9 Gaussian Naive Bayes,10 Gaussian Processes,11 K-Nearest Neighbors,12 Logistic Regression,13 Multilayer Perceptron,14 Random Forest,15 Quadratic Discriminant Analysis16).

  • 33.

    Load the required packages.

import os

import pandas

import numpy as np

import random

import sys

from classification import classify_leave_one_out_cv

from preprocessing import imputation_scaling, SupervisedSelector

from sklearn.ensemble import RandomForestClassifier, AdaBoostClassifier

from sklearn.svm import SVC

from sklearn.gaussian_process import GaussianProcessClassifier

from sklearn.model_selection import GridSearchCV

from sklearn import metrics

from sklearn.tree import DecisionTreeClassifier

from sklearn.linear_model import LogisticRegression

from sklearn.neighbors import KNeighborsClassifier

from sklearn.neural_network import MLPClassifier

from sklearn.naive_bayes import GaussianNB

from sklearn.discriminant_analysis import QuadraticDiscriminantAnalysis

from sklearn.dummy import DummyClassifier

from sklearn.gaussian_process.kernels import RBF, DotProduct, RationalQuadratic, WhiteKernel, Matern

  • 34.

    Load the data.

data = pd.read_csv(f"data_cleaned.csv", index_col=0)

endpoint = "disease_severity"

X = data.drop(endpoint, axis=1)

y = data[endpoint]

  • 35.

    Read in variables selected by feature selection.

> sel_variables = pd.read_csv("iterativeBoruta.txt", header=None)[0].tolist()

  • 36.

    Perform data imputation for features with a degree of data missingness of ≤ 15%.

num_columns = X.select_dtypes(include=[‘"float64"]).columns

bin_columns = X.select_dtypes(include=["int64"]).columns

cat_columns = X.select_dtypes(include=["object"]).columns

preprocessor = imputation_scaling(num_columns, bin_columns, cat_columns, X)

X_imputed = preprocessor.fit_transform(X) X_imputed = SupervisedSelector(preprocessor, sel_variables).transform(X_imputed)

  • 37.

    Create a selection of machine learning models to train.

models = {‘svc’: SVC(probability=True),

 ‘rfc’: RandomForestClassifier(),

 ‘gpr’: GaussianProcessClassifier(),

 ‘abc’: AdaBoostClassifier(base_estimator = DecisionTreeClassifier(random_state = 11, max_features = "auto", class_weight = "balanced",max_depth = None)),

 ‘log’: LogisticRegression(),

 ‘knn’: KNeighborsClassifier(),

 ‘mlp’: MLPClassifier(),

 ‘gnb’: GaussianNB(),

 ‘qda’: QuadraticDiscriminantAnalysis(),

 ‘mcl’: DummyClassifier(strategy="most_frequent}

  • 38.

    Define the hyperparameter selection to tune models.

Note: Hyperparameters determine the structure and complexity of models. Machine learning models must be hyperparameter tuned to optimize their performance and generalize well to unseen data, as the default hyperparameters may not be suitable for every dataset. Proper tuning ensures that the model achieves the best possible accuracy and avoids issues like overfitting or underfitting. The selection of hyperparameters trained is different for every model architecture, so carefully checking of scikit-learn documentation is advised.

grids = {‘rfc’:{

‘n_estimators’: [100, 300, 1000],

 ‘max_depth’: [2,4,6],

 ‘max_features’: [2,4,6],

 ‘ccp_alpha’: list(np.linspace(0, 0.025, 2)),

 },

 ‘svc’:{‘C’: [0.1, 1, 10, 100],

 ‘gamma’: [1, 0.1, 0.01, 0.001, 0.0001],

 ‘kernel’: [‘rbf’, ‘poly’, ‘linear’]

 },

 ‘gpr’:{‘kernel’:[1∗RBF(), 1∗DotProduct(), 1∗Matern(), 1∗RationalQuadratic(), 1∗WhiteKernel()]},

 ‘abc’:{"base_estimator__criterion" : ["gini", "entropy"],

 "base_estimator__splitter" : ["best", "random"],

 "n_estimators": [1, 2]

 },

 ‘log’:{‘penalty’: [‘l1’,‘l2’], ‘C’: [0.001,0.01,0.1,1,10,100,1000]},

 ‘knn’:{‘n_neighbors’: list(range(1, 15)),

 ‘weights’: [‘uniform’, ‘distance’],

 ‘metric’: [‘euclidean’, ‘manhattan’]},

 ‘mlp’:{‘solver’: [‘adam’],

 ‘max_iter’: [50, 100, 200],

 ‘alpha’: 10.0 ∗∗ -np.arange(0, 5),

 ‘hidden_layer_sizes’: [(random.randrange(15, 41), random.randrange(5, 16)) for i in range(5)],

 },

 ‘gnb’: {‘var_smoothing’: np.logspace(-9,9, num=100)},

 ‘qda’: {‘reg_param’: (0.00001, 0.0001, 0.001,0.01, 0.1),

 ‘store_covariance’: (True, False),

 ‘tol’: (0.0001, 0.001,0.01, 0.1)},

 ‘mcl’: {}}

  • 39.

    Perform leave-one-out cross validation for each combination of model type and scenario (scenario being combination of different set of features.).

    Within each cross validation split.
    • a.
      For optimal performance, optimize model hyperparameters during training. The scoring metric evaluates the performance of the cross-validated model on the test set. Due to a slight class imbalance in the dataset, balanced accuracy was chosen as metric.
      • i.
        Scoring metric: balanced accuracy.
    • b.
      Assess and evaluate predictive power of each classification model.
      • i.
        Evaluation metrics: balanced classification accuracy, F1-score, precision, recall, and ROC-AUC score. For their explanation, please refer to Table 5.
    • c.
      Measure the importance of each individual feature by applying an iterative permutation-based feature importance assessment15 with 100 iterations.
      • i.
        Average feature importance of all 100 iterations.
        resultsPath = "results"
         for model in models.keys():
         df_before = pd.DataFrame()
         df_features = pd.DataFrame()
         df_importances = pd.DataFrame()
         saveIndivdualPred = True
         clf = GridSearchCV(models[model], grids[model], scoring='balanced_accuracy', verbose=1, cv=5, n_jobs=-1)
        result = classify_leave_one_out_cv(
         clf,
         X_imputed,
         y,
         model=model,
         save_to = resultsPath + f"/{model}", select_features=True,
         permutation_repeats=100,
         scale_features = False,
         saveIndivdualPred = saveIndivdualPred,
         logfile = "log_clinical")
         result['model'] = model
         df_before = df_before.append(result['df_results'], ignore_index=True)
         df_importances = df_importances.append(result['importances_df'], ignore_index=True)
         if saveIndivdualPred:
         df_indPred = pd.DataFrame()
         df_indPred = df_indPred.append(result["df_indPred"], ignore_index=True)
         del result['df_results'], result['importances_df'], result['df_indPred']
         df_features = df_features.append(result, ignore_index=True)
         df_before.to_csv((resultsPath+f"/prediction_cv_test_{model}.csv"), index=False)
        df_features.to_csv((resultsPath+f"/features_test_{model}.csv"), index=False) df_importances.to_csv((resultsPath+f"/importances_test_{model}.csv"), index=False) df_indPred.to_csv((resultsPath+f"/individualPredictions_test_{model}.csv"), index=False)

Table 5.

Evaluation metrics for binary classification

Evaluation metric Description Equation
Balanced classification accuracya Represents the proportion of accurately predicted cases relative to the total number of cases. TPTP+FN+TNTN+FP2
Precision Represents the proportion of predicted positives that actually belong to the predicted class. TPTP+FP
Recall Represents the proportion of true positive predictions in the actual positive cases. TPTP+FN
F1-score Harmonic mean of precision and recall. 2×precision×recallprecision+recall
AUC A measure of a classification models’ capability to differentiate between different classes (AUC = 1, ideal classification model; AUC = 0.5, equivalent to a random guess). TPR=TPTP+FN
FPR=FPFP+TN

Evaluation metrics are used to assess model performance. Their definition is based on a confusion matrix consisting of true positive (TP), true negative (TN), false positive (FP) and false negative (FN).17 The classifier with the best performance is assigned the value 1, while the classifier with the worst performance is assigned the value 0.

a

Compared to classification accuracy, balanced classification accuracy is a more appropriate metric for the uneven distribution of data between classes, as it treats the number of correct predictions for each class separately and averages the result.

Expected outcomes

The experimental protocol should result in successful extraction and quantification of 10 cholesterol-related sterols in human blood serum samples, including cholesterol. For more details on range of their concentration, please refer to Kočar et al.1 Expected sterol concentrations measured at hospital admission of COVID-19 patients are listed in Table 6.

Table 6.

Concentrations of sterol intermediates in human serum samples in patients with mild and severe COVID-19 measured at hospital admission

Sterol name Concentration [ng/mL]
Mild (N = 14) Severe (N = 48)
lanosterol 48.34 ± 100.98 19.95 ± 11.77
24,25-dihydrolanosterol 4.15 ± 3.83 40.46 ± 228.38
T-MAS 31.18 ± 11.46 36.13 ± 11.16
dihydro-T-MAS 38.83 ± 38.33 48.17 ± 44.41
zymosterol 234.29 ± 146.37 196.70 ± 107.56
zymostenol 593.27 ± 263.01 546.95 ± 352.53
24-dehydrolathosterol 51.84 ± 30.73 36.94 ± 20.31
lathosterol 781.71 ± 677.21 974.93 ± 524.65
desmosterol 538.41 ± 548.88 302.72 ± 112.58
cholesterol 1156.96 ± 385.43 1112.98 ± 271.54

Data are represented as mean ± SD. Mild, patients with mild course of COVID-19; Severe, patients with severe course of COVID-19; SD, standard deviation. This data was previously published by our group.1

The computational pipeline presented should yield an imputed dataset without missingness in independent or dependent variables. The subsequently employed unsupervised iterative variable selection derives a robust subset of variables essential for prediction of the outcome of interest (i.e., disease severity). Furthermore, the designed prediction pipeline provides researchers with nine different hyperparameter-tuned machine learning models, evaluated using five different performance metrics. Lastly, the permutation-based feature importance employed should give a stable estimate of the importance of individual variables. Evaluation metrics results for all tested classifiers are shown in Table 7.

Table 7.

Evaluation metrics for all tested classifiers

Clinical
Model precision recall f1 accuracy AUC
RFC 0.934 0.948 0.941 0.902 0.949
GPR 0.914 0.948 0.930 0.884 0.944
ABC 0.910 0.910 0.910 0.854 0.755
LOG 0.920 0.940 0.930 0.884 0.942
KNN 0.896 0.836 0.865 0.787 0.706
MLP 0.839 0.933 0.883 0.799 0.625
GNB 0.976 0.910 0.942 0.909 0.955
QDA 0.953 0.903 0.927 0.884 0.920
MCL 0.817 1.000 0.899 0.817 0.500
T1 sterols
Model precision recall f1 accuracy AUC
RFC 0.849 0.963 0.902 0.829 0.664
GPR 0.821 0.993 0.899 0.817 0.587
ABC 0.832 0.813 0.823 0.713 0.540
LOG 0.825 0.985 0.898 0.817 0.595
KNN 0.831 0.843 0.837 0.732 0.555
MLP 0.807 0.873 0.839 0.726 0.410
GNB 0.836 0.948 0.888 0.805 0.501
QDA 0.841 0.948 0.891 0.811 0.557
MCL 0.817 1.000 0.899 0.817 0.500
Clinical + T1 sterols
Model precision recall f1 accuracy AUC
RFC 0.926 0.940 0.933 0.890 0.950
GPR 0.920 0.940 0.930 0.884 0.945
ABC 0.947 0.925 0.936 0.896 0.846
LOG 0.907 0.948 0.927 0.878 0.933
KNN 0.888 0.828 0.857 0.774 0.683
MLP 0.844 0.970 0.903 0.829 0.596
GNB 0.983 0.866 0.921 0.878 0.935
QDA 0.961 0.910 0.935 0.896 0.935
MCL 0.817 1.000 0.899 0.817 0.500

ABC, AdaBoost; GNB, Gaussian Naive Bayes; GPR, Gaussian Processes; KNN, K-Nearest Neighbors; LOG, Logistic Regression; MCL, Majority Classifier; MLP, Multilayer Perceptron; RCF, Random Forest; QDA, Quadratic Discriminant Analysis. Previously published by our group,1 Supplementary material.

Limitations

The protocol was designed and optimized to extract and quantify sterol intermediates from the post-squalene part of cholesterol synthesis in human serum samples from patients hospitalized due to COVID-19. One of the limitations of this protocol is that cholesterol must be measured separately in a new independent run, as its serum concentration is significantly higher compared to concentrations of other sterols, which prolongs the duration of the protocol. The second limitation is that we have focused on sterols where standards are commercially available and have not yet focused on sterols whose structure awaits to be determined.

While our prediction models and decision support system demonstrate promising results, it is essential to acknowledge their limitations to ensure their appropriate and informed application. The accuracy of our prediction models may vary based on the dataset and the specific parameters used for training. Although the Leave-One-Out Cross Validation (LOOCV) with iterative permutation-based feature importance procedure utilized in this study distinguishing itself through its high generalizability and robustness, it is crucial to understand that the accuracy can fluctuate depending on the complexity and nature of the data being analyzed. The trust level associated with the evaluated metrics is primarily based on the validation and cross-validation results. While the system aims to minimize errors, it is not infallible and may occasionally produce inaccurate predictions. Therefore, a main obstacle that should be addressed is the lack of validation of the models described above in an external patient cohort.

Due to the LOOCV employed, the calculations suggested in this protocol are computationally expensive and might require the use of high performance computing (HPC) to be conducted in a timely manner.

Troubleshooting

Problem 1

LC-MS/MS analysis, Chromatographic separation: Cholesterol concentrations being significantly higher than those of other sterols.

Potential solution

Perform LC-MS/MS analysis of cholesterol separately from other sterol intermediates. Dilute your samples if necessary – in our case a 5-fold dilution of the samples was used. Optimize the sample injection volume.

Problem 2

Feature selection: Unsupervised variable selection yields no or too few variables.

Potential solution

In case the BorutaPy algorithm is too stringent and yields too few variables, the percentile threshold used to select variables can be lowered (parameter: “perc”); more information can be found in the documentation of BorutaPy (https://github.com/scikit-learn-contrib/boruta_py).

Problem 3

Classification models: Lacking robustness of computational predictions.

Potential solution

Despite the thorough internal validation procedure suggested in this protocol, computational results may fluctuate. While this is to be expected to a certain degree due to the probabilistic nature of machine learning algorithms, this effect can be inflated due to e.g., insufficient sample sizes collection or heavy class imbalance. If the computational resources allow it, we suggest to repeat model training with different random number initializations and subsequent calculations of confidence intervals for predictions. If lack of robustness persists, we advise to explore the impact of variable variance on model performances through e.g., sensitivity analysis and ablation studies.

Resource availability

Lead contact

Further information and requests for resources and reagents should be directed to and will be fulfilled by the lead contact, Damjana Rozman (damjana.rozman@mf.uni-lj.si).

Technical contact

Technical questions on executing this protocol should be directed to and will be answered by the technical contact, Eva Kočar and Sonja Katz (eva.kocar@mf.uni-lj.si and sonja.katz@wur.nl).

Materials availability

This study did not generate new unique reagents.

Data and code availability

The source code and fitted models are available on a public GitHub, copyrighted under the MIT License (https://github.com/sonjakatz/covid_sterols_ML; https://doi.org/10.5281/zenodo.12167403).

Acknowledgments

We thank P. Bogovič, G. Turel, and F. Strle for providing the samples and clinical data and T. Blagus, V. Dolžan, P. Nassib, and J. Stojnić for the help with sample collection.

This work was funded by the Slovenian Research and Innovation Agency (ARIS) program grants P1-0390 and P2-0359 and the PhD grant for young researchers (to E.K.). S.K. was supported by the European Union’s Horizon 2020 research and innovation program under the Marie Sklodowska-Curie (grant agreement no. 860895 TranSYS). We also acknowledge the support by the Network of Research and Infrastructure Centres of University of Ljubljana (MRIC-UL-CFGBC, IP-0022; financed by the ARIS) and infrastructure program ELIXIR-SI (financed by the European Regional Development Fund and by the Ministry of Education, Science, and Sport of the Republic of Slovenia). The funding sources played no role in the study design, data analysis, or writing the manuscript.

Author contributions

Conceptualization, D.R., E.K., and C.S.; methodology, E.K., M. Moškon, Ž.P., and S.K.; software, M. Moškon, Ž.P., and S.K.; formal analysis, E.K., M. Moškon, Ž.P., and S.K.; resources, D.R., M. Moškon, and M. Mraz; data curation, S.K. and E.K.; writing – original draft, E.K. and S.K.; writing – review and editing, D.R., M. Moškon, C.S., and T.R.; visualization, E.K.; supervision, D.R., M. Moškon, and M. Mraz; project administration, D.R., M. Moškon, and E.K.; funding acquisition, D.R., M. Mraz, and V.A.P.M.d.S.

Declaration of interests

The authors declare no competing interests.

Contributor Information

Miha Moškon, Email: miha.moskon@fri.uni-lj.si.

Damjana Rozman, Email: damjana.rozman@mf.uni-lj.si.

References

  • 1.Kočar E., Katz S., Pušnik Ž., Bogovič P., Turel G., Skubic C., Režen T., Strle F., Martins dos Santos V.A.P., Mraz M., et al. COVID-19 and cholesterol biosynthesis: Towards innovative decision support systems. iScience. 2023;26 doi: 10.1016/j.isci.2023.107799. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Skubic C., Vovk I., Rozman D., Križman M. Simplified LC-MS method for analysis of sterols in biological samples. Molecules. 2020;25 doi: 10.3390/molecules25184116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Brown A.J., Sharpe L.J. In: Lipoproteins and Membranes. Sixth Edition. Ridgway N.D., McLeod R.S., editors. Elsevier; 2016. Chapter 11 - Cholesterol Synthesis; pp. 327–358. [DOI] [Google Scholar]
  • 4.Skubic C., Rozman D. In: Mammalian Sterols: Novel Biological Roles of Cholesterol Synthesis Intermediates, Oxysterols and Bile Acids. Rozman D., Gebhardt R., editors. Springer International Publishing; 2020. Sterols from the Post-Lanosterol Part of Cholesterol Synthesis: Novel Signaling Players; pp. 1–22. [DOI] [Google Scholar]
  • 5.Kandutsch A.A., Russell A.E. Preputial gland tumor sterols. 3. A metabolic pathway from lanosterol to cholesterol. J. Biol. Chem. 1960;235:2256–2261. doi: 10.1016/s0021-9258(18)64608-3. [DOI] [PubMed] [Google Scholar]
  • 6.Pedregosa F., Varoquaux G., Gramfort A., Michel V., Thirion B., Grisel O., Blondel M., Prettenhofer P., Weiss R., Dubourg V., et al. Scikit-learn: Machine Learning in Python. J. Mach. Learn. Res. 2011;12:2825–2830. [Google Scholar]
  • 7.Kursa M.B., Rudnicki W.R. Feature selection with the boruta package. J. Stat. Softw. 2010;36:1–13. doi: 10.18637/jss.v036.i11. [DOI] [Google Scholar]
  • 8.Python implementations of the Boruta all-relevant feature selection method. 2023. https://github.com/scikit-learn-contrib/boruta_py
  • 9.Thongkam J., Xu G., Zhang Y. IEEE World Congress on Computational Intelligence; 2008. AdaBoost Algorithm with Random Forests for Predicting Breast Cancer Survivability. IEEE International Joint Conference on Neural Networks; pp. 3062–3069. [DOI] [Google Scholar]
  • 10.Rish I. An empirical study of the naive Bayes classifier. IJCAI 2001 workshop on empirical methods in artificial intelligence. International Joint Conference on Artificial Intelligence (IJCAI) 2001:41–46. [Google Scholar]
  • 11.Rasmussen C.E. Gaussian Processes in Machine Learning. Lect. Notes Comput. Sci. 2004;3176:63–71. doi: 10.1007/978-3-540-28650-9_4. [DOI] [Google Scholar]
  • 12.Kramer O. Dimensionality Reduction with Unsupervised Nearest Neighbors. Springer; 2013. K-Nearest Neighbors; pp. 13–23. [DOI] [Google Scholar]
  • 13.Dreiseitl S., Ohno-Machado L. Logistic regression and artificial neural network classification models: a methodology review. J. Biomed. Inform. 2002;35:352–359. doi: 10.1016/S1532-0464(03)00034-0. [DOI] [PubMed] [Google Scholar]
  • 14.Gardner M.W., Dorling S.R. Artificial neural networks (the multilayer perceptron) - a review of applications in the atmospheric sciences. Atmos. Environ. 1998;32:2627–2636. doi: 10.1016/S1352-2310(97)00447-0. [DOI] [Google Scholar]
  • 15.Breiman L. Random forest. Mach. Learn. 2001;45:5–32. doi: 10.1023/A:1010933404324. [DOI] [Google Scholar]
  • 16.Tharwat A. Linear vs. quadratic discriminant analysis classifier: a tutorial. International Journal of Applied Pattern Recognition. 2016;3:145. doi: 10.1504/ijapr.2016.079050. [DOI] [Google Scholar]
  • 17.Luque A., Carrasco A., Martín A., de las Heras A. The impact of class imbalance in classification performance metrics based on the binary confusion matrix. Pattern Recogn. 2019;91:216–231. doi: 10.1016/j.patcog.2019.02.023. [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The source code and fitted models are available on a public GitHub, copyrighted under the MIT License (https://github.com/sonjakatz/covid_sterols_ML; https://doi.org/10.5281/zenodo.12167403).


Articles from STAR Protocols are provided here courtesy of Elsevier

RESOURCES