Skip to main content
PLOS One logoLink to PLOS One
. 2024 Apr 2;19(4):e0299600. doi: 10.1371/journal.pone.0299600

Machine learning evaluation for identification of M-proteins in human serum

Alexandros Sopasakis 1, Maria Nilsson 2, Mattias Askenmo 2, Fredrik Nyholm 2, Lillemor Mattsson Hultén 2,3, Victoria Rotter Sopasakis 2,4,*
Editor: John Adeoye5
PMCID: PMC10986985  PMID: 38564628

Abstract

Serum electrophoresis (SPEP) is a method used to analyze the distribution of the most important proteins in the blood. The major clinical question is the presence of monoclonal fraction(s) of antibodies (M-protein/paraprotein), which is essential for the diagnosis and follow-up of hematological diseases, such as multiple myeloma. Recent studies have shown that machine learning can be used to assess protein electrophoresis by, for example, examining protein glycan patterns to follow up tumor surgery. In this study we compared 26 different decision tree algorithms to identify the presence of M-proteins in human serum by using numerical data from serum protein capillary electrophoresis. For the automated detection and clustering of data, we used an anonymized data set consisting of 67,073 samples. We found five methods with superior ability to detect M-proteins: Extra Trees (ET), Random Forest (RF), Histogram Grading Boosting Regressor (HGBR), Light Gradient Boosting Method (LGBM), and Extreme Gradient Boosting (XGB). Additionally, we implemented a game theoretic approach to disclose which features in the data set that were indicative of the resulting M-protein diagnosis. The results verified the gamma globulin fraction and part of the beta globulin fraction as the most important features of the electrophoresis analysis, thereby further strengthening the reliability of our approach. Finally, we tested the algorithms for classifying the M-protein isotypes, where ET and XGB showed the best performance out of the five algorithms tested. Our results show that serum capillary electrophoresis combined with decision tree algorithms have great potential in the application of rapid and accurate identification of M-proteins. Moreover, these methods would be applicable for a variety of blood analyses, such as hemoglobinopathies, indicating a wide-range diagnostic use. However, for M-protein isotype classification, combining machine learning solutions for numerical data from capillary electrophoresis with gel electrophoresis image data would be most advantageous.

Introduction

The analysis of fractional proteins in serum and urine (SPEP and UPEP) is included in standardized care in case of suspicion of haematological malignant chronic disease, such as multiple myeloma. Accurate diagnosis is an essential component of optimal cancer care. M-protein (also called myeloma protein or paraprotein) denotes an antibody or a fragment of an antibody that is produced in abnormal amounts by a pre-malignant or malignant plasma cell and is used as a marker for haematological malignant diseases such as myeloma as well as the asymptomatic disease monoclonal gammopathy of undetermined significance (MGUS). The size of the M-protein fraction reflects the degree of disease severity, and the level is regularly checked to follow treatment efficiency. Serum protein electrophoresis (SPEP) separates serum globulins based on their physical properties and is a diagnostic tool used to identify M-protein. In the case of malignancy, normal polyclonal immunoglobulin (Ig) production is often suppressed, and the M-protein is identified as an unusually distinct peak or atypical curve appearance within the gamma globulin fraction or beta globulin fraction, depending on the type and quantity of M-protein present.

Present-time SPEP is a semi-automated process, but involves defined laboring steps, and may be affected by intra- and inter-observer variability. The need for this type of analysis is expected to grow with an expanding population and an increased number of elderly people, which entails enhanced cancer risk, further increasing the requirement for medical expertise.

Recent studies show that machine learning can be used to assess protein electrophoresis by examining protein glycan patterns to follow up tumor surgery [1]. For example, changes in N-glycosylation patterns were analyzed using a computer-assisted machine learning method for interpreting serum proteins after surgical lung tumor resection [1]. The classification analysis resulted in a panel of N-glycans, which could be used to follow up on the effects of surgical resection of lung tumors [1]. In addition, decision tree based machine learning has been used in interpreting amino acid patterns in plasma as well as in studies based on raw signal data, such as electrocardiogram features in patients with cardiac arrhythmia, EEG signals in hearing test processes and protein mass spectra signals in serum from patients with non small cell lung cancer [2–5]. Furthermore, other types of machine learning tools have successfully been implemented in nuclear medicine and radiology for assessment of Positron emission tomography-computed tomography PET-CT or PET/CT images from 100 different organs [6,7]. Assessment using the artificial intelligence (AI)-based analysis tool Recomia showed an accuracy coefficient of 0.93 [7].

Implementing machine learning algorithms as a clinical diagnostic decision support for medical expertise in the analysis of fractionated serum proteins provides a possibility to improve quality, safety and efficiency for patients, physicians and biomedical staff and save medical resources. Thus, in this study, we investigated and compared several different decision tree algorithms for the detection of M-protein in 67,073 serum samples from 32,490 patients. Raw numerical signal data from the capillary SPEP was used and the different algorithms were evaluated for their detection accuracy as well as used for identification and verification of critical features in the data toward diagnosis.

Materials and methods

This study was approved by the Swedish Ethical Review Authority, project ID 2021–03301, and performed in accordance with the Helsinki Declaration (as revised in 2013) and aligns with the STARD recommendations for reporting diagnostics studies. The study is a retrospective study with anonymized human clinical data with no requirement for informed consent. The research data are by law subject to the Public Access to Information and Secrecy Act (SFS 2009:400)

Serum protein separation

The serum samples used for this study were submitted to the laboratory for serum protein analysis in the Region Västra Götaland, Sweden, during 2015–2020. The data was accessed for research on July 5th, 2021. All samples were analyzed on both an automated capillary electrophoresis system (Capillarys 2 Flex Piercing, SEBIA, Issy-les-Moulineaux, France), for quantification, as well as by a semi-automated agarose gel technology system (serum gel electrophoresis) (Hydrasys 2 System and PENTA PN 1260 kit, SEBIA, Issy-les-Moulineaux, France) [8,9] to not miss small fractions of M-protein. For samples where a suspected M-protein was present (based on the results from the capillary SPEP and gel SPEP), immunofixation (Hydrasys 2 System and Hydrasys IF kit, SEBIA, Issy-les-Moulineaux, France) was performed to determine the type of M-protein. M-proteins were quantified based on the capillary zone electrophoresis (CZE) (S1 Fig) using the SEBIA software (Phoresis, SEBIA, Issy-les-Moulineaux, France).

Data processing

For the training and classification of M-proteins we used the labeled anonymized data set described above and in the Result section. All data were assessed by medical experts who determined whether M-protein was present or not based on CZE from capillary gel electrophoresis (S1 Fig) and gel pictures, and, when required, immunofixation gel pictures.

The data was randomly divided into 70% for training and 30% for testing. Using the training data, we initially taught 26 decision trees methods to detect the M-proteins. These 26 decision trees, which are part of the python sklearn package LightGBM, can quickly ascertain which methods are best for detecting M-proteins. Subsequently, we picked five of those algorithms indicated by LightGBM as performing best and trained individual python algorithms for each of those. For the individual training we randomly divided the data into 90% for training and 10% for testing. To train each of these decision trees, the algorithm minimizes the misclassification cost and specifically, in the case of boosted methods, errors of previous trees in fitting the data. During training, only the training set was used to exclusively train the models, avoiding exposure to the test set. This precaution guards against overfitting as well as data leakage, preserving the models’ generalization ability. We also note that some of the models trained, such as for instance XGB, incorporate mechanisms to mitigate overfitting, ensuring robust generalization. Pruning and bagging techniques were implemented to effectively reduce the size of the resulting trees and to reduce the variance in the results of the final trees. To evaluate the performance of each tree we used metrics, such as accuracy, balanced accuracy, ROC-AUC and F1-score. The hyperparameters for all algorithms implemented in this work are set to their default values.

We installed python and relevant software libraries and programs using a docker environment with the following specifications:

Python 3.8.10, gcc 9.4.0, keras 2.6.0, open_cv 4.5.4, pyforest 1.1.0. Pandas 1.0.5 with Matplotlib 3.4.3 was used for data upload and processing. All decision trees were also based on this same version of the sklearn library. The specific parameters used for each of these trees are detailed in the Supplementary Materials and Methods section. Feature importance was based on the Shapley values from the python package shap version 0.37.0.

To evaluate the ability of the algorithm to determine the M-protein isotypes, we used precision scores and recall scores, as well as confusion matrix analysis.

Results

For this study we used a total of 67,073 patient serum samples from 32,940 individuals of which 15,684 samples were positive for M-protein. The average age was 57±21.5 years. The gender distribution was 15,164 men and 17,776 women. The patient demographics and distribution of monoclonal isotypes are shown in Fig 1, Tables 1 and 2 and S2 Fig.

Fig 1. Patient demographics and M-protein isotype distribution.

Fig 1

(A), Age distribution and number of individuals within age quartiles for the individuals included in the study, total n = 32,940. Age is shown as mean ± standard deviation. (B), M-protein isotype distribution among the M-protein positive samples included in the study, n = 15,684.

Table 1. M-protein isotype distribution among the M-protein positive samples included in the study.

M-protein isotype Number of M-protein positive samples %
IgG 10,538 67.2
IgM 2,215 14.1
IgA 1,745 11.1
IgD 9 0.06
FLC 404 2.6
IgG + IgM 60 0.4
IgG + IgA 33 0.2
IgM + IgA 6 0.04
IgG + IgD 1 0.01
IgG + FLC 488 3.1
IgA + FLC 102 0.7
IgM + FLC 81 0.5
IgD + FLC 2 0.01
Total 15,684  

Table 2. Distribution of concentrations for IgG, IgA and IgM M-proteins, divided into three groups: <1 g/l, 1–5 g/l and >5 g/l.

M-protein concentration <1 g/l 1–5 g/l >5 g/l
Number of IgG M-protein samples 16 2,822 7,700
Number of IgA M-protein samples 3 396 1,346
Number of IgM M-protein samples 3 663 1,549
Total IgG+IgA+IgM M-protein samples 22 3,881 10,595

Decision tree methods detect M-proteins with high accuracy

In order to compare the accuracy of different decision tree-based methods to detect the existence of M-proteins in the serum samples, we evaluated 26 methods by training them for identification of M-proteins on 50,391 samples and then testing them on 16,797 unseen samples. The methods showing the highest accuracy were Extra Trees (ET), Random Forest (RF), Histogram Grading Boosting Regressor (HGBR), Light Gradient Boosting Method (LGBM), and Extreme Gradient Boosting (XGB), resulting in accuracy scores between 0.962 and 0.977 (Table 3). Further classification testing using receiver operating characteristic (ROC) area under the curve (AUC) calculations and F1 score calculations confirmed superior ability of these methods to classify unseen data (Table 3 and Fig 2). The ROC AUC scores reached as high as 0.993 (Table 3 and Fig 2) and the F1 scores reached as high as 0.977 (Table 3). Accuracy scores, ROC AUC values and F1 scores for the remaining methods tested are listed in S1 Table. The computational cost of training all the algorithms was in total less than 20 minutes in a CPU Intel i7-9700E 8-core processor. Once trained each algorithm was able to produce results in real time.

Table 3. The top five decision tree methods for identification of M-protein.

Classifier Accuracy Balanced Accuracy ROC AUC F1 Score
ET 0.977 0.950 0.993 0.977
RF 0.973 0.942 0.989 0.973
HGBR 0.962 0.942 0.988 0.961
LGBM 0.962 0.902 0.987 0.961
XGB 0.965 0.912 0.985 0.964

ET = Extra Trees, RF = Random Forest, HGBR = Histogram Grading Boosting Regressor, Light LGBM = Gradient Boosting Method, and XGB = Extreme Gradient Boosting.

Fig 2. Receiver operating characteristic (ROC) calculations.

Fig 2

ROC curves for the five most successful decision tree algorithms: Extra Trees (ET), Random Forest (RF), Histogram Grading Boosting Regressor (HGBR), Light Gradient Boosting Method (LGBM), and Extreme Gradient Boosting (XGB). A less successful algorithm (ADA) is included as a reference. The inset shows the corresponding area under the curve (AUC) values.

M-protein detection by decision tree methods can be verified to the gamma and beta fractions

The majority of machine learning algorithms are unable to communicate how they produce a particular outcome, a phenomenon known as the “black box” [10]. However, this knowledge is important for our purposes in terms of understanding how the input data leads to the M-protein diagnosis as well as verifying that the algorithm is processing the input data correctly.

We aimed to resolve this issue using Shapley Additive explanations (SHAP) [11,12]. This method allows us to understand which of the input data features are most important in order to discriminate between samples containing M-proteins and samples not containing M-proteins. Practically, this task is equivalent to optimizing the classification outcome of the algorithm under all possible combinations of input data features. A successful outcome was one which, for our application, correctly predicted the presence of M-protein. The CZE input data was divided into 300 features, each corresponding to the migration time on the x-axis of the CZE (Fig 3A). The results verified the gamma globulin fraction and part of the beta-2 globulin fraction as the most important for detecting M-protein (Figs 3A and S3), corresponding to the features with migration times between 240 and 262 seconds in the CZE (Fig 3B), thereby verifying the results and further strengthening the reliability of the top five algorithms. Furthermore, using SHAP values also enabled us to compute an M-protein probability score for each patient (Fig 3C), which can be used to predict the likelihood of M-protein being present for a specific patient. For comparison, we also performed the same analysis with a classic feature importance approach, which gave limited results (S4 Fig). Whereas the classic approach primarily identifies features corresponding to the gamma fraction, SHAP identifies features for both the beta and gamma fractions (Figs 3C and S4).

Fig 3. Calculations for Shapley Additive explanations (SHAP).

Fig 3

(A), Capillary zone electrophoresis (CZE) from a sample containing M-protein. The x-axis is divided into 300 time points (features), representing the migration time during SPEP. The most important features for detecting M-protein based on the SHAP calculations are highlighted in yellow, clearly verifying the gamma globulin fraction and part of the beta-2 globulin fraction as the most important for detecting M-protein. (B), The ten most important features (from a total of 300) for detection of M-protein. Each feature value corresponds to the migration time on the x-axis of the CZE. The x-axis shows the mean SHAP value. (C), SHAP values for a sample with no M-protein (top) and with M-protein (bottom). Red numbers indicate positively correlated features for detecting M-protein and blue numbers indicate negatively correlated features. Probability scores for the presence of M-protein in the samples are shown in bold numbers. Features with their corresponding SHAP values are indicated below each panel.

Extra trees and random forest show the best success rate in classifying M-protein isotypes

We next wanted to test the ability of the top five scoring methods in the initial analysis above (ET, RF, HGBR, LGBM, and XGB) to correctly determine the isotype of the M-protein present in a sample. The scoring for small M-spikes was poorer than larger ones as expected (data not shown), but to what degree is impossible to evaluate with the low number of samples we have in the small size category in our data set. For this purpose, we had to exclude the samples with free light chain M-protein, IgD M-protein, and samples with more than one type of M-protein in our data set, since these samples were too few to be properly trained and tested by the algorithms. Out of the 6,704 samples used for testing, the ET algorithm was most successful in scoring IgG M-proteins, with an F1 score of 0.87 (Table 4). The Random forest and XGB algorithms reached an F1 score of 0.44 for IgA M-proteins (Table 4), whereas XGB reached an F1 score of 0.35 for IgM proteins.

Table 4. Ability of the five best algorithms to determine the M-protein isotypes.

Extra trees Total true number in the testing population Precision score Recall score F1 score
Healthy 5,300 (79.1%) 0.96 0.99 0.98
IgG isotype 1,020 (15.2%) 0.84 0.9 0.87
IgA isotype 185 (2.7%) 0.95 0.29 0.44
IgM isotype 199 (3.0%) 0.95 0.2 0.32
Random forest Total true number in the testing population Precision score Recall score F1 score
Healthy 5,300 (79.1%) 0.95 0.99 0.97
IgG isotype 1,020 (15.2%) 0.81 0.87 0.84
IgA isotype 185 (2.7%) 0.98 0.25 0.40
IgM isotype 199 (3.0%) 0.91 0.15 0.25
HGBR Total true number in the testing population Precision score Recall score F1 score
Healthy 5,300 (79.1%) 0.93 1.00 0.96
IgG isotype 1,020 (15.2%) 0.81 0.75 0.78
IgA isotype 185 (2.7%) 0.97 0.21 0.34
IgM isotype 199 (3.0%) 0.76 0.08 0.15
LGBM Total true number in the testing population Precision score Recall score F1 score
Healthy 5,300 (79.1%) 0.93 1.00 0.96
IgG isotype 1,020 (15.2%) 0.81 0.76 0.79
IgA isotype 185 (2.7%) 0.97 0.21 0.34
IgM isotype 199 (3.0%) 0.84 0.08 0.15
XGB Total true number in the testing population Precision score Recall score F1 score
Healthy 5,300 (79.1%) 0.94 1.00 0.97
IgG isotype 1,020 (15.2%) 0.83 0.77 0.80
IgA isotype 185 (2.7%) 0.95 0.29 0.44
IgM isotype 199 (3.0%) 0.62 0.24 0.35

Total number of samples in the test population = 6,704.

Precision score: Proportion of correctly predicted M-protein isotypes in the test sample population.

Recall score: Proportion of correctly predicted M-protein isotypes out of the true isotype populations.

F1 score = the harmonic mean of the precision and the recall scores.

To evaluate the isotype classification performance of the algorithms in more detail, a confusion matrix analysis was performed for each of the five algorithms (Table 5). This analysis verified ET and XGB as the most successful algorithms in scoring IgG, IgA and IgM, and, in addition, showed that ET had fewer false negative scores compared to the other four algorithms. Taken together, out of the 6,704 samples in the test population, ET and XGB showed a higher classification success rate compared to RF, HGBR, and LGBM.

Table 5. Confusion matrices showing the isotype classification performance for the five best algorithms.

Extra trees Healthy IgG isotype IgA isotype IgM isotype
Healthy 5273 25 2 0
IgG isotype 101 916 1 2
IgA isotype 111 21 53 0
IgM isotype 27 133 0 39
Proportion false positives 4% 5% 5% 16%
Random forest Healthy IgG isotype IgA isotype IgM isotype
Healthy 5261 38 1 0
IgG isotype 132 885 0 3
IgA isotype 108 31 46 0
IgM isotype 29 141 0 29
Proportion false positives 5% 2% 9% 19%
HGBR Healthy IgG isotype IgA isotype IgM isotype
Healthy 5280 20 0 0
IgG isotype 245 770 0 5
IgA isotype 124 23 38 0
IgM isotype 44 138 1 16
Proportion false positives 7% 3% 24% 19%
LGBM Healthy IgG isotype IgA isotype IgM isotype
Healthy 5277 23 0 0
IgG isotype 238 778 1 3
IgA isotype 123 24 38 0
IgM isotype 51 132 0 16
Proportion false positives 7% 3% 16% 19%
XGB Healthy IgG isotype IgA isotype IgM isotype
Healthy 5275 22 2 1
IgG isotype 211 781 0 28
IgA isotype 102 30 53 0
IgM isotype 43 107 1 48
Proportion false positives 6% 5% 38% 17%

Total number of samples in the test population = 6,704.

Rows show actual classification numbers.

Columns show predicted isotype values.

Boxes highlighted in green represent true positives.

Remaining values for each column represent false positives.

Remaining values for each row represent false negatives.

Example: Out of a total of 6,704 samples, 5,273 samples were correctly classified as healthy by the Extra tree algorithm. 101 samples were falsely classified as healthy, but were positive for IgG M-protein. 2 healthy samples were falsely classified as IgA M-protein.

Discussion

SPEP is a standard screening method for evaluating immunoglobulin patterns. A specific group of diseases, gammopathies, gives rise to an atypical peak in the histogram gamma globulin or beta globulin regions following serum electrophoresis due to the presence of M-protein (or paraprotein), caused by abnormal proliferation of a single clone of plasma cells.

Recent reports have shown that machine learning can be a useful decision support tool for various clinical analyses including examining protein glycan patterns to follow up the effects of surgical resection of lung tumors [1], interpreting amino acid patterns in plasma as well as in studies using raw signal data rather than data related to concentration of specific molecules [3–5]. Thus, in the current study, we aimed to investigate and compare several different machine learning algorithms for M-protein identification in human serum samples following SPEP.

Using 26 different decision tree algorithms we trained and tested them on our data set, which included 67,073 patient serum samples, to evaluate which methods were superior in accurately determining the presence of M-protein. Five methods, Extra Trees, Random Forest, Histogram Grading Boosting Regressor, Light Gradient Boosting Method, and Extreme Gradient Boosting, were found to generate highly accurate identification scores, ranging from 96% to 98% accuracy and ROC AUCs of 0.99 for all methods. Chabrun et al recently reported an AI application for SPEP analysis, based on four deep-learning models that included numerous parameters [13]. Their AI application generated M-protein classification accuracy scores of 91.2% - 93.1% although the ROC AUC reached as high as 0.99. The discrepancy in accuracy score between our model and the model by Chabrun et al is likely most attributable to less precise annotations due to difficulty in clearly defining M-spikes in some SPEP samples, a common problem when interpreting conditions with restricted heterogeneity in immunoglobulin fractions. Our data was based on samples analyzed by multiple assays, which allowed robust results and thereby more clearly detected and defined M-spikes.

Complex machine learning algorithms, including deep learning models, as well as methods like XGBoost and LightGBM, share a common limitation referred to as the ’black box’ phenomenon, where understanding the internal decision-making process leading to outcomes becomes challenging [10]. This often leads to uncertainty regarding the reliability of the method used. In addition, the “black box” phenomenon prevents from identifying novel patterns and properties within the data set that could contribute to better understand and interpret underlying features of a disease. Explanation methods, such as Shapley Additive explanations (SHAP), can be employed to enhance the interpretability of these models by providing insights into feature importance and contributing factors. Other groups have tried to combine deep learning algorithms with decision trees [14,15] or applied natural language processing approaches [16] to classify and predict multiple myeloma or other diseases, all with accuracy scores of <93%. Combining deep neural networks and decision trees requires heavy processing of input data and substantial tuning of model parameters which can explain the lower predictor accuracy and F1 scores making these types of models difficult to apply in practice.

While decision tree methods can offer increased transparency due to their rule-based nature, it is important to note that this transparency might diminish as we move to more complex algorithms like ensemble models, including XGBoost (XGB) and LightGBM (LGBM). These ensemble methods can introduce additional complexity that could impact the interpretability of the models and can make it harder to identify suitable values of their hyperparameters. By utilizing decision tree algorithms based on the raw numerical data generated from capillary gel electrophoresis, we were able to elucidate the features responsible for the detection of M-proteins using Shapley Additive explanations (SHAP) [11,12], thereby verifying the gamma globulin fraction and part of the beta-2 globulin fraction as the M-protein hallmarks for the algorithm outcome (M-protein of IgA type is generally found in the beta-2 zone). This is a crucial element of our approach, not only for verifying already known patterns and features, but particularly for identifying possible novel patterns when using the method for other applications. Elucidating previously unidentified patterns and features in a dataset could largely improve the understanding and interpretation of the analysis results, thereby increasing the sensitivity and quality of the analysis. Ultimately, this leads to better care for the patient.

Taking advantage of machine learning methods as a support tool for the evaluation and determination of blood proteins creates possibilities to not only increase the accuracy of the results and safety for the patient, but possibly also extensive savings of resources. By using machine learning methods to eliminate negative samples (e.g., samples with no M-protein present), the medical and biomedical staff can focus on the positive samples (e.g., determining the type of M-protein present in the sample). Implementing SHAP calculations further adds value in this context by facilitating predictions and introduction of appropriate cut off margins. Furthermore, using machine learning methods as a support tool can help reduce the amount of unnecessary additional analyses which are often performed due to uncertainty and/or inexperience of the person analyzing the initial electrophoresis result.

A limitation with our approach was the low accuracy scores for identifying the specific isotype of M-proteins present. In the context of identifying specific M-protein isotypes, it is important to consider the potential impact of class imbalance within the data set. The relatively low accuracy scores observed in our approach may be attributed to the presence of certain M-protein isotypes that are inherently less prevalent within the population. This imbalance in class distribution can pose challenges for accurate classification. Decision tree methods are generally effective in tackling class imbalance within the data set [17] by their inherent ability to prioritize minority classes [18], thanks to impurity-based splitting criteria like the Gini index or entropy or ensemble methods [19] like Random Forests or gradient boosting. However, to mitigate this issue and enhance the model’s ability to discern rarer isotypes, techniques such as SMOTE (Synthetic Minority Over-sampling Technique) could be explored. By generating synthetic instances of underrepresented isotypes, SMOTE aims to balance class proportions and improve model performance. The incorporation of such methods could help to further address the impact of class imbalance and potentially lead to more accurate identification of specific M-protein isotypes.

Other likely contributing factors for the low accuracy scores for identifying specific M-protein isotypes is that samples with small M-protein fractions often are difficult to distinguish in the CZE. Furthermore, abnormal patterns other than those resulting from the presence of M-proteins could confound the results generated by a machine learning tool. Such abnormal patterns could be identified in future work using anomaly detection methods. For a more accurate scoring of M-protein isotypes, a combination of decision tree methods and deep learning algorithms of immunofixation electrophoresis, such as that recently described by Hu et al [20], would likely be more advantageous.

The incidence of samples with free light chains in our data set is low compared to what is generally observed and reported in the literature. Light chain multiple myeloma (LCMM) accounts for approximately 15% of all cases of multiple myeloma [21] and free light chain M-proteins are often observed in small quantities in serum when analyzed by SPEP. Since M-proteins of very small concentrations are difficult to detect in the CZE, it is important to take into account that this may have influenced our results.

In conclusion, we found that decision tree machine learning algorithms can identify the presence of M-proteins in human serum following SPEP with high accuracy, using routinely collected laboratory data. In addition, we were able to verify the gamma globulin and beta-2 globulin fractions as the hallmark for the algorithm outcome, further enhancing the reliability of the method. The use of machine learning algorithms for the prediction of overall prognosis is a promising new area that can support doctors with clinical decisions, improve diagnosis accuracy and safety for the patients as well as markedly reduce the volume of resources needed.

Supporting information

S1 Fig. Capillary zone electrophoresis (CZE) outlining the different serum protein fractions in a normal sample and in a sample containing M-protein.

(TIF)

pone.0299600.s001.tif (67.6MB, tif)
S2 Fig. Concentration range of free light chains among the samples in the data set.

(TIF)

pone.0299600.s002.tif (99.6MB, tif)
S3 Fig. SHAP calculation results from nine samples containing M-protein and three samples with no M-protein.

(TIF)

pone.0299600.s003.tif (67.6MB, tif)
S4 Fig. Classic feature importance approach based on the trained Extra Trees Classifier.

(TIF)

pone.0299600.s004.tif (99.6MB, tif)
S1 Table. Accuracy scores, ROC AUC values and F1 scores for the remaining methods tested.

(TIF)

pone.0299600.s005.tif (97.3MB, tif)

Abbreviations

CZE

capillary zone electrophoresis

ET

Extra trees

RF

Random Forest

HGBR

Histogram Grading Boosting Regressor

LGBM

–Light Gradient Boosting Method

XGB

Extreme Gradient Boosting

CE

capillary electrophoresis

ROC

Receiver Operating Characteristic

AUC

Area under the curve

SHAP

Shapley Additive explanations

Ig

Immunoglobulin

SPEP

Serum Protein Electrophoresis

UPEP

Urine Protein Electrophoresis

Data Availability

No - some restrictions will apply; The data have been uploaded in a public data repository, the Swedish National Data Service (https://snd.gu.se/en), with metadata and documentation made available according to the FAIR principles, and it has been assigned a permanent identifier (DOI): https://doi.org/10.5878/a2aa-kt50. To access the actual data files, a request management system is available to file a formal request to the University of Gothenburg for the data at https://snd.gu.se/en or snd@snd.gu.se. The code can be accessed through github at https://github.com/a0s6044/ProteinElectrophoresis.

Funding Statement

This work is supported by the department of Clinical Chemistry, Sahlgrenska University Hospital, Gothenburg, Sweden (VRS, LMH, MN, FN, MA). There is no other specific funding for this work. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

References

  • 1.Meszaros B, Jarvas G, Kun R, Szabo M, Csanky E, Abonyi J, et al. Machine Learning Based Analysis of Human Serum N-glycome Alterations to Follow up Lung Tumor Surgery. Cancers (Basel). 2020;12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Wilkes EH, Emmett E, Beltran L, Woodward GM, Carling RS. A Machine Learning Approach for the Automated Interpretation of Plasma Amino Acid Profiles. Clin Chem. 2020;66:1210–8. doi: 10.1093/clinchem/hvaa134 [DOI] [PubMed] [Google Scholar]
  • 3.Alarsan FI, Younes M. Analysis and classification of heart diseases using heartbeat features and machine learning algorithms. J Big Data-Ger. 2019;6. [Google Scholar]
  • 4.Kucukakarsu M, Kavsaoglu AR, Alenezi F, Alhudhaif A, Alwadie R, Polat K. A Novel Automatic Audiometric System Design Based on Machine Learning Methods Using the Brain’s Electrical Activity Signals. Diagnostics (Basel). 2023;13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Monari E, Casali C, Cuoghi A, Nesci J, Bellei E, Bergamini S, et al. Enriched sera protein profiling for detection of non-small cell lung cancer biomarkers. Proteome Sci. 2011;9:55. doi: 10.1186/1477-5956-9-55 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Borrelli P, Larsson M, Ulen J, Enqvist O, Tragardh E, Poulsen MH, et al. Artificial intelligence-based detection of lymph node metastases by PET/CT predicts prostate cancer-specific survival. Clin Physiol Funct Imaging. 2021;41:62–7. doi: 10.1111/cpf.12666 [DOI] [PubMed] [Google Scholar]
  • 7.Tragardh E, Borrelli P, Kaboteh R, Gillberg T, Ulen J, Enqvist O, et al. RECOMIA-a cloud-based platform for artificial intelligence research in nuclear medicine and radiology. EJNMMI Phys. 2020;7:51. doi: 10.1186/s40658-020-00316-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Larsson A, Hansson LO. Analysis of inflammatory response in human plasma samples by an automated multicapillary electrophoresis system. Clin Chem Lab Med. 2004;42:1396–400. doi: 10.1515/CCLM.2004.260 [DOI] [PubMed] [Google Scholar]
  • 9.Larsson A, Hansson LO. Comparison between a second generation automated multicapillary electrophoresis system with an automated agarose gel electrophoresis system for the detection of M-components. Ups J Med Sci. 2008;113:65–72. doi: 10.3109/2000-1967-219 [DOI] [PubMed] [Google Scholar]
  • 10.Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence. 2019;1:206–15. doi: 10.1038/s42256-019-0048-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Lundberg SM LS. A Unified Approach to Interpreting Model Predictions. Neural Information Processing Systems. Long Beach, CA 2017. p. 1–10. [Google Scholar]
  • 12.Lundberg SM EG, Lee SI. Consistent Individualized Feature Attribution for Tree Ensembles. arXiv. https://arxiv.org/abs/1802.038882018.
  • 13.Chabrun F, Dieu X, Ferre M, Gaillard O, Mery A, Chao de la Barca JM, et al. Achieving Expert-Level Interpretation of Serum Protein Electrophoresis through Deep Learning Driven by Human Reasoning. Clin Chem. 2021;67:1406–14. doi: 10.1093/clinchem/hvab133 [DOI] [PubMed] [Google Scholar]
  • 14.Park DJ, Park MW, Lee H, Kim YJ, Kim Y, Park YH. Development of machine learning model for diagnostic disease prediction based on laboratory tests. Sci Rep. 2021;11:7567. doi: 10.1038/s41598-021-87171-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Yan W, Shi H, He T, Chen J, Wang C, Liao A, et al. Employment of Artificial Intelligence Based on Routine Laboratory Results for the Early Diagnosis of Multiple Myeloma. Front Oncol. 2021;11:608191. doi: 10.3389/fonc.2021.608191 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Ryu JH, Zimolzak AJ. Natural Language Processing of Serum Protein Electrophoresis Reports in the Veterans Affairs Health Care System. JCO Clin Cancer Inform. 2020;4:749–56. doi: 10.1200/CCI.19.00167 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Mo Ahsan Ahmad UK, Faizan Ansari, Nafees Mo, Ravindranath Sawane. Comparison of Various Machine Learning Techniques Based on Variable Selection under Imbalanced Data. International Journal of Engineering Development and Research. 2022;10:60–70. [Google Scholar]
  • 18.Rodríguez JJ, Díez-Pastor J.F García-Osorio C. Ensembles of Decision Trees for Imbalanced Data. In: Sansone C, Kittler J., Roli F., editor. Multiple Classifier Systems. Berlin, Heidelberg: Springer; 2011. p. 76–85. [Google Scholar]
  • 19.Polikar R. Ensemble Learning. In: Zhang C, Ma Y, editor. Ensemble Machine Learning. New York, NY: Springer; 2012. p. 1–34. [Google Scholar]
  • 20.Hu H, Xu W, Jiang T, Cheng Y, Tao X, Liu W, et al. Expert-Level Immunofixation Electrophoresis Image Recognition based on Explainable and Generalizable Deep Learning. Clin Chem. 2023;69:130–9. doi: 10.1093/clinchem/hvac190 [DOI] [PubMed] [Google Scholar]
  • 21.Rafae A, Malik MN, Abu Zar M, Durer S, Durer C. An Overview of Light Chain Multiple Myeloma: Clinical Characteristics and Rarities, Management Strategies, and Disease Monitoring. Cureus. 2018;10:e3148. doi: 10.7759/cureus.3148 [DOI] [PMC free article] [PubMed] [Google Scholar]

Decision Letter 0

John Adeoye

10 Aug 2023

PONE-D-23-17568Machine learning evaluation for identification of M-proteins in human serumPLOS ONE

Dear Dr. Rotter Sopasakis,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please submit your revised manuscript by Sep 24 2023 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

John Adeoye

Academic Editor

PLOS ONE

Journal requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2. Please provide additional details regarding participant consent. In the ethics statement in the Methods and online submission information, please ensure that you have specified (1) whether consent was informed and (2) what type you obtained (for instance, written or verbal, and if verbal, how it was documented and witnessed). If your study included minors, state whether you obtained consent from parents or guardians. If the need for consent was waived by the ethics committee, please include this information.

If you are reporting a retrospective study of medical records or archived samples, please ensure that you have discussed whether all data were fully anonymized before you accessed them and/or whether the IRB or ethics committee waived the requirement for informed consent. If patients provided informed written consent to have data from their medical records used in research, please include this information

3. Please note that PLOS ONE has specific guidelines on code sharing for submissions in which author-generated code underpins the findings in the manuscript. In these cases, all author-generated code must be made available without restrictions upon publication of the work. Please review our guidelines at https://journals.plos.org/plosone/s/materials-and-software-sharing#loc-sharing-code and ensure that your code is shared in a way that follows best practice and facilitates reproducibility and reuse.

4. We note that the grant information you provided in the ‘Funding Information’ and ‘Financial Disclosure’ sections do not match.

When you resubmit, please ensure that you provide the correct grant numbers for the awards you received for your study in the ‘Funding Information’ section.

5. Thank you for stating the following financial disclosure:

"This work is supported by the department of Clinical Chemistry, Sahlgrenska University Hospital, Gothenburg, Sweden (VRS, LMH, MN, FN, MA). There is no other specific funding for this work."         

Please state what role the funders took in the study.  If the funders had no role, please state: "The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript."

If this statement is not correct you must amend it as needed.

Please include this amended Role of Funder statement in your cover letter; we will change the online submission form on your behalf.

6. We note that you have indicated that data from this study are available upon request. PLOS only allows data to be available upon request if there are legal or ethical restrictions on sharing data publicly. For more information on unacceptable data access restrictions, please see http://journals.plos.org/plosone/s/data-availability#loc-unacceptable-data-access-restrictions.

In your revised cover letter, please address the following prompts:

    a) If there are ethical or legal restrictions on sharing a de-identified data set, please explain them in detail (e.g., data contain potentially sensitive information, data are owned by a third-party organization, etc.) and who has imposed them (e.g., an ethics committee). Please also provide contact information for a data access committee, ethics committee, or other institutional body to which data requests may be sent.

   b) If there are no restrictions, please upload the minimal anonymized data set necessary to replicate your study findings as either Supporting Information files or to a stable, public repository and provide us with the relevant URLs, DOIs, or accession numbers. For a list of acceptable repositories, please see http://journals.plos.org/plosone/s/data-availability#loc-recommended-repositories.

We will update your Data Availability statement on your behalf to reflect the information you provide.

7. We note that you have stated that you will provide repository information for your data at acceptance. Should your manuscript be accepted for publication, we will hold it until you provide the relevant accession numbers or DOIs necessary to access your data. If you wish to make changes to your Data Availability statement, please describe these changes in your cover letter and we will update your Data Availability statement to reflect the information you provide.

8. Please include captions for your Supporting Information files at the end of your manuscript, and update any in-text citations to match accordingly. Please see our Supporting Information guidelines for more information: http://journals.plos.org/plosone/s/supporting-information

Additional Editor Comments:

The reviewers both agree that the dataset used for ML model construction is valuable but both cited the poor methodological approach of the study. In view of this, the authors are advised to revise their ML modelling methodology extensively in line with contemporary standards to allow for further consideration of their manuscript in the Journal.

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Partly

**********

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #2: No

**********

3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: Yes

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

**********

5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: Please see the attached file for my comments, in addition with the summary of my review below:

The authors performed capillary & gel electrophoresis and immunofixation for more to 60k serum samples. Based on their data, they classified their samples into either “absence of M-protein” or “presence of M-protein”, and classified the subtype of M-protein, when relevant. Next, they trained various ML classifiers to automate those classification tasks, by feeding those with raw numerical data of the SPEP (Serum Protein Electrophoresis) curves, i.e. the value of each point of the SPEP curve.

The authors present an extremely valuable dataset: more than 60k samples, annotated thanks to three biological assays: capillary SPEP, gel SPEP and immunofixation; with the interpretation from laboratory experts. This alone is a solid argument supporting the publication of their results. It should be noted that the data might be described better, for instance by 1) indicating the size of M-spikes in their curves and the proportion of small/medium/large M-spikes, for instance; and 2) the proportion of other abnormal patterns which may mask or be confounded with M-proteins (artifacts (fibrinogen, iodinated contrast agents…), beta-gamma bridging). Furthermore, it is not clear if all samples were analyzed with all 3 assays: capillary SPEP, gel SPEP and immunofixation. A table may be useful here. Finally, the choice of the authors to partition their data into a training and test sets only, without any validation set, despite their number of samples, is at major risk of data leakage: there is an urge in addressing or at least discussing this point in the manuscript.

The main downside of this study, in my opinion, is the machine learning methodology used here, which is far from state of the art standards. SPEP data, is, by definition, signal data. Multiple works have highlighted the superiority of DL (deep learning) methods, mostly but not limited to CNN & transformers, to such data. However, the word “signal” is never used throughout the manuscript and DL models are absent from the models trained by the authors. The question of why the authors chose to elude this major point stays unanswered at the end of this manuscript. The ML (meachine learning) knowledge of the authors does not seem to be an obstacle to the use of DL models, based on the expertise they demonstrate in ML and DL, largely discussing the upsides and downsides of various ML models including DL ones. Furthermore, training/inference speed should not be an issue either, since 60k samples times 300 points per curve is a relatively small amount of data to process (the authors state that 20 minutes were needed to train all 26 models in this study).

The use of SHAP values, here, further highlights this problem. The models seem to simply “look at high values” in the beta & gamma regions and deduce if there is an M-spike in the sample. However, M-spikes are not always characterized by high values, and small M-spikes may only be visually detected, by the expert, thanks to an abnormal qualitative pattern, rather than by observing high quantities in the beta/gamma fractions (e.g., shoulder in the beta fraction). This is only possible with ML if treating the data as a signal, e.g., by using convolutional layers. Since the authors do not inform about the exact patterns observed in their dataset (see my previous paragraph about data), it is impossible to predict their models’ expected behavior on such samples. Unfortunately, those samples are crucial, and highly responsible for the fact that SPEP is not yet fully automated in modern laboratories.

Finally, the authors imply that their methodology may have other applications. This is highly doubtful, as they are few examples of biological assays outputting highly standardized/aligned signal data, which may give such robust results with the methodology they use. Indeed, the previous studies the authors cite to support their work have been using this kind of ML tools to make predictions based on concentrations deducted from signal data, rather than the raw data themselves.

In total, this study seems highly promising, thanks notably to a highly valuable dataset, but which should be analyzed using state of the art machine learning tools to be considered in today’s literature.

Reviewer #2: This is a very interesting paper with a clear clinical question ( The major clinical question is the presence of monoclonal fraction(s) of antibodies (M-protein/paraprotein), which is essential for the diagnosis and follow-up of hematological diseases, such as multiple myeloma) . They also have a very nice dataset.

Regarding the adopted evaluation procedures, there is an important step that is missing: how did the authors choose the hyperparameters? Usually , one performs a gridsearch using k-fold in the training set as GridSeachCV form sk-learn does. If you use the test set to determine the best set of hyperparameters, the results woulb be biased.

it would be nice to compare the results provided by SHAP with standard feature importance that several tree-based algorithms can provide.

In the discussion , the authors say “Deep learning algorithms are highly suitable and powerful for image analysis but exhibit some disadvantages. One such weakness is the “black box” phenomenon, the inability to explain how the outcome result was achieved”. One can say the same thing about XGB or LGBM. For deep learning , there are several explanation methods that could be employed, including SHAP. Please correct this statement.

In the discussion the authors also say “The use of decision tree methods can be more transparent compared to deep learning algorithms, if implemented as we propose here, but work best for numerical series of data rather than data retrieved from images”. Decision trees are more transparent, but this is not true for the other algorithms derived from them as XGB. It is also very hard to tune XGB or LGBM hyperparameters, because there are several of them(see the documentation)

The authors say that the limitation of their approach was the relatively low accuracy

scores for identifying the specific isotype of M-protein present, probably due to the low number of certain M-protein isotypes in the population. What the authors had was an imbalanced dataset. This could be correct using SMOTE techniques.

**********

6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: Yes: Floris CHABRUN

Reviewer #2: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

Attachment

Submitted filename: Reviewing PONE.pdf

pone.0299600.s006.pdf (404KB, pdf)
PLoS One. 2024 Apr 2;19(4):e0299600. doi: 10.1371/journal.pone.0299600.r002

Author response to Decision Letter 0


16 Nov 2023

We thank the Reviewers for their constructive and insightful comments and suggestions, which we believe have considerably improved our manuscript. We have fully addressed the reviewers’ comments and concerns in a point-by-point response, submitted as requested as a "Response to Reviewers" file. All changes made to the manuscript are shown with track changes in the “Revised Article with Changes Highlighted” file. We hope that the reviewing process finds our revised manuscript acceptable for publication in PLOS ONE.

Attachment

Submitted filename: Response to Reviewers.pdf

pone.0299600.s007.pdf (572.5KB, pdf)

Decision Letter 1

John Adeoye

9 Feb 2024

PONE-D-23-17568R1Machine learning evaluation for identification of M-proteins in human serumPLOS ONE

Dear Dr. Rotter Sopasakis,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please submit your revised manuscript by Mar 25 2024 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

John Adeoye

Academic Editor

PLOS ONE

Journal Requirements:

Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice.

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #1: (No Response)

Reviewer #3: All comments have been addressed

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #3: Yes

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #3: Yes

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #3: Yes

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #3: Yes

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: First, I want to thank the authors for carefully studying and replying to each and every one of my comments, and for the thorough work I believed they have undertaken for improving their manuscript. In my opinion, all major concerns regarding this study have been addressed.

Particularly, the current version of their manuscript now enables a detailed understanding of the exact data used by the authors and potentially accessible to the community, notably due to the addition of Tables 1 and 2 and confirmation that all samples were double checked with both CZE and gel SPEP. This confirms the high value of those data.

I still have a few minor comments to the authors, which may or may not trigger modifications in the final manuscript, to their discretion:

- New tables 1 and 2 are interesting and I feel are nice additions to the manuscript. In Table 2, I am not sure about the use of the word “tertiles”, which implies that selected cut-offs, namely 1g/L and 5g/L, would divide the dataset into three subsets of equal size (1 third of the samples), which is not the case here.

- Several arguments against the use of deep learning in this study seem fallacious in my opinion. For instance:

o the assumed ability of tree models to better handle class imbalance due to Gini index/entropy overlooks the fact that cross entropy is one of the most widely used loss functions in deep learning; for the very same reason.

o the authors imply that convolutional layers would not be suitable to the analysis of 306-point-wide traces reshaped to 17x18 images. Analyzing small images is perfectly performed by CNN, for instance on countless MNIST examples online (28x28 images). Furthermore, 306-wide traces can (and should here) be reshaped to (306,1) arrays/tensors to avoid jeopardizing spatial information, as described in several previous works analyzing 1-dimensional signal. Finally, authors could even use 1d-conv layers to avoid reshaping raw traces if reshaping itself is a concern.

o I can understand that the authors refuse to use DL compared to tree models, and I think the “no free lunch theorem” sufficiently justifies this strategy. But in my opinion, the specific arguments cited above are misleading and should be removed from the final version.

- Authors corrected the name of the python package they used from “shapely” to “Shapely” version 1.7.1. I would like to make sure this is not a mistake: Figure 3 highly resembles plots obtained with the “SHAP” package based on “ShapLEy” (not “ShapELy”) values. Furthermore, SHAP is also cited multiple times by the author. ShapELy v 1.7.1 seems to be a Python package related to geometric analysis rather than feature importance. Can the authors confirm there is no typo here?

- Knowing that the authors, even unknowingly, complied with recommendations such as STARD is extremely positive. I think the readership’s trust would highly benefit from citing this in the Methods section, though this is not mandatory.

- The table cited in the “Response to reviewers” file, depicting the performance of models according to the M-spike concentration (<1g/L, 1-5g/L, >5g/L) is really interesting, and in my opinion its addition in the manuscript or supplementary material, or at least a sentence in the Results would be of interest for the readers of this work.

Reviewer #3: This is an interesting study that demonstrates the analysis of a large set of electrophoresis data using machine learning for the identification of M-proteins in serum. Two referees have thoroughly assessed the paper and provided useful feedback which has been carefully addressed by the authors. While I agree that the ML algorithms used are not overly sophisticated, their applicability to the problem makes the procedure easy to implement. It would be good if this could be tested on 'new' electrophoresis data from various sources to test the transferability of the method, but I am aware this is beyond the scope of the current study. The dataset is extensive and relevant and I think the paper will make a good addition to the literature.

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: Yes: Floris Chabrun

Reviewer #3: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

PLoS One. 2024 Apr 2;19(4):e0299600. doi: 10.1371/journal.pone.0299600.r004

Author response to Decision Letter 1


10 Feb 2024

Reviewer #1: First, I want to thank the authors for carefully studying and replying to each and every one of my comments, and for the thorough work I believed they have undertaken for improving their manuscript. In my opinion, all major concerns regarding this study have been addressed.

Particularly, the current version of their manuscript now enables a detailed understanding of the exact data used by the authors and potentially accessible to the community, notably due to the addition of Tables 1 and 2 and confirmation that all samples were double checked with both CZE and gel SPEP. This confirms the high value of those data.

We thank the reviewer for carefully and thoroughly assessing our manuscript and our changes and responses in the first round of revision. We have addressed the new comments as seen below in a point-by-point response. We hope that our responses have resolved any remaining uncertainties that the reviewer had.

I still have a few minor comments to the authors, which may or may not trigger modifications in the final manuscript, to their discretion:

- New tables 1 and 2 are interesting and I feel are nice additions to the manuscript. In Table 2, I am not sure about the use of the word “tertiles”, which implies that selected cut-offs, namely 1g/L and 5g/L, would divide the dataset into three subsets of equal size (1 third of the samples), which is not the case here.

We have removed the word “tertile” from Table 2 as requested and instead added “three groups: <1 g/l, 1-5 g/l and >5 g/l”.

- Several arguments against the use of deep learning in this study seem fallacious in my opinion. For instance:

o the assumed ability of tree models to better handle class imbalance due to Gini index/entropy overlooks the fact that cross entropy is one of the most widely used loss functions in deep learning; for the very same reason.

o the authors imply that convolutional layers would not be suitable to the analysis of 306-point-wide traces reshaped to 17x18 images. Analyzing small images is perfectly performed by CNN, for instance on countless MNIST examples online (28x28 images). Furthermore, 306-wide traces can (and should here) be reshaped to (306,1) arrays/tensors to avoid jeopardizing spatial information, as described in several previous works analyzing 1-dimensional signal. Finally, authors could even use 1d-conv layers to avoid reshaping raw traces if reshaping itself is a concern.

o I can understand that the authors refuse to use DL compared to tree models, and I think the “no free lunch theorem” sufficiently justifies this strategy. But in my opinion, the specific arguments cited above are misleading and should be removed from the final version.

At the reviewer’s request, we have now removed the paragraph with our arguments against deep learning methods in our setting from the discussion.

- Authors corrected the name of the python package they used from “shapely” to “Shapely” version 1.7.1. I would like to make sure this is not a mistake: Figure 3 highly resembles plots obtained with the “SHAP” package based on “ShapLEy” (not “ShapELy”) values. Furthermore, SHAP is also cited multiple times by the author. ShapELy v 1.7.1 seems to be a Python package related to geometric analysis rather than feature importance. Can the authors confirm there is no typo here?

We thank the reviewer for correction. Indeed, we are using the SHAP (SHapley Additive exPLanations) python library package to perform feature importance analysis and not the Shapely 1.7.1 library. This typo is now corrected in the manuscript. The following incorrect text "Feature importance was based on the Shapely values from the Shapely 1.7.1 library" has been replaced by "Feature importance was based on the Shapley values from the python package shap version 0.37.0”.

- Knowing that the authors, even unknowingly, complied with recommendations such as STARD is extremely positive. I think the readership’s trust would highly benefit from citing this in the Methods section, though this is not mandatory.

We thank the reviewer for the suggestion. A citation has now been added to the Materials and Methods section that states that our study aligns with the STARD recommendations for reporting diagnostic studies.

- The table cited in the “Response to reviewers” file, depicting the performance of models according to the M-spike concentration (<1g/L, 1-5g/L, >5g/L) is really interesting, and in my opinion its addition in the manuscript or supplementary material, or at least a sentence in the Results would be of interest for the readers of this work.

We have added a sentence in the Results section as requested.

Reviewer #3: This is an interesting study that demonstrates the analysis of a large set of electrophoresis data using machine learning for the identification of M-proteins in serum. Two referees have thoroughly assessed the paper and provided useful feedback which has been carefully addressed by the authors. While I agree that the ML algorithms used are not overly sophisticated, their applicability to the problem makes the procedure easy to implement. It would be good if this could be tested on 'new' electrophoresis data from various sources to test the transferability of the method, but I am aware this is beyond the scope of the current study. The dataset is extensive and relevant and I think the paper will make a good addition to the literature.

We thank the reviewer for carefully and thoroughly assessing our manuscript and our changes and responses in the first round of revision. We are very grateful that Reviewer #3 took on the task as a new reviewer and appreciate the work the reviewer put in as well as the comments. We hope to access more electrophoresis data in the future from other sources to further test our method as suggested, but as the reviewer already pointed out, this is beyond the scope of the current study.

Attachment

Submitted filename: Response to reviewers.pdf

pone.0299600.s008.pdf (129.3KB, pdf)

Decision Letter 2

John Adeoye

14 Feb 2024

Machine learning evaluation for identification of M-proteins in human serum

PONE-D-23-17568R2

Dear Dr. Rotter Sopasakis,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice for payment will follow shortly after the formal acceptance. To ensure an efficient process, please log into Editorial Manager at http://www.editorialmanager.com/pone/, click the 'Update My Information' link at the top of the page, and double check that your user information is up-to-date. If you have any billing related questions, please contact our Author Billing department directly at authorbilling@plos.org.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

John Adeoye

Academic Editor

PLOS ONE

Additional Editor Comments (optional):

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #1: All comments have been addressed

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: The authors have carefully updated their manuscript and addressed all interrogations I had in the previous versions, and in my opinion is suitable for publication in its current form.

I would like to thank the authors for their time and for carefully rewriting some parts of their manuscript that may have lacked clarity in the past.

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: Yes: Floris Chabrun

**********

Acceptance letter

John Adeoye

25 Mar 2024

PONE-D-23-17568R2

PLOS ONE

Dear Dr. Rotter Sopasakis,

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now being handed over to our production team.

At this stage, our production department will prepare your paper for publication. This includes ensuring the following:

* All references, tables, and figures are properly cited

* All relevant supporting information is included in the manuscript submission,

* There are no issues that prevent the paper from being properly typeset

If revisions are needed, the production department will contact you directly to resolve them. If no revisions are needed, you will receive an email when the publication date has been set. At this time, we do not offer pre-publication proofs to authors during production of the accepted work. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few weeks to review your paper and let you know the next and final steps.

Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

If we can help with anything else, please email us at customercare@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Dr. John Adeoye

Academic Editor

PLOS ONE

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 Fig. Capillary zone electrophoresis (CZE) outlining the different serum protein fractions in a normal sample and in a sample containing M-protein.

    (TIF)

    pone.0299600.s001.tif (67.6MB, tif)
    S2 Fig. Concentration range of free light chains among the samples in the data set.

    (TIF)

    pone.0299600.s002.tif (99.6MB, tif)
    S3 Fig. SHAP calculation results from nine samples containing M-protein and three samples with no M-protein.

    (TIF)

    pone.0299600.s003.tif (67.6MB, tif)
    S4 Fig. Classic feature importance approach based on the trained Extra Trees Classifier.

    (TIF)

    pone.0299600.s004.tif (99.6MB, tif)
    S1 Table. Accuracy scores, ROC AUC values and F1 scores for the remaining methods tested.

    (TIF)

    pone.0299600.s005.tif (97.3MB, tif)
    Attachment

    Submitted filename: Reviewing PONE.pdf

    pone.0299600.s006.pdf (404KB, pdf)
    Attachment

    Submitted filename: Response to Reviewers.pdf

    pone.0299600.s007.pdf (572.5KB, pdf)
    Attachment

    Submitted filename: Response to reviewers.pdf

    pone.0299600.s008.pdf (129.3KB, pdf)

    Data Availability Statement

    No - some restrictions will apply; The data have been uploaded in a public data repository, the Swedish National Data Service (https://snd.gu.se/en), with metadata and documentation made available according to the FAIR principles, and it has been assigned a permanent identifier (DOI): https://doi.org/10.5878/a2aa-kt50. To access the actual data files, a request management system is available to file a formal request to the University of Gothenburg for the data at https://snd.gu.se/en or snd@snd.gu.se. The code can be accessed through github at https://github.com/a0s6044/ProteinElectrophoresis.


    Articles from PLOS ONE are provided here courtesy of PLOS

    RESOURCES