Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 Oct 30.
Published in final edited form as: Cancer Epidemiol Biomarkers Prev. 2025 Dec 1;34(12):2259–2266. doi: 10.1158/1055-9965.EPI-25-1032

SMAGS-LASSO: A Novel Feature Selection Method for Sensitivity Maximization in Early Cancer Detection

Hamid Khoshfekr Rudsari 1, Sara Khorami-Sarvestani 2, Johannes F Fahrmann 2, James P Long 1, Samir Hanash 2, Kim-Anh Do 1, Ehsan Irajizad 1,*
PMCID: PMC12570486  NIHMSID: NIHMS2115402  PMID: 40996315

Abstract

Background:

Sensitivity and specificity are foundational metrics for cancer detection tools. However, most machine learning algorithms prioritize overall accuracy during optimization, which fails to align with clinical priorities of early detection. We aim to develop a feature selection machine learning algorithm while maximizing sensitivity at a given specificity.

Methods:

We developed SMAGS-LASSO, a machine learning algorithm that combines our developed Sensitivity Maximization at a Given Specificity (SMAGS) framework with L1 regularization for feature selection. This approach simultaneously optimizes sensitivity at user-defined specificity thresholds while performing feature selection. SMAGS-LASSO utilizes a custom loss function with L1 regularization and multiple parallel optimization techniques. We used train-test splits and cross-validation, comparing against LASSO and Random Forest using sensitivity and AUC metrics. We evaluated our method on synthetic datasets and real-world protein colorectal cancer biomarker data.

Results:

In synthetic datasets designed to contain strong signals for both sensitivity and specificity, SMAGS-LASSO significantly outperformed standard LASSO, achieving sensitivity of 1.00 (95% CI: 0.98–1.00) compared to 0.19 (95% CI: 0.13–0.23) for LASSO at 99.9% specificity. In colorectal cancer data, SMAGS-LASSO demonstrated 21.8% improvement over LASSO (p-value = 2.24E-04) and 38.5% over Random Forest (p-value = 4.62E-08) at 98.5% specificity while selecting the same number of biomarkers.

Conclusions:

SMAGS-LASSO enables development of minimal biomarker panels that maintain high sensitivity at predefined specificity thresholds, offering superior performance for early cancer detection.

Impact:

This method provides a promising approach for early cancer detection and other medical diagnostics requiring sensitivity-specificity optimization.

1-. Introduction

In clinical diagnostics, particularly for diseases with low prevalence such as cancer, the effective classification of patients into healthy control and disease groups represents a critical challenge. While numerous metrics have been developed to evaluate classification performance, including accuracy, sensitivity, specificity, and area under the receiver operating characteristic (ROC) curve (AUC), each addresses specific aspects of model performance. While sensitivity, specificity, and AUC maintain their statistical integrity across different class distributions, accuracy can be misleading in imbalanced datasets—a common challenge in early disease detection scenarios. Sensitivity (true-positive rate) and specificity (true-negative rate) stand as particularly important metrics in early cancer detection as has been observed in prior studies [1], [2], [3]. Sensitivity measures a model’s ability to correctly identify positive cases, while specificity reflects its capacity to correctly classify negative cases. In early cancer detection and risk assessment application, these metrics take on heightened significance: high sensitivity is essential to minimize missed cancer diagnoses, while high specificity helps avoid unnecessary clinical procedures in healthy individuals that can lead to physical, psychological, and financial burdens [4]. Traditional classification methods, such as logistic regression with maximum likelihood estimation, are designed to optimize overall accuracy and do not explicitly prioritize sensitivity—an essential objective in early cancer detection. To address this challenge, we previously developed SMAGS (Sensitivity Maximization at A Given Specificity) [1], which modified the standard logistic regression loss function to optimize sensitivity at a predetermined specificity threshold. While SMAGS demonstrated promising results with smaller feature sets, its application to high-dimensional biomarker data such as multiple protein or peptide biomarkers, remains to be explored.

Feature selection becomes crucial not only for computational efficiency but also for developing interpretable models that can identify the most informative biomarkers for early disease detection. Zhang [5] proposed a feature selection method based on global sensitivity analysis to determine the most relevant feature subsets and improve prediction performance of machine learning models. Efimov and Sulieman [6] introduced a Sobol sensitivity-based approach for feature selection that efficiently evaluates feature importances without retraining the model. Similarly, Wang et al. [7] proposed a method that maximizes independent classification information, addressing the critical balance between feature relevance to the target class and feature redundancy. Garcia-Nieto et al. [8] specifically addressed cancer diagnosis by introducing a multi-objective genetic algorithm for gene selection of microarray datasets, which performs feature selection from the perspective of both sensitivity and specificity. Their approach demonstrated the value of considering these metrics simultaneously in the feature selection process. For scenarios with limited labeled data, Xu et al. [9] proposed a discriminative semi-supervised feature selection method based on manifold regularization, which maximizes classification margins while exploiting the underlying data geometry. In the context of regularization methods, Li et al. [10] investigated the adversarial robustness of LASSO-based feature selection, highlighting the importance of stability in feature selection approaches.

Existing approaches have several limitations if we want to maximize sensitivity for applications such as early cancer detection. First, many feature selection methods focus on maximizing overall accuracy or AUC rather than optimizing sensitivity at clinically relevant specificity thresholds, which is crucial for cancer screening. Second, methods that do consider partial AUC or weighted sensitivity/specificity often treat feature selection as a separate preprocessing step rather than integrating it into the optimization process. Third, approaches that use regularization for feature selection typically employ standard loss functions that don’t directly address the sensitivity-specificity tradeoff. Finally, many existing methods lack robust cross-validation procedures specifically designed to maintain the desired specificity constraints while selecting optimal regularization parameters.

In this paper, we introduce SMAGS-LASSO, a novel machine learning method that integrates the SMAGS framework with L1 regularization (LASSO) [11] to simultaneously maximize sensitivity at a given specificity while performing effective feature selection. The method employs a custom loss function that combines sensitivity optimization with L1 regularization, dynamically adjusting the classification threshold based on a specified specificity percentile. SMAGS-LASSO utilizes multiple optimization techniques processed in parallel to explore the parameter space comprehensively, ensuring robust convergence. To determine the optimal regularization parameter (λ), we implement a cross-validation procedure that selects the value minimizing classification error while maintaining the desired specificity threshold.

In the paper, we explain the objective function, optimization and cross-validation algorithms in Section 2. We then present the results of SMAGS-LASSO for synthetic data and real datasets of protein biomarkers and compare them with SMAGS and LASSO methods in Section 3. In Section 4, we discuss the advantages and limitations of our approach, exploring its clinical implications for early cancer detection and identifying directions for future research. Finally, we conclude with a summary of our contributions and their potential impact on biomarker discovery for cancer diagnostics.

2-. Methods

Here, we propose a novel approach called SMAGS- LASSO, which extends traditional LASSO regression to maximize sensitivity while maintaining sparsity in the coefficient vector. This approach is particularly beneficial in biomedical and clinical contexts where correctly identifying positive cases (high sensitivity) at a given specificity is critical.

2-A. Problem Formulation

Consider a binary classification problem with a feature matrix X Rn×p and binary outcome vector y {0, 1}n, where n is the number of observations and p is the number of features. The SMAGS-LASSO method aims to find a sparse coefficient vector β Rp and intercept β0 R that maximize sensitivity while controlling for specificity through a pre- specified parameter.

2-A-1. Objective Function

Our objective function differs from traditional LASSO by directly optimizing sensitivity rather than likelihood or mean squared error. Therefore, we have the following:

maxβ,β0i=1nyi^.yii=1nyi-λβ1,Subject to1-yT(1-y^)1-yT(1-y)SP, (1)

where the first part of (1) is the proportion of true positive predictions among all positive cases, λ is the regularization parameter that controls the level of sparsity and β1 is the L1-norm of the coefficient vector, encouraging sparsity. In (1), SP is the given specificity and y^i is the predicted class for observation i, determined by

yi^=I(σ(xiTβ+β0)>θ), (2)

where σz=11+e-z is the sigmoid function, and θ is a threshold parameter determined adaptively to control the specificity level. We then select features as non-zero if the absolute value of each individual coefficient exceeds 5% of the largest coefficient’s absolute value in our method.

2-B. Optimization Procedure

The SMAGS-LASSO optimization is challenging due to the non-differentiable nature of both the sensitivity and the L1 penalty. We employ a multi-pronged optimization strategy using several algorithms:

  1. Initialize coefficients using a standard logistic regression model,

  2. Apply multiple optimization algorithms (Nelder-Mead[12], BFGS [13], CG [13], L-BFGS-B [14]) with varying tolerance levels,

  3. Select the model with the highest sensitivity among the converged solutions.

This approach leverages parallel processing to efficiently explore multiple optimization paths.

2-C. Cross Validation Framework

To select the optimal regularization parameter λ, we implemented a cross-validation procedure specifically designed for our sensitivity-maximizing objective. The procedure:

  1. Creates k-fold partitions of the data (k = 5 by default),

  2. Evaluates a sequence of λ values on each fold,

  3. Measures performance using sensitivity mean squared error (MSE) metric:
    MSEsensitivity=1-i=1nyi^yii=1nyi2, (3)
  4. Tracks the norm ratio βλ1β1 to quantify sparsity.

The norm ratio provides an interpretable measure of the model’s sparsity, where βλ1 is the L1-norm of the coefficient vector at regularization parameter λ, and βλ1 is the L1-norm of the coefficient vector from a full (unregularized) model.

The cross-validation process selects the λ value that minimizes the sensitivity MSE, effectively finding the most regularized model that maintains high sensitivity.

2-D. Evaluation Framework

We employed a comprehensive evaluation strategy to assess SMAGS-LASSO performance against established methods including standard LASSO, unregularized SMAGS, and Random Forest. All experiments used 80/20 stratified train-test splits to maintain balanced class representation and ensure robust performance assessment. An open-source implementation is available at github.com/khoshfekr1994/SMAGS.LASSO.

2-D-1. Synthetic and Real-World Datasets

Our evaluation strategy follows a two-pronged approach. First, we engineered a synthetic dataset with parameters specifically designed to demonstrate the capabilities of our method when compared with SMAGS and LASSO approaches. To rigorously evaluate our SMAGS-LASSO method against SMAGS and standard LASSO approaches, we engineered three synthetic datasets with distinct signal patterns. Each dataset comprised 2,000 samples (1,000 per class) with 100 features. For all experiments, we employed an 80/20 train-test split and set a high specificity target (SP = 99.9%) to simulate scenarios where false positives must be minimized. Our baseline dataset, Synthetic Dataset 1 (referred as No Signal), consisted of features drawn from a normal distribution (mean=50, SD=5) with no differential expression pattern between classes. The heatmap of this synthetic data is shown in Supplementary Figure 1(a). This homogeneous dataset served as a control to assess feature selection performance in the absence of true signals. In Synthetic Dataset 2 (referred as Unidirectional Signal), we introduced localized signal patterns within the first 50 features. Each signal feature contained a small region of 20 samples with elevated expression (mean=100, SD=2) against a low background (mean=2, SD=2). The heatmap of this synthetic data is shown in Supplementary Figure 1(b). These high-expression regions were positioned sequentially within cases, creating a pattern of sparse but strong unidirectional signals. In Synthetic Dataset 3 (referred as Bidirectional Signal), we designed a more complex dataset with bidirectional signals to challenge the feature selection capabilities of the models. The first 50 features exhibited high positive expression (mean=100, SD=2) in the case class (for sensitivity improvement), while features 51–100 displayed negative expression (mean=−100, SD=2) in the control class (for specificity improvement). The heatmap of this synthetic data is shown in Supplementary Figure 1(c). These opposing signals create a complex pattern where standard methods may struggle to maintain sensitivity when constrained by high specificity requirements.

Subsequently, to evaluate SMAGS-LASSO in a real clinical context, we applied our method to a real-world colorectal cancer dataset [15]. Colorectal cancer remains one of the most common and lethal malignancies worldwide, with an estimated 2.2 million new cases and 1.1 million deaths anticipated by 2030, underscoring the urgent need for more effective diagnostic and prognostic tools [16]. All colorectal cases are relatively early stages (before they metastasized) when surgical intervention could potentially be curative. The case population has a median age of 64.16 years (range: 22–93 years) with balanced gender representation (54.6% male, 45.4% female). Healthy controls had a median age of 49.30 years (range: 17–87.6 years) with similar gender distribution to cases (53.5% male, 46.5% female) and no known history of cancer or chronic diseases. The dataset contains measurements of 39 serum-based protein biomarkers from both cancer cases and controls, with highly imbalanced class distributions as detailed in Table 1. We included the full list of 39 protein biomarkers in Supplementary Table 1. The target specificity was set to 98.5% for colorectal cancer [17] reflecting the clinical requirement for minimizing false positives in cancer screening applications.

Table 1.

Summary Statistics of the Utilized CancerSeek Dataset

Cancer Type Total Samples (%males, %females) Cases (%males, %females) Controls (%males, %females) Age Distribution (Mean, Range)
Colorectum 1200 (53.83% males, 46.17% females) 388 (54.64% males, 45.36% females) 812 (53.45% males, 46.55% females) Cases: 64.16 (22.00–93.00) Controls: 49.30 (17.00–87.62)
*

Note: The dataset contains measurements for 39 biomarkers.

2-D-2. Performance Metrics

We evaluated methods using sensitivity at predetermined specificity thresholds, AUC, and number of selected features. Optimal regularization parameters (λ) were determined using k-fold cross-validation (k=5) by minimizing sensitivity mean squared error. The cross-validation procedure selected the λ value that achieved the highest sensitivity while maintaining the desired specificity constraint.

2-D-3. Statistical Analysis

We calculated 95% confidence intervals for all performance metrics using bootstrap resampling (R=1,000 iterations). Statistical significance of performance differences between methods was assessed using paired bootstrap tests. P-values were calculated by comparing the distribution of performance differences across bootstrap samples, with significance set at α = 0.05.

For each evaluation, we ensured fair comparison by selecting the same number of features across methods when possible. For Random Forest feature selection, we employed Mean Decrease in Gini impurity scores followed by logistic regression on selected features.

2-E. Data Availability Statement

The synthetic data generated in this study as well as the real-world protein biomarker data are publicly available in https://github.com/khoshfekr1994/SMAGS.LASSO.

3-. Results

In this section, we present the findings from our novel feature selection approach utilizing SMAGS-LASSO and benchmark its performance against previously established methods: SMAGS and LASSO.

3-A. Synthetic Data

3-A-1. Synthetic Dataset 1 (No Signal)

This dataset consisted of features drawn from a normal distribution (mean=50, SD=5) with no differential expression pattern between classes. In this scenario, all three methods, LASSO (λ = 0.005), SMAGS, and SMAGS-LASSO (λ = 0.06), showed comparable and relatively poor performance at the target specificity (Figures (1a) and (1b)). The comparable number of selected features (LASSO: 64, SMAGS-LASSO: 65) allows for fair comparison between the methods, confirming that none can extract meaningful patterns from noise when constrained to high specificity.

Figure 1.

Figure 1.

Performance comparison on synthetic datasets. ROC curves for (a,b) Dataset 1 (no signal), (c,d) Dataset 2 (unidirectional signal), and (e,f) Dataset 3 (bidirectional signal), shown for both training and test sets. The vertical dashed line indicates the target specificity of 99.9%. Numbers in parentheses indicate the number of features selected by each method.

3-A-2. Synthetic Dataset 2 (Unidirectional Signal)

Here, with an introduced distinct signal patterns (signal feature contained a small region of 20 samples with elevated expression (mean=100, SD=2) against a low background (mean=2, SD=2)), all three methods—LASSO (λ = 0.01), SMAGS, and SMAGS-LASSO (λ = 1.35)—achieved perfect sensitivity (1.0) at the target specificity of 99.9%, as shown in Figures (1c) and (1d). This dataset verified that SMAGS-LASSO functions can correctly identify all cases under favorable conditions where signals are strong and consistent. Notably, both LASSO and SMAGS-LASSO identified a similar number of features (50), which corresponds to the true number of signal features, while SMAGS selected all 100 features, indicating less parsimony in feature selection.

3-A-3. Synthetic Dataset 3 (Bidirectional Signal)

Here, we imposed the first 50 features exhibited high positive expression (mean=100, SD=2) in the case class (for sensitivity improvement), while features 51–100 displayed negative expression (mean=−100, SD=2) in the control class (for specificity improvement). As shown in Figures (1e) and (1f), SMAGS-LASSO (λ = 1.05) demonstrated significantly higher sensitivity at the target specificity (99.9%) compared to LASSO (λ = 0.0495). While both methods identified the same number of features (50), SMAGS-LASSO consistently maintained perfect sensitivity (1.0) in both training and test sets, effectively capturing both positive and negative signal patterns. In contrast, LASSO achieved moderate sensitivity (0.31 in training, 0.19 in test), suggesting difficulty in handling complext signals under high specificity constraints. The unregularized SMAGS method also achieved perfect sensitivity, but no feature selection has been done.

This dataset illustrates the key advantage of our SMAGS-LASSO approach: it maintains both high specificity and sensitivity in complex signal scenarios while achieving feature parsimony comparable to LASSO. Such capabilities are particularly valuable in applications like clinical biomarker discovery, where identifying the minimal set of truly discriminative features is critical.

3-B. Model Performance of Colorectal Cancer Biomarker Dataset

Applying the method to real dataset (colorectal cancer) demonstrates superior performance of SMAGS-LASSO compared to LASSO in terms of maximizing sensitivity at the target specificity level in both training (0.813, 95% CI: 0.720–0.864) and testing (0.795, 95% CI: 0.459–0.877) sets at SP = 0.985. Figure 2 shows ROC curves comparing SMAGS-LASSO versus LASSO for the colorectal cancer biomarker dataset in both training (Figure 2-a) and testing (Figure 2-b) sets. We see 21.8% increase in sensitivity at SP = 0.985 for SMAGS-LASSO vs LASSO in test set (p-value = 2.24E-04) and 54.2% increase in sensitivity in training set (p-value < 0.0001). We used cross-validation for determining the optimal λ value of λ = 0.161 for the colorectal cancer for SMAGS-LASSO where it selects 3 biomarkers––DKK1, GDF-15 and OPG. Supplementary Figure 2 illustrates the cross-validation process for determining the optimal λ value for colorectal cancer for SMAGS-LASSO. The sensitivity MSE plot in Supplementary Figure 1 shows clear minima, indicating robust model selection. We then use λ = 0.085 for LASSO which selects 3 biomarkers for comparison with SMAGS-LASSO. Table 2 presents detailed AUC and sensitivity performances for the colorectal cancer across SMAGS-LASSO and LASSO. Table 2 shows that SMAGS-LASSO also outperforms LASSO in terms of AUC with having an AUC of 0.97 (95% CI: 0.959–0.980) in training set and 0.94 (95% CI: 0.901–0.970) in test set which represent an increase of 7.4% in training set (p-value < 0.0001) and 4.8% in test set (p-value < 0.0001).

Figure 2.

Figure 2.

ROC curves for colorectal cancer biomarker detection on (a) training data and (b) test data. This figure compares the performance of SMAGS-LASSO versus LASSO methods. Both methods selected 3 biomarkers, with λ = 0.086 for LASSO and λ = 0.161 for SMAGS-LASSO.

Table 2.

Performances of the models

Model Dataset AUC (95% CI) Sensitivity at Specificity (95% CI) #Selected Biomarkers (total = 39) λ
LASSO Training 0.896 (0.876–0.916) 0.271 (0.200–0.433) 3 0.85
Test 0.892 (0.844–0.936) 0.577 (0.162–0.683)
Random Forest Training 0.873 (0.850–0.895) 0.287 (0.213–0.422) 3
Test 0.854 (0.797–0.902) 0.410 (0.050–0.583)
SMAGS-LASSO Training 0.970 (0.958–0.980) 0.813 (0.719–0.864) 3 0.161
Test 0.940 (0.901–0.973) 0.795 (0.459–0.876)
*

Note:

Specificity SP = 0.985 [17]

ΔSensitivity SMAGS-LASSO and LASSO in training set = 0.542 (p-value < 0.0001)

ΔSensitivity SMAGS-LASSO and Random Forest in training set = 0.526 (p-value < 0.0001)

ΔSensitivity SMAGS-LASSO and LASSO in test set = 0.218 (p-value = 2.24E-04)

ΔSensitivity SMAGS-LASSO and Random Forest in test set = 0.385 (p-value = 4.617E-08)

ΔAUC SMAGS-LASSO and LASSO in training set = 0.074 (p-value < 0.0001)

ΔAUC SMAGS-LASSO and Random Forest in training set = 0.097 (p-value < 0.0001)

ΔAUC SMAGS-LASSO and LASSO in test set = 0.048 (p-value < 0.0001)

ΔAUC SMAGS-LASSO and Random Forest in test set = 0.086 (p-value < 0.0001)

We also performed Random Forest method to compare with SMAGS-LASSO for feature selection. We trained a RF model with 500 trees and importance scoring enabled, then selected the top three features based on Mean Decrease in Gini impurity scores. To ensure a fair comparison with SMAGS-LASSO (which selected 3 features), we limited our RF feature selection to the same number of features. We then trained logistic regression using only these three RF-selected features for final predictions. Table 2 shows the results of detailed AUC and sensitivity at SP = 0.985 for this RF feature selection approach. SMAGS-LASSO demonstrated superior performance with an increase of 21.8% in training (p-value < 0.0001) and 38.5% (p-value < 0.0001) in test set sensitivity at SP = 0.985 compared to the RF approach. SMAGS-LASSO also showed increased AUC performance versus RF––an increase of 9.7% in training (p-value < 0.0001) and 8.6% in test set (p-value < 0.0001). We show the ROC curves for comparing SMAGS-LASSO versus LASSO and RF in Supplementary Figure 3.

4-. Discussion

Our study introduces SMAGS-LASSO, a novel feature selection method that integrates the SMAGS framework with L1 regularization to maximize sensitivity at predefined specificity thresholds while maintaining feature parsimony. The results from both synthetic datasets and real-world colorectal cancer biomarker data demonstrate several key advantages of our approach.

SMAGS-LASSO consistently outperformed LASSO method in terms of maximizing sensitivity at target specificity. This was particularly evident in the bidirectional synthetic dataset, where SMAGS-LASSO maintained perfect sensitivity while LASSO struggled with the complex signal patterns. Using publicly available colorectal cancer biomarker dataset SMAGS-LASSO outperformed both LASSO and RF in sensitivity at SP = 0.985 with more than 50% increase in training set and 20% increase in test set. We also saw an increase in AUC performance of more than 7% in training set and more than 5% in test set compared with both LASSO and RF methods.

Although blood-based colorectal cancer-specific antigens have been extensively investigated, only two serum biomarkers—carcinoembryonic antigen (CEA) and carbohydrate antigen 19–9 (CA 19–9)—are currently used in clinical practice [18]. However, both markers characterized by low sensitivity and specificity [19], highlighting the pressing need for more robust and accessible biomarkers capable of reliably detecting colorectal cancer and monitoring its progression. In this context, SMAGS-LASSO identified three promising serum-based biomarkers: Dickkopf-related protein 1 (DKK1), growth differentiation factor 15 (GDF-15), and TNF receptor superfamily member 11b (OPG/TNFRSF11B). Notably, these markers are biologically interconnected through major cancer-related signaling pathways, including WNT/β-catenin and RANK/RANKL [20], [21]. Previous studies have also reported these proteins as biomarkers for early detection of colorectal cancer [22], [23], [24] which further supports the robustness and relevance of SMAGS-LASSO.

The computational efficiency of SMAGS-LASSO represents another significant advantage. By integrating feature selection directly into the sensitivity optimization process, our method eliminates the need for separate preprocessing steps. The parallel optimization approach efficiently explores multiple solution paths, increasing the probability of finding the global optimum. The cross-validation framework successfully identified optimal regularization parameters across different cancer types, as evidenced by the clear minima in the sensitivity MSE curves.

Despite these promising results, several limitations of our approach warrant discussion. First, the sensitivity MSE metric used in cross-validation may not fully capture the complexity of model performance in imbalanced datasets. While it effectively guides regularization parameter selection toward models with high sensitivity, incorporating measures of specificity variation could provide more robust model selection, particularly in highly imbalanced clinical datasets. The current implementation relies on multiple optimization algorithms run in parallel, which increases computational requirements. Future work could explore more efficient optimization techniques specifically designed for non-differentiable objectives with L1 regularization, such as proximal gradient methods or coordinate descent algorithms that have proven effective for LASSO-type problems [25].

The applicability of SMAGS-LASSO may vary across cancer types depending on existing screening infrastructure and underlying biology. For cancers with established screening (colonoscopy, mammography) [26,27], blood-based panels must demonstrate clear advantages and may require higher specificity thresholds to avoid unnecessary procedures. Conversely, cancers lacking effective screening programs [28] may tolerate lower specificity to maximize detection. Additionally, cancers with distinct molecular subtypes such as breast cancer [29] may benefit from subtype-specific models rather than pan-cancer approaches, as training on heterogeneous populations could dilute biomarker signals. Future work should investigate whether separate SMAGS-LASSO models for cancer subtypes yield superior performance and whether optimal regularization parameters vary with the strength of underlying genetic or protein signatures.

The integration of feature selection directly into sensitivity optimization has significant clinical implications, particularly for early cancer detection. By identifying minimal sets of biomarkers that maintain high sensitivity at clinically acceptable specificity thresholds, SMAGS-LASSO could enable more cost-effective and practically implementable screening programs.

For colorectal cancer, which is typically diagnosed at advanced stages with poor prognosis, the identification of three biomarkers achieving 79.5% sensitivity at 98.5% specificity represents a promising advancement. This suggests that SMAGS-LASSO could contribute to developing more effective multi-marker screening panels for colorectal cancer and potentially other cancers.

The reduction in feature set size––from 39 to 3 biomarkers in the colorectal cancer example––without compromising performance has practical benefits beyond computational efficiency. Smaller biomarker panels reduce testing costs, simplify clinical implementation, and potentially decrease measurement noise. This parsimony is particularly valuable in resource-constrained healthcare settings where comprehensive biomarker testing may not be feasible.

4-A. Conclusion

In conclusion, we have developed SMAGS-LASSO, a novel feature selection method that integrates sensitivity maximization with L1 regularization for early cancer detection. Our method demonstrates superior performance compared to standard LASSO and Random Forest approaches, identifying biologically relevant biomarkers achieving 79.5% sensitivity at 98.5% specificity for colorectal cancer. While limitations exist and further validation is needed, SMAGS-LASSO offers a promising framework for biomarker discovery where sensitivity-specificity tradeoffs are critical. Future work will focus on addressing limitations discussed earlier and validating the method across larger and more diverse clinical datasets to advance precision medicine and early detection strategies.

Supplementary Material

1
2
3
4

5-. Acknowledgements

Supported by National Institute of Health (NIH) Grant Nos. U01CA271888 (S. Hanash.), RP160693; (K.A. Do); Specialized Programs of Research Excellence (SPOREs) (P50CA140388; K.A. Do, J.P. Long.); Center for Clinical and Translational Science (CCTS) (TR000371; K.A. Do, J.P. Long.); and all authors received the generous philanthropic contributions to The University of Texas MD Anderson Cancer Center Moon Shots Program and the Lyda Hill Foundation.

Footnotes

The authors declare no potential conflicts of interest.

References

  • [1].Ghasemi SM, Gu C, Fahrmann JF, Hanash S, Do K, Long JP et al. , “A Novel Sensitivity Maximization at a Given Specificity Method for Binary Classifications,” Cancer Prev Res (Phila), vol. 18, no. 3, pp. 117–123, Mar. 2025, doi: 10.1158/1940-6207.CAPR-24-0236. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [2].Irajizad E, Kenney A, Tang T, Vykoukal J, Murage E, Dennison JB et al. , “A blood-based metabolomic signature predictive of risk for pancreatic cancer,” Cell Rep Med, vol. 4, no. 9, p. 101194, Sep. 2023, doi: 10.1016/j.xcrm.2023.101194. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [3].Fahrmann JF, Marsh T, Irajizad E, Patel N, Murage E, Vykoukal J et al. , “Blood-Based Biomarker Panel for Personalized Lung Cancer Risk Assessment,” Journal of Clinical Oncology, vol. 40, no. 8, pp. 876–883, Mar. 2022, doi: 10.1200/JCO.21.01460. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [4].Crosby D, Bhatia S, Brindle KM, Coussens LM, Dive C, Emberton M et al. , “Early detection of cancer,” Science (1979), vol. 375, no. 6586, Mar. 2022, doi: 10.1126/SCIENCE.AAY9040. [DOI] [Google Scholar]
  • [5].Zhang P, “A novel feature selection method based on global sensitivity analysis with application in machine learning-based prediction model,” Appl Soft Comput, vol. 85, p. 105859, Dec. 2019, doi: 10.1016/J.ASOC.2019.105859. [DOI] [Google Scholar]
  • [6].Efimov D and Sulieman H, “Sobol Sensitivity: A Strategy for Feature Selection,” Springer Proceedings in Mathematics and Statistics, vol. 190, pp. 57–75, 2017, doi: 10.1007/978-3-319-46310-0_4. [DOI] [Google Scholar]
  • [7].Wang J, Wei JM, Yang Z, and Wang SQ, “Feature selection by maximizing independent classification information,” IEEE Trans Knowl Data Eng, vol. 29, no. 4, pp. 828–841, Apr. 2017, doi: 10.1109/TKDE.2017.2650906. [DOI] [Google Scholar]
  • [8].García-Nieto J, Alba E, Jourdan L, and Talbi E, “Sensitivity and specificity based multiobjective approach for feature selection: Application to cancer diagnosis,” Inf Process Lett, vol. 109, no. 16, pp. 887–896, Jul. 2009, doi: 10.1016/J.IPL.2009.03.029. [DOI] [Google Scholar]
  • [9].Xu Z, King I, Lyu MRT, and Jin R, “Discriminative semi-supervised feature selection via manifold regularization,” IEEE Trans Neural Netw, vol. 21, no. 7, pp. 1033–1047, Jul. 2010, doi: 10.1109/TNN.2010.2047114. [DOI] [PubMed] [Google Scholar]
  • [10].Li F, Lai L, and Cui S, “On the Adversarial Robustness of LASSO Based Feature Selection,” IEEE Transactions on Signal Processing, vol. 69, pp. 5555–5567, 2021, doi: 10.1109/TSP.2021.3115943. [DOI] [Google Scholar]
  • [11].Tibshirani R, “Regression Shrinkage and Selection Via the Lasso,” J R Stat Soc Series B Stat Methodol, vol. 58, no. 1, pp. 267–288, Jan. 1996, doi: 10.1111/J.2517-6161.1996.TB02080.X. [DOI] [Google Scholar]
  • [12].Nelder JA and Mead R, “A Simplex Method for Function Minimization,” Comput J, vol. 7, no. 4, pp. 308–313, Jan. 1965, doi: 10.1093/COMJNL/7.4.308. [DOI] [Google Scholar]
  • [13].Nocedal J and Wright SJ “Sequential Quadratic Programming,” Numerical Optimization, pp. 526–573, Jun. 1999, doi: 10.1007/0-387-22742-3_18. [DOI] [Google Scholar]
  • [14].Byrd RH, Lu P, Nocedal J, and Zhu C, “A Limited Memory Algorithm for Bound Constrained Optimization,” SIAM Journal on Scientific Computing, vol. 16, no. 5, pp. 1190–1208, 1995, doi: 10.1137/0916069. [DOI] [Google Scholar]
  • [15].Cohen JD, Li L, Wang Y, Thoburn C, Afsari B, Danilova L et al. , “Detection and localization of surgically resectable cancers with a multi-analyte blood test,” Science (1979), vol. 359, no. 6378, pp. 926–930, Feb. 2018, doi: 10.1126/science.aar3247. [DOI] [Google Scholar]
  • [16].Vacante M, Borzì AM, Basile F, and Biondi A, “Biomarkers in colorectal cancer: Current clinical utility and future perspectives,” World Journal of Clinical Cases, vol. 6, no. 15, pp. 869–881, Dec. 2018, doi: 10.12998/WJCC.V6.I15.869. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [17].Imperiale TF, Ransohoff DF, Itzkowitz SH, Levin TR, Lavin P, Lidgard GP et al. , “Multitarget Stool DNA Testing for Colorectal-Cancer Screening,” New England Journal of Medicine, vol. 370, no. 14, pp. 1287–1297, Apr. 2014, doi: 10.1056/NEJMOA1311194. [DOI] [PubMed] [Google Scholar]
  • [18].Hauptman N and Glavač D, “Colorectal Cancer Blood-Based Biomarkers,” Gastroenterol Res Pract, vol. 2017, no. 1, p. 2195361, Jan. 2017, doi: 10.1155/2017/2195361. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [19].Lakemeyer L, Sander S, Wittau M, Henne-Bruns D, Kornmann M, and Lemke J, “Diagnostic and Prognostic Value of CEA and CA19–9 in Colorectal Cancer,” Diseases 2021, vol. 9, no. 1, p. 21, Mar. 2021, doi: 10.3390/DISEASES9010021. [DOI] [Google Scholar]
  • [20].Wang Y, Liu Y, Huang Z, Chen X, and Zhang B, “The roles of osteoprotegerin in cancer, far beyond a bone player,” Cell Death Discov, vol. 8, no. 1, p. 252, Dec. 2022, doi: 10.1038/s41420-022-01042-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [21].Fujita KI and Janz S, “Attenuation of WNT signaling by DKK-1 and −2 regulates BMP2-induced osteoblast differentiation and expression of OPG, RANKL and M-CSF,” Mol Cancer, vol. 6, no. 1, pp. 1–13, Oct. 2007, doi: 10.1186/1476-4598-6-71. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [22].Tsukamoto S, Ishikawa T, Lida S, Ishiguro M, Mogushi K, Mizushima H et al. , “Clinical significance of osteoprotegerin expression in human colorectal cancer,” Clinical Cancer Research, vol. 17, no. 8, pp. 2444–2450, Apr. 2011, doi: 10.1158/1078-0432.CCR-10-2884. [DOI] [PubMed] [Google Scholar]
  • [23].Chen X, Zeng Q, Yin L, Yan B, Wu C, Feng J et al. , “Enhancing immunotherapy efficacy in colorectal cancer: targeting the FGR-AKT-SP1-DKK1 axis with DCC-2036 (Rebastinib),” Cell Death & Disease 2025 16:1, vol. 16, no. 1, pp. 1–15, Jan. 2025, doi: 10.1038/s41419-024-07263-8. [DOI] [Google Scholar]
  • [24].Wallin U, Glimelius N, Jirström K, Darmanis S, Nong RY, Pontén F et al. , “Growth differentiation factor 15: A prognostic marker for recurrence in colorectal cancer,” Br J Cancer, vol. 104, no. 10, pp. 1619–1627, May 2011, doi: 10.1038/BJC.2011.112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [25].Beck A, T.-S.M. journal on imaging sciences, and undefined 2009, “A fast iterative shrinkage-thresholding algorithm for linear inverse problems,” SIAM Journal on Imaging Sciences, vol. 2, no. 1, pp. 183–202, 2009, doi: 10.1137/080716542. [DOI] [Google Scholar]
  • [26].Davidson KW, Barry MJ, Mangione CM, Cabana M, Caughey AB, Davis EM et al. , “Screening for Colorectal Cancer: US Preventive Services Task Force Recommendation Statement,” JAMA, 2021;325(19):1965–1977, doi: 10.1001/jama.2021.6238. [DOI] [PubMed] [Google Scholar]
  • [27].Nicholson WK, Silverstein M, Wong JB, Barry MJ, Chelmow D, Coker TR et al. , “Screening for Breast Cancer: US Preventive Services Task Force Recommendation Statement,” JAMA, 2024;331;(22):1918–1930, doi: 10.1001/jama.2024.5534. [DOI] [PubMed] [Google Scholar]
  • [28].Waleleng BJ, Adiwinata R, Wenas NT, Haroen H, Rotty L, Gosal F et al. , “Screening of pancreatic cancer: Target population, optimal timing and how?,” Ann Med Surg (Lond), 2022. Nov 5;84:104814. doi: 10.1016/j.amsu.2022.104814. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [29].Perou CM, Sørlie T, Eisen MB, van de Rijn M, Jeffrey SS, Rees CA et al. , “Molecular portraits of human breast tumors,” Nature, vol. 406, no. 6797, pp. 747–752, Aug. 2000, doi: 10.1038/35021093. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1
2
3
4

Data Availability Statement

The synthetic data generated in this study as well as the real-world protein biomarker data are publicly available in https://github.com/khoshfekr1994/SMAGS.LASSO.

RESOURCES