Skip to main content
NPJ Systems Biology and Applications logoLink to NPJ Systems Biology and Applications
. 2025 Dec 12;11:137. doi: 10.1038/s41540-025-00614-x

Feature learning augmented with sampling and heuristics (FLASH) improves model performance and biomarker identification

Shivam Kumar 1, Abhinav Agarwal 1, Samrat Chatterjee 1,✉
PMCID: PMC12700939  PMID: 41387731

Abstract

Big biological datasets, such as gene expression profiles, often contain redundant features that degrade model performance and limit generalization across independent datasets with complexities like class imbalance and hidden sub-clusters. To overcome challenges, we present ‘FLASH’, a novel feature selection method combining filtration and heuristic-based systematic elimination. FLASH generates random samples and computes p-values for each feature using multiple statistical tests (t-test, ANOVA, Wilcoxon Rank-Sum, Brunner–Munzel, Mann–Whitney). Features are scored by aggregating significant p-values across samples. The coefficient from the machine learning model with the highest accuracy on the filtered features is used to rank them. Recursive elimination with cross-validation systematically removes features while monitoring accuracy. The final subset is selected based on the highest performance during elimination, to achieve effective feature selection. We show that our method preserves predictive performance on independent datasets. Our comprehensive evaluation across diverse datasets showed that FLASH outperforms the compared feature selection methods dRFE, Mutual information, MRMR, ElasticNet, NeuralNet, Permutation test and SAGA within the scope of our tested datasets and evaluation settings. Additionally, features selected by FLASH demonstrated greater biological relevance, as evidenced by higher overlap with disease-associated genes from DisGeNET in an independent dataset.

Subject terms: Computational biology and bioinformatics, Mathematics and computing

Introduction

In the era of modern biomedical technologies, rapid advancements and innovations have accumulated vast amounts of biological omics data1–3. These datasets consist of thousands of features (p) for each sample (n), providing valuable insights into various biological phenomena. However, the number of available samples is very small due to challenges in recruiting participants. The number further narrows down in the case of omic studies due to the high cost associated with data generation4,5. This high cost leads to fewer samples with omics data, which comes under the big data category, causing smaller n with larger p. Analyzing these omics data with “large p, small n” poses significant challenges, such as identifying features associated with specific phenotypes or diseases6.

The curse of dimensionality may occur when training a model with a significantly smaller number of samples using all the omic features7. Feature selection algorithms offer a potential solution to this challenge by identifying a subset of features significantly associated with phenotypes of interest8. These algorithms aim to reduce the dimensionality of the data, improving computational efficiency and enhancing the interpretability of the results. The need for more accurate and efficient feature selection tools is evident, especially in biomedical research. For instance, the PAM50 gene panel has proven effective in accurately defining major breast cancer subtypes based on the expression profiles of 50 genes9. Similarly, the commercial availability of the 21-gene panel OncoType DX has contributed significantly to estimating the likelihood of recurrence in early-stage female ER+ breast cancer10. There is an increasing demand for guided diagnostic and treatment approaches11–13. These tests’ success depends on identifying disease signatures, which capture disease patterns robustly in large populations. So, developing advanced feature selection algorithms becomes more crucial for enhancing patient care and overall outcomes. Feature selection is a challenging task, as it falls under the category of problems in computational complexity theory called NP-hard problems (nondeterministic polynomial time), which requires exponential time to find the global optimal solution14. As a result, existing feature selection algorithms rely on heuristic rules to find local optimal solutions, leading to variations in their performance across different datasets. These rules can be categorized into three classes. Filter methods utilize the intrinsic properties of the data to select the optimal feature subset without relying on any learning algorithm15. Wrapper methods employ a learning algorithm to evaluate different feature subsets and evolve towards an optimal solution through multiple iterations16. Embedded methods perform feature filtration during model training and perform simultaneous optimization with feature elimination and classification17.

Several algorithms have been developed to address the feature selection problem in this context. The AltWOA enhances the conventional Whale Optimization Algorithm by introducing altruism within the whale population, resulting in more effective global optimization18. Similarly, enCFS employs ensemble techniques and sampling to enhance feature selection’s discrimination and stability19. The VNLHHO offers a metaheuristic approach that balances global exploration and local exploitation for gene feature extraction20. The MBAO combines a filtering approach with an efficient wrapper to select the most informative genes21. The MGRFE leverages an embedded integer-coded genetic algorithm for minimal gene combinations with maximal information22. The RIFS2D algorithm challenges conventional wisdom by showcasing the predictive potential of low-ranked features, even when top-ranked features fall short23. Numerous approaches rely on wrapper methods that evaluate and compare subsets of features according to their predictive value. Like many other feature selection methods, they tend to be computationally intensive and require substantial time and resources24. These computational limitations motivate the need for techniques that improve efficiency and scalability when handling large-scale or high-dimensional data. Most methods use a filter score like a t-test, Fisher score, Wilcoxon test correlation, etc. It requires users to choose a predefined cut-off value depending on the nature of the data. The elimination step requires users to select a minimal number of features as a model parameter. In our opinion, the algorithm should choose the minimum number of features. One common drawback among these methods is the absence of sampling, which can limit their ability to handle imbalanced or high-dimensional datasets effectively. By incorporating sampling techniques, we can achieve a more balanced and representative view of the data, which helps reduce variance and improves the stability of feature selection–particularly in complex or high-dimensional settings. When these sampling strategies are embedded directly within feature selection methods, they enable the identification of features that consistently perform well across various data subsets. This consistency mitigates the model’s sensitivity to data fluctuations and reduces the risk of overfitting. As a result, it enhances generalizability, boosts computational efficiency, and yields more robust and accurate feature subsets for downstream analysis and modeling tasks.

In this paper, we propose a feature selection algorithm that employs a sampling technique to identify features that are consistently informative across subsets of the data, to enhance the classification of new samples. This algorithm aims to provide more robust results in complex datasets with class imbalance and many features. The proposed feature selection algorithm, FLASH, which stands for Feature learning augmented with sampling and heuristics, is a hybrid of filter and recursive elimination methods. First, we created the filter scores by applying statistical tests to large randomly sampled subsets from the data. Random sampling helps reduce bias toward specific patterns in the dataset and improves the generalizability of the selected features to unseen or external datasets. We employed five statistical tests to ensure robustness across different distributional assumptions one-way ANOVA25, rank-sum test25, Mann–Whitney U test25, t-test25, and the Brunner–Munzel test26. Random sample selection from the population reduces bias towards the data under investigation and increases its chance for global application on other data27. This global application increases the scope of our algorithm by mitigating the potential impact of outliers or specific characteristics present in a non-randomly selected sample, thereby enhancing the chance of improved accuracy when applied to a general population28. Through extensive evaluation across diverse datasets, we demonstrate the effectiveness and efficiency of our proposed algorithm, FLASH, in selecting informative features associated with the phenotype of interest. To assess the biological relevance of the selected features, we benchmarked FLASH by comparing its outputs on an independent dataset against known disease-associated genes curated in the DisGeNET database. This comparison highlights FLASH’s ability to generate reliable and meaningful biomedical insights.

Results

The impact of FLASH on different ML algorithms

This study evaluates a feature subset by assessing its classification performance using multiple representative classifiers employing a repeated stratified 10-fold cross-validation approach. The metric employed, called maximum accuracy (mAcc), represents the highest classification accuracy achieved among these classifiers. A total of six popular classifiers are used to evaluate the given feature subset (Fig. 1). These classifiers include support vector machine (SVM), Random Forest, XGBoost, logistic regression (LR), decision tree (DTree), and K nearest neighbor (KNN). The result is derived from the model performance on the selected feature sets. By leveraging the diverse capabilities of these classifiers, a comprehensive assessment of the feature subset’s classification performance is achieved, providing valuable insights for feature selection in the studied context.

Fig. 1. Performance comparison over multiple ML algorithms.

Fig. 1

A mAcc, B F1-Score.

Among the evaluated classifiers, SVM, KNN, and Logistic Regression consistently demonstrated superior performance in compared to non-linear methods such as Decision Trees, particularly regarding accuracy and F1-score. These findings suggest that the effectiveness of FLASH features may vary depending on the choice of ML algorithm, highlighting the importance of specific characteristics and requirements of the dataset when selecting an appropriate ML approach.

Ablation study: FLASH boosts performance beyond baseline

The filtering approach requires fewer computational resources than RFECV and helps identify initial features29. Since applying RFECV directly to a dataset with numerous features is inefficient29, a preliminary filtering step becomes essential. However, it is important to note that additional efforts are required to systematically optimize the final set of features for the classification algorithm. The evaluation of an ML model is required at every iteration after elimination to capture the best subset based on accuracy. This process allows the selection of features to contribute to the classification performance.

In some cases, the mAcc has reached its peak value of 1, indicating good classification performance (Fig. 2). The improvement becomes more significant when the F1-score is considered. For instance, in the case of the GSE755 dataset, the F1-Score initially measured below 0.4 but experienced a substantial jump after embedding RFECV, resulting in an F1-Score of 0.77. These findings highlight the effectiveness and impact of incorporating RFECV in enhancing classification performance, particularly regarding accuracy and precision. The performance has also been compared with the baseline, which contains the complete set of features. We observed that using all the features led to a lower model accuracy. Literature also supports the observation that including redundant information may hinder the learning behavior of the model30–32.

Fig. 2. Importance of feature selection.

Fig. 2

Performance comparison of filter score with and without RFECV over baseline (using all features). Results are shown for two metrics: A mAcc, B F1-Score.

Association between FLASH selected feature and outcome label

To assess the interpretability and coherence of FLASH-selected features. We visualized the correlation coefficient of the top-ranked genes concerning class labels across 11 datasets. Each dataset exhibits a distinct distribution of gene-phenotypic association as shown in (Fig. 3). For the remaining five datasets, refer to (Supplementary Fig. 4).

Fig. 3. Correlation profiles of FLASH-selected genes across six transcriptomics datasets.

Fig. 3

Each plot shows horizontal barplots of correlation coefficients between top-ranked genes and class labels. The supplementary material (Supplementary Table 3) provides the Gene Name/ID list plotted along the x-axis for each dataset.

GSE37751 and GSE40419 exhibited strongly negative correlation peaks (ρ < 0.6), which could be potential tumor promoters in the disease state. GSE18864 displayed multiple transcripts with high positive correlation (ρ > 0.4). These positively correlated genes may represent candidates for tumor-suppressive roles or markers of normal physiological states.

Glioblastoma_2 presents a sparse but high magnitude correlation. It could be the presence of oncogenic drivers. These findings underscore FLASH’s strength: to adapt feature selection to the statistical structure of each dataset. While still covering biologically meaningful signals across diverse conditions.

Benchmarking FLASH with contemporary feature selection algorithms

The proposed algorithm FLASH’s performance is compared with recently and widely used feature selection algorithms like MRMR33, Mutual information (Mut-Info)34, dRFE35, and SAGA36. We also compared FLASH’s performance with intrinsic ML models such as a neural network and ElasticNet37 to demonstrate the necessity of extrinsic feature selection in imbalanced datasets. Additionally, we used a permutation test38, a non-parametric method that assesses the significance of observed group differences. The test shuffles labels to build a reference distribution and compute a p-value, without relying on data distribution assumptions. The repeated cross-validation approach and the metrics,mAcc and F1 score, were also used.

The FLASH showed better results than other algorithms, reaching over 90% mean accuracy on 8 out of 11 datasets. In the Illumina HiSeq datasets, the accuracy exceeded 95%. The value sometimes reaches nearly 100% (see Fig. 4A). For instance, in the GSE99309 dataset, the other algorithms achieved ~0.65 mAcc, while FLASH attained a good mAcc 1.

Fig. 4. Comparing FLASH with other feature selection methods.

Fig. 4

All algorithms except NeuralNet, ElasticNet, and Permutation test have explicit feature selection methods. NeuralNet, ElasticNet, and Permutation test have intrinsic feature selection methods. Results are presented in four metrics: A mAcc, B F1-Score, C Precision, D Recall.

It is noteworthy that FLASH was performed on both simple and complex datasets. For example, in GSE18864, a less complex dataset, other feature selection algorithms35 achieved an mAcc of 1, and FLASH also demonstrated an mAcc of 1. In the GSE35725 dataset, FLASH exhibits a significantly better result than other methods. The datasets GSE233242 and GSE40419, based on Illumina HiSeq, showed superior performance with the FLASH algorithm, achieving an average accuracy of 98.55%. For a detailed table, refer to (Supplementary Table 4). These findings show better performance of the FLASH compared to the other methods across multiple datasets, demonstrating its effectiveness in achieving high classification accuracy on different datasets.

Comparison of FLASH features with other algorithms

To compare the selected gene sets of FLASH with methods like dRFE, Mut_info, MRMR, and SAGA. We performed the Jaccard similarity analysis across 11 datasets, as shown in (Fig. 5). For details on the total number of features selected, please refer to (Supplementary Table 5). Jaccard index between genes selected by FLASH and those selected by different algorithms. FLASH showed low similarity with most algorithms across several datasets. Reinforcing its tendency to select a more parsimonious and informative feature set.

Fig. 5. Comparing features selected by FLASH with other feature selection methods.

Fig. 5

Jaccard index between genes selected by FLASH and other methods.

The Jaccard matrix shows distinct overlap trends, indicating that dRFE shared a moderate Jaccard similarity with FLASH as well as with datasets like Glioblastoma_2, GSE19159, and GSE5764, from which we can suggest that recursive feature elimination methods may identify comparable biologically significant features. Mutual Information and SAGA consistently showed low overlap across the datasets. It shows a difference in their feature selection approaches. FLASH recovers stable and biologically plausible features while offering unique gene sets that differ meaningfully from traditional methods. It potentially contributes novel insights into disease-specific molecular signatures.

To examine the characteristics of selected features, we evaluated their effect size distributions using Cohen’s d. Analysis of Cohen’s d39 showed variation in effect size distributions across different feature selection methods. For a detailed table, refer to (Supplementary Table 5). Mut_info and mRMR often selected features with large effect sizes (Cohen’s d > 0.8), while FLASH and SAGA usually selected features with moderate effect sizes (Cohen’s d 0.2–0.8). In Glioblastoma_2, FLASH showed an enrichment of strong-effect features (Cohen’s d > 0.8) and achieved notably high accuracy compared to other feature selection methods, reaching a maximum of 100%. In GSE99039, FLASH prioritized features with moderate effect sizes (~0.5) while reaching a maximum accuracy of 100%. FLASH did not focus only on strong univariate signals, but it also detected dispersed patterns of variation, and balanced statistical evidence with biological relevance by selecting features that represented broad distributions.

FLASH improves ML model performance on independent datasets

We conducted a case study using two breast cancer datasets, Illumina GSE5219440,41 and microarray GSE2142242,43 to evaluate the generalizability and, in contrast, the specificity of our selected features. This analysis also positions our method as a potential biomarker identification tool for gene expression datasets. The feature set identified from the breast cancer datasets GSE233242 and GSE5764 was applied to predict the disease categories of two independent datasets, GSE52194 and GSE21422. The performance of FLASH was evaluated on multiple metrics, in the present data, and the accuracy obtained by FLASH is comparable with other methods, (Fig. 6A). However, FLASH’s significance lies in its ability to capture features related to the system on which it is applied.

Fig. 6. Evaluating FLASH on an independent dataset.

Fig. 6

A Performance comparison of features selected from breast cancer data on independent datasets. B Matched breast cancer-related genes (from DisGeNet) and selected genes for different methods.

To further validate the biological relevance of the selected features, we compared them with breast cancer-related genes from the DisGeNet database44. FLASH-identified features exhibited a 70% overlap for GSE52194 and an 83% overlap for GSE223242 with disease-associated genes from the DisGeNet database, significantly higher than the overlap achieved by the other methods (Fig. 6B). This high overlap underscores the effectiveness of FLASH in identifying biologically meaningful features pertinent to diabetes, thereby demonstrating its potential as a robust tool for biomarker discovery in gene expression studies.

Discussion

Building predictive models using molecular signatures is crucial in biomedical research. There are many applications of such models, like diagnosis, prognosis, etc., for example, MammaPrint45, Decipher46, VeriStrat47 etc., have been developed for predicting recurrence and survival in breast and lung cancer. Unfortunately, the number of such tests is very low and is only available for a few diseases. The limited availability is due to the dimensionality of the data, which limits the models’ training ability and makes them less robust. Feature selection provides a possible solution to this problem by finding a smaller subset from a large pool of features. Identification of such a meaningful subset of features improves model interpretability and computational efficiency.

In this study, we have introduced a feature selection algorithm, FLASH, that combines elements of both filter and recursive elimination methods. The fusion approach first filters the identified candidate genes with significant discriminatory power. First, we created thousands of data snapshots to filter the genes by randomly sampling subsets from the original data. Then, we implemented a t-test, One-way ANOVA, Wilcoxon Rank-Sum, Brunner–Munzel, and Mann–Whitney U on each randomly drawn set and obtained p-values for every feature in the respective random set. The average of statistically significant p-values (<0.01) was taken after transforming them to -log10(p-value) to obtain a score which describes the character of a feature being significant in most of the data snapshots. The sampled significant score ensures the reliability and validity of the selected features. Subsequently, we utilize recursive elimination to refine the feature set, systematically removing less informative features and evaluating the cross-validated accuracy at each iteration. We selected the One-way ANOVA test among the five statistical methods based on accuracy scores, using the Kneedle algorithm to identify the optimal cutoff. We have used twenty gene expression datasets from the literature to evaluate FLASH with other well-established and recent algorithms like dRFE, MRMR, SAGA, etc. These datasets have various degrees of complexity, such as large feature sizes, smaller samples, and an imbalanced distribution of classes/ labels. Considering the selected datasets, first, we have demonstrated the importance of feature selection by comparing the predictive performance of the feature subset over the complete feature set. The result showed a gain in accuracy in all the datasets, and in eight out of eleven datasets, this gain was more than 90%. In the evaluation datasets, Gliboblastoma_2, GSE35725 and GSE18864 have an accuracy of 100% previously reported in the literature22,23,35. However, datasets like GSE233242, GSE5764, GSE40419, and GSE99039 have an average accuracy of approximately 90%, with limited consensus among different studies. Some of the datasets’ reported accuracy was below 70%, making it very challenging to translate them at a clinical level. In our experimentation with such challenging datasets, FLASH achieved more than 90% accuracy in 8 out of 11 datasets. Other metrics, such as precision and recall, which are relevant to understanding category-wise prediction results, are also consistent.

Our case study on a diabetic dataset underscores our approach’s practical utility for disease-associated signature identification. The other prediction algorithms have shown around 90% accuracy in the original dataset22,23,35. However, after evaluating the diabetes ML model on a new dataset as a test set, its performance dropped to 50%, equivalent to random prediction. Nonetheless, FLASH maintained its predictive accuracy with training and testing data around 90%. We have also observed that the selected features matched well with the known breast cancer-associated genes from the DisGeNet database.

In conclusion, this work introduced FLASH, an algorithm that addresses the limitations of existing feature selection methods by incorporating a sampling technique. FLASH enhances model performance, yielding better results on independent datasets. This method offers a new feature selection strategy that can generate robust biomarkers for disease diagnosis.

Methods

Dataset description

We have used 19 gene expression datasets to evaluate the feature selection algorithm, including three Illumina HiSeq datasets: GSE233242, GSE40419, and GSE134878. These datasets have complex characteristics such as high-class imbalance, large feature vs sample ratio and poor predictions by previous feature selection methods. Six datasets are widely recognized within the research community22,23,35 as standard benchmarks for evaluating feature selection methodologies. The description of the dataset is given in Table 1.

Table 1.

The datasets used in this study with defined characteristics

Name Characteristics Category No. of features Ref.
ALL-AML Category imbalance ALL: 47, AML: 25 7129 63
ALL2 Category imbalance relapse: 65, non-relapse: 35 12625 63
ALL3 Category imbalance non-mdr: 101, mdr: 24 12625 63
Breast Large feature size non-relapse: 51, relapse: 46 24481 64
CNS Category imbalance survivor: 39, non-survivor: 21 7129 65
Colon Category imbalance tumor: 40, non-tumor: 22 2000 66

Breast cancer

(GSE19159)

Category imbalance relapse: 111, non-relapse: 57 2905 67

Autistic

(GSE25507)

Large feature size autism: 82, non-autism: 64 54613 68

Lung cancer

(GSE30219)

Category imbalance+ early-stage: 198, late-stage: 93 54675 69

Diabetes

(GSE35725)

Large feature size T1D: 57, Healthy: 44 54675 70

Myeloma

(GSE755)

Category imbalance lesion: 137, non-lesion: 36 12625 71

Parkinson

(GSE99039)

Large feature size non-IPD: 233, IPD: 205 54675 72
Glibolastoma_2 Category imbalance glioblastoma: 77, non-tumor: 23 15435 73

Human Breast

(GSE37751)

Large feature size Tumor: 61, Non-tumor: 47 33297 74–76

Human Breast

(GSE18864)

Category imbalance+

Large feature size

treated: 60, non-treated: 24 54675 77–79

Human Breast

(GSE5764)

Large feature size non-cancer: 20, cancer: 10 54675 80

Human Breast

(GSE233242)

Large feature size tumor: 43, normal: 43 15044 81

Lung Cancer

(GSE40419)

Large feature size tumor: 86, normal: 77 36742 82

Essential Tremor

(GSE134878)

Large feature size tremor: 33, control: 21 25931 83

Additionally, we selected independent test datasets GSE5219440,41 and GSE2142242,43, which share characteristics with GSE223242 and GSE5764, respectively, both related to breast cancer. This dataset was used to validate the trained diabetes model. All the experiments were conducted on a computing server with 128 GB of system memory and a 48-core Intel Xeon CPU (2.20 GHz).

The feature learning augmented with sampling and heuristics (FLASH) algorithm has been established in the context of binary classification problems. Such problems have two groups of samples: positive and negative. The predictive model uses a certain set of features and an ML classifier to predict the groups using unknown/ new feature values.

FLASH employs an embedded approach comprising two components (Fig. 7). The first component is a filtration method that selects an initial pool of features, also known as candidate features. The second component of the algorithm is the recursive feature elimination with cross-validation (RFECV), which can handle univariate and multicollinearity48,49.

Fig. 7. Design of feature learning framework FLASH.

Fig. 7

The dataset block represents the initial data, and the color indicates the group/class distribution. The framework has three major phases: sampling, feature scoring, and feature learning. Sampling involves drawing a random set with repetition while maintaining a similar group representation as the original data. Feature scoring calculates the filter score by performing a significance test on the random set. The final step, feature learning, includes RFECV to generate an optimal set of features from the filtered pool.

The proposed algorithm: FLASH

The two-step approach of obtaining the filter score Fscore for each feature is first to randomly draw subsets from the complete dataset using pseudo code (See Algorithm 1). This subset selection strategy employs a stratified hybrid bootstrap sampling approach, ensuring that each class is proportionally represented in every resampled subset. In every iteration, we begin by selecting a portion of the samples from each class without replacement to preserve the diversity and uniqueness of the dataset. The remaining samples are then chosen with replacement, allowing us to introduce a deliberate level of redundancy in a controlled manner50,51.

To enhance the robustness of statistical assessment scores, varied data distributions. We applied five significance tests to each dataset: one-way ANOVA, rank-sum test, Mann–Whitney U test, t-test, and Brunner–Munzel test52,53. For each test, p-values were computed for all features across K = 1000 randomized subsets. The filter score for each feature is derived by transforming the p-values using a negative log transformation pmod = −log10(p-value)54. The transformed score for the ith subset is denoted as pmodi.

To calculate the Fscore, we consider only significant p-values, i.e., only those i’s for which pmodi>2. These pmodi’s are added and then divided by the total number of random samples.

So, for K randomization, we have

Fscore=∑i=1K1 {pmodi>2}K 1

Thus, we calculated the fraction of significant features over the total number of features to assess the influence of stratified random sampling. The fraction is sorted from lowest to highest value (Fig. 8). This fraction is sorted in ascending order, and the observed increase in filtration suggests inflation of type 1 error. A considerable variation is observed in the fraction of the significant features for different random sampling, with the highest in breast cancer data (0.53) and the lowest in GSE30219(0.22) and ALL-AML (0.20).

Fig. 8. Fraction of significant features for different random sampling.

Fig. 8

The fraction of significant features is plotted for different random samplings. Each panel shows results for a different dataset.

This large variation supports our hypothesis that substructures present in the data could influence the choice of significant features. Thus, removing features based on sampled p-values would improve the feature filtration for the downstream process.

Algorithm 1

Pseudo code to calculate p-values for sampled subsets using five statistical tests

Require: Dataset D with S samples, N features, and class labels Y; number of iterations K; sampling ratio α ∈ (0, 1)

Ensure: Matrices TTEST, WILCOX, ANOVA, BM, MWU∈RN×K storing p-values

 1: Partition D into class-wise subsets: D0, D1, …, DC where C is the number of classes

 2: Initialize matrices: TTEST, WILCOX, ANOVA, BM, MWU ← 0N×K

 3: for i = 1 to K do

 4: Dsub←∅

 5: for each class c ∈ {0, …, C} do

 6: Sc ← ∣Dc∣

 7: Draw ⌊α ⋅ Sc⌋ samples from Dcwithout replacement→Duniquec

 8: Draw Sc − ⌊α ⋅ Sc⌋ samples from Dcwith replacement→Drepc

 9: Dsub←Dsub∪ (Duniquec∪Drepc)

10: end for

11: for each feature j = 1 to N do

12: Compute t-test p-value → TTEST[j, i]

13: Compute Wilcoxon Rank-Sum p-value → WILCOX[j, i]

14: Compute One-way ANOVA p-value → ANOVA[j, i]

15: Compute Brunner–Munzel p-value → BM[j, i]

16: Compute Mann–Whitney U p-value → MWU[j, i]

17: end for

18: end for

The core principle of FLASH lies in training a supervised model using the initial screened features based on a predefined percentage cut-off on Fscore. This initial step filters and selects statistically significant features. We used these features to train classification models, including Decision Tree, K-Nearest Neighbours (KNN), Support Vector Machine (SVM), Logistic Regression, Random Forest, and XGBoost. The absolute value of the model’s coefficient with maximum accuracy ranked the features. By iterating through these ranked features, we systematically eliminated lower-ranked features from the set using recursive feature elimination with cross-validation (RFECV) and continued until the last feature. During this process, the model’s accuracy was monitored at each step, and the set with the highest model performance was selected as the optimal set. RFECV inherently includes an early stopping mechanism to optimize feature selection. However, in FLASH, we want to ensure that the optimal set corresponds to maximum accuracy, so we iterate through the entire feature set. If multiple optimal sets have the same accuracy, then FLASH choose the set with fewer features. This dynamic elimination strategy ensures that the final selected features are statistically significant and optimally contribute to the model’s predictive performance. (See Algorithm 2). The pseudo-code 2 outlines the step-by-step feature selection process within the FLASH framework.

Algorithm 2

Pseudo code to perform stepwise feature elimination.

 1: Input: Samples with N features

 2: Output: An optimal set of features Ffinal for maximum accuracy

 3: S ← Featuresinitial

 4: Initialize Fmaster ← {}

 5: fori = Num_Features to 1 do

 6: Build predictive model on S

 7: Srank ← Sorted features based on current model’s coefficient values from high to low

 8: fleast ← Srank[i]

 9: Snew ← S − {fleast}

10: Sacc ← accuracy using cross-validation on Snew

11: Store (Snew, Sacc) to Fmaster

12: S ← Snew

13: end for

14: Ffinal ← element from Fmaster with Sacc maximum among all sets

Quantifying model effectiveness

The number of positive samples is denoted by Np, and the negative samples by Nn. The precision (P) is defined as the proportion of predicted positives that were actually correct, while recall (R) is defined as the proportion of actual positive samples that were correctly identified55. The formulas for P and R are P=TPTP+FP and R=TPTP+FN, where TP represents the number of correctly predicted positive samples, and TN represents the number of correctly predicted negative samples. The overall prediction accuracy, denoted as Acc, is calculated as TP+TNNp+Nn, and mACC is the maximum accuracy achieved by any ML algorithm, similar to the approach used in previous studies56–58. To estimate a more robust and balanced predictive performance, we also used the F1-Score (F1)59, which is defined as the harmonic mean of precision and recall, i.e., F1=2*P*RP+R. We have used 10 times repeated 10-fold cross-validation to obtain the model predictions withf these metrics60,61.

Evaluation of statistical filters and recursive elimination strategy

To evaluate our algorithm, we assessed the performance of five statistical tests: one-way ANOVA25, rank-sum test25, Mann–Whitney U test25, t-test25, and the Brunner–Munzel test26 across 8 datasets, as we can refer to the (Supplementary Table 1). We computed the classification accuracy after feature selection and calculated the proportion of datasets where the accuracy exceeded a 95% threshold. One-way ANOVA test achieved the highest score of 0.6875, exceeding 95% accuracy in the datasets. It was followed by the t-test, which scored 0.63. Mann–Whitney U, Brunner–Munzel, and Ranksum tests recorded lower scores, all falling below 0.59, as seen in the (Supplementary Table 2. Based on this, we used a One-way ANOVA for the subsequent analysis.

To determine the optimal feature pool size, we applied the kneedle algorithm62 across all five statistical tests to identify the “elbow point”. The algorithm identified 6% as the optimal test cutoff (Supplementary Fig. 2). We present the results using the five statistical tests at 6% (Fig. 9A) and ANOVA tests in (Fig. 9B) demonstrate a 6% cutoff yielded across the eight datasets. To evaluate the robustness of our cutoff, we calculated the mean accuracy for different p-values (<0.05, and <0.001) (Supplementary Fig. 3). The mean accuracy for features with p-values <0.001 was lower than that observed for other significance thresholds and was excluded from further analysis. We found that the mean accuracies for p-value thresholds of <0.05 and <0.01 were comparable, with both showing maximum accuracy of ~6%. We adopted a p-value cutoff of <0.01 for subsequent accuracy calculations to ensure the selection of robust features while maintaining statistical relevance.

Fig. 9. Parameter calibration in selected datasets.

Fig. 9

A Comparison of different tests at 6% obtained using mAcc. B Comparison of different cut-off percentages obtained from Fscore using mAcc. C Comparison of different elimination steps for RFECV using mAcc.

We experimented with feature elimination step sizes ranging from 1 to 5 in the RFECV process. A step size of 1 showed the most stable and optimal performance across all eight datasets (Fig. 9C). These findings lead to our final parameter selection for all subsequent experiments.

Supplementary information

Acknowledgements

We thank Dr. Anna Gambin from the University of Warsaw, Poland, for her valuable feedback and insightful comments on this work. S.K.'s research is supported by the THSTI PhD fellowship; A.A.'s research is supported by a G. N. Ramachandran Fellowship from the Dept. of Biotechnology, Govt. of India.

Author contributions

S.K.: Conceptualization, Data curation, Formal analysis, Methodology, Visualization, Writing original draft. A.A.: Formal analysis, Methodology, Visualization, Data curation, Writing original draft. S.C.: Conceptualization, Methodology, Writing review and editing, Supervision, funding.

Data availability

The datasets analyzed in this study are publicly available and have associated GEO identifiers and references. These can be accessed through the GEO database. We have provided access to datasets not associated with a GEO identifier in the GitHub repository https://github.com/samrat-lab/FLASH/.

Code availability

The scripts used in this work are available at https://github.com/samrat-lab/FLASH/.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Supplementary information

The online version contains supplementary material available at 10.1038/s41540-025-00614-x.

References

  • 1.Kreitmaier, P., Katsoula, G. & Zeggini, E. Insights from multi-omics integration in complex disease primary tissues. Trends Genet.39, 46–58 (2023). [DOI] [PubMed] [Google Scholar]
  • 2.Li, Y. & Ning, K. Biomedical applications: The need for multi-omics. In Methodologies of Multi-Omics Data Integration and Data Mining: Techniques and Applications, 13–31 (Springer, 2023).
  • 3.Yang, L., Yang, Y., Huang, L., Cui, X. & Liu, Y. From single-to multi-omics: future research trends in medicinal plants. Brief. Bioinforma.24, bbac485 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Brooks, T. G., Lahens, N. F., Mrčela, A. & Grant, G. R. Challenges and best practices in omics benchmarking. Nat. Rev. Genet.25, 326–339 (2024). [DOI] [PubMed] [Google Scholar]
  • 5.Neagu, A.-N. et al. Omics-based investigations of breast cancer. Molecules28, 4768 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Liao, J. G. & Chin, K.-V. Logistic regression for disease classification using microarray data: model selection in a large p and small n case. Bioinformatics23, 1945–1951 (2007). [DOI] [PubMed] [Google Scholar]
  • 7.Kumar Myakalwar, A. et al. Less is more: Avoiding the LIBS dimensionality curse through judicious feature selection for explosive detection. Sci. Rep.5, 13169 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Liu, H. et al. Evolving feature selection. IEEE Intell. Syst. 20, 64–76 (2005). [Google Scholar]
  • 9.Chen, Y., Gu, Y., Hu, Z. & Sun, X. Sample-specific perturbation of gene interactions identifies breast cancer subtypes. Brief. Bioinform.22, bbaa268 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Buus, R. et al. Molecular drivers of onco DX, prosigna, EndoPredict, and the breast cancer index: A TransATAC study. J. Clin. Oncol.39, 126–135 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Curigliano, G. et al. Incorporating clinicopathological and molecular risk prediction tools to improve outcomes in early hr+/her2–breast cancer. NPJ Breast Cancer9, 56 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Lim, C. X. et al. Healthcare professionals’ and consumers’ knowledge, attitudes, perspectives, and education needs in oncology pharmacogenomics: A systematic review. Clin. Transl. Sci.16, 2467–2482 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Krystel-Whittemore, M., Tan, P. H. & Wen, H. Y. Predictive and prognostic biomarkers in breast tumours. Pathology56, 186–191 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.MotieGhader, H., Masoudi-Sobhanzadeh, Y., Ashtiani, S. H. & Masoudi-Nejad, A. mRNA and microRNA selection for breast cancer molecular subtype stratification using meta-heuristic based algorithms. Genomics112, 3207–3217 (2020). [DOI] [PubMed] [Google Scholar]
  • 15.Bommert, A., Welchowski, T., Schmid, M. & Rahnenführer, J. Benchmark of filter methods for feature selection in high-dimensional gene expression survival data. Brief. Bioinform23, bbab354 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Jović, A., Brkić, K. & Bogunović, N. A review of feature selection methods with applications. In 38th international convention on information and communication technology, electronics and microelectronics (MIPRO), 1200–1205 (2015).
  • 17.Pirgazi, J., Alimoradi, M., Esmaeili Abharian, T. & Olyaee, M. H. An efficient hybrid filter-wrapper metaheuristic-based gene selection method for high dimensional datasets. Sci. Rep.9, 18580 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Kundu, R., Chattopadhyay, S., Cuevas, E. & Sarkar, R. AltWOA: Altruistic whale optimization algorithm for feature selection on microarray datasets. Comput. Biol. Med.144, 105349 (2022). [DOI] [PubMed] [Google Scholar]
  • 19.Wang, A., Liu, H., Yang, J. & Chen, G. Ensemble feature selection for stable biomarker identification and cancer classification from microarray expression data. Comput. Biol. Med.142, 105208 (2022). [DOI] [PubMed] [Google Scholar]
  • 20.Qu, C. et al. Improving feature selection performance for classification of gene expression data using harris hawks optimizer with variable neighborhood learning. Brief. Bioinform22, bbab097 (2021). [DOI] [PubMed] [Google Scholar]
  • 21.Pashaei, E. Mutation-based binary aquila optimizer for gene selection in cancer classification. Comput. Biol. Chem.101, 107767 (2022). [DOI] [PubMed] [Google Scholar]
  • 22.Peng, C. et al. MGRFE: Multilayer recursive feature elimination based on an embedded genetic algorithm for cancer classification. IEEE ACM Trans. Comput. Biol. Bioinform.18, 621–632 (2021). [DOI] [PubMed] [Google Scholar]
  • 23.Gao, S. et al. RIFS2D: A two-dimensional version of a randomly restarted incremental feature selection algorithm with an application for detecting low-ranked biomarkers. Comput. Biol. Med.133, 104405 (2021). [DOI] [PubMed] [Google Scholar]
  • 24.Pudjihartono, N., Fadason, T., Kempa-Liehr, A. W. & O’Sullivan, J. M. A review of feature selection methods for machine learning-based disease risk prediction. Front. Bioinforma.2, 927312 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Hazra, A. & Gogtay, N. Biostatistics series module 3: comparing groups: numerical variables. Indian J. Dermatol.61, 251–260 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Brunner, E. & Munzel, U. The nonparametric behrens-fisher problem: asymptotic theory and a small-sample approximation. Biometrical J.42, 17–25 (2000). [Google Scholar]
  • 27.Ahmed, S. K. How to choose a sampling technique and determine sample size for research: A simplified guide for researchers. Oral. Oncol. Rep.12, 100662 (2024). [Google Scholar]
  • 28.Lohr, S. L. Sampling: Design and Analysis (Chapman and Hall/CRC, 2021).
  • 29.Mangal, A. & Holm, E. A. A comparative study of feature selection methods for stress hotspot classification in materials. Integrating Mater. Manuf. Innov.7, 87–95 (2018). [Google Scholar]
  • 30.Danasingh, A. A. G. S., Subramanian, Aa. B. & Epiphany, J. L. Identifying redundant features using unsupervised learning for high-dimensional data. SN Appl. Sci.2, 1367 (2020). [Google Scholar]
  • 31.Lü, X., Meng, L., Chen, C. & Wang, P. Fuzzy removing redundancy restricted boltzmann machine: Improving learning speed and classification accuracy. IEEE Trans. Fuzzy Syst.28, 2495–2509 (2019). [Google Scholar]
  • 32.Zhang, B. & Cao, P. Classification of high dimensional biomedical data based on feature selection using redundant removal. PloS one14, e0214406 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Ding, C. & Peng, H. Minimum redundancy feature selection from microarray gene expression data. J. Bioinforma. Comput. Biol.3, 185–205 (2005). [DOI] [PubMed] [Google Scholar]
  • 34.Kraskov, A., Stögbauer, H. & Grassberger, P. Estimating mutual information. Phys. Rev. E-Stat. Nonlinear Soft Matter Phys.69, 066138 (2004). [DOI] [PubMed] [Google Scholar]
  • 35.Han, Y., Huang, L. & Zhou, F. A dynamic recursive feature elimination framework (dRFE) to further refine a set of OMIC biomarkers. Bioinformatics37, 2183–2189 (2021). [DOI] [PubMed] [Google Scholar]
  • 36.Marjit, S., Bhattacharyya, T., Chatterjee, B. & Sarkar, R. Simulated annealing aided genetic algorithm for gene selection from microarray data. Comput. Biol. Med.158, 106854 (2023). [DOI] [PubMed] [Google Scholar]
  • 37.Zou, H. & Hastie, T. Regularization and variable selection via the elastic net. J. R. Stat. Soc. Ser. B Stat. Methodol.67, 301–320 (2005). [Google Scholar]
  • 38.Arboretti, R., Barzizza, E., Biasetton, N. & Disegna, M. A review of multivariate permutation tests: Findings and trends. J. Multivariate Anal207, 105421 (2025). [Google Scholar]
  • 39.Cohen, J. Statistical Power Analysis for the Behavioral Sciences (Routledge, 2013).
  • 40.Eswaran, J. et al. Transcriptomic landscape of breast cancers through mrna sequencing. Sci. Rep.2, 264 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Horvath, A. et al. Novel insights into breast cancer genetic variance through rna sequencing. Sci. Rep.3, 2256 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Kretschmer, C., Conradi, A., Kemmner, W. & Sterner-Kock, A. Latent transforming growth factor binding protein 4 (ltbp4) is downregulated in mouse and human dcis and mammary carcinomas. Cell. Oncol.34, 419–434 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Kretschmer, C. et al. Identification of early molecular markers for breast cancer. Mol. cancer10, 15 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Piñero, J. et al. The disgenet knowledge platform for disease genomics: 2019 update. Nucleic Acids Res.48, D845–D855 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Haan, J. C. et al. Mammaprint and blueprint comprehensively capture the cancer hallmarks in early-stage breast cancer patients. Genes Chromosomes Cancer61, 148–160 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Jairath, N. K. et al. A systematic review of the evidence for the decipher genomic classifier in prostate cancer. Eur. Urol.79, 374–383 (2021). [DOI] [PubMed] [Google Scholar]
  • 47.Koc, M. A. et al. Molecular and translational biology of the blood-based veristrat® proteomic test used in cancer immunotherapy treatment guidance. J. Mass Spectrom. Adv. Clin. lab30, 51–60 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Misra, P. & Yadav, A. S. Improving the classification accuracy using recursive feature elimination with cross-validation. Int. J. Emerg. Technol.11, 659–665 (2020). [Google Scholar]
  • 49.Chan, J. Y.-L. et al. A correlation-embedded attention module to mitigate multicollinearity: An algorithmic trading application. Mathematics10, 1231 (2022). [Google Scholar]
  • 50.Atenafu, E. G., Hamid, J. S., Stephens, D., To, T. & Beyene, J. A small p-value from an observed data is not evidence of adequate power for future similar-sized studies: A cautionary note. Contemp. Clin. trials30, 155–157 (2009). [DOI] [PubMed] [Google Scholar]
  • 51.Efron, B. & Tibshirani, R. J. An Introduction to the Bootstrap (Chapman and Hall/CRC, 1994).
  • 52.Bui, P. H. D., Nguyen, L. Y. B., Ngo, L. D. & Nguyen, H. T. T-test-based feature selection on dna microarrays gene expression data for leukemia classification. In International Conference on Industrial, Engineering and Other Applications of Applied Intelligent Systems, 207–218 (Springer, 2025).
  • 53.Koul, N. & Manvi, S. S. Feature selection from gene expression data using simulated annealing and partial least squares regression coefficients. Glob. Transit. Proc.3, 251–256 (2022). [Google Scholar]
  • 54.Rotimi, S. O. et al. Gene expression profiling analysis reveals putative phytochemotherapeutic target for castration-resistant prostate cancer. Front. Oncol.9, 714 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Van Rijsbergen, C. J. Foundation of evaluation. J. Documentation30, 365–373 (1974). [Google Scholar]
  • 56.Clifford, G. D. et al. Recent advances in heart sound analysis. Physiological Meas.38, E10–E25 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Ren, Y. et al. Gender specificity improves the early-stage detection of clear cell renal cell carcinoma based on methylomic biomarkers. Biomark. Med.12, 607–618 (2018). [DOI] [PubMed] [Google Scholar]
  • 58.Guo, D., Li, J., Jiang, S.-H., Li, X. & Chen, Z. Intelligent assistant driving method for tunnel boring machine based on big data. Acta Geotechnica17, 1019–1030 (2022). [Google Scholar]
  • 59.Grandini, M., Bagli, E. & Visani, G. Metrics for multi-class classification: an overview. Preprint at https://arxiv.org/abs/2008.05756 (2020).
  • 60.Conti Bellocchi, M. C. et al. Development and validation of a risk score for prediction of clinical success after duodenal stenting for malignant gastric outlet obstruction. Expert Rev. Gastroenterol. Hepatol.16, 393–399 (2022). [DOI] [PubMed] [Google Scholar]
  • 61.Moore, J. H. & Williams, S. M. New strategies for identifying gene-gene interactions in hypertension. Ann. Med.34, 88–95 (2002). [DOI] [PubMed] [Google Scholar]
  • 62.Satopaa, V., Albrecht, J., Irwin, D. & Raghavan, B. Finding a” kneedle” in a haystack: Detecting knee points in system behavior. In 2011 31st International Conference on Distributed Computing Systems Workshops, 166–171 (IEEE, 2011).
  • 63.Chiaretti, S. et al. Gene expression profile of adult t-cell acute lymphocytic leukemia identifies distinct subsets of patients with different response to therapy and survival. Blood103, 2771–2778 (2004). [DOI] [PubMed] [Google Scholar]
  • 64.Dabba, A., Tari, A., Meftali, S. & Mokhtari, R. Gene selection and classification of microarray data method based on mutual information and moth flame algorithm. Expert Syst. Appl.166, 114012 (2021). [Google Scholar]
  • 65.Pomeroy, S. L. et al. Prediction of central nervous system embryonal tumour outcome based on gene expression. Nature415, 436–442 (2002). [DOI] [PubMed] [Google Scholar]
  • 66.Alon, U. et al. Broad patterns of gene expression revealed by clustering analysis of tumor and normal colon tissues probed by oligonucleotide arrays. Proc. Natl. Acad. Sci. USA96, 6745–6750 (1999). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Gravier, E. et al. A prognostic DNA signature for T1T2 node-negative breast cancer patients. Genes Chromosomes Cancer49, 1125–1134 (2010). [DOI] [PubMed] [Google Scholar]
  • 68.Alter, M. D. et al. Autism and increased paternal age related changes in global levels of gene expression regulation. PLoS One6, e16715 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Rousseaux, S. et al. Ectopic activation of germline and placental genes identifies aggressive metastasis-prone lung cancers. Sci. Transl. Med.5, 186ra66 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Levy, H. et al. Transcriptional signatures as a disease-specific and predictive inflammatory biomarker for type 1 diabetes. Genes Immun.13, 593–604 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Tian, E. et al. The role of the wnt-signaling antagonist DKK1 in the development of osteolytic lesions in multiple myeloma. N. Engl. J. Med.349, 2483–2494 (2003). [DOI] [PubMed] [Google Scholar]
  • 72.Shamir, R. et al. Analysis of blood-based gene expression in idiopathic parkinson disease. Neurology89, 1676–1683 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Sun, L. et al. Neuronal and glioma-derived stem cell factor induces angiogenesis within the brain. Cancer Cell9, 287–300 (2006). [DOI] [PubMed] [Google Scholar]
  • 74.Putluri, N. et al. Pathway-centric integrative analysis identifies rrm2 as a prognostic marker in breast cancer associated with poor survival and tamoxifen resistance. Neoplasia16, 390–402 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Tang, W. et al. Correction: Integrated proteotranscriptomics of breast cancer reveals globally increased protein-mrna concordance associated with subtypes and survival. Genome Med.17, 69 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Terunuma, A. et al. Myc-driven accumulation of 2-hydroxyglutarate is associated with breast cancer prognosis. J. Clin. Investig.124, 398–412 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Juul, N. et al. Assessment of an rna interference screen-derived mitotic and ceramide pathway metagene as a predictor of response to neoadjuvant paclitaxel for primary triple-negative breast cancer: a retrospective analysis of five clinical trials. lancet Oncol.11, 358–365 (2010). [DOI] [PubMed] [Google Scholar]
  • 78.Li, Y. et al. Amplification of laptm4b and ywhaz contributes to chemotherapy resistance and recurrence of breast cancer. Nat. Med.16, 214–218 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Silver, D. P. et al. Efficacy of neoadjuvant cisplatin in triple-negative breast cancer. J. Clin. Oncol.28, 1145–1153 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Turashvili, G. et al. Novel markers for differentiation of lobular and ductal invasive breast carcinomas by laser microdissection and microarray analysis. BMC cancer7, 55 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Li, S.-Y. et al. Tumor circadian clock strength influences metastatic potential and predicts patient prognosis in luminal a breast cancer. Proc. Natl. Acad. Sci.121, e2311854121 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Seo, J.-S. et al. The transcriptional landscape and mutational profile of lung adenocarcinoma. Genome Res.22, 2109–2119 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Martuscello, R. T. et al. Gene expression analysis of the cerebellar cortex in essential tremor. Neurosci. Lett.721, 134540 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

The datasets analyzed in this study are publicly available and have associated GEO identifiers and references. These can be accessed through the GEO database. We have provided access to datasets not associated with a GEO identifier in the GitHub repository https://github.com/samrat-lab/FLASH/.

The scripts used in this work are available at https://github.com/samrat-lab/FLASH/.


Articles from NPJ Systems Biology and Applications are provided here courtesy of Nature Publishing Group

RESOURCES