Skip to main content
. 2026 Jun 19;27(12):5552. doi: 10.3390/ijms27125552
Algorithm 1. Cross-validation, Cross-MAS feature selection, and held-out classification workflow
Input: Dengue expression matrix and metadata; chikungunya expression matrix and metadata; common protein-coding gene set shared between the two datasets; number of folds K = 3; candidate numbers of selected genes per Cross-MAS category k ∈ {1, 2, 3, 4, 5}.
Step 1. Define stratified folds. Independently partition the dengue and chikungunya datasets into three folds. For dengue, assign one severe dengue case to each fold because only three severe cases were available. For classification, group dengue with warning signs and severe dengue samples into a single dengue class. For chikungunya, distribute healthy and chikungunya samples across the three folds as evenly as possible.
Step 2. Define cross-validation iterations. The three cross-validation iterations are:
  • Fold12: folds 1 and 2 for training; fold 3 for held-out testing.

  • Fold13: folds 1 and 3 for training; fold 2 for held-out testing.

  • Fold23: folds 2 and 3 for training; fold 1 for held-out testing.

Step 3. Initialize storage. Initialize empty objects to store fold-specific ranked genes, selected feature panels, training predictions, held-out test predictions, and fold-level performance metrics.
Step 4. Run cross-validation. For each cross-validation iteration i ∈ {Fold12, Fold13, Fold23}:
4.1. Define training and held-out test subsets. Use the two folds specified by iteration i as the training subset and the remaining fold as the held-out test subset, separately for dengue and chikungunya.
4.2. Perform dataset-specific differential expression and MAS ranking before cross-dataset integration. At this stage, the dengue and chikungunya datasets remain separate with no batch-effect correction.
  • Within the dengue training subset, compare dengue samples with dengue healthy controls using GLMQL-MAS.

  • Within the chikungunya training subset, compare chikungunya samples with chikungunya healthy controls using GLMQL-MAS.

  • Rank significant upregulated and downregulated genes in each dataset using the MAS score defined in Equation (1).

4.3. Apply Cross-MAS feature integration within matched training partitions.
  • Match the dengue and chikungunya training partitions corresponding to iteration i.

  • Apply the Cross-MAS ranking function defined in Equation (2) separately to the upregulated and downregulated ranked gene lists.

  • Partition genes into six categories: dengue-specific upregulated, dengue-specific downregulated, chikungunya-specific upregulated, chikungunya-specific downregulated, shared upregulated, and shared downregulated.

4.4. Select fold-specific feature panels.
  • For each candidate value k ∈ {1, 2, 3, 4, 5}, select the top k genes from each of the six Cross-MAS categories.

  • Combine these genes to form the selected feature panel for iteration i and candidate k.

  • Store the selected genes for iteration i and candidate k.

4.5. Apply batch correction only for downstream PCA and classification.
  • After feature selection is complete, align the matched dengue and chikungunya training subsets to form the fold-specific training matrix for classification.

  • Align the matched dengue and chikungunya held-out subsets to form the fold-specific held-out test matrix.

  • Apply batch-effect correction only at this downstream modeling stage for PCA visualization and multinomial logistic regression.

4.6. Fit PCA using training data only.
  • Subset the batch-corrected training expression matrix to the selected genes.

  • Fit principal component analysis using only the training samples.

  • Extract PC1 and PC2 from the training samples.

4.7. Train the classifier using training data only.
  • Train a multinomial logistic regression model using training PC1 and PC2 as predictors.

  • Class labels are healthy, dengue, and chikungunya.

4.8. Evaluate the held-out test subset.
  • Subset the held-out test expression matrix to the same selected genes.

  • Project held-out test samples onto the PCA axes learned from the training subset.

  • Use the trained multinomial logistic regression model to predict held-out test labels.

4.9. Store fold-level outputs. Store selected genes, training predictions, held-out test predictions, balanced accuracy, macro F1, and confusion matrix results for iteration i and candidate k.
Step 5. Aggregate outputs across folds. After all three cross-validation iterations are completed, aggregate the stored outputs across i ∈ {Fold12, Fold13, Fold23} for each candidate value of k. For each k, combine training predictions across folds and combine held-out test predictions across folds. Calculate aggregated balanced accuracy, macro F1, and confusion matrices for training and held-out test performance. These aggregated results are used to summarize model performance across candidate feature-panel sizes.
Output: Fold-specific ranked gene groups, selected feature panels, training and held-out test predictions, and aggregated classification performance metrics.
End