Skip to main content
Nature Communications logoLink to Nature Communications
. 2026 Aug 11;17:9644. doi: 10.1038/s41467-026-76429-z

CAKR: commutative algebra k-mer representations for genomics

Faisal Suwayyid 1,2, Yuta Hozumi 2,6, Mushal Zia 2, JunJie Wee 2, Hongsong Feng 3, Guo-Wei Wei 2,4,5,✉
PMCID: PMC13554176  PMID: 42711309

Abstract

Despite the availability of various sequence analysis models, comparative genomic analysis remains a challenge in genomics, genetics, and phylogenetics. Commutative algebra, a fundamental tool in algebraic geometry and number theory, has rarely been used in data and biological sciences. In this study, we introduce commutative algebra k-mer representations as a nonlinear algebraic framework for analyzing genomic sequences. This representation bridges commutative algebra, algebraic topology, combinatorics, and machine learning to establish a mathematical framework for comparative genomic analysis. We evaluate its effectiveness on three tasks including genetic variant classification, phylogenetic tree reconstruction, and viral classification, typically requiring alignment-based, alignment-free, and machine-learning approaches, respectively. In this work, we show that commutative algebra k-mer representations outperform five state-of-the-art sequence analysis methods across twelve primary datasets, with two additional supplementary fragment-placement benchmarks, especially in viral classification, and maintain relatively stable predictive accuracy as dataset size increases, underscoring scalability and robustness.

Subject terms: Machine learning, Computational models


Comparative genomic analysis remains a fundamental challenge in genomics, genetics and phylogenetics, motivating the development of new mathematical frameworks for sequence comparison. The authors introduce a commutative algebra k-mer representation that integrates concepts from algebra, topology, combinatorics and machine learning, and evaluate its application to genetic variant classification, phylogenetic tree reconstruction and viral classification.

Introduction

Comparative genomics examines genetic variation across species and populations to study evolution, identify functional elements, assess diversity, and reconstruct phylogeny1,2. Comparative analysis of genomic sequences includes phylogenetic inference, functional annotation, and phenotype classification3,4. These analyses require a genome space, a metric space whose points represent genomes and whose distances capture biologically meaningful similarities. An effective metric should reflect structural, functional, and evolutionary relationships, enabling robust comparison and downstream analyses.

Traditional approaches rely on alignment-based methods that identify substitutions, insertions, and deletions through global or local optimization. Tools such as Clustal Omega5, MAFFT6, and MUSCLE7 are effective for closely related sequences8. However, their computational cost scales poorly with sequence length and dataset size, and their accuracy deteriorates for highly divergent sequences, limiting their applicability in phylogenetics7.

Alignment-free methods address these limitations by mapping sequences to fixed-length vectors9, enabling scalability10,11 and whole-genome analysis12,13. Frequency-based approaches use k-mer count vectors14, while others employ entropy15, Lempel-Ziv16, or Kolmogorov complexity17 measures. Model-based alignment-free approaches, such as CAFE18, further derive explicit evolutionary distances from k-mer frequency statistics under substitution assumptions. Although efficient, many neglect positional or structural information, limiting their performance in genetic variant analysis10.

Recent advances aim to enrich such feature representations. The natural vector method (NVM) encodes statistical moments of (k)-mer positions19,20. Chaos game representation (CGR) maps sequences into fractal images21,22, while Fourier power spectrum (FPS) analysis extracts dominant periodicities23,24. Multi-scale integration has also been achieved via fuzzy integrals25,26. Persistent challenges include sensitivity to parameter choices, such as (k), weights, and dimensions27,28, and a limited capacity to detect biologically significant variants. These challenges motivate novel computational approaches to genomics.

Commutative algebra, the study of commutative rings, ideals, and modules, underpins key areas of modern mathematics, including algebraic geometry, number theory, and homological algebra. Despite its foundational role in pure mathematics, it has hardly been applied in science and technology due to its abstractness and lack of metric. Recently, Suwayyid and Wei introduced multi-scale analysis to commutative algebra, enabling the potential application of abstract nonlinear algebra to data science and learning 29.

In this work, we introduce commutative algebra to genomics. By leveraging persistent Stanley–Reisner theory (PSRT)29, we propose commutative algebra k-mer representations (CAKR) to integrate k-mer representations of sequences30 with persistence modules arising from Stanley–Reisner constructions. CAKR is evaluated on twelve primary datasets: one for genetic variant classification, seven for phylogenetic tree reconstruction, and four for viral classification derived from the National Center for Biotechnology Information (NCBI) Virus database. In addition, we include two supplementary fragment-placement benchmarks for barcode-based influenza query matching and comparison with SEPP.

We systematically compare CAKR against five state-of-the-art alignment-free methods: the natural vector method (NVM)31, the Markov k-string model (MKS)32, Feature Frequency Profile with Jensen–Shannon (FFP-JS) divergence and with Kullback–Leibler (FFP-KL) divergence 28,33, and Fourier power spectrum (FPS)23, and an alignment-based method, MAFFT6. Section “Results” presents the experimental results. Section “Discussion” discusses performance, limitations, and generalization. Section “Methods” describes the proposed CAKR methodology and introduces a new purity metric for evaluating the quality of phylogenetic tree reconstruction.

Results

An overview of CAKR

To evaluate the effectiveness of the proposed CAKR method, illustrated by the workflow in Fig. 1, we consider three key applications: genetic variant classification, phylogenetic tree reconstruction, and viral classification. Genetic variant classification is an important application in genetics and bioinformatics. It tracks genetic variations in selected gene lists associated with specific diseases, phenotypes, and populations, for which alignment methods are typically favored. Additionally, phylogenetic tree reconstruction of genetic sequences plays a fundamental role in elucidating evolutionary relationships both among and within species, where alignment-free methods have advantages3,4. Finally, viral classification is a general machine learning approach for the genetic analysis and prediction of unknown viral sequences. It is challenging to design a unified approach for these diverse tasks.

Fig. 1. Illustration of the CAKR workflow.

Fig. 1

Given a query sequence, k-mers are first extracted. For each k-mer, the set of its occurrence positions within the sequence is treated as a sequence of integers. The persistent facet numbers associated with these sequences of integers are then computed and used to represent the corresponding k-mer. The feature vectors of all k-mers of the same size are concatenated to construct an algebraic representation. Pairwise distances between these sequences are subsequently defined and used for tasks such as genome variant classification, phylogenetic tree reconstruction, and genome classification. SARS-CoV-2 artwork adapted from NIAID NIH BIOART Source BIOART-00046549 (public domain; courtesy of NIAID); figure assembled by the authors using Inkscape. CAKR, commutative algebra k-mer representation; UPGMA, unweighted pair group method with arithmetic mean; K-NN, K-nearest neighbors. The symbol r denotes the filtration radius. Nucleotide colors denote A (teal), C (purple), G (blue), and T (yellow); the remaining colors distinguish schematic components and carry no quantitative meaning.

Genetic variant classification

The emergence of SARS-CoV-2 variants during the COVID-19 pandemic posed critical challenges for monitoring viral evolution, guiding public health measures, and informing vaccine and therapeutic design34. Variant differences can be subtle, often involving only a few mutations in the spike protein receptor-binding domain35, despite genome lengths of approximately 29.9 kb. For the genetic variant classification task, we analyzed 44 complete SARS-CoV-2 genomes from GISAID30,36,37. This dataset follows the curated benchmark of Li et al.36, where representative complete genomes were selected to cover major variant lineages (alpha, beta, gamma, delta, lambda, Mu, GH/490R, and Omicron) while maintaining a compact evaluation set. Each genome is labeled according to its assigned lineage, and branches are annotated and color-coded accordingly in our phylogenetic trees. We employed CAKR with k = 6, as determined by our data-driven scale selection framework, which identifies an optimal k-mer size through a weighted consensus of coverage, information-sparsity, and stability-adjusted entropy criteria applied to the dataset. Because SARS-CoV-2 genomes are highly similar across lineages, we use a slightly larger k to capture subtle variant-defining patterns while remaining computationally practical. NVM was run with k = 5, while FFP-KL, FFP-JS, and MKS used k = 3; FPS required no k-mer parameter. Among six alignment-free methods, only CAKR achieved perfect clustering, producing well-defined monophyletic groups (Fig. 2a and Supplementary Fig. 1) and accurately resolving inter-variant relationships. FFP-JS, FFP-KL, and MKS attained moderate accuracy but mis-clustered some delta, gamma, GH/490R, and Omicron samples, whereas NVM and FPS failed to produce meaningful phylogenetic structure. Performance was quantified using the label-based purity metric given by Eq. 26, which measures the proportion of samples with identical labels grouped under a common ancestor. Overall, CAKR achieved the highest purity score among all methods (Fig. 2d).

Fig. 2. Illustration and comparison of the proposed CAKR model with other methods in variant classification and phylogenetic tree reconstruction.

Fig. 2

a CAKR classification of SARS-CoV-2 variants. b, c CAKR phylogenetic tree reconstructions of the mammalian mitochondrial genome and HRV datasets. d–f Comparison of prediction accuracies of different methods on the SARS-CoV-2, mammalian mitochondrial, and HRV datasets, respectively. CAKR commutative algebra k-mer representation, NVM natural vector method, FFP-JS feature frequency profile with Jensen–Shannon divergence, FFP-KL feature frequency profile with Kullback–Leibler divergence, FPS Fourier power spectrum, MKS Markov k-string model, HRV human rhinovirus. Branch and label colors in (a–c) denote the known variant or taxonomic groups shown by the corresponding labels. In d–f, the darker bar indicates the highest-performing method, whereas the paler bars indicate the remaining methods. The methods compared in (d–f) are identified beneath the corresponding bars. Source data are provided as a Source Data file.

Phylogenetic tree reconstruction

Phylogenetic reconstruction is central to elucidating evolutionary relationships among taxa3,4. We assessed CAKR on six benchmark datasets from ref. 30, encompassing complete genomes and gene sequences with established taxonomic labels. Sequence lengths range from ~2000 nt for influenza HA genes to ~17,000 nt for mammalian mitochondrial and Ebola virus genomes, and up to several megabases for bacterial genomes.

Phylogenetic trees were inferred using CAKR and compared with those from representative alignment-free methods. Following30, we set k = 3 for FFP-KL, FFP-JS, and MKS models, k = 5 for NVM, and selected k = 4 for CAKR using our data-driven k-mer scale selection framework; FPS required no k-mer parameter. All trees were built with the UPGMA algorithm and visualized in Interactive Tree Of Life (iTOL) v638. As a supplementary sensitivity analysis, we also compared CAKR trees built with UPGMA and Neighbor-joining (NJ). UPGMA achieved perfect purity on the rhinovirus, mammalian, and SARS-CoV-2 datasets, whereas NJ reduced purity for rhinovirus and mammalian data while matching UPGMA on SARS-CoV-2 (Supplementary Section 7 and Supplementary Fig. 15).

In the mammalian mitochondrial dataset, the goal was to assess how accurately each method recovers clades consistent with known host species classifications. CAKR achieved perfect concordance, producing monophyletic clades for all major mammalian orders, including delineated Primates, Cetacea, and Artiodactyla. It also preserved internal coherence within Carnivora and Perissodactyla (Fig. 2b and Supplementary Fig. 3). Furthermore, CAKR attained the highest purity score among all methods (Fig. 2e).

Other methods showed varying levels of discordance: FPS partially recovered major orders but split Artiodactyla; NVM misplaced several taxa and fragmented Carnivora; FFP-JS and FFP-KL preserved many lineages but mis-grouped smaller orders and split Carnivora; and MKS produced the weakest resolution, fragmenting multiple orders entirely.

In the HRV dataset, CAKR separates HRV-B and HRV-C into coherent clusters and places HEV as an outgroup (Fig. 2c), while HRV-A shows less consistent clustering, with a small number of HRV-A taxa occurring near the HRV-B region of the tree. On the other hand, FPS achieved perfect classification, cleanly separating HRV-A, HRV-B, and HRV-C clades, with the outgroup (HEV) distinctly isolated (Supplementary Fig. 4). NVM, FFP-JS, and FFP-KL produced moderate results, partially recovering the three clades but misplacing several HRV-A genomes within HRV-B subtrees, indicating reduced sensitivity to closely related subtypes. MKS performed poorest, failing to recover a coherent subgroup structure and producing extensive intermixing among clades, along with poor outgroup resolution.

In the HEV dataset, CAKR, FFP-JS, FFP-KL, and NVM all achieved perfect genotypic clustering (Supplementary Fig. 5). MKS misclassified one Group 3 genome, while FPS performed worst, failing to separate Groups 3 and 4, though correctly clustering Group 1.

In the influenza HA dataset, CAKR, FFP-JS, and FFP-KL all achieved perfect subtype classification (Supplementary Fig. 6), though FFP-JS and FFP-KL grouped H3N2 and H2N2 under a shared node, unlike CAKR. NVM failed to cluster H1N1 cohesively and misclassified one H2N2 sequence, while MKS misclassified an H1N1. FPS performed worst, correctly clustering only the H7N9 subtype.

In the ebolavirus dataset, all methods correctly separated the five known species (Supplementary Fig. 7). NVM, FFP-JS, FFP-KL, and MKS distinguished epidemic lineages except for the 1996 outbreak, with NVM producing the longest inter-clade branches. FFP-JS and FFP-KL showed shorter branches, indicating weaker sensitivity, while MKS grouped EBOV and RESTV under a common ancestor. FPS failed to resolve the epidemic-level structure within EBOV.

To evaluate performance on longer sequences, we applied CAKR to 30 complete bacterial genomes ranging from 0.9 to 6.5 Mb. Despite the substantial increase in sequence size and complexity, all methods except NVM and FPS recovered correct family-level groupings (Supplementary Fig. 8). Overall, CAKR consistently delivered accurate and biologically coherent phylogenies, outperforming existing alignment-free methods across diverse genomic datasets. A qualitative cross-dataset summary of phylogenetic reconstruction behavior is provided in Supplementary Table 2.

To quantify phylogenetic consistency throughout our experiments, we use the purity metric in Eq. 26 as our primary evaluation measure (Fig. 2). Purity summarizes the extent to which sequences sharing the same taxonomic label form coherent (approximately monophyletic) groups in the inferred tree, providing a label-based assessment aligned with known classifications.

In addition, we report the normalized Robinson–Foulds (nRF) distance between CAKR trees and the MAFFT-based reference trees as a complementary, topology-only measure of tree-to-tree similarity. Averaged across datasets, CAKR attained an nRF of 0.39. Further details are provided in Supplementary Section 3.

To further assess discrimination in a narrow-clade setting, we applied CAKR and the competing baselines to two closely related Salmonella datasets. In a focused dataset of 294 genomes, comprising S. enterica subspecies arizonae (n = 35), diarizonae (n = 17), enterica (n = 194), and salamae (n = 48), CAKR achieved the highest average purity (0.8885), exceeding MKS (0.8831), FFP-KL and FFP-JS (both 0.7417), FPS (0.6920), and NVM (0.5938). We then expanded the analysis to a broader dataset of 519 closely related Salmonella genomes. On this more extensive narrow-clade benchmark, MKS and FFP remained competitive, while NVM and FPS exhibited substantially reduced resolution. Moreover, CAKR again achieved the highest purity, outperforming MKS and FFP by at least 3.9%.

Viral classification

In this task, we conduct viral classification experiments on the four NCBI datasets described in Supplementary Section 1, following the experimental design outlined by Hozumi et al.30. In particular, we adopt the same comparative benchmarking strategies to ensure consistency. The labels assigned to viral sequences in the NCBI Virus Database are regularly revised, as the International Committee on Taxonomy of Viruses (ICTV) continuously updates viral classifications based on new scientific findings. This ongoing process reflects the complexity of the classification problem and emphasizes the importance of using methods that remain reliable despite changes in biological taxonomy. A summary of the dataset versions, filtering criteria, and sample sizes is provided in Supplementary Table 1. Moreover, many of the NCBI genomes analyzed in this study already include ambiguous nucleotides (e.g., N, R, W, and Y), reflecting uncertainty or partial information that commonly arises in real sequencing/assembly pipelines. The stable performance observed across these datasets therefore provides empirical evidence that CAKR is robust to a moderate level of incompleteness.

Two classification tasks are considered. The first employs a 1-nearest neighbor (1-NN) classifier on feature representations derived from the persistent facet ideals featurization, following the methodology introduced by Sun et al.31. A test sample is deemed correctly classified if its nearest neighbor, under the algebraic distance metric, shares the same viral family label. This protocol models realistic scenarios in which newly sequenced viral genomes are annotated based on proximity to previously characterized reference genomes.

The second task involves a 5-nearest neighbors (5-NN) classifier, using the same experimental protocol described in ref. 30. Performance is assessed via 5-fold cross-validation repeated over 30 random seeds to ensure statistical robustness. To control for class imbalance and maintain classification reliability, the evaluation is restricted to viral families with at least 15 representative sequences. In both tasks, we empirically observed that increasing the value of k in the k-mer algebraic representations generally leads to a decline in model performance. This behavior is expected, as larger values of k result in more distinct representations for each genome, irrespective of their biological classification, thereby reducing the amount of shared structural information that can be effectively used for grouping. The model exhibits consistently strong performance for k = 3, 4, and 5, with k = 4 selected for use in both tasks.

For each dataset, stratified 5-fold cross-validation was conducted using 30 independent random seeds to obtain performance metrics. Classification performance was assessed using accuracy (ACC), balanced accuracy (BA), macro-F1 score (F1), recall, and precision. All metrics were computed using the macro-averaging scheme to ensure equal weight across viral families, regardless of class imbalance.

All methods exhibit a consistent decrease in accuracy from 2020 to 2024 across both classification tasks. As summarized in Table 1, our model exhibits strong predictive performance under the 1-nearest neighbor (1-NN) classification protocol. On the NCBI 2020 dataset, consisting of 6993 samples, the model achieved an accuracy of 0.932. For the larger NCBI 2022 dataset comprising 11,428 samples, an accuracy of 0.920 was obtained. For the NCBI 2024 dataset, we evaluated performance under two settings: for the filtered set of 12,154 samples, an accuracy of 0.891 was obtained, while for the complete set containing 13,645 samples, the model achieved an accuracy of 0.892.

Table 1.

One-nearest-neighbor viral classification accuracy

Data CAKR NVM FFP-JS FFP-KL FPS MKS
NCBI 2020 0.932 0.879 0.862 0.862 0.732 0.734
NCBI 2022 0.920 0.875 0.870 0.870 0.732 0.735
NCBI 2024 0.891 0.829 0.825 0.826 0.656 0.637
NCBI 2024 All 0.892 0.825 0.832 0.832 0.647 0.647

CAKR was compared with NVM, FFP-JS, FFP-KL, FPS, and MKS on four NCBI viral genome datasets. Values report 1-NN classification accuracy. Source data are provided as a Source Data file. Boldface indicates the highest accuracy within each row.

Furthermore, using 5-nearest neighbors classification with 5-fold cross-validation, as summarized in Fig. 3 and Supplementary Table 3, our proposed model demonstrates higher predictive performance compared to existing state-of-the-art approaches across all datasets. Specifically, on the NCBI 2020 dataset, our model achieved an accuracy of 0.913. On the larger NCBI 2022 dataset, the model attained an accuracy of 0.902. On the filtered NCBI 2024 dataset, the model achieved an accuracy of 0.876, while on the full NCBI 2024 dataset containing all samples, an accuracy of 0.876 was recorded. These results underscore the robustness and effectiveness of our method in modeling complex viral genome spaces and establishing reliable predictive frameworks for viral classification tasks.

Fig. 3. Performance of CAKR in 5-nearest-neighbor viral classification.

Fig. 3

CAKR was compared with five alignment-free methods, NVM, FFP-JS, FFP-KL, FPS, and MKS, on four NCBI viral genome datasets: NCBI 2020, NCBI 2022, NCBI 2024, and NCBI 2024 All. Classification was performed using a 5-nearest-neighbor classifier with stratified 5-fold cross-validation repeated over 30 random seeds. Performance was evaluated using accuracy (ACC), balanced accuracy (BA), macro-F1 score, recall, and precision. Macro-averaged scores are reported to account for class imbalance among viral families. CAKR achieved the highest score across all datasets and evaluation metrics. CAKR commutative algebra k-mer representation, NVM natural vector method, FFP-JS feature frequency profile with Jensen–Shannon divergence, FFP-KL feature frequency profile with Kullback–Leibler divergence, FPS Fourier power spectrum, MKS Markov k-string model, NCBI National Center for Biotechnology Information, ACC accuracy, BA balanced accuracy, F1 macro-averaged F1 score. Panel colors distinguish the four dataset versions: blue, gold, green, and pink represent NCBI 2020, NCBI 2022, NCBI 2024, and NCBI 2024 All, respectively. Within each panel, the darker bar in each metric group indicates the highest-performing method, whereas the paler bars indicate the remaining methods. Method names are indicated beside the bars in (a) and apply in the same order to (b–d). Source data are provided as a Source Data file.

As a further assessment of robustness, we quantified the stability of method rankings using Kendall’s coefficient of concordance (W). By treating evaluation metrics as judges within each dataset and datasets as judges within each metric, we consistently observed high agreement (W = 0.965–0.982 and W = 0.964–0.979, respectively; Supplementary Tables 4 and 5). The associated χ2 tests confirmed that this concordance is statistically significant, indicating that the relative ordering of methods is largely stable with respect to both the choice of performance metric and the specific dataset version. In particular, the top-ranked position of CAKR remained unchanged across all metrics and datasets, demonstrating that the observed performance hierarchy is statistically robust rather than metric- or dataset-dependent. For further details, see Supplementary Section 8.

Discussion

Overall performance

Across all benchmark evaluations, CAKR demonstrates consistently strong overall performance relative to the five baseline alignment-free methods. On six phylogenetic tree–construction datasets, it achieves the highest overall tree accuracy, with the purity index Eq. 26 exceeding that of every competitor by at least 4 percentage points on average (absolute difference; Fig. 2). On the four extensive NCBI collections curated in refs. 27,30,31, CAKR attains the top macro-averaged scores for accuracy (ACC), balanced accuracy (BA), F1, recall, and precision, surpassing the runner-up by 4–7 percentage points across all datasets (Fig. 3).

This improvement stems from CAKR’s explicit encoding of the spatial distribution of k-mers through locality-sensitive features, in contrast to most alignment-free methods, which rely primarily on k-mer frequency distributions with limited positional or spatial information. This positional awareness enables CAKR to resolve subtle yet biologically consequential sequence variants, thereby enhancing performance in large-scale viral classification, as well as in phylogenetic tree reconstruction and genetic variant classification.

Comparable to alignment methods for variant classification

Alignment-based tools, such as MAFFT, are highly effective for within-species variant inference, where sequences are sufficiently similar and share reliable positional homology. As divergence increases or when heterogeneous genomic content reduces the stability or interpretability of a single global alignment across taxa, alignment-based comparisons may become less robust; in such remote-homology regimes, profile-HMM approaches (e.g., HMMER) are often used to detect conserved gene/protein-family relationships. To compare CAKR and MAFFT under conditions favorable to alignment methods, we analyzed three benchmark datasets, such as SARS-CoV-2, mammalian mitochondrial genomes, and human rhinovirus (HRV). The trees produced by CAKR and MAFFT show strong agreement: both recover the principal SARS-CoV-2 variant groups (alpha, beta, gamma, delta, lambda, Mu, GH/490R, and Omicron), correctly group Primates, Carnivora, and Cetacea in the mammalian set, and delineate HRV-A, HRV-B, HRV-C with the HEV outgroup (Fig. 2 and Supplementary Fig. 9). Hence, CAKR attains alignment-level agreement in regimes where alignment is presumed strongest for variant inference, while retaining robustness and efficiency for heterogeneous or multi-species data. On the other hand, CAKR can directly compare the similarity of entirely different sequences, such as those of HRV, HIV, and SARS-CoV-2, for which MAFFT, Clustal Omega, and other alignment methods do not work.

Robustness

CAKR maintains high accuracy as the dataset size and quality vary. In a 5-NN setting, its accuracy declines only modestly, from 91.3% on the NCBI 2020 dataset consisting of 6993 sequences to 87.6% on the 13,645-sequence NCBI 2024 All dataset, exhibiting similar stability when error-containing reads are included. A 1-NN evaluation shows the same pattern, with performance decreasing from 93.2% to 89.2% across the same data progression. Empirically, k ∈ {3, 4, 5, 6} suffices for strong results, with k = 4 giving the best single-model accuracy; combining these values in an ensemble further boosts performance with negligible additional cost.

In addition to dataset-scale robustness, Supplementary Section 10 evaluates robustness to incomplete and fragmented sequence inputs. Under simulated truncation, contiguous deletion, and fragmentation, CAKR retained strong clustering purity for most benchmark datasets, with the retained fraction of sequence information having the largest effect on performance. The average-purity trends across k under simulated incompleteness are summarized in Supplementary Fig. 19. A complementary barcode-based fragment placement experiment with external influenza queries further showed that partial sequences can still be matched to the correct viral type and homologous gene segment in an alignment-free manner.

To assess whether performance is driven mainly by simple compositional effects, we further performed sequence-bias control experiments involving low-complexity masking, repeat masking, GC normalization, and mono- and dinucleotide shuffling. Across datasets, the strongest degradations occurred under shuffling, supporting the view that CAKR captures higher-order sequence organization rather than simple composition alone (Supplementary Section 11 and Supplementary Figs. 20–25).

Interpretability of CAKR

Rational learning requires explainable features and an interpretable neural network design. Figure 4 illustrates the interpretability of CAKR through two representative examples, demonstrating how the persistent Stanley–Reisner theory (PSRT) encodes the filtration structure via the algebraic decomposition of facet ideals. Panels (a–d) depict a synthetic example of eight points uniformly placed on a circle of radius 1. As the filtration parameter increases, simplicial complexes are constructed by adding a k-simplex whenever all its vertices lie within the closed ball centered at at least one of its vertices. Panel (b) shows the barcodes of persistent facet ideals: P0 encodes vertex lifespans (isolation), P1 tracks edge persistence until subsumed into higher simplices, and P2, P3 capture 2- and 3-simplices respectively. The barcodes reveal uniform connectivity and regular geometric structure in this example.

Fig. 4. Illustrative example of persistent homology and persistent Stanley–Reisner invariants.

Fig. 4

a A filtration of a simplicial complex arising from an octagon. b Persistent facet ideal barcodes derived from the same filtration, encoding combinatorial face-level activity rather than homology. c, d Persistent f-vector and h-vector curves, respectively, derived from the same filtration. e The N1-U.S-P primer sequence, with the positions of the nucleotide C marked. f Persistent facet ideal barcodes of these positions, reflecting the activity of the persistent facet ideals under the induced filtration. g, h Persistent f-vector and h-vector curves, respectively, derived from the filtration of the sequence. The symbol r denotes the filtration radius; Pi denotes the collection of persistent i-dimensional facet ideals; and fk and hk denote the kth components of the persistent f- and h-vectors, respectively. In a, colors distinguish vertices, edges, and higher-dimensional simplices. In b, f, barcode colors distinguish P0,P1,P2, and P3, as indicated by the in-panel key. In c, d, g, h, curve colors distinguish the indexed fk- and hk-components indicated by the in-panel labels. Nucleotide colors in (e) denote A (teal), C (purple), G (green), and T (yellow). Source data are provided as a Source Data file.

Panels (e, f) present a biologically motivated example, examining the distribution of the 1-mer C in the N1-U.S-P primer. Each occurrence of C is treated as an integer in 1D space. Equal-length bars in P0 reflect regularly spaced cytosines, while longer bars indicate isolated ones. Notably, one isolated C connects to its nearest neighbor at filtration radius 5, and its corresponding edge vanishes at radius 7, suggesting nearby higher-order interactions. The brief lifespans in P2 indicate that 2-simplices (triangles) are formed quickly and filled soon thereafter, implying tight local clustering among cytosines.

Panels (c) and (g) show the persistent f-vectors, fk(r), quantifying the number of active k-simplices at filtration value r, while Panels (d) and (h) show the corresponding h-vectors, hk(r), measuring the incremental additions of independent k-facets. In both examples, increases in f1 signal edge formation, and any growth in fk for k ≥ 2 is necessarily preceded by growth in f1, making f1 a necessary precursor for higher-order structure. The h-vectors further reveal when such structures are nontrivial versus when they are absorbed into larger simplices.

Altogether, CAKR’s interpretability stems from its ability to encode k-mer spatial configurations as algebraic signatures derived from the evolution of facet ideals across the filtration. The persistent f- and h-vectors provide structured summaries of how generators of these ideals emerge and interact at varying scales. In particular, they capture the combinatorial and geometric regularity of k-mer distributions, offering insights into clustering, isolation, and interaction patterns among sequence motifs in both synthetic and biological settings.

Generalizability

Because CAKR encodes a sequence purely as a word over a finite alphabet, it extends naturally beyond the DNA datasets analysed here. In practice, we transcribed RNA sequences to their DNA equivalents, but the same framework also accommodates amino-acid strings for protein sequence modeling. More broadly, any categorical sequence data built from a limited alphabet of symbols can be modeled by the proposed commutative-algebraic constructions. The proposed CAKR therefore provides a versatile mathematical foundation for sequence analysis and, by extension, for data science tasks involving symbolic or ordered data.

Although all datasets considered in this work consist of single-segment viral or bacterial genomes, the CAKR representation is defined for arbitrary DNA strings and therefore extends directly to multi-chromosomal organisms. For such genomes, one may either (i) featurize and aggregate the CAKR descriptors of all chromosomes to obtain a single organism-level representation for phylogeny and classification or (ii) analyze chromosomes individually to investigate chromosome-specific structure. In the organism-level setting, genomes from the same species would generally be expected to cluster together in the CAKR feature space. In the chromosome-level setting, chromosomes from the same species may still appear closer to one another because they often share similar compositional and mutational biases, and therefore exhibit broadly similar low-order k-mer distributions. Moreover, CAKR leverages both k-mer content and locality/positional information, which can further promote within-species similarity when such genomic signatures are shared. However, the method does not assume that all chromosomes from a given species must form a single tight cluster. It can therefore capture biologically meaningful heterogeneity associated with sex chromosomes, structural variation, repeats, or horizontal transfer.

Extended comparisons

Several well-established methods provide useful additional points of comparison with CAKR. Among alignment-free approaches, we considered CVTree39 and Mash40, while among widely used phylogenetic inference frameworks we included IQ-TREE 3 and RAxML-NG41,42. We evaluated these methods on representative benchmark datasets, including SARS-CoV-2, mammalian mitochondrial genomes, and rhinoviruses where applicable.

For the SARS-CoV-2 dataset, the purity scores were 0.9028 for CVTree (Supplementary Fig. 10), 0.9028 for Mash (Supplementary Fig. 11), 0.9600 for IQ-TREE 3 (Supplementary Fig. 2a), and 0.9028 for RAxML-NG (Supplementary Fig. 2b). On the mammalian dataset, CVTree and Mash achieved purity scores of 0.8438 and 0.9531, respectively, whereas on the rhinovirus dataset their purity scores were 0.7778 and 0.6485, respectively. These results show that both additional alignment-free baselines and standard phylogenetic inference frameworks can perform strongly on selected datasets, with IQ-TREE 3 performing particularly well on SARS-CoV-2. At the same time, CAKR remains competitive while maintaining a fully alignment-free, sequence-representation-based framework. The corresponding phylogenetic trees produced by CVTree, Mash, IQ-TREE 3, and RAxML-NG are shown in Supplementary Figs. 2, 10, and 11, respectively, and the detailed CAKR-SEPP fragment-placement comparison is provided in Supplementary Section 10 and Source Data file.

In addition, because fragmentary sequence placement is an important practical setting that is not fully captured by whole-tree reconstruction alone, we also compared CAKR in the Supplementary Information with SEPP, a standard HMM-based phylogenetic placement framework43. In the influenza B PB1 fragment-placement experiment, the two frameworks showed substantial agreement at the level of ranked candidate placements: exact Top 1 agreement was observed for 6 of 15 fragmentary queries, and 11 of 15 fragments showed overlap between the top three CAKR candidates and the top two SEPP placements. Since the two approaches rely on different scoring systems, the comparison is most meaningful in terms of agreement in ranked placements rather than raw score magnitudes. These results indicate that CAKR compares favorably with established alternatives across both whole-sequence phylogenetic reconstruction and fragment-placement settings, while retaining the practical advantage of being fully alignment-free. Detailed results for this comparison are provided in Supplementary Section 10 and the Source Data file.

Computational efficiency

Although CAKR consistently outperformed the competing methods in phylogenetic accuracy and viral classification, it remains computationally practical for genome-scale analyses. As expected from its richer algebraic construction, CAKR is more demanding at the feature-extraction stage. Thus, the practical trade-off is a higher one-time featurization cost in exchange for improved downstream phylogenetic and classification performance. For fixed k, all methods exhibit similar runtime growth with respect to sequence length, while CAKR incurs a higher per-sequence featurization cost due to Vietoris–Rips complex construction and the computation of persistent facet invariants. For example, facet-based CAKR featurization increases from ≈4.6 × 10−3 s at length 10 to ≈2.8 × 102 s at length 3 × 106. These timings refer to featurization only; pairwise distances are subsequently computed in the resulting feature space, where the low dimensionality at small k keeps distance calculations inexpensive. Complete wall-clock runtime comparisons against NVM, FFP-JS, FFP-KL, FPS, MKS, and MAFFT across multiple sequence lengths and k-mer sizes are provided in Supplementary Section 5 and Supplementary Fig. 12.

In terms of memory, the final CAKR feature representation is compact for the parameter choices used in this study. For N DNA sequences, k-mer size k, F filtration values, and maximum facet dimension dmax, the dense feature matrix has size N×((dmax+1)F4k). In our experiments, we used dmax=0 and F = 3, so the storage requirement is 24N4k bytes in double precision. Consequently, for the values k = 3, 4, 5, and 6 used in the benchmarks, the final stored feature arrays remain modest for the phylogenetic datasets and manageable for the larger viral classification datasets. The main computational burden is therefore the one-time featurization step, rather than the storage of the resulting features or the subsequent pairwise distance calculations. Detailed memory estimates are provided in Supplementary Section 5.

Overall, CAKR computes richer topological and algebraic descriptors than simpler k-mer frequency methods, which makes the initial feature-extraction stage more expensive. However, the resulting fixed-length representations are compact, inexpensive to compare, and can be reused for downstream distance calculation, clustering, tree reconstruction, and classification. Thus, CAKR is particularly useful when improved phylogenetic or classification accuracy is prioritized over minimizing featurization time, while simpler k-mer methods may remain preferable for very rapid exploratory screening.

Incomplete sequences

Our method is primarily designed for global, alignment-free comparison of complete sequences, since it relies on the overall k-mer distribution and the associated positional and combinatorial structure across the sequence. For this reason, CAKR is not intended as a direct replacement for local alignment, short-read mapping, or contig assembly methods, which require explicit base-level local matching.

At the same time, to assess applicability beyond ideal full-length inputs, we performed additional analyses on incomplete genomes, fragmented sequences, and fragment placement-style inference (Supplementary Section 10). In a simulated incompleteness experiment, benchmark datasets were degraded through truncation, contiguous deletion, and fragmentation to mimic partial recovery and draft-quality assemblies. Across most datasets, clustering purity remained strong over a wide range of degradation settings, with the retained fraction of sequence information emerging as the dominant determinant of performance, whereas fragmentation alone did not produce a consistent collapse in clustering quality. In particular, Ebolavirus, rhinovirus, HEV, and influenza HA gene remained comparatively robust, whereas mammalian mitochondrial genomes showed more moderate sensitivity. For SARS-CoV-2, purity remained low across nearly all incomplete settings, indicating that the difficulty in this case is driven more by the intrinsic similarity of the sequences rather than by incompleteness itself.

We further evaluated a barcode-based fragment placement setting using external influenza query sequences. In this proof-of-concept experiment, CAKR barcodes computed from partial or complete query genes were compared against a reference barcode library derived from influenza genome segments. Each of the 14 external query sequences attained its highest cosine similarity to the correct viral type and the corresponding homologous gene segment (Source Data file), indicating that CAKR can retain meaningful discriminatory information even for previously unseen partial inputs.

These results show that CAKR does not rely exclusively on perfectly assembled full genomes and can remain effective under substantial truncation, deletion, and fragmentation, while also supporting fragment placement-style inference in an alignment-free manner. Nevertheless, when incompleteness becomes substantial, the resulting representations may deviate more strongly from those of the corresponding complete genomes, potentially introducing variability in clustering, distance estimation, and downstream inference. Large-scale benchmarking on true low-coverage assemblies or shotgun metagenomic datasets remains an important direction for future work.

Methods

In this section, we provide an overview of persistent Stanley–Reisner theory. Then, we provide the construction of the k-mer algebraic representations of sequences. We also propose a new purity metric to assessing the performance of phylogenetic tree reconstruction tools.

Persistent Stanley–Reisner theory

Persistent Stanley–Reisner theory is a novel framework for algebraic data analysis, leveraging tools from commutative algebra. Unlike traditional topological data analysis, which emphasizes geometric and topological features, such as loops and voids, through persistent homology44, persistent Stanley–Reisner theory focuses on the algebraic and combinatorial structure of simplicial complexes, using invariants derived from commutative algebra29. A filtration process is then applied to these complexes to track the evolution and persistence of such features across multiple spatial or geometric scales. This approach introduces algebraic invariants, such as persistent h-vectors, f-vectors, graded Betti numbers, and facet ideals, thus providing a new algebraic perspective within the broader framework of algebraic data analysis.

Persistent Stanley–Reisner structures over a filtration

Let k be a field, and let Δ be a simplicial complex on the finite vertex set V = {x1,…,xn}. Suppose f:Δ→R is a monotone function, i.e., f(τ) ≤ f(σ) whenever τ ⊆ σ, which induces an increasing filtration

f~:=Δtt∈R,whereΔt:=σ∈Δ∣f(σ)≤t.

Let S = k[x1,…,xn] be the standard graded polynomial ring over k, and for each t∈R, define the Stanley–Reisner ideal of Δt as

It:=xi1⋯xir∣{xi1,…,xir}∉Δt⊆S,

with corresponding Stanley–Reisner ring

k[Δt]:=S/It.

Since the filtration is increasing, the subcomplexes satisfy Δs ⊆ Δt for s ≤ t, which implies a descending chain of monomial ideals:

Is⊇Itforalls≤t.

Each ideal It admits a canonical primary decomposition indexed by the facets of Δt:

It=⋂σ∈F(Δt)Pσ,wherePσ:=(xi∣xi∉σ). 1

We refer to the collection Pt:={Pσ∣σ∈F(Δt)} as the facet ideals of Δt.

To capture the dimension-wise structure, we stratify by face dimension: for each i ≥ 0, where we define

Pit:=Pσ∈Pt∣dim(σ)=i, 2

so that

Pt=⨆i=0dim(Δt)Pit.

We define persistence algebraically as follows: a facet ideal Pσ∈Pit is said to persist to level t′>t if Pσ∈Pit′. The set of such persistent i-dimensional primes is

Pit,t′:=Pit∩Pit′. 3

The corresponding facet persistent number is given by

Fit,t′:=Pit,t′, 4

which records the number of i-dimensional prime components common to Δt and Δt′.

The collection {Fit,t′}i,t,t′ serves as a combinatorial invariant encoding the persistence of prime facets in the Stanley–Reisner filtration, providing an algebraic analogue of topological barcodes in persistent homology.

Persistent graded Betti numbers of Stanley–Reisner rings

Let k be a field and S = k[x1,…,xn] the standard graded polynomial ring. For each filtration level t∈R, the Stanley–Reisner ring k[Δt] ≔ S/It inherits a natural Z-graded S-module structure and admits a minimal graded free resolution:

⋯→⨁jS(−j)βi,j(k[Δt])→⋯→k[Δt]→0, 5

where βi,j(k[Δt]):=dimkToriS(k[Δt],k)j are the graded Betti numbers.

Hochster’s formula relates these graded Betti numbers to the topological Betti numbers of the induced subcomplexes:

βi,j+i(k[Δt])=∑W⊆V∣W∣=j+irankH~j−1(ΔWt;k), 6

where H~j−1(ΔWt;k) denotes the (j − 1)-st reduced simplicial homology group over k, and ΔWt:={σ∈Δt∣σ⊆W} is the subcomplex induced on the vertex set W ⊆ V45.

In particular, Hochster’s formula can be reformulated in terms of the (non-reduced) Betti numbers of induced subcomplexes. For each integer i ≥ 0, the following identities hold:

βi,i+1(k[Δt])=∑W⊆V∣W∣=i+1β0(ΔWt)−1, 7
βi,i+j(k[Δt])=∑W⊆V∣W∣=i+jβj−1(ΔWt),forallj≥2, 8

where ΔWt denotes the subcomplex of Δt induced on the vertex subset W ⊆ V, and βr(ΔWt) denotes the r-th Betti number of ΔWt with coefficients in k.

To refine this in a persistent setting, for t≤t′, we define the persistent graded Betti number

βi,i+jt,t′(k[Δ]):=∑W⊆V∣W∣=i+jrankιj−1t,t′:H~j−1(ΔWt)→H~j−1(ΔWt′), 9

where ιj−1t,t′ is the homomorphism on reduced homology induced by inclusion. This provides a multigraded algebraic refinement of classical persistent Betti numbers, encoding both topological persistence and the combinatorial properties of the evolving homology classes.

In the special case where ∣W∣ = ∣V∣, the persistent graded Betti number reduces to

βi,∣V∣t,t′=β∣V∣−i−1t,t′,

recovering the classical persistent Betti number of homological degree ∣V∣ − i − 1. More generally, the family {βi,i+jt,t′}i,j encodes a richer multiscale invariant that interpolates between algebraic and topological persistence.

Persistent f- and h-vectors

Let (Δt)t∈R be a filtration of a finite (d − 1)-dimensional simplicial complex Δ, induced by some face function f:Δ→R. For each fixed level t, the complex Δt consists of those faces σ ∈ Δ with f(σ) ≤ t. At each level t, one may associate the classical combinatorial invariants of face counts and their derived quantities.

The f-vector of Δt is defined as

f(Δt)=f−1t,f0t,f1t,…,fd−1t,

where fit denotes the number of i-dimensional faces in Δt, f−1t=1 by convention, and d=d(t)=dim(Δt)+1. The associated h-vector is defined as h(Δt)=(h0t,…,hdt), where

hmt=∑j=0md(t)−jm−j(−1)m−jfj−1t,form=0,…,d(t), 10

and hmt=0 for all m > d(t).

This transformation is invertible, with the inverse relation given by

fm−1t=∑i=0md(t)−im−ihit,form=0,…,d(t), 11

where d(t)=dim(Δt)+1.

To extend these invariants to the persistent setting, one replaces the classical Betti numbers with the persistent graded Betti numbers βi,jt,t′ defined over filtration levels t≤t′29. This enables a multiscale, combinatorial interpretation of how face structures persist across different filtration levels.

Let (Δt)t∈R be a filtration of a simplicial complex Δ. The persistent h-vector between levels t≤t′ is defined as

hmt,t′:=∑j=0mn−d(t′)+m−j−1m−j∑i=0j(−1)iβi,jt,t′,form=0,…,d(t′), 12

where βi,jt,t′ denotes the persistent graded Betti numbers of k[Δ] over [t,t′], and let d(t′)=dim(Δt′)+1.

The corresponding persistent f-vector is then defined via the inverse transformation:

fm−1t,t′:=∑i=0md(t′)−im−ihit,t′,form=0,…,d(t′). 13

These persistent vectors capture how the combinatorial structure of Δ evolves through the filtration, blending face enumeration with homological persistence. In contrast to the classical static f- and h-vectors, their persistent counterparts reflect the dynamic appearance and disappearance of faces and their relations across multiple scales, providing richer algebraic-combinatorial invariants for analysis.

We now consider the following simplifications, which will play a central role in the k-mer algebraic representation framework. These observations refine the relationship between persistent h-vectors and persistent graded Betti numbers, serving to streamline computations in applications involving Vietoris–Rips complexes derived from sequence data.

Let βi,jt,t′ denote the persistent graded Betti numbers of the Stanley–Reisner ring k[Δ] over the filtration interval [t,t′], as defined in Eq. 9. To streamline notation, we set

Bj:=∑i=0j(−1)iβi,jt,t′. 14

Alongside this, one introduces the coefficients

αj(m):=n−d(t′)+m−j−1m−j, 15

which appear in the linear transformation relating the h-vector of a simplicial complex to its graded Betti numbers.

It follows that the persistent h-vector component hmt,t′ satisfies the identity

hmt,t′=∑j=0mαj(m)Bj,foreachm∈N. 16

Additional structural identities among the persistent Betti numbers further simplify this formula. In particular, it is known that

β0,0t,t′=1,βi,it,t′=0foralli≥1,β0,jt,t′=0forallj≥1,βi,jt,t′=0foralli>j.

Consequently, one obtains

B0=β0,0t,t′=1,

and for each j ≥ 1, the alternating sum simplifies to

Bj=∑i=1j−1(−1)iβi,jt,t′.

k-mer algebraic representations of sequences

In this section, we review the k-mer representation framework introduced by Hozumi et al.30, which provides a foundational method for embedding sequences as collections of integer sequences in a geometric space. Let A be a finite alphabet and let k > 0 be an integer. A k-mer over A is a word x=x1x2⋯xk∈Ak. Given a fixed k-mer x∈Ak, we define the k-mer indicator function δx:Ak→{0,1} by

Δx(y)=1ify=x0otherwise 17

Given a sequence S=s1s2⋯sN∈AN, we define the set of positions at which the k-mer x occurs in S as

Sx=i∈[1,N−k+1]∣δx(sisi+1⋯si+k−1)=1. 18

The corresponding pairwise distance matrix Dx=(dijx∣i,j∈Sx) is defined by

dijx=∣i−j∣,foralli,j∈Sx. 19

These distance matrices serve as the input for persistent Stanley–Reisner computations. Specifically, for each k-mer x∈Ak, the corresponding sequence of integers Sx⊂R gives rise to a family of Stanley–Reisner algebraic feature vectors computed over a filtration interval [r0, r1]. For filtration values r,r′∈[r0,r1] with r≤r′, we define

vxr,r′=vir,r′(x)i∈N,

where vir,r′(x) denotes a persistent invariant of dimension i, such as the f-vector, h-vector, or facet number, associated with the Vietoris–Rips complex built from Sx.

To simplify notation, we restrict to the diagonal case r=r′, and denote the resulting feature vector by

vx:=vi(x)i∈N.

For a fixed integer k > 0, the full representation of the sequence S∈AN is given by the concatenation of these vectors over all k-mers:

vSk:=vxx∈Ak=vi(x)x∈Aki∈N,

which we refer to as the k-mer algebraic representation of S at level k. This construction yields a feature vector indexed jointly by algebraic dimension i and k-mer x∈Ak.

Algebraic genetic distances

To compare two sequences S1∈AN1 and S2∈AN2, we define a family of weighted Euclidean metrics that aggregate Stanley–Reisner algebraic information across both the algebraic dimensions and k-mer lengths. Let ak,i ≥ 0 denote a non-negative weight assigned to homological dimension i at scale k. The dimension- and scale-weighted algebraic distance is defined by

dv(S1,S2):=∑k=1K∑i=0Dkak,i⋅vS1,ik−vS2,ik2, 20

where vS,ik:=vi(x)∣x∈Ak is the vector of dimension-i features computed over all k-mers in S, and Dk is the maximum dimension considered for k.

There are various strategies for selecting the weights ak,i. A common choice in the literature is to assign exponentially decaying weights across scales, for example ak,i = 2−(k − 1), which emphasizes shorter k-mers while retaining contributions from larger scales. To additionally incorporate the homological dimension, one may use

ak,i=2−(i⋅K+k−1), 21

which imposes exponential decay both in the scale k and in the dimension i.

In contrast to these monotone decay schemes, we adopt a data-driven weighting centered around a preferred scale k*, selected as described in “Choice of k-mer Size and Feature Representation.” Specifically, we assign the largest weight to k* and impose exponential decay as k moves away from this scale in either direction. This leads to the choice

ak,i=2−∣k−k*∣⋅2−i∑ℓ=1K∑j=0Dℓ2−∣ℓ−k*∣⋅2−j, 22

where the normalization ensures that the weights sum to one.

This construction preserves the multiscale nature of the representation while concentrating the contribution around the empirically most informative k-mer length. In particular, scales near k* receive the highest weight, while both smaller and larger k-values are exponentially downweighted, thereby balancing resolution and stability in the resulting metric.

Within the CAKR framework, three distinct types of persistent algebraic features are employed to define pairwise distances between sequences. Specifically, the distance df(S1, S2) is derived from the f-vector curves, where the feature vector v is computed from these curves; the distance dh(S1, S2) is defined analogously using h-vector curves; and the distance dF(S1,S2) is based on the facet count vectors associated with the underlying filtration.

The final composite distance, integrating these three feature types, is given by

d(S1,S2):=df(S1,S2)+dh(S1,S2)+dF(S1,S2), 23

which captures a broad range of persistent characteristics across multiple features and filtration levels. This composite metric constitutes the core of the CAKR approach to alignment-free sequence comparison.

In the applications considered in this work, we restrict our attention to a single feature representation, namely, the facet vector curves, and employ a fixed window length k for k-mers.

Choice of k-mer size and feature representation

Rather than fixing a single k-mer size a priori, we adopt a data-driven strategy to determine an optimal scale k* for each dataset. For each candidate k, we construct the feature matrix X(k)∈Rn×4k and evaluate three complementary criteria capturing structural, informational, and stability properties of the representation.

Specifically, we consider: (i) a coverage-based criterion that identifies a structural scale via the elbow of the coverage curve, (ii) an information-sparsity score that balances entropy with feature sparsity, and (iii) a stability-adjusted entropy that penalizes contributions from rare (singleton) features. These criteria capture distinct aspects of the trade-off between expressiveness and robustness across k.

Across datasets, these criteria do not yield a consistent ordering. In particular, the information-sparsity score tends to favor smaller values of k, while the stability-adjusted entropy closely follows the structural elbow. To reconcile these differences, we introduce a scoring framework that evaluates each criterion relative to the coverage-based structural reference, rewarding earlier selections while penalizing large deviations. This induces a ranking of the criteria, from which we derive weights reflecting their relative importance.

The global scale k* is then selected via a weighted consensus rule combining the three criteria. Full definitions of the criteria, the scoring function, and the weighting procedure are provided in the Supplementary Information (Supplementary Section 9). Representative dataset-specific scale-selection curves for the SARS-CoV-2, mammalian mitochondrial, and rhinovirus datasets are shown in Supplementary Figs. 16–18.

In addition to the persistent facet count vector, we also tested the f- and h-vector representations. Their performances were found to be comparable; however, the facet representation yielded more stable and compact features, whereas the f- and h-vectors tended to produce larger components.

For the Vietoris–Rips filtration, we employed thresholds at 0, 4k, 2 × 4k,…, ensuring adequate coverage of k-mer co-occurrence patterns. A relatively small number of filtration steps was chosen to maintain computational efficiency and limit memory usage, while still achieving stable performance across datasets.

Further empirical analyses of the effect of k-mer size across methods, including direct 1-NN accuracy curves and cumulative-best phylogenetic performance summaries over k, are provided in Supplementary Section 6 and Supplementary Figs. 13, 14.

Computational simplifications of the persistent h-vectors and f-vectors

Vietoris–Rips simplicial complexes arising from the k-mer algebraic representations possess a structural property that significantly simplifies their algebraic analysis. Specifically, many of the persistent graded Betti numbers vanish in higher homological degrees, which reduces the complexity of computations involving persistent h-vectors.

Let (X, d) be a finite metric space (or a finite set of points in a metric space). For a scale parameter ε ≥ 0, the Vietoris–Rips complex VRε(X) is the abstract simplicial complex whose vertex set is X, and where a finite subset σ = {x0,…,xp} ⊆ X spans a p-simplex whenever all pairwise distances are bounded by ε, i.e.,

d(xi,xj)≤εforall0≤i<j≤p.

Equivalently, VRε(X) is the clique (flag) complex of the proximity graph on X with an edge between x and y whenever d(x, y) ≤ ε. We refer to ref. 46 for background and standard properties.

As ε increases, these complexes form a nested family

VRε1(X)⊆VRε2(X)wheneverε1≤ε2,

which is called the Vietoris–Rips filtration. Persistent homology is then computed from the induced maps on homology along this filtration, tracking the birth and death of topological features across scales.

The Vietoris–Rips construction is widely used in topological data analysis because it is entirely determined by pairwise distances and is therefore straightforward to build from a distance matrix; moreover, it serves as a computationally efficient proxy for other distance-based complexes (e.g., Čech-type constructions) while retaining meaningful multi-scale topological information46.

The following proposition formalizes this observation and highlights its relevance to k-mer algebraic representations:

Proposition 1

Let Δ = Δt denote the Vietoris–Rips complex at scale t associated with a sequence X⊂R. Then for every subset W ⊆ V and every j ≥ 2, the persistent Betti numbers satisfy

βj−1t,t′(ΔW)=0.

As a consequence,

βi,i+jt,t′=0foralli≥1,j≥2,

and the only potentially nonzero contributions occur in degree shifts of one, namely

βi,i+1t,t′=∑W⊆V∣W∣=i+1β0t,t′(ΔW)−1.

Therefore, the alternating sum of persistent Betti numbers at total degree j simplifies to

Bj:=∑i=0j(−1)iβi,jt,t′=(−1)j−1βj−1,jt,t′forallj≥1.

In particular, the persistent h-vector expression in Eq. 16 becomes

hmt,t′=α0(m)+∑j=1mαj(m)(−1)j−1βj−1,jt,t′,withh0t,t′=1.

To establish Proposition 1, we prove a more general structural result concerning Vietoris–Rips complexes over sequences in R. Let X⊆R be a finite sequence, and let VRϵ(X) denote the Vietoris–Rips complex at scale ϵ.

Each facet F ⊆ VRϵ(X) admits a unique minimal element x=inf(F)∈X. Moreover, if another facet G ⊆ VRϵ(X) satisfies inf(G)=x, then necessarily G = F. That is, the minimal element uniquely determines the facet. Consequently, the assignment x ↦ Fx, where Fx denotes the unique facet with minimal element x, satisfies

Fx=Fy⇔x=y.

In particular, the collection of facets is in bijective correspondence with the set of the minimal elements of the facets, and hence can be linearly ordered by their infima:

Fx≤Fy⇔x≤y.

Proposition 2

Let X⊆R be a finite sequence. Then for all q ≥ 1,

Hq(VRϵ(X))=0.

Proof.

Let VRϵ(X)=⋃i=1nFxi, where the facets Fxi are ordered such that x1 < x2 < ⋯ < xn.

We proceed by induction on n.

Base case: when n = 1, VRϵ(X) is a single simplex, which is contractible. Therefore, Hq = 0 for all q ≥ 1.

Inductive step: assume the result holds for n − 1 facets, where n > 1. Let:

K1=⋃i=1n−1Fxi,K2=Fxn,K=K1∪K2.

Note that K1, K2, and K1 ∩ K2 are all simplicial complexes. Notice that the vertices of K1 ∩ K2 lie within the interval [xn, xn − 1 + ϵ], whose length is at most ϵ, the intersection K1 ∩ K2 is either empty or a simplex.

By the Mayer–Vietoris sequence, we obtain the long exact sequence in homology:

⋯→Hq(K1∩K2)→Hq(K1)⊕Hq(K2)→Hq(K)→Hq−1(K1∩K2)→⋯

By the inductive hypothesis, Hq(K1) = 0 for all q ≥ 1, and since K2 is a simplex, Hq(K2) = 0 as well. Furthermore, K1 ∩ K2 is either a simplex or empty, so:

Hq(K1∩K2)=0forallq≥1.

Thus, the exact sequence reduces to:

0→Hq(K)→0⇒Hq(K)=0forallq≥2.

To analyze H1(K), we consider:

0→H1(K)→H0(K1∩K2)→H0(K1)⊕H0(K2)→H0(K)→0.

If K1∩K2=∅, then H1(K) = 0 since K = K1⊔K2 is a disjoint union of two contractible subcomplexes. Thus, the result holds in this case.

Suppose instead that K1∩K2≠∅. Then K1 ∩ K2 is a simplex and hence contractible, in particular connected. Since K2 is also a simplex, it is connected and contractible. Moreover, the inclusion of K2 into K = K1 ∪ K2 does not change the number of connected components, so K and K1 have the same number of components. Therefore, the canonical map

H0(K1)⊕H0(K2)→H0(K)

has kernel of dimension one, namely dimH0(K2)=1. Since K1 ∩ K2 is connected, the induced map

H0(K1∩K2)→H0(K1)⊕H0(K2)

is injective. Consequently, in the Mayer–Vietoris sequence, the connecting homomorphism

H1(K)→H0(K1∩K2)

must be the zero map. It follows that H1(K) = 0 in this case as well. By induction on the number of simplices, we conclude that

Hq(VRϵ(X))=0forallq≥1.

This structural property does not extend to higher-dimensional ambient spaces X⊂Rd for d ≥ 2; for instance, consider the Vietoris–Rips complex formed from the vertices of a regular hexagon in R2. Therefore, Proposition 1 is a consequence of the special linear ordering available in one-dimensional point clouds. This leads to a significant simplification in the computation of persistent Betti numbers arising from k-mer algebraic representations.

Given a simplicial complex Δ, its 1-skeleton induces an undirected graph G(Δ) with vertex set V and edge set

E(Δ):={u,v}⊆V∣{u,v}∈Δ,∄w∈Vwithu<w<v.

Equivalently,

E(Δ)={u,v}∈Δ∣(u,v)∩V=∅.

When Δ is a Vietoris–Rips complex built on k-mer representations in R, the persistent Betti numbers of Δ are entirely determined by the topology of the associated graph G(Δ). This follows directly from Proposition 1. We formalize this relationship in the following theorem:

Theorem 1

Let Δ be a Vietoris–Rips simplicial complex on a finite sequence X⊂R, and let G(Δ) = (V, E(Δ)) denote its 1-skeleton. Then the persistent Betti numbers of Δ satisfy:

βi,i+1(G(Δ))=βi,i+1(Δ),andβi,i+j(G(Δ))=βi,i+j(Δ)=0forallj≥2.

Purity metric for assessing the performance of phylogenetic tree reconstruction methods

We introduce a purity metric for assessing monophyly in phylogenetic trees. Let S be a finite set of size n = ∣S∣, and let P={S1,S2,…,Sk} be a partition of S into disjoint subsets such that ⋃i=1kSi=S. We define the purity of the partition P as

purity(P)=∑i=1k∣Si∣n2. 24

This quantity reflects the degree to which elements are concentrated within the subsets of the partition. A higher purity indicates that the majority of elements reside in a small number of large subsets, while a lower purity corresponds to a more evenly distributed partition.

To illustrate, consider several representative scenarios. If the partition is perfect in the sense that all elements are grouped into a single subset, i.e., k = 1, then

purity(P)=nn2=1,

which is the maximal possible value. If the partition consists of n singleton subsets (i.e., ∣Si∣ = 1 for all i), then

purity(P)=∑i=1n1n2=1n,

which is minimal. If the partition consists of two equal-sized subsets, each of size n/2, then

purity(P)=2122=12.

Finally, if one subset dominates the partition, for example, with ∣S1∣ = n − 1 and ∣S2∣ = 1, then

purity(P)=n−1n2+1n2=1−2(n−1)n2,

which approaches 1 as n → ∞, but is strictly less than 1 for any finite n.

Let S be a finite set of leaf nodes in a phylogenetic tree, and let each element of S be assigned a categorical label (e.g., species, clade, or functional class). For each label ℓ, let S(ℓ) ⊆ S denote the set of leaves with label ℓ, and let n(ℓ) = ∣S(ℓ)∣ be the number of such leaves.

To assess the purity of the tree with respect to label ℓ, we identify all maximal subtrees whose leaves are exclusively labeled ℓ. These subtrees define a partition Pℓ={S1,S2,…,Sk} of S(ℓ), where each Si ⊆ S(ℓ) is the set of leaves in a pure subtree.

The purity of the label ℓ is then defined as:

purity(Pℓ)=∑i=1k∣Si∣n(ℓ)2, 25

where the numerator ∣Si∣ denotes the size of a pure subtree and the denominator normalizes by the total number of leaves of label ℓ.

A purity of 1.0 indicates that all leaves of label ℓ are perfectly clustered under a single subtree (i.e., monophyletic), while a lower purity reflects fragmentation of that label across multiple subtrees. Averaging the purity scores across all labels provides an overall measure of the taxonomic coherence of the tree:

avg_purity=1∣L∣∑ℓ∈Lpurity(Pℓ), 26

where L is the set of all unique labels in the tree.

This approach is beneficial for evaluating the extent to which a phylogenetic tree respects known groupings, such as taxonomic families or functional clusters, without requiring an explicit reference.

Software and computational environment

The CAKR analyses were implemented in Python using NumPy 1.26.4, scikit-learn 1.4.2, SciPy 1.13.1, GUDHI 3.10.1, Biopython 1.84, pandas 2.2.2, and ETE Toolkit 3.1.3. The exact source-code release used in this study is CAKR v1.0.0, archived on Zenodo at https://doi.org/10.5281/zenodo.21426540. Comparative analyses additionally used MAFFT, IQ-TREE 3, RAxML-NG, Mash, CVTree3, SEPP, and Interactive Tree Of Life (iTOL) v6. Phylogenetic trees were constructed using the UPGMA algorithm and, in the supplementary sensitivity analyses, the neighbor-joining algorithm.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

Reporting Summary (2.5MB, pdf)

Source data

Source Data (52.1KB, xlsx)

Acknowledgements

We gratefully acknowledge all data contributors, i.e., the authors and their originating laboratories responsible for obtaining the specimens, and their submitting laboratories for generating the genetic sequence and metadata and sharing via the GISAID Initiative, on which this research is based. F.S. gratefully acknowledges the support of King Fahd University of Petroleum and Minerals.

Author contributions

F.S. designed the method and study, wrote the code, performed the computational studies, wrote the first draft, and revised the manuscript. Y.H. designed the method and study, collected the data, performed the computational studies, and revised the manuscript. M.Z. designed the study, wrote the code, performed computational studies, and revised the manuscript. J.J.W. revised the manuscript. H.F. wrote code, prepared figures, and revised the manuscript. G.-W.W. designed the study, conceptualized and supervised the project, acquired funding, and revised the manuscript.

Peer review

Peer review information

Nature Communications thanks Timothy Rozday and other anonymous reviewers for their contribution to the peer review of this work. A peer review file is available.

Funding

This work was supported in part by NIH grants R01AI164266 and R35GM148196, the MSU Research Foundation, the University of Georgia, and the Georgia Research Alliance.

Data availability

The genomic sequence data analyzed in this study were obtained from NCBI, including GenBank and the NCBI Virus resource, and from GISAID for the SARS-CoV-2 dataset. GenBank accession.version identifiers are provided in Supplementary Tables 8–13, and the GISAID accession identifiers for the 44 SARS-CoV-2 genomes are provided in Supplementary Table 7. The underlying GISAID sequences and metadata are not redistributed because of GISAID access restrictions and must be obtained directly from GISAID by registered users in accordance with its terms of use. The curated and processed datasets, excluding the underlying GISAID files, are available from Zenodo at https://doi.org/10.5281/zenodo.1875792847. Additional data are provided in the Supplementary Information and Source Data file. Further supporting materials not included in these files are available from the corresponding author upon request. Source data are provided with this paper.

Code availability

The source code used to implement the CAKR framework and perform the comparative analyses reported in this study is publicly available on GitHub at https://github.com/FaisalSuwayyid/CAKL. The version used in this study, CAKR v1.0.0, has been archived on Zenodo at https://doi.org/10.5281/zenodo.2142654048.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Supplementary information

The online version contains Supplementary material available at https://doi.org/10.1038/s41467-026-76429-z.

References

  • 1.Rubin, G. M. et al. Comparative genomics of the eukaryotes. Science287, 2204–2215 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Frazer, K. A., Pachter, L., Poliakov, A., Rubin, E. M. & Dubchak, I. VISTA: computational tools for comparative genomics. Nucleic Acids Res.32, W273–W279 (2004). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Nei, M. Phylogenetic analysis in molecular evolutionary genetics. Annu. Rev. Genet.30, 371–403 (1996). [DOI] [PubMed] [Google Scholar]
  • 4.Bellgard, M. I., Itoh, T., Watanabe, H., Imanishi, T. & Gojobori, T. Dynamic evolution of genomes and the concept of genome space. Ann. N. Y. Acad. Sci.870, 293–300 (1999). [DOI] [PubMed] [Google Scholar]
  • 5.Sievers, F. et al. Fast, scalable generation of high-quality protein multiple sequence alignments using Clustal Omega. Mol. Syst. Biol.7, 539 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Katoh, K. & Standley, D. M. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol. Biol. Evol.30, 772–780 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Edgar, R. C. MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res.32, 1792–1797 (2004). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Patiño-Galindo, J. Á et al. Recombination and lineage-specific mutations linked to the emergence of SARS-CoV-2. Genome Med.13, 124 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Vinga, S. Editorial: alignment-free methods in computational biology. Brief. Bioinform.15, 341–342 (2014). [DOI] [PubMed] [Google Scholar]
  • 10.Zielezinski, A., Vinga, S., Almeida, J. & Karlowski, W. M. Alignment-free sequence comparison: benefits, applications, and tools. Genome Biol.18, 1–17 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Bonham-Carter, O., Steele, J. & Bastola, D. Alignment-free genetic sequence comparisons: a review of recent approaches by word analysis. Brief. Bioinform.15, 890–905 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Bernard, G., Chan, C. X. & Ragan, M. A. Alignment-free microbial phylogenomics under scenarios of sequence divergence, genome rearrangement and lateral genetic transfer. Sci. Rep.6, 28970 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Zielezinski, A. et al. Benchmarking of alignment free sequence comparison methods. Genome Biol.20, 144 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Blaisdell, B. E. A measure of the similarity of sets of sequences not requiring sequence alignment. Proc. Natl. Acad. Sci. USA83, 5155–5159 (1986). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Tribus, M. & McIrvine, E. C. Energy and information. Sci. Am.225, 179–190 (1971). [Google Scholar]
  • 16.Otu, H. H. & Sayood, K. A new sequence distance measure for phylogenetic tree construction. Bioinformatics19, 2122–2130 (2003). [DOI] [PubMed] [Google Scholar]
  • 17.Li, M. & Vitányi, P. An Introduction to Kolmogorov Complexity and its Applications, 3rd edn, Vol. 3 (Springer, 2008).
  • 18.Lu, Y.-Y. et al. CAFE: aCcelerated Alignment-FrEe sequence analysis. Nucleic Acids Res.45, W554–W559 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Yu, C. et al. Real time classification of viruses in 12 dimensions. PLoS ONE8, e64328 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Deng, M., Yu, C., Liang, Q., He, R. L. & Yau, S. S.-T. A novel method of characterizing genetic sequences: genome space with biological distance and applications. PLoS ONE6, e17293 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Jeffrey, H. J. Chaos game representation of gene structure. Nucleic Acids Res.18, 2163–2170 (1990). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Randić, M., Novič, M. & Plavšić, D. Milestones in graphical bioinformatics. Int. J. Quantum Chem.113, 2413–2446 (2013). [Google Scholar]
  • 23.Hoang, T. et al. A new method to cluster DNA sequences using Fourier power spectrum. J. Theor. Biol.372, 135–145 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Yin, C., Chen, Y. & Yau, S. S.-T. A measure of DNA sequence similarity by Fourier transform with applications on hierarchical clustering. J. Theor. Biol.359, 18–28 (2014). [DOI] [PubMed] [Google Scholar]
  • 25.Saw, A. K. et al. Alignment-free method for DNA sequence clustering using fuzzy integral similarity. Sci. Rep.9, 3753 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Yu, C., Liang, Q., Yin, C., He, R. L. & Yau, S. S.-T. A novel construction of genome space with biological geometry. DNA Res.17, 155–168 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Yu, H. & Yau, S. S.-T. The optimal metric for viral genome space. Comput. Struct. Biotechnol. J.23, 2083–2096 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Sims, G. E., Jun, S.-R., Wu, G. A. & Kim, S.-H. Alignment-free genome comparison with feature frequency profiles (FFP) and optimal resolutions. Proc. Natl. Acad. Sci. USA106, 2677–2682 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Suwayyid, F. & Wei, G.-W. Persistent Stanley–Reisner theory. Found. Data Sci.8, 287–312 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Hozumi, Y. & Wei, G.-W. Revealing the shape of genome space via k-mer topology. SIAM J. Life Sci.1, 310–332 (2026). [Google Scholar]
  • 31.Sun, N. et al. Geometric construction of viral genome space and its applications. Comput. Struct. Biotechnol. J.19, 4226–4234 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Qi, J., Wang, B. & Hao, B.-I. Whole proteome prokaryote phylogeny without sequence alignment: a k-string composition approach. J. Mol. Evol.58, 1–11 (2004). [DOI] [PubMed] [Google Scholar]
  • 33.Jun, S.-R., Sims, G. E., Wu, G. A. & Kim, S.-H. Whole-proteome phylogeny of prokaryotes by feature frequency profiles: an alignment-free method with optimal feature resolution. Proc. Natl. Acad. Sci. USA107, 133–138 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Tao, K. et al. The biological and clinical significance of emerging SARS-CoV-2 variants. Nat. Rev. Genet.22, 757–773 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Chen, J. & Wei, G. W. Omicron BA.2 (B.1.1.529.2): high potential for becoming the next dominant variant. J. Phys. Chem. Lett.13, 3840–3849 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Li, X., Zhou, T., Feng, X., Yau, S.-T. & Yau, S. S.-T. Exploring geometry of genome space via Grassmann manifolds. Innovation5, 100677 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Khare, S. et al. GISAID’s role in pandemic response. China CDC Wkly.3, 1049–1051 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Letunic, I. & Bork, P. Interactive Tree Of Life (iTOL) v6: recent updates to the phylogenetic tree display and annotation tool. Nucleic Acids Res.52, W78–W82 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Zuo, G. & Hao, B. CVTree3 web server for whole-genome-based and alignment-free prokaryotic phylogeny and taxonomy. Genom. Proteom. Bioinform.13, 321–331 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Ondov, B. D. et al. Mash: fast genome and metagenome distance estimation using minHash. Genome Biol.17, 132 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Wong, T. K. F. et al. IQ-TREE 3: phylogenomic inference software using complex evolutionary models. Mol. Biol. Evol.43, msag117 (2026). [DOI] [PMC free article] [PubMed]
  • 42.Kozlov, A. M., Darriba, D., Flouri, T., Morel, B. & Stamatakis, A. RAxML-NG: a fast, scalable and user-friendly tool for maximum likelihood phylogenetic inference. Bioinformatics35, 4453–4455 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Mirarab, S., Nguyen, N. & Warnow, T. SEPP: saté-enabled phylogenetic placement. in Biocomputing, 247–258 (World Scientific, 2012). [DOI] [PubMed]
  • 44.Su, Z. et al. Topological data analysis and topological deep learning beyond persistent homology: a review. Artif. Intell. Rev.59, 58 (2025). [DOI] [PMC free article] [PubMed]
  • 45.Bruns, W. & Herzog, J.Cambridge Studies in Advanced Mathematics. in Cohen-Macaulay Rings, Vol. 39 (Cambridge University Press, 1998).
  • 46.Edelsbrunner, H. & Harer, J. L. Mathematical Surveys and Monographs. in Computational Topology: An Introduction, Vol. 69 (American Mathematical Society, 2010).
  • 47.Suwayyid, F. et al. Data for CAKR: commutative algebra k-mer representations for genomics. Zenodo 10.5281/zenodo.18757928 (2026). [DOI] [PubMed]
  • 48.Suwayyid, F. et al. CAKR: commutative algebra k-mer representations for genomics, version 1.0.0. Zenodo 10.5281/zenodo.21426540 (2026). [DOI] [PubMed]
  • 49.NIAID Visual & Medical Arts. SARS-CoV-2. NIAID NIH BioArt Source, BioArt image 465; public domain; courtesy of NIAID. bioart.niaid.nih.gov/bioart/465 (2024).

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Reporting Summary (2.5MB, pdf)
Source Data (52.1KB, xlsx)

Data Availability Statement

The genomic sequence data analyzed in this study were obtained from NCBI, including GenBank and the NCBI Virus resource, and from GISAID for the SARS-CoV-2 dataset. GenBank accession.version identifiers are provided in Supplementary Tables 8–13, and the GISAID accession identifiers for the 44 SARS-CoV-2 genomes are provided in Supplementary Table 7. The underlying GISAID sequences and metadata are not redistributed because of GISAID access restrictions and must be obtained directly from GISAID by registered users in accordance with its terms of use. The curated and processed datasets, excluding the underlying GISAID files, are available from Zenodo at https://doi.org/10.5281/zenodo.1875792847. Additional data are provided in the Supplementary Information and Source Data file. Further supporting materials not included in these files are available from the corresponding author upon request. Source data are provided with this paper.

The source code used to implement the CAKR framework and perform the comparative analyses reported in this study is publicly available on GitHub at https://github.com/FaisalSuwayyid/CAKL. The version used in this study, CAKR v1.0.0, has been archived on Zenodo at https://doi.org/10.5281/zenodo.2142654048.


Articles from Nature Communications are provided here courtesy of Nature Publishing Group

RESOURCES