Skip to main content
Briefings in Functional Genomics logoLink to Briefings in Functional Genomics
. 2026 Sep 28;25:elag008. doi: 10.1093/bfgp/elag008

Supervised clustering of bacterial promoter identifies two groups with different relevant positions at −10

Paulo Cambranis-Boldo 1, Gustavo Sganzerla Martinez 2,3, André Borges Farias 4, Ernesto Perez-Rueda 5,✉
PMCID: PMC13617569  PMID: 42803614

Abstract

Promoters are DNA sequences responsible for the specific recognition of the transcriptional machinery in all the biological systems. In bacteria, two main groups of promoters have been described, the σ70-like family, which includes the housekeeping σ70 and the alternative σ38, σ32, σ28, and σ24; and the σ54 family. However, promoter sequences may differ even inside the same family, and this classification may not fully capture the functional diversity of groups. In this work, we explored and classified a collection of bacterial promoters associated with six σ factor families in the bacterium Escherichia coli K-12 using a Supervised clustering workflow that uses Shapley values. From this analysis, we identified two subgroups of sequences differentiated by the conservation of the −9 to −7 positions upstream the Transcription Start Site, suggesting that this region may be employed in classification frameworks. Finally, a detailed analysis of σ70 promoter sequences identify five clusters with a conserved signal in the −10, but with different sequence composition. In summary, these signatures could provide more insights into the sigma factor-promoter interaction and the separation of the DNA double-strand.

Keywords: bacterial promoters, sigma factors, machine learning, supervised clustering, E. coli K-12

Introduction

Transcription regulation is a central process for the overall gene expression of bacterial organisms, where RNA polymerase (RNApol) associated with sigma (σ) factors and promoter sequences play a pivotal role. Promoters are DNA sequences located upstream of the transcription start site (TSS) and are specifically recognized by the RNApol σ factors in bacterial transcription [1]. The specific recognition of σ factors towards the DNA binding site ensures the proper mRNA synthesis in response to diverse environmental conditions [2–5]. In general, promoters have been broadly classified into two major groups [2, 4, 6]: the σ70-like family, which includes the housekeeping σ70 and the alternative σ38, σ32, σ28, and σ24; and the σ54 family. However, promoter sequences may differ even inside the same family [1, 2, 7], and this classification may not fully capture the sequence and functional diversity of groups. An alternative approach to grouping promoters involves analyzing their physicochemical properties, such as DNA duplex stability (DDS) [8, 9]. Such a strategy has been successfully used in other models which report accurate prediction of sequence function [10–12] including characterization and prediction of transcription factor binding sites (TFBSs) [13]. Furthermore, it has been described that variability in the promoter sequence, their information content, as well as presence or absence of motif elements and spacers between them, modulate promoter activity [14–17].

The DDS signatures of promoters associated with different sigma families in the bacterium Escherichia coli K-12 display a distinctive pattern: a local instability located around 10 base pairs upstream the TSS [8]. This feature arises from the presence of the −10 motif in the canonical σ70 promoter, also known as the Pribnow box (consensus sequence TATAAT), which is rich in thymine and adenine [3]. Additionally, other motifs, such as the −10 extended, −35, and discriminator motifs, are embedded within some promoter sequences in this family [3, 4]. For the σ54 promoters, two conserved motifs at −12 and −24 positions have been identified [18]. Features like DDS, which are rooted in the fundamental physicochemical properties of DNA, may serve as suitable candidates for representing promoter sequences in classification frameworks [19].

Machine learning (ML) has emerged as a powerful tool for the identification of bacterial promoters, TFBS, and transcription factors (TFs), among other tasks in bioinformatics. In this regard, a recent analysis to characterize and predict promoters employs an XGBoost algorithm to detect the distinct duplex stability energy signature [20]; whereas, Gradient Boosting, hybrid convolutional/long and short memory network models, have been implemented for the detection and classification of long terminal repeat sequences in plants [21]; and ML-based transcriptome analysis have also been used to infer the relative activity of gene regulatory Modulons in Salmonella Typhimurium [22]. Convolutional neural network (CNN) models, trained on TF binding changes in the promoters, can predict the changes in splicing patterns in metazoa [23]; whereas BertSNR integrates sequence-level and token-level information to identify TFBS at single-nucleotide resolution has already been described [24]. The application of ML workflows to biological data is opening a new era in which vast datasets generated by sequencing and “omics” technologies can be leveraged to enhance our understanding of biological systems.

Therefore, biological data has been explored with unsupervised ML algorithms as well. In this regard, DNA sequence data represents a challenge to identify patterns due to its high dimensionality and low variance of some features. An alternative to this heterogeneity is supervised clustering, which leverages labeled data and has several advantages over the traditional unsupervised clustering, such as learning subclasses and enhancing classification tasks [25]. In this context, Shapley values, which is a framework for providing interpretability of ML-based models, was implemented for the discovery of subgroups [26], measure the contribution of a specific variable [27–29], allowing a deeper understanding of model predictions and providing insights into the importance of features and their local and global contributions in classification tasks.

In this work, we explore and classify a collection of bacterial promoters from six sigma factor families in E. coli K-12 using a Supervised clustering workflow, suggesting that the signatures identified could provide more insights into the σ factor-promoter interaction and the separation of the DNA double-strand.

Materials and methods

Promoter sequences used for analysis

A total of 3217 promoter sequences associated with the six sigma families (σ70, σ38, σ32, σ28, σ24, and σ54) and related with 1901 genes of E. coli K-12 were downloaded from RegulonDB v12.0 [30]. In general, the two hexameric motifs centered at or near −10 and –35, and their variants, such as those promoters overlapping the TSS, those promoters where the −35 is absent, and those promoters with an extension of the −10 or extended promoters, are contained in a region of 81 nucleotides (−60 to +20) relative to the TSS. Therefore, these sequences were edited to have a length size of 81 pb, i.e. 60 positions upstream of the TSS or +1, and 20 downstream. This range of base-pair positions encodes the information needed to classify bacterial promoters [17, 31, 32]. In a second step, we excluded redundancy and those sequences with more than one sigma factor associated, leaving a dataset of 2938 promoters representing 91% of the original dataset (Table 1). Finally, the negative sequences (non-promoter sequences) consist of randomly generated DNA strings with the same length sizes. To this end, a Python script was used, by considering the probability of each base to appear in the original dataset. Figure 1.

Table 1.

Sequences obtained from the RegulonDB included in this work.

Sigma Family Total of sequences Number of sequences filtered*
σ70 1911 1781
σ54 95 90
σ38 219 132
σ32 324 286
σ28 147 138
σ24 521 511
Total 3217 2938

*Those sequences were considered in posterior analysis. The negative datasets are equal in number to positive sequences.

Figure 1.

The pipeline for promoter analysis based on supervised clustering.

Research methodology flowchart.

Sequence parametrization and normalization

DDS was used to encode the sequence information into a machine-readable format [19, 32]. This encoding was performed by assigning a DDS value to every dinucleotide in the sequences being converted, using a Python (version 3.11.6) script. The rationale behind this conversion considers the number of hydrogen bonds between the nucleotides, two in the adenine (A) and thymine (T), and three in the cytosine (C) and guanine (G), affecting the chemical stability of the DNA molecule [19]. The conversion of a string of length i results in a vector of size i-1 that represents each sequence. An example of this process is shown in Supplementary material S1. Additionally, DDS data was normalized using the min-max algorithm available in Scikit-learn Python library (version ​​1.6.1). Therefore, the conversion of promoter DNA sequences into DDS depicts the AT rich elements identified in promoter sequences [8]. Figure 1. While alternative encoding strategies (including one-hot encoding, k-mer frequencies, or physicochemical property matrices) are compatible with the downstream pipeline, DDS was selected because it encodes a continuous biophysical property that has direct mechanistic relevance to promoter function, as local duplex stability governs DNA strand separation during transcription initiation. Consequently, Shapley importance scores derived from DDS-encoded features can be interpreted in terms of thermodynamic contributions at specific sequence positions, providing a biologically grounded explanation for the clustering patterns observed, an interpretive level that positional identity encodings such as one-hot do not inherently afford.

Model training and evaluation

Random Forest (RF) algorithms (Python with Scikit-learn 1.6.1, with parameters detailed in Supplementary Material S2) were selected for implementation due to their low computational cost, and because the post-hoc analysis by SHAP Python library has an optimized method for tree-based models known as TreeSHAP (see below). Therefore, the implementation of Shapley value calculation is exact, whereas others (KernelSHAP, DeepSHAP) are only approximations. In this regard, tree-based models have been also used successfully to classify promoters using the DDS values as features [8]. The RF algorithms were trained using the following strategies:

  • a) Positives-vs-negatives approach: the model was trained to distinguish promoter sequences (positive dataset) from non-promoter sequences (negative dataset) that were generated randomly using a Python script, which gave equal likelihood to each nucleotide to appear, meant to act as proxies for non-promoter regions (Table 2).

  • b) One-vs-rest approach: The model was trained to distinguish a family of sigma promoters from the rest by considering a family of interest as a positive and the rest as negatives. For instance, in the first iteration, all sequences labeled as σ70 are considered as positive, while σ54, σ38, σ32, σ28 and σ24 are considered negative. In the next iteration, σ54 would be considered positive and the rest of the families as negative.

Table 2.

Promoter datasets used in this work.

Sigma factor Number of promoters Negative dataset
σ70 1500 promoters 1500 non-promoters
All families 90 promoters from each of the six families 540 non-promoters
One-vs-rest set 90 promoters from each of the six families 450 promoters, i.e. 90 promoters per each of the remaining six families.

Non-promoters sequences were randomly generated, as it was previously described.

Regardless of the strategy, care was given to keep classes balanced. For example, the “All families” dataset has 90 sequences of each promoter family for a total of 540 positive sequences, which were contrasted with 540 negative sequences. Table 2.

To describe the promoter features, the whole dataset was considered. Using this approach, the classification model achieves perfect or near perfect performance, since it “memorizes” the features that distinguish each class. This approach, which does not consider the classic train-test split of the dataset, is necessary to obtain the main results in our workflow. Classification is not the main objective of this work, but rather, to discover new subgroups in promoter sequences classified as true positives. Nevertheless, the models were evaluated following standard ML practices with an 80–20 train-test split, assessing their ability to capture and generalize the underlying feature relationships. Performance metrics, including precision, recall (sensitivity), accuracy, F1 score, and the area under the ROC curve (AUC), were computed using built-in functions from scikit-learn version 1.6.1 (Supplementary materials S3–S4).

Two models were trained for different purposes within this workflow. The SHAP model was trained on the complete dataset without a train-test split, as previously described; this model achieves a near-perfect or perfect classification because it has access to all examples, and its sole purpose is to generate SHAP feature attributions for the subgroup discovery analysis. In this regard, classification performance is not a goal of this model, and its accuracy is not reported as a performance metric. The evaluation model was trained on 80% of the data and tested on the held-out 20%, following standard ML practice. This model was exclusively used to characterize the difficulty of the classification task and to compare the Random Forest against alternative classifiers. Performance metrics such as accuracy, precision, recall, F1 score, and AUC (reported in Supplementary S3–S5) only refer to this evaluation model and should not be interpreted as measures of the quality of the SHAP analysis. These metrics from Random Forest were compared with four other alternative classifiers: Extra Trees, AdaBoost, Gradient Boosting, and Support Vector Machine (with RBFkernel), all implemented in scikit-learn with default hyperparameters.

Model training

To train the Random Forest algorithm, the six sigma families of promoter sequences were considered, as detailed in Table 2. The “all families” and the “One-vs-rest” training sets have a number of sequences that is a factor of 90, since that is the number of sequences available in the σ54 family. On the other hand, the σ70 family is the most numerous, allowing us to test the workflow using only their own sequences to find subclasses, therefore, we selected a set of 1500 sequences to train the model.

Shapley value calculation and analysis

Shapley values were calculated using the SHAP library for Python (version 0.46.0) [29] for the models trained on the“σ70,” “All families,” and “One-vs-rest” sets. A vector of Shapley values is obtained for each sequence used in the training of the model, explaining the model’s decision-making process. Then, the average absolute value per feature was calculated, considering all the sequences in the dataset. The result is an 80-dimensional vector containing mean absolute Shapley values per class. When the training is carried out in a positive versus negative approach, two sets of mean absolute Shapley values are obtained, one for the positives and one for the negatives. However, the positive and negative sets are mirror images of each other and therefore only the set for the positive class (i.e. promoters) was considered. On the other hand, when the model was trained on a one-vs-rest approach, many sets of Shapley values are obtained, one per class. Shapley values represent the importance of each feature for each prediction and are either positive or negative.

Dimensionality reduction and clustering algorithms

Dimensionality reduction was performed on the sets of Shapley values to enable clustering, and then, visualize it in two dimensions. For this purpose, the UMAP algorithm (Python package version 0.5.9) [33] was used due to its capacity to conserve and represent the distance of objects in high dimensional space. The array of Shapley values has the same dimensionality as the original DDS data, i.e. 80 dimensions. UMAP was applied to reduce the dimensionality to 4 dimensions, and then the DBSCAN algorithm (included in the scikit-learn Python package) was used to identify clusters [34–36]. The parameters of DBSCAN were determined empirically by visual inspection of the clusters obtained in dimensionality reduction and held constant throughout the workflow (epsilon was set 1.5 and minimum samples to 10). The ability to define clusters depends on the n-neighbors parameter of the UMAP algorithm, which was determined by iterating over a wide range of values: 1000, 500, 250, 100, 50, 25, and 10 (Supplementary materials S5–S6). At a certain point, decreasing the n-neighbors parameter of the algorithm does little to enhance clustering and resolution. This was the criterium used to select the n-neighbors parameter in the present work.

Statistical tests

To evaluate differences in DDS values among groups, the Kruskal–Wallis test was employed due to the non-parametric distribution of the data. Post hoc comparisons were conducted using Dunn’s test to identify specific group differences. All statistical analyses were performed using the R programming language (version 4.3.3). Variance was computed using the standard formula. A significance threshold of P ≤ .05 was applied for all statistical tests.

Robustness of Shapley values to negative set composition

As the negative datasets in the present work are randomly generated, the question of whether the results presented here are spurious. To confirm the consistency of results, the complete pipeline for the all-families dataset was re-executed five times using a different random seed (seeds 42, 123, 149, 30, and 512) for the generation of the negative sequences (random noise). Algorithms that also required a random seed such as Random Forest and UMAP always used the same seed (seed 37), across runs. For each seed, we computed a Pearson correlation coefficient for the mean absolute SHAP profile of all samples of the run versus the primary analysis (the run that is reported in the present work). Cluster assignment consistency was quantified using the Adjusted Rand Index (ARI) computed on positive sequences only. All analyses were carried out in Python. Additionally, a negative dataset consisting of randomly selected coding regions from the E. coli reference genome was employed, having similar properties to that of randomly generated sequences (Supplementary material S7). For simplicity, and to maintain comparability with similar research in this field [20], this work used only randomly generated negative sequences.

Weblogos and information content analysis

The weblogo and information content were carried out using the Logomaker Python library (0.8.7) [37] with default parameters. Weblogos are representations of Position Weight Matrices, which indicate what DNA bases are expected in certain positions.

Results

Machine learning algorithm selection and evaluation

The performance of all tree-based algorithms was comparable across all datasets and consistently outperformed SVM (Supplementary material S8). The accuracy and AUC improved significantly in the smaller datasets, showing that the All families dataset is a challenging classification task due to the variety of samples within it. As previously mentioned, RF was selected not only for its good performance, but also because the SHAP Python library provides the exact TreeSHAP implementation for Random Forest, in contrast to the approximate kernel-based methods required for SVM and boosting ensembles. The results hereon are derived from RF.

Calculation of DNA duplex stability and Shapley values

To determine the DDS as a signature for the six different promoter families in E. coli, an analysis of their variance of DDS is displayed in Fig. 2. From this analysis, we identified that all the σ70 families exhibit low relative variance downstream of −10; whereas σ54 has a conserved DDS value in −12. A second peak that could correspond with the −35 emerges prominently in σ28, and σ32 promoters. Additional energy patterns around the −10, i.e. a drop-then-rise, in σ28 and σ32 promoters, or a single peak, such as in the σ70 and σ54 promoters, were identified.

Figure 2.

Duplex Stability (DDS) as a signature for the six different promoter families in E. coli.

Mean DDS values of six σ promoter families. Deep blue color represents positions where DDS values were the least variables. In these positions, most families displayed variance values lower than 0.15 DDS units. In yellow, the positions where the DDS was less conserved are indicated. See text for details.

DDS are features calculated directly from the DNA sequence of the promoters while Shapley values are calculated from an ML model trained on these sequences. The present workflow depends on the calculation of Shapley values to be able to distinguish clusters (Supplementary material S9). To use Shapley values to gain insight from data, the model that is used to calculate them was trained on the whole dataset.

The alignment of the Shapley values with the most prominent features of the DDS plot is noticeable in Fig. 3. For every positive-vs-negative dataset, the peak in Shapley values (importance) coincides with the peak of DDS at positions −11, −10 and −9 (Fig. 3a) as well as −8, −7 and −6 (Fig. 3b). This finding indicates that a high value (less negative) in DDS around this position is an important feature for the model to assign a true positive class.

Figure 3.

Alignment of the Shapley values with the most prominent features of the DDS.

Mean DDS values with variation (blue and yellow) and mean absolute Shapley values (red) calculated from datasets of  Table 2. For all sets, the −10 was the most important for classification of (a) σ70 family, and (b) all σ families.

We further hypothesized that important segments of the promoter for the biological function of the sequence have small variations in their DDS energy signatures. The plots of mean DDS in Fig. 3, show variable color depending on the normalized variance in the sequences. From this, we identified some relevant peaks that emerge when variations are relatively small. For all the datasets, the region with the least variation corresponds to the −10 motif, a feature that the models deem important for assigning a positive label.

The Shapley values calculated in this workflow depend on the negative dataset used to train the model. Despite this, our analysis of Shapley value sensitivity indicates that DDS signals in promoters have sufficiently distinct features such that no matter the seed used to generate the negatives, the same features are used to distinguish clusters, and the Shapley values are strongly correlated. Details on this analysis are available in Supplementary materials S10–S11.

Clustering of promoter signals

To identify new subgroups in the datasets, a DBSCAN clustering algorithm was implemented. Although in many cases the clusters are distinguishable by eye, an algorithm is needed to assign labels and further explore data. DBSCAN detected and labeled clusters that are sufficiently separated from each other (Fig. 4). The grouping that the sequences display is only due to the proximity in their Shapley values in a high dimensional space. The naming convention for the discovered clusters was C0, C1, C2, and so on.

Figure 4.

DBSCAN clusters promoters identificacion.

Clusters identified in the (a) σ70 and (b) all families datasets. The data points are colored according to which family they belong, with the gray data points indicating non-promoters. The circle represents the cluster to which the points were assigned, according to the DBSCAN algorithm. Clusters made up of only negative sequences were left out of further analysis.

As the variety of sigma families increased in the datasets, from only σ70 to all families, we expected to find at least one cluster per family. Figure 4 shows that positive sequences were well distanced from negative ones, both σ70 + σ54, and in all the families. In both assays, the positive sequences were clustered into two separate groups. Further analysis of these clusters revealed that the features in which they differ as detected by the models are those with the highest mean absolute Shapley values, which were between the −10 and −5 positions.

The labels obtained by the DBSCAN algorithm were used to visualize the DDS signature of each cluster, displayed as a sequence logo (Figs 5 and 6).

Figure 5.

Cluster identification based on the DDS signature.

Cluster C0 (P or Peak) and C1 (V or valley) found in the all-families dataset. The mean DDS signature of P promoters has a peak from position −9 to −5 that has low variability. Furthermore, this cluster has a high information content in position −7 and − 6. In contrast, V promoters have a lower information content and no DDS peak.

Figure 6.

Motifs and DDS values identificated in the σ70 promoter dataset.

DDS and motifs of clusters discovered in the σ70 promoter dataset. The weblogo is obtained from the sequences (positive only) in each cluster, and it is aligned with the DDS signature within the same cluster. The normalized variance coincides with the position with the most information content. A conserved G/T pair is responsible for a drop in the DDS plot.

From the dataset that included all families, two clusters were identified: a cluster rich in AT/TA, between the −9 and −5 positions with up to 1 bit of information, and a second cluster where the information content is much lower at 0.1 bits in that position (Fig. 5). We named these subgroups of promoters as P (for the peak in energy) and V (for a valley). The AT/TA dinucleotides are the pair with the least DDS energy, which manifests in the plot as a peak with low variance. The promoters without this characteristic have, in turn, more stable values on average.

In relation to the σ70 sequences, five clusters were identified. Figure 6 summarizes the DDS signatures and the sequence logos. In this regard, σ70 promoter sequences were clustered depending on the presence and position of a conserved C/G pair. This feature causes the mean DDS signal in the −10 motif to have a particular shape, for instance, Cluster C0 lacks the C/G feature at all, causing the signal to be a simple peak with the majority TAAA motif. In contrast, cluster C1 has the C/G just upstream of the T/A rich region, causing the mean DDS to be lower, then increase in the subsequent positions. Cluster C2 and Cluster C4 are close together but separate from other clusters (Fig. 4a) due to their −10 value being low, however, Cluster C1 and Cluster C4 in that the sequence logo is similar although shifted one position to the left, and the DDS signature being a dip, then a peak.

One-vs-rest analysis

The One-vs-rest analysis enforces the by-family classification of sigma promoter sequences to the ML model. The model is tasked with differentiating each family from the rest, and thus the Shapley values highlight the more important features for this task (Fig. 7).

Figure 7.

Family classification of sigma promoter by Shapley values.

Shapley values of all six sigma families. These sets of Shapley values were obtained from a one-vs-all classification task and represent the importance of each position to distinguish that class from all others. The localization of high Shapley value close −10 suggests that families possess a characteristic DDS value in this region of biological significance.

Shapley values peaked at or around position −10 across all families. Additionally, a second peak was observed at position −24 in σ54 promoters. Dimensionality reduction applied to these sets of Shapley values resulted in the separation of each family (Supplementary material S12). For some families, such as σ70, σ54, and σ32, this process revealed the presence of two clusters. The primary determinant of cluster separation was the characteristics of the −10 motif Supplementary Material S13, S14 & S15). The cases of σ70 and σ32 further support our earlier findings in the “All Families” dataset, highlighting the existence of promoters with either high or low DDS values at position −10. In contrast, for σ54, the separation was driven by the conservation of the −24 motif, indicating the presence of both conserved and degenerate σ54 promoters.

Considering the task of computationally predicting the family of a new sigma promoter, we explored how Shapley values can help in determining the most important features for classification. For all classification tasks, we noticed that Shapley values pointed to the position with critical biological activity (−10 motif, discriminator motif), and we hypothesized that these features hold differences in their values that may allow for classification of sequences. The Krustal-Wallis test was performed on all the features and the value of the chi squared statistic was found to have a correlation with the mean absolute Shapley value of the position, thus reinforcing the notion that Shapley values point to the values that are different between the families (Supplementary material S16). We tested the values of DDS in these positions using Dunn’s test and found that most families present significant differences in their distributions (Supplementary material S17).

Discussion and conclusions

Shapley values, a concept from game theory, point to features of biological significance when inferred from a model that has been trained to distinguish promoter from non-promoter sequences. The region spanning from −10 to the TSS is important in vivo due to being the site where the DNA double-strand is dissociated to allow for the RNApol complex to initiate transcription, a process mediated by the σ factor. As the −35 is inessential to initiate transcription in some cases [38], the −10 is the most conserved feature among all promoter sequences in E. coli, and thus, our model has placed the highest Shapley values to this region.

It has been demonstrated that the greatest contributor to the variance of promoter strength is the −10 motif [7]. The clusters discovered in this work are most likely to have differences in their strength (activity), as they differ mostly in the −10 box. Likeness to the canonical −10 motif (TATAAT) could suggest a strong promoter, and we expect the discovered P promoters to be more conserved than the V subgroups. In this regard, the low structural stability (high DDS values), around −10 (P/V-like pattern), suggest that this region is central for the transcription initiation that involve isomerization of polymerase and separation of the promoter DNA around the transcription start; i.e. they could be associated with the open complex formation. Indeed, the −10 element (−12TATAAT−7) is recognized as both double-stranded DNA for the T:A bp at position −12 and as nontemplate, single-stranded DNA from positions −11 to −7. The single-stranded sequences at positions −11 to −7 as well as the −5 contribute to later steps in transcription initiation that involve isomerization of polymerase and separation of the promoter DNA around the TSS. Therefore, the double-stranded elements may be used in various combinations to yield an effective promoter [39].

The σ70 subgroups discovered in our clustering analysis differ in their DDS signature. Some groups have a highly conserved and stable micro motifs followed by unstable base pairs, while others present more complex combinations. These signatures could provide more insights into the sigma factor-promoter interaction and the separation of the DNA double-strand. Furthermore, we suggest that the mechanics of this interaction must be different between these new subgroups.

In this regard, we compared the promoter activity in the clusters for the case of σ70 against experimental information recently described [40], as it is the dataset that yields the most subclusters. From this analysis, we identified that (in average) 20% of the promoters per cluster are active; and the cluster 6, followed by the cluster 8 and 2, contains the largest amount of active promoters with the greatest similarity to the −10 consensus motif, suggesting that they are available to be specifically recognized and transcribed by σ70 (Supplementary material S18). Therefore, P and V values could suggest different activity, however further experimental evidence is necessary.

In the one-vs-rest analysis, the σ54 family stood out, as the Shapley values concentrated in position −24, where the important CG motif, associated with specific σ54 binding, is found (Supplementary material S19). It seems that when the algorithm is tasked with classifying promoters into each family, the Shapley values calculated from it point to the most critical features for classification. It has been observed that some families have a very well conserved −35 motif, while others rely mostly on −10 motif for activation of transcription. Finally, we identified distinct peaks of the −35 signal in families σ28 and σ32; however, Shapley values were low in these positions; probably due to the high variation in the −35 hexamer. This leaves the hexamer at −10 as the most important for distinguishing each family. Indeed, position −8, where the discriminator would be, is where all six families are most statistically different.

The n-neighbors parameter of UMAP controls the trade-off between local and global structure in the embedding and, consequently, the granularity at which clusters are resolved. It does not control its composition. Across the range explored in Supplementary materials S5–S6, the same sequences consistently co-occurred within nested groupings: small values of n-neighbors emphasize local neighborhoods and partition the embedding into a larger number subclusters, while larger values progressively merge these subclusters into broader groups without reassigning their members. The cluster assignments are therefore stable in composition and differ only in the level of subdivision displayed. Because the present analysis targets the most general organization of the sequence space, we report results at the value of n-neighbors that yield the smallest stable number of clusters; smaller values can be used to inspect substructure within these groups but do not alter the conclusions drawn here.

Finally, the information encoded in DDS properties to classify and discover new groups within a dataset of promoter sequences from E. coli K-12. Using DDS values instead of promoter sequence, we were able to identify a new classification for promoters: Peak and Valley promoter families. The workflow leverages the patterns deemed relevant by a ML algorithm to discover new subgroups.

Key Points

  • We explored and classified a collection of bacterial promoters associated with six sigma factor families in the bacterium Escherichia coli K-12 using a supervised clustering workflow.

  • We found that Shapley values could indicate biological activity in sequences, as well as indicating the features that are more relevant for distinguishing different subgroups.

  • We identified two subgroups of promoters differentiated by the conservation of the −9 to −7 positions, relative to the TSS.

  • These signatures could provide more insights into the DNA-promoter interaction and the separation of the DNA double-strand.

Supplementary Material

Supplementary_material_elag008

Acknowledgements

We acknowledge Israel Sanchez-Dominguez, Israel Josué Novelo Zel, and Raul Galindo for technical support. There was no additional external funding received for this study. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

Paulo Cambranis-Boldo is a graduate student at Universidad Nacional Autónoma de México, focuses on bioinformatics, machine learning, and data science.

Dr. Gustavo Sganzerla Martinez is Postdoctoral Fellow at Dalhousie University, and specialized in bioinformatics, machine learning, and data science.

Dr. André Borges Farias is Postdoctoral fellow at Laboratório Nacional de Computação Científica – LNCC-Brazil, specializing in structural biology, bioinformatics, machine learning, and data science.

Dr. Ernesto Pérez-Rueda is Professor at Universidad Nacional Autónoma de México, specializing in gene regulation, bioinformatics, machine learning, and data science.

Contributor Information

Paulo Cambranis-Boldo, Instituto de Investigaciones en Matemáticas Aplicadas y en Sistemas, Unidad Académica del Estado de Yucatán, Universidad Nacional Autónoma de México, Mérida, Yucatán, 97302, México.

Gustavo Sganzerla Martinez, Microbiology and Immunology, Dalhousie University, 5850 College Street, Halifax, Nova Scotia, B3H4H7, Canada; Department of Basic Medical Sciences, Shantou University Medical College, Shantou, 515041, Guangdong, China.

André Borges Farias, Laboratório Nacional de Computação Científica - LNCC, Avenida Getúlio Vargas, Petrópolis, Rio de Janeiro 25651075, Brazil.

Ernesto Perez-Rueda, Instituto de Investigaciones en Matemáticas Aplicadas y en Sistemas, Unidad Académica del Estado de Yucatán, Universidad Nacional Autónoma de México, Mérida, Yucatán, 97302, México.

Author contributions

Paulo Cambranis-Boldo (Formal analysis [equal], Investigation [equal], Methodology [equal], Writing—original draft [equal], Writing—review & editing [equal]), Gustavo Sganzerla Martinez (Formal analysis [equal], Investigation [equal], Writing—review & editing [equal]), André Borges Farias (Methodology [equal], Writing—original draft [equal], Writing—review & editing [equal]), Ernesto Perez Rueda (Formal analysis [equal], Funding acquisition [equal], Investigation [equal], Methodology [equal], Resources [equal], Supervision [equal], Writing—original draft [equal], Writing—review & editing [equal]).

Funding

This work was supported by the Dirección General de Asuntos del Personal Académico-Universidad Nacional Autónoma de México [IN-206326 to E. P-R.]. P.C-B was supported by the Secretaría de Ciencia, Humanidades, Tecnología e Innovación (SECIHTI) with a Master’s scholarship (4067674).

Conflicts of interest

GSM is a shareholder of the company BioForge Canada Limited. BioForge Canada Limited is a company that uses bioinformatics in immunological approaches to the monitoring, prevention, and treatment of infectious diseases. The author disclose sthat the interests of BioForge Canada Limited had no impact in this study.

Data availability

The research data are available as Supplementary materials S1–S16.

References

  • 1.Hertz GZ, Stormo GD. Escherichia coli promoter sequences: analysis and prediction. Methods Enzymol 1996;273:30–42. 10.1016/S0076-6879(96)73004-5. [DOI] [PubMed] [Google Scholar]
  • 2. Browning  DF, Busby  SJW. The regulation of bacterial transcription initiation. Nat Rev Microbiol  2004;2:57–65. 10.1038/nrmicro787. [DOI] [PubMed] [Google Scholar]
  • 3. Feklístov  A, Sharon  BD, Darst  SA. et al.  Bacterial sigma factors: a historical, structural, and genomic perspective. Annu Rev Microbiol  2014;68:357–76. 10.1146/annurev-micro-092412-155737. [DOI] [PubMed] [Google Scholar]
  • 4. Paget  M. Bacterial sigma factors and anti-sigma factors: structure, function and distribution. Biomolecules  2015;5:1245–65. 10.3390/biom5031245. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Lee  DJ, Minchin  SD, Busby  SJW. Activating transcription in bacteria. Annu Rev Microbiol  2012;66:125–52. 10.1146/annurev-micro-092611-150012. [DOI] [PubMed] [Google Scholar]
  • 6. Bush  M, Dixon  R. The role of bacterial enhancer binding proteins as specialized activators of σ54-dependent transcription. Microbiol Mol Biol Rev  2012;76:497–529. 10.1128/mmbr.00006-12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. LaFleur  TL, Hossain  A, Salis  HM. Automated model-predictive design of synthetic promoters to control transcriptional profiles in bacteria. Nat Commun  2022;13:5159. 10.1038/s41467-022-32829-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Martinez  GS, De Ávila  S, Silva  E. et al.  DNA structural and physical properties reveal peculiarities in promoter sequences of the bacterium Escherichia coli K-12. SN Appl Sci  2021;3:740. 10.1007/s42452-021-04713-2. [DOI] [Google Scholar]
  • 9. Martinez  GS, Pérez-Rueda  E, Sarkar  S. et al.  Machine learning and statistics shape a novel path in archaeal promoter annotation. BMC Bioinformatics  2022;23:171. 10.1186/s12859-022-04714-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Rangannan  V, Bansal  M. Identification and annotation of promoter regions in microbial genome sequences on the basis of DNA stability. J Biosci  2007;32:851–62. 10.1007/s12038-007-0085-1. [DOI] [PubMed] [Google Scholar]
  • 11. Shahmuradov  IA, Mohamad Razali  R, Bougouffa  S. et al.  bTSSfinder: a novel tool for the prediction of promoters in cyanobacteria and Escherichia coli. Bioinformatics  2017;33:334–40. 10.1093/bioinformatics/btw629. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. De Avila  S, Silva  E, Echeverrigaray  S. et al.  BacPP: bacterial promoter prediction—a tool for accurate sigma-factor specific assignment in enterobacteria. J Theor Biol  2011;287:92–9. 10.1016/j.jtbi.2011.07.017. [DOI] [PubMed] [Google Scholar]
  • 13. Borges Farias  A, Sganzerla Martinez  G, Galán-Vásquez  E. et al.  Predicting bacterial transcription factor binding sites through machine learning and structural characterization based on DNA duplex stability. Brief Bioinform  2024;25:bbae581. 10.1093/bib/bbae581. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Bharanikumar  R, Premkumar  KAR, Palaniappan  A. PromoterPredict: sequence-based modelling of Escherichia coli σ70 promoter strength yields logarithmic dependence between promoter strength and sequence. PeerJ  2018;6:e5862. 10.7717/peerj.5862. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Lagator  M, Sarikas  S, Steinrueck  M. et al.  Predicting bacterial promoter function and evolution from random sequences. eLife  2022;11:e64543. 10.7554/eLife.64543. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Deal  C, De Wannemaeker  L, De Mey  M. Towards a rational approach to promoter engineering: understanding the complexity of transcription initiation in prokaryotes. FEMS Microbiol Rev  2024;48:fuae004. 10.1093/femsre/fuae004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Shultzaberger  RK, Chen  Z, Lewis  KA. et al.  Anatomy of Escherichia coli σ70 promoters. Nucleic Acids Res  2007;35:771–88. 10.1093/nar/gkl956. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Barrios  H, Valderrama  B, Morett  E. Compilation and analysis of sigma(54)-dependent promoter sequences. Nucleic Acids Res  1999;27:4305–13. 10.1093/nar/27.22.4305. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Friedel  M, Nikolajewa  S, Sühnel  J. et al.  DiProDB: a database for dinucleotide properties. Nucleic Acids Res  2009;37:D37–40. 10.1093/nar/gkn597. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Martinez  GS, Perez-Rueda  E, Kumar  A. et al.  CDBProm: the comprehensive directory of bacterial promoters. NAR Genomics Bioinforma  2024;6:lqae018. 10.1093/nargab/lqae018. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Horvath  J, Jedlicka  P, Kratka  M. et al.  Detection and classification of long terminal repeat sequences in plant LTR-retrotransposons and their analysis using explainable machine learning. BioData Min  2024;17:57. 10.1186/s13040-024-00410-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Park  JY, Jang  M, Lee  S-M. et al.  Unveiling the novel regulatory roles of RpoD-family sigma factors in salmonella Typhimurium heat shock response through systems biology approaches. PLoS Genet  2024;20:e1011464. 10.1371/journal.pgen.1011464. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Lin  T-C, Tsai  C-H, Shiau  C-K. et al.  Predicting splicing patterns from the transcription factor binding sites in the promoter with deep learning. BMC Genomics  2024;25:830. 10.1186/s12864-024-10667-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Luo  H, Tang  L, Zeng  M. et al.  BertSNR: an interpretable deep learning framework for single-nucleotide resolution identification of transcription factor binding sites based on DNA language model. Bioinforma Oxf Engl  2024;40:btae461. 10.1093/bioinformatics/btae461. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Eick CF, Zeidat N, Zhao Z. Supervised clustering – algorithms and benefits. In: Proceedings of the 16th IEEE International Conference on Tools with Artificial Intelligence (ICTAI 2004), Boca Raton, FL, Los Alamitos, CA, USA: IEEE Computer Society. 2004, 774–6. 10.1109/ICTAI.2004.111 [DOI] [Google Scholar]
  • 26. Cooper  A, Doyle  O, Bourke  A. Supervised Clustering for Subgroup Discovery: An Application to COVID-19 Symptomatology, in Machine Learning and Principles and Practice of Knowledge Discovery in Databases, vol. 1525, Kamp  M, Koprinska  I, Bibal  A, Bouadi  T, Frénay  B, Galárraga  L, Oramas  J, Adilova  L, Graça  G. et al., Eds., in Communications in Computer and Information Science, vol. 1525., Cham: Springer International Publishing, 2021, pp. 408–22. 10.1007/978-3-030-93733-1_29. [DOI] [Google Scholar]
  • 27. Merrick  L, Taly  A. The Explanation Game: Explaining Machine Learning Models Using Shapley Values, in Machine Learning and Knowledge Extraction, vol. 12279, Holzinger  A, Kieseberg  P, Tjoa  AM, Weippl  E, Eds., in Lecture Notes in Computer Science, vol. 12279., Cham: Springer International Publishing, 2020, pp. 17–38. 10.1007/978-3-030-57321-8_2. [DOI] [Google Scholar]
  • 28.Rozemberczki B, Watson L, Bayer P et al. The Shapley value in machine learning. In: De Raedt L (ed). Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence (IJCAI-22), Vienna, Austria: International Joint Conferences on Artificial Intelligence Organization. 2022, 5572–9. 10.24963/ijcai.2022/778 [DOI] [Google Scholar]
  • 29. Lundberg  SM, Lee  S-I. A unified approach to interpreting model predictions. Proc 31st Int Conf Neural Inf Process Syst  2017;30;4765–74. [Google Scholar]
  • 30. Salgado  H, Gama-Castro  S, Lara  P. et al.  RegulonDB v12.0: a comprehensive resource of transcriptional regulation in E. Coli K-12. Nucleic Acids Res  2023;52:D255–64. 10.1093/nar/gkad1072. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Lloréns-Rico  V, Lluch-Senar  M, Serrano  L. Distinguishing between productive and abortive promoters using a random forest classifier in mycoplasma pneumoniae. Nucleic Acids Res  2015;43:3442–53. 10.1093/nar/gkv170. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. SantaLucia  J, Hicks  D. The thermodynamics of DNA structural motifs. Annu Rev Biophys Biomol Struct  2004;33:415–40. 10.1146/annurev.biophys.32.110601.141800. [DOI] [PubMed] [Google Scholar]
  • 33.McInnes L, Healy J, Saul N et al. UMAP: uniform manifold approximation and projection. J Open Source Softw 2018;3:861. 10.21105/joss.00861. [DOI] [Google Scholar]
  • 34. Burr  T, Mitchell  J, Kolb  A. et al.  DNA sequence elements located immediately upstream of the −10 hexamer in Escherichia coli promoters: a systematic study. Nucleic Acids Res  2000;28:1864–70. 10.1093/nar/28.9.1864. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Barne  KA, Bown  JA, Busby  SJ. et al.  Region 2.5 of the Escherichia coli RNA polymerase sigma70 subunit is responsible for the recognition of the ‘extended-10’ motif at promoters. EMBO J  1997;16:4034–40. 10.1093/emboj/16.13.4034. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Kumar  A, Malloch  RA, Fujita  N. et al.  The minus 35-recognition region of Escherichia coli sigma 70 is inessential for initiation of transcription at an ‘extended minus 10’ promoter. J Mol Biol  1993;232:406–18. 10.1006/jmbi.1993.1400. [DOI] [PubMed] [Google Scholar]
  • 37. Tareen  A, Kinney  JB. Logomaker: beautiful sequence logos in python. Bioinformatics  2020;36:2272–4. 10.1093/bioinformatics/btz921. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Ester  M, Kriegel  H-P, Sander  J. et al.  A density-based algorithm for discovering clusters in large spatial databases with noise, in Proceedings of the Second International Conference on Knowledge Discovery and Data Mining, in KDD’96. Portland, Oregon: AAAI Press, 1996, pp. 226–31. [Google Scholar]
  • 39. Hook-Barnard  IG, Hinton  DM. Transcription initiation by mix and match elements: flexibility for polymerase binding to bacterial promoters. Gene Regul Syst Biol  2007;1:275–93. 10.1177/117762500700100020. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Urtecho  G, Insigne  KD, Tripp  AD. et al.  Genome-wide Functional Characterization of Escherichia coli Promoters and Sequence Elements Encoding Their Regulation. 2023;12:RP92558. 10.7554/eLife.92558.1. [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary_material_elag008

Data Availability Statement

The research data are available as Supplementary materials S1–S16.


Articles from Briefings in Functional Genomics are provided here courtesy of Oxford University Press

RESOURCES