Skip to main content
BMC Genomics logoLink to BMC Genomics
. 2025 Oct 24;26:957. doi: 10.1186/s12864-025-12162-z

Machine learning prediction of bacterial optimal growth temperature from protein domain signatures reveals thermoadaptation mechanisms

Haida Liu 1,#, Geyi Zhu 1,#, Lijuan Chen 2, Hui Ye 3, Yunhua Zhang 4,, Guomin Han 1,3,
PMCID: PMC12553261  PMID: 41136894

Abstract

Cultivating the vast majority of uncultured microbes requires knowledge of their physiological preferences, particularly optimal growth temperature (OGT). We present a machine learning approach that utilizes protein domain frequencies from bacterial genomes to predict OGT across a wide continuous range (1–83 °C). Our Random Forest model, trained on a dataset of 1,498 genomes, achieved high predictive accuracy (R²=0.853 on test data, 82.4% of predictions within a ± 10 °C error margin), substantially advancing current capabilities and offering a practical tool to guide cultivation experiments. Analysis of the model identified key protein domain signatures associated with thermal adaptation. The enrichment of domains related to polyamine metabolism, the tRNA methyltransferase family, and CRISPR-Cas systems was positively correlated with higher OGTs, providing genomic evidence for their roles in thermotolerance. Conversely, domains involved in redox homeostasis, transport, and nucleic acid binding were more abundant at lower temperatures. These findings not only facilitate targeted cultivation efforts but also deepen our understanding of the molecular strategies bacteria employ to thrive across diverse thermal niches.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12864-025-12162-z.

Keywords: Comparative genomics, Optimal growth temperature (OGT), Machine learning, Predictive modeling, Protein domains, Bacterial thermoadaptation

Introduction

Temperature is a critical environmental factor that significantly influences microbial growth and reproduction. Temperature affects these processes by modulating enzymatic activities, cell membrane fluidity, nutrient uptake, and ion exchange [1]. Extremes of temperature, both high and low, can lead to enzyme denaturation, membrane disruption, impaired nutrient acquisition, and ion imbalance, ultimately inhibiting microbial proliferation [1]. While many microorganisms thrive within moderate temperature ranges [2], certain species exhibit remarkable adaptations to extreme thermal conditions. For instance, Psychromonas ingrahamii maintains viability at temperatures as low as −12 °C [3], whereas Methanopyrus kandleri strain 116 can proliferate at temperatures reaching 122 °C [4]. Microbial organisms are broadly categorized into psychrophiles, mesophiles, and thermophiles based on their optimal growth temperature ranges. The differential response of microorganisms to temperature variations is largely attributed to genome-encoded molecular adaptations within cellular components, including lipids, nucleic acids, and proteins [510]. These adaptations ensure the functional integrity of macromolecules across diverse thermal environments. Specifically, cold adaptation often involves increased unsaturated fatty acid content to maintain membrane fluidity, while thermophilic adaptations include higher proportions of saturated fatty acids and enhanced DNA stability [6, 8, 9, 11].

The advent of high-throughput genomics has created a stark imbalance: while genomic data expands exponentially, our knowledge of microbial phenotypes, including optimal growth temperature (OGT), remains comparatively scarce. This gap hinders efforts to cultivate the vast “unculturable majority” and limits our ability to understand microbial contributions to biogeochemical cycles and biotechnology. Consequently, predicting microbial phenotypes directly from genomic data has emerged as a major goal in modern microbiology [12].

Predicting microbial physiological traits from genomes is a rapidly developing field. Various genomic features, such as codon usage [13], GC content [14], k-mer distributions [15], CRISPR-Cas content [16], and protein domain frequencies [17, 18], have been combined with multiple machine learning models (e.g., Random Forest, Support Vector Machines, Convolutional Neural Networks) and proven effective for predicting phenotypes like pH preference [19], oxygen requirements [20], growth rate (minimal doubling time) [13], and growth temperature [15]. For instance, tools based on codon usage bias, such as gRodon [13] and its improved hybrid version Phydon [21], can accurately estimate growth rates and have been used to predict this trait for as many as 111,349 microbial species. Hu et al. (2022) suggested a positive correlation between the OGT of prokaryotes and their genomic GC content [14]. However, another study argued that a higher GC content in structural RNAs and an increased proportion of specific amino acids in proteins correlated with higher growth temperatures, reflecting the need for enhanced molecular stability [22]. Based on k-mer distributions, Wang et al. (2022) developed CnnPOGTP, a CNN model that captures genomic compositional features without gene annotation, achieving high-accuracy OGT prediction [15]. Liu et al. (2023) observed that predicted minimal doubling times are positively correlated with the spacer number and other parameters of CRISPR-Cas systems, such as array number and Cas gene cluster number [16].

Protein domain frequencies have also been used to predict other bacterial traits, such as motility, oxygen requirement, or spore formation [18]. Similarly, Ramoneda et al. (2023) developed a predictive framework based on protein domain frequencies by first inferring bacterial pH preferences from large-scale biogeographical data and then using these inferred phenotypes to identify a specific set of 56 functional genes whose presence or absence correlated with acid- or alkali-tolerance [19]. Recently, Koblitz et al. (2025) utilized machine learning models to accurately predict multiple bacterial phenotypes, including oxygen preference and temperature [23]. Our team has also accurately predicted the three primary oxygen preference types in bacteria using machine learning, and preliminary experimental validation has suggested that four candidate sequences possess the identified gene functions [20]. Previous studies have explored various genomic features, such as ribosomal RNA sequences, GC content, and k-mer distributions, to predict bacterial growth temperatures [2427]. However, these approaches either achieve high-precision OGT prediction based on k-mer distributions or simplify OGT prediction into a classification task (e.g., psychrophile vs. mesophile), without simultaneously conducting an in-depth and detailed investigation into the underlying molecular mechanisms.

In this study, we address the limitations of existing methods by constructing a predictive model based on the protein domain composition within bacterial genomes. Our primary objectives are to (i) develop an accurate, continuous-variable prediction model, (ii) identify the specific protein domain signatures most strongly associated with temperature adaptation, and (iii) provide new insights into the molecular basis of thermotolerance.

Methods

Bacterial optimal growth temperature dataset

We utilized a comprehensive database of bacterial optimal growth temperatures curated by Sato et al. (2020) and accessible at https://togodb.org/db/tempura [25]. This database, generated through manual literature review and cross-referencing, provided the ground truth for our machine learning approach. Additionally, we supplemented this dataset with over 60,000 temperature phenotypes queried from the BacDive database [28].

Genome data acquisition and processing

As of August 2025, the National Center for Biotechnology Information (NCBI) RefSeq database, accessible via its FTP server (ftp://ftp.ncbi.nlm.nih.gov/genomes/refseq/bacteria/assembly_summary.txt.), contained approximately 430,000 bacterial reference genomes [29]. By matching species names against our curated dataset of organisms with known optimal growth temperatures (OGTs), we selected 1,498 bacterial samples. These selected samples spanned an OGT range of 1 °C to 83 °C, and their corresponding protein sequences were subsequently downloaded using custom scripts.

Protein sequences from the 1,498 genomes were annotated for domain content using pfam_scan.pl against the Pfam-A hidden Markov model (HMM) database (version 33.0) [30]. The command used was: pfam_scan.pl -fasta input.fa -dir [Pfam-A.hmm_database_location] > pfam.outfile. The frequency of each annotated protein domain was quantified per genome to construct a protein domain frequency matrix, which served as the feature dataset. The structure of this dataset is detailed in Supplementary Table 1.

Machine learning model training and selection

The dataset was partitioned into training set (75%) and a testing set (25%) using a 3:1 ratio. To select the optimal algorithm, a suite of models (e.g., xgboost, svm, kknn, rvm, ctree, cubist, cv.glmnet, rpart and Random Forest) were trained and evaluated on the training set using 10-fold cross-validation. Based on its superior performance (see Results), the Random Forest algorithm was selected for the final model.

The Random Forest model was implemented with key hyperparameters set to balance performance and efficiency. The number of trees (ntree) was set to 1000 to ensure model stability and reliable feature importance metrics while maintaining computational tractability. For post-hoc analysis, feature importance was calculated (importance = TRUE) to rank domain contributions, and the sample proximity matrix was generated (proximity = TRUE). These latter two parameters do not influence the predictive performance of the model.

Model evaluation

The predictive performance of the final trained model was assessed on the independent, held-out test set. We calculated Pearson’s correlation coefficient (r), its associated p-value, and the coefficient of determination (R²) to quantify the correlation between predicted and observed OGTs. As a complementary metric, the percentage of test set samples with predicted temperatures falling within ± 10 °C of the actual temperatures was calculated. Furthermore, to evaluate the model’s ability to predict variations in optimal growth temperature among different strains of the same species, we compiled an additional test set. This set was created by screening genomic and sequence databases to identify strains of the same species with differing reported optimal or cultivation temperatures, which were then used for a targeted prediction assessment.

Analysis of molecular mechanisms underlying bacterial temperature adaptation

Following model selection, feature importance analysis, incorporating cross-validation techniques, was performed to identify key protein domains associated with bacterial temperature adaptation. Protein domains with high importance scores were selected for further investigation into their roles in determining optimal growth temperature preferences.

First, to mitigate biases from non-unique OGT values, the feature vectors of samples with identical OGTs were averaged, following Zhang et al. (2018) [31]. Second, to account for phylogenetic non-independence, we applied the Phylogenetic Generalized Least Squares (PGLS) method [32]. A phylogenetic tree was constructed from 16 S rDNA sequences using Clustal Omega (v1.2.4) [33] and FastTree (v2.2) [34], which was then used in R to calculate Pagel’s lambda and the phylogenetically corrected p-values for the correlations.

Inter-group variation analysis

Principal Coordinates Analysis (PCoA) was employed to visualize the dissimilarity or similarity among data samples based on their protein domain composition. Bray-Curtis distances were calculated using the ‘vegan’ R package [35]. To statistically test for significant differences between thermal groups, Analysis of Similarities (ANOSIM) was also performed using the ‘vegan’ package.

Results

Performance of the random forest model for OGT prediction

A Random Forest model trained on a dataset of 1,498 bacterial samples demonstrated a statistically significant correlation between predicted and actual OGTs (Pearson’s r = 0.826, p = 6.83e-88, R² = 0.853; Fig. 1A). The practical accuracy of the model was evaluated on the independent test set across different error tolerances. 82.4% of predictions fell within a ± 10 °C margin and 55.9% within a ± 5 °C margin of the observed values (Fig. 1B, Supplementary Table 2). The corresponding accuracies on the training set were 79.8% and 53.4%, respectively.

Fig. 1.

Fig. 1

Performance Evaluation of the Random Forest Model for Optimal Growth Temperature (OGT) Prediction Note:A Scatter plot assessing the correlation between predicted and observed OGT on the independent test set. B Distribution of prediction errors (Predicted OGT - Observed OGT) for the model on the independent test set

Evaluation of the model’s performance on strains of the same species with different optimal or cultivation temperatures revealed that the predicted temperatures among these strains generally varied by less than 1 °C (Supplementary Table 3). This indicates that while the protein domain-based approach can accurately predict OGTs over a broad inter-species range, it has limited sensitivity for resolving finer-scale temperature preference differences that may exist between closely related strains.

Comparison of multiple machine learning algorithms

We further evaluated and compared the performance of eight different machine learning algorithms in addition to Random Forest. Scatter plots for each model show the relationship between observed and predicted values, with a smoothed curve indicating the trend (Fig. 2A). All models demonstrated a positive correlation, suggesting they could capture underlying patterns in the data to some extent, with R² values ranging from 0.66 to 0.823. Among them, CTree, Cubist, and CV showed better fitting effects, with their data points more concentrated around the trend line and relatively higher R² values, the highest being CV with an R² of 0.823. However, the Random Forest model’s R² of 0.853 was still superior to all other tested models. The extremely high statistical significance (p = 6.83e-88) further confirms the model’s powerful and reliable predictive capability for OGT. Consequently, the Random Forest model was selected for subsequent analyses.

Fig. 2.

Fig. 2

Comparative Performance Analysis of Different Machine Learning Algorithms for OGT Prediction Note: A Scatter plots evaluating the correlation between predicted and observed OGT for various algorithms on the independent test set. B Comparison of prediction accuracy across algorithms on the test set. Accuracy is defined as the percentage of samples predicted within a ± 10 °C error margin. Algorithm abbreviations: xgboost, eXtreme Gradient Boosting; svm, Support Vector Machine; kknn, k-Nearest Neighbors; rvm, Relevance Vector Machine; ctree, Conditional Inference Tree; cubist, Cubist algorithm; cv.glmnet, Elastic Net Regularization; rpart, Recursive Partitioning and Regression Trees; RandomForest, Random Forest algorithm

Identification of key protein domains via feature importance

To identify protein domains most predictive of OGT, feature selection was performed using the Random Forest model. Ten-fold cross-validation with the rfcv function in R indicated that a subset of 19 protein domains yielded the minimum cross-validation error, representing the most informative features (Fig. 3A). Feature importance was ranked based on the percentage increase in mean squared error when a feature was permuted. Domains such as ABC_tran_Xtn, HRDC, and SelR were among the highest-ranking features, indicating a strong association with bacterial OGT prediction (Fig. 3B).

Fig. 3.

Fig. 3

Identification of Key Protein Domain Features using the Random Forest Model Note: A Ten-fold cross-validation curve illustrating model error as a function of the number of features included, ranked by importance. The minimum indicates the optimal feature set size (19). B Feature importance ranking of the top 19 protein domains based on the percentage increase in Mean Squared Error (%IncMSE) upon permutation. C PGLS-corrected heatmap visualizing the relationship between the abundance of key domains and OGT across the temperature gradient

Correlation between protein domain abundance and optimal growth temperature

The relative abundance of the top 19 most important protein domains was analyzed across the OGT spectrum. Phylogenetically corrected correlation analysis (PGLS) revealed distinct and significant patterns (Fig. 3C, Supplementary Fig. 1). Ten domains (ABC_tran_Xtn, HRDC, SelR, DEAD, Sulfate_transp, UPF0176_N, YjdM, RQC, RF3_C, VF530) exhibited a significant negative correlation with OGT, being more abundant in psychrophiles and mesophiles. Conversely, nine domains (including AdoMet_dc, Spermine_synt_N, DUF2905, Spermine_synth, Radical_SAM, Cas_Cas6, RAMPs, GCD14, UPF0004) showed a significant positive correlation with OGT, being more prevalent in thermophiles. Functional annotations (Table 1) revealed that domains enriched at high temperatures are associated with polyamine synthesis (Spermine_synt_N, Spermine_synth), CRISPR-Cas systems (Cas_Cas6, RAMPs). Domains enriched at lower temperatures are involved in processes such as protein repair (SelR, DEAD), nucleic acid binding (HRDC), and transport (ABC_tran_Xtn).

Table 1.

Functions of protein domains with high contribution scores

graphic file with name 12864_2025_12162_Tab1_HTML.jpg

Domain_name, protein domain name; Putative function, domain function; Red font indicates protein domains and their functions that may be related to high temperatures.

Protein domain composition differs across thermal groups

The overall variation in protein domain composition among bacteria adapted to different thermal environments was assessed. Calculation of Bray-Curtis distances indicated greater dissimilarity between thermophilic samples and both psychrophilic and mesophilic samples, compared to the distance between psychrophiles and mesophiles (Supplementary Fig. 2). Analysis of Similarities (ANOSIM) confirmed that the protein domain composition differed significantly between the thermal groups (R = 0.588, p = 0.001; Fig. 4A). Principal Coordinates Analysis (PCoA) based on Bray-Curtis distances visually supported this finding, showing a distinct separation of the thermophile cluster from the more closely associated psychrophile and mesophile clusters (Fig. 4B). These results demonstrate that adaptation to different temperature regimes corresponds to significant shifts in the overall protein domain profile of bacterial genomes.

Fig. 4.

Fig. 4

Analysis of Microbial Community Differentiation Based on Protein Domain Composition Across Thermal Groups Note: A Analysis of Similarities (ANOSIM) results comparing Bray-Curtis dissimilarities within and between the defined thermal groups. The ANOSIM R statistic (0.588) and p-value (0.001) indicate significant differentiation. B Principal Coordinates Analysis (PCoA) plot based on Bray-Curtis dissimilarities calculated from protein domain frequencies, illustrating the separation of samples from different thermal groups (psychrophile, mesophile, thermophile)

Discussion

A significant fraction of microbial life remains uncultured, estimated at ~ 87% of bacterial/archaeal genera [36], hindering our understanding of their physiology and ecological roles. Cultivation often relies on laborious trial-and-error, as optimal conditions are unknown. While advanced cultivation strategies exist [37], predicting key parameters like OGT directly from genomic data could substantially increase cultivation success rates and accelerate microbial discovery. The increasing availability of genomic sequences provides the necessary data foundation for developing such predictive models.

Previous studies have demonstrated the potential of using genomic features to predict bacterial phenotypes. In the context of OGT, various approaches have been developed. For example, the work of Koblitz et al. (2025) simplified the prediction to a classification problem (“thermophile” or not) rather than a precise value [22]. In contrast, Wang et al. (2022) developed CnnPOGTP, a Convolutional Neural Network (CNN) model that achieved high accuracy in predicting continuous OGTs (R² = 0.891 and 0.880 on validation and test sets, respectively) using annotation-free k-mer features [15]. Sauer and Wang (2019) also achieved accurate results (R² = 0.835 on the test set) using a model trained on over 1,000 genome-derived features, including GC content, k-mer frequencies, and amino acid composition [26]. Our study, which develops a model based on protein domain frequencies, has also achieved accurate OGT prediction. Our Random Forest model achieved a high accuracy (R² = 0.853) on an independent test set across a broad range (1–83 °C), with 82.4% of predictions falling within a ± 10 °C error margin. Our predictive accuracy is comparable to that of CnnPOGTP and slightly superior to the results of Sauer and Wang (2019). However, our evaluation of the model on different strains of the same species revealed that it performs poorly in predicting intra-species OGT differences, likely due to the scarcity of such data in the training set. Nevertheless, considering that the typical growth breadth of many bacteria is around 30 °C and the OGT is often not the midpoint [2], the 82.4% accuracy within a ± 10 °C margin suggests our predictions are sufficiently precise to effectively guide initial cultivation attempts for uncharacterized species.

The choice of machine learning algorithm significantly impacted prediction performance, with notable differences observed between methods. In contrast, Random Forest demonstrated superior performance, accurately capturing the continuous OGT distribution with the highest correlation and accuracy metrics on the test set, evidenced by a well-fitted regression line and narrow confidence intervals (Fig. 2). Notably, xgboost performed well in a previous pH prediction study [19], underscoring the task-specific nature of algorithm suitability and the importance of evaluating multiple methods.

Beyond prediction, our model allowed identification of protein domains strongly associated with OGT. Using feature importance analysis, we identified 19 domains whose genomic abundance significantly correlated (positively or negatively) with bacterial OGT (Fig. 3). These specific domain abundance shifts contribute to a larger pattern: the overall protein domain composition significantly differs between thermal groups, with thermophiles distinctly separating from the more closely related psychrophiles and mesophiles (ANOSIM R = 0.588, p = 0.001; Fig. 4). This aligns with established knowledge that thermal adaptation involves genome-level molecular changes [9, 10]. Domains positively correlated with OGT included those involved in polyamine biosynthesis (Spermine_synt_N, Spermine_synth, AdoMet_dc), tRNA modification (GCD14), and CRISPR-Cas systems (Cas_Cas6, RAMPs). Conversely, domains negatively correlated with OGT included those related to redox/protein homeostasis (SelR, DEAD), transport (ABC_tran_Xtn), and nucleic acid binding (HRDC). Notably, our findings partially overlap with those of Koblitz et al. (2025), who also identified domains like Spermine/spermidine synthase, CRISPR associated protein Cas1, and S-adenosylmethionine decarboxylase as important features for distinguishing thermophiles from non-thermophiles, although they did not further explore their potential functions [23].

We acknowledge that our initial machine learning approach did not account for the phylogenetic structure of the species, which may have biased the results. However, based on a reviewer’s suggestion, we introduced PGLS correction in the downstream analysis of domain importance, which likely enhances the reliability of our findings. A positive correlation was observed between the copy number of spermine synthase (Spermine_synt_N, Spermine_synth domains) genes and OGT, suggesting increased thermotolerance is associated with a higher capacity for spermine synthesis. Polyamines are crucial for various cellular processes, and hyperthermophiles, similar to their archaeal counterparts, require long-chain and branched polyamines for macromolecular stability at high temperatures [38]. The concurrent positive correlation of genes containing these domains and S-adenosylmethionine decarboxylase (AdoMet_dc domain), a key enzyme in polyamine biosynthesis [5], strongly supports the hypothesis that enhanced polyamine metabolism is a significant genomic signature of adaptation to high temperatures.

Nucleic acids are highly susceptible to denaturation and misfolding at high temperatures, severely impacting RNA functional stability. Thermophilic and hyperthermophilic microorganisms have evolved a suite of synergistic mechanisms to mitigate these challenges, including extending mRNA half-life, optimizing rRNA/tRNA tertiary structures, and utilizing polyamines for stabilization [3941]. Among these strategies, chemical modification of tRNA is particularly critical for stabilizing local conformations and maintaining translational efficiency. GCD14 is a subunit of the tRNA methyltransferase complex required for the 1-methyladenosine modification and maturation of initiator methionyl-tRNA [42]. Extensive research has established the central role of tRNA methylation in thermoadaptation. For example, the absence of the m¹A58 methyltransferase TrmI in Thermus thermophilus leads to a growth defect at 80 °C [43], and TrmH-catalyzed Gm18 modification requires polyamines to maintain substrate RNA conformation at high temperatures [44]. Our analysis revealed a positive correlation between the number of GCD14 domain members in bacterial genomes and their thermotolerance, collectively suggesting that the tRNA methyltransferase family, particularly members containing the GCD14 domain, is a key factor in the evolution of microbial adaptation to high temperatures.

Furthermore, we observed a positive correlation between bacterial thermotolerance and the abundance of genes encoding several domains associated with CRISPR-Cas domains (Cas_Cas6, RAMPs). These systems defend against viruses and mobile genetic elements [45]. Cas6, a key endoribonuclease for crRNA maturation, belongs to the RAMP superfamily, which includes many Cas proteins also found enriched in thermophiles in our analysis [46]. Intriguingly, the activity of some Cas enzymes, like Cas6 from the hyperthermophile Pyrococcus furiosus, is strongly temperature-dependent [46]. A large-scale study found that ~ 90% of thermophiles (vs. 45% of mesophiles) possess complete CRISPR-Cas systems, suggesting a strong positive correlation [16, 47, 48]. Another study showed that various Cas proteins in Thermus thermophilus are upregulated in response to salt and oxidative stress, implying a broader role in stress response [48]. Collectively, this suggests that the expansion of CRISPR-Cas systems may not only be related to genome stability and defense but may also play a more general role in survival under the extreme stress of high-temperature environments. Of course, correlation does not always imply causation; this association may reflect the ecological context of high-temperature environments rather than a direct thermotolerance strategy. This finding provides important clues for exploring the non-canonical functions of CRISPR-Cas in thermoadaptation, a question that warrants further experimental investigation.

Conclusion

This study demonstrates that a Random Forest model utilizing genomic protein domain frequencies can accurately predicts bacterial optimal growth temperature across a wide range (1–83 °C) with high fidelity (R²=0.853 on test data; 82.4% accuracy within ± 10 °C), offering a valuable tool to guide cultivation efforts for uncultured microbes. Feature analysis revealed distinct protein domain profiles linked to thermal adaptation, providing genomic evidence reinforcing the roles of polyamines and tRNA modification pathways in thermotolerance, while also highlighting a significant and intriguing association between CRISPR-Cas systems and life at high temperatures.

Supplementary Information

Authors’ contributions

H.L.: Investigation, Methodology, Writing – original draft. G.Z.: Investigation, Methodology, Writing – original draft. L.C.: Data curation, Funding acquisition. H.Y.: Validation. Y.Z.: Project administration, Supervision, Writing – review & editing. G.H.: Conceptualization, Funding acquisition, Methodology, Resources, Writing – review & editing. All authors read and approved the final manuscript.

Funding

This work was supported by grants from the National Natural Science Foundation of China (No. 32172769), the Germplasm Resource Bank Construction Project of the Anhui Provincial Department of Agriculture and Rural Affairs, and the Joint Research Project for Elite Maize Varieties of Anhui Province.

Data availability

All key data, analysis scripts, and important results, including feature importance lists and test set predictions, generated during this study are publicly available in the GitHub repository: [https://github.com/gmh007/Bacterial-Optimal-Growth-Temperature-Prediction].

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Haida Liu and Geyi Zhu are contributed equally to this work.

Contributor Information

Yunhua Zhang, Email: yunhua9681@163.com.

Guomin Han, Email: guominhan@ahau.edu.cn.

References

  • 1.Ginez LD, Osorio A, Vázquez-Ramírez R, Arenas T, Mendoza L, Camarena L, et al. Changes in fluidity of the E. coli outer membrane in response to temperature, divalent cations and polymyxin-B show two different mechanisms of membrane fluidity adaptation. FEBS J. 2022;289:3550–67. 10.1111/febs.16358. [DOI] [PubMed] [Google Scholar]
  • 2.Tortora GJ, Funke BR, Case CL. Microbiology: An Introduction. 13th ed. Pearson; 2019.
  • 3.Auman AJ, Breezee JL, Gosink JJ, Kämpfer P, Staley JT. Psychromonas ingrahamii sp. nov., a novel gas vacuolate, psychrophilic bacterium isolated from Arctic Polar sea ice. Int J Syst Evol Microbiol. 2006;56:1001–7. 10.1099/ijs.0.64068-0. [DOI] [PubMed] [Google Scholar]
  • 4.Takai K, Nakamura K, Toki T, Tsunogai U, Miyazaki M, Miyazaki J, et al. Cell proliferation at 122 degrees C and isotopically heavy CH4 production by a hyperthermophilic methanogen under high-pressure cultivation. Proc Natl Acad Sci U S A. 2008;105:10949–54. 10.1073/pnas.0712334105. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Bauvois C, Jacquamet L, Huston AL, Borel F, Feller G, Ferrer J-L. Crystal structure of the cold-active aminopeptidase from colwellia psychrerythraea, a close structural homologue of the human bifunctional leukotriene A4 hydrolase. J Biol Chem. 2008;283:23315–25. 10.1074/jbc.M802158200. [DOI] [PubMed] [Google Scholar]
  • 6.Berry ED, Foegeding PM. Cold temperature adaptation and growth of microorganisms †. J Food Prot. 1997;60:1583–94. 10.4315/0362-028X-60.12.1583. [DOI] [PubMed] [Google Scholar]
  • 7.Coker JA. Recent advances in understanding extremophiles. F1000Res. 2019;8:F1000 Faculty Rev-1917. 10.12688/f1000research.20765.1.
  • 8.Collins T, Margesin R. Psychrophilic lifestyles: mechanisms of adaptation and biotechnological tools. Appl Microbiol Biotechnol. 2019;103:2857–71. 10.1007/s00253-019-09659-5. [DOI] [PubMed] [Google Scholar]
  • 9.Jamroze A, Perugino G, Valenti A, Rashid N, Rossi M, Akhtar M, et al. The reverse gyrase from pyrobaculum calidifontis, a novel extremely thermophilic DNA topoisomerase endowed with DNA unwinding and annealing activities. J Biol Chem. 2014;289:3231–43. 10.1074/jbc.M113.517649. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Suutari M, Laakso S. Microbial fatty acids and thermal adaptation. Crit Rev Microbiol. 1994;20:285–328. 10.3109/10408419409113560. [DOI] [PubMed] [Google Scholar]
  • 11.Siliakus MF, van der Oost J, Kengen SWM. Adaptations of archaeal and bacterial membranes to variations in temperature, pH and pressure. Extremophiles. 2017;21:651–70. 10.1007/s00792-017-0939-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Karlsen ST, Rau MH, Sánchez BJ, Jensen K, Zeidan AA. From genotype to phenotype: computational approaches for inferring microbial traits relevant to the food industry. FEMS Microbiol Rev. 2023;47:fuad030. 10.1093/femsre/fuad030. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Weissman JL, Hou S, Fuhrman JA. Estimating maximal microbial growth rates from cultures, metagenomes, and single cells via codon usage patterns. Proc Natl Acad Sci U S A. 2021;118:e2016810118. 10.1073/pnas.2016810118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Hu E-Z, Lan X-R, Liu Z-L, Gao J, Niu D-K. A positive correlation between GC content and growth temperature in prokaryotes. BMC Genomics. 2022;23:110. 10.1186/s12864-022-08353-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Wang S, Li G, Liao Z, Cao Y, Yun Y, Su Z, et al. CnnPOGTP: a novel CNN-based predictor for identifying the optimal growth temperatures of prokaryotes using only genomic k-mers distribution. Bioinformatics. 2022;38:3106–8. 10.1093/bioinformatics/btac289. [DOI] [PubMed] [Google Scholar]
  • 16.Liu Z-L, Hu E-Z, Niu D-K. Investigating the relationship between CRISPR-cas content and growth rate in bacteria. Microbiol Spectr. 2023;11:e0340922. 10.1128/spectrum.03409-22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Jensen DB, Vesth TC, Hallin PF, Pedersen AG, Ussery DW. Bayesian prediction of bacterial growth temperature range based on genome sequences. BMC Genomics. 2012;13(Suppl 7):3. 10.1186/1471-2164-13-S7-S3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Lingner T, Mühlhausen S, Gabaldón T, Notredame C, Meinicke P. Predicting phenotypic traits of prokaryotes from protein domain frequencies. BMC Bioinformatics. 2010;11:481. 10.1186/1471-2105-11-481. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Ramoneda J, Stallard-Olivera E, Hoffert M, Winfrey CC, Stadler M, Niño-García JP, et al. Building a genome-based understanding of bacterial pH preferences. Sci Adv. 2023;9:eadf8998. 10.1126/sciadv.adf8998. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Wan S, Liu H, Zhu G, Geng Y, Li W, Chen L, et al. Decoding oxygen preference: machine learning discovers functional genes in bacteria. Genomics. 2025;111095. 10.1016/j.ygeno.2025.111095. [DOI] [PubMed] [Google Scholar]
  • 21.Xu L, Zakem E, Weissman JL. Improved maximum growth rate prediction from microbial genomes by integrating phylogenetic information. Nat Commun. 2025;16:4256. 10.1038/s41467-025-59558-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Hurst LD, Merchant AR. High guanine–cytosine content is not an adaptation to high temperature: a comparative analysis amongst prokaryotes. Proc R Soc Lond B. 2001;268:493–7. 10.1098/rspb.2000.1397. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Koblitz J, Reimer LC, Pukall R, Overmann J. Predicting bacterial phenotypic traits through improved machine learning using high-quality, curated datasets. Commun Biol. 2025;8:897. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Edirisinghe JN, Goyal S, Brace A, Colasanti R, Gu T, Sadhkin B et al. Machine Learning-Driven phenotype predictions based on genome annotations. 2023. 10.1101/2023.08.11.552879.
  • 25.Sato Y, Okano K, Kimura H, Honda K. TEMPURA: database of growth temperatures of usual and rare prokaryotes. Microbes Environ. 2020;35:ME20074. 10.1264/jsme2.ME20074. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Sauer DB, Wang D-N. Predicting the optimal growth temperatures of prokaryotes using only genome derived features. Bioinformatics. 2019;35:3224–31. 10.1093/bioinformatics/btz059. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.van den Elzen A, Helena-Bueno K, Brown CR, Chan LI, Melnikov SV. Ribosomal proteins can hold a more accurate record of bacterial thermal adaptation compared to rRNA. Nucleic Acids Res. 2023;51:8048–59. 10.1093/nar/gkad560. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Schober I, Koblitz J, Sardà Carbasse J, Ebeling C, Schmidt ML, Podstawka A, et al. BacDive in 2025: the core database for prokaryotic strain data. Nucleic Acids Res. 2025;53:D748–56. 10.1093/nar/gkae959. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Sayers EW, Cavanaugh M, Frisse L, Pruitt KD, Schneider VA, Underwood BA, et al. GenBank 2025 update. Nucleic Acids Res. 2025;53:D56–61. 10.1093/nar/gkae1114. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Mistry J, Chuguransky S, Williams L, Qureshi M, Salazar GA, Sonnhammer ELL, et al. Pfam: the protein families database in 2021. Nucleic Acids Res. 2021;49:D412–9. 10.1093/nar/gkaa913. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Zhang J, Zhang N, Liu Y-X, Zhang X, Hu B, Qin Y, et al. Root microbiota shift in rice correlates with resident time in the field and developmental stage. Sci China Life Sci. 2018;61:613–21. 10.1007/s11427-018-9284-4. [DOI] [PubMed] [Google Scholar]
  • 32.Freckleton RP, Harvey PH, Pagel M. Phylogenetic analysis and comparative data: a test and review of evidence. Am Nat. 2002;160:712–26. 10.1086/343873. [DOI] [PubMed] [Google Scholar]
  • 33.Sievers F, Higgins DG. Clustal omega for making accurate alignments of many protein sequences. Protein Sci. 2018;27:135–45. 10.1002/pro.3290. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Price MN, Dehal PS, Arkin AP. Fasttree 2–approximately maximum-likelihood trees for large alignments. PLoS ONE. 2010;5:e9490. 10.1371/journal.pone.0009490. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Dixon P. VEGAN, a package of R functions for community ecology. J Veg Sci. 2003;14:927–30. 10.1111/j.1654-1103.2003.tb02228.x. [Google Scholar]
  • 36.Lloyd KG, Steen AD, Ladau J, Yin J, Crosby L. Phylogenetically novel uncultured microbial cells dominate Earth microbiomes. mSystems. 2018;3:e00055–18. 10.1128/mSystems.00055-18. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Lewis WH, Tahon G, Geesink P, Sousa DZ, Ettema TJG. Innovations to culturing the uncultured microbial majority. Nat Rev Microbiol. 2021;19:225–40. 10.1038/s41579-020-00458-8. [DOI] [PubMed] [Google Scholar]
  • 38.Michael AJ. Polyamine function in archaea and bacteria. J Biol Chem. 2018;293:18693–701. 10.1074/jbc.TM118.005670. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Jegousse C, Yang Y, Zhan J, Wang J, Zhou Y. Structural signatures of thermal adaptation of bacterial ribosomal RNA, transfer RNA, and messenger RNA. PLoS ONE. 2017;12:e0184722. 10.1371/journal.pone.0184722. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Paz A, Mester D, Baca I, Nevo E, Korol A. Adaptive role of increased frequency of polypurine tracts in mRNA sequences of thermophilic prokaryotes. Proc Natl Acad Sci U S A. 2004;101:2951–6. 10.1073/pnas.0308594100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Hori H, Kawamura T, Awai T, Ochi A, Yamagami R, Tomikawa C, et al. Transfer RNA modification enzymes from thermophiles and their modified nucleosides in tRNA. Microorganisms. 2018;6:110. 10.3390/microorganisms6040110. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Anderson J, Phan L, Cuesta R, Carlson BA, Pak M, Asano K, et al. The essential Gcd10p–Gcd14p nuclear complex is required for 1-methyladenosine modification and maturation of initiator methionyl-tRNA. Genes Dev. 1998;12:3650–62. 10.1101/gad.12.23.3650. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Guelorget A, Roovers M, Guérineau V, Barbey C, Li X, Golinelli-Pimpaneau B. Insights into the hyperthermostability and unusual region-specificity of archaeal Pyrococcus Abyssi tRNA m1A57/58 methyltransferase. Nucleic Acids Res. 2010;38:6206–18. 10.1093/nar/gkq381. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Hori H, Terui Y, Nakamoto C, Iwashita C, Ochi A, Watanabe K, et al. Effects of polyamines from thermus thermophilus, an extreme-thermophilic eubacterium, on tRNA methylation by tRNA (Gm18) methyltransferase (TrmH). J Biochem. 2016;159:509–17. 10.1093/jb/mvv130. [DOI] [PubMed] [Google Scholar]
  • 45.Li T, Yang Y, Qi H, Cui W, Zhang L, Fu X, et al. CRISPR/Cas9 therapeutics: progress and prospects. Signal Transduct Target Ther. 2023;8:36. 10.1038/s41392-023-01309-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Sokolowski RD, Graham S, White MF. Cas6 specificity and CRISPR RNA loading in a complex CRISPR-cas system. Nucleic Acids Res. 2014;42:6532–41. 10.1093/nar/gku308. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Weinberger AD, Wolf YI, Lobkovsky AE, Gilmore MS, Koonin EV. Viral diversity threshold for adaptive immunity in prokaryotes. mBio. 2012;3:e00456–12. 10.1128/mBio.00456-12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Lan X-R, Liu Z-L, Niu D-K. Precipitous increase of bacterial CRISPR-cas abundance at around 45° C. Front Microbiol. 2022;13:773114. 10.3389/fmicb.2022.773114. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

All key data, analysis scripts, and important results, including feature importance lists and test set predictions, generated during this study are publicly available in the GitHub repository: [https://github.com/gmh007/Bacterial-Optimal-Growth-Temperature-Prediction].


Articles from BMC Genomics are provided here courtesy of BMC

RESOURCES