Skip to main content
Communications Chemistry logoLink to Communications Chemistry
. 2025 Nov 7;8:340. doi: 10.1038/s42004-025-01717-6

Machine learning-driven prediction of substrates for enzymes introducing or removing protein post-translational modifications

Nashira H Ridgeway 1, Anand Chopra 1, Valentina Lukinović 1, Michal Feldman 2, François Charih 1,3, Dan Levy 2, James R Green 3, Kyle K Biggar 1,
PMCID: PMC12594832  PMID: 41203923

Abstract

The exploration of post-translational modifications (PTMs) within the proteome is pivotal for advancing our understanding of disease and the function of cancer therapeutics. However, identifying genuine sites of PTMs introduced or removed by an enzyme of interest amid numerous candidates is challenging. We present a machine learning (ML)-driven search method, which combines ML with enzyme-mediated modification of complex peptide arrays to predict unexplored PTM sites for an enzyme of interest. Experimental validation confirmed that this approach correctly predicted 37-43% of proposed PTM sites, unveiling candidate sites of the methyltransferase SET8 and the deacetylases SIRT1-7. Our approach marks an important performance increase over traditional in vitro methods across separate enzyme classes. Mass spectrometry analysis confirmed the dynamic methylation status of several predicted SET8 substrates, and the deacetylation of 64 unique sites identified for SIRT2. This method has also revealed changes in SET8-regulated substrate network among breast cancer missense mutations, collectively revealing insight into differential enzyme function in disease. By disentangling the substrate features that dictate PTM-inducing enzyme specificity, this approach demonstrates potential in uncovering enzyme-substrate networks within PTM pathways.

Subject terms: Enzymes, Cheminformatics, Computational chemistry


Understanding post-translational modifications (PTMs) is crucial for advancing disease research and cancer therapeutics, yet identifying specific enzyme-induced PTM sites remains challenging. Here, the authors introduce a machine learning-driven method combined with high-throughput peptide array synthesis, significantly enhancing PTM site prediction accuracy and revealing insights into enzyme-substrate networks, with implications for cancer research.

Introduction

In the current era, in which the human genome1 has been decoded for nearly two decades, unraveling the functional intricacies of the vast majority of human proteins remains a mystery. This challenge predominantly arises from the influence of post-translational modifications (PTMs), which are reversible chemical alterations with the potential to profoundly shape a function of a modified protein. With more than 500 distinct PTMs identified to date, the functional proteome transcends the approximately 20,000 proteins encoded by the human genome. This dynamic process involves covalently attaching functional chemical groups, such as methyl or acetyl, to specific amino acid (AA) residues within proteins. PTMs, which are mediated by specific modifying enzymes, can lead to substantial alterations in a protein’s activity, stability, and folding2. Notably, proteins leverage PTMs as a cellular messaging system, responding to various external cues and stressors. This dynamic interplay between proteins and PTMs facilitates cellular adaptation to diverse environments (i.e., ensuring the maintenance of cellular homeostasis). However, the dysregulation of PTMs has been implicated in conditions ranging from cancer to inflammatory and immune disorders, highlighting their critical role in both health and disease3,4.

Central to understanding PTMs is recognizing the interaction networks between the modifying enzymes and their corresponding PTM-modified substrates (i.e., enzyme-substrate networks). This recognition illuminates the extent to which a modifying enzyme impacts the proteome. Yet, conventional methods of discovery face challenges that have historically hindered the growth of these substrate networks. Peptide arrays and mass spectrometry (MS) analysis, although valuable, have their own set of limitations and biases5,6. Peptide arrays offer high-throughput representation of protein segments, but fall short of capturing the full scope of PTM function as new modification sites are identified7,8. Regardless, peptide arrays have commonly been applied to detail the binding site and search for unreported substrates of PTM-inducing enzymes914. In contrast, MS analysis provides a comprehensive view of cellular mechanics, but often requires affinity or chemical enrichment steps which may be challenging for certain PTMs, particularly for the discovery of lysine methylation and lysine methyltransferase (KMT) substrates as protein methylation lacks modification-specific pan-affinity reagents1,15,16.

Amid these challenges, the integration of artificial intelligence has emerged as a beneficial approach to dissecting protein function and substrate selection for PTM-inducing enzymes. Although generalized in silico methods to predict where a class of PTM may occur have paved the way1719, the application of deep learning, as seen in MusiteDeep20, is a notable progression forward. However, the prediction of specific substrates for PTM-inducing enzymes remains a starkly unexplored frontier2123. Furthermore, a majority of methods employ in silico approaches alone, frequently limiting their performance by their reliance on databases that can be poorly representative for the enzyme under study or of uncertain quality24. Often machine learning (ML)-based PTM prediction methods that are enzyme-specific require details about structure or the metabolic networks of the enzyme24,25. To address these concerns some investigations have combined in vitro experiments with in silico predictions, exemplified by approaches such as the Bayesian framework employed to characterize the substrates of protein-tyrosine phosphatase, PTP1B, using protein–protein interaction prediction26. Recently, Vinogradov et. al. applied mRNA display to develop a deep learning algorithm to profile substrate selection for a serine dehydratase and a cysteine/serine cyclodehydratase27. The creative approach employs mRNA display to generate a large, randomised dataset to train a ML model, however this technique is restricted to PTMs that may be chemoselectively biotinylated, and the randomised library is not biologically relevant27. In this study, we transcend traditional techniques by adopting an approach which applies high throughput in vitro peptide array experiments to generate ML models, referred to as our ML-hybrid approach, to enzyme-substrate identification that demonstrated utility across two classes of enzymes. We show that this paradigm can successfully identify enzyme-catalyzed PTM sites for different protein modification groups, including lysine methylation (SET8) and (de)acetylation (SIRT1-7) modifying enzymes.

Since the identification of substrates for KMT enzymes has lagged the discovery of lysine methylation sites, the identification of these substrates serves as a representative model in this study. The previously identified recognition site of SET8, a member of the Su(var)3-9, Enhancer of zeste, Trithorax-homology (SET) family of KMT enzymes, manifests a notable specificity towards lysines located in unfolded regions of proteins11, positioned ±4 AAs from the central lysine. This heightened specificity poses challenges in distinguishing SET8 methylation sites solely from peptide arrays, as the involvement of biophysical features beyond a simple sequential representation is anticipated in substrate recognition. Considering these complexities, SET8 emerges as a prime candidate for the development of a systematic ML-based approach to substrate identification. SET8 mono-methylates histone H4 lysine 20 (H4-K20), an event implicated in DNA damage repair, DNA replication, and cell cycle control28,29. SET8 also targets non-histone proteins, including K382 in the C-terminal protein domain of the p53 tumor suppressor and K248 of the proliferating cell nuclear antigen (PCNA)11,23,30. Additional substrates have been proposed, including UHRF1-K385 and α-tubulin-K31131,32. SET8 has been shown to be overexpressed in bladder cancer, non-small cell and small cell lung carcinomas, pancreatic cancer, leukemia, and other diseases29.

To broadly apply our ML-hybrid approach in delineating substrate specificity across a separate enzyme class and better evaluate model performance, we investigate the substrate networks associated with the sirtuin (SIRT) family of nicotinamide adenine dinucleotide (NAD + )-dependent deacetylases. Comprising seven homologs denoted as SIRT1 to SIRT7, the SIRT family plays a pivotal role in diverse physiological processes, including inflammation, glucose and lipid metabolism, oxidative stress response, cell apoptosis, autophagy, cell proliferation, as well as cell migration and invasion33. In the context of cancer, the SIRT family has been implicated in a wide spectrum of malignancies, autoimmune disorders, cardiovascular diseases, and respiratory disorders33. Consequently, the identification of substrates governing the function of the enzymes in this family holds substantial relevance.

Here, we improved on conventional in vitro and in silico techniques of substrate discovery by employing a unique ML approach trained on a peptide representation of the current modified methyl-lysine and acetyl-lysine proteomes. Unlike most ML predictors of PTMs, our “hybrid” approach begins with the experimental generation of enzyme-specific training data, rather than relying purely on an online database, in which just a handful of substrates may exist for the enzyme of interest and often lack negative examples34. By chemically synthesizing a representative PTM proteome using peptide arrays, then subjecting them to in vitro enzymatic activity, we can characterize enzymatic PTM activity in a facile way for virtually any classification of PTM. With the development of a ML model, augmented by generalised PTM-specific prediction17,35, we created ML-hybrid ensemble models unique to each enzyme that demonstrate enhanced predictive accuracy in cell models (Fig. 1). Moreover, the enzyme-substrate networks produced by ML-hybrid ensemble models specific to each SIRT family member reveals unreported pathways of conserved and enzyme-specific interaction.

Fig. 1. Graphical overview of the hybrid in vitro and in silico methods employed to generate a ML model for the prediction of PTMs by specific enzymes.

Fig. 1

A The generation of data is completed with peptide arrays that consist of the known sites of the PTM of interest and B an enzymatic construct. C ML models with data balancing are applied and assessed for performance on the resulting imbalanced dataset with metrics such as F-score, precision, and recall. D Finally, the resulting model is applied to the proteome to search for unique substrates of the PTM of interest, expanding the known substrate network of the enzyme of interest.

This ML-hybrid ensemble method demonstrates improved performance over conventional approaches and demonstrates the potential for broader applicability by its success within two classes of PTM-inducing enzymes. The use of high-throughput experiments to generate data for unique ML models specific to a PTM-inducing enzyme enhances the capacity predict the substrates of such enzymes, potentially streamlining the discovery of enzyme activity.

Results

SET8 expression and conventional substrate prediction

To accurately compare the efficacy of the ML-hybrid ensemble approach to conventional techniques, the in vitro method of permutation array-based prediction was first applied to identify potential substrates for SET8. Hence, an array-based permutation motif was generated, using the well-characterised substrate of histone H4-K207,11. The H4-K20 sequence was mutated ±4 AAs outside of the central methylated lysine (denoted as ‘X’; GGAXXXXKXXXXNIQ) and synthesized onto a peptide array. Next, a well-characterized and highly active SET8193-352 construct36 was expressed and purified (Supplementary Fig. 1A). Methyltransferase activity of the construct was confirmed using the canonical H4-K20 peptide (GGAKRHRKVLRDNIQ) (Supplementary Fig. 1B). The synthesized array of peptides, or permutation array was then exposed to the active SET8 construct to identify sequence variants susceptible to methylation (Supplementary Fig. 2A, Supplementary Data 1). The activity of each peptide SPOT in the permutation array was quantified through relative densitometry and analyzed using the motif-generating software PeSA2.0. This analysis produced the following motif: [KPGCHIVD]XH[RVIKYSAHML]K[IVT]L[RDLGI]X (Supplementary Fig. 2B)37. PeSA2.0 assigns weighted values to AAs at each position based on their observed capacity to be methylated, enabling a detailed representation of the enzyme’s substrate specificity37. Notably, these results closely resemble a prior motif determined in another study of SET8, as depicted by Kudithipudi et. al.11. A search of the known methyl-lysine proteome (Supplementary Data 2) was performed using the PeSA2.0-generated permutation motif. Methyl-lysine sites that scored above the normalized score cutoff of 0.5 were accepted as positive, relative to a perfect motif match (assigned a score of 1), which yielded 346 hits (Supplementary Data 3). Of these candidate substrate hits, just 26 peptides were validated as being methylated by SET8 in vitro (using peptide arrays and the SET8 construct), indicating a method precision rate (i.e., [True Positive] / [True Positive + False Positive]) of 7.5% in this enriched methyl-lysine proteome dataset (Supplementary Data 3).

Next, we applied the generated motif to identify unreported SET8 substrates from a dataset of predicted surface-exposed lysine residues of unknown methylation status (i.e., surface-exposed lysine proteome, generated with NetSurfP1.0). This approach is purposed to reveal the applicability of these approaches to the exploration of enzyme substrates beyond those that are currently known to be modified (e.g., undiscovered methylation sites, missense cancer mutations, etc.). Using the scoring matrix from the permutation motif array (Supplementary Fig. 2B), a search of the surface-exposed lysine proteome was performed and yielded 15,961 sites contained within 2,424 proteins (Supplementary Data 4). A randomly selected subset of these 15,961 sites (n = 100) tested in vitro identified two positive hits, indicating a relatively poor precision rate of 2%.

SET8 training set generation

To apply an effective ML model, the initial dataset used to train the model must provide sufficient samples of the positive case and the negative case, ideally in equal amounts38,39. A randomly sampled subset (n = 100) of sites from the surface-exposed lysine proteome were tested for SET8 activity using peptide arrays and identified no sites of SET8 methylation. This observation is indicative of a specific recognition motif for SET8-mediated methylation, and emphasizes the need for an improved approach for generating training data representative of sites positive for modification11. To efficiently enhance the likelihood of detecting positives, the human known methyl-lysine proteome was obtained from PhosphoSitePlus, a database of PTMs34. This dataset contains modified lysines, including mono-, di-, and tri-methylation. PTM-inducing enzymes such as KMTs often contain conserved catalytic domains and act upon methylatable histone and non-histone substrates, meaning the known methyl-lysine proteome should contain an enriched number of substrates for SET8; in contrast, the methylated status of the surface-exposed lysine proteome remains unknown30.

A total of 4593 peptides from the known methyl-lysine proteome were synthesized as peptide SPOT arrays and tested for in vitro SET8 methyltransferase activity34. Specifically, of the 4593 peptides tested, 213 were deemed to be positive for SET8 methylation (Supplementary Data 3). The 213 positive sites were identified across 179 proteins, indicating that several proteins harbored multiple SET8 methylation sites. To date, the commonly accepted substrates of SET8 are H4-K20, P53-K382, and PCNA-K24830. The 213 sites identified in vitro with the targeted search alone holds the potential to expand upon these four substrates.

SET8 base model fitting and fine-tuning

In order to generate features, the precursor to any ML model, the training data for SET8 was numerically encoded with the application of Molecular ACCess System (MACCS) keys, one-hot encoding of peptide sequences, and ProtDCal molecular descriptions (Supplementary Materials S3.1)3842. For the 15 AA peptides representative of the lysine methylome, the resulting set contains 483 features. To effectively validate the final ML model created, a holdout set of the features was withheld for final validation (as outlined in Supplementary Materials S1). Considering the relatively small size of the remaining features, now the ML training data, K-fold cross-validation was applied, in which SET8 training data was further split into (1) training and (2) testing sets repeatedly to effectively assess model fitting and prevent overfitting (Supplementary Materials S1)43. The F-score was selected as the best measurement of the predictive performance of the model considering the imbalanced representation of positives within the dataset, defined by the harmonic mean of precision and recall44,45.

Fit to the training data, a linear discriminant analysis, along with random oversampling of the positive class (i.e., sites positive for SET8 methylation), attained the highest F-score of 0.1346,47. An m-threshold analysis was then performed using the holdout set of never-seen data to effectively assess the effect of the decision threshold on performance (Fig. 2A). Additionally, both the precision-recall and receiver operating characteristic (ROC) curves were generated with the holdout set (Fig. 2B). Metrics for the threshold that optimized F-score resulted in an F-score of 0.24, a precision of 0.2, recall of 0.31 ([True Positive] / [True Positive + False Negative]), and specificity of 0.83 ([True Negative] / [False Positive + True Negative]). The metrics further demonstrate the benefit of the F-score, and describe our positive identification rate, rather than using specificity, which is falsely inflated by the large proportion of negative identifications within the training dataset. The accurate classification of negatives is particularly challenging as many PTMs dynamically change over time due to various molecular events within the cell. It is important to note that due to the nature of either metric, improvements in precision will result in a decrease in recall, and vice-versa; a careful balance between the two is often desirable for optimal ML performance44. Feature importance was calculated for the selected SET8 model (Supplementary Fig. 3), hereafter referred to as the base model. This revealed the importance of the sequential one-hot encoding throughout the sequence, demonstrating a proclivity for cysteine and histidine, and an aversion to tryptophan (Supplementary Fig. 3).

Fig. 2. SET8 ML model assessment and preliminary experimental verification.

Fig. 2

A Threshold analysis of the basic SET8 ML model, including performance metrics of precision, recall, and specificity. Optimized base ML threshold is indicated at 0.72, as determined by maximized F-score. B The top graph is the precision-recall curve for the base ML model, with the receiver-operating-characteristic (ROC) curve on the bottom. C Threshold analysis of the ML-hybrid ensemble model, as in A. Selected threshold for the optimization of model performance is indicated at 0.82 in red. D Progression of model predictions and SET8 substrates validated with peptide array experiments. E Comparison of methods, including a random search, the SET8 recognition motif generated with permutation arrays, and the ML-hybrid ensemble model applied to the proteome dataset comprising surface-exposed lysine. F Amino acid composition comparison for the ML-hybrid ensemble model predictions of SET8 methylation sites. Confirmed substrate sites of SET8 methylation shown below, for H4-K20 and p53-K382. G SAFE mapping of ML-hybrid ensemble model predictions to the HuRI interactome, clustered for associated biological processes.

SET8 ML-hybrid ensemble model construction

The ability of a lysine residue to undergo methylation is a prerequisite for any newly predicted SET8 substrate. To enhance the performance of the SET8 substrate prediction, or base model, a composite or ensemble model was constructed using MethylSight, the current state-of-the-art generalized predictor of lysine methylation17. In the inaugural study, MethylSight identified 51 distinctive sites of histone methylation, and 89% of the sites were confirmed to physically exhibit methylated lysine17. Much like the SET8 training data previously described, MethylSight uses ProtDCal to characterize the 15-AA-long peptides centered on a central lysine17. Hence, it is well suited for integration with the SET8 substrate prediction model using stacked ensemble learning.

As with the initial model fitting, a holdout set was withheld for final testing, and stratified K-fold cross-validation was applied on the remaining data to assess the performance of each model (Supplementary Materials S1). Here two features were applied: the SET8 substrate prediction score (described above); and the MethylSight score (i.e., the likelihood of methylation). The F-score was optimized with the application of a logistic regression model and SVM SMOTE oversampling4850. The simplicity of logistic regression was reflected in the singular hyperparameter of 100 max iterations determined from the tuning process48. Applying the holdout set, an F-score of 0.12 was selected based on inspection of the precision-recall trade-off, for the ensemble model at a threshold cutoff of 0.82, resulting in values of 0.25 for precision, 0.08 for recall, and 0.98 for specificity. Although the F-score decreased slightly, the efficacy of the model towards the unseen data within the proteome (i.e. model generalizability) has increased with the incorporation of MethylSight through the ensemble approach. This is effectively demonstrated by the approximate three-fold increase in performance of the ensemble model (27.8%; 139 hits / 500 tested predictions) when compared to the base ML model alone (10.4%; 52 hits / 500 tested predictions) upon validating positive substrate predictions for SET8 methyltransferase activity using peptide SPOT arrays, as previous described (Supplementary Fig. 4). A comparison of performance metrics with classification threshold is illustrated in Fig. 2C. Given the performance increase gained from the integration of methyl-lysine prediction into the ensemble model (hereafter referred to as the SET8 ML-hybrid ensemble model), the investigation proceeded with this hybrid model (Supplementary Fig. 4; Supplementary Data 5 and 6).

Proteome-wide prediction of SET8 substrates

Using our SET8 ML-hybrid ensemble model, we predicted an additional 2367 sites of SET8 methylation from the surface-exposed lysine proteome. Experimental validation of these 2367 predicted positive sites of SET8 methylation was completed by testing each site (represented as 15 AA peptides) for in vitro methylation using peptide SPOT arrays. Of these predictions, 885 sites permitted in vitro SET8 methyltransferase activity, representing a validated precision of 37.4% (Fig. 2D, E). The precision of this method is much improved over the 0% validated precision of the random search method and the 2% validated precision determined with the permutation array within the same surface-exposed lysine proteome. Additionally, a notable improvement was gained over the 4.7% validation rate generated from the extensive search of the highly enriched known methyl-lysine proteome dataset (Fig. 2E). To gain insight into potential sequence conservation among SET8 substrates, an analysis of the sequence composition of the 885 sites of SET8 methylation identified by the ML-hybrid ensemble model demonstrates that SET8 substrate selection may be more permissive than previously proposed, as each AA position flanking the central lysine is predicted to accept an array of different AAs beyond those within H4-K20 and p53-K382 (Fig. 2F). This visual depiction of sequential representation directly demonstrates the lack of confirmed substrates for SET8, and the benefit of the ML-hybrid ensemble approach in decoding enzyme-substrate networks. The SET8 ML-hybrid ensemble model proved to be 100% accurate in identifying a subset (n = 362) of predicted negative, lowest-scoring sites, as verified by peptide array experiments with SET8 (Supplementary Data 7). Based on these findings, our SET8 ML-hybrid ensemble model improves on the traditional substrate identification approach.

Modified residues must be accessible to be enzymatically altered, and filtering residues based on a higher predicted surface accessibility may improve predictive precision. To determine the effect of the predicted surface exposure across the annotated recognition sites, NetSurfP1.0 was applied to the full-length proteins containing the lysine-centric sites of all 2367 predictions from the SET8 ML-hybrid ensemble model. An increased amount of predicted surface exposure for the central lysine within the full-length sequence corresponds to a slight increase in precision (37.3 to 42.0%) (Supplementary Fig. 5), however the number of experimentally validated SET8 ML-hybrid substrates drastically decreases with this change (885 to 402). This finding indicates that increasing the cutoff beyond the recommended value of 0.2 for NetSurfP1.0-predicted surface exposure51 within the experimental set will incur additional false positives. Hence, the chosen parameters effectively reflect surface accessibility in the context of suitable enzyme accessibility.

The 2367 positive SET8 methylation sites predicted by the SET8 ML-hybrid ensemble model within the surface-exposed lysine dataset were contained within 1203 proteins. To investigate the enriched biological functions of the 1203 modified proteins (i.e., predicted SET8 substrate network), clustering analysis with GO annotations was performed using the spatial analysis of functional enrichment (SAFE) approach (Fig. 2G)52,53. The HuRI proteome was selected for protein mapping because of its quality and high-confidence protein interactions54. A shared theme among the enriched biological processes determined is involvement in cell homeostasis, regulation, and control of the cell cycle (Fig. 2G). Given the established involvement of SET8 with these cellular events, mediated through known substrates, the possibility that SET8 might participate in such processes through the methylation of other substrates identified by our SET8 ML-hybrid ensemble model is bolstered29,30,36,55. Other affiliated processes include mRNA and RNA polyadenylation. Regulation through the polyadenylation of mRNA has been associated with other SET-domain-containing methyltransferases, specifically SET1 and SET2, through histone methylation56. The involvement of SET8 in transcription modulation may also implicate it in polyadenylation regulation; however, a direct connection has not been reported29. In summary, the substrates generated by the ML-hybrid ensemble model provide the potential to unveil unexplored functional narratives for SET8 and its role(s) in disease.

The efficacy of the SET8 ML-hybrid ensemble model is further demonstrated by the progression from the surface-exposed lysine residues (a total of 145,379 sites) to the 2367 predictions (Supplementary Data 8 and 9), resulting in 885 in vitro validated sites (Fig. 2D). Functional analysis of the 885 validated sites predicted for SET8 activity, mapped to the HuRI proteome using SAFE, also reveals an association with signaling and cellular processes, as shown in Fig. 2G. Specifically, these validated sites are linked to EGF signaling and protein ubiquitination, including processes mediated by the proteasome (Supplementary Fig. 6).

An additional investigation was performed to identify the effect of missense mutations found in breast cancer cells on SET8 activity.

Cell-based validation of SET8 substrate candidates

A comprehensive understanding of enzyme substrate selection is crucial for guiding cellular identification studies and accelerating experimental validation. To validate SET8-influenced cellular methyl-lysine events, we used parallel reaction monitoring MS to assess the in vitro SET8 substrates identified by our SET8 ML-hybrid ensemble model in a targeted manner.

To streamline the number of methylation sites monitored, we generated an isolation list from the 885 sites validated in vitro which focused on primary and secondary SET8 interactors, as defined by the STRING database57 (Supplementary Data 10). The resulting network (Fig. 3A) comprised 44 proteins, containing 75 predicted SET8 methylation sites that were subsequently confirmed to demonstrate in vitro SET8 methyltransferase activity using peptide SPOT arrays. Of these 75 sites, it was predicted that 32 sites would create tryptic peptides in silico; these were targeted for MS monitoring in SET8 overexpressed HCT116 cells29,58 (Fig. 3B; Supplementary Materials S3.2). Overexpression of SET8 was compared with untreated empty vector control conditions, as knockout of SET8 and subsequent loss of H4K20 monomethylation severely impacts cellular proliferation29,58. Of the 32 monitored sites, only nine were reliably detectable, and elevated levels of mono-methylation were observed in three (33%) of these substrates: SETD1B-K41, KAT6A-K314, and PRDM12-K269 (n = 1 per condition, based on 3 pooled biological samples) (Fig. 3C, 3D). The remaining sites did not positively respond to SET8 overexpression (Supplementary Fig. 7). In conclusion, the ML-hybrid ensemble model can identify substrates of possible SET8 methylation activity, as confirmed by site-targeted MS monitoring.

Fig. 3. MS analysis of the validated predictions of SET8 methyltransferase activity, generated by the ML-hybrid ensemble model.

Fig. 3

A Primary (predicted SET8 interactors) and secondary interactors (a predicted interactor of a SET8-interacting protein) of SET8, as represented in the STRING database. Circles represent MS-monitored sites. B Western blot analysis of the overexpression of SET8 within HCT116 cells, compared with wild type. C Targeted MS results of SETD1B-K41me1, including the peptide sequence described in the isolation list. Alternate line colors are representative of individual daughter ions detected within the parental peptide sequence. D Averaged peak intensities relative of SETD1B-K41me1, KAT6A-K314me1, and PRDM12-K269me1 from SET8 overexpression (SET8 OE), compared with wild type (NT) (n = 3 biological samples, pooled). Small minor peaks in the chromatogram can result from background noise, isobaric or co-eluting compounds, or unexpected fragment ions.

Redefining SET8 substrate networks in cancer

Elevated expression of SET8 is linked to a high mortality rate in patients with breast cancer, as reported in Fig. 4A, modified from Liu et al.57,59. However, the behavior of SET8 in cancer remains unclear, and further investigation is required to uncover the functional role(s) SET8 plays in tumorigenesis. As cancer-associated mutations continue to diversify, mutation datasets serve as a valuable resource with which to elucidate the effect mutations have on protein structure and function. In the case of missense mutations, they may cause the gain or loss of methylatable lysine or make changes to neighboring residues, which then dictate the suitability of these sites for SET8 methylation57.

Fig. 4. Prediction of SET8 substrates among missense mutations in breast cancer.

Fig. 4

A Kaplan-Meier plot depicting patient survival with high (green) and low (red) SET8 expression in breast cancer (hazard ratio is 5.90 and p-value is 0.009). Modified from Liu et. al., 2016. B Comparison of predicted SET8 sites of lysine methyltransferase activity between healthy and mutated cancerous states. Squares at the end of each row correspond to the color code applied in (C) which is a visualization of cancer-mutated and healthy SET8 predicted ML scores for each site analyzed. D Nuclear excision repair pathway initiated by DNA damage. Mutation of ERCC4 involved in the critical excision step is predicted to invoke SET8 methyltransferase activity by the ML-hybrid ensemble model. Gap filling by PCNA, a known substrate of SET8 in normal cells, is included.

To explore the possibility of gain or loss of SET8 substrates in breast cancer, missense mutations were downloaded from the COSMIC database (v.96) and applied to the human proteome. Of the initial mutations, 9438 either occurred within seven AAs of a lysine residue (e.g., any residue) or resulted in the gain or loss of an individual lysine, directly impacting the creation or loss of a potential methylation site (Supplementary Data 11). The corresponding unmutated sites, except sites in which a lysine did not previously exist, were also assembled. Application of the SET8 ML-hybrid ensemble model to normal and breast cancer datasets predicted that most of the mutations (94.6%) would not affect SET8’s methylation behavior toward the site, likely because their structure is not changed dramatically by a single AA mutation. In contrast, 4.0% (376) of mutations resulted in a predicted gain of SET8 methylation, and 0.7% (62) resulted in a loss (Fig. 4B). Of the 4.0% of sites predicted to gain SET8 substrate status, 46.8% (176) were the result of the mutation introducing a new lysine that is itself predicted to be methylated by SET8 (Fig. 4C). MCODE clustering analysis of the total set of mutations revealed a directed subset of predicted substrate interactions that were highly interconnected with SET8 (Supplementary Fig. 11). Mutations within the subset that resulted in a gain of predicted SET8 methylation were investigated for involvement in pathways implicated in breast cancer. Of particular interest was XPF (encoded by ERCC4), a protein associated with the vital cellular process of DNA damage repair60. Specifically, the XPF-S352A mutation led to the prediction of a unreported SET8 methylation site at XPF-K350. As detailed in Fig. 4D, XPF is directly involved in DNA damage repair, including nucleotide excision, double-strand break, and interstrand cross-link repair pathways61. Gap filling is then proceeded by PCNA, a known substrate of SET830, further implicating SET8 in the NER pathway62. Interestingly, SET8 has been Implicated In DNA repair previously, specifically in 53BPI/BRCA1 double-stranded DNA repair through histone H4K20 mono-methylation62. Beyond breast cancer, our SET8 ML-hybrid ensemble model was also applied to missense mutations present in pancreatic cancer (COSMIC database, v.96) (Supplementary Fig. 12A, B; Supplementary Data 12).

Employment of the ML-hybrid ensemble method to the SIRT1-7 family

To demonstrate the broad applicability of our ML-hybrid approach to predict enzyme substrates, the procedures for the fitting of an ensemble model were carried out for each of the seven SIRT members using published data obtained from peptide microarrays5. MuSiteDeep35 was employed for the generalized PTM prediction component of the ensemble model. Notably the base model training dataset contained 6801 peptides of 13 AA in length, representing the acetyl-lysine proteome5, resulting in the generation of 443 descriptive features. The number of deacetylation sites (i.e., positive and negative ratio) differed for each SIRT, as well as the best-fit model and meta-model (Supplementary Fig. 8, Supplementary Materials S3.3). The surface-exposed lysine proteome was employed once more to predict SIRT substrates; however sequences were limited to 13 AAs to reflect the training data. The shared predicted substrates between each SIRT enzyme (Supplementary Data 13) are demonstrated by a Circos plot, highlighting the high degree of substrate overlap of substrates within the top 10% of predicted substrates (Fig. 5A), and a quantitative analysis of the shared substrates within all predictions is demonstrated within an UpSet plot (Fig. 5B). Our predictive ML-hybrid approach revealed unique insights into the SIRT1-7 family, including overlap between specific SIRT enzymes (e.g., the high level of substrates shared between SIRT2 and SIRT3; Fig. 5B) and substrate features that impact the selection of conserved and specific substrates (Supplementary Fig. 8).

Fig. 5. Sirtuin family ML model assessment and experimental prediction analysis.

Fig. 5

A Circos plot demonstrating the overlap between the top 10% of predicted substrates for the 7 sirtuins, as generated by 7 unique ML-hybrid ensemble models. B An UpSet plot to illustrate the totality of the predictions as generated by the ML-hybrid ensemble models for the sirtuin family. The size of each group of shared predictions is represented by the intersection size found on the x-axis. C Threshold analysis of the ML-hybrid ensemble model for SIRT2, including performance metrics of precision, recall, and specificity. The selected threshold of optimized model performance is 0.50, as indicated in red. D Precision-recall curve atop the ROC curve for the ML-hybrid ensemble model predictive of SIRT2 deacetylation. E MS confirmation of predicted sites of SIRT2 deacetylation as generated by the SIRT2 ML-hybrid ensemble model from HCT116 cells overexpressing SIRT2.

To extend our prediction of SIRT substrates to cellular validation, we looked at the performance of our ML-hybrid ensemble model to confirm reported SIRT2 cellular substrates from HCT116 cells over-expressing FLAG-SIRT2-WT or catalytically inactive FLAG-SIRT2-H187Y SIRT2 as previously identified by tandem MS63. Our SIRT2 ML-hybrid ensemble model (Fig. 5C, D, Supplementary Fig. 9) was able to successfully reconfirm 64 of 149 previously confirmed SIRT2 substrates (Supplementary Data 14), indicating an precision rate of 43.0% (Fig. 5E) and demonstrates the ability of our model to be generalized to deacetylase enzyme classes with high experimental precision.

Discussion

Our ML-hybrid ensemble model represents an advancement in the identification of substrates for PTM-inducing enzymes, offering a more efficient and scalable approach compared to existing methods. This strategy is successfully demonstrated for two enzyme classes and may be readily extended to other PTMs. A key strength of this approach lies in its capacity to accurately characterize PTM-specific subsets of the proteome using peptide arrays, which allows for the profiling of several thousand sites in parallel (Fig. 1). This capability provides an improvement in identifying unreported enzyme substrates over traditional array-based methods5,64.

Although many models can predict PTM sites, few are capable of identifying the specific enzyme responsible17,21,65. In contrast, our ML-hybrid ensemble model overcomes this limitation65,66 by targeting enzyme-specific substrates, as demonstrated with SET8 and SIRT2. These models achieve experimentally validated precision rates of 37−43%, far surpassing the 2.0% precision achieved by traditional permutation arrays used to identify enzyme-substrate recognition motifs. The high precision of the approach makes it an ideal tool for substrate identification studies, facilitating more targeted and focused validation efforts.

Despite the limitations imposed by the reduced amount of labeled PTM-specific datasets, which constrain model complexity, our approach remains computationally efficient. The ML-hybrid ensemble method proves capable of effectively identifying potential substrates of PTM-inducing enzymes with limited experimentally validated data, presenting the potential to rapidly expand the substrate space for PTM-inducing enzymes. Deep learning algorithms, including neural networks, are not suitable for enzyme-substrate PTM prediction due to these limitations, but our approach aligns with the current understanding that deep learning is more appropriate for generalized PTM predictions, as exemplified by MusiteDeep65. Notably, our method uses abbreviated peptide sequences, which may not effectively represent full-length enzyme-substrate interactions, especially when considering that the folded structure may prevent access to the residues that present potential PTM activity. This constraint is consistent across conventional array-based experiments, which employ similar methods to validate results11, yet the goal of such methods as well as our approach is to better guide in vivo investigations through targeted, scored lists of potential substrates. Moreover, although the datasets of known modified sites used to generate the data for our method carry bias due to research-driven interest and experimental methods that often target known motifs, our ML-hybrid model identified potential sites for SET8 without clear motif bias (Fig. 2F). In spite of these constraints, the high number of predictions generated by our model, over 2300 for SET8, has led to 885 validated sites, greatly expanding the pool of candidates for MS monitoring11.

The targeted MS monitoring of predicted sites of SET8 methylation in HCT116 cells revealed the increased monomethylation of SETD1B-K41, KAT6A-K314, and PRDM12-K269, suggesting SET8’s involvement in gene activation and regulatory pathways (Fig. 3). Specifically, SETD1B-K41me1 levels suggest a possible connection between SET8 and the COMPASS complex, of which SETD1B is a part67. Additionally, this observation hints at a potential connection between H3K4 mono-methylation and SET8 through SETD1B and the COMPASS complex. KAT6A, commonly localized to CpG islands, may be influenced by SET8 mono-methylation at K26968, potentially impacting gene regulation. Furthermore, the association of both KAT6A and SETD1B with CXXC1 (CpG-binding protein) implicates SET8 in gene activation68,69. The regulatory domain binding factor PRDM12 was observed to demonstrate elevated K269me1 levels with SET8 overexpression70. Further investigation is needed to determine the significance of these findings, particularly for substrates like CAMKMT-K49 and SETDB1-K732, which did not exhibit methylation changes in response to SET8 overexpression (Supplementary Fig. 7) and may indicate either, (1) sites that cannot be methylated by SET8 or (2) substrates already maximally methylated in cells.

Our model’s ability to explore relatively uncharted datasets and reveal unique insights is exemplified by its application to breast cancer-specific SET8 enzyme-substrate networks. To help reveal potential insights into the role(s) SET8 plays in breast cancer, our SET8 ML-hybrid ensemble model was applied to explore the creation and loss of potential substrates to help annotate a predicted breast-cancer-specific SET8 enzyme-substrate network. When applied to cancer mutation datasets, the ML-hybrid ensemble model has specific advantages. The sensitivity of the model is evident when comparing oncogenic and healthy proteome scores, which provide a unique perspective on oncogenic mutations. For example, the predicted gain of SET8 substrate methylation due to a missense mutation sheds light on potential proteins of interest, particularly the association of SET8 with breast cancer within the NER pathway through XPF (Fig. 4)61. This implicates SET8 in the DNA repair process, employing PCNA, another substrate of SET8, for gap filling61. The XPF-S352A mutation and subsequent SET8 methylation at XPF-K350 could be hypothesized to affect the NER pathway, such that DNA damage repair may not be completed. Interestingly, the K350 site falls within the XPF-SLX4 interaction region71, an interaction associated with the Fanconi anemia DNA repair pathway and defects in SLX4 function have been linked to breast cancer72. Beyond direct involvement with cancer, the mutations yielding a predicted change in SET8 activity may also serve to affect a reader site for other proteins involved in cancer, as in PTM crosstalk7376. Such sites may not directly possess a methyl-binding protein (associated reader), but could serve to regulate neighboring PTMs, or yield potentially allosteric effects which alter the proliferation of breast cancer73,75.

The broader applicability of our approach is suggested by its success with another enzyme class. Using applied data balancing methods for our SIRT models5 we were able to achieve notably high precision rates, with MS-validated precision for SIRT2 reaching 43.0%; a comparable but slight improvement over the 41.2% precision determined by the testing set (Supplementary Materials S3.2), advocating for the performance of the other SIRT models. The conserved and enzyme-specific substrates within each SIRT network further provides a unique insight into potential pathways of activity (Fig. 5). Interestingly, SIRT2 and SIRT3 share almost twice the substrates as any other combination of SIRT. This finding is supported by prior investigation that uncovered that the pair may compensate for each other when the other is deficient77. The second largest overlap in predicted substrates occurs between SIRT5 and SIRT7. As the most recently discovered member, SIRT7 is relatively underrepresented in literature, however some associations between SIRT5 and SIRT7 exist. The two have been implicated within disease states such as cardiac hypertrophy and inflammatory bowel diseases, as well as the regulation of the NF-κB inflammation signalling pathway33. The shared specificity and substrate overlap captured by the ML-hybrid ensemble approach aligns well with prior investigations suggesting the compensation of SIRTs for one another33,77.

Potential improvements to our method focus on the ML component of the pipeline. The density of the SPOT peptide arrays is correlated with substrate preference and may improve performance when incorporated through regression-based ML models as opposed to the classification models currently used. Alternatively, the incorporation of the density as a feature within the current framework will also be investigated. Recent updates to the representations of peptides as features for ML will form the basis of future work to improve our approach, including those generated by protein language models such as ESM-2 and ProtBERT considering both have recently been employed for peptide-based studies78,79.

In conclusion, our ML-hybrid ensemble approach offers a refined method for the prioritized identification of enzyme-substrate networks and the features of substrates that influence selection. Although some challenges exist, including the inherently limited dataset size of known PTM sites and an unavoidable class imbalance, we have implemented strategies to overcome these limitations, offering a robust alternative to traditional methods to enable the discovery of substrates for PTM-modifying enzymes. The successful validation of predicted SET8, and SIRT2 substrates from MS data underscores the reliability of our approach, outperforming traditional methods and opens alternative avenues for studying enzyme-substrate selection in cellular processes. Moving forward, the ML-hybrid ensemble method holds promise for expanding our understanding of enzyme-substrate networks, PTM dynamics, and protein function modulation.

Methods

Peptide synthesis

Peptide SPOT arrays were synthesized to commercial aminated cellulose membranes (Intavis Inc.) with standard Fmoc (N-(9-fluorenyl)methoxycarbonyl) chemistry automatically using a Multipep synthesizer (Intavis Inc.)80. Resulting arrays contained approximately 2 nmol peptide per SPOT, separated from the membrane surface by a flexible C-terminal 6-aminohexanoic acid linker. Post-treatment involved the cleavage of the protective groups from the side-chains with an acidic solution (51% water, 47.5% trifluoroacetic acid, and 1.5% tri-isopropylsilane)80,81. Arrays were then washed extensively in ethanol and dried for future use. Dried arrays were stored desiccated at 4 °C.

Protein expression and purification

A bacterial expression construct of human SET8 encoding the catalytic site residues 191-352 was cloned into the parallel expression vector pHIS2 using BamHI and XhoI restriction sites55. The construct was transformed in BL21 DE3 Escherichia coli, and after propagation were induced with 0.3 mM isopropyl β-D-1-thiogalacttopyranoside overnight at 16 °C55. Following soluble protein extraction, affinity column purification of the lysate was completed with 500 μL HisPurTM Ni-NTA Resin (ThermoFisher Scientific, Cat# 88221). Fractions were eluted using P500 buffer (50 mM NaHPO4 (pH 7), 500 mM NaCl, 10% glycerol, 0.05% TritonX-100, 1 mM DTT, 500 mM Imidazole) and then dialyzed into a storage buffer (20 mM tris pH 7.5-8, 200 mM NaCl, 10% glycerol, 1 mM DTT). Protein purity was assessed by SDS-PAGE using a 12% polyacrylamide gel (Supplementary Fig. 1A) stained with Coomassie Brilliant Blue G-250 stain. Protein concentration was determined with a Bradford Assay80,82. Prior to storage, activity was confirmed using the Methyltransferase-Glo Assay to assess in vitro activity with H4-K20 peptide (GGAKRHRKVLRDNIQ) (Supplementary Fig. 1B), according to the manufacturer’s specifications (Promega, Cat# V7601). Recombinant SET8 protein was then snap frozen and stored at -80 °C.

Training dataset generation

The annotated Human lysine methylome was obtained from PhosphoSitePlus (accessed February 5th, 2020). This dataset contained sequence information for each lysine methylation site represented as 15 AA peptides, centered on the methylated lysine (i.e., position 8)34,81,83. The data was then cleaned using Python 3 and the Pandas package for efficient isolation of human protein sites of lysine methylation8285. Briefly, sites that were within 7 AAs from the beginning or end of the sequence were padded with alanine residues, and any duplicate sequences were removed (i.e., methylation events that occur within a conserved 15 AA sequence, but among unique proteins). The resulting dataset yielded a total of 4593 Human lysine methylation sites (Supplementary Data 2) which were synthesized on peptide arrays.

Peptide arrays were then methylated using 1 μM SET8191-325 with 5 μCi/mL S-[methyl-3H]-Adenosyl-L-methionine (SAM) in methylation buffer (50 mM Tris pH 8.5, 2 mM MgCl2, 10 μM DTT) overnight at room temperature84,86. After methylation, arrays were washed 6×3 min in buffer (100 mM NH4HCO3, 1% SDS) and then a 7% 2,5-Diphenyloxazole solution in ethanol was sprayed generously over the array and left to air dry. This procedure was repeated a total of three times. Dried peptide arrays were then exposed to intensifying screens (Dupont, Cronex Lighting Plus) at room temperature for two weeks and imaged using a Typhoon™ FLA 7000 IP phosphorimager (General Electric). Peptide methylation was determined by densitometry using the Protein Array Analyzer (v.1.1.c) toolset for Image J (v.1.53) to obtain SPOT densitometry from the array images. The average of 4 control spots (H4-K20 peptide) were present at the corners of all methylated peptide arrays and used to normalize densitometry across the nine arrays. SET8-methylated peptides were determined by any peptide with normalized signal > 2 SD local SPOT background (p < 0.05).

The peptide substrate dataset for the sirtuin family was adapted from Rauh et al., 20135. Peptides which produced a signal ratio score of 1 or greater were classified as positive, and positions with masked AA residues were replaced with alanine (Supplementary Data 15).

Biochemical Feature Generation Peptide libraries were represented numerically through one-hot encoding, a common method of vectorizing peptides38,39. A matrix describing AA position and identity with a 1 or 0 is applied to the sequence, resulting in 300 features for SET8, and 260 for the sirtuins (Eq. 1).

xAA,pos=1ifseqpos==AA0ifelse,forpos=1,N,AA={A,C,D,,W} 1

To encode molecular structure, the Molecular Access System (MACCS) Keys of 166 binary fingerprints were applied40. Created to describe molecular structure, MACCS Keys include predefined atom symbols, bond types, and atom properties40. The RDKit package for Python (www.rdkit.org) was applied to generate MACCS Keys from the site sequence, representing the molecular structure for the AAs present.

Aggregate molecularly descriptive features were generated for the peptide sequence through ProtDCal (v4.5)42. The 17 metrics selected included molecular weight, hydrophobicity, isoelectric point, free energy, and Levitt’s Probability for various protein conformation amongst other descriptive sequence-based metrics42 (Supplementary Materials S3.1). Combined, the sequential information provided by one-hot encoding, along with the unique molecular elements represented by the MACCS keys and the molecularly descriptive features determined by ProtDCal resulted in 483 features for SET8, and 443 for the sirtuins. Each feature accurately captures a structural component from sequence alone, enabling the use of varied lengths of peptides beyond sites representative of lysine, such as arginine, or any of the other 20 essential AAs.

Machine learning – model fitting

All procedures related to model fitting and data balancing were completed in Python 3 using the Scikit-Learn and Imbalanced-learn packages82,8488. Pandas and Numpy were applied for data handling and storage, and plots were generated with Matplotlib83,85,8790.

Class imbalance was present in all samples, although it ranged in degree for each enzyme studied. The SET8-methylated lysine peptide array data provided a class imbalance of 213 positives within the 4593 sites tested for SET8 activity. Imbalance, along with the relatively small size of the training data, was taken into careful consideration when selecting ML models, and the allocation of data45. Additional information regarding the methodology of data allocation is present within the Supplementary Materials S1.

For proper comparison to a baseline model, the dummy classifier which simply randomly classifies data input was first fit to the data. Applying the F-score as the primary selection criteria and a model complexity considerate of the size and imbalance of our dataset as the secondary selection criteria, the selected models were linear discriminant analysis (LDA) for SET8 and SIRT2, though the decision tree classifier was selected for some members of the SIRT family (Supplementary Materials S3.3)46.

Class imbalance in which the positive case composed less than 30% of the training dataset was addressed with data balancing or sampling methods, which was the case for all datasets except SIRT5 (38.4% positive, 61.6% negative). Both over and undersampling techniques were tested to improve F-score with our models, and the best performing data balancing approach was selected (Supplementary Materials S3.3). In these approaches, a selection of the positive class is replicated within the dataset, increasing the overall percentage of positive values85,87. Alternatively, negative values are removed to reduce the overall percentage of the negative case within the training dataset. Additional information regarding the K-fold cross validation applied, hyperparameter tuning, and optimizing the decision threshold of the model is available within S1 of the Supplementary Materials.

Machine Learning – Ensemble Learning

To improve the overall performance of the models, ensemble learning algorithms were explored. In ensemble algorithms, a secondary ML predictor is applied in parallel. The lysine methylation predictor MethylSight was implemented as the secondary method for the SET8 model, as it generally predicts sites of lysine methylation; a dependent variable for a lysine to be a SET8 substrate17. MethylSight uses support vector ML, and was validated experimentally to predict previously unidentified sites of methylation, including the identification of KDM5B substrate H2B-K43me217. For the sirtuin family, the generalized deep learning predictor MusiteDeep was applied as the secondary predictor. The N6-acetyl-lysine predictive function of MusiteDeep was employed to generally identify sites of lysine acetylation, much like MethylSight for SET8. An easily-accessible generalized PTM predictor, MusiteDeep outperformed other representative ML and deep learning algorithms for N6-acetyl-lysine prediction35.

The ensemble methods explored included hard voting, soft voting, and stacking (Supplementary Fig. 10, Supplementary Materials 1). Each method takes in the probability score output by either predictor that a site is methylated. The hard voting method classifies through most votes, or scores above our specified positive cutoff thresholds for the positive class by either model89,91. In soft voting, classification is determined through the mean of probabilities by either model89,91. Finally, stacking utilizes the scores of each model as features of a third model, along with dataset balancing as previously outlined89,91. Each method was tested, and F-scores were generated with repeated stratified K-fold cross-validation. As it was available, a holdout set from the MethylSight investigation was employed to provide the go absolute unbiased final validation of each ensemble method tested for the SET8 model17. No holdout set was obtained for MuSiteDeep, therefore the associated investigations employed the use of a holdout set representative of 10% of the experimental data.

Experimental dataset generation and evaluation

The human proteome was obtained from UniProt, and analyzed with NetSurfP1.0 to generate values for relevant surface accessibility51,90,92. Lysine residues with a relevant surface accessibility value of over 0.2 were selected (the NetSurfP1.0 threshold for surface exposure), and the surrounding (either ±7 or ±6 from central lysine) AAs were isolated from the full protein sequence to form the experimental datasets. Any sites less than 7 or 6 AAs from the beginning or end of a protein sequence were padded with alanine. Additionally, any duplicated sequences or sequences found within the training dataset were dropped.

The SET8 ensemble learning model was applied to the surface exposed set and resulted in 2367 positive predictions of previously unidentified lysine methylation sites that are also SET8 substrates (Supplementary Data 8). These values were mapped to the human binary protein interactome (HuRI), and abundant biological processes were showcased with the Spatial Analysis of Functional Enrichment (SAFE) package for Python 352,54,82,84. All predicted positive sites were tested individually via peptide SPOT array experiments for SET8 in vitro methylation, as previously described (p < 0.05). The resulting seven SIRT models were applied and predicted various unreported sites of deacetylation (Supplementary Data 13). The commonality of predictions between each SIRT was illustrated through Circos and UpSet plots as generated by the pyCircos and UpSetPlot packages for Python 382,84,9194.

Cell line transfection

Human colorectal carcinoma HCT116 cells were cultured in Dulbecco’s modified Eagle’s medium (DMEM, Gibco Cat# 11965092) supplemented with 10% heat-inactivated FBS and 1% penicillin/streptomycin at 37 °C, 5% CO2. Cell lines were regularly tested for Mycoplasma pneumoniae using polymerase chain reaction (PCR) assay. For SET8 overexpression experiments, cells were transfected with 10 μg pcDNA3.1 HA-SET8 using jetOPTIMUS transfection reagent (Polyplus, Cat# 101000025), following manufacturer instruction. The cells were harvested 48 h after transfection and cell pellets were snap frozen and stored at -20°C until further use. Cell pellets were lysed in protein extraction buffer containing 50 mM tris (pH 8), 150 mM NaCl, 1 mM EDTA, 10% glycerol, 0.5% NP-40, and protease inhibitors (1 mM PMSF, 10 μM E-64, 1 μM Pepstatin, 1 μM Leupeptin, and 1 mM sodium orthovanadate). Protein concentration was determined by Pierce BCA Protein Assay (Thermo Fisher Scientific) and samples were stored at -80°C until use. Conditions of SIRT2 overexpression in HCT116 cells are described previously63.

Mass spectrometry

A total of 500 µg of soluble protein lysate from WT control human colorectal carcinoma HCT116 cells (denoted WT) or HCT116or cells overexpressing HA-SET8 (described above; denoted SET8 OE) was digested with Arg-C (Roche, 11370529001). First, protein lysates were diluted in Arg-C digestion buffer to final concentrations of 100 mM Tris-HCl, 10 mM CaCl2 (pH 7.6), followed by standard reduction and alkylation at room temperature. Briefly, proteins were first reduced with 3 mM TCEP for 45 minutes, alkylated with 15 mM iodoacetamide (IAA) for 60 minutes, followed by quenching of any unused IAA by incubation with 20 mM DTT for 45 minutes. Next, Arg-C proteases were added to a final ratio of 1:200 protein-to-protease ratio. The digestion occurred overnight at 37  oC with end-to-end rotation. Resultant peptides were desalted C18 Spin Columns (Thermo Fisher Scientific, 89870) and C18-ZipTip (Millipore Sigma, ZTC18S096) according to manufacturer’s instructions, dried by SpeedVac, and resuspended in 20 μL of MS-grade water supplemented with 0.1% formic acid. Samples were stored at -80°C until use.

To assess mono-methylation levels of identified in vitro SET8 substrates, parallel reaction monitoring (PRM)-MS was performed using a Q-Exactive Plus hybrid quadrupole-orbitrap mass spectrometer, at the John L. Holmes Mass Spectrometry Facility at the University of Ottawa, as previously described95. Briefly, the mass spectrometer was set to perform a full MS scan (350-1,600 m/z) at a resolution of 70,000 (at 200 m/z), followed by PRM scans at a resolution of 17,500 (at 200 m/z). PRM-MS scanning was guided by an isolation list built in Skyline Software88 (Supplementary Materials S3.4)94,96. The isolation list was generated by first identifying the primary interactors of SET8, isolated from the STRING human interactome dataset using Python 3 and the Pandas package8285,95,97. The interactors of the primary interactors of SET8 (i.e., secondary interactors) were also included in this list. Proteins which contained the predicted and validated SET8 in vitro methylation sites from peptide SPOT array experiments were selected from within the primary and secondary interactors of SET8, and the resulting network was mapped with the NetworkX package for Python98. The verified sites contained within this sub-network were then used to generate an isolation list for targeted MS analysis. To quantify relative mono-methylation levels of detectable target peptides, Skyline software was used to design an isolation list and was used to derive the average total peak areas for modified target substrates as well as other unmodified peptides along the protein sequence. The average total peak of each modification site of interest was generated from three pooled biological samples. Pooled samples were used to enhance detection efficiency. The average peak intensity from these samples was then divided by that of all other reliably detected unmodified peptides within a given parental protein (PeptideAtlas repository (project ID: PASS05848)). The MS-verified sites for SIRT2 were obtained from the dataset provided in Zhang et. al., 202263.

Application to the cancer proteome

Targeted screen mutant repositories were obtained for breast cancer from v96 of the COSMIC database (cancer.sanger.ac.uk)99. In Python 3, the data were cleaned to isolate missense mutations, which were then applied to the full human proteome obtained from UniProt84,92. Mutations which occur within ±7 AAs from either (1) a lysine, or (2) a position mutated to a lysine, were accepted within the oncoproteome set. An accompanying dataset for the healthy proteome was generated. The ML-hybrid ensemble model for SET8 methylation prediction was applied to both the healthy and cancerous dataset. The resulting scores were compared, yielding several potential states for mutations; (1) mutations predicted to have a null effect on SET8 methylation (i.e., no change from healthy population), (2) mutations predicted to remove SET8 methylation activity (i.e., loss of SET8 substrate from healthy population), and (3) mutations predicted to induce SET8 methylation (i.e., gain of SET8 from health population). The human interactome from the STRING database was applied to the mutated proteins and the probability scores from the SET8 ML-hybrid ensemble model were scaled to resemble STRING scores and included within the interaction list as well. The network was mapped with Cytoscape (v3.9.1), and molecular subcomplexes were isolated using the Molecular Complex Detection cluster generator from the clusterMaker (v2.0) toolset100102. The subcomplex which contained SET8 was further investigated for proteins implicated in pathways associated with the cancer type, and resulting pathways were represented in Adobe Illustrator.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

42004_2025_1717_MOESM2_ESM.pdf (135.1KB, pdf)

Description of Additional Supplementary Files

Supplementary Data 1 (13.7KB, xlsx)
Supplementary Data 2 (357.3KB, xlsx)
Supplementary Data 3 (44.2KB, xlsx)
Supplementary Data 4 (776.1KB, xlsx)
Supplementary Data 5 (52.4KB, xlsx)
Supplementary Data 6 (50.9KB, xlsx)
Supplementary Data 7 (39.8KB, xlsx)
Supplementary Data 8 (7.4MB, xlsx)
Supplementary Data 9 (193.9KB, xlsx)
Supplementary Data 10 (17.8KB, xlsx)
Supplementary Data 11 (883.2KB, xlsx)
Supplementary Data 12 (134.4KB, xlsx)
Supplementary Data 13 (18.3MB, xlsx)
Supplementary Data 14 (20.6KB, xlsx)
Supplementary Data 15 (348.9KB, xlsx)
Reporting Summary (77.3KB, pdf)

Acknowledgements

We thank the Natural Science and Engineering Research Council of Canada (NSERC) (Grant No. RGPIN-2023-04651) to K.K.B. We extend our gratitude to Ruofan Wang and Ali H. Shukri for their contributions to the mass spectrometry analyses. We thank Lucas Ridgeway for infographic and visual dissemination consultation as per universal accessibility guidelines, and Leanne Ridgeway for copy editing.

Author contributions

The experiments were conceived by K.K.B. and N.H.R. Peptide array and protein expression experiments along with the related data processing was carried out by N.H.R, with support from M.F., D.L., and K.K.B. Mass spectrometry experiments were carried out by N.H.R., A.C. and V.L., with the associated data processing provided by A.C. with support from K.K.B. The machine learning experiments were conceived by N.H.R. and J.R.G. and performed and analyzed by N.H.R. MethylSight data was provided by F.C. The data analysis and visualisations were performed by N.H.R. with support from J.R.G. and K.K.B. The paper was written by N.H.R., D.L., J.R.G. and K.K.B.

Peer review

Peer review information

Communications Chemistry thanks Albert Jeltsch and the other, anonymous, reviewers for their contribution to the peer review of this work.

Data availability

The data conveying all results reported may be readily obtained within the github repository https://github.com/nashirag/ML-Hybrid_Ensemble_Method, or through Zenodo (10.5281/zenodo.17082245). Source data for all figures is accessible within the same repository, including all SPOT peptide array densitometry figures are available. All mass spectrometry data is available through the PeptideAtlas repository (project ID: PASS05848).

Code availability

The Python code implemented to produce all results may be accessed at https://github.com/nashirag/ML-Hybrid_Ensemble_Method.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Supplementary information

The online version contains supplementary material available at 10.1038/s42004-025-01717-6.

References

  • 1.Brandi, J., Noberini, R., Bonaldi, T. & Cecconi, D. Advances in enrichment methods for mass spectrometry-based proteomics analysis of post-translational modifications. J. Chromatogr. A1678, 463352 (2022). [DOI] [PubMed] [Google Scholar]
  • 2.Deribe, Y. L., Pawson, T. & Dikic, I. Post-translational modifications in signal integration. Nat. Struct. Mol. Biol.17, 666–672 (2010). [DOI] [PubMed] [Google Scholar]
  • 3.Liu, J., Qian, C. & Cao, X. Post-translational modification control of innate immunity. Immunity45, 15–30 (2016). [DOI] [PubMed] [Google Scholar]
  • 4.Qian, M. et al. Targeting post-translational modification of transcription factors as cancer therapy. Drug Discov. Today25, 1502–1512 (2020). [DOI] [PubMed] [Google Scholar]
  • 5.Rauh, D. et al. An acetylome peptide microarray reveals specificities and deacetylation substrates for all human sirtuin isoforms. Nat. Commun.4, 2327 (2013). [DOI] [PubMed] [Google Scholar]
  • 6.Merbl, Y. & Kirschner, M. W. Large-scale detection of ubiquitination substrates using cell extracts and protein microarrays. Proc. Natl Acad. Sci. USA.106, 2543–2548 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Moore, K. E. & Gozani, O. An unexpected journey: Lysine methylation across the proteome. Biochim. Biophys. Acta (BBA) Gene Reg. Mech.1839, 1395–1403 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Polo, S. et al. A single motif responsible for ubiquitin recognition and monoubiquitination in endocytic proteins. Nature416, 451–455 (2002). [DOI] [PubMed] [Google Scholar]
  • 9.Rathert, P., Zhang, X., Freund, C., Cheng, X. & Jeltsch, A. Analysis of the substrate specificity of the dim-5 histone lysine methyltransferase using peptide arrays. Chem. Biol.15, 5–11 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Rathert, P. et al. Protein lysine methyltransferase G9a acts on non-histone targets. Nat. Chem. Biol.4, 344–346 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Kudithipudi, S., Dhayalan, A., Kebede, A. F. & Jeltsch, A. The SET8 H4K20 protein lysine methyltransferase has a long recognition sequence covering seven amino acid residues. Biochimie94, 2212–2218 (2012). [DOI] [PubMed] [Google Scholar]
  • 12.Kudithipudi, S., Kusevic, D., Weirich, S. & Jeltsch, A. Specificity analysis of protein lysine methyltransferases using SPOT peptide arrays. J. Vis. Exp. 99, e52203 (2014). [DOI] [PMC free article] [PubMed]
  • 13.Chopra, A., Willmore, W. G. & Biggar, K. K. Insights into a cancer-target demethylase: substrate prediction through systematic specificity analysis for KDM3A. Biomolecules12, 641 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Bradley, D. et al. The substrate quality of CK2 target sites has a determinant role on their function and evolution. Cell Syst.15, 544–562.e8 (2024). [DOI] [PubMed] [Google Scholar]
  • 15.Mitchell, C. J. et al. Unbiased identification of substrates of protein tyrosine phosphatase ptp-3 in C. elegans. Mol. Oncol.10, 910–920 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Yu-Ying, Y., Markus, G. & Howard, H. C. Identification of lysine acetyltransferase p300 substrates using 4-pentynoyl-coenzyme A and bioorthogonal proteomics. Bioorg. Med. Chem. Lett.21, 4976–4979 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Biggar, K. K. et al. Proteome-wide prediction of lysine methylation leads to identification of H2BK43 methylation and outlines the potential methyllysine proteome. Cell Rep.32, 107896 (2020). [DOI] [PubMed] [Google Scholar]
  • 18.Jamal, S., Ali, W., Nagpal, P., Grover, A. & Grover, S. Predicting phosphorylation sites using machine learning by integrating the sequence, structure, and functional information of proteins. J. Transl. Med.19, 218 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Kiemer, L., Bendtsen, J. D. & Blom, N. NetAcet: prediction of N-terminal acetylation sites. Bioinformatics21, 1269–1270 (2005). [DOI] [PubMed] [Google Scholar]
  • 20.Neely, B. A. et al. Toward an integrated machine learning model of a proteomics experiment. J. Proteome Res.22, 681–696 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Deng, W. et al. GPS-PAIL: prediction of lysine acetyltransferase-specific modification sites from protein sequences. Sci. Rep.6, 39787 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Wu, Z., Lu, M. & Li, T. Prediction of substrate sites for protein phosphatases 1B, SHP-1, and SHP-2 based on sequence features. Amino Acids46, 1919–1928 (2014). [DOI] [PubMed] [Google Scholar]
  • 23.Wang, X. et al. UbiBrowser 2.0: a comprehensive resource for proteome-wide known and predicted ubiquitin ligase/deubiquitinase–substrate interactions in eukaryotic species. Nucleic Acids Res.50, D719–D728 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Smith, K., Rhoads, N. & Chandrasekaran, S. Protocol for CAROM: a machine learning tool to predict post-translational regulation from metabolic signatures. STAR Protoc.3, 101799 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Lanouette, S. et al. Discovery of substrates for a SET domain lysine methyltransferase predicted by multistate computational protein design. Structure23, 206–215 (2015). [DOI] [PubMed] [Google Scholar]
  • 26.Ferrari, E. et al. Identification of new substrates of the protein-tyrosine phosphatase PPT1B by bayesian integration of proteome evidence. J. Biol. Chem.286, 4173–4185 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Vinogradov, A. A., Chang, J. S., Onaka, H., Goto, Y. & Suga, H. Accurate models of substrate preferences of post-translational modification enzymes from a combination of mRNA display and deep learning. ACS Cent. Sci.8, 814–824 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Fang, J. et al. Purification and functional characterization of SET8, a nucleosomal histone H4-lysine 20-specific methyltransferase. Curr. Biol.12, 1086–1099 (2002). [DOI] [PubMed] [Google Scholar]
  • 29.Milite, C. et al. The emerging role of lysine methyltransferase SETD8 in human diseases. Clin. Epigenet8, 102 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Biggar, K. K., Wang, Z. & Li, S. S.-C. SnapShot: Lysine methylation beyond histones. Mol. Cell68, 1016–1016.e1 (2017). [DOI] [PubMed] [Google Scholar]
  • 31.Zhang, H. et al. SET8 prevents excessive DNA methylation by methylation-mediated degradation of UHRF1 and DNMT1. Nucleic Acids Res.47, 9053–9068 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Chin, H. G. et al. The microtubule-associated histone methyltransferase SET8, facilitated by transcription factor LSF, methylates α-tubulin. J. Biol. Chem.295, 4748–4759 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Wu, Q.-J. et al. The sirtuin family in health and disease. Sig Transduct. Target Ther.7, 402 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Hornbeck, P. V. et al. PhosphoSitePlus, 2014: mutations, PTMs and recalibrations. Nucleic Acids Res.43, D512–D520 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Wang, D. et al. MusiteDeep: a deep-learning based webserver for protein post-translational modification site prediction and visualization. Nucleic Acids Res.48, W140–W146 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Yin, Y. et al. SET8 recognizes the sequence RHRK20VLRDN within the N terminus of histone H4 and mono-methylates lysine 20. J. Biol. Chem.280, 30025–30031 (2005). [DOI] [PubMed] [Google Scholar]
  • 37.Topcu, E., Ridgeway, N. H. & Biggar, K. K. PeSA 2.0: A software tool for peptide specificity analysis implementing positive and negative motifs and motif-based peptide scoring. Comput. Biol. Chem.101, 107753 (2022). [DOI] [PubMed] [Google Scholar]
  • 38.Yang, K. K., Wu, Z. & Arnold, F. H. Machine-learning-guided directed evolution for protein engineering. Nat. Methods16, 687–694 (2019). [DOI] [PubMed] [Google Scholar]
  • 39.Erjavac, I., Kalafatovic, D. & Mauša, G. Coupled encoding methods for antimicrobial peptide prediction: how sensitive is a highly accurate model?. Artif. Intell. Life Sci.2, 100034 (2022). [Google Scholar]
  • 40.Durant, J. L., Leland, B. A., Henry, D. R. & Nourse, J. G. Reoptimization of MDL keys for use in drug discovery. J. Chem. Inf. Comput. Sci.42, 1273–1280 (2002). [DOI] [PubMed] [Google Scholar]
  • 41.Ruiz-Blanco, Y. B., Paz, W., Green, J. & Marrero-Ponce, Y. ProtDCal: A program to compute general-purpose-numerical descriptors for sequences and 3D-structures of proteins. BMC Bioinforma.16, 162 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Romero-Molina, S., Ruiz-Blanco, Y. B., Green, J. R. & Sanchez-Garcia, E. ProtDCal-Suite: A web server for the numerical codification and functional analysis of proteins. Protein Sci.28, 1734−1743 (2019). [DOI] [PMC free article] [PubMed]
  • 43.Szeghalmy, S. & Fazekas, A. A comparative study of the use of stratified cross-validation and distribution-balanced stratified cross-validation in imbalanced learning. Sensors23, 2333 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Burkov, A. The Hundred-PageMachine Learning Book Hard Cover ed. edition, Vol. 160 (Andriy Burkov, Polen, 2019).
  • 45.Brownlee, J. Imbalanced Classification with Python: Choose Better Metrics, Balance Skewed Classes, and Apply Cost-Sensitive Learning. (Machine Learning Mastery, 2021).
  • 46.Izenman, A. J. Linear discriminant analysis. In Modern Multivariate Statistical Techniques (ed. Izenman, A. J.) 237–280 (Springer New York, 2013).
  • 47.Kamalov, F., Leung, H.-H. & Cherukuri, A. K. Keep it simple: random oversampling for imbalanced data. In 2023Advances in Science and Engineering Technology International Conferences (ASET) 1–4 (IEEE, Dubai, 2023).
  • 48.Wright, R. E. Logistic regression. Read. Underst. Multivar. Stat.217, 244 (1995). [Google Scholar]
  • 49.Chawla, N. V., Bowyer, K. W., Hall, L. O. & Kegelmeyer, W. P. SMOTE: Synthetic minority over-sampling technique. jair16, 321–357 (2002). [Google Scholar]
  • 50.Nguyen, H. M., Cooper, E. W. & Kamei, K. Borderline over-sampling for imbalanced data classification. Int. J. Knowl. Eng. Soft Data Paradig.3, 4–21 (2009). [Google Scholar]
  • 51.Petersen, B., Petersen, T. N., Andersen, P., Nielsen, M. & Lundegaard, C. A generic method for assignment of reliability scores applied to solvent accessibility predictions. BMC Struct. Biol.9, 51 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Baryshnikova, A. Spatial Analysis of Functional Enrichment (SAFE) in Large Biological Networks. In Computational Cell Biology (eds von Stechow, L. & Santos Delgado, A.) 249–268 (Springer New York, 2018). [DOI] [PubMed]
  • 53.The Gene Ontology, C. onsortium et al. The gene ontology resource: enriching a GOld mine. Nucleic Acids Res.49, D325–D334 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Luck, K. et al. A reference map of the human binary protein interactome. Nature580, 402–408 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Couture, J.-F., Collazo, E., Brunzelle, J. S. & Trievel, R. C. Structural and functional analysis of SET8, a histone H4 Lys-20 methyltransferase. Genes Dev.19, 1455–1465 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Kaczmarek Michaels, K., Mohd Mostafa, S., Ruiz Capella, J. & Moore, C. L. Regulation of alternative polyadenylation in the yeast saccharomyces cerevisiae by histone H3K4 and H3K36 methyltransferases. Nucleic Acids Res.48, 5407–5425 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Liu, B. et al. A functional single nucleotide polymorphism of SET8 is prognostic for breast cancer. Oncotarget7, 34277–34287 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Jørgensen, S. et al. The histone methyltransferase SET8 is required for S-phase progression. J. Cell Biol.179, 1337–1345 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Yang, C., Wang, K., Zhou, Y. & Zhang, S.-L. Histone lysine methyltransferase SET8 is a novel therapeutic target for cancer treatment. Drug Discov. Today26, 2423–2430 (2021). [DOI] [PubMed] [Google Scholar]
  • 60.Bogliolo, M. et al. Mutations in ERCC4, encoding the DNA-repair endonuclease XPF, cause fanconi anemia. Am. J. Hum. Genet.92, 800–806 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Faridounnia, M., Folkers, G. & Boelens, R. Function and interactions of ERCC1-XPF in DNA damage response. Molecules23, 3205 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Xu, L. et al. Roles for the methyltransferase SETD8 in DNA damage repair. Clin. Epigenet14, 34 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Zhang, H. et al. Quantitative proteomic analysis of the lysine acetylome reveals diverse SIRT2 substrates. Sci. Rep.12, 3822 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Levy, D. et al. A proteomic approach for the identification of novel lysine methyltransferase substrates. Epigenetics Chromatin4, 19 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Meng, L. et al. Mini-review: recent advances in post-translational modification site prediction based on deep learning. Comput. Struct. Biotechnol. J.20, 3522–3532 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Schwartz, D. Prediction of lysine post-translational modifications using bioinformatic tools. Essays Biochem.52, 165–177 (2012). [DOI] [PubMed] [Google Scholar]
  • 67.Shilatifard, A. The COMPASS Family of histone H3K4 methylases: mechanisms of regulation in development and disease pathogenesis. Annu. Rev. Biochem.81, 65–95 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Weber, L. M. et al. The histone acetyltransferase KAT6A is recruited to unmethylated CpG islands via a DNA binding winged helix domain. Nucleic Acids Res.51, 574–594 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Shinsky, S. A., Monteith, K. E., Viggiano, S. & Cosgrove, M. S. Biochemical reconstitution and phylogenetic comparison of human SET1 family core complexes involved in histone methylation. J. Biol. Chem.290, 6361–6375 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Rienzo, M. et al. PRDM12 in health and diseases. IJMS22, 12030 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Hashimoto, K., Wada, K., Matsumoto, K. & Moriya, M. Physical interaction between SLX4 (FANCP) and XPF (FANCQ) proteins and biological consequences of interaction-defective missense mutations. DNA Repair35, 48–54 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Bakker, J. L. et al. Analysis of the novel fanconi anemia gene SLX4 / FANCP in familial breast cancer cases. Hum. Mutat.34, 70–73 (2013). [DOI] [PubMed] [Google Scholar]
  • 73.Grimes, M. et al. Integration of protein phosphorylation, acetylation, and methylation data sets to outline lung cancer signaling networks. Sci. Signal.11, eaaq1087 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Leutert, M., Entwisle, S. W. & Villén, J. Decoding post-translational modification crosstalk with proteomics. Mol. Cell. Proteom.20, 100129 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Yang, J., Hu, Y., Zhang, B., Liang, X. & Li, X. The JMJD Family histone demethylases in crosstalk between inflammation and cancer. Front. Immunol.13, 881396 (2022). [DOI] [PMC free article] [PubMed]
  • 76.Ikram, S. et al. The SMYD3-MAP3K2 signaling axis promotes tumor aggressiveness and metastasis in prostate cancer. Sci. Adv.9, eadi5921 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Heinonen, T. et al. Dual deletion of the sirtuins SIRT2 and SIRT3 impacts on metabolism and inflammatory responses of macrophages and protects from endotoxemia. Front. Immunol.10, 2713 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Chen, L. T. et al. Target sequence-conditioned design of peptide binders using masked language modeling. Nat. Biotechnol.10.1038/s41587-025-02761-2 (2025). [DOI] [PMC free article] [PubMed]
  • 79.Rathore, A. S., Kumar, N., Choudhury, S., Mehta, N. K. & Raghava, G. P. S. Prediction of hemolytic peptides and their hemolytic concentration. Commun. Biol.8, 176 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Chopra, A. et al. A peptide array pipeline for the development of spike-ACE2 interaction inhibitors. Peptides158, 170898 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Hilpert, K., Winkler, D. F. & Hancock, R. E. Cellulose-bound peptide arrays: preparation and applications. Biotechnol. Genet. Eng. Rev.24, 31–106 (2007). [DOI] [PubMed] [Google Scholar]
  • 82.Bradford, M. M. A rapid and sensitive method for the quantitation of microgram quantities of protein utilizing the principle of protein-dye binding. Anal. Biochem.72, 248–254 (1976). [DOI] [PubMed] [Google Scholar]
  • 83.Hornbeck, P. V. et al. PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Res.40, D261–D270 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Rossum, G. V. & Drake, F. L. Python 3 Reference Manual. (CreateSpace, 2009).
  • 85.McKinney, W. Data structures for statistical computing in python. In Proceedings of the 9th Python in Science Conference 56–61 (IEEE, 2010).
  • 86.Rowe, E. M. & Biggar, K. K. An optimized method using peptide arrays for the identification of in vitro substrates of lysine methyltransferase enzymes. MethodsX5, 118–124 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Pedregosa, F. et al. Scikit-learn: machine learning in python. J. Mach. Learn. Res. 12, 2825−2830 (2018).
  • 88.Lemaitre, G., Nogueira, F. & Aridas, C. K. Imbalanced-learn: a python toolbox to tackle the curse of imbalanced datasets in machine learning. J. Mach. Learn. Res. 18, 1−5 (2016).
  • 89.Harris et al. Array programming with NumPy. Nature585, 357–362 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Hunter, J. D. Matplotlib: a 2D graphics environment. Comput. Sci. Eng.9, 90–95 (2007). [Google Scholar]
  • 91.Dietterich, T. G. Ensemble methods in machine learning. In Multiple Classifier Systems, 1–15 (Springer Berlin Heidelberg, 2000).
  • 92.The UniProt Consortium et al. UniProt: The universal protein knowledgebase in 2023. Nucleic Acids Res.51, D523–D531 (2023). [DOI] [PMC free article] [PubMed]
  • 93.Krzywinski, M. et al. Circos: An information aesthetic for comparative genomics. Genome Res.19, 1639–1645 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Lex, A., Gehlenborg, N., Strobelt, H., Vuillemot, R. & Pfister, H. UpSet: Visualization of intersecting sets. IEEE Trans. Vis. Comput. Graph.20, 1983–1992 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Rosario, F. J. et al. Placental remote control of fetal metabolism: trophoblast mTOR signaling regulates liver IGFBP-1 phosphorylation and IGF-1 bioavailability. IJMS24, 7273 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96.MacLean, B. et al. Skyline: an open source document editor for creating and analyzing targeted proteomics experiments. Bioinformatics26, 966–968 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Szklarczyk, D. et al. STRING v10: protein–protein interaction networks, integrated over the tree of life. Nucleic Acids Res.43, D447–D452 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Hagberg, A. A., Schult, D. A. & Swart, P. J. Exploring network structure, dynamics, and function using NetworkX. In Proceedings of the 7th Python in Science Conference (eds Varoquaux, G., Vaught, T. & Millman, J.) 11–15 (Pasadena, CA USA, 2008).
  • 99.Tate, J. G. et al. COSMIC: The catalogue of somatic mutations in cancer. Nucleic Acids Res.47, D941–D947 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 100.Shannon, P. et al. Cytoscape: A software environment for integrated models of biomolecular interaction networks. Genome Res.13, 2498–2504 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101.Morris, J. H. et al. clusterMaker: a multi-algorithm clustering plugin for cytoscape. BMC Bioinforma.12, 436 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 102.Bader, G. D. & Hogue, C. W. An automated method for finding molecular complexes in large protein interaction networks. BMC Bioinforma.4, 2 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

42004_2025_1717_MOESM2_ESM.pdf (135.1KB, pdf)

Description of Additional Supplementary Files

Supplementary Data 1 (13.7KB, xlsx)
Supplementary Data 2 (357.3KB, xlsx)
Supplementary Data 3 (44.2KB, xlsx)
Supplementary Data 4 (776.1KB, xlsx)
Supplementary Data 5 (52.4KB, xlsx)
Supplementary Data 6 (50.9KB, xlsx)
Supplementary Data 7 (39.8KB, xlsx)
Supplementary Data 8 (7.4MB, xlsx)
Supplementary Data 9 (193.9KB, xlsx)
Supplementary Data 10 (17.8KB, xlsx)
Supplementary Data 11 (883.2KB, xlsx)
Supplementary Data 12 (134.4KB, xlsx)
Supplementary Data 13 (18.3MB, xlsx)
Supplementary Data 14 (20.6KB, xlsx)
Supplementary Data 15 (348.9KB, xlsx)
Reporting Summary (77.3KB, pdf)

Data Availability Statement

The data conveying all results reported may be readily obtained within the github repository https://github.com/nashirag/ML-Hybrid_Ensemble_Method, or through Zenodo (10.5281/zenodo.17082245). Source data for all figures is accessible within the same repository, including all SPOT peptide array densitometry figures are available. All mass spectrometry data is available through the PeptideAtlas repository (project ID: PASS05848).

The Python code implemented to produce all results may be accessed at https://github.com/nashirag/ML-Hybrid_Ensemble_Method.


Articles from Communications Chemistry are provided here courtesy of Nature Publishing Group

RESOURCES