Summary
Synonymous variations, once considered irrelevant to gene regulation or disease progression, are now known to play a significant role in human diseases. We developed a one-stop resource, comprehensive database for deleterious synonymous variation prediction (CDsyn) to assess their significance and disease relevance. CDsyn encompasses six categories of information, including existing prediction scores, conservation scores, translation efficiency, sequence information, population frequency and other annotation information. With the assistance of CDsyn, we emphasize the significance of splicing mutation prediction score in identifying pathogenic synonymous mutations. Furthermore, we demonstrated the feasibility of developing a pathogenicity prediction algorithm for synonymous variations with CDsyn. In two practical tests, the CDsyn-comprehensive approach (along with its internal methods) outperformed both the SynMICdb and dbDSM databases, as well as the InterVar tool. CDsyn facilitates the prioritization of deleterious variants in clinical sequencing contexts and can enhance the understanding of synonymous variations in relation to human diseases.
Subject areas: mathematical biosciences, biological database, genomic analysis, sequence analysis
Graphical abstract

Highlights
-
•
A comprehensive annotation database of human synonymous variation
-
•
CDsyn enables direct development for deleterious synonymous variation prediction models
-
•
User-friendly web server and extensive annotation source meet diverse needs
Mathematical biosciences; biological database; genomic analysis; sequence analysis
Introduction
Synonymous variations are one of the most abundant DNA mutations in human genomics, which have been shown to play crucial roles in approximately 90 diseases.1,2 However, synonymous mutations have been overlooked in many past studies because of their “silent” attribute at the protein level. Nonetheless, with the deeper insights into mutations and their disease-causing mechanisms, synonymous mutations have returned to the center stage of scientific research.1 Synonymous variations are believed to contribute to tumor development by affecting the adjacent splice sites, changing RNA stability and RNA folding, and disrupting protein folding.3,4 Furthermore, Zhang et al. experimentally demonstrated that most synonymous mutations in representative yeast genes are strongly non-neutral.5,6 Xie et al. provided evidence that synonymous mutations could promote tumorigenesis by disrupting m6A-dependent mRNA metabolism.7 Wei et al. revealed functional synonymous mutations in human cells through prime editor-based high-throughput screening.8 These recent studies have underlined the pivotal role of synonymous mutations in human health. Consequently, prioritizing mechanistic investigations into synonymous mutations is imperative, given their underappreciated contributions to disease pathogenesis.9
Large-scale databases significantly contribute to elucidating the mechanism behind disease pathogenesis. In the field of non-synonymous variation research, dbNSFP is a popular database that aids in variants studies.10,11,12,13 It developed from version 1.0 to version 4.0 and is still updating. More importantly, dbNSFP accelerated the development of genomics and precision medicine to some extent. The dbNSFP database has provided comprehensive functional annotations, allowing the construction of a large number of algorithms and prediction models relatively easily. At the genomics level, the WGSA database, a variant annotation database of all kinds of alterations, has also been developed.14 Subsequently, dbDSM, a database of deleterious synonymous mutations, emerged.15 However, current databases exhibit some limitations. First, there is a shortage of a comprehensive database of synonymous variations. For example, WGSA contained only one synonymous mutation annotation label, which was derived from the usDSM method. Genomics researchers cannot obtain all information related to synonymous mutations in a one-stop platform. Second, synonymous mutations may cause disease by affecting splicing, but the features related to splicing mutations are either not included in the existing databases or are only described in a rudimentary manner. These databases mainly contain common features, for example, algorithm results, conservation scores, and allele frequencies of large cohorts. Third, the number of synonymous mutations collected in existing databases is incomplete. Although dbDSM collected possible damaging synonymous mutations at the two-thousand, it did not integrate with other known annotation resources. Another famous database, SynMICdb, contains no more than 700,000 pieces of data.4 Fourth, the existing databases are not compatible with the multiple genome versions. Though hg38 gradually became the main genome version in many research projects, chm13v2 is the most complete genome version to date, and the hg19 version is still adopted by multiple databases and information sources (Table S1). Thus, there is an urgent need to build a comprehensive and up-to-date database of synonymous mutations to meet all kinds of needs in precision medicine.
In this study, we establish a comprehensive database for deleterious synonymous variation prediction, CDsyn, which includes as many existing algorithm prediction results as possible, especially splicing mutation prediction features. Additionally, evolution conservation scores, sequence information, translation efficiency, allele frequencies of well-known large cohorts, and other features, such as classification tags of the InterVar and ClinVar, are included. To our knowledge, CDsyn is the most comprehensive database designed especially for synonymous variations. It is a one-stop platform that contains plentiful information about how to evaluate synonymous mutations. Because CDsyn possesses complete synonymous mutations identified in the RefSeq gene, it can be used directly to obtain all kinds of annotation results, eliminating the need to frequently search for variations on many websites or installing software locally first. Since CDsyn provides richly annotated resources matched to the three most commonly used genome versions, we could utilize it to develop algorithms in genomics. More importantly, we firmly believe that the use of the CDsyn database is beneficial for identifying pathogenic synonymous mutations in the human genome, which can help elucidate the role of synonymous mutations in the occurrence and progression of diseases and promote the rapid development of precision medicine.
Results
Database contents
CDsyn categorized six types of information for synonymous variation annotation (Figure 1) and provided 23,427,958 synonymous variation annotated by ANNOVAR in total. The number dropped to 23,388,493 in hg19 and 23,339,382 in chm13v2 (Figure 2 and Table S2). The distribution of synonymous variations in various chromosomes in the three reference genome versions is depicted in Figure 2. Apparently, the hg38 version contains the most comprehensive synonymous mutation data. By providing comprehensive genomic reference version annotations (e.g., hg19 and hg38), the CDsyn database accommodates diverse query requirements across different genomic coordinate systems, thus ensuring compatibility with broad research applications.
Figure 1.
The data structure of CDsyn
The CDsyn is comprised of pathogenicity prediction scores, evolutionary conservation, translation efficiency, sequence information, population frequency, and other information.
Figure 2.
The distribution of synonymous mutations at the chromosome level in three genome reference versions is included in the CDsyn
The hg38 version of the genome is the main reference version, and the hg19 and chm13v2 versions are converted based on the hg38 version.
See also Figures S1 and S2.
Available algorithms are collected in CDsyn, which includes the general prediction methods, the specific pathogenicity prediction methods of synonymous variations, and the splicing mutation prediction methods. The coverage of each algorithm is shown in Figure 3. CADD is a unique algorithm that has complete coverage. Except for Spidex, fathmm_xf_noncoding, DDIG, DANN, Syntool, and dbscSNV, the coverage of the remaining algorithms is over 90%. As no specific method has complete coverage, developing a specific model with 100% coverage is needed. Figure 4 shows the frequency of rank scores from 19 algorithms in different ranges. An interesting observation is that the distribution of rank scores from AbSplice and SpliceAI is uneven, as depicted in Figure S3 and Table S3.
Figure 3.
The coverage of the prediction results of 19 algorithms recorded in CDsyn
The data points exhibiting a coverage rate below 90 are rendered in light gray coloration.
Figure 4.
Distributions of the rank scores of 27 prediction results
See also Figure S3.
The correlation between rank scores of deleteriousness prediction and conservation
CDsyn provided 27 prediction scores from 19 methods and 10 conservation scores from 5 algorithms. To analyze the relationship among those methods, we computed Pearson’s correlation coefficients between each pair. Fathmm_xf_noncoding, dbscSNV_ADA_SCORE, and dbscSNV_RF_SCORE were not adopted for analysis because they possess a large amount of missing data. We observed that most general prediction methods have a weak correlation (0–0.4) with splicing mutation prediction methods or evolutionary conservation scoring methods. In synonymous mutation-specific prediction algorithms, only PrDSM has a medium correlation with Phast_cons.mammalian, Phast_cons.vertebrate, Phylop.vertebrate, and Gerp++gt2 (0.4–0.6). CADD, Eigen, fathmm.MKL_coding, fathmm.MKL_noncoding, and three general prediction methods have a moderate relation with evolutionary conservation scoring methods. In addition, DDIG has a medium negative correlation with spidex_dpsi_Z score (−0.6 to −0.4) (Figure 5A and Table S4). Using 1- correlation coefficients as the distance measure, we constructed a UPGMA (unweighted pair-group method with arithmetic mean) dendrogram of the rank scores (Figure 5B).13 Interestingly, seven of nine splicing mutation prediction scores formed a cluster, while the remaining two scores, synonymous mutation prediction scores, and conservation scores formed another cluster. Notably, values derived from distinct algorithmic categories (e.g., splicing effect predictors vs. conservation metrics) exhibited no co-localization on the terminal phylogenetic branches. This mutually exclusive pattern supports the feasibility of constructing a composite prediction model for synonymous variation pathogenicity through systematic integration of splicing mutation potential and evolutionary conservation scores.
Figure 5.
The correlation between deleteriousness prediction scores and conservation scores
(A) Pearson’s correlation coefficients between rank scores of deleteriousness prediction and conservation.
(B) UPGMA dendrogram of the deleteriousness prediction scores and conservation scores.
UPGMA, unweighted pair-group method with arithmetic mean.
The performance of various prediction tools in four testing datasets
To find a prediction model with the best performance in screening for damaging synonymous mutations in CDsyn, we compared 27 prediction scores from 19 approaches. Six prediction scores were not selected because they have too little data to match the test data. In the silico prediction of splicing mutation, dbscSNV demonstrated an excellent performance in 1st evaluation dataset and 3rd evaluation dataset (Figures 6A and 6C). However, the prediction coverage of dbscSNV was small; thus, we suppose it is not appropriate for pathogenic synonymous mutation prediction as a solo tool. Taken together, SilVA and EnDSM are the outstanding performance synonymous mutation-specific tools, and fathmm_xf_coding is the general algorithm that had a relatively superior performance in the four evaluation datasets. As a splicing mutation prediction tool, AbSplice had the highest AUC value with the highest coverage of prediction (Figure 6).
Figure 6.
Performance of the prediction tools of CDsyn in four evaluation datasets
(A) 1st evaluation dataset.
(B) 2nd evaluation dataset.
(C) 3rd evaluation dataset.
(D) 4th evaluation dataset.
See also Figure S4.
Given that InterVar utilizes a series of evidence to evaluate the pathogenicity of variants, we also executed the InterVar on four evaluation datasets (Figure S4). However, the InterVar could not identify the pathogenic synonymous mutations in the evaluation datasets, possibly because it is dependent on enough evidence to make a correct decision. It indirectly suggested that establishing the prediction model may ameliorate this difficult situation.
Synonymous and splicing mutation features demonstrate high predictive performance in the testing datasets
Through multiple analyses of the prediction ability of the six category features with four testing datasets, we discovered that synonymous mutation prediction scores and splicing mutation prediction scores have relatively high AUC and stable performance (Figure 7). In the process of constructing the ridge model, the above two class features were validated with a large contribution of positive correlation (Figure S5A).
Figure 7.
Utilization of the four evaluation datasets to evaluate the ability of six class features to recognize the deleterious synonymous mutations separately
See also Figures S5 and S6.
What’s more, to make fuller use of pathogenic mutations and benign variants for training predictive models, we analyzed the prediction results of training data with different P:B (pathogenic mutations vs. benign variants) ratios. Under the precondition of maintaining the balance between training data volume and model performance, it is plausible to speculate that data with a P:B ratio of 1.0:1.0 are appropriate as a training set (Figures S5B and S6), which has been adopted in several research.16,17,18
CDsyn-comprehensive approach exhibits a superior capacity in screening pathogenic synonymous mutations
In the task of utilizing the CDsyn database, prioritizing pathogenic synonymous mutations in the ABCA4 and ATP7B genes (Tables 1 and 2 and Tables S5 and S6), we compared the CDsyn-comprehensive approach with selected top-performing prediction tools (fathmm_xf_coding, EnDSM, and AbSplice), an ACMG-AMP-based clinical interpretation tool (InterVar), and other synonymous variation databases (SynMICdb and dbDSM), all of which can be used to evaluate the impact of synonymous variations. SynMICdb achieved 100% accuracy in the tests on two genes. However, it covered only 9.5% and 11.1% of the target variants, respectively. Regrettably, those variants were absent from dbDSM, and InterVar classified all of the variants as benign. As shown in Tables 1 and 2, the prediction tools from CDsyn (fathmm_xf_coding, EnDSM, and AbSplice) exhibited excellent performance. Importantly, the CDsyn-comprehensive approach demonstrates superior capacity compared to individual methods in screening pathogenic synonymous mutations, although one VUS variant was excluded during rate calculation.
Table 1.
CDsyn-related methods outperform the other three tools on ABCA4’s synonymous variations
| CDsyn-comprehensive approach | fathmm_xf_codinga | EnDSMa | Absplicea | InterVar | SynMICdb | dbDSM | |
|---|---|---|---|---|---|---|---|
| Coverage of resultc | 100% | 100% | 100% | 100% | 100% | 9.50% | 0% |
| P(True)/P:VUS:B(Predicted)d | 11/(10:0:1) | 11/(8:0:3) | 11/(8:0:3) | 11/(7:0:4) | 11/(0:0:11) | 11/(2:0:0) | 11/(0:0:0) |
| B(True)/P:VUS:B(Predicted)d | 10/(1:1:8) | 10/(3:0:7) | 10/(6:0:4) | 10/(1:0:9) | 10/(0:0:10) | 10/(0:0:0) | 10/(0:0:0) |
| TPR | 91% | 72.7% | 72.7% | 63.6% | 0% | 100% | – |
| TNR | 88.9%b | 70% | 40% | 90% | 100% | 0% | – |
| Precisione | 91%b | 72.7% | 57.1% | 87.5% | – | 100% | – |
| Accuracye | 90%b | 71.4% | 57.1% | 76.2% | 47.6% | 100% | – |
TPR, true positive rate; TNR, true negative rate.
The algorithms collected in CDsyn.
The VUS variant is removed when calculating the rate.
The percentage of variants included in a method (or database) relative to the total variants provided in the original study.
“P(True)/P:VUS: B(Predicted)” describes scenarios where truly pathogenic mutations are classified by the predictor as pathogenic (P), variants of unknown significance (VUS), or benign (B). Conversely, “B(True)/P:VUS: B(Predicted)” refers to scenarios where truly benign mutations are classified as pathogenic (P), VUS, or benign (B).
Precision and accuracy are computed based on the available variants for each method (or database).
Table 2.
CDsyn-related methods outperform the other three tools on ATP7B’s synonymous variations
| CDsyn-comprehensive approach | fathmm_xf_codinga | EnDSMa | Absplicea | InterVar | SynMICdb | dbDSM | |
|---|---|---|---|---|---|---|---|
| Coverage of resultc | 100% | 100% | 100% | 100% | 100% | 11.10% | 0% |
| P(True)/P:VUS:B(Predicted)d | 4/(1:0:3) | 4/(1:0:3) | 4/(2:0:2) | 4/(1:0:3) | 4/(0:0:4) | 4/(1:0:0) | 4/(0:0:0) |
| B(True)/P:VUS:B(Predicted)d | 5/(0:1:4) | 5/(1:0:4) | 5/(1:0:4) | 5/(3:0:2) | 5/(0:0:5) | 5/(0:0:0) | 5/(0:0:0) |
| TPR | 25% | 25% | 50% | 25% | 0% | 100% | – |
| TNR | 100%b | 80% | 80% | 40% | 100% | 0% | – |
| Precisione | 100%b | 50% | 66.7% | 25% | – | 100% | – |
| Accuracye | 62.5%b | 55.6% | 66.7% | 33.3% | 55.6% | 100% | – |
TPR, true positive rate; TNR, true negative rate.
The algorithms collected in CDsyn.
The VUS variant is removed when calculating the rate.
The percentage of variants included in a method (or database) relative to the total variants provided in the original study.
“P(True)/P:VUS: B(Predicted)” describes scenarios where truly pathogenic mutations are classified by the predictor as pathogenic (P), variants of unknown significance (VUS), or benign (B). Conversely, “B(True)/P:VUS: B(Predicted)” refers to scenarios where truly benign mutations are classified as pathogenic (P), VUS, or benign (B).
Precision and accuracy are computed based on the available variants for each method (or database).
Discussion
CDsyn is a comprehensive database for all potential synonymous mutations in the human genome and their functional predictions. Moreover, it can be queried with variants in three genome versions (hg19, hg38, and chm13v2). CDsyn was developed for two main purposes. The first is to facilitate the filtering or prioritizing of synonymous variations in studies of mapping rare Mendelian disease genes using the whole genome/exome-sequencing approach. The second is to facilitate developing a new deleteriousness prediction model of synonymous mutations. In our research, we not only analyzed the performance of existing prediction tools and developed a new model for predicting harmful synonymous mutations that depend on CDsyn but also explored how to utilize the CDsyn in distinguishing the pathogenic synonymous mutations from benign synonymous variations. Given the considerable value and significance attributed to CDsyn, we plan to continue updating the database to have the most extensive collection of synonymous mutations and more disease-related annotation information.
However, several aspects need to be improved in CDsyn. Firstly, synonymous mutations have been speculated to affect human health by way of miRNA-mediated gene regulation.19MiRNA-related features could be used to enrich the information of CDsyn. Secondly, owing to access restrictions, some well-known tools, such as Trap and regSNPs-splicing, have not been integrated into CDsyn. Thirdly, the synonymous variations collected in CDsyn are based on the annotation of RefSeq genes, and, likely, some synonymous variants defined by other gene databases are not included in CDsyn. These are the aspects that we will take into account for future research. As the evaluation of CDsyn’s performance using synonymous variations of defined pathogenicity in real clinical scenarios was limited to a small number of variants from available reports, larger datasets are necessary to validate our conclusions.
The remarkable progress of deep learning technology has profoundly transformed the field of biology, exemplified by groundbreaking achievements such as AlphaFold’s precise prediction of protein tertiary structures and AlphaMissense’s exceptional performance in predicting missense mutation pathogenicity.20,21 It is worth emphasizing that deep learning is poised to revolutionize the prediction of pathogenicity of synonymous mutations. Although synonymous mutations do not alter amino acid sequences, they modify DNA sequence information and consequently affect RNA sequences. Given that synonymous mutations may induce functional changes through indirect mechanisms—including impacts on RNA processing, stability, or translation efficiency—accurate prediction of RNA spatial structures will significantly enhance our understanding of the pathogenic mechanisms underlying synonymous mutations and improve prediction accuracy.1 While notable RNA spatial structure prediction algorithms such as ARES, RhoFold+, trRosettaRNA, and Nufold have emerged, the absence of a comprehensive RNA 3D structural database persists due to the inherent complexity of RNA structures and the scarcity of experimental data.22,23,24,25,26 Future breakthroughs in this domain will not only deepen insights into the deleterious effects of synonymous mutations but also substantially elevate the precision of predictive methodologies, thereby unraveling the intricate mechanisms underlying human diseases. We will closely monitor and systematically integrate emerging research findings into the CDsyn database to ensure its sustained comprehensiveness and user-friendliness.
In conclusion, we present CDsyn, a significant synonymous mutation database. It not only contains a variety of features about the pathogenicity of synonymous variations but also includes diverse information from multiple disease-related sources. Using this source, the deleterious effects of synonymous mutation can be evaluated comprehensively and scientifically. To facilitate the usage of CDsyn, we provide a user-friendly web-based database, while providing the latest version of CDsyn for batch-fast query. More importantly, we explored the worth of CDsyn in the pathogenicity prediction of synonymous mutations. It is exciting to find that CDsyn is suitable for developing prediction tools for synonymous mutations. We believe CDsyn will serve as a comprehensive and easily accessible tool for functional annotations and predictions of synonymous mutations and help medical geneticists and bioinformatics scholars to uncover the pathogenesis mechanism of synonymous mutations.
Limitations of the study
Our study integrated annotation sources only from genomics and transcriptomics. In the future, epigenomic data should be incorporated into CDsyn to enhance database comprehensiveness. Given the crucial role of RNA structure in evaluating the impact of synonymous variations, we will actively follow cutting-edge advances in RNA structural research and integrate relevant findings into CDsyn.
Resource availability
Lead contact
Further information and requests for resources should be directed to and will be fulfilled by the lead contact, Bing Zeng (bingzengvip@126.com).
Materials availability
This study did not generate new, unique data.
Data and code availability
The CDsyn can be queried online or downloaded from the URL (http://www.cdsyn.com), and the complete data can be obtained from Zenodo. A detailed list of datasets and other items is provided in the following link: CDsyn: https://zenodo.org/records/18320261 (https://doi.org/10.5281/zenodo.18320261).
Code and any additional information required in this paper are available from the lead contact upon request.
STAR★Methods
Key resources table
Method details
Data sources and processing
The version of the human reference genome is GCA_000001405.15; the genes and their corresponding synonymous variations are annotated by ANNOVAR (2020-06-07).56 Although the hg38 version is the main version used in several studies, we still match the variants in the hg19 version and chm13v2 version under CrossMap (v0.7.0), because the hg19 version genome is still used by several algorithms, and the chm13v2 version genome is the latest version that is believed to become mainstream in the future. The CDsyn includes six sections: pathogenicity prediction score, evolutionary conservation, translation efficiency, sequence information, population frequency, and other information (Figure 1). The pathogenicity prediction section can also be divided into three parts: synonymous mutation-specific pathogenicity prediction methods, general prediction methods, and splicing mutation pathogenicity prediction methods.
The existing synonymous mutation-related algorithms are comprised of CADD, DANN, fathmm-MKL, fathmm_xf, PhD-SNPg, DDIG, frDSM, EnDSM, usDSM, PrDSM, SilVA, and Syntool.16,17,18,44,45,47,48,49,50,51,52,61 The first five tools are general prediction methods of deleterious variants, and damaging synonymous mutations are included. The rest of the tools are synonymous mutation-specific pathogenicity prediction tools. Synvep, Spidex, SpliceAI, AbSplice, and dbscSNV are splicing mutation-related algorithms.53,54,62,63,64 Five algorithms (Trap, Predict-SNP2, regSNPs-splicing, MutationTaster2, and SeDSM) are not integrated into our result because they are not available or cannot obtain information in batches.65,66,67,68,69 These prediction results are all collected from the office website or computed by the original algorithm. For AbSplice, we select the maximum of AbSplice_DNA, delta_logit_psi, and delta_psi to describe the result of this method. Delta_psi is one of the prediction scores from SpliceAI and is also recorded in AbSplice. Therefore, we only used the score to represent SpliceAI. For Synvep, given the multiple transcripts of one gene, we compute the maximum, minimum, and mean of the Synvep score to represent this algorithm. To make sure these prediction results are comparable, we create a rank score for each prediction result as in the previous paper.13 Besides prediction methods, evolutionary conservation features are collected to measure the importance of sequence elements. The scores of phastcons, phylop, fitcons, and mappability were obtained from the ‘myvariant’ R package (1.32.0), the gerp++gt2 was downloaded from ANNOVAR, and SilVA provides the score of #GERP++ (Table S7).
Sequence information is defined as features extracted from DNA or RNA sequences. Most of the features in sequence information are from SilVA, including CpG_exon, dMES, etc. Rmsk and TfbsConsSites are two sources downloaded in ANNOVAR.70,71 The former represents the repeat sequences in human genomes, and the variations in repeat regions are supposed to be neutral. The latter is the archive of conserved sites in transcription factor binding sites. If a synonymous variation is associated with tfbsConsSites, the variant tends to be damaging. The translation efficiency is a class of features related to codon bias and autocorrelation,54 it consists of Codon Adaptation Index (CAI), Fraction of Optimal Codons (fracOpt), Codon Usage Bias (CUB), Intrinsic Codon Deviation Index (ICDI), Synonymous Codon Usage Order (SCUO), Codon Autocorrelation Measure (CAM), Change of Frequency measure (CF), and tRNA Adaptation Index (tAI). Those features are computed by codonBiasMetrics (0.1.0), and the context sequence is set at 50 bases. Another two features, #RSCU and dRSCU, are obtained from SilVA.
There are various databases about large cohort studies and disease studies. We collected those databases to establish the relationship between synonymous variations and disease. Population frequencies are significant features used to evaluate the pathogenicity of the mutations. Thus, we incorporated the allele frequencies of gnomad (V4.1), 1000 Genomes, ExAC, and ChinaMAP into CDsyn (Figure S1).27,40,41,42 Some other disease-related data are uniformly classified into the category of “other information” by us. GWAS catalog and GRASP2, two databases collecting information about polymorphic sites and diseases, are integrated into CDsyn.28,29 ClinVar (20240716) is an archive that reports the relationships between human variations and phenotypes, with supporting evidence.30 dbDSM (1.0) is another damaging synonymous variation database,15 and VariSNP (2017-02-16) is a genomic resource for neutral variants, which is deprived of dbSNP.31 De novo mutations are possibly damaging, and Gene4Denovo is an integrated database for de novo mutations, so we add this source to CDsyn.36 COSMIC(V100) and ICGC(V28) are two well-known databases in the cancer field,37,38 since synonymous mutations may play a driver mutation in cancer, adding this information to the database makes it more valuable.3,4 The remaining information included in other information is avsnp151 and InterVar (20210727).32,39 Different from the pathogenicity prediction computational methods, InterVar is a tool for the clinical interpretation of genetic variants by the ACMG-AMP 2015 guidelines,72 the label provided by InterVar probably offers us a more comprehensive opinion. The results of InterVar are computed with the default parameters (Figure S2). Clinically pathogenic variants designated for subsequent analysis were annotated with the “Classification” label and added to the CDsyn database.
The four evaluation datasets
To compare the performance of the existing deleteriousness prediction tools, we constructed four evaluation datasets. As partial algorithms collected in CDsyn were developed before 2018, and VariSNP was used as training data for some algorithms, we used the ClinVar(20180101–20240716) as a test dataset for evaluating the performance of the prediction model.30,31 Furthermore, the intersection between VariSNP and ClinVar was removed. The filtered data from ClinVar was constructed as ‘first test data for evaluation (1st evaluation data)’. ‘Second test data for evaluation (2nd evaluation data)’ and ‘third test data for evaluation (3rd evaluation data)’ were obtained from the usDSM project, as the test data have been generated in recent years.18 The balanced testing dataset from the usDSM project was also used, and named ‘fourth test data for evaluation (4th evaluation data)’. The statistical information of databases and datasets used in our research is listed in Table S8.
The normalization of various features
To explore the value of different features from CDsyn in distinguishing pathogenic synonymous mutation from benign variants, we created a rank score for each feature first, so that the features of the same category are comparable to each other.13 We calculated the rank score of the prediction score of synonymous mutations, and the features of conservation, sequence information, and translation efficiency were also calculated. Then, we used the BPCAfill method in the R package ‘pcaMethods’ (1.94.0) to interpolate the missing data of rank score from pathogenicity prediction and translation efficiency,10,58,73,74 and we used the mean value to interpolate the missing value of conservation,75 sequence information. After that, the mean rank score of each category feature was used to represent that class of features. For population frequency, we used the raw value and interpolated the missing value as zero. We adopted 36 sub-features to describe the impact of population frequency. If the sub-feature is less than 0.001, we score the sub-feature 1/36; if the sub-feature is between 0.001 and 0.01, we score the sub-feature 1/72; if the sub-feature is between 0.01 and 0.05, we score the sub-feature −1/72, and if the sub-feature is greater than 0.05, we score the sub-feature −1/36. At last, the sum of scores of diverse population frequencies is utilized to describe the population frequency. We use the scores calculated from different logics to predict the pathogenicity of synonymous mutations and use the R package ‘pROC’ (1.18.5) to acquire the area under the curve (AUC). The test datasets employed in this section are the same as those used in the evaluation of existing algorithms.
Use CDsyn to develop a prediction method for deleterious synonymous mutations
To evaluate the feasibility and utility of developing synonymous mutation pathogenicity prediction methods using the CDsyn database, we constructed a prediction model employing ridge regression through the R package ‘glmnet’ (4.18). The mean rank scores of pathogenicity prediction, conservation, sequence information, translation efficiency, and population frequency were used as input features. ClinVar (20240716), HGMD (201602), and dbDSM (1.0) were used as the training set.15,30,33 The synonymous variations in ClinVar that were labeled with ‘reviewed by expert panel', ‘criteria provided, multiple submitters, no conflicts' were selected, and the mutations reported as ‘Pathogenic/Likely pathogenic' were regarded as positive data, while the variations reported as ‘Benign/Likely benign' were regarded as negative data. In HGMD, we chose mutations labeled with ‘DM/DM?' as positive data. In dbDSM (1.0), we chose the 300 mutations that were used in many studies as positive data.18,49,50,76 The positive data of three databases comprised pathogenic mutations of the training set (1301) after removing duplicated data, and after removing the mutations that appeared in pathogenic mutations, we defined the negative data of ClinVar as a ‘full benign set’ (114,338). Despite pathogenic synonymous mutations being far less than benign synonymous variations, we want to make the fullest use of these variations with known pathogenicity to train the model. Therefore, we analyze the performance of the prediction model trained on data with various ratios of B:P, ranging from 0.2 to 2, with an increasing value of 0.2. The benign variations were random sampling from the ‘full benign set’.
Since the composition of our training set is similar to the dataset of usDSM, we used the two testing datasets from usDSM18 and used the balanced set from the second testing set as the third testing set. Lastly, through the same filtering process to avoid type I circularity, we removed the intersection between training data and the third testing set to acquire the independent testing set (P: B = 25: 26).
Two case tests regarding CDsyn
To verify the practical utility of the CDsyn database, we systematically collected synonymous variations of ABCA4 and ATP7B from the literature.34,35,77 Subsequently, the fathmm_xf_coding, EnDSM, and Absplice, which possessed outstanding ability in the previous test, were chosen to recognize the deleterious synonymous mutations in two studies individually. More importantly, CDsyn-comprehensive approach, two established databases (dbDSM and SynMICdb), and InterVar were used to prioritize the deleterious synonymous variations.4,15,32 CDsyn-comprehensive approach is a method that utilizes all existing prediction algorithms of CDsyn except Spidex, fathmm_xf_noncoding, DDIG, Syntool, and dbscSNV, as their prediction coverage is less than 90%. In CDsyn-comprehensive approach, we firstly obtained the evaluation result of the synonymous mutations in specific synonymous mutation prediction tools, general mutation prediction tools, and splicing mutation prediction tools separately. A synonymous mutation is labeled as deleterious by synonymous mutation-specific prediction tools or general mutation prediction tools if it was predicted to be harmful in at least four of six tools, or it was predicted to be harmful by splicing mutation prediction tools in at least three of four tools.72 The minimal threshold of the synonymous mutation categorized as benign was the same as before. If a synonymous mutation neither belongs to deleterious variants nor benign variants, it is considered a variant of uncertain significance (VUS). Then, synonymous mutations were classified as deleterious if meeting either of the following criteria: (a) being annotated as deleterious by at least two distinct method categories above; (b) being labeled as deleterious by one method category and non-benign by another. The classification criteria for benign variants followed an analogous approach to those established for pathogenic variants. For other methods, the synonymous mutation with a rank score greater than 0.5 is deemed as deleterious in this section of study.
Database implementation
The website of the CDsyn database (http://www.cdsyn.com) was constructed under the Django open source framework (https://www.djangoproject.com/). All data are compiled into a compressed text file hg38_CDsyn_v1.1.txt.gz, available for download via the URL https://zenodo.org/records/18320261.78 The data can be used to annotate variations directly by ANNOVAR. The data dictionary for supplementary tables is listed in Table S9.
Quantification and statistical analysis
All analyses were performed in Python 3.10, Perl 5.32.1, and R 4.2.3. No statistical analysis was performed.
Acknowledgments
The authors would like to gratefully acknowledge the data contributors and all who offered help for our study.
The work was supported by the National Natural Science Foundation of China (No. 62206164), the Natural Science Foundation of Guangdong Province, China (2024A1515012870), Hunan Provincial Natural Science Foundation of China (2023JJ70007, 2023JJ70039, 2023JJ40001, 2024JJ9007, and 2025JJ90249), the Science and Technology Innovation Program of Hunan Province (2024RC3249), Clinic Research Foundation of Aier Eye Hospital Group (Grant No. AMF2301D42, AGK2301D09, AGF2301D29, AGF2301D33, AGF2301D34, and AIM2507D02), Scientific Research Program of Xiangjiang Philanthropy Foundation (SKY25081 and SKY25084), Shenzhen Medical Research Funding (A2303042), and Shenzhen Science and Technology Program (JCYJ20230807143413027, JCYJ20240813165100001, JCYJ20250604185713018, and JCYJ20220530164601003).
Author contributions
B.Z., B.Q., and W.D. designed the research; B.Z. collected the data; B.Z. and S.Z. performed data analysis and wrote the manuscript; X.Z, Z.L, C.Z, H.L, J.L., and Y.L. calibrated the data; B.Z. designed and wrote the code of the website; D.L., W.D., and N.A. revised the manuscript.
Declaration of interests
The authors declare no competing interests.
Published: February 5, 2026
Footnotes
Supplemental information can be found online at https://doi.org/10.1016/j.isci.2026.114918.
Contributor Information
Bing Zeng, Email: bingzengvip@126.com.
Dongcheng Liu, Email: ldc2025@163.com.
Weiwei Dai, Email: daiweiwei@aierchina.com.
Bo Qin, Email: qinbo2024@ext.jnu.edu.cn.
Supplemental information
References
- 1.Lin B.C., Katneni U., Jankowska K.I., Meyer D., Kimchi-Sarfaty C. In silico methods for predicting functional synonymous variants. Genome Biol. 2023;24:126. doi: 10.1186/s13059-023-02966-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Sauna Z.E., Kimchi-Sarfaty C. Understanding the contribution of synonymous mutations to human disease. Nat. Rev. Genet. 2011;12:683–691. doi: 10.1038/nrg3051. [DOI] [PubMed] [Google Scholar]
- 3.Supek F., Miñana B., Valcárcel J., Gabaldón T., Lehner B. Synonymous mutations frequently act as driver mutations in human cancers. Cell. 2014;156:1324–1335. doi: 10.1016/j.cell.2014.01.051. [DOI] [PubMed] [Google Scholar]
- 4.Sharma Y., Miladi M., Dukare S., Boulay K., Caudron-Herger M., Groß M., Backofen R., Diederichs S. A pan-cancer analysis of synonymous mutations. Nat. Commun. 2019;10 doi: 10.1038/s41467-019-10489-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Shen X., Song S., Li C., Zhang J. Synonymous mutations in representative yeast genes are mostly strongly non-neutral. Nature. 2022;606:725–731. doi: 10.1038/s41586-022-04823-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Shen X., Song S., Li C., Zhang J. Further evidence for strong non-neutrality of yeast synonymous mutations. Mol. Biol. Evol. 2024;41 doi: 10.1093/molbev/msae224. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Lan Y., Xia Z., Shao Q., Lin P., Lu J., Xiao X., Zheng M., Chen D., Dou Y., Xie Q. Synonymous mutations promote tumorigenesis by disrupting m6A-dependent mRNA metabolism. Cell. 2025;188:1828–1841.e15. doi: 10.1016/j.cell.2025.01.026. [DOI] [PubMed] [Google Scholar]
- 8.Niu X., Tang W., Liu Y., Mo B., Yu Y., Liu Y., Wei W. Prime editor-based high-throughput screening reveals functional synonymous mutations in human cells. Nat. Biotechnol. 2025 doi: 10.1038/s41587-025-02710-z. [DOI] [PubMed] [Google Scholar]
- 9.Zeng Z., Bromberg Y. Predicting Functional Effects of Synonymous Variants: A Systematic Review and Perspectives. Front. Genet. 2019;10:914–915. doi: 10.3389/fgene.2019.00914. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Liu X., Jian X., Boerwinkle E. dbNSFP: A lightweight database of human nonsynonymous SNPs and their functional predictions. Hum. Mutat. 2011;32:894–899. doi: 10.1002/humu.21517. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Liu X., Jian X., Boerwinkle E. dbNSFP v2.0: A Database of Human Non-synonymous SNVs and Their Functional Predictions and Annotations. Hum. Mutat. 2013;34:E2393–E2402. doi: 10.1002/humu.22376. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Liu X., Wu C., Li C., Boerwinkle E. dbNSFP v3.0: A One-Stop Database of Functional Predictions and Annotations for Human Non-synonymous and Splice Site SNVs. Hum. Mutat. 2016;37:235–241. doi: 10.1002/humu.22932. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Liu X., Li C., Mou C., Dong Y., Tu Y. dbNSFP v4: a comprehensive database of transcript-specific functional predictions and annotations for human nonsynonymous and splice-site SNVs. Genome Med. 2020;12:103–108. doi: 10.1186/s13073-020-00803-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Liu X., White S., Peng B., Johnson A.D., Brody J.A., Li A.H., Huang Z., Carroll A., Wei P., Gibbs R., et al. WGSA: An annotation pipeline for human genome sequencing studies. J. Med. Genet. 2016;53:111–112. doi: 10.1136/jmedgenet-2015-103423. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Wen P., Xiao P., Xia J. dbDSM: A manually curated database for deleterious synonymous mutations. Bioinformatics. 2016;32:1914–1916. doi: 10.1093/bioinformatics/btw086. [DOI] [PubMed] [Google Scholar]
- 16.Cheng N., Li M., Zhao L., Zhang B., Yang Y., Zheng C.H., Xia J. Comparison and integration of computational methods for deleterious synonymous mutation prediction. Brief. Bioinform. 2020;21:970–981. doi: 10.1093/bib/bbz047. [DOI] [PubMed] [Google Scholar]
- 17.Livingstone M., Folkman L., Yang Y., Zhang P., Mort M., Cooper D.N., Liu Y., Stantic B., Zhou Y. Investigating DNA-RNA-and protein-based features as a means to discriminate pathogenic synonymous variants. Hum. Mutat. 2017;38:1336–1347. doi: 10.1002/humu.23283. [DOI] [PubMed] [Google Scholar]
- 18.Tang X., Zhang T., Cheng N., Wang H., Zheng C.H., Xia J., Zhang T. UsDSM: A novel method for deleterious synonymous mutation prediction using undersampling scheme. Brief. Bioinform. 2021;22 doi: 10.1093/bib/bbab123. [DOI] [PubMed] [Google Scholar]
- 19.Wang Y., Qiu C., Cui Q. A large-scale analysis of the relationship of synonymous SNPs changing microRNA regulation with functionality and disease. Int. J. Mol. Sci. 2015;16:23545–23555. doi: 10.3390/ijms161023545. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Tunyasuvunakool K., Adler J., Wu Z., Green T., Zielinski M., Žídek A., Bridgland A., Cowie A., Meyer C., Laydon A., et al. Highly accurate protein structure prediction for the human proteome. Nature. 2021;596:590–596. doi: 10.1038/s41586-021-03828-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Cheng J., Novati G., Pan J., Bycroft C., Žemgulytė A., Applebaum T., Pritzel A., Wong L.H., Zielinski M., Sargeant T., et al. Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science. 2023;381:eadg7492. doi: 10.1126/science.adg7492. [DOI] [PubMed] [Google Scholar]
- 22.Townshend R.J.L., Eismann S., Watkins A.M., Rangan R., Karelina M., Das R., Dror R.O. Geometric deep learning of RNA structure. Science. 2021;373:1047–1051. doi: 10.1126/science.abe5650. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Shen T., Hu Z., Sun S., Liu D., Wong F., Wang J., Chen J., Wang Y., Hong L., Xiao J., et al. Accurate RNA 3D structure prediction using a language model-based deep learning approach. Nat. Methods. 2024;21:2287–2298. doi: 10.1038/s41592-024-02487-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Wang W., Feng C., Han R., Wang Z., Ye L., Du Z., Wei H., Zhang F., Peng Z., Yang J. trRosettaRNA: automated prediction of RNA 3D structure with transformer network. Nat. Commun. 2023;14 doi: 10.1038/s41467-023-42528-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Kagaya Y., Zhang Z., Ibtehaz N., Wang X., Nakamura T., Punuru P.D., Kihara D. NuFold: end-to-end approach for RNA tertiary structure prediction with flexible nucleobase center representation. Nat. Commun. 2025;16:881. doi: 10.1038/s41467-025-56261-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Li Y., Zhang C., Feng C., Pearce R., Lydia Freddolino P., Zhang Y. Integrating end-to-end learning with deep geometrical potentials for ab initio RNA structure prediction. Nat. Commun. 2023;14 doi: 10.1038/s41467-023-41303-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Cao Y., Li L., Xu M., Feng Z., Sun X., Lu J., Xu Y., Du P., Wang T., Hu R., et al. The ChinaMAP analytics of deep whole genome sequences in 10,588 individuals. Cell Res. 2020;30:717–731. doi: 10.1038/s41422-020-0322-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Welter D., MacArthur J., Morales J., Burdett T., Hall P., Junkins H., Klemm A., Flicek P., Manolio T., Hindorff L., Parkinson H. The NHGRI GWAS Catalog, a curated resource of SNP-trait associations. Nucleic Acids Res. 2014;42:1001–1006. doi: 10.1093/nar/gkt1229. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Eicher J.D., Landowski C., Stackhouse B., Sloan A., Chen W., Jensen N., Lien J.P., Leslie R., Johnson A.D. GRASP v2.0: An update on the Genome-Wide Repository of Associations between SNPs and Phenotypes. Nucleic Acids Res. 2015;43:D799–D804. doi: 10.1093/nar/gku1202. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Landrum M.J., Lee J.M., Benson M., Brown G.R., Chao C., Chitipiralla S., Gu B., Hart J., Hoffman D., Jang W., et al. ClinVar: Improving access to variant interpretations and supporting evidence. Nucleic Acids Res. 2018;46:D1062–D1067. doi: 10.1093/nar/gkx1153. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Schaafsma G.C.P., Vihinen M. VariSNP, A benchmark database for variations from dbSNP. Hum. Mutat. 2015;36:161–166. doi: 10.1002/humu.22727. [DOI] [PubMed] [Google Scholar]
- 32.Li Q., Wang K. InterVar: Clinical Interpretation of Genetic Variants by the 2015 ACMG-AMP Guidelines. Am. J. Hum. Genet. 2017;100:267–280. doi: 10.1016/j.ajhg.2017.01.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Stenson P.D., Mort M., Ball E.V., Evans K., Hayden M., Heywood S., Hussain M., Phillips A.D., Cooper D.N. The Human Gene Mutation Database: towards a comprehensive repository of inherited mutation data for medical research, genetic diagnosis and next-generation sequencing studies. Hum. Genet. 2017;136:665–677. doi: 10.1007/s00439-017-1779-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Kaltak M., Corradi Z., Collin R.W.J., Swildens J., Cremers F.P.M. Stargardt disease-associated missense and synonymous ABCA4 variants result in aberrant splicing. Hum. Mol. Genet. 2023;32:3078–3089. doi: 10.1093/hmg/ddad129. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Xu W.Q., Wang R.M., Dong Y., Wu Z.Y. Pathogenicity of Intronic and Synonymous Variants of ATP7B in Wilson Disease. J. Mol. Diagn. 2023;25:57–67. doi: 10.1016/j.jmoldx.2022.10.002. [DOI] [PubMed] [Google Scholar]
- 36.Zhao G., Li K., Li B., Wang Z., Fang Z., Wang X., Zhang Y., Luo T., Zhou Q., Wang L., et al. Gene4Denovo: An integrated database and analytic platform for de novo mutations in humans. Nucleic Acids Res. 2020;48:D913–D926. doi: 10.1093/nar/gkz923. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Forbes S.A., Beare D., Boutselakis H., Bamford S., Bindal N., Tate J., Cole C.G., Ward S., Dawson E., Ponting L., et al. COSMIC: Somatic cancer genetics at high-resolution. Nucleic Acids Res. 2017;45:D777–D783. doi: 10.1093/nar/gkw1121. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Zhang J., Bajari R., Andric D., Gerthoffert F., Lepsa A., Nahal-Bose H., Stein L.D., Ferretti V. The International Cancer Genome Consortium Data Portal. Nat. Biotechnol. 2019;37:367–369. doi: 10.1038/s41587-019-0055-9. [DOI] [PubMed] [Google Scholar]
- 39.Sherry S.T., Ward M.H., Kholodov M., Baker J., Phan L., Smigielski E.M., Sirotkin K. DbSNP: The NCBI database of genetic variation. Nucleic Acids Res. 2001;29:308–311. doi: 10.1093/nar/29.1.308. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Lek M., Karczewski K.J., Minikel E.V., Samocha K.E., Banks E., Fennell T., O’Donnell-Luria A.H., Ware J.S., Hill A.J., Cummings B.B., et al. Analysis of protein-coding genetic variation in 60,706 humans. Nature. 2016;536:285–291. doi: 10.1038/nature19057. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Auton A., Abecasis G.R., Altshuler D.M., Durbin R.M., Bentley D.R., Chakravarti A., Clark A.G., Donnelly P., Eichler E.E., Flicek P., et al. A global reference for human genetic variation. Nature. 2015;526:68–74. doi: 10.1038/nature15393. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Karczewski K.J., Weisburd B., Thomas B., Solomonson M., Ruderfer D.M., Kavanagh D., Hamamsy T., Lek M., Samocha K.E., Cummings B.B., et al. The ExAC browser: Displaying reference data information from over 60 000 exomes. Nucleic Acids Res. 2017;45:D840–D845. doi: 10.1093/nar/gkw971. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Zhao H., Sun Z., Wang J., Huang H., Kocher J.P., Wang L. CrossMap: A versatile tool for coordinate conversion between genome assemblies. Bioinformatics. 2014;30:1006–1007. doi: 10.1093/bioinformatics/btt730. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Rentzsch P., Witten D., Cooper G.M., Shendure J., Kircher M. CADD: Predicting the deleteriousness of variants throughout the human genome. Nucleic Acids Res. 2019;47:D886–D894. doi: 10.1093/nar/gky1016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Quang D., Chen Y., Xie X. DANN: A deep learning approach for annotating the pathogenicity of genetic variants. Bioinformatics. 2015;31:761–763. doi: 10.1093/bioinformatics/btu703. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Shihab H.A., Gough J., Cooper D.N., Stenson P.D., Barker G.L.A., Edwards K.J., Day I.N.M., Gaunt T.R. Predicting the Functional, Molecular, and Phenotypic Consequences of Amino Acid Substitutions using Hidden Markov Models. Hum. Mutat. 2013;34:57–65. doi: 10.1002/humu.22225. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Rogers M.F., Shihab H.A., Mort M., Cooper D.N., Gaunt T.R., Campbell C. FATHMM-XF: Accurate prediction of pathogenic point mutations via extended features. Bioinformatics. 2018;34:511–513. doi: 10.1093/bioinformatics/btx536. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Capriotti E., Fariselli P. PhD-SNPg: A webserver and lightweight tool for scoring single nucleotide variants. Nucleic Acids Res. 2017;45:W247–W252. doi: 10.1093/nar/gkx369. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Wang H., Sun J., Liu M., Zheng C.H., Xia J., Cheng N. frDSM: An Ensemble Predictor With Effective Feature Representation for Deleterious Synonymous Mutation in Human Genome. IEEE/ACM Trans. Comput. Biol. Bioinform. 2023;20:371–377. doi: 10.1109/TCBB.2022.3167468. [DOI] [PubMed] [Google Scholar]
- 50.Cheng N., Wang H., Tang X., Zhang T., Gui J., Zheng C.H., Xia J. An Ensemble Framework for Improving the Prediction of Deleterious Synonymous Mutation. IEEE Trans. Circuits Syst. Video Technol. 2022;32:2603–2611. doi: 10.1109/TCSVT.2021.3063145. [DOI] [Google Scholar]
- 51.Buske O.J., Manickaraj A., Mital S., Ray P.N., Brudno M. Identification of deleterious synonymous variants in human genomes. Bioinformatics. 2013;29:1843–1850. doi: 10.1093/bioinformatics/btt308. [DOI] [PubMed] [Google Scholar]
- 52.Zhang T., Wu Y., Lan Z., Shi Q., Yang Y., Guo J. Syntool: A novel region-based intolerance score to single nucleotide substitution for synonymous mutations predictions based on 123,136 individuals. BioMed Res. Int. 2017;2017:1–5. doi: 10.1155/2017/5096208. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Wagner N., Çelik M.H., Hölzlwimmer F.R., Mertes C., Prokisch H., Yépez V.A., Gagneur J. Aberrant splicing prediction across human tissues. Nat. Genet. 2023;55:861–870. doi: 10.1038/s41588-023-01373-3. [DOI] [PubMed] [Google Scholar]
- 54.Zeng Z., Aptekmann A.A., Bromberg Y. Decoding the effects of synonymous variants. Nucleic Acids Res. 2021;49:12673–12691. doi: 10.1093/nar/gkab1159. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Mark A. myvariant: Accesses MyVariant.info variant query and annotation services. 2023. https://myvariant.info/
- 56.Wang K., Li M., Hakonarson H. ANNOVAR: Functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res. 2010;38 doi: 10.1093/nar/gkq603. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Robin X., Turck N., Hainard A., Tiberti N., Lisacek F., Sanchez J.C., Müller M. pROC : an open-source package for R and S+ to analyze and compare ROC curves. BMC Bioinf. 2011;12:77. doi: 10.1186/1471-2105-12-77. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Stacklies W., Redestig H., Scholz M., Walther D., Selbig J. pcaMethods - A bioconductor package providing PCA methods for incomplete data. Bioinformatics. 2007;23:1164–1167. doi: 10.1093/bioinformatics/btm069. [DOI] [PubMed] [Google Scholar]
- 59.Friedman J., Hastie T., Tibshirani R. Regularization Paths for Generalized Linear Models via Coordinate Descent. J. Stat. Softw. 2010;33:1–22. doi: 10.18637/jss.v033.i01. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Ionita-Laza I., Mccallum K., Xu B., Buxbaum J.D. A spectral approach integrating functional genomic annotations for coding and noncoding variants. Nat. Genet. 2016;48:214–220. doi: 10.1038/ng.3477.A. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Shihab H.A., Rogers M.F., Gough J., Mort M., Cooper D.N., Day I.N.M., Gaunt T.R., Campbell C. An integrative approach to predicting the functional effects of non-coding and coding sequence variation. Bioinformatics. 2015;31:1536–1543. doi: 10.1093/bioinformatics/btv009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Xiong H.Y., Alipanahi B., Lee L.J., Bretschneider H., Yuen R.K.C., Hua Y., Gueroussov S., Hamed S., Hughes T.R., Morris Q., et al. The human splicing code reveals new insights into the genetic determinants of disease. Science. 2015;347 doi: 10.1126/science.1254806. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Jaganathan K., Kyriazopoulou Panagiotopoulou S., McRae J.F., Darbandi S.F., Knowles D., Li Y.I., Kosmicki J.A., Arbelaez J., Cui W., Schwartz G.B., et al. Predicting Splicing from Primary Sequence with Deep Learning. Cell. 2019;176:535–548.e24. doi: 10.1016/j.cell.2018.12.015. [DOI] [PubMed] [Google Scholar]
- 64.Jian X., Boerwinkle E., Liu X. In silico prediction of splice-altering single nucleotide variants in the human genome. Nucleic Acids Res. 2014;42:13534–13544. doi: 10.1093/nar/gku1206. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Wang L., Zhang T., Yu L., Zheng C.H., Yin W., Xia J., Zhang T. Deleterious synonymous mutation identification based on selective ensemble strategy. Brief. Bioinform. 2023;24 doi: 10.1093/bib/bbac598. [DOI] [PubMed] [Google Scholar]
- 66.Schwarz J.M., Cooper D.N., Schuelke M., Seelow D. Mutationtaster2: Mutation prediction for the deep-sequencing age. Nat. Methods. 2014;11:361–362. doi: 10.1038/nmeth.2890. [DOI] [PubMed] [Google Scholar]
- 67.Zhang X., Li M., Lin H., Rao X., Feng W., Yang Y., Mort M., Cooper D.N., Wang Y., Wang Y., et al. regSNPs-splicing: a tool for prioritizing synonymous single-nucleotide substitution. Hum. Genet. 2017;136:1279–1289. doi: 10.1007/s00439-017-1783-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Bendl J., Musil M., Štourač J., Zendulka J., Damborský J., Brezovský J. PredictSNP2: A Unified Platform for Accurately Evaluating SNP Effects by Exploiting the Different Characteristics of Variants in Distinct Genomic Regions. PLoS Comput. Biol. 2016;12:1–18. doi: 10.1371/journal.pcbi.1004962. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Gelfman S., Wang Q., McSweeney K.M., Ren Z., La Carpia F., Halvorsen M., Schoch K., Ratzon F., Heinzen E.L., Boland M.J., et al. Annotating pathogenic non-coding variants in genic regions. Nat. Commun. 2017;8 doi: 10.1038/s41467-017-00141-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Jurka J. Repbase Update: A database and an electronic journal of repetitive elements. Trends Genet. 2000;16:418–420. doi: 10.1016/S0168-9525(00)02093-X. [DOI] [PubMed] [Google Scholar]
- 71.Lenhard B., Wasserman W.W. TFBS: Computational framework for transcription factor binding site analysis. Bioinformatics. 2002;18:1135–1136. doi: 10.1093/bioinformatics/18.8.1135. [DOI] [PubMed] [Google Scholar]
- 72.Richards S., Aziz N., Bale S., Bick D., Das S., Gastier-Foster J., Grody W.W., Hegde M., Lyon E., Spector E., et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genet. Med. 2015;17:405–424. doi: 10.1038/gim.2015.30. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Dong C., Wei P., Jian X., Gibbs R., Boerwinkle E., Wang K., Liu X. Comparison and integration of deleteriousness prediction methods for nonsynonymous SNVs in whole exome sequencing studies. Hum. Mol. Genet. 2015;24:2125–2137. doi: 10.1093/hmg/ddu733. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.Li C., Zhi D., Wang K., Liu X. MetaRNN: differentiating rare pathogenic and rare benign missense SNVs and InDels using deep learning. Genome Med. 2022;14 doi: 10.1186/s13073-022-01120-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Oba S., Sato M.A., Takemasa I., Monden M., Matsubara K.I., Ishii S. A Bayesian missing value estimation method for gene expression profile data. Bioinformatics. 2003;19:2088–2096. doi: 10.1093/bioinformatics/btg287. [DOI] [PubMed] [Google Scholar]
- 76.Shi F., Yao Y., Bin Y., Zheng C.H., Xia J. Computational identification of deleterious synonymous variants in human genomes using a feature-based approach. BMC Med. Genomics. 2019;12:12. doi: 10.1186/s12920-018-0455-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Cornelis S.S., Runhart E.H., Bauwens M., Corradi Z., De Baere E., Roosing S., Haer-Wigman L., Dhaenens C.M., Vulto-van Silfhout A.T., Cremers F.P.M. Personalized genetic counseling for Stargardt disease: Offspring risk estimates based on variant severity. Am. J. Hum. Genet. 2022;109:498–507. doi: 10.1016/j.ajhg.2022.01.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.ZENG B. CDsyn: A Comprehensive Database for Deleterious Human Synonymous SNVs Prediction at Zenodo. 2025. [DOI] [PMC free article] [PubMed]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The CDsyn can be queried online or downloaded from the URL (http://www.cdsyn.com), and the complete data can be obtained from Zenodo. A detailed list of datasets and other items is provided in the following link: CDsyn: https://zenodo.org/records/18320261 (https://doi.org/10.5281/zenodo.18320261).
Code and any additional information required in this paper are available from the lead contact upon request.







