Abstract
Alzheimer's disease (AD) is a progressive neurodegenerative disease characterized by progressive cognitive decline. Over 200 pathogenic mutations in amyloid-β precursor protein (APP), presenilin-1 (PSEN1), and presenilin-2 (PSEN2), have been implicated in AD. Yet, many rare and common variants have not been completely classified as protective or benign, risk-modifiers, or pathogenic, which is important for research on the disease mechanisms and discovery of treatment methods. The majority of these variants are missense mutations, and there is an active need for computational approaches to accurately predict their molecular consequences. AlphaMissense (AM) is a novel technology that uses population frequency data along with structural and sequential contexts from AlphaFold to predict the pathogenicity of missense mutations. Herein, we sought to evaluate the capabilities of AM on 114 variants of unknown significance (VUS), including 56 missense variants of PSEN1, 25 of APP, and 33 of PSEN2 by benchmarking its prediction against their respective Aβ isoform levels in vitro, respectively. We found that the AM scores correlated moderately well with the critical Aβ42/Aβ40 biomarker and Aβ40 levels in the transmembrane proteins compared to weaker correlations in traditional approaches, including Combined Annotation Dependent Depletion (CADD) v1.7, evolutionary model of variant effect (EVE), and Evolutionary Scale Modeling-1b (ESM-1B). Yet, there were non-significant correlations identified with Aβ42 levels in all models. Furthermore, we found that AM does not rely completely on structural contexts from AlphaFold2, as it accurately predicted the effects of known variants on residues with a low predicted local distance difference test (pLDDT) score. Additionally, based on the receiver operating characteristic-area under the curve analysis (ROC-AUC), we found that AM retained a high performance on 263 validated variants of these amyloidogenic genes, and performed the greatest compared to other models for the 114 VUS. We believe this is the first study to provide comprehensive characterization and validation of AM in comparison to the widely utilized pathogenicity scoring models for VUS involved in proteins implicated in AD.
Keywords: Alzheimer's disease, AlphaMissense, PSEN1, PSEN2, Pathogenicity prediction
Highlights
-
•
AM has a moderate correlation with Aβ42/Aβ40 levels in VUS of APP, PSEN1, and PSEN2.
-
•
This is the first study to validate AM against functional assays of VUS.
-
•
There is a weak correlation for Aβ42 and Aβ40 levels among the computational tools.
-
•
AM can accurately predict variants on intrinsically disordered regions of proteins.
1. Introduction
Alzheimer's disease (AD) is the most common form of dementia, for which the hallmarks include neuronal loss, accumulation of Aβ plaque and neurofibrillary tangles [1]. However, early-onset AD (EAOD) constitutes only a small portion of patients, estimated at 41.2 per 100,000 individuals [2]. In this group of patients, mutations in amyloid-β precursor protein (APP), presenilin-1 (PSEN1), and presenilin-2 (PSEN2) cause autosomal dominant EAOD, where it is estimated to include 5.3 per 100,000 at risk individuals [2]. As of March 2025, there have been 92 variants of PSEN1, 91 of PSEN2, and 117 of APP, culminating in 300 total variants in known literature from the Alzforum database server (https://www.alzforum.org/). While many of these variants have been classified as benign, or protective, risk-inducing, or pathogenic polymorphisms, a number of variants remain uncertain ([3]; [19,20]). The high ambiguity in many of these variants is that pathogenicity is evaluated through segregation analysis, where a specific family has several generations of clinically-confirmed AD compared to those without the variant remaining healthy, older individuals, and many patients have limited or no related family history. In a recent study by Hsu et al. [4], a pathogenicity algorithm was developed using genetics, traditional bioinformatic predictions, and in vitro analyses for improving evaluation of 90 variants of unknown significance (VUS). This clinical classification algorithm was further improved upon and used on a larger sample of 53 VUS by Marsh et al. [5]. However, even with these major developments, these VUS cannot be evaluated in clinical trials for AD unless the gold standard segregation data is available.
Therefore, there is an urgent need for high-quality methods to systematically characterize the severity of APP, PSEN1, and PSEN2 mutations. Recently, a novel deep-learning model, AlphaMissense (AM), by Google DeepMind trained on population frequency data and structural and sequential contexts from an AlphaFold-based system (Chang et al., 2023). On the ClinVar dataset of variants, AM obtained a 90 % precision and outperformed both PolyPhen-2 [21] and SIFT [22], which are the two bioinformatic tools employed by Hsu et al. [4] and Marsh et al. [5]. Yet, even with an impressive predictive accuracy, there have been limited validation studies attempting to characterize the strengths and limitations of this innovation [24]. In a recent study by McDonald et al. [6], the pathogenicity of variants on the cystic fibrosis transmembrane conductance regulator gene (CFTR) were compared against protein function in vitro (channel function, trafficking and folding competency, pharmacological response), where predictions were moderately correlated with function. In another study, AM was used to evaluate VUS on 59 genes associated with Mendelian disorders [7]. AM has also been used in combination with Phenolyzer for predicting the effects of missense variants in a patient [8]. Other than McDonald et al. [6], to our knowledge, there has been no studies that compare the abilities of AM with its structural and evolutionary training to learn the pathogenicity of variants on protein function with supporting in vitro data. This study was inspired by McDonald et al. [6] to evaluate the abilities of AM compared to other computational approaches in characterizing the missense variants of these transmembrane proteins whose impacts remain largely unknown. Herein, we carried out AM predictions and other state-of-the-art deep learning scoring against in vitro Aβ levels of 56 missense variants from PSEN1, 25 from APP, and 33 from PSEN2, respectively. Secondly, we benchmarked the capabilities of these models against 263 variants that have received a complete classification from prior literature.
2. Materials and methods
2.1. In vitro analyses
The source data of Aβ levels for the PSEN1 and PSEN2 variants was directly derived from Hsu et al. [4] and Marsh et al. [5], where the exact experimental protocol was used for developing the cell-based assays. Both studies used mouse neuroblastoma cells (N2A) with endogenous Psen1 and Psen2 knocked out by CRISPR/Cas9 (N2APS1/PS2 KO) for evaluating PSEN1 and PSEN2 variants, while N2A in APP variants [4,5]. Data has been reported for 38 variants from Marsh et al. [5] and 22 variants of PSEN1 from Hsu et al. [4], but our analyses did not include D40del, QR127G, G378fs, and dE9, culminating in 56 variants. Likewise, for PSEN2, 14 variants from Marsh et al. [5] and 23 from Hsu et al. [4], but E126fs, K306fs, and H162 N were not included and N141I was repeated, for a total of 33 variants. H162 N is not the naturally occurring residue found in the isoforms in UniProt, and is missing in the Alzforum database as of April 23, 2025. For APP, 3 variants were reported from Marsh et al. [5] and 25 from Hsu et al. [4], but K670 N/M671L and G708G were excluded and L723P was repeated, totaling in 25 variants. Furthermore, for our analyses involving validated variants of PSEN1, PSEN2, and APP, we obtained these classifications from the Alzforum Database (n = 263), which has been displayed in the Supplementary Material.
2.2. 3D structures and pathogenicity predictions
AM scores were obtained from the AlphaFold Structure Database, where predictions were made on the open-source AF2 model (Accession: AF-P49768-F1-v4; AF-P49810-F1-v4). The AM source data given from the heatmap was then extracted into an external file for follow-up analyses. Scores above 0.56 were considered pathogenic and those below 0.34 were considered benign, respectively. For Combined Annotation Dependent Depletion (CADD) scoring, we used the Protvar (https://www.ebi.ac.uk/ProtVar/) database to obtain the predictions, and similarly extracted the data to an external file (Rentzsch et al. [9]; Stephenson et al. [10]. CADD scores below 15 were considered benign and those above were classified as pathogenic. Scoring from the evolutionary model of variant effect (EVE) (https://evemodel.org/) was set at values greater than 0.65 were pathogenic and those below 0.35 were benign, respectively [11]. Likewise, for Evolutionary Scale Modeling-1B (ESM-1B) the scoring of the negative log-likelihood ratio (-LLR) is the following: LLR <7 is benign and -LLR >8 is pathogenic, respectively (https://huggingface.co/spaces/ntranoslab/esm_variants) (Brandes et al., 2023). The source data for all pathogenicity prediction algorithms was recorded in an external file and have been provided in the Supplementary Material.
2.3. Statistical analyses
All statistical analyses and visualizations were performed in GraphPad (version 10.4.1) (GraphPad Software, La Jolla, CA, USA). For linear correlational analyses, we calculated Pearson's and Spearman's correlation coefficients, and determined whether the predicted slope was significantly non-zero. We have reported our analyses of all data values provided without consideration of outliers using the ROUT method at Q = 1 %, and later measured their effects. All data has been presented as the mean ± standard error of the mean (SEM). Unless stated otherwise, a p-value less than 0.05 was statistically significant. The p-values for evaluating significance of a non-zero slope have been shown in all the figures, and those for Spearman's ρ and the correlation coefficient (r) have been provided in the GraphPad file in the Supplementary Material.
2.4. ROC-AUC analyses
In our receiver operating characteristic-area under the curve analysis (ROC-AUC), we followed a similar procedure reported from McDonald et al. [6]. The pathogenicity predictions were conducted in pairwise comparison through a binary classification model. For instance, when evaluating a confirmed pathogenic variant, a correct prediction assigns the value ‘1’, and if benign or ambiguous would be assigned ‘0’. Due to the sample of variants and the inherent nature of VUS, we had limited our analyses to variants that have been characterized as benign or pathogenic, and removed ambiguous or uncertain ground truths from the Alzforum. Additionally, predictions of ‘probably’ or ‘likely’ were assigned one of the two binary categories. Not all computational tools follow the same scaling of probabilities involved in classifications, especially CADD, which limited our analyses to only binary classification. Another important consideration was ‘risk factor’ predictions reported from Marsh et al. [5] and Hsu et al. [4] that we had considered as benign when performing analysis because it has a lower priority over variants that are ‘probably’ pathogenic in nature. Risk factors have biological variability, where some may have greater influence than others, but for the sake of this analysis, these variants were considered benign. Using an in-house Python script with the scikit-learn library by Pedregosa et al. [25], we formulated a binary classification of the variants following the described scoring methods for each computational model in Section 2.2. Our in-house Python script has been deposited in the GitHub repository described in the Supplementary Material. We treated one class of variables (i.e. pathogenic, benign) as positive and the remaining as negative, thereby being able to avoid consideration of ambiguous predictions from AM and other models, forming two models. For the VUS, the classifications defined by Hsu et al. [4] and Marsh et al. [5] were assumed as the ground truths.
For complete transparency of our datasets, we must note that there are duplicates in both the validated variants following ACMG-AMP guidelines from Alzforum (n = 263), and those from Marsh et al. [5] and Hsu et al. [4] (n = 114). We performed our analyses ‘as is’ without artificially altering these datasets in order to capture the complete abilities of these tools in real-world contexts, where preliminary clinical classifications often can differ from guideline criterias. Our hope of this evaluation was to determine how well these computational approaches are to clinical evaluations compared to the validated guidelines. The source data used for these analyses have been reported in Supplementary Tables 1 and 2
3. Results
3.1. AlphaMissense is reliable in predicting variants on intrinsically disordered domains
Before evaluating the in vitro data against AM scoring predictions, we first determined the accuracy of AlphaFold2 (AF2) in predicting the structures of APP, PSEN1, and PSEN2 [12]. We had originally assumed that the AF2 structure had influence in AM predictions as the predictions for the human proteome are based on it in the AlphaFold Structure Database [23]. From APP, we initially considered this transmembrane protein to be ineligible for analysis because the structure had a low global predicted distance difference test (pLDDT) of 67.45 compared to PSEN1 of 72.45 and PSEN2 of 71.45, respectively. However, we performed a case study involving 263 mutations with a known pathogenicity according to the American College of Medical Genetics and Genomics and the Association for Molecular Pathology (ACMG-AMP) guidelines from the Alzforum Database. This included 34 variants from APP, 200 from PSEN1, and 29 from PSEN2. All data has been displayed in Table 1 and Supplementary Table 2, where the overall accuracy of AM was 91.00 %, 70.59 %, and 89.66 % for pathogenic PSEN1, APP, and PSEN2 variants, respectively. Likewise, among benign variants, the accuracy of AM was 75.00 %, 72.00 %, and 81.82 % for benign PSEN1, APP, and PSEN2 variants, respectively. Within the sample of analyses, a significant number of variants were found on residues with low pLDDT scores below 70, including 17 from APP, 31 from PSEN1, and 11 from PSEN2, respectively. From variants on disordered regions, 64.71 % of variants were correctly predicted on APP, 93.55 % for PSEN1, and 100 % for PSEN2, culminating in a 72.86 % overall accuracy. We further describe these results in Section 3.3. Surprisingly, even with poor structural contexts provided from AF2, AM is still able to make accurate predictions of these disordered residues. This is a key validation of AM similar to those found in disordered regions of the CFTR gene [6]. We believe that the AF2 structural models for the 3 proteins evaluated in this study are the most accurate homology template as shown in the SWISS-MODEL because it is derived from resolved experimental structures. Therefore, we began our correlational analyses of AM against the Aβ in vitro data for APP, PSEN1, and PSEN2.
Table 1.
Binary classification of prediction models against confirmed variants. Pathogenic and benign models for each computational tool have been displayed for PSEN1 (n = 200), APP (n = 34), and PSEN2 (n = 29). Each variant included has received complete evaluation under the ACMG-AMP guidelines. All benchmarks were calculated using the scikit-learn library by Pedregosa et al. (2011) with an in-house Python script. ∗Given a prediction was not available for R35Q from EVE, this variant was removed from evaluations. ∗∗The T18 M variant was also not available from EVE, and therefore was removed from calculations.
| PSEN1(n = 200) | Accuracy | Precision | Recall | Specificity | f1-Score | AUC |
| AM-Pathogenic | 0.9100 | 1.0000 | 0.9062 | 1.0000 | 0.9508 | 0.9531 |
| CADD-Pathogenic | 0.9650 | 0.9648 | 1.0000 | 0.1250 | 0.9821 | 0.5625 |
| ESM-Pathogenic | 0.9550 | 0.9946 | 0.9583 | 0.8750 | 0.9761 | 0.9167 |
| EVE-Pathogenic∗ | 0.7789 | 1.0000 | 0.7708 | 1.0000 | 0.8706 | 0.8854 |
| AM-Benign | 0.9500 | 0.4444 | 1.0000 | 0.9479 | 0.6154 | 0.9740 |
| CADD-Benign | 0.9650 | 1.0000 | 0.1250 | 1.0000 | 0.2222 | 0.5625 |
| ESM-Benign | 0.9750 | 0.6364 | 0.8750 | 0.9792 | 0.7368 | 0.9271 |
| EVE-Benign∗ | 0.9146 | 0.2727 | 0.8571 | 0.9167 | 0.4138 | 0.8869 |
| APP (n = 34) | ||||||
| AM-Pathogenic | 0.7059 | 0.9167 | 0.5500 | 0.9286 | 0.6875 | 0.7393 |
| CADD-Pathogenic | 0.5882 | 0.5882 | 1.0000 | 0.0000 | 0.7407 | 0.5000 |
| ESM-Pathogenic | 0.8235 | 0.8889 | 0.8000 | 0.8571 | 0.8421 | 0.8286 |
| EVE-Pathogenic | 0.5294 | 0.7500 | 0.3000 | 0.8571 | 0.4286 | 0.5786 |
| AM-Benign | 0.7353 | 0.6316 | 0.8571 | 0.6500 | 0.7273 | 0.7536 |
| CADD-Benign | 0.5882 | 0.0000 | 0.0000 | 1.0000 | 0.0000 | 0.5000 |
| ESM-Benign | 0.8824 | 0.9167 | 0.7857 | 0.9500 | 0.8462 | 0.8679 |
| EVE-Benign | 0.6765 | 0.5882 | 0.7143 | 0.6500 | 0.6452 | 0.6821 |
| PSEN2 (n = 29) | ||||||
| AM-Pathogenic | 0.8966 | 0.7692 | 1.0000 | 0.8421 | 0.8696 | 0.9211 |
| CADD-Pathogenic | 0.4138 | 0.3704 | 1.0000 | 0.1053 | 0.5405 | 0.5526 |
| ESM-Pathogenic | 0.7241 | 0.5625 | 1.0000 | 0.6313 | 0.6923 | 0.7658 |
| ∗∗EVE-Pathogenic | 0.7500 | 0.6667 | 0.6000 | 0.8333 | 0.6316 | 0.7167 |
| AM-Benign | 0.8276 | 1.0000 | 0.7368 | 1.0000 | 0.8485 | 0.8684 |
| CADD-Benign | 0.4138 | 1.0000 | 0.1053 | 1.0000 | 0.1905 | 0.5526 |
| ESM-Benign | 0.6897 | 1.0000 | 0.5263 | 1.0000 | 0.6897 | 0.7632 |
| ∗∗EVE-Benign | 0.7500 | 0.8667 | 0.7222 | 0.8000 | 0.7879 | 0.7611 |
3.2. AlphaMissense retains the strongest correlation with Aβ levels for PSEN1 variants
After curating all pathogenicity predictions from the models, we performed correlational analyses of these scores against the Aβ levels for 56 variants of PSEN1. This protein is a part of the γ-secretase complex that regulates cleavage of APP, PSEN1 mutations promote the production of longer forms of Aβ peptides. Particularly, Aβ42 is produced and introduces greater risk of plaque formation and its accumulation consistent with AD pathology. In diagnostics, the elevated Aβ42/Aβ40 ratio is hypothesized to be a biomarker for AD, though contraversial and presented with inconsistencies in the general cutoff [29]. In Fig. 1A–B, we have displayed the trends of AM predictions on the three-dimensional structure (3D) and the heatmap. First, we evaluated the Aβ42/Aβ40 ratios among the computational approaches (Fig. 1C). With AM, we found that the model held the strongest positive correlation with Aβ42/Aβ40 (ρ = 0.6097, r = 0.3864) as the predicted pathogenicity increased. This performance was very close but slightly lower with EVE (ρ = 0.5561, r = 0.3834), and even weaker with ESM-1b (ρ = 0.5218, r = 0.3814). Interestingly, we found that CADD scoring was the weakest correlated (ρ = 0.3592, r = 0.1542) with the ratio compared to all other metrics. The predicted slope for AM was significant (p = 0.0033), ESM was significant (p = 0.0037), EVE was significant (p = 0.0039), but CADD was not (p = 0.2550) from 0. We also investigated outliers that were identified in the regression, and how they influenced the overall trend. Using the ROUT method at Q = 1 %, we identified 7 outliers, including L174R, L241P, T245P, Y288H, D385E, S390 N, and P436A in all 4 computational models. With removal of these outliers, we performed linear regression to identify changes in the observed trends. The correlation of AM (ρ = 0.4853, r = 0.49), ESM (ρ = 0.4273, r = 0.3722), and EVE (ρ = 0.4502, r = 0.3792) weakened, but those of CADD (ρ = 0.3548, r = 0.2891) strengthened slightly and was significantly non-zero (p = 0.0439), but the overall trend remained the same as prior. In Marsh et al. [5], it was found that all outliers except for Y288H had significantly increased Aβ42/Aβ40 driven by reduction of both Aβ42 and Aβ40 levels, and rare in gnomAD population frequency database, so they were directly classified as pathogenic variants. Y288H was also classified as pathogenic, but the increased ratio can be directly attributed to decreased Aβ40 levels. All models correctly predicted that these 7 outliers were pathogenic, where the average was 0.994 for AM, 14.68 for ESM, 0.963 for EVE, and 27.4 for CADD, respectively. Overall, the correlational analysis for the Aβ42/Aβ40 ratio has been displayed.
Fig. 1.
Pathogenicity prediction scoring against in vitro Aβ levels of PSEN1 variants. (a) The ribbon structure of PSEN1 colored by the scaling of predicted pathogenicity of AM is visualized with a (b) heatmap of the predicted AM score for all residues. (c) Aβ42/Aβ40 levels of all PSEN1 variants against the scoring of each predictive model is shown, including correlational statistics. The red dotted line was used as an indicator for variants with unusually high Aβ42/Aβ40 levels. The same visualizations are shown for (d) Aβ42 levels and (e) Aβ40 along with all statistics. Correlation coefficients (r) and Spearman Coefficient (ρ) have been shown in each correlational analyses. The p-values shown for each correlation are whether the predicted slope was significantly non-zero. Dotted lines at y = 1 represent WT value for in vitro Aβ levels.
Next, we evaluated the correlation of Aβ42 levels among the PSEN1 variants against the pathogenicity prediction models, and the results have been visualized in Fig. 1D. As described previously with the Aβ42/Aβ40 ratios, the levels of Aβ42 and Aβ40 can either both or individually drive the change from the wild-type (WT) levels, which could be challenging for any predictive model to accurately gauge the general trend for individual Aβ peptides. For AM, there was a weak positive correlation (ρ = 0.2004, r = 0.0681) compared to a negative correlation of ESM (ρ = −0.1254, r = −0.09), EVE (ρ = −0.2162, r = −0.1276), and CADD (ρ = −0.0879, r = −0.0864). The slope for AM was not significant from zero (p = 0.6179), CADD was not significant (p = 0.5263), ESM was not significant (p = 0.5074), and EVE was significant (p = 0.3530). Generally, all models were unable to accurately predict the in vitro Aβ42 levels, although there is a non-significant negative correlation observed.
Lastly, we calculated the correlation of Aβ40 levels among the variants against the model scores shown in Fig. 1E. We found that there was a modest negative correlation with AM (ρ = −0.5768, r = −0.4898), CADD (ρ = −0.3458, r = −0.3444), ESM (ρ = −0.4695, r = −0.4732), and EVE (ρ = −0.5502, r = −0.5151). The predicted slopes of all models were significant from zero (p < 0.0001). Overall, considering the trends of Aβ42 not having significant changes from the WT levels, the marked decrease in Aβ40 with increasing predicted pathogenicity, and therefore accounts for the modest positive correlation observed in the ratio. In both key categories of Aβ42/Aβ40 ratio and Aβ40, AM had a stronger correlation with in vitro Aβ levels compared to other deep learning-based approaches and the traditional CADD model.
3.3. AlphaMissense outperforms other predictions for PSEN2, though not significantly for APP
Next, we evaluated the 25 APP VUS variants against the scoring for each pathogenicity prediction model, where the results of the analysis are displayed in Fig. 2. The trends in AM scores are visualized on the 3D structure and heatmap in Fig. 2A–B. In the key Aβ42/Aβ40 ratio, AM and CADD were found to have the highest correlation among the models, but were found to not be significant in either models (Spearman's r; AM, p = 0.1392. CADD, p = 0.1336) shown in Fig. 2C. With Aβ42 levels, AM was also not significantly correlated (ρ = 0.2908, p = 0.1585) (Fig. 2D). Interestingly, we found that ESM had the highest correlation with Aβ42 levels (ρ = 0.3654) among the models despite not being significant (p = 0.0725). This sharp increase is attributed to the two outliers of T719 N and L723P that were influential points that shifted the predicted slope, and after its removal caused the correlation to become weaker than AM (ρ = 0.1848). Lastly, with Aβ40 levels, there was a general moderately negative correlation found against higher pathogenicity predictions (Fig. 2E). However, none of the models were found to have a correlation significantly different from no correlation similar to the results obtained for correlations against Aβ42 levels. We believe that the non-significant results received for APP VUS is likely caused by the low sample size, and is comprehensively described in Section 4.
Fig. 2.
Pathogenicity prediction scoring against in vitro Aβ levels of APP variants. (a) The ribbon structure of APP colored by the scaling of predicted pathogenicity of AM is visualized with a (b) heatmap of the predicted AM score for all residues. (c) Aβ42/Aβ40 levels of all APP variants against the scoring of each predictive model is shown, including correlational statistics. The red dotted line was used as an indicator for variants with unusually high Aβ42/Aβ40 levels. The same visualizations are shown for (d) Aβ42 levels and (e) Aβ40 along with all statistics. Correlation coefficients (r) and Spearman Coefficient (ρ) have been shown in each correlational analyses. The p-values shown for each correlation are whether the predicted slope was significantly non-zero. Dotted lines at y = 1 represent WT value for in vitro Aβ levels.
Furthermore, we performed similar analyses with PSEN2 VUS, where the results are shown in Fig. 3, and trends of AM scores in the 3D structure and sequence shown in Panel A-B. Intriguingly, we found that the correlation of Aβ42/Aβ40 ratio biomarker was significant, moderately positively correlated (ρ = 0.4262, p = 0.0134) while the remaining models had relations that were not significant (Fig. 3C). Though AM correctly predicted that the N141Y and N141I were pathogenic and had a ratio greater than 10, we performed analysis without these two points and found that the relation weakened and was not significant (ρ = 0.3106, p = 0.0890). Still, AM performed better than EVE that correctly identified the two outliers as pathogenic, but were still weaker and not significant in the correlation including these two points (ρ = 0.3511, p = 0.0528). Next, we evaluated the Aβ42 levels against the predictive models, and found that there was a significant, moderately positive correlation found in AM (ρ = 0.4598, r = 0.5235), ESM (ρ = 0.3797, r = 0.4342), and EVE (ρ = 0.4444, r = 0.4678), but not CADD (ρ = 0.2971, r = 0.2856) (Fig. 3D). Finally, we evaluated the Aβ40 levels, and similarly found that the correlation between the predictive models and in vitro data was not significant but was negatively correlated (Fig. 3E).
Fig. 3.
Pathogenicity prediction scoring against in vitro Aβ levels of PSEN2 variants. (a) The ribbon structure of PSEN2 colored by the scaling of predicted pathogenicity of AM is visualized with a (b) heatmap of the predicted AM score for all residues. (c) Aβ42/Aβ40 levels of all PSEN2 variants against the scoring of each predictive model is shown, including correlational statistics. The red dotted line was used as an indicator for variants with unusually high Aβ42/Aβ40 levels. The same visualizations are shown for (d) Aβ42 levels and (e) Aβ40 along with all statistics. Correlation coefficients (r) and Spearman Coefficient (ρ) have been shown in each correlational analyses. The p-values shown for each correlation are whether the predicted slope was significantly non-zero. Dotted lines at y = 1 represent WT value for in vitro Aβ levels.
3.4. AlphaMissense outperform other approaches in VUS and validated benchmarks
Continuing from our obtained datasets in Section 3.1, we performed an ROC-AUC analyses on validated variants of the three proteins from the Alzforum and similarly the estimated pathogenicity of VUS from Marsh et al. [5] and Hsu et al. [4]. The results of the models from the Alzforum dataset are reported in Table 1. For validated variants of PSEN1 and PSEN2, AM held the strongest performance in distinguishing benign and pathogenic mutations with the highest AUC. However, with APP, AM was conservative in its predictions as it has a low false positive rate, and held the second highest performance to ESM-1b. Generally, EVE predictions were moderately high but still lower than both AM and ESM. Another key finding was that CADD struggles to identify benign variants and is biased towards detecting pathogenic variants, attributing to its lower metrics among the other three models. Next, we evaluated predictions against VUS that were given preliminary classifications by prior studies, where the results have been displayed in Table 2. The overall trend is that all models had a lower performance (∼0.9–∼0.6–0.7 AUC) compared to the validated variants. This finding appears similar to the AUC of these approaches against de novo variants reported by Cheng et al. [13] (shown in Fig. 2C of the publication) around 0.6 to 0.8 AUC. VUS are naturally more challenging to predict because of ambiguous clinical information, lesser evolutionary information, and their novelty, which is logical for the lower performance. However, there is a possibility of differences between clinical classifications and the guideline criterias that can describe this discrepancy. AM demonstrated superior performance against other approaches, holding the highest AUC in both benign and pathogenic variants of all three proteins, and ESM-1b and EVE finished closely behind. CADD had consistently ranked the lowest and even had an performance that was akin to random chance for VUS classified as benign variants of APP. Overall, AM is useful for making functional predictions of variants compared to other deep learning tools and traditional approaches like CADD, which will ultimately be beneficial for characterizing novel variants.
Table 2.
Binary classification of prediction models among VUS. Pathogenic and benign models for each computational tool have been displayed for PSEN1 (n = 56), APP (n = 25), and PSEN2 (n = 33). Each variant has been derived from those studied by Hsu et al. (2020) and Marsh et al. (2025). All benchmarks were calculated using the scikit-learn library by Pedregosa et al. (2011) with an in-house Python script. ∗Given a prediction was not available for R35Q from EVE, this variant was removed from evaluations.
| PSEN1 (n = 56) | Accuracy | Precision | Recall | Specificity | f1-Score | AUC |
| AM-Pathogenic | 0.7500 | 0.8333 | 0.7895 | 0.6667 | 0.8108 | 0.7281 |
| CADD-Pathogenic | 0.6786 | 0.6786 | 1.0000 | 0.0000 | 0.8085 | 0.5000 |
| ESM-Pathogenic | 0.7679 | 0.6429 | 0.5000 | 0.8684 | 0.5625 | 0.6842 |
| EVE-Pathogenic∗ | 0.6545 | 0.8065 | 0.6579 | 0.6471 | 0.7246 | 0.6525 |
| AM-Benign | 0.7321 | 0.5882 | 0.5556 | 0.8158 | 0.5714 | 0.6857 |
| CADD-Benign | 0.6786 | 0.0000 | 0.0000 | 1.0000 | 0.0000 | 0.5000 |
| ESM-Benign | 0.7500 | 0.6429 | 0.5000 | 0.8684 | 0.5625 | 0.6842 |
| EVE-Benign∗ | 0.7091 | 0.5294 | 0.5294 | 0.7895 | 0.5294 | 0.6594 |
| APP (n = 25) | ||||||
| AM-Pathogenic | 0.7200 | 0.5000 | 0.7143 | 0.7222 | 0.5882 | 0.7183 |
| CADD-Pathogenic | 0.2800 | 0.2800 | 1.0000 | 0.0000 | 0.4375 | 0.5000 |
| ESM-Pathogenic | 0.6400 | 0.4167 | 0.7143 | 0.6111 | 0.5263 | 0.6627 |
| EVE-Pathogenic | 0.7200 | 0.5000 | 0.4286 | 0.8333 | 0.4615 | 0.6310 |
| AM-Benign | 0.6800 | 0.8571 | 0.6667 | 0.7143 | 0.7500 | 0.6905 |
| CADD-Benign | 0.2800 | 0.0000 | 0.0000 | 1.0000 | 0.0000 | 0.5000 |
| ESM-Benign | 0.6000 | 0.8333 | 0.5556 | 0.7143 | 0.6667 | 0.6349 |
| EVE-Benign | 0.5600 | 0.7692 | 0.5556 | 0.5714 | 0.6452 | 0.5635 |
| PSEN2 (n = 33) | ||||||
| AM-Pathogenic | 0.8182 | 0.6667 | 0.8000 | 0.8261 | 0.7273 | 0.8130 |
| CADD-Pathogenic | 0.3636 | 0.3226 | 1.0000 | 0.0870 | 0.4878 | 0.5435 |
| ESM-Pathogenic | 0.6364 | 0.4444 | 0.8000 | 0.5652 | 0.5714 | 0.6826 |
| EVE-Pathogenic | 0.7273 | 0.5385 | 0.7000 | 0.7391 | 0.6087 | 0.7196 |
| AM-Benign | 0.7879 | 0.9444 | 0.7391 | 0.9000 | 0.8293 | 0.8196 |
| CADD-Benign | 0.3636 | 1.0000 | 0.0870 | 1.0000 | 0.1600 | 0.5435 |
| ESM-Benign | 0.6061 | 0.9167 | 0.4783 | 0.9000 | 0.6286 | 0.6891 |
| EVE-Benign | 0.6061 | 0.8571 | 0.5217 | 0.8000 | 0.6486 | 0.6609 |
4. Discussion
In this study, we provide a critical characterization and validation of AM in comparison to other deep-learning approaches in predicting the pathogenicity of 114 VUS in APP, PSEN1, and PSEN2 using in vitro functional assay data from prior experimental studies. Using correlational analyses, we found that AM was able to decipher the key Aβ42/Aβ40 biomarker in pathogenic variants among PSEN1 and PSEN2, and had a stronger correlation compared to other traditional and computational approaches. However, all models had generally failed to correlate with individual Aβ42 and Aβ40 levels of each VUS. Secondly, we found that AM does not rely exclusively on structural contexts to make accurate predictions as it was able to characterize ∼70 % validated variants found on low pLDDT residues. Lastly, from an ROC-AUC analysis, we found that AM had a very high AUC above 0.9 close to other deep-learning approaches on a dataset of 263 validated variants. Despite having the highest AUC (∼0.7) among the other predictive models on the 114 VUS dataset, AM still had a lower predictive power compared to its performance against the ACMG-AMP guidelines.
A major finding of this study that remains inconclusive is how individual Aβ42 and Aβ40 levels were not strongly correlated with predicted pathogenicity. All models had resulted in a weak negative correlation, though not statistically significant for Aβ40 levels. However, this pattern was not uniform in Aβ42 levels that were positively correlated in APP while negatively in PSEN1 and PSEN2. The molecular mechanisms by which PSEN leads to AD is broadly debated, but the two theories include (1) the amyloid hypothesis that variants initiate disease pathogenesis by increasing Aβ42 levels [28], and (2) the presenilin hypothesis that the mutations leads to loss-of-function and eventually dementia [27]. However, Sun et al. [14] demonstrated that the mutations suppress Aβ production and catalytic γ-secretase activity, where 104 variants led to a marked decrease in Aβ40 and Aβ42 production. This study also found that the key Aβ42/Aβ40 ratio was not uniformly increased (13 of the 96 pathogenic variants evaluated had a decrease). It appears that marked decreases in Aβ40 levels significantly more than Aβ42 levels appeared to elevate this ratio [26]. The findings from this study support this conclusion as Aβ40 levels were negatively correlated with predicted pathogenicity than Aβ42 that did not have a clear trend. Though the individual decrease in Aβ40 levels was not significant, the consistent Aβ42 levels may have led to a stronger and significant positive correlation to appear with the Aβ42/Aβ40 ratio. By including a greater sample of variants in the future, it may be possible to detect significance with individual Aβ40 levels, but the current trends observed in this work still support these previous findings [14].
Another key finding from this work is that AM scoring does not necessarily rely only on AlphaFold-derived predictions, and is capable of making correct predictions without supporting structural evidence. In our prior work using AF2 to determine the structural, stability, and functional effects of proposed AD-causing and Nasu-Hakola disease (NHD)-causing coding variants in TREM2, we had initially believed that residues with a low-confidence pLDDT and located on the surface of the protein (RSA >25 %) were generally predicted to be benign [15]. AD-causing variants of TREM2 are located on the surface of the receptor and impact ligand-binding affinity, but were predicted by AM to be benign variants [16]. Likewise, NHD-causing variants are buried in the structure of TREM2 and lead to complete loss-of-function [16], and therefore predicted to be pathogenic by AM [15]. When conducting the presented study, we initially considered AM would make a prediction similar to those as TREM2, where residues on the surface would be expected to be benign, and therefore a failed prediction. However, we found that AM had a strong accuracy in identifying both pathogenic and benign variants. We hypothesize that AM is using sequential contexts from evolutionary conservation and population frequencies to correctly discern the impacts of these variants without the direct need of structural information. In McDonald et al. [6], a similar pattern was also observed in the CFTR gene, where AM made a correct prediction despite having low-confidence structural data. Our study provides stronger evidence for this observation through the 51 accurate predictions of 70 known variants.
In the future, we hope to perform larger-scale analyses of AM against many genes to comprehensively evaluate its accuracy to predicting functional effects of missense variants. Due to the inherent nature of VUS, there was a low sample size of variants and their respective in vitro data to be correlated in this study. Initially, we had considered using the in vitro Aβ data provided from Petit et al. [17], but due to differing methods compared to Marsh et al. [5] and Hsu et al. [4], we did not combine this biological data. It must also be denoted that even with a clinical classification algorithm developed using orthogonal genetic, bioinformatic, and in vitro data, the pathogenicity of these variants remains uncertain [4]. Multiple pathogenicity prediction models are recommended to obtain a full characterization of the missense mutations given their differing training, and further studies are needed especially for those based on deep-learning approaches in the future. We recommend the development of ensemble approaches using the orthogonal deep-learning tools (AM, EVE, Primate-AI, etc), as these models may achieve a stronger predictive accuracy to the functional effects of variants, similar to those previously completed with REVEL [18]. To our knowledge, this is the first study to compare the scores of pathogenicity prediction tools against biophysical data of VUS and classified variants in crucial amyloidogenic genes. We have also further benchmarked the capabilities of AM in making accurate functional predictions despite poor structural contexts. Lastly, our correlational analyses may be relevant for the discussion of the presenilin hypothesis and the implication of pathogenic mutations in EAOD.
CRediT authorship contribution statement
Joshua Pillai: Writing – review & editing, Writing – original draft, Visualization, Validation, Software, Resources, Methodology, Investigation, Formal analysis, Conceptualization. Sophia Liu: Methodology, Investigation, Formal analysis, Data curation. Kijung Sung: Validation, Supervision, Project administration, Conceptualization. Linda Shi: Writing – review & editing, Writing – original draft, Supervision, Resources, Project administration, Conceptualization. Chengbiao Wu: Writing – review & editing, Writing – original draft, Investigation, Conceptualization.
Funding information
This material was based upon work supported by a gift from Beckman Laser Institute Inc. to LS. Special thanks to the private donors to our University of California, San Diego (UCSD) Institute for Engineering in Medicine, Biophotonics Technology Center: Dr. Shu Chien from UCSD Bioengineering, Dr. Lizhu Chen from CorDx Inc., Dr. Xinhua Zheng, David & Leslie Lee for their generous donations.
Declaration of competing interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Footnotes
Supplementary data to this article can be found online at https://doi.org/10.1016/j.bbrep.2025.102049.
Contributor Information
Linda Shi, Email: zshi@ucsd.edu.
Chengbiao Wu, Email: chw049@ucsd.edu.
Appendix A. Supplementary data
The following is the Supplementary data to this article:
Data availability
Source data has been provided in the Supplementary Data.
References
- 1.Bellelli F., Angioni D., Arosio B., Vellas B., De Souto Barreto P. Hallmarks of aging and Alzheimer's Disease pathogenesis: paving the route for new therapeutic targets. Ageing Res. Rev. 2025;106 doi: 10.1016/j.arr.2025.102699. [DOI] [PubMed] [Google Scholar]
- 2.Campion D., et al. Early-onset autosomal dominant alzheimer disease: prevalence, genetic heterogeneity, and mutation spectrum. Am. J. Hum. Genet. 1999;65:664–670. doi: 10.1086/302553. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Benitez B.A., et al. The PSEN1, P.E318G variant increases the risk of Alzheimer's disease in APOE-Ε4 carriers. PLoS Genet. 2013;9 doi: 10.1371/journal.pgen.1003685. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Hsu S., et al. Systematic validation of variants of unknown significance in APP, PSEN1 and PSEN2. Neurobiol. Dis. 2020;139 doi: 10.1016/j.nbd.2020.104817. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Marsh J.A., et al. Evaluating pathogenicity of variants of unknown significance in APP, PSEN1, and PSEN2. Neurotherapeutics. 2025 doi: 10.1016/j.neurot.2025.e00527. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.McDonald E.F., Oliver K.E., Schlebach J.P., Meiler J., Plate L. Benchmarking AlphaMissense pathogenicity predictions against cystic fibrosis variants. PLoS One. 2024;19 doi: 10.1371/journal.pone.0297560. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Kurtovic-Kozaric A., et al. Comprehensive evaluation of AlphaMissense predictions by evidence quantification for variants of uncertain significance. Front. Genet. 2024;15 doi: 10.3389/fgene.2024.1487608. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Wang T., et al. Predicting valuable missense variants with AlphaMissense in a multiple pulmonary infection patient. Clin. Case Rep. 2024;12 doi: 10.1002/ccr3.8453. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Rentzsch P., Witten D., Cooper G.M., Shendure J., Kircher M. CADD: predicting the deleteriousness of variants throughout the human genome. Nucleic Acids Res. 2018;47:D886–D894. doi: 10.1093/nar/gky1016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Stephenson J.D., et al. ProtVar: mapping and contextualizing human missense variation. Nucleic Acids Res. 2024;52:W140–W147. doi: 10.1093/nar/gkae413. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Frazer J., et al. Disease variant prediction with deep generative models of evolutionary data. Nature. 2021;599:91–95. doi: 10.1038/s41586-021-04043-8. [DOI] [PubMed] [Google Scholar]
- 12.Jumper J., et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583–589. doi: 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Cheng J., et al. Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science. 2023;381 doi: 10.1126/science.adg7492. [DOI] [PubMed] [Google Scholar]
- 14.Sun L., Zhou R., Yang G., Shi Y. Vol. 114. Proceedings of the National Academy of Sciences; 2016. (Analysis of 138 Pathogenic Mutations in Presenilin-1 on the in Vitro Production of Aβ42 and Aβ40 Peptides by γ-secretase). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Pillai J., Sung K., Wu C. Predicting the impact of missense mutations on an unresolved protein's stability, structure, and function: a case study of Alzheimer's disease‐associated TREM2 R47H variant. Comput. Struct. Biotechnol. J. 2025;27:564–574. doi: 10.1016/j.csbj.2025.01.024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Kober D.L., et al. Neurodegenerative disease mutations in TREM2 reveal a functional surface and distinct loss-of-function mechanisms. Elife. 2016;5 doi: 10.7554/eLife.20391. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Petit D., et al. Aβ profiles generated by Alzheimer's disease causing PSEN1 variants determine the pathogenicity of the mutation and predict age at disease onset. Mol. Psychiatr. 2022;27:2821–2832. doi: 10.1038/s41380-022-01518-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Ioannidis N.M., et al. REVEL: an ensemble method for predicting the pathogenicity of rare missense variants. Am. J. Hum. Genet. 2016;99:877–885. doi: 10.1016/j.ajhg.2016.08.016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Sassi C., et al. Investigating the role of rare coding variability in Mendelian dementia genes (APP , PSEN1 , PSEN2 , GRN , MAPT , and PRNP) in late-onset Alzheimer's disease. Neurobiol. Aging. 2014;35:2881.e1–2881.e6. doi: 10.1016/j.neurobiolaging.2014.06.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Guerreiro R.J., et al. Genetic screening of Alzheimer's disease genes in Iberian and African samples yields novel mutations in presenilins and APP. Neurobiol. Aging. 2008;31:725–731. doi: 10.1016/j.neurobiolaging.2008.06.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Adzhubei I., Jordan D.M., Sunyaev S.R. Predicting functional effect of human missense mutations using PolyPhen‐2. Curr. Protocols Human Genetics. 2013;76 doi: 10.1002/0471142905.hg0720s76. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Ng P.C. SIFT: predicting amino acid changes that affect protein function. Nucleic Acids Res. 2003;31:3812–3814. doi: 10.1093/nar/gkg509. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Varadi M., et al. AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic Acids Res. 2021;50:D439–D444. doi: 10.1093/nar/gkab1061. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Curtis D. Assessment of ability of AlphaMissense to identify variants affecting susceptibility to common disease. Eur. J. Hum. Genet. 2024;32:1419–1427. doi: 10.1038/s41431-024-01675-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Pedregosa, et al. SciKit-learn: machine learning in Python. J. Mach. Learn. Res. 2011:2825–2830. [Google Scholar]
- 26.Kelleher R.J., Shen J. Presenilin-1 mutations and Alzheimer's disease. Proc. Natl. Acad. Sci. 2017;114:629–631. doi: 10.1073/pnas.1619574114. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Shen J., Kelleher R.J. The presenilin hypothesis of Alzheimer's disease: evidence for a loss-of-function pathogenic mechanism. Proc. Natl. Acad. Sci. 2006;104:403–409. doi: 10.1073/pnas.0608332104. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Hardy J., Selkoe D.J. The Amyloid hypothesis of Alzheimer's disease: progress and problems on the road to therapeutics. Science. 2002;297:353–356. doi: 10.1126/science.1072994. [DOI] [PubMed] [Google Scholar]
- 29.Selkoe D.J., Hardy J. The amyloid hypothesis of Alzheimer's disease at 25 years. EMBO Mol. Med. 2016;8:595–608. doi: 10.15252/emmm.201606210. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Source data has been provided in the Supplementary Data.



