Fig. 4. Performance evaluation of sequence-only VESM models on clinical VEP.
a, Benchmarking of VESM models against 24 other PLMs and VEP methods on an independent, publicly available ClinVar benchmark from ProteinGym (2.57 × 104 benign and 2.68 × 104 pathogenic variants across 2,227 genes). Models are color-coded based on the additional sources of information they use. Top: global AUC is shown for each model. Error bars correspond to s.d. of the ROC-AUC scores centered around the mean (estimated by bootstrapping; n = 50 label-balanced resamples of 6,000 variants). Bottom: the corresponding ROC curve is shown for top-performing models of each class. b, A boxplot of the ClinVar AUC per gene for n = 259 genes that have at least ten pathogenic or benign labeled variants. c, The percentage of ClinVar variants that can be confidently classified by each model. The prediction accuracy and the total number of annotated variants are computed by varying the classification threshold for each method. The average and a two-standard-deviation interval across n = 10 bootstraps are shown (Methods). d, A comparison with AlphaMissense7 on a recent (03/25) release of the ClinVar dataset. A boxplot of the AUC achieved by each model across a range of MAF filtering thresholds from gnomAD v4 (Methods). e, A performance evaluation on a subset of the ClinVar dataset that excludes all human variants that have been used for training AlphaMissense (MAF >1×10−5 in gnomAD v2). The error bars correspond to s.d. of the ROC-AUC scores centered around the mean (estimated by bootstrapping; n = 100 label-balanced resamples of 10,000 variants). f, Calibrated binary classification metrics evaluating the performance of each model on the same MAF-filtered ClinVar dataset used in e. The boxplots show: center line (median); box limits (Q1, Q3); whiskers (±1.5× interquartile range); points (outliers).
