Skip to main content
. 2026 Mar 30;23(4):772–784. doi: 10.1038/s41592-026-03050-9

Fig. 6. VESM models accurately quantify missense variant severity across continuous clinical phenotypes in UK Biobank.

Fig. 6

a, A correlation between variant-level VESM-3B predictions and single-variant association effect sizes (Genebass β coefficients) across 153 gene–phenotype pairs. The circle size represents missense SKAT-O significance (gene-level association P value from Genebass); color indicates gene-level pLoF effect size (burden test; Genebass). Prominent outliers are highlighted. b, Performance comparison of VESM models against AlphaMissense, ESM base models and an allele-frequency baseline across 103 strongly associated gene–phenotype pairs (Methods). Performance is summarized by average association strength (−log10 regression P value (top)) and stratified by phenotype category (bottom). Regression-derived P values were computed using Pearson’s product-moment correlation (two-sided test; scipy.stats.pearsonr); no adjustment for multiple comparisons was applied. c, A Pearson correlation between VESM-3B variant-level predictions and single-variant association effect sizes for gene–phenotype pairs with regression derived P value < 0.1 (two-sided test; as in b). Phenotypes are grouped by phenotype category (lipid metabolism, liver function and renal function), and the color denotes pLoF burden test effect size.