Skip to main content
. 2026 Mar 30;23(4):772–784. doi: 10.1038/s41592-026-03050-9

Fig. 3. Iterative averaging co-distillation effectively compresses the ESM family into a single PLM.

Fig. 3

a, Single model versus ensemble performance on Balanced ClinVar and ProteinGym DMS benchmarks. Each point represents a pairwise ensemble obtained by averaging the predictions of two individual models (base ESM ensembles in gray, co-distilled ensembles in blue). Contour lines indicate performance density. Single model performances are annotated for reference. b, An overview of the subsequent rounds of averaging co-distillation. c, The performance across co-distillation rounds for the four participating models on Balanced ClinVar (left) and ProteinGym DMS (right). All models show monotonic improvement across rounds, with ESM2-3B converging to match the ensemble’s performance (dashed line) on both benchmarks after round 3, yielding VESM-3B. d, LLR distillation of VESM-3B into smaller-parameter models. Left: an overview of the knowledge distillation process (Methods). Right: percent improvement of the resulting VESM models over their corresponding base ESM models across both benchmarks.