Fig. 2. Maximum confidence co-distillation of the ESM family substantially improves individual ESM models.
a, An overview of the co-distillation framework; for each protein, all models make predictions, which are then combined into a single 20xL matrix by taking the element-wise minimum (maximum confidence per variant). Each individual model is then updated by calculating a loss with respect to its own output (Methods). b, A performance evaluation of individual co-distilled models compared to their baseline on ClinVar (2,400 genes with a total of ~27K pathogenic and benign variants, class balanced per gene (left)) and ProteinGym (120 organismal fitness and activity DMS assays (right)). The dashed red lines indicate the performance of ESMIN. c, An ablation study varying the percentage of proteins used as training data during co-distillation. The full dataset (100%, including the validation split) contains 18,683 human proteins. Ablations were made by first excluding all proteins that are similar to the ones used in the benchmark (5,279 proteins, >30% sequence similarity) and then sequentially eliminating proteins from the remaining set (nested subsets). d, The performance of learned model embeddings in nine downstream tasks (supervised fine-tuning on experimental measurements). Task-independent (frozen) embeddings were extracted from the base and co-distilled versions of ESM2-650M and used to train task-specific models (Methods). The table reports the mean and s.d. across three random seeds; statistically significant differences (two-sided paired t-test; P < 0.05) are highlighted in bold.
