Skip to main content
. 2026 Mar 30;23(4):772–784. doi: 10.1038/s41592-026-03050-9

Extended Data Fig. 3. Max confidence co-distillation.

Extended Data Fig. 3

(a) Recovery of conserved domain signals after co-distillation. For KRAB (n = 519 proteins) and BRICHOS (n = 48), the violin plots (median, centre line; interquartile range, box; whiskers, 1.5× IQR) show the average LLR (mutational sensitivity) across domain positions for base models versus their co-distilled counterparts (round 1 denoted as _r1). Final round co-distilled models (denoted as _r3) are also included for reference. All co-distilled models are able to recover signals missed by their own bases (see Fig. 1b). (b) LLR heatmaps for ZFP57 (KRAB) and ITM2B (BRICHOS); while base ESM1b and ESM2-650M models miss one domain each (as highlighted in Fig. 1a), their co-distilled versions can identify both. (c) Model composition ablation for max-confidence co-distillation (round-1) using only the top-3, the top-8, the default 11-model set, or all 12 (incl. 15B). Performance on Balanced ClinVar and DMS fitness/activity shows that performance saturates once a moderately diverse medium-to-large subset is included. (d) Minimum (max-confidence) versus average aggregation under the top-3 and 11-model settings. Minimum aggregation yields consistent performance gains, especially with more diverse model sets.