Skip to main content
Nature Communications logoLink to Nature Communications
. 2026 Jun 22;17:7841. doi: 10.1038/s41467-026-74517-8

Allostery is a widespread cause of loss-of-function variant pathogenicity

Xiaotian Liao 1,2, Ben Lehner 1,3,4,5,✉
PMCID: PMC13439481  PMID: 42331831

Abstract

Allosteric communication between non-contacting sites in proteins plays a fundamental role in biological regulation and drug action. While allosteric gain-of-function variants are known drivers of oncogene activation, the broader importance of allostery in genetic disease and protein evolution is less clear. Here, we introduce a comparative framework that disentangles functional disruption by mutations from protein destabilization. Applying this framework across diverse datasets—ranging from paired experimental measurements of abundance and activity to proteome-wide comparisons of evolutionary fitness and biophysical stability predictions—we provide evidence that allostery is a widespread cause of loss-of-function variant pathogenicity in human genetic diseases. In addition, our analyses reveal a conserved distance-dependent decay of allosteric mutational effects outside of protein active sites. As an important mechanism of pathogenicity, allostery needs to be better mapped, understood, and predicted across the human proteome.

Subject terms: Genetics, Systems biology, Proteomics, Structural biology


Mutations at sites distant from a protein’s active site can disrupt function through allosteric mechanisms. Here, authors present evidence that allostery is a widespread and underappreciated cause of loss-of-function variant pathogenicity in human diseases.

Introduction

Allostery, the regulation of a protein’s activity through changes at sites distant from an active site, underpins protein regulation1–3. Allostery allows proteins to function as molecular switches and microprocessors, with their enzymatic activities and molecular interactions quantitatively controlled by small-molecule binding, post-translational modifications, and macromolecular interactions. Allostery is central to the control of metabolic pathways, signal transduction, growth control and neural information processing in both normal physiology and disease1,3–8. Moreover, many effective therapeutics bind allosteric sites and have allosteric mechanisms of action, including drugs targeting ion channels, receptors, kinases, and transcription factors2,9–13.

The importance of allostery for human genetic disease is, however, less well established5,14,15. While the role of allostery in gain-of-function (GOF) diseases is well-documented, exemplified by driver mutations in KRAS16,17, JAK218,19, and EGFR20,21, its contribution to loss-of-function (LOF) human genetic disease remains underexplored15,22,23. While it has been estimated that ~40-60% of LOF variants cause disease by destabilizing the protein fold (reducing abundance)24–28, a substantial fraction retain stability but fail to function27,29–37. These variants likely disrupt specific functional sites - not only orthosteric active sites but also distal allosteric positions, which form the dynamic networks and energetic pathways essential for signal propagation through the protein structure. We term these variants ‘functional mutations’ to distinguish them from variants affecting protein stability.

The potential scale of such distal functional disruption is underscored by recent comprehensive mutagenesis studies, which have generated the first systematic experimental maps of allosteric communication throughout proteins35–39. These near-complete atlases for a handful of proteins have revealed that many mutations—even those distant from active sites and binding sites—can exert measurable allosteric effects on protein function. For example, 114 amino acid substitutions in the oncoprotein KRAS allosterically inhibit binding to the oncoprotein effector RAF137. Even in small protein domains such as SH3 domains, many variants have small but measurable allosteric effects35 that, in combination, can strongly impair binding to interaction partners40. These first allosteric maps suggest that allostery is widespread and highlight the need to reconsider the importance of allostery for variant pathogenicity and molecular evolution.

Predicting these energetic networks computationally has traditionally relied on identifying co-evolution41–44. Because allosteric communication requires energetic coupling between distant residues, a mutation at one site creates evolutionary pressure for compensatory mutations at coupled sites. Conventional statistical methods, such as Statistical Coupling Analysis (SCA), and Direct Coupling Analysis (DCA), exploited this signal to map “sectors” and pairs of co-evolving residues that correspond to functional allosteric networks45–48.

Recent advances in deep learning have scaled this logic to the entire proteome49–51. Protein Language Models (pLMs) like ESM-1v49 are trained on millions of sequences and learn to predict the probability of an amino acid given its context. Crucially, these models implicitly learn “total evolutionary fitness”, capturing not just folding constraints, but also potentially the long-range statistical dependencies (allostery) and binding interactions important for survival. However, pLMs do not distinguish the mechanism by which a variant is deleterious. To disentangle these mechanisms, we require models such as ThermoMPNN52, which are trained on biophysical data and predict the folding free energies of mutations.

Here, we integrate these complementary computational models to identify pathogenic variants that act through functional disruption rather than stability loss. We first validate this framework using clinically important proteins (PTEN, GCK) where paired experimental measures of abundance and activity allow us to distinguish variants that impair molecular function from those that reduce protein abundance. We then extend this analysis to essential DNA repair proteins (BRCA1, RAD51C, BAP1) profiled by saturation genome editing, demonstrating that predicted stability helps isolate functional mutations. Next, we turn to comprehensive experimental allosteric maps of binding domains (PSD-95-PDZ3, GRB2-SH3) to confirm that the residual variance between predicted ESM-1v fitness and measured folding stability captures changes in binding free energy. Finally, we scale this analysis to six diverse human enzymes to demonstrate the universality of these observations. Across multiple analyses and structurally diverse proteins, these LOF allosteric effects are enriched close to protein active sites and decay over distance. Our results suggest that allostery is a widespread mechanism of variant pathogenicity that needs to be better understood and predicted across the human proteome.

Results

Pathogenic variants often act through mechanisms beyond stability

From a structural perspective, missense mutations that impair protein function can act through three primary mechanisms: (1) globally destabilizing the protein fold, (2) directly disrupting the active or binding site, or (3) indirectly perturbing function through distal or allosteric effects. To estimate the relative contribution of stability to LOF pathogenicity, we combined proteome-wide stability predictions, large-scale experimental abundance measurements, and clinical variant annotations.

Evaluation of ClinVar variants

We first used ThermoMPNN (TMPNN), a stability-prediction model built on the ProteinMPNN53 inverse folding network and trained on large-scale experimental data27, to predict changes in folding free energy (ΔΔGf) for all possible missense mutations across 19,288 human proteins. This analysis encompassed 187,634,329 variants, including 102,056 annotated in the ClinVar database54,55. As expected, pathogenic variants were significantly more likely to be predicted as destabilizing than benign variants (48.4% vs 8.8%, odds ratio (OR) = 9.752, p < 2.2e-16, Fisher’s exact test (FET)). Predicted stability changes are thus a very useful, albeit imperfect, predictor of pathogenicity (Receiver Operating Characteristic – Area Under the Curve (ROC-AUC) = 0.698, Precision–Recall Area Under the Curve (PR-AUC) = 0.564, n = 32,147 pathogenic, n = 69,909 benign variants). Nevertheless, a large proportion of pathogenic variants are predicted to be stable. Even when restricting the analysis to the protein core–where mutations are most likely to disrupt folding56 (OR = 10.315, p < 2.2e-16, FET)–we found that while 67.2% of pathogenic variants are destabilizing, 30.5% of pathogenic variants are predicted to preserve wild-type level stability (Fig. 1A).

Fig. 1. Pathogenic variants often act through mechanisms beyond stability.

Fig. 1

A Classification of ClinVar missense variants into three categories: WT-like (light gray), destabilizing (dark gray), and stabilizing (blue), based on their impact on protein stability/abundance. Percentages and total counts are shown for all ClinVar variants included in the proteome analysis (left), protein core variants (middle-left), variants from ‘Domainome 1.0’ (middle-right), and ‘VAMP’-seq (right), stratified by benign and pathogenic clinical labels. TMPNN = ThermoMPNN. B Schematic of the “activity-abundance residual” framework used to isolate functional effects. A LOESS regression (black curve) defines the baseline relationship between protein abundance (or predicted stability) and activity (or predicted fitness). Variants tracking the diagonal are “abundance-driven,” where functional loss is commensurate with destabilization. In contrast, “activity-driven” variants (red) deviate significantly from this trend, retaining high abundance/stability but losing function. The vertical distance from the trendline (residual) quantifies this abundance-independent functional effect. The table (right) lists the specific combinations of experimental (flask icon) and computational (network icon) metrics used to calculate these residuals across the different datasets in this study (Figs. 2–5). LOESS Locally Estimated Scatterplot Smoothing. AbundancePCA = Abundance Protein Fragment Complementation Assay.

We then validated these computational findings using large-scale experimental datasets. We analyzed the ‘Domainome’ and ‘VAMP-seq’ (Variant Abundance by Massively Parallel sequencing) datasets (Supplemental Table 2), which measure cellular protein abundance for mutations throughout hundreds of domains and several full-length proteins. These experimental data closely mirrored the predictions: 52% of pathogenic variants in protein domains and 55.8% of pathogenic variants in full-length proteins were measured to disrupt abundance (Fig. 1A). Consequently, while stability is predictive of pathogenicity, approximately half of all clinically pathogenic missense variants retain near wild-type-level stability. This indicates that a substantial fraction of human genetic disease is driven by mechanisms beyond fold destabilization.

Stability explains only a modest fraction of evolutionary fitness

To quantify the extent to which stability accounts for total variant fitness, we compared stability values against ESM-1v, a protein language model that captures “total fitness” based on evolutionary sequence variation49.

Across the entire proteome, ThermoMPNN-predicted stability explained only a modest proportion of the variance in ESM-1v scores (Fig. S1C, median R² = 0.224, interquartile range (IQR) 0.125–0.298). Importantly, this result was not specific to ESM-1v; we observed similar results when comparing stability against AlphaMissense (Fig. S1C), confirming that state-of-the-art variant effect predictors capture pathogenic mechanisms distinct from pure thermodynamic stability loss. To ensure this “missing variance” was not simply due to surface-exposed sites, we repeated the analysis for buried residues only (relative solvent accessible surface area (rSASA) < 0.2, excluding disordered and predicted binding-site residues). Even in the protein core, stability still only partially explained predicted pathogenicity (Fig. S1C, median R² = 0.158, IQR 0.096–0.232). Notably, this relationship varied by protein class: stability explained less variance in regulatory proteins such as voltage-gated ion channels (~10%) compared to ~23–26% in ribosomal proteins, immunoglobulins, and T-cell receptors (Fig. S1F).

This observation was further confirmed by experimental data. We compared ESM-1v scores against experimentally measured abundance changes from the ‘Mega-scale’ (protease sensitivity) and ‘Domainome’ (cellular abundance) datasets. Consistently, experimentally quantified abundance changes explained only a modest fraction of the variance in evolutionary fitness (Fig. S1C, median R² = 0.252 and 0.218, respectively). Similarly, for proteins with high-quality VAMP-seq data, abundance changes explained a median of only 26% of the variance in ESM-1v scores. These results demonstrate that protein stability is insufficient to explain evolutionary fitness: a large fraction of variants is deleterious while preserving wild-type stability/abundance.

Isolating functional effects via activity-abundance residuals

To systematically distinguish functional defects from abundance-driven effects, we formulated a framework based on “activity-abundance residuals” (Fig. 1B). We modeled the global relationship between protein abundance and activity using LOESS (locally estimated scatterplot smoothing) regression. The trendline defines the baseline expectation: the typical activity retained by a variant given its abundance.

Visualizing variants relative to this LOESS fit reveals two distinct LOF populations. “Abundance-driven variants” track the LOESS diagonal, where loss of activity is commensurate with destabilization. In contrast, “activity-driven variants” deviate significantly from this trend, clustering in the lower-right quadrant where abundance remains high but activity is lost. We reasoned that the residual—the vertical distance between a variant’s observed activity and the LOESS fit—effectively isolates abundance-independent functional effects. Large negative residuals identify variants where loss of function is disproportionate to the loss of abundance, thereby highlighting disrupted active sites, binding interfaces, or long-range allosteric effects.

Clinical pathogenic variants in PTEN and GCK show allosteric decay

To validate that “activity-abundance residuals” specifically isolate functional and allosteric sites, we analyzed two clinically important proteins where high-quality measurements of both abundance and activity are available: the lipid phosphatase PTEN (Fig. 2A) and the glucose-sensing kinase glucokinase (GCK) (Fig. 2C). PTEN germline variants are associated with a spectrum of disorders including Cowden syndrome and other cancer predisposition syndromes57. GCK, on the other hand, acts as the body’s glucose sensor; LOF variants cause Maturity-Onset Diabetes of the Young (GCK-MODY), while GOF variants cause hyper insulinemic hypoglycemia58–60.

Fig. 2. Measured activity vs measured abundance in PTEN and GCK.

Fig. 2

A Crystal structure of PTEN (PDB: 1D5R) colored by the median activity-abundance residuals. L(+)-tartaric acid (TLA), located near the catalytic site, is shown in orange. Regions with large negative residuals (red) indicate sites with high abundance but low activity. B Scatter plots of measured protein abundance (‘VAMP’-seq) versus phosphatase activity (yeast complementation) for 4,839 PTEN variants. A LOESS regression curve (black line) models the baseline abundance-activity relationship. Residuals from this fit quantify deviations from the trend. Left: All variants. Center: ClinVar Benign variants (circles). Right: ClinVar Pathogenic variants (triangles). Dashed gray lines indicate wild-type abundance (0.71) and activity (−1.11) values. Two canonical pathogenic mutations at the active site, C124R and R130G, are highlighted. C Crystal structure of GCK (PDB: 1V4S) colored by the median activity-abundance residuals. Glucose (GLC-orange) and the allosteric activator (MRK-dark cyan) are highlighted. D Distributions of activity–abundance residuals for ClinVar benign, pathogenic, and abundance-matched control variants in PTEN (left) and GCK (right). For all distributions, the violins show the kernel density of all values. Within each violin, an embedded boxplot indicates the median (center horizontal line; values labelled) and the interquartile range (IQR), with whiskers extending to 1.5× the IQR; an overlaid vertical line spans the 2.5th–97.5th percentiles (central 95% of values). To generate the abundance-matched null distribution, variants were first binned by their measured protein abundance (0.5-unit window). For each pathogenic variant, a residual score was randomly sampled from the pool of non-pathogenic variants residing in the same abundance bin (1000 bootstrap iterations). Deviation of pathogenic variants from this null expectation was assessed via a one-sided empirical P-value. Comparisons between Benign and Pathogenic variants were performed using a one-sided Wilcoxon rank-sum test. E Distance-dependent decay of functional effects. The absolute magnitude of the activity-abundance residual was plotted against the minimal side-chain heavy atom distance to the substrate. An exponential decay model (black curve representing the least-squares fit) was fitted to variants with negative activity-abundance residuals, with 95% bootstrap confidence intervals shown as gray shading. Three-star ClinVar pathogenic variants (variants reviewed by an expert panel) that reside outside of orthosteric sites are marked as black triangles and labeled.

We applied our residual framework to datasets combining ‘VAMP’-seq (abundance) with yeast-based functional assays for 4,839 PTEN variants61,62 and 8,255 GCK variants33,60. By fitting a LOESS curve to the abundance-activity landscape, we defined the expected activity for a given stability (Fig. 2B, Fig. S2A). In both proteins, we observed the expected “abundance-driven” population tracking the diagonal. However, a distinct population of variants deviated significantly, retaining high abundance but exhibiting severely compromised function (large negative residuals).

As expected, many pathogenic ClinVar variants reduced protein abundance. However, many others reduced enzymatic activity more than expected given their abundance levels (Fig. 2D). Mapping these residuals onto the protein structures revealed their enrichment in both active sites and distal allosteric regions. In PTEN, high-residual variants localized to the catalytic core and surface-exposed residues, including the WPD, P, and TI loops that form the active site pocket (Fig. 2A). In GCK, pathogenic variants clustered not only in the glucose-binding pocket but also in helix 13, a known allosteric regulator of activity63,64 (Fig. 2C). In contrast, clinically benign variants generally maintained wild-type-like abundance and activity and showed minimal residuals (Fig. 2D).

We next asked whether these functional mutational effects follow a physical pattern consistent with allosteric propagation—the transmission of energetic changes from one site in a protein to another, which provides a mechanistic basis for how mutations distant from active sites can still impact function65,66. If mutations act solely by locally disrupting active sites, functional effects should be binary (present at the active site, absent elsewhere). However, in the presence of allostery, energetic perturbations should propagate through the scaffold, decaying with distance from the functional site. This distance-dependent allosteric decay is visible in the small number of experimental allosteric maps generated to date35–39.

To test this, we plotted the absolute abundance-corrected activity residuals of each variant against its distance from the bound ligand (Fig. 2E). In both PTEN and GCK, we observed a clear decay: variants closest to the orthosteric sites had the strongest effects, and the typical effect size dropped by half over distances of 14.9 Å (PTEN) and 21.2 Å (GCK). This distance-dependent pattern was statistically significant in both proteins. Linear regression of the log-transformed absolute residuals against distance yielded negative slopes (PTEN: slope = –0.034, P = 8.78e-08; GCK: slope = –0.045, P < 2e-16).

Both PTEN and GCK contain clinically pathogenic variants that clustered in the active site and interface regions—as expected—but many were also located at distal positions (Fig. 2E). To test whether such non-orthosteric pathogenic variants exert functional effects beyond abundance loss, we compared their activity–abundance residuals to those of abundance-matched, non-pathogenic controls. In both orthosteric and non-orthosteric regions, pathogenic variants showed significantly larger residuals than abundance-matched controls (Fig. S2B), indicating greater loss of activity than expected from abundance effects alone. Many pathogenic variants in PTEN and GCK, therefore, likely act through allosteric mechanisms.

Clinically pathogenic variants in saturation genome editing data

While the analysis of PTEN and GCK shows that many pathogenic variants could act allosterically, paired experimental measurements of abundance and activity are not available for most human proteins. We therefore asked whether this framework could be generalized by substituting measured abundance with predicted thermodynamic stability.

To test this, we examined three essential DNA repair proteins–BRCA1, RAD51C, and BAP1–for which high-quality functional data are available from Saturation Genome Editing (SGE) screens67–69. SGE provides a direct readout of cellular fitness but does not explicitly measure protein abundance. We hence used ThermoMPNN as a computational proxy for structural stability and compared it against the experimental SGE functional scores (Fig. 3, Fig. S3, 4).

Fig. 3. Measured function score vs predicted thermoMPNN stability in SGE proteins.

Fig. 3

A Crystal structure of the BRCA1-BRCT domain (PDB: 1T29). Residues are colored by the median residuals. BACH1 peptide is highlighted in yellow. B Scatter plots of experimental function (SGE score) versus predicted ΔΔGf for 1,837 BRCA1 variants. A LOESS regression (black line) defines the baseline relationship. Left: All variants. Center: ClinVar Benign variants (circles). Right: ClinVar Pathogenic variants (triangles). Example pathogenic variants C39Y and M1775R are highlighted. C Crystal structure of the BRCA1-RING domain complexed with BARD1-RING domain (yellow) (PDB: 1JM7). D Violin plots comparing SGE-ThermoMPNN ΔΔGf residuals between benign and pathogenic variants for BRCA1, RAD51C, and BAP1. For all distributions, the violins show the kernel density of all values. Within each violin, an embedded boxplot indicates the median (center horizontal line; values labelled) and the interquartile range (IQR), with whiskers extending to 1.5× the IQR; an overlaid vertical line spans the 2.5th–97.5th percentiles (central 95% of values). Comparisons between Benign and Pathogenic variants were performed using a one-sided Wilcoxon rank-sum test. E Distance-dependent decay of functional residuals. The absolute magnitude of the residual was plotted against the minimal side-chain heavy atom distance to the ligands (BACH1 for BRCA1-BRCT; BARD1-RING for BRCA1-RING). An exponential decay model (black curve representing the least-squares fit) was fitted to variants with negative residuals, with 95% bootstrap confidence intervals shown as gray shading. Three-star ClinVar pathogenic variants (variants reviewed by an expert panel) that reside outside of orthosteric sites are marked as black triangles and labeled.

Consistent with our previous observations, we found that predicted stability is necessary but insufficient to explain pathogenicity. In BRCA1 (Fig. 3A–C), benign variants were mostly predicted to be stable. However, pathogenic variants split into two distinct populations: those that are severely destabilizing (high ThermoMPNN ΔΔGf) and those that are predicted to be highly stable yet are functionally impaired (low SGE score, low ThermoMPNN ΔΔGf). Consequently, predicted stability showed only a moderate correlation with experimental function across all three proteins (Fig. S4A, Spearman’s ρ = 0.40, 0.40, 0.25, respectively), hinting that a large fraction of pathogenic variants in these proteins likely arise from other mechanisms.

We next calculated the “SGE-ThermoMPNN ΔΔGf residuals” to isolate these functional effects (Fig. 3B, Fig. S3B). In BRCA1, large negative residuals clustered around known functional interfaces. In the BRCT domain, residuals highlighted the phosphorpeptide-binding groove, while in the RING domain, high-residual variants mapped to the zinc ion binding and BARD1 heterodimerization interface (Fig. 3A, C). Quantitatively, pathogenic variants also exhibited significantly larger negative residuals (Fig. 3D).

Importantly, as observed in PTEN and GCK, these functional residuals exhibited the characteristic distance-dependent decay (Fig. 3E, Fig. S3C). In the BRCA1-RING domain, for example, the magnitude of the residuals was the highest at the dimerization interface and zinc ion binding sites and dampened as the distance from the functional site increased.

Finally, we leveraged these high-quality SGE datasets to validate the reliability of our computational models. To determine whether the “missing variance” between predicted stability and fitness reflects genuine biology or model noise, we compared ESM-1v fitness scores directly against the SGE “ground truth” (Fig. S4). To quantify this, we decomposed the variance into the biological signal (ESM-1v–ThermoMPNN ΔΔGf) and the predictive noise (SGE –ESM-1v). We found that the biological signal is consistently larger in magnitude than the model noise (Fig. S4B). Further, while the biological signal exhibits a clear distance-dependent decay from orthosteric sites–consistent with physical propagation–the predictive noise is spatially uniform. This shows that the residuals capture a robust, spatially coherent functional signal rather than random model error.

Experimental allosteric maps confirm distance-dependent decay

To confirm that these “activity-abundance residuals” specifically isolate binding and allosteric effects, we turned to experimental allosteric maps–datasets where the impact of every mutation on both folding stability (ΔΔGf) and binding affinity (ΔΔGb) has been explicitly measured. Recent advances have enabled the generation of such comprehensive maps for individual proteins and protein domains35–39. In these systems, mutations that alter binding energy without directly contacting the ligand must, by definition, act through an allosteric mechanism. To evaluate this, we focused on two domains with complete mutational maps: the third PDZ domain of PSD-95 and the second SH3 domain of GRB2 (Fig. 4A).

Fig. 4. Predicted ESM-1v score vs measured stability in PSD-95-PDZ3 and GRB2-SH3.

Fig. 4

A Crystal structures of PSD-95-PDZ3 bound to CRIPT (A, PDB: 1BE9) and GRB2-SH3 bound to GAB2 (PDB: 2VWF), colored by the median ESM-1v–measured ΔΔGf residual for each site. B Scatter plots between predicted ESM-1v and experimentally measured folding stability (top: PSD-95-PDZ3, bottom: GRB2-SH3). A LOESS regression (black line) defines the baseline stability-function relationship. C Distributions of ESM-1v–measured ΔΔGf residuals stratified by functional class: orthosteric (direct binding contacts), allosteric (experimentally defined distal functional sites), and others. For all distributions, the violins show the kernel density of all values. Within each violin, an embedded boxplot indicates the median (center horizontal line; values labelled) and the interquartile range (IQR), with whiskers extending to 1.5× the IQR; an overlaid vertical line spans the 2.5th–97.5th percentiles (central 95% of values). Comparisons were performed using a one-sided Wilcoxon rank-sum test. D Distance-dependent decay of functional effects. The absolute magnitude of the residual was plotted against the minimal distance to the ligand. An exponential decay model (black curve representing the least-squares fit) was fitted to variants with negative residuals, with 95% bootstrap confidence intervals shown as gray shading. Variants are colored by site type (orange: orthosteric; blue: allosteric; gray: other).

As observed in our proteome-wide analyses, folding stability changes alone explained only a modest fraction of predicted pathogenicity: R² = 0.214 for PSD-95 and R² = 0.359 for GRB2 (Fig. S5A). To isolate the stability-independent functional effects, we fitted a LOESS curve between experimentally measured ∆∆Gf and ESM-1v predicted pathogenicity for each domain and computed residuals, quantifying the extent to which mutations disrupt function beyond what would be expected from their impact on folding stability alone (Fig. 4B).

As expected, the largest absolute median residuals were observed at orthosteric binding residues and with experimentally annotated allosteric mutations (Fig. 4C). These effects also followed a spatial pattern consistent with allosteric propagation: absolute residual magnitudes were highest near the binding interface and decayed with increasing 3D distance from the ligand (PSD-95-PDZ3: slope = -0.134, P = 6.74e-07; GRB2-SH2: slope = -0.109, P = 1.04e-05). Visualizing the ESM-1v–measured ∆∆Gf residual further illustrates this distance dependence of predicted functional effects away from the binding interface (Fig. 4D).

We next evaluated experimentally measured ∆∆Gb values to directly assess allosteric effects. For residues outside the binding interface, changes in ∆∆Gb isolate the allosteric component of binding disruption. These values explained 30.2% of ESM-1v variance in PSD-95 and 4.5% in GRB2 (Fig. S5A). The smaller contribution of ∆∆Gb to pathogenicity in GRB2-SH3 likely reflects the domain’s very small size and reduced allostery.

Importantly, binding free energy changes (∆∆Gb) also showed a distance-dependent decay from the ligand35–39 (Fig. S5B). Across residues, ∆∆Gb values correlated with ESM-1v–measured ∆∆Gf residuals (Fig. S5C), with a stronger relationship in GRB2 (Spearman’s ρ = –0.67) than in PSD-95 (ρ = –0.32). Moreover, models incorporating both ∆∆Gf and ∆∆Gb explained more pathogenicity variance than models using stability alone—even after excluding mutations within the binding interface (Supplemental Tables 4–7). This further demonstrates the allosteric contribution to pathogenicity is independent of changes in fold stability.

Functional effects decay with distance

The results from experimental data above suggest a general physical principle: deleterious mutations often disrupt function by breaking long-range energetic networks rather than locally destabilizing the fold. To test the universality of this principle, we extended our analysis to a set of six structurally and functionally diverse human allosteric proteins (Fig. 5, Supplemental Table 3). For each allosteric protein in our test set, we compared predicted pathogenicity (ESM-1v) to predicted folding stability changes (ΔΔGf) from ThermoMPNN. We then fitted a LOESS curve and computed residuals to quantify variant effects not attributable to stability changes (Fig. 5B, Fig. S6A).

Fig. 5. Predicted ESM-1v score vs predicted thermoMPNN stability in allosteric enzymes.

Fig. 5

A Crystal structure of PDK1 (PDB: 4RQK) colored by the median ESM-1v–ThermoMPNN ΔΔGf residuals. The ATP substrate is shown in orange, and the allosteric inhibitor RS1 is labeled in dark cyan. B Scatter plot of ThermoMPNN-predicted ΔΔGf versus ESM-1v scores for missense mutations in PDK1. A LOESS curve (black line) models the trend, with residuals quantifying the deviation between predicted ΔΔGf and predicted ESM-1v fitness. C Per-residue relationship between the absolute median ESM-1v–ThermoMPNN ΔΔGf residual and the minimal side chain heavy atom distance to the ATP substrate. An exponential decay model (black curve representing the least-squares fit) was fitted, with 95% bootstrap confidence intervals shown as gray shading. Sites are colored by functional class: orthosteric sites (orange), allosteric sites contacting RS1 (cyan triangles), and other sites (sage). D Distributions of per-site absolute median residuals for orthosteric, allosteric, and other sites in PDK1. For all distributions, the violins show the kernel density of all values. Within each violin, an embedded boxplot indicates the median (center horizontal line; values labelled) and the interquartile range (IQR), with whiskers extending to 1.5× the IQR; an overlaid vertical line spans the 2.5th–97.5th percentiles (central 95% of values). Comparisons were performed using a one-sided Wilcoxon rank-sum test. E Crystal structures of diverse allosteric enzymes colored by median ESM-1v-ThermoMPNN ΔΔGf residuals. Proteins shown: Glucokinase (GCK) (PDB: 1V4S), Protein Tyrosine Phosphatase 1B (PTP1B) (PDB: 1PTY), Caspase-1 (CASP1) (PDB: 2HBQ), Checkpoint Kinase 1 (CHK1) (PDB: 2E9N), and Isocitrate Dehydrogenase 1 (IDH1) (PDB:6IO0). Bound substrates—glucose (GLC), O-phosphotyrosine (PTR), peptide-like inhibitor (PI), CHK1 inhibitor (76A), and NDP—are shown in orange. Known allosteric regulators (MRK, AOU) are highlighted in dark cyan.

Across all six proteins, many variants exhibited large residuals, indicating functional effects not accounted for by changes in fold stability. ESM-1v and ΔΔGf were moderately correlated (Spearman’s ρ = –0.25 to –0.67), with enrichment of large residuals at annotated orthosteric sites and distal regulatory positions where known allosteric regulators bind (Fig. 5, Fig. S6).

We next asked whether these stability-independent functional effects follow a spatial pattern consistent with allosteric propagation35–38,65,66. In all proteins, the absolute median residuals were strongest for residues close to the active site, with the predicted functional effects of mutations decaying with 3D distance from the ligand (Fig. 5C, Fig. S6B). This trend was statistically significant: in a linear model regressing the log-transformed absolute residuals on distance, increasing distance was significantly associated with weaker functional effects (Supplemental Table 3). Across proteins, the half-distances for decay (d1/2) are comparable to the distance-dependent decay of allosteric mutational effects observed in experimental allosteric maps35–39 (Supplemental Table 3).

We also compared the magnitudes of ESM-1v–ThermoMPNN ΔΔGf residuals at orthosteric sites, distal sites where known allosteric regulators bind, and the rest of the positions. In three out of the six proteins, absolute median residuals were significantly higher at both orthosteric and known allosteric regulator contact sites than at other positions (Fig. 5D, Fig. S6C). In PTP1B, CASP1 and IDH1, allosteric regulator contact sites also showed larger absolute residuals on average (Fig. S6C), although the difference was not statistically significant. It is important to note that this definition of allosteric sites is conservative, as it is based only on the binding sites of known allosteric regulators. Most allosteric proteins will harbor additional allosteric sites that remain untested.

Notably, the distance-dependent decay is not uniform. Some positions exhibit residuals much larger than expected for their distance from the active site. We refer to these as allosteric hotspots—residues with larger functional impact than expected. These hotspots include binding sites of known allosteric regulators, such as inhibitors of protein tyrosine phosphatase 1B (PTP1B) and compound RS1, which targets pyruvate dehydrogenase kinase 1 (PDK1)70,71. These findings mirror the distance-dependent decay of mutational effects observed in experimental allosteric maps—including decay from binding interfaces in protein interaction domains35, in KRAS37, and in activity-altering mutations in the Src kinase36, as well as in the allosteric phytohormone receptor PYL138. They are also similar to the decay of abundance corrected mutational effects in SGE genes (Fig. 3) and the two protein binding domains (Fig. 4) presented above.

Finally, benchmarking our computational residuals against diverse experimental datasets33,35,60–62 confirms a consistent and significant correlation (Fig. S7). These results suggest that allosteric mechanisms potentially underlie a substantial and underappreciated component of pathogenicity in missense variants, and that the magnitude of these effects decays in a conserved, distance-dependent manner from functional sites.

Discussion

Taken together, our results indicate that allostery is an important cause of LOF variant pathogenicity. Although many pathogenic variants destabilize proteins24–28, changes in stability only partially account for the pathogenicity of variants. Both state-of-the-art computational predictions (Figs. 1,5), and direct experimental measurements (Figs. 2–4) indicate that the indirect, allosteric effects of mutations on activity are a frequent mechanism of pathogenicity. In both PTEN and GCK, for example, many known pathogenic variants have no or weak effects on stability but strong experimentally measured allosteric effects on activity.

To date, comprehensive experimental allosteric maps have only been generated for a handful of proteins35–39, and a limitation of our study is the extrapolation of conclusions from these proteins to the whole proteome. Nevertheless, results from our proteome-wide analyses using state-of-the-art computational predictors are highly consistent with the conclusions from the experimental allosteric maps. Ultimately, it will be important in future work to scale up the production of experimental allosteric maps, particularly for proteins that harbor large numbers of clinically observed variants33,60,62. Furthermore, we have only experimentally considered protein-protein binding and enzyme catalytic activity in this study; future efforts should develop methods to quantify allosteric effects on additional molecular activities, such as DNA binding and small-molecule interactions72. A comprehensive understanding of pathogenic mechanisms will ultimately require experimental tools capable of capturing allosteric regulation across the full spectrum of disease-relevant molecular activities.

Additional caveats should be noted. First, ESM-1v is a proxy for variant pathogenicity and evolutionary fitness derived from protein sequences; while it captures many aspects of mutational impact on protein function, it is not a direct measure of clinical pathogenicity. Consequently, variants with large ESM-1v-∆∆Gf residuals that are currently not classified as pathogenic likely represent a combination of unannotated variants, subclinical functional changes that do not manifest as disease, and artifacts of the predictive models. Second, our computational analyses and the equivalent experimental data aim to quantify allostery, but they do not identify the underlying causal molecular mechanisms, which likely involve changes in both conformation and protein dynamics3,7,73–75.

However, despite this diversity of possible mechanisms, a conserved principle of allosteric communication appears to be the distance-dependent decay that we have highlighted here. An approximately exponential distance-dependent decay of functional mutational effects is seen in all seven experimental allosteric maps published to date35–39 and it is also observed in the PTEN and GCK allosteric maps that we have presented here (Fig. 2). It is also visible when plotting computational predictions of variant fitness after accounting for predicted stability effects (Fig. 5). Distance-dependent decay has been observed for evolutionary conservation in enzyme active sites, for chemical shift perturbations in nuclear magnetic resonance (NMR) experiments, and for energetic couplings in molecular dynamic simulations65,66,76. We believe, therefore, that distance-dependent allosteric decay is likely to be a conserved general principle of protein biophysics.

More broadly, the activity–abundance residual framework establishes a foundation for systematically mapping functional communication across the proteome. By quantifying the prevalence and governing principles of distance-dependent decay at scale, this approach will not only reveal the extent of allostery in human proteins but also identify candidate pockets most amenable to therapeutic modulation.

Pervasive and distance-dependent allostery has important implications for human genetics. First, many LOF variants are likely to have allosteric mechanisms of action. Second, our data predict that pathogenic variants are more likely to have allosteric mechanisms when they are located structurally close to active sites. Third, and conversely, drugs targeting non-orthosteric protein pockets are more likely to have strong allosteric effects when they target pockets closer to a protein’s active site. In future work, it will be important to further refine the rules of allostery, allowing improved mechanistic interpretation of variant pathogenicity and prediction of the most effective sites to target to therapeutically modulate protein function.

Methods

Data processing for human proteome-wide predictions

Proteome-wide ESM-1v scores were obtained from Zenodo (May 2025)77, with detailed methods described in the accompanying publication78. AlphaFold predicted structures were downloaded from the AlphaFold Protein Structure Database (June 2025)79 and filtered based on criteria used in the reference24, excluding proteins shorter than 16 amino acids or longer than 2700 amino acids. This yielded a final dataset of 20,144 unique, unfragmented human proteins, of which 19,288 had matching ESM-1v scores and were retained for downstream analyses.

Stability changes upon mutation (ΔΔGf) were predicted for all possible missense mutations using ThermoMPNN v1.0.052 with default settings applied to the AlphaFold-predicted structures.

Protein core positions were identified based on relative solvent-accessible surface area (rSASA), calculated using DSSP with BioPython with default settings80. Residues were classified as “core” if they were not predicted to be disordered or part of a binding site and had rSASA <0.2, following criteria from previous publications81. Binding and active sites were annotated using proteome-wide predictions from the PIONEER interactome82.

Proteins were grouped into functional classes according to the “Function and compartment-based” classification provided by The Human Protein Atlas83–85. Clinical variant annotations were obtained from the supplementary materials described above24. Variants were originally extracted from ClinVar54,55,86 and processed accordingly26, retaining only single-nucleotide variants with at least a one-star review status and mapped to GRCh38. For this study, variants labeled pathogenic or likely pathogenic were grouped as “pathogenic”, while those labeled benign or likely benign were grouped as “benign”. Variants with other labels or synonymous mutations were excluded. This yielded 69,909 benign and 32,147 pathogenic clinically annotated variants. ThermoMPNN-predicted missense mutations were classified as: destabilizing if ΔΔGf > 1 kcal/mol, stabilizing if ΔΔGf < –0.5 kcal/mol, following the thresholds defined in the original publication52.

A subset of allosteric proteins, listed in Supplemental Table 3, was selected for targeted analyses. These included all full-length human proteins present in two curated datasets: six proteins from the 20-protein benchmark set used to validate the allosteric network model Ohm87, and one additional protein, isocitrate dehydrogenase 1 (IDH1), included from the reference88. The enzyme prothrombin was excluded because its allosteric regulator functions as an activator rather than an inhibitor. For glucokinase (GCK), annotations of allosteric sites were based on literature reports rather than crystal structure contacts, as its allosteric regulator (MRK) is also an activator. Orthosteric and allosteric residues were defined based on direct atomic contacts with bound substrates or modulators as observed in crystal structures. Residues were considered in contact if their atoms exhibited van der Waals overlap with atoms of the ligand, as determined using the ChimeraX VDW overlap method89.

Data processing for mega-scale, domainome 1.0, VAMP-seq, and SGE datasets

ESM-1v scores for variants in the ‘Mega-scale’ dataset were obtained from the ProteinGym database90. For the ‘Domainome 1.0’ dataset, ESM-1v scores were extracted from the supplementary materials of the original publication28. ESM-1v scores for ‘VAMP’-seq and SGE variants were derived from the proteome-wide ESM-1v dataset described above. ‘VAMP’-seq abundance fitness data were collected for nine full-length human proteins: PRKN, OCT, CYP2C9, TPMT, ASPA, NUDT15, VKOR, PTEN, and GCK (Supplemental Table 2)33,61,91–96. For the GCK dataset, we utilized experimentally measured functional scores and excluded all imputed values.

AlphaMissense predictions were downloaded from the Zenodo repository [https://zenodo.org/records/10813168] in November 2025, and a logit transformation was applied to the raw AlphaMissense scores. ESM-251 scores were generated using the 3-billion parameter model (esm2_t36_3B_UR50D); consistent with the ESM-1v protocol, variant effects were quantified as the log-likelihood ratio between the mutant and wild-type amino acids at a masked position. Vertebrate PhyloP scores were extracted from the CADD97 website [https://cadd.gs.washington.edu/download] in January 2025 and aggregated by amino acid substitution to obtain the median conservation scores for each missense variant.

Clinical variant annotations for both the ‘Domainome 1.0’ and ‘VAMP’-seq proteins were obtained from the supplementary data of the reference described above24, following the same filtering criteria.

For SGE data, to directly compare the biological signal against predictive noise, we rank-normalized both the experimental SGE scores and ESM-1v predictions to a common 0–1 scale. We then calculated the ‘predictive noise’ as the difference between these normalized values for each variant (Normalized SGE – Normalized ESM) and analyzed whether this deviation followed a spatially coherent, distance-dependent decay characteristic of allostery or the spatial uniformity expected of random model error.

Variant classification

  • ‘Domainome 1.0’ variants were classified based on criteria defined in the original publication. Mutations were considered stabilizing if they had a statistically significant stabilizing effect (FDR_stabilizing < 0.1) and a scaled fitness > 0.3. Conversely, mutations were classified as destabilizing if they had a significant destabilizing effect (FDR_destabilizing < 0.1) and a scaled fitness <0.

  • ‘VAMP’-seq variants were categorized based on abundance thresholds defined by ProteinGym90. Mutations falling below the defined abundance cutoff were considered destabilizing.

  • ‘Mega-scale’ variants were classified using thresholds established in the original study27. Note that the ΔΔGf directionality in this dataset is different relative to others: mutations with ΔΔGf > 1 kcal/mol were classified as stabilizing, while those with ΔΔGf < –1 kcal/mol were classified as destabilizing.

Data processing for experimental allosteric maps

Analyses for PSD-95-PDZ3 and GRB2-SH3 were performed using variants with corresponding ESM-1v scores, obtained from the human proteome-wide prediction dataset. Orthosteric sites were defined based on published structural annotations35. Variants were classified as allosteric if they occurred outside orthosteric sites and exhibited binding ΔΔG (ΔΔGb) values greater than or equal to the mean ΔΔGb of orthosteric-site mutations. Evolutionary coupling scores were obtained from the EVcouplings server, and for each position, we calculated the maximum coupling strength (cn) to any residue within the defined orthosteric interface44,98.

For PTEN and GCK, analyses were based on variants with both abundance and activity fitness measurements. In PTEN, orthosteric site annotations were manually curated from the literature99–105. These included the catalytic loops (arginine loop, residues 35–49; WPD loop, 88–98; P loop, 123–130; TI loop, 159–171), calcium-binding regions (residues 200–212, 226–238, and 258–268), the Cα2 loop (residues 327–335), the membrane-binding helix (residues 151–174), and the PI(4,5)P₂-binding motif (PBM, residues 1–15).

In GCK, orthosteric sites were defined based on annotated ligand-contacting residues described in previous structural studies106–109. These included glucose-binding residues (positions 151–153, 168–169, 204–206, 225–231, 254–258, 287, and 290) and ATP-binding residues (positions 78–85, 295–296, 331–333, 336, and 410–416).

For BRCA1, we defined orthosteric sites within the RING domain as residues contacting BARD1-RING (3, 6–8, 10–11, 14–15, 17–18, 21–22, 38, 76–79, 81–82, 85, 88–89, 92–93, 96–97) or coordinating Zn2+ ions (24, 27, 39, 41, 44, 47, 61, 64). In the BRCA1-BRCT domain, sites were defined as residues comprising the BACH1 binding interface (1655–1658, 1690–1691, 1696, 1698–1702, 1740–1741, 1774–1775, 1811, 1813, 1836, 1839).

For RAD51C, orthosteric sites were defined as the ATP-binding Walker A and Walker B motifs (residues 125–132 and 238–242) and the nuclear localization sequence110 (residues 366–370). For BAP1, orthosteric sites were defined as the catalytic residues111 (85, 91, 169, 184).

For PTEN, GCK, and the SGE genes, updated ClinVar data were downloaded from the NCBI ClinVar website (https://www.ncbi.nlm.nih.gov/clinvar/) in January 2026 and processed using the same filtering criteria described above.

Statistical analysis and data visualization

Linear regression models were used to analyze proteome-wide variant effect predictions, large-scale abundance datasets, and individual protein domains with experimentally mapped allosteric mutations. To evaluate the relative contributions of folding ΔΔGf and binding ΔΔGb to ESM-1v scores, both ordinary linear regression and linear mixed-effects models were employed. Linear regression was performed using the lm() function from the stats package in R, while mixed-effects models were fitted using the lmer() function from the lme4 package112,113. As results from both approaches were highly consistent (Supplemental Tables 4–7), linear regression outcomes are reported for interpretability. Adjusted R² values were extracted directly from the summary output of the fitted linear models in R. To estimate variability and provide robust confidence intervals, adjusted R² was also bootstrapped using 1000 resampling iterations. Model diagnostics included residual and prediction analysis, and model comparisons were performed using ANOVA, with F-statistics and p-values reported.

LOESS residual analyses were performed using only mutations located outside orthosteric sites. LOESS fitting was conducted with the loess() function in R, using a span of 0.7 and the family = “symmetric” setting. For all proteins, distances were calculated as the minimum distance between any side-chain heavy atom of the residue and the ligand or orthosteric site.

To generate abundance-matched controls, we performed bootstrap resampling using all non-pathogenic variants as the sampling pool. Variants were grouped into fixed-width abundance bins (bin size = 0.5) spanning the full range of experimentally measured abundance scores. For each bootstrap iteration (n = 1000), we generated a matched null dataset by replacing each pathogenic variant with a randomly sampled non-pathogenic variant from the same abundance bin. Statistical significance between pathogenic variants and the abundance-matched null model was assessed using an empirical p-value, defined as the fraction of bootstrap iterations where the median residual of the null distribution was lower than or equal to the observed median residual of the pathogenic variants. Additionally, differences between benign and pathogenic variants were assessed using a Wilcoxon rank-sum test.

To quantify the distance-dependent decay of mutational effects, nonlinear least-squares regression was applied to the LOESS residuals. The allosteric decay model included two parameters: the decay rate (−b) and the x-intercept (a). This fitting was performed using only mutations with negative residuals that fell below the LOESS curve. Nonlinear regression was conducted with the nls() function in R, using initial parameter values of a = 1 and b = 0.1. 95% confidence intervals were calculated by bootstrapping the model parameter with 1000 replicates.

All statistical analyses were performed in R version 4.5.0113. Experimental schematics were created using the NIH BioArt Source, while all other plots were generated using the ggplot2 package version 4.0.3114. Three-dimensional protein structure visualizations were rendered using ChimeraX version 1.6.189.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

Reporting Summary (84.6KB, pdf)

Source data

Source Data (24.8MB, xlsx)

Acknowledgements

We thank Taylor Mighell, Mike Thompson, Maximilian R. Stammnitz, Albert Escobedo, Taraneh Zarin, Antoni Beltran, Xianghua Li, Amit Sud, and other former and current members of the Lehner lab for feedback and discussions. We thank the Sanger HumGen/GenGen Informatics for their assistance with the computational infrastructure.

Author contributions

X.L. performed all analyses. X.L. and B.L. conceived the project, designed analyses, and wrote the manuscript.

Peer review

Peer review information

Nature Communications thanks Lea Starita and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. A peer review file is available.

Funding

Work in the lab of B.L. was funded by Wellcome (Grant reference: 220540/Z/20/A, ‘Wellcome Sanger Institute Quinquennial Review 2021-2026’), a European Research Council (ERC) Advanced (883742) grant, the Spanish Ministry of Science and Innovation (LCF/PR/HR21/52410004, EMBL Partnership, Severo Ochoa Center of Excellence), Agència de Gestió d’Ajuts Universitaris i de Recerca (AGAUR, 2021 SGR 01226), and the CERCA Program/Generalitat de Catalunya. X.L is funded by the Wellcome Sanger Institute PhD Graduate program.

Data availability

Proteome-wide predictions of protein stability changes and corresponding ESM-1v scores generated in this study have been deposited in the Zenodo database under accession code 18386427 as Supplementary Table 1. Raw ThermoMPNN predictions of protein stability changes generated in this study are available at Zenodo under accession code 18381534. Large-scale experimental measurements of protein abundance changes across human domains, including datasets from the ‘Mega-scale’ and ‘Domainome’ studies, are accessible via their respective supplementary materials. ‘VAMP’-seq abundance data are available from MaveDB [10.1186/s13059-019-1845-6] and ProteinGym [10.1101/2023.12.07.570727]90,115. Experimental ∆∆G values for PSD95-PDZ3 and GRB2-SH3 domains, and experimental measurements of abundance and activity-based fitness for PTEN and GCK, and SGE scores, are available through the original studies cited in the references. PDB codes include 1D5R, 1V4S, 1T29, 1JM7, 1BE9, 2VWF, 4RQK, 1PTY, 2HBQ, 2E9N, 6IO0. Source data are provided with this paper.

Code availability

All code and metadata to reproduce the analyses are hosted on GitHub [https://github.com/lehner-lab/allostery_pathogenicity].

Competing interests

B.L. is a founder and shareholder of ALLOX. The remaining author declares no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Supplementary information

The online version contains supplementary material available at 10.1038/s41467-026-74517-8.

References

  • 1.Fenton, A. W. Allostery: an illustrated definition for the “second secret of life.” Trends Biochem. Sci.33, 420–425 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Guarnera, E. & Berezovsky, I. N. Allosteric drugs and mutations: chances, challenges, and necessity. Curr. Opin. Struct. Biol.62, 149–157 (2020). [DOI] [PubMed] [Google Scholar]
  • 3.Motlagh, H. N., Wrabl, J. O., Li, J. & Hilser, V. J. The ensemble nature of allostery. Nature508, 331–339 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Wu, N., Barahona, M. & Yaliraki, S. N. Allosteric communication and signal transduction in proteins. Curr. Opin. Struct. Biol.84, 102737 (2024). [DOI] [PubMed] [Google Scholar]
  • 5.Nussinov, R. & Tsai, C.-J. Allostery in disease and in drug discovery. Cell153, 293–305 (2013). [DOI] [PubMed] [Google Scholar]
  • 6.Goodey, N. M. & Benkovic, S. J. Allosteric regulation and catalysis emerge via a common route. Nat. Chem. Biol.4, 474–482 (2008). [DOI] [PubMed] [Google Scholar]
  • 7.Gunasekaran, K., Ma, B. & Nussinov, R. Is allostery an intrinsic property of all dynamic proteins? Proteins57, 433–443 (2004). [DOI] [PubMed] [Google Scholar]
  • 8.Tsai, C.-J. & Nussinov, R. A unified view of “how allostery works”. PLoS Comput. Biol.10, e1003394 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Sivcev, S., Kudova, E. & Zemkova, H. Neurosteroids as positive and negative allosteric modulators of ligand-gated ion channels: P2X receptor perspective. Neuropharmacology234, 109542 (2023). [DOI] [PubMed] [Google Scholar]
  • 10.Keov, P., Sexton, P. M. & Christopoulos, A. Allosteric modulation of G protein-coupled receptors: a pharmacological perspective. Neuropharmacology60, 24–35 (2011). [DOI] [PubMed] [Google Scholar]
  • 11.Wu, P., Clausen, M. H. & Nielsen, T. E. Allosteric small-molecule kinase inhibitors. Pharmacol. Ther.156, 59–68 (2015). [DOI] [PubMed] [Google Scholar]
  • 12.Yu, Y., Yu, Q. & Zhang, X. Allosteric inhibition of HIF-2α as a novel therapy for clear cell renal cell carcinoma. Drug Discov. Today24, 2332–2340 (2019). [DOI] [PubMed] [Google Scholar]
  • 13.Mighell, T. L. & Lehner, B. A small molecule stabilizer rescues the surface expression of nearly all missense variants in a GPCR. Nat Struct. Mol. Biol. 32, 2429–2440 (2025). [DOI] [PMC free article] [PubMed]
  • 14.Sheik Amamuddy, O. et al. Integrated computational approaches and tools for allosteric drug discovery. Int. J. Mol. Sci.21, 847 (2020). [DOI] [PMC free article] [PubMed]
  • 15.Shen, Q. et al. Proteome-scale investigation of protein allosteric regulation perturbed by somatic mutations in 7000 cancer genomes. Am. J. Hum. Genet.100, 5–20 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Ostrem, J. M., Peters, U., Sos, M. L., Wells, J. A. & Shokat, K. M. K-Ras(G12C) inhibitors allosterically control GTP affinity and effector interactions. Nature503, 548–551 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Kwan, A. K., Piazza, G. A., Keeton, A. B. & Leite, C. A. The path to the clinic: a comprehensive review on direct KRASG12C inhibitors. J. Exp. Clin. Cancer Res.41, 27 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Bandaranayake, R. M. et al. Crystal structures of the JAK2 pseudokinase domain and the pathogenic mutant V617F. Nat. Struct. Mol. Biol.19, 754–759 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Hubbard, S. R. Mechanistic insights into regulation of JAK2 tyrosine kinase. Front. Endocrinol.8, 361 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Sharma, S. V., Bell, D. W., Settleman, J. & Haber, D. A. Epidermal growth factor receptor mutations in lung cancer. Nat. Rev. Cancer7, 169–181 (2007). [DOI] [PubMed] [Google Scholar]
  • 21.Yun, C.-H. et al. The T790M mutation in EGFR kinase causes drug resistance by increasing the affinity for ATP. Proc. Natl. Acad. Sci. USA105, 2070–2075 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Nussinov, R., Yavuz, B. R. & Jang, H. Allostery in disease: anticancer drugs, pockets, and the tumor heterogeneity challenge. J. Mol. Biol.437, 169050 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Nussinov, R. & Jang, H. The value of protein allostery in rational anticancer drug design: an update. Expert Opin. Drug Discov.19, 1071–1085 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Cagiada, M., Jonsson, N. & Lindorff-Larsen, K. Decoding molecular mechanisms for loss-of-function variants in the human proteome. Elife14, RP108160 (2025).
  • 25.Jänes, J. et al. Predicted mechanistic impacts of human protein missense variants. bioRxiv10.1101/2024.05.29.596373 (2024).
  • 26.Tiemann, J. K. S., Zschach, H., Lindorff-Larsen, K. & Stein, A. Interpreting the molecular mechanisms of disease variants in human transmembrane proteins. Biophys. J.122, 2176–2191 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Tsuboyama, K. et al. Mega-scale experimental analysis of protein folding stability in biology and design. Nature620, 434–444 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Beltran, A., Jiang, X., Shen, Y. & Lehner, B. Site-saturation mutagenesis of 500 human protein domains. Nature1, 10 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Backwell, L. & Marsh, J. A. Diverse molecular mechanisms underlying pathogenic protein mutations: beyond the loss-of-function paradigm. Annu. Rev. Genomics Hum. Genet.23, 475–498 (2022). [DOI] [PubMed] [Google Scholar]
  • 30.Taipale, M. Disruption of protein function by pathogenic mutations: common and uncommon mechanisms. Biochem. Cell Biol.97, 7 (2019). [DOI] [PubMed]
  • 31.Notbohm, J. & Perica, T. Biochemistry and genetics are coming together to improve our understanding of genotype to phenotype relationships. Curr. Opin. Struct. Biol.89, 102952 (2024). [DOI] [PubMed] [Google Scholar]
  • 32.Cagiada, M. et al. Understanding the origins of loss of protein function by analyzing the effects of thousands of variants on activity and abundance. Mol. Biol. Evol.38, 3235–3246 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Gersing, S. et al. Characterizing glucokinase variant mechanisms using a multiplexed abundance assay. Genome Biol.25, 98 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Høie, M. H., Cagiada, M., Beck Frederiksen, A. H., Stein, A. & Lindorff-Larsen, K. Predicting and interpreting large-scale mutagenesis data using analyses of protein stability and conservation. Cell Rep.38, 110207 (2022). [DOI] [PubMed] [Google Scholar]
  • 35.Faure, A. J. et al. Mapping the energetic and allosteric landscapes of protein binding domains. Nature604, 175–183 (2022). [DOI] [PubMed] [Google Scholar]
  • 36.Beltran, A., Naqvi, M. M., Faure, A. J. & Lehner, B. The allosteric landscape of the Src kinase. Sci. Adv.12, eaea2726 (2026). [DOI] [PMC free article] [PubMed]
  • 37.Weng, C., Faure, A. J., Escobedo, A. & Lehner, B. The energetic and allosteric landscape for KRAS inhibition. Nature626, 643–652 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Stammnitz, M. R. & Lehner, B. The genetic architecture of an allosteric hormone receptor. Nat. Commun.17, 4735 (2026). [DOI] [PMC free article] [PubMed]
  • 39.Mighell, T. L. & Lehner, B. GPCR-MAPS: high-resolution functional and allosteric mapping of G protein-coupled receptor activation and bias. bioRxiv10.1101/2025.05.30.656974 (2025).
  • 40.Escobedo, A., Voigt, G., Faure, A. J. & Lehner, B. Genetics, energetics, and allostery in proteins with randomized cores and surfaces. Science389, eadq3948 (2025). [DOI] [PubMed]
  • 41.Halabi, N., Rivoire, O., Leibler, S. & Ranganathan, R. Protein sectors: evolutionary units of three-dimensional structure. Cell138, 774–786 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Xie, J., Zhang, W., Zhu, X., Deng, M. & Lai, L. Coevolution-based prediction of key allosteric residues for protein function regulation. Elife12, 81850 (2023). [DOI] [PMC free article] [PubMed]
  • 43.Yan, L., Ravasio, R., Brito, C. & Wyart, M. Architecture and coevolution of allosteric materials. Proc. Natl. Acad. Sci. USA114, 2526–2531 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Marks, D. S. et al. Protein 3D structure computed from evolutionary sequence variation. PLoS One6, e28766 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Teşileanu, T., Colwell, L. J. & Leibler, S. Protein sectors: statistical coupling analysis versus conservation. PLoS Comput. Biol.11, e1004091 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Salinas, V. H. & Ranganathan, R. Coevolution-based inference of amino acid interactions underlying protein function. Elife7, e34300 (2018). [DOI] [PMC free article] [PubMed]
  • 47.Morcos, F. et al. Direct-coupling analysis of residue coevolution captures native contacts across many protein families. Proc. Natl. Acad. Sci. Usa.108, E1293–E1301 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Gueudré, T., Baldassi, C., Zamparo, M., Weigt, M. & Pagnani, A. Simultaneous identification of specifically interacting paralogs and interprotein contacts by direct coupling analysis. Proc. Natl. Acad. Sci. USA113, 12186–12191 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Meier, J. et al. Language models enable zero-shot prediction of the effects of mutations on protein function. bioRxiv10.1101/2021.07.09.450648 (2021).
  • 50.Cheng, J. et al. Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science381, eadg7492 (2023). [DOI] [PubMed] [Google Scholar]
  • 51.Lin, Z. et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science379, 1123–1130 (2023). [DOI] [PubMed] [Google Scholar]
  • 52.Dieckhaus, H., Brocidiacono, M., Randolph, N. Z. & Kuhlman, B. Transfer learning to leverage larger datasets for improved prediction of protein stability changes. Proc. Natl. Acad. Sci. USA121, e2314853121 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Dauparas, J. et al. Robust deep learning–based protein sequence design using ProteinMPNN. Science378, 49–56 (2022). [DOI] [PMC free article] [PubMed]
  • 54.Landrum, M. J. et al. ClinVar: public archive of interpretations of clinically relevant variants. Nucleic Acids Res.44, D862–D868 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Landrum, M. J. et al. ClinVar: improvements to accessing data. Nucleic Acids Res.48, D835–D844 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Tokuriki, N. & Tawfik, D. S. Stability effects of mutations and protein evolvability. Curr. Opin. Struct. Biol.19, 596–604 (2009). [DOI] [PubMed] [Google Scholar]
  • 57.Pilarski, R. PTEN hamartoma tumor syndrome: a clinical overview. Cancers11, 844 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Froguel, P. et al. Close linkage of glucokinase locus on chromosome 7p to early-onset non-insulin-dependent diabetes mellitus. Nature356, 162–164 (1992). [DOI] [PubMed] [Google Scholar]
  • 59.Das, S. K. et al. Polymorphisms in the glucokinase-associated, dual-specificity phosphatase 12 (DUSP12) gene under chromosome 1q21 linkage peak are associated with type 2 diabetes. Diabetes55, 2631–2639 (2006). [DOI] [PubMed] [Google Scholar]
  • 60.Gersing, S. et al. A comprehensive map of human glucokinase variant activity. Genome Biol.24, 97 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Mighell, T. L., Evans-Dutson, S. & O’Roak, B. J. A saturation mutagenesis approach to understanding PTEN lipid phosphatase activity and genotype-phenotype relationships. Am. J. Hum. Genet.102, 943–955 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Matreyek, K. A. et al. Multiplex assessment of protein variant abundance by massively parallel sequencing. Nat. Genet.50, 874–882 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Kamata, K., Mitsuya, M., Nishimura, T., Eiki, J.-I. & Nagata, Y. Structural basis for allosteric regulation of the monomeric allosteric enzyme human glucokinase. Structure12, 429–438 (2004). [DOI] [PubMed] [Google Scholar]
  • 64.Larion, M. & Miller, B. G. 23-Residue C-terminal alpha-helix governs kinetic cooperativity in monomeric human glucokinase. Biochemistry48, 6157–6165 (2009). [DOI] [PubMed] [Google Scholar]
  • 65.Rajasekaran, N., Sekhar, A. & Naganathan, A. N. A universal pattern in the percolation and dissipation of protein structural perturbations. J. Phys. Chem. Lett.8, 4779–4784 (2017). [DOI] [PubMed] [Google Scholar]
  • 66.Rajasekaran, N., Suresh, S., Gopi, S., Raman, K. & Naganathan, A. N. A general mechanism for the propagation of mutational effects in proteins. Biochemistry56, 294–305 (2017). [DOI] [PubMed] [Google Scholar]
  • 67.Findlay, G. M. et al. Accurate classification of BRCA1 variants with saturation genome editing. Nature562, 217–222 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Olvera-León, R. et al. High-resolution functional mapping of RAD51C by saturation genome editing. Cell187, 5719–5734.e19 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Waters, A. J. et al. Saturation genome editing of BAP1 functionally classifies somatic and germline variants. Nat. Genet.56, 1434–1445 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Rettenmaier, T. J. et al. A small-molecule mimic of a peptide docking motif inhibits the protein kinase PDK1. Proc. Natl. Acad. Sci. USA111, 18590–18595 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Wiesmann, C. et al. Allosteric inhibition of protein tyrosine phosphatase 1B. Nat. Struct. Mol. Biol.11, 730–737 (2004). [DOI] [PubMed] [Google Scholar]
  • 72.Li, X. & Lehner, B. T. F.-MAPS: fast high-resolution functional and allosteric mapping of DNA-binding proteins. bioRxiv10.1101/2025.10.20.683418 (2025).
  • 73.Laskowski, R. A., Gerick, F. & Thornton, J. M. The structural basis of allosteric regulation in proteins. FEBS Lett.583, 1692–1698 (2009). [DOI] [PubMed] [Google Scholar]
  • 74.Hilser, V. J., Wrabl, J. O. & Motlagh, H. N. Structural and energetic basis of allostery. Annu. Rev. Biophys.41, 585–609 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Tzeng, S.-R. & Kalodimos, C. G. Protein dynamics and allostery: an NMR view. Curr. Opin. Struct. Biol.21, 62–67 (2011). [DOI] [PubMed] [Google Scholar]
  • 76.Jack, B. R., Meyer, A. G., Echave, J. & Wilke, C. O. Functional sites induce long-range evolutionary constraints in enzymes. PLoS Biol.14, e1002452 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Livesey, B. ESM-1v predictions for human uniprot IDs. Zenodo 10.5281/ZENODO.8301668 (2023).
  • 78.Livesey, B. J. & Marsh, J. A. Variant effect predictor correlation with functional assays is reflective of clinical classification performance. Genome Biol. 26, 104 (2025). [DOI] [PMC free article] [PubMed]
  • 79.Varadi, M. et al. AlphaFold protein structure database: massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic Acids Res.50, D439–D444 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Rost, B. & Sander, C. Conservation and prediction of solvent accessibility in protein families. Proteins20, 216–226 (1994). [DOI] [PubMed] [Google Scholar]
  • 81.Hanson, J., Yang, Y., Paliwal, K. & Zhou, Y. Improving protein disorder prediction by deep bidirectional long short-term memory recurrent neural networks. Bioinformatics33, 685–692 (2017). [DOI] [PubMed] [Google Scholar]
  • 82.Xiong, D. et al. A structurally informed human protein–protein interactome reveals proteome-wide perturbations caused by disease mutations. Nat. Biotechnol.43, 1510–1524 (2025). [DOI] [PMC free article] [PubMed]
  • 83.Uhlén, M. et al. Proteomics. Tissue-based map of the human proteome. Science347, 1260419 (2015). [DOI] [PubMed] [Google Scholar]
  • 84.Thul, P. J. et al. A subcellular map of the human proteome. Science356, 3321 (2017). [DOI] [PubMed]
  • 85.Uhlen, M. et al. A pathology atlas of the human cancer transcriptome. Science357, 2507 (2017). [DOI] [PubMed]
  • 86.Landrum, M. J. et al. ClinVar: public archive of relationships among sequence variation and human phenotype. Nucleic Acids Res.42, D980–D985 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Wang, J. et al. Mapping allosteric communications within individual proteins. Nat. Commun.11, 3862 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Kannan, G. R., Hie, B. L. & Kim, P. S. Single-sequence, structure free allosteric residue prediction with protein language models. bioRxiv10.1101/2024.10.03.616547 (2024).
  • 89.Meng, E. C. et al. UCSF ChimeraX: tools for structure building and analysis. Protein Sci.32, e4792 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Notin, P. et al. ProteinGym: large-scale benchmarks for protein fitness prediction and design. Adv. Neural Inf. Process. Syst. 10, 570707 (2023).
  • 91.Clausen, L. et al. PRKN-linked familial Parkinson’s disease: cellular and molecular mechanisms of disease-linked variants. Cell. Mol. Life Sci.81, 223 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92.Yee, S. W. et al. The full spectrum of SLC22 OCT1 mutations illuminates the bridge between drug transporter biophysics and pharmacogenomics. Mol. Cell84, 1932–1947 (2024). [DOI] [PMC free article] [PubMed]
  • 93.Amorosi, C. J. et al. Massively parallel characterization of CYP2C9 variant enzyme activity and abundance. Am. J. Hum. Genet.108, 1735–1751 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Grønbæk-Thygesen, M. et al. Deep mutational scanning reveals a correlation between degradation and toxicity of thousands of aspartoacylase variants. Nat. Commun.15, 4026 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Suiter, C. C. et al. Massively parallel variant characterization identifies NUDT15 alleles associated with thiopurine toxicity. Proc. Natl. Acad. Sci. USA117, 5394–5401 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96.Chiasson, M. A. et al. Multiplexed measurement of variant abundance and activity reveals VKOR topology, active site and human variant impact. Elife9, 58026 (2020). [DOI] [PMC free article] [PubMed]
  • 97.Rentzsch, P., Witten, D., Cooper, G. M., Shendure, J. & Kircher, M. CADD: predicting the deleteriousness of variants throughout the human genome. Nucleic Acids Res.47, D886–D894 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Hopf, T. A. et al. Sequence co-evolution gives 3D contacts and structures of protein complexes. Elife3, 03430 (2014). [DOI] [PMC free article] [PubMed]
  • 99.Masson, G. R. & Williams, R. L. Structural mechanisms of PTEN regulation. Cold Spring Harb. Perspect. Med.10, a036152 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 100.Shenoy, S. S., Nanda, H. & Lösche, M. Membrane association of the PTEN tumor suppressor: electrostatic interaction with phosphatidylserine-containing bilayers and regulatory role of the C-terminal tail. J. Struct. Biol.180, 394–408 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101.Nanda, H., Heinrich, F. & Lösche, M. Membrane association of the PTEN tumor suppressor: neutron scattering and MD simulations reveal the structure of protein-membrane complexes. Methods77–78, 136–146 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 102.Thiriet, M. Intracellular Signaling Mediators in the Circulatory and Ventilatory Systems. (Springer, 2012).
  • 103.Irvine, W. A., Flanagan, J. U. & Allison, J. R. Computational prediction of amino acids governing protein-membrane interaction for the PIP cell signaling system. Structure27, 371–380.e3 (2019). [DOI] [PubMed] [Google Scholar]
  • 104.Lee, J. O. et al. Crystal structure of the PTEN tumor suppressor: implications for its phosphoinositide phosphatase activity and membrane association. Cell99, 323–334 (1999). [DOI] [PubMed] [Google Scholar]
  • 105.Serebriiskii, I. G. et al. Comprehensive characterization of PTEN mutational profile in a series of 34,129 colorectal cancers. Nat. Commun.13, 1618 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 106.DiStefano, M. T. et al. Curating clinically relevant transcripts for the interpretation of sequence variants. J. Mol. Diagn.20, 789–801 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107.Beck, T. & Miller, B. G. Structural basis for regulation of human glucokinase by glucokinase regulatory protein. Biochemistry52, 6232–6239 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108.Gloyn, A. L. et al. Insights into the structure and regulation of glucokinase from a novel mutation (V62M), which causes maturity-onset diabetes of the young. J. Biol. Chem.280, 14105–14113 (2005). [DOI] [PubMed] [Google Scholar]
  • 109.Beer, N. L. et al. Insights into the pathogenicity of rare missense GCK variants from the identification and functional characterization of compound heterozygous and double mutations inherited in CIS. Diab. Care35, 1482–1484 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 110.Rein, H. L. et al. Comprehensive RAD51C ovarian cancer variant analysis uncouples homologous recombination and replicative functions. Nat. Commun.16, 6539 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111.Peng, H. et al. Familial and somatic BAP1 mutations inactivate ASXL1/2-mediated allosteric regulation of BAP1 deubiquitinase by targeting multiple independent domains. Cancer Res.78, 1200–1213 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 112.Bates, D., Mächler, M., Bolker, B. & Walker, S. Fitting linear mixed-effects models Usinglme4. J. Stat. Softw. 67, 1–48 (2015).
  • 113.Team, R. D. C. & Others. R: a language and environment for statistical computing. R. Found. Stat. Comput.1, 409 (2016). [Google Scholar]
  • 114.Wickham, H. ggplot2. Wiley Interdiscip. Rev. Comput. Stat. 3, 180–185 (2011).
  • 115.Esposito, D. et al. MaveDB: an open-source platform to distribute and interpret data from multiplexed assays of variant effect. Genome Biol.20, 223 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Reporting Summary (84.6KB, pdf)
Source Data (24.8MB, xlsx)

Data Availability Statement

Proteome-wide predictions of protein stability changes and corresponding ESM-1v scores generated in this study have been deposited in the Zenodo database under accession code 18386427 as Supplementary Table 1. Raw ThermoMPNN predictions of protein stability changes generated in this study are available at Zenodo under accession code 18381534. Large-scale experimental measurements of protein abundance changes across human domains, including datasets from the ‘Mega-scale’ and ‘Domainome’ studies, are accessible via their respective supplementary materials. ‘VAMP’-seq abundance data are available from MaveDB [10.1186/s13059-019-1845-6] and ProteinGym [10.1101/2023.12.07.570727]90,115. Experimental ∆∆G values for PSD95-PDZ3 and GRB2-SH3 domains, and experimental measurements of abundance and activity-based fitness for PTEN and GCK, and SGE scores, are available through the original studies cited in the references. PDB codes include 1D5R, 1V4S, 1T29, 1JM7, 1BE9, 2VWF, 4RQK, 1PTY, 2HBQ, 2E9N, 6IO0. Source data are provided with this paper.

All code and metadata to reproduce the analyses are hosted on GitHub [https://github.com/lehner-lab/allostery_pathogenicity].


Articles from Nature Communications are provided here courtesy of Nature Publishing Group

RESOURCES