Abstract
Protein kinase activation is driven by conformational changes across multiple structural components, including the conserved Asp–Phe–Gly (DFG) motif, but whether these transitions follow a universal mechanism remains unclear. Here we combine over 8.3 milliseconds of distributed unbiased molecular dynamics simulations with Markov state models (MSMs) to compare the conformational landscapes of the ABL1, EGFR and MET kinase domains. To maximize unbiased sampling of functionally relevant conformational space, we use a transfer seeding strategy that steers AlphaFold2 models derived from homologous templates to sample MET conformational states absent from available experimental databases. We find that related DFG-motif geometries separate into distinct kinetic networks. These shared structural states are connected by kinase-specific activation pathways with different regulatory elements controlling the slowest step of activation. Our findings reveal that the shared nomenclature masks distinct transition mechanisms between kinase domains, revealing new regions critical for activity and targetable conformations for inhibitor design.
1. Introduction
Protein kinases are central regulators of eukaryotic cellular signalling. More than 500 protein kinases are encoded in the human genome, accounting for approximately 2% of human genes (1). Kinases control diverse processes including cell growth, proliferation, differentiation, and apoptosis through phosphorylation of protein substrates (2; 3). Given their central role in all aspects of cell life, kinase activity is tightly regulated through mechanisms such as autoinhibition, phosphorylation, and allosteric modulation (4; 5; 6). Disruption of their activity by mutation or dysregulated expression is often associated with many human diseases, such as a variety of cancers, emphasizing kinases as one of the most important target classes in drug discovery (7; 8; 9; 10).
Kinase activity is largely rooted in its catalytic kinase domain (KD), a highly conserved domain across the protein kinase family that mediates phosphorylation activity through conformational changes. KDs have a bilobal structure, consisting of a smaller N-terminal lobe and a larger C-terminal lobe, with the ATP-binding site located in the cleft between them (Supplementary Fig. S2). Several conserved structural motifs have been found to play central roles in kinase catalysis and regulation (11). Briefly, the glycine-rich loop (G-loop) in the N-lobe helps position and stabilize ATP, and the β3 strand contains a conserved lysine that forms a salt bridge with the αC helix to organize the active-site architecture. In the C-lobe, the activation loop (A-loop) is a flexible segment that controls substrate access (12). At the N-terminus of the A-loop lies the highly conserved Asp-Phe-Gly (DFG) motif, whose Asp residue coordinates a divalent magnesium ion that helps position the ATP phosphates for catalysis (11; 13; 14; 15). Together, conserved elements organize the catalytic machinery required for phosphate transfer, and their coordinated rearrangements provide a structural basis by which kinases switch between functionally distinct conformational states (4; 11; 15).
Kinase activity is not determined by a single static structure but by the dynamic equilibrium between active and inactive states. States are commonly distinguished by characteristic arrangements of the DFG motif, the αC helix, and the A-loop (Supplementary Fig. S2). Modi and Dunbrack have determined that three main conformational states—DFG-in, DFG-out, and DFG-inter—can be defined according to two distances between these structural elements (D1 and D2, Fig. 1b) (14; 16; 17). They further noted how these main states can be subdivided according to specific values of φ, ψ and χ1 angles within the DFG motif. The DFG-in BLAminus state corresponds most closely to the active kinase conformation where catalytic metal ions bind (see Methods) (14). Overall, the DFG motif is a central indicator for describing kinase conformational ensembles in which active-like, intermediate, and inactive-like states coexist.
Figure 1: DFG motif conformational states of kinase domains.

(a) Structural overlay of human ABL kinase in three conformations: DFG-in (PDB: 2f4j, purple), DFG-inter (PDB: 4xey, salmon), and DFG-out (PDB: 1opl, yellow). Inset: zoomed-in view of the DFG motif. D1 and D2 refer to distances between αC helix-M290-Cα and DFG-F382-Cζ, and between β3-K271-Cα and DFG-F382-Cζ, respectively. (b) Empirical partitioning of DFG states based on D1 and D2. (c) Aggregate 1235.39 μs simulation initiated from a DFG-in conformation (PDB: 2wgj, purple point) projected onto D1 and D2, showing confinement to the DFG-in basin.
Ultimately, kinase activity is determined by the relative populations of these aforementioned conformational states. (4; 17; 18; 19; 20). Changes in the active–inactive equilibrium are central to normal regulation, for example, through phosphorylation or ligand binding, but can also contribute to diseases. Indeed, several clinically observed kinase mutations bias the conformational ensemble towards the active state, resulting in dysregulated signalling and oncogenic activation (21; 22; 23). Conversely, kinase inhibitors exploit these conformational preferences through distinct binding modes (10; 24). Type I inhibitors bind competitively in the ATP pocket, whereas type II inhibitors extend into an adjacent hydrophobic back pocket that is typically exposed in the DFG-out inactive conformation. By stabilizing inactive states, type II inhibitors are broadly observed to achieve improved selectivity (25; 26). Allosteric inhibitors bind outside the ATP pocket and can stabilize distinct inactive or autoinhibited conformations, offering opportunities for high specificity and reduced resistance (27; 28). Thus, understanding how genetic variants, post-translational modifications, and ligand binding modulate kinase function requires a quantification of how these events and mutations modulate kinase structural ensembles (22; 29; 30; 31). Previous efforts have highlighted that some protein families follow common activation mechanisms, although it remains unclear if tyrosine kinase domains share a common activation ensemble (32). An ensemble-level characterization expands opportunities for structure-based drug design. The discovery of cryptic or transient states opens new opportunities for selective targeting to modulate kinase activity and potentially overcome resistance (33; 34; 35; 36; 37).
Despite their importance, kinase conformational ensembles remain challenging to characterize. Experimental techniques such as X-ray crystallography and cryo-EM provide high-resolution structural snapshots but typically capture static or averaged structures and commonly observe the most populated conformational states (38). Solution methods such as NMR and hydrogen–deuterium exchange can report on conformational dynamics but often provide limited atomistic detail for low-population intermediates (39). Molecular dynamics (MD) simulations provide a complementary route to probe kinase dynamics at atomic resolution, yet functionally important transitions, such as DFG flips and coupled A-loop or αC helix rearrangements, can occur on millisecond timescales and may remain inaccessible to conventional equilibrium simulations initiated from a single starting structure (15; 40; 41; 42; 43). Consistently, our own single-starting-point unbiased simulations struggle to sample the ensemble (Fig. 1c), showing that even 1.24 ms of unbiased simulations initiated from an experimentally determined DFG-in conformation remains trapped in this state. Enhanced-sampling methods can accelerate barrier crossing and discover rare states, but quantitative interpretation of the resulting ensembles can depend on the chosen collective variables, biasing protocol, and reweighting procedure (35; 44; 45; 46; 47). As a result, no singular protocol has observed multiple activation transitions across many kinase domains, and it remains unclear whether the activation pathway follows a conserved mechanism across kinases.
Markov state models (MSMs) provide a framework to address this challenge. By integrating information from multiple unbiased MD simulations, MSM analysis enables the identification of stable states and estimating transition timescales between them. Kinases are well-suited to this approach since many experimental structures spanning active and inactive states are available (12), providing seeds for sampling relevant conformational states (35; 48; 49; 50). However, for structurally under-represented kinases, experimental structures may not cover all relevant functional states. AI-based structure predictors offer a collective-variable-independent way to generate diverse structural seeds without first carrying out expensive exploratory simulations. In particular, AlphaFold2 (AF2) can be guided by template structures, enabling structural information from well-characterized homologues to be transferred onto a target sequence (51; 52; 53). For conserved protein families such as kinases and GPCRs, homologous structures annotated by conformational state can be used to generate physically plausible starting models in states that are poorly represented in the target’s experimental structural ensemble (43; 53; 54; 55). We refer to this strategy as transfer seeding, where homologue-derived conformational templates generate key target-specific seeds for subsequent unbiased MD and MSM analysis.
Here, we characterized the equilibrium conformational landscapes and activation pathways of three apo tyrosine kinase domains, ABL1 (which we refer to as ABL), EGFR, and MET, using MSMs constructed from unbiased MD simulations. We generated large aggregate unbiased simulation datasets comprising 0.592 ms for ABL, 1.836 ms for EGFR using PDB seeds, and 5.938 ms for MET using PDB and AF2-seeded structures on the Folding@home distributed computing platform (37) (Supplementary Fig. S6). By sampling each kinase with a consistent, unbiased simulation and MSM analysis protocol, our simulations enable direct comparison of their conformational ensembles. Although the three kinases sample similar active and inactive conformational states, we find that their equilibrium populations differ substantially. The three kinase activation pathways also proceed through distinct intermediate states and interconvert on timescales ranging from approximately 100 μs to 6 ms. We further find that the A-loops define kinase-specific conformational ensembles that shape the transition pathways. Our work establishes a systematic biophysical framework for comparing kinase activation landscapes, enabling mechanistic discovery of how conserved structural elements give rise to kinase-specific conformational dynamics.
2. Results
2.1. ABL and EGFR share common DFG conformational states but differ in populations and kinetics
We extensively sampled kinase conformational ensembles using a diverse set of PDB seeds from two well-represented kinase domains, ABL and EGFR (Supplementary Table S1, Fig. S4, S5). Projection onto the D1 and D2 distances showed broad coverage of the DFG-in, DFG-inter, and DFG-out regions for both ABL and EGFR kinase domains (Fig. 2a, Fig. S4, S5), from which we built kinase-specific MSMs (Fig. 2b). We identified six macrostates in the ABL and EGFR MSMs, respectively (see Methods). To relate these macrostates to experimentally observed kinase conformations, we labelled them with the kinome-wide DFG motif classification introduced by Modi and Dunbrack (14) (see Methods), thereby comparing metastable states derived from MSMs with structural classes observed across the kinome.
Figure 2: MSMs of ABL and EGFR explore similar DFG states but differ in population and landscapes.
(a) DFG-state classification of MD simulation frames based on D1 and D2 distances. Each MD frame is a points coloured by assignment to DFG-in (purple), DFG-inter (salmon), DFG-out (yellow), or undefined (gray). Initial seeding structures are marked as white circles. (b) ABL and EGFR free energy landscapes estimated from MSM equilibrium distribution projected onto tIC1 and tIC2. Macrostates are labelled with sizes proportional to their populations. (c) Pie charts show the decomposition of each macrostate into Dunbrack states. Labels report macrostate equilibrium population and Dunbrack state fractions.
Our MSMs show that ABL and EGFR access a shared set of conformational states, but assign different equilibrium populations to these shared states, which include the active-like DFG-in BLAminus state, the SRC-like inactive DFG-in BLBplus state (16; 56), the DFG-out BBAminus state, and the DFG-inter BABtrans state (Fig. 2c). In their respective six-state coarse-grained MSMs, most macrostates are dominated by a single Dunbrack class (Fig. 2c), indicating that the kinome-wide structural nomenclature captures a substantial part of the slow conformational heterogeneity resolved by the simulations. In apo ABL, the equilibrium ensemble is dominated by macrostate 6 (49.4%), which is primarily composed of DFG-out BBAminus conformations and resembles the inactive arrangement stabilized by type II inhibitors. The second most populated ABL macrostate, state 5 (24.2%), corresponds to the DFG-in BLAminus ensemble, which is characteristic of catalytically active kinase structures (17). The EGFR equilibrium ensemble is dominated by DFG-in conformations, with state 6 (51.9%) corresponding primarily to the BLBplus ensemble. This state adopts a SRC-like inactive conformation, in which the DFG-D remains oriented towards the ATP-binding site but the regulatory β3-Lys–αC-Glu salt bridge is disrupted and the A-loop is semi-closed compared with the active catalytic geometry (Supplementary Fig. S16). The second most populated EGFR macrostate 5 (39.9%) corresponds to the active-like BLAminus ensemble. Ensemble images of each macrostate are provided in Supplementary Figs. S8 and S9.
The correspondence between DFG structural labels and MSM macrostates is strong but not absolute. Specifically, we observe that kinetic proximity between macrostates cannot be directly inferred by DFG labels alone. In other words, the DFG labels are insufficient to predict what state any kinase domain is most likely to transition into, further highlighting the need for considering the complete ensemble. We quantified the interconversion timescales between macrostates using mean first-passage times (MFPTs), defined as the average time required for trajectories starting in one state to reach a target state for the first time (Fig. S13). In ABL, macrostate 6 contains both a DFG-out BBAminus substate and a DFG-in BLBplus substate. These two substates interconvert rapidly, with an MFPT of 68.2±0.20 μs, whereas the transition from the BLBplus-containing macrostate 6 to the active-like BLAminus macrostate 5 occurs on a much slower timescale, with an MFPT of 2.39±0.20 ms (Supplementary Fig. S14). Therefore, the BBAminus-to-BLBplus transition represents a local, rapidly reversible rearrangement within this inactive basin. Similarly, ABL macrostate 4 contains kinetically mixed DFG-out BBAminus and DFG-inter BABtrans conformations (Supplementary Fig. S15). These conformations share an open A-loop but differ mainly through the dihedrals of DFG motifs and intermittently formed β3-Lys–αC-Glu salt bridge. We found this macrostate to resemble A-loop-closed structures bound to type I inhibitors in PDB (57; 58; 59). In summary, our MSM analysis shows that ABL and EGFR occupy a common DFG-centred conformational space, but differ substantially in how this space is thermodynamically populated and kinetically connected.
2.2. Transfer seeding expands coverage of the MET kinase DFG-state ensemble
MET is less represented in the experimental structural record (Fig. 3c, S17). Indeed, there are no MET DFG-inter structures available in the PDB, leaving a gap between the experimentally observed DFG-in and DFG-out regions (Fig. 3a, S17). Consistent with the rarity of DFG transitions on conventional MD timescales, simulations initiated only from this limited PDB ensemble are unlikely to identify missing intermediate states de novo.
Figure 3: Complete sampling and MSM construction for MET kinase are enabled by transfer seeding.

(a) Number of MET PDB structures (solid) and AF2-derived seeds (hatched) used for simulations. (b) Simulation snapshots projected onto D1 and D2 distances, coloured by DFG-state assignment as in Fig. 2. Initial seeding structures are marked to indicate PDB seeds (white circles) and AF-derived seeds (white triangles). (c) MET free energy landscape estimated from MSM equilibrium distribution projected onto tIC1 and tIC2. Macrostates are labelled with sizes proportional to their populations. (d) Pie charts show the decomposition of each macrostate into Dunbrack states. Labels report macrostate equilibrium population and Dunbrack state fractions.
We first tested whether missing MET DFG states could be recovered from the DFG-in basin using enhanced sampling approaches, including accelerated MD, Gaussian accelerated MD, umbrella sampling, metadynamics, and adaptive sampling (see Methods). Several of these approaches recovered DFG-inter and DFG-out conformations, indicating that the experimentally missing DFG-inter region is physically accessible. However, the apparent free-energy landscapes differed substantially between protocols, particularly in the relative depths of DFG basins and transition regions (Supplementary Fig. S1). We attribute this to a combination of biasing and reweighting uncertainty, incomplete barrier crossing, and limitations of low-dimensional collective variables. Meanwhile, some conformational changes sampled in unbiased MD involve coupled rearrangements and are not cleanly described as motion along D1 and D2 alone. Therefore, these biased simulations, without an in-depth consideration of multiple viable collective variables (54), are mainly for state-discovery, but are insufficient for quantitative estimates of the equilibrium ensemble. Details of enhanced sampling results can be found in the Supplementary Information.
To systematically generate a MET ensemble in a similar manner to ABL and EGFR without biasing physics, we built upon the principle of adaptive seeding where many parallel simulations are initialized from diverse conformations (35). Rather than relying on enhanced sampling to obtain seeds, we provided AF2 with annotated kinase structures from KinCore as conformational templates for MET modelling (see Methods, Supplementary Fig. S11). This template supervision was essential because untargeted AF2 modelling alone could not recover DFG states other than DFG-in (Supplementary Fig. S3). Notably, AF2 is most reliable at generating the desired templated conformation relative to other available co-folding methods (Supplementary Fig. S12).
We refer to this supervised structural generation as transfer seeding, in which structural information from ABL and EGFR is transferred onto the MET sequence with modern co-folding methods. Our transfer seeding approach successfully generates MET kinase structures in Dunbrack states not represented in the MET PDB ensemble and not obtainable by AF2 alone (Fig. 3a). To minimally augment the seeding of the MET kinase landscape without overly biasing the starting sampling away from experimental conformations, we provide only two structures per Dunbrack state (Fig. 3a).
Subsequent simulations and MSM analysis show that the absence of DFG-inter conformations in MET reflects incomplete experimental coverage rather than an inaccessible region of the MET conformational landscape. The union of PDB- and AF2-seeded simulations populated a connected DFG-in, DFG-inter, and DFG-out landscape (Fig. 3b). The resulting MET MSM shows that the DFG-inter conformations form part of this connected equilibrium conformational landscape with equilibrium population in states 1, 3, and 5 (Fig. 3c,d).
As with ABL and EGFR, Dunbrack state labels alone do not fully describe the MET conformational landscape because existing PDB structures may not capture the conformational heterogeneity of the A-loop. In the MET MSM, macrostates dominated by the same DFG label are not necessarily structurally equivalent (Fig. 3c,d). This is most apparent within the DFG-in ensemble, which separates into three distinct kinetic basins. State 6 (6.5%) corresponds to the active-like BLAminus ensemble, whereas states 9 (39.9%) and 8 (23.7%) both adopt BLBplus conformations but differ substantially in their A-loop ensembles (Supplementary Fig. S10). Similarly, DFG-out macrostates are all labelled as BBAminus but display distinct A-loop arrangements (Supplementary Fig. S10). Our results indicate that MET conformational heterogeneity extends beyond the DFG motif, with the A-loop contributing substantially to the separation of major kinetic basins within each DFG class.
MFPTs confirm that the structurally distinct macrostates 8 and 9 are kinetically separate from one another despite sharing the same Dunbrack state label (Supplementary Fig. S13). These states interconvert rapidly, with MFPTs of 12.0±0.6 μs and 19.0±0.7 μs in the forward and reverse directions, respectively, corresponding structurally to fast local A-loop exchange within the dominant DFG-in ensemble (Fig. S10). In contrast, transitions from states 8 and 9 to the active-like BLAminus state 6 occur on a much slower millisecond timescale, whereas escape from state 6 back to states 8 and 9 is approximately an order of magnitude faster. These observations indicate that the active-like MET basin is kinetically disfavoured in the apo form. However, we note that our understanding of these kinetic barriers is limited to the apo form and that the presence of cofactors and substrates may alter this transition probability.
2.3. Activation transitions follow kinase-specific pathways
While we are simulating the apo form, previous simulation efforts highlight that the MD is capable of capturing key conformational intermediates in activation pathways present across thermal equilibrium (35; 36; 43; 60). We extended these approaches to compute and compare activation pathways and intermediates for ABL, EGFR, and MET kinase domains. To capture activation pathways for ABL, EGFR, and MET, we use using transition path theory (61) by decomposing reactive trajectories from DFG-out (BBAminus) to DFG-in (BLAminus) into reactive fluxes on MSM microstates (Fig. 4). For each kinase domain, we report the highest-flux pathway and define its rate-limiting step as the transition with the smallest flux. Transition path ensembles are provided in Supplementary Fig. S21. To systematically identify structural changes accompanying the rate-limiting step, we trained Random Forest classifiers to distinguish pre- versus post-rate-limiting conformations and extracted the most discriminating structural features (see Methods).
Figure 4: Rate-limiting steps in the DFG-out to DFG-in transition.
Analysis for ABL (a), EGFR (b), and MET (c). Left: mean and standard deviation of top-ranked features from the random forest classifier along the transition pathway. The rate-limiting step is indicated by black arrows; pre- and post-rate-limiting states are shown as blue and red markers respectively. Right: superposition of pre-rate-limiting (blue) and post-rate-limiting (red) conformations, overlaid on the full ensemble (grey). Insets highlight key regulatory elements.
The ABL highest-flux pathway contains two steps, with the BBAminus-to-BLBplus transition being the flux bottleneck (Fig. 4a). The rearrangement is initiated by αC helix swinging out, disrupting the β3-Lys–αC-Glu salt bridge. This opens the back pocket and creates space for the A-loop to rotate towards a SRC-like orientation, an inactive DFG-in conformation, with the two-turn helix packed against the αC helix. Meanwhile, the G-loop transitions from a short helical structure to a more extended turn towards the ATP binding site. The subsequent step transitions from the SRC-like ensemble to the BLAminus ensemble, where the A-loop rotates further to the open conformation, the αC helix returns inward to re-form the salt bridge, and the G-loop returns to a kinked conformation.
The EGFR highest-flux pathway reveals a distinct multi-step transition (Fig. 4b, Supplementary Fig. S21b). It initiates with rearrangement of the X-DFG residue, converting from the DFG-out BBAminus to the inactive DFG-in intermediate ABAminus state. The subsequent ABAminus-to-BLAminus transition is the rate-limiting step, characterized by XDFG motif backbone dihedral reorganization accompanied by subtle A-loop rearrangement, while the A-loop remains open throughout. Minimal structural changes are observed in the G-loop or αC helix in the rate-limiting step. Notably, the SRC-like BLBplus ensemble does not participate in the top net-flux pathways, indicating that it forms an off-pathway basin that mostly interconverts with the active-like BLAminus ensemble.
In MET, the highest-flux pathway proceeds from the BBAminus ensemble (state 7) through two BABtrans intermediates (states 1 and 3) before reaching the BLAminus state 6 (Fig. 4c, Supplementary Fig. S21c). The A-loop RMSD to the active state decreases progressively from 21.9 Å (state 7) to 11.4 Å (state 1) and to 3.6 Å (state 3). The rate-limiting transition from state 7 to state 1 is associated with the initial opening of the A-loop, which forms helical elements interacting with the αC helix while the DFG motif adopts BABtrans. This is accompanied by β3-Lys–αC-Glu salt bridge formation and G-loop extension into the space released by A-loop opening. The subsequent transition from state 1 to state 3 involves further opening of the A-loop and dissociation of the helical elements, after which the final transition from state 3 to state 6 completes the BABtrans-to-BLAminus rearrangement. Notably, the path ensemble also shows lower-flux branches connecting both states 1 and 3 directly to the sink state 6, indicating that although the route through states 7, 1, 3, and 6 is dominant, a more direct transition from state 1 to state 6 is also accessible (Supplementary Fig. S21c).
2.4. State-dependent activation loop exposure identifies phosphorylation-competent intermediates
The A-loop is a dynamic regulatory element in kinases. Its flexibility allows it to control substrate access via conformational change, but also causes the A-loop to often remain unresolved or partially resolved in experimental crystal structures (62). A-loop phosphorylation provides a common regulatory switch by stabilizing catalytically competent conformations (63), although the conformational state in which phosphorylation sites, sometimes called phosphosites, become accessible can differ between kinases. To examine coupling between phosphorylation site accessibility and A-loop dynamics in our MSMs, we characterized the A-loop ensemble of each macrostate using per-residue secondary-structure labels assigned by DSSP (see Methods) and relative solvent-accessible surface area (rSASA), normalized to a range of 0 to 1 per residue.
In the active-like BLAminus ensembles, the phosphosites ABL Tyr393 (63) and MET Tyr1235 (64) show related local packing. DSSP analysis indicates that these sites are embedded in locally structured segments with β strands and turns, rather than in fully disordered loops (Supplementary Figs. S18 and S20). In ABL and MET, phosphosite side chains are partially shielded by neighbouring bulky or aliphatic residues, resulting in low solvent exposure in the active-like BLAminus ensemble (Fig. 5a,c).
Figure 5: A-loop phosphorylation-site exposure is kinase specific, with Tyr393 (ABL) and Tyr1235 (MET) being more accessible in intermediate ensembles than in the active-like ensemble.
Analysis for ABL (a), EGFR (b), and MET (c). rSASA profiles (left column) and representative A-loop conformations (right column) are shown for the active-like ensemble (colored purple, right) and, where active-state (purple) phosphosites are buried, a more exposed intermediate ensemble (colored orange, right). Lines show mean rSASA and shaded regions indicate standard deviation for the sampled conformations; horizontal dashed lines mark rSASA values of 0.1 and 0.5 to indicate complete burial or complete solvent exposure; vertical dashed lines mark phosphosite residues. Phosphosite residues are shown as pink surfaces in the structure. Surrounding residues are shown in grey surfaces.
This state-dependent exposure suggests that the final apo active-like ensemble may not be the main phosphorylation-competent conformation. For MET, this interpretation is consistent with experimental evidence that A-loop phosphorylation is sequential, with Tyr1235 phosphorylated before Tyr1234 (65; 66). In the highest-flux pathway, MET activation proceeds from the BBAminus inactive state through BABtrans state 1 before reaching the active-like BLAminus state 6. In state 1, Tyr1235 is more exposed (rSASA 0.44±0.09), whereas Tyr1234 remains largely buried (rSASA 0.10±0.09). Tyr1234 becomes more exposed only after progression towards the active-like state 6 (rSASA 0.35 ± 0.06). These observations suggest that the BABtrans ensemble may represent a pre-active, phosphorylation-competent intermediate for the first MET A-loop phosphorylation event at Tyr1235, which promotes subsequent phosphorylation at Tyr1234 and A-loop rearrangement towards the active-like ensemble.
The same principle of state-dependent phosphosite accessibility might apply to ABL. Tyr393 is less solvent-exposed in the active-like BLAminus state 5 (rSASA 0.07±0.04) than in the BLBplus ensemble in state 6 (rSASA 0.61±0.09), where Tyr393 points towards solvent (Fig. 5a). This ensemble is also involved in the highest-flux pathway as the post-rate-limiting step. We therefore hypothesize that ABL Tyr393 phosphorylation is preferentially enabled by a BLBplus-like pre-active intermediate that presents Tyr393 in a more solvent-accessible geometry, rather than by the final apo active-like BLAminus state.
EGFR presents a contrasting picture. In the active-like BLAminus ensemble, the A-loop remains more open and predominantly coil-like around Tyr869. The Tyr869 side chain points outward towards solvent and is therefore highly exposed (Fig. 5b). This is consistent with EGFR’s non-canonical regulatory mechanism, in which catalytic activation is primarily mediated by asymmetric kinase-domain dimerisation rather than by a required A-loop phosphorylation switch (67; 68). In this context, Tyr869 exposure may reflect regulatory accessibility for cross-talk phosphorylation rather than a prerequisite step for forming the active catalytic state.
Given these observations from pre-phosphorylation apo ensembles, we hypothesize that A-loop phosphorylation sites are exposed in kinase-specific conformational windows along the activation pathway such that phosphorylation may select or stabilize transient pre-active ensembles rather than marking the final active state. This hypothesis may be tested by combining conformation-selective perturbations with HDX-MS (69), NMR (70; 71), or time-resolved phosphosite mass spectrometry (72) to determine whether phosphosite exposure and phosphorylation order track enrichment of the predicted intermediates.
3. Conclusions
Our simulations and MSMs show that kinase activation does not proceed through a single conserved mechanism shared across tyrosine kinases, but through kinase-specific transitions embedded within broader regulatory ensembles. Although ABL, EGFR and MET can access DFG-in, DFG-inter and DFG-out geometries, their equilibrium populations, kinetic connectivity and dominant activation routes differ substantially. Common structural labels can mask such distinct dynamical mechanisms. While the Dunbrack nomenclature provides an interpretable framework for comparing kinase conformations, macrostates with similar DFG assignments can differ in kinetic connectivity. The A-loop is also a major determinant of kinase-specific behaviour, shaping both transition pathways and phosphorylation-site exposure. Together, our results support an ensemble view of kinase activation in which functional states are defined not only by conserved motif geometry, but also by how regulatory elements are thermodynamically populated and kinetically connected.
This mechanistic heterogeneity has direct implications for interpreting kinase regulation and for conformation-selective inhibitor design. The accessibility of DFG-out states, the stability of inactive intermediates, and the coupling between the DFG motif, αC helix, G-loop, and A-loop differ between kinases. These features suggest that inhibitor selectivity is shaped by the full conformational network through which active and inactive states are populated and exchanged, rather than by the presence of a nominal DFG-in or DFG-out structure alone. For example, the SRC-like ABL intermediate and the semi-open MET activation-loop ensemble illustrate how transient or sparsely populated states may create kinase-specific opportunities for stabilizing inactive conformations or redirecting activation pathways.
Finally, our work demonstrates how transfer seeding offers a practical route to generate kinase ensembles and build MSMs under a consistent and systematic protocol, even when available experimental structures incompletely cover the relevant conformational landscape. For MET, homologue-guided AF2 models expanded the starting ensemble beyond the available experimentally determined structures and enabled sampling of connected DFG-inter and DFG-out regions. This strategy complements adaptive and enhanced-sampling approaches by using conformational information already present in related family members to initialise unbiased simulations. The present models are limited to isolated apo KDs, while ligand binding, phosphorylation, mutations and regulatory domains may reshape both populations and pathways (29; 30; 31; 73; 74). Ultimately, we envision that extending this framework to perturbed ensembles and linking functional insights with ensemble and kinetic descriptions will open new avenues for the design of next-generation kinase inhibitors with improved selectivity and affinity.
4. Methods
4.1. Structural seed generation
For each target kinase domain (KD) we assembled a pool of starting conformations (“seeds”) intended to span the DFG-in, DFG-inter, and DFG-out regions of the Modi–Dunbrack D1 vs. D2 landscape (14). Seeds were drawn from two sources: (i) experimentally determined structures available in the Protein Data Bank (PDB-based seeding), and (ii) AlphaFold2 (AF2) models generated with ColabFold using curated template sets biased towards specific Dunbrack substates (AF2-based, or “transfer seeding”). All seeds were processed through an identical structure-preparation and solvation pipeline so that downstream simulations differed only in their starting coordinates. Scripts for seed generation are available at https://github.com/sukritsingh/kinase-seeding.
4.2. PDB-based seed generation
Apo KD seeds for ABL1 (KLIFS kinase ID 392) and EGFR (KLIFS kinase ID 406) were generated from every structure available for the corresponding kinase in KLIFS (75). For each kinase, the full set of PDB entries, chains, and alternate models was queried with the opencadd KLIFS remote interface and filtered to retain only entries with reported resolution. Each (PDB ID, chain, alternate location) tuple was then passed as a PDBProtein component to the KinoML OEKLIFSKinaseApoFeaturizer (76), which uses the OpenEye Spruce toolkit with the RCSB loop database (rcsb_spruce.loop_db) to strip ligands, waters, and cofactors, model missing loops, repair missing side chains, assign protonation states at pH 7.4, add caps, and return an apo KD. Because the individual structures cover different residue ranges, all apo models for a given kinase were cropped to the common residue window (maximum of the per-structure minimum residue, minimum of the per-structure maximum residue), and any structure with internal gaps in that window was discarded; the retained structures were then re-capped with ACE/NME and residues were renumbered using assign_caps and update_residue_identifiers so that all seeds for a given kinase shared an identical topology. The resulting per-PDB apo structures constitute the PDB-based seeds: 83 seeds for ABL1 and 400 seeds for EGFR.
4.3. Transfer seeding of MET via AF2
For MET (KLIFS kinase ID 446, UniProt P08581), the PDB ensemble is dominated by DFG-in conformations. To augment this ensemble, we used homologue-guided AF2 modelling (“transfer seeding”) (51; 52; 77) with MSA subsampling and stochastic dropout enabled. MET KD sequences were modelled using (i) the default ColabFold MSA server (MMseqs2), (ii) stochastic dropout enabled at inference, (iii) num_recycles = 3, and (iv) two models per prediction, to generate a minimal set of AF2 seeds of MET kinase.
Custom template sets derived from the KinCore database (16), sorted by Dunbrack state labels, were used to bias each prediction towards a specific DFG state. The sorted template library is deposited at https://osf.io/spu2a/, and was created by downloading and re-sorting all structures into a directory tree keyed by those labels (one sub-directory per spatial DFG state, and within each, one sub-directory per Dunbrack substate). To bias AF2 towards a chosen state for MET, the templates directory was pointed at the corresponding labelled sub-directory, so that only structures carrying that Dunbrack label were available as templates. Separate AF2 runs were executed for each Dunbrack state template pool. As a minimal set to supplement existing PDB structures, two models from each Dunbrack state were created, giving a total of 16 AF2-based seeds, and were pooled with the native MET PDB structures. This pooled set of 117 PDB-based structures and 16 AF2-seed structures, a total of 133 structures, was then used as the initial MET kinase simulation seeds. The combined set of 133 seeds was processed through the same KinoML apo-kinase featurizer and common-residue cropping described above to yield a single, topology-consistent set of MET seeds. A sample demonstration of using the KinCore-sorted template library to do transfer seeding is available as a Colab notebook at https://colab.research.google.com/drive/1Dcb5ZouYJQfyq1np-TBmKJE8DOAbA9dU.
4.4. Molecular dynamics simulations
All seeds (PDB-based for ABL1/EGFR; PDB- and AF2-based for MET) were parameterized and equilibrated with an identical OpenMM 8.0 protocol (78), using AMBER ff14SB for protein (79) and TIP3P for water. Each apo KD was placed in a cubic periodic box with solvent padding on all sides (minimum 1.2 nm) using Modeller.addSolvent, neutralised, and brought to 0.15 M NaCl. To ensure consistent solvation NPT ensembles across conformations, and that conformational differences did not alter the solvent padding, the number of solvent molecules was held consistent at the maximum number of solvent molecules across all conformations. In other words, the largest possible number of solvent molecules designated by the padding cutoff was applied to all conformations. Bonds to hydrogen were constrained, water was held rigid, and hydrogen-mass repartitioning (HMR, hydrogen mass = 4.0 amu) was applied so that a 4 fs Langevin timestep could be used. Nonbonded interactions were treated with PME at a 1.0 nm real-space cutoff. Systems were energy-minimized, then assigned initial velocities from a Maxwell–Boltzmann distribution at 310 K (setVelocitiesToTemperature(310)) and propagated for 5 ns in the NPT ensemble (310 K, 1 atm) with a Monte Carlo Barostat and the Langevin Integrator with a BAOAB-like splitting (collision rate 1.0 ps−1, constraint tolerance 10−5). These equilibrated structures were launched in full on the Folding@home platform as production runs (37). Production runs were launched with 1000 NPT trajectories per starting structure (“per seed”) using OpenMM 8.0.0 with CUDA mixed precision. Integrator settings, thermostat, barostat, force field, and timestep were identical to equilibration. Example simulation scripts are provided at https://github.com/sukritsingh/distributed-sampling.
4.5. DFG conformational state assignment
Simulation snapshots of ABL, EGFR, and MET were assigned to DFG-in, DFG-inter, or DFG-out spatial classes by clustering the D1 and D2 distance space using HDBSCAN (80). Each snapshot was then further annotated using the Modi–Dunbrack DFG conformational nomenclature (14). In this scheme, states are labelled as BLAminus, BLAplus, ABAminus, BLBminus, BLBplus, BLBtrans, BABtrans, or BBAminus. The first three letters denote the Ramachandran basins occupied by the backbone dihedrals of the X-DFG, DFG-Asp, and DFG-Phe residues, where X-DFG is the residue immediately preceding the DFG motif. The suffix minus, plus, or trans denotes the rotamer of DFG-Phe. For each snapshot, a seven-dimensional torsion vector,
was compared with the reference centroid of each Dunbrack state using the periodic cosine distance
| (1) |
Snapshots were assigned to the Dunbrack state with the smallest distance. Snapshots for which all state-centroid distances exceeded 1 were classified as noise.
4.6. Markov state model construction
Trajectories shorter than 100 ns were discarded. Structural features (see Supplementary Table S2) were extracted, subjected to tICA (lag 100 ns, top 20 components), and discretized into 1000 microstates by k-means clustering. Transition counts were computed using effective counts with pseudo-count regularization. Reversible maximum-likelihood MSMs and Bayesian MSMs were estimated using deeptime (81). The choice of MSM parameters was validated by the implied timescale convergence (Supplementary Fig. S7) and kept consistent for ABL, EGFR, and MET. Macrostates were coarse-grained by PCCA+. Reactive pathways were computed using transition path theory (82). Analysis and plotting scripts are available at https://github.com/meyresearch/kinase_analysis.
4.7. Enhanced sampling simulations
Accelerated MD (83), Gaussian accelerated MD (84), umbrella sampling (20×20 grid on D1/D2), adaptive sampling (85), and metadynamics (86) were deployed on Folding@home and HPC resources targeting MET KD. Reweighting was performed by Miao cumulant expansion (aMD, GaMD) (87), MBAR (umbrella sampling) (88), or summation of Gaussians (in metadynamics) (89). Full details are provided in the Supplementary Methods.
Supplementary Material
6. Acknowledgements
We thank the citizen-scientists of Folding@home for donating their computing resources that enabled the distributed simulations (projects 16497–16499, 17601–17605, 17645, 17646, 17649, 17650). This work used resources from the High-Performance Computing Group at Memorial Sloan Kettering Cancer Center. The authors are grateful to the MSKCC DigITs and HPC team, especially Jamie Cheong, Lohit Valleru, and Monica Chakradeo for their assistance with high-performance computing resources. S.S. is a Damon Runyon Quantitative Biology Fellow from the Damon Runyon Cancer Research Foundation (DRQ-14–22) and acknowledges support from an NCI Pathway to Independence Award for Outstanding Early-Stage Postdoctoral Researchers (NCI K99 CA286801). The authors are thankful to John Chodera and Markus Seeliger for their helpful thoughts and feedback.
Footnotes
8 Competing interests
The authors declare no competing interests.
9 Disclaimer
The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
5. Data availability
All analysis scripts and plotting scripts for MSM construction and analysis are available online at https://github.com/meyresearch/kinase_analysis. MSMs and related models are deposited on Zenodo at https://doi.org/10.5281/zenodo.20544389. The KinCore-sorted template library is hosted at https://osf.io/spu2a/. All scripts for setup and equilibration of seed structures, as well as the trimmed seed structures, are available at https://github.com/sukritsingh/kinase-seeding. Simulation setup and distributed-sampling scripts are available online at https://github.com/sukritsingh/distributed-sampling, including scripts describing simulation parameters for unbiased simulations. Scripts for the framework of adaptive sampling on Folding@home are available at https://github.com/sukritsingh/fahdaptive. The raw MD simulation trajectories for ABL, EGFR and MET comprise approximately 0.5, 1.5 and 2.3 TB of data, respectively. Owing to the size of these datasets and the lack of online storage capacity, the raw trajectories are not deposited in a public repository but are freely available upon any request.
References
- [1].Piovesan A. et al. Human protein-coding genes and gene feature statistics in 2019. BMC Research Notes 12, 315 (2019). URL 10.1186/s13104-019-4343-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [2].Trembley J. H. et al. Protein kinase CK2 – diverse roles in cancer cell biology and therapeutic promise. Molecular and Cellular Biochemistry 478, 899–926 (2023). URL 10.1007/s11010-022-04558-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [3].Manning B. D. & Toker A. AKT/PKB Signaling: Navigating the Network. Cell 169, 381–405 (2017). URL https://www.sciencedirect.com/science/article/pii/S0092867417304130. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [4].Huse M. & Kuriyan J. The Conformational Plasticity of Protein Kinases. Cell 109, 275–282 (2002). URL https://www.sciencedirect.com/science/article/pii/S0092867402007419. [DOI] [PubMed] [Google Scholar]
- [5].Taylor S. S. & Kornev A. P. Protein kinases: evolution of dynamic regulatory proteins. Trends in Biochemical Sciences 36, 65–77 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [6].Hubbard S. R. & Miller W. T. Receptor tyrosine kinases: mechanisms of activation and signaling. Current opinion in cell biology 19, 117–123 (2007). URL https://pmc.ncbi.nlm.nih.gov/articles/PMC2536775/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [7].Cicenas J., Zalyte E., Bairoch A. & Gaudet P. Kinases and Cancer. Cancers 10, 63 (2018). URL https://pmc.ncbi.nlm.nih.gov/articles/PMC5876638/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [8].Tomuleasa C. et al. Therapeutic advances of targeting receptor tyrosine kinases in cancer. Signal Transduction and Targeted Therapy 9, 201 (2024). URL https://www.nature.com/articles/s41392-024-01899-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [9].Roskoski R. Properties of FDA-approved small molecule protein kinase inhibitors: A 2025 update. Pharmacological Research 216, 107723 (2025). URL https://www.sciencedirect.com/science/article/pii/S1043661825001483. [DOI] [PubMed] [Google Scholar]
- [10].Cohen P., Cross D. & Jänne P. A. Kinase drug discovery 20 years after imatinib: progress and future directions. Nature Reviews Drug Discovery 20, 551–569 (2021). URL https://www.nature.com/articles/s41573-021-00195-4. Number: 7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [11].Arter C., Trask L., Ward S., Yeoh S. & Bayliss R. Structural features of the protein kinase domain and targeted binding by small-molecule inhibitors. The Journal of Biological Chemistry 298, 102247 (2022). URL https://pmc.ncbi.nlm.nih.gov/articles/PMC9382423/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [12].Möbitz H. The ABC of protein kinase conformations. Biochimica et Biophysica Acta (BBA) - Proteins and Proteomics 1854, 1555–1566 (2015). URL https://www.sciencedirect.com/science/article/pii/S1570963915000904. [DOI] [PubMed] [Google Scholar]
- [13].Anderson B. et al. How many kinases are druggable? A review of our current understanding. Biochemical Journal 480, 1331–1363 (2023). URL 10.1042/BCJ20220217. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [14].Modi V. & Dunbrack R. L. Defining a new nomenclature for the structures of active and inactive kinases. Proceedings of the National Academy of Sciences 116, 6818–6827 (2019). URL https://www.pnas.org/doi/full/10.1073/pnas.1814279116. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [15].Meng Y., Lin Y. l. & Roux B. Computational Study of the “DFG-Flip” Conformational Transition in c-Abl and c-Src Tyrosine Kinases. The Journal of Physical Chemistry. B 119, 1443–1456 (2015). URL https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4315421/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [16].Modi V. & Dunbrack R. L. Jr. Kincore: a web resource for structural classification of protein kinases and their inhibitors. Nucleic Acids Research 50, D654–D664 (2022). URL 10.1093/nar/gkab920. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [17].Faezov B. & Dunbrack R. L. AlphaFold2 models of the active form of all 437 catalytically competent human protein kinase domains (2023). URL https://www.biorxiv.org/content/10.1101/2023.07.21.550125v2. Pages: 2023.07.21.550125 Section: New Results. [Google Scholar]
- [18].Azam M., Seeliger M. A., Gray N. S., Kuriyan J. & Daley G. Q. Activation of tyrosine kinases by mutation of the gatekeeper threonine. Nature Structural & Molecular Biology 15, 1109–1118 (2008). URL https://www.nature.com/articles/nsmb.1486. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [19].Shan Y. et al. A conserved protonation-dependent switch controls drug binding in the Abl kinase. Proceedings of the National Academy of Sciences 106, 139–144 (2009). URL https://www.pnas.org/doi/abs/10.1073/pnas.0811223106. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [20].Dar A. C. & Shokat K. M. The evolution of protein kinase inhibitors from antagonists to agonists of cellular signaling. Annual Review of Biochemistry 80, 769–795 (2011). [DOI] [PubMed] [Google Scholar]
- [21].Nussinov R., Tsai C.-J. & Jang H. A New View of Activating Mutations in Cancer. Cancer Research 82, 4114–4123 (2022). URL 10.1158/0008-5472.CAN-22-2125. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [22].Sutto L. & Gervasio F. L. Effects of oncogenic mutations on the conformational free-energy landscape of EGFR kinase. Proceedings of the National Academy of Sciences 110, 10616–10621 (2013). URL https://www.pnas.org/doi/10.1073/pnas.1221953110. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [23].Chan W. W. et al. Conformational control inhibition of the BCR-ABL1 tyrosine kinase, including the gatekeeper T315I mutant, by the switch-control inhibitor DCC-2036. Cancer cell 19, 556–568 (2011). URL https://pmc.ncbi.nlm.nih.gov/articles/PMC3077923/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [24].Roskoski R. Classification of small molecule protein kinase inhibitors based upon the structures of their drug-enzyme complexes. Pharmacological Research 103, 26–48 (2016). URL https://www.sciencedirect.com/science/article/pii/S1043661815301298. [DOI] [PubMed] [Google Scholar]
- [25].Vijayan R. S. K. et al. Conformational Analysis of the DFG-Out Kinase Motif and Biochemical Profiling of Structurally Validated Type II Inhibitors. Journal of Medicinal Chemistry 58, 466–479 (2015). URL https://pmc.ncbi.nlm.nih.gov/articles/PMC4326797/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [26].Wang X., DeFilippis R. A., Yan W., Shah N. P. & Li H.-y. Overcoming Secondary Mutations of Type II Kinase Inhibitors. Journal of Medicinal Chemistry 67, 9776–9788 (2024). URL 10.1021/acs.jmedchem.3c01629. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [27].Mingione V. R., Paung Y., Outhwaite I. R. & Seeliger M. A. Allosteric regulation and inhibition of protein kinases. Biochemical Society transactions 51, 373–385 (2023). URL https://pmc.ncbi.nlm.nih.gov/articles/PMC10089111/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [28].Wang B. et al. An overview of kinase downregulators and recent advances in discovery approaches. Signal Transduction and Targeted Therapy 6, 423 (2021). URL https://www.nature.com/articles/s41392-021-00826-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [29].Pegram L. M. et al. Activation loop dynamics are controlled by conformation-selective inhibitors of ERK2. Proceedings of the National Academy of Sciences 116, 15463–15468 (2019). URL https://www.pnas.org/doi/10.1073/pnas.1906824116. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [30].Palhano Zanela T. M., Woudenberg A., Romero Bello K. G. & Underbakke E. S. Activation loop phosphorylation tunes conformational dynamics underlying Pyk2 tyrosine kinase activation. Structure 31, 447–454.e5 (2023). Place: London, England : 1993. [DOI] [PubMed] [Google Scholar]
- [31].Silva G. M. d., Lam K., Dalgarno D. C. & Rubenstein B. M. Compound Mutations in the Abl1 Kinase Cause Inhibitor Resistance by Shifting DFG Flip Mechanisms and Relative State Populations (2024). URL https://www.biorxiv.org/content/10.1101/2024.05.23.595569v1. Pages: 2024.05.23.595569 Section: New Results. [Google Scholar]
- [32].Vithani N. et al. G protein activation occurs via a largely universal mechanism. The Journal of Physical Chemistry B 128, 3554–3562 (2024). URL https://doi.org/10.1021/acs.jpcb.3c07028. 10.1021/acs.jpcb.3c07028. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [33].Nussinov R., Yavuz B. R. & Jang H. Allostery in Disease: Anticancer Drugs, Pockets, and the Tumor Heterogeneity Challenge. Journal of Molecular Biology 437, 169050 (2025). URL https://www.sciencedirect.com/science/article/pii/S0022283625001160. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [34].Seeliger M. A. et al. Equally Potent Inhibition of c-Src and Abl by Compounds that Recognize Inactive Kinase Conformations. Cancer Research 69, 2384–2392 (2009). URL 10.1158/0008-5472.CAN-08-3953. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [35].Sun X., Singh S., Blumer K. J. & Bowman G. R. Simulation of spontaneous G protein activation reveals a new intermediate driving GDP unbinding. eLife 7, e38465 (2018). URL 10.7554/eLife.38465. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [36].Bhakat S., Vats S., Mardt A. & Degterev A. Generalizable Protein Dynamics in Serine-Threonine Kinases: Physics is the key (2025). URL https://www.biorxiv.org/content/10.1101/2025.03.06.641878v1. Pages: 2025.03.06.641878 Section: New Results. [Google Scholar]
- [37].Zimmerman M. I. et al. SARS-CoV-2 simulations go exascale to predict dramatic spike opening and cryptic pockets across the proteome. Nature Chemistry 1–9 (2021). URL https://www.nature.com/articles/s41557-021-00707-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [38].Camille S. et al. Structural determination of small proteins by cryo-EM using a coiled coil module strategy. Scientific Reports 15, 36800 (2025). URL https://www.nature.com/articles/s41598-025-20608-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [39].Kay L. E. New Views of Functionally Dynamic Proteins by Solution NMR Spectroscopy. Journal of Molecular Biology 428, 323–331 (2016). URL https://www.sciencedirect.com/science/article/pii/S0022283615006932. [DOI] [PubMed] [Google Scholar]
- [40].Xie T., Saleh T., Rossi P. & Kalodimos C. G. Conformational states dynamically populated by a kinase determine its function. Science 370, eabc2754 (2020). URL https://www.science.org/doi/10.1126/science.abc2754. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [41].Braun E. et al. Best Practices for Foundations in Molecular Simulations [Article v1.0]. Living journal of computational molecular science 1, 5957 (2019). URL https://pmc.ncbi.nlm.nih.gov/articles/PMC6884151/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [42].Narayan B. et al. The Transition Between Active and Inactive Conformations of Abl Kinase Studied by Rock Climbing and Milestoning. Biochimica et biophysica acta. General subjects 1864, 129508 (2020). URL https://pmc.ncbi.nlm.nih.gov/articles/PMC7012767/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [43].Meller A., Bhakat S., Solieva S. & Bowman G. R. Accelerating Cryptic Pocket Discovery Using AlphaFold. Journal of Chemical Theory and Computation 19, 4355–4363 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [44].Valsson O., Tiwary P. & Parrinello M. Enhancing Important Fluctuations: Rare Events and Metadynamics from a Conceptual Viewpoint. Annual Review of Physical Chemistry 67, 159–184 (2016). URL https://www.annualreviews.org/content/journals/10.1146/annurev-physchem-040215-112229. [DOI] [PubMed] [Google Scholar]
- [45].Bussi G. & Laio A. Using metadynamics to explore complex free-energy landscapes. Nature Reviews Physics 2, 200–212 (2020). URL https://www.nature.com/articles/s42254-020-0153-0. [Google Scholar]
- [46].Zimmerman M. I., Porter J. R., Sun X., Silva R. R. & Bowman G. R. Choice of Adaptive Sampling Strategy Impacts State Discovery, Transition Probabilities, and the Apparent Mechanism of Conformational Changes. J. Chem. Theory Comput. 14, 5459–5475 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [47].Singh S. & Hanson S. Running and analyzing massively parallel molecular simulations (2024). URL https://chemrxiv.org/engage/chemrxiv/article-details/67099e0712ff75c3a12a09a2.
- [48].Huang X., Bowman G. R., Bacallado S. & Pande V. S. Rapid equilibrium sampling initiated from nonequilibrium data. Proceedings of the National Academy of Sciences 106, 19765–19769 (2009). URL https://www.pnas.org/doi/10.1073/pnas.0909088106. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [49].Shamsi Z., Moffett A. S. & Shukla D. Enhanced unbiased sampling of protein dynamics using evolutionary coupling information. Scientific Reports 7, 12700 (2017). URL https://www.nature.com/articles/s41598-017-12874-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [50].Wan H. & Voelz V. A. Adaptive Markov state model estimation using short reseeding trajectories. The Journal of Chemical Physics 152, 024103 (2020). URL https://pmc.ncbi.nlm.nih.gov/articles/PMC7047717/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [51].Jumper J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021). URL https://www.nature.com/articles/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [52].Mirdita M. et al. ColabFold: making protein folding accessible to all. Nature Methods 19, 679–682 (2022). URL https://www.nature.com/articles/s41592-022-01488-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [53].Heo L. & Feig M. Multi-state modeling of G-protein coupled receptors at experimental accuracy. Proteins: Structure, Function, and Bioinformatics 90, 1873–1885 (2022). URL https://onlinelibrary.wiley.com/doi/abs/10.1002/prot.26382. _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/prot.26382. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [54].Vani B. P., Aranganathan A. & Tiwary P. Exploring Kinase Asp-Phe-Gly (DFG) Loop Conformational Stability with AlphaFold2-RAVE. Journal of Chemical Information and Modeling 64, 2789–2797 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [55].Otten L., Leung J. M., Chong L. T. & Zuckerman D. M. Rectifying ai-generated protein structure ensembles for equilibrium using physics-based computations. bioRxiv (2026). URL https://www.biorxiv.org/content/early/2026/04/03/2026.03.24.714034. https://www.biorxiv.org/content/early/2026/04/03/2026.03.24.714034.full.pdf. [Google Scholar]
- [56].Levinson N. M. et al. A Src-Like Inactive Conformation in the Abl Tyrosine Kinase Domain. PLOS Biology 4, e144 (2006). URL https://journals.plos.org/plosbiology/article?id=10.1371/journal.pbio.0040144. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [57].Kwarcinski F. E. et al. Conformation-Selective Analogues of Dasatinib Reveal Insight into Kinase Inhibitor Binding and Selectivity. ACS Chemical Biology 11, 1296–1304 (2016). URL https://pubs.acs.org/doi/10.1021/acschembio.5b01018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [58].Levinson N. M. & Boxer S. G. Structural and Spectroscopic Analysis of the Kinase Inhibitor Bosutinib and an Isomer of Bosutinib Binding to the Abl Tyrosine Kinase Domain. PLOS ONE 7, e29828 (2012). URL https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0029828. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [59].Nagar B. et al. Structural Basis for the Autoinhibition of c-Abl Tyrosine Kinase. Cell 112, 859–871 (2003). URL https://www.sciencedirect.com/science/article/pii/S0092867403001946. [DOI] [PubMed] [Google Scholar]
- [60].Bowman G. R., Ensign D. L. & Pande V. S. Enhanced Modeling via Network Theory: Adaptive Sampling of Markov State Models. Journal of Chemical Theory and Computation 6, 787–794 (2010). URL 10.1021/ct900620b. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [61].E. W. & Vanden-Eijnden E. Towards a Theory of Transition Paths. Journal of Statistical Physics 123, 503–523 (2006). URL 10.1007/s10955-005-9003-9. [DOI] [Google Scholar]
- [62].McClendon C. L., Kornev A. P., Gilson M. K. & Taylor S. S. Dynamic architecture of a protein kinase. Proceedings of the National Academy of Sciences 111, E4623–E4631 (2014). URL https://www.pnas.org/doi/abs/10.1073/pnas.1418402111. https://www.pnas.org/doi/pdf/10.1073/pnas.1418402111. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [63].Dorey K. et al. Phosphorylation and structure-based functional studies reveal a positive and a negative role for the activation loop of the c-Abl tyrosine kinase. Oncogene 20, 8075–8084 (2001). [DOI] [PubMed] [Google Scholar]
- [64].Longati P., Bardelli A., Ponzetto C., Naldini L. & Comoglio P. M. Tyrosines1234–1235 are critical for activation of the tyrosine kinase encoded by the MET proto-oncogene (HGF receptor). Oncogene 9, 49–57 (1994). [PubMed] [Google Scholar]
- [65].Grädler U. et al. Biophysical and structural characterization of the impacts of MET phosphorylation on tepotinib binding. The Journal of Biological Chemistry 299, 105328 (2023). URL https://pmc.ncbi.nlm.nih.gov/articles/PMC10654029/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [66].Chiara F., Michieli P., Pugliese L. & Comoglio P. M. Mutations in the met Oncogene Unveil a “Dual Switch” Mechanism Controlling Tyrosine Kinase Activity *. Journal of Biological Chemistry 278, 29352–29358 (2003). URL https://www.jbc.org/article/S0021-9258(20)84439-1/abstract. [DOI] [PubMed] [Google Scholar]
- [67].Stamos J., Sliwkowski M. X. & Eigenbrot C. Structure of the epidermal growth factor receptor kinase domain alone and in complex with a 4-anilinoquinazoline inhibitor. The Journal of Biological Chemistry 277, 46265–46272 (2002). [DOI] [PubMed] [Google Scholar]
- [68].Hubbard S. R. The juxtamembrane region of EGFR takes center stage. Cell 137, 1181–1183 (2009). [DOI] [PubMed] [Google Scholar]
- [69].Sheetz J. B., Lemmon M. A. & Tsutsui Y. Chapter Eleven - Dynamics of protein kinases and pseudokinases by HDX-MS. In Jura N. & Murphy J. M. (eds.) Pseudokinases, vol. 667 of Methods in Enzymology, 303–338 (Academic Press, 2022). URL https://www.sciencedirect.com/science/article/pii/S0076687922001288. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [70].Julien M. et al. Multiple Site-Specific Phosphorylation of IDPs Monitored by NMR. Methods in Molecular Biology 2141, 793–817 (2020). [DOI] [PubMed] [Google Scholar]
- [71].Kumar G. S., Page R. & Peti W. Chapter Seven - Preparation of Phosphorylated Proteins for NMR Spectroscopy. In Wand A. J. (ed.) Biological NMR Part A, vol. 614 of Methods in Enzymology, 187–205 (Academic Press, 2019). URL https://www.sciencedirect.com/science/article/pii/S0076687918302519. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [72].Cheng H. et al. Inferring kinase–phosphosite regulation from phosphoproteome-enriched cancer multi-omics datasets. Briefings in Bioinformatics 26, bbaf143 (2025). URL 10.1093/bib/bbaf143. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [73].Li F., Fahie M. A., Gilliam K. M., Pham R. & Chen M. Mapping the conformational energy landscape of Abl kinase using ClyA nanopore tweezers. Nature Communications 13, 3541 (2022). URL https://www.nature.com/articles/s41467-022-31215-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [74].Cotto-Rios X. M. et al. Inhibitors of BRAF dimers using an allosteric site. Nature Communications 11, 4370 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [75].Kanev G. K., de Graaf C., Westerman B. A., de Esch I. J. P. & Kooistra A. J. KLIFS: an overhaul after the first 5 years of supporting kinase research. Nucleic Acids Research 49, D562–D569 (2021). URL 10.1093/nar/gkaa895. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [76].Castro R. L.-R. d. et al. The Journey of Data: Lessons Learned in Modeling Kinase Affinity, Selectivity, and Resistance [Article v1.0]. Living Journal of Computational Molecular Science 6, 3875–3875 (2025). URL https://livecomsjournal.org/index.php/livecoms/article/view/v6i1e3875. [Google Scholar]
- [77].Heo L. & Feig M. Multi-state modeling of G-protein coupled receptors at experimental accuracy. Proteins: Structure, Function, and Bioinformatics 90, 1873–1885 (2022). URL https://onlinelibrary.wiley.com/doi/abs/10.1002/prot.26382. _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/prot.26382. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [78].Eastman P. et al. OpenMM 8: Molecular Dynamics Simulation with Machine Learning Potentials. The Journal of Physical Chemistry B 128, 109–116 (2024). URL 10.1021/acs.jpcb.3c06662. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [79].Maier J. A. et al. ff14SB: Improving the Accuracy of Protein Side Chain and Backbone Parameters from ff99SB. Journal of Chemical Theory and Computation 11, 3696–3713 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- [80].McInnes L., Healy J. & Astels S. hdbscan: Hierarchical density based clustering. Journal of Open Source Software 2, 205 (2017). URL https://joss.theoj.org/papers/10.21105/joss.00205. [Google Scholar]
- [81].Hoffmann M. et al. Deeptime: a Python library for machine learning dynamical models from time series data. Machine Learning: Science and Technology 3, 015009 (2021). URL 10.1088/2632-2153/ac3de0. [DOI] [Google Scholar]
- [82].Metzner P., Schütte C. & Vanden-Eijnden E. Transition Path Theory for Markov Jump Processes. Multiscale Modeling & Simulation 7, 1192–1219 (2009). URL https://epubs.siam.org/doi/10.1137/070699500. [Google Scholar]
- [83].Hamelberg D., Mongan J. & McCammon J. A. Accelerated molecular dynamics: A promising and efficient simulation method for biomolecules. The Journal of Chemical Physics 120, 11919–11929 (2004). URL 10.1063/1.1755656. [DOI] [PubMed] [Google Scholar]
- [84].Miao Y., Feher V. A. & McCammon J. A. Gaussian Accelerated Molecular Dynamics: Unconstrained Enhanced Sampling and Free Energy Calculation. Journal of Chemical Theory and Computation 11, 3584–3595 (2015). URL https://pubs.acs.org/doi/10.1021/acs.jctc.5b00436. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [85].Zimmerman M. I. & Bowman G. R. FAST Conformational Searches by Balancing Exploration/Exploitation Trade-Offs. J. Chem. Theory Comput. 11, 5747–5757 (2015). [DOI] [PubMed] [Google Scholar]
- [86].Laio A. & Parrinello M. Escaping free-energy minima. Proceedings of the National Academy of Sciences 99, 12562–12566 (2002). URL http://www.pnas.org/content/99/20/12562.full. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [87].Miao Y. et al. Improved Reweighting of Accelerated Molecular Dynamics Simulations for Free Energy Calculation. Journal of Chemical Theory and Computation 10, 2677–2689 (2014). URL 10.1021/ct500090q. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [88].Shirts M. R. & Chodera J. D. Statistically optimal analysis of samples from multiple equilibrium states. The Journal of Chemical Physics 129, 124105 (2008). URL https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2671659/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [89].Tribello G. A., Bonomi M., Branduardi D., Camilloni C. & Bussi G. PLUMED 2: New feathers for an old bird. Computer Physics Communications 185, 604–613 (2014). URL http://linkinghub.elsevier.com/retrieve/pii/S0010465513003196. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All analysis scripts and plotting scripts for MSM construction and analysis are available online at https://github.com/meyresearch/kinase_analysis. MSMs and related models are deposited on Zenodo at https://doi.org/10.5281/zenodo.20544389. The KinCore-sorted template library is hosted at https://osf.io/spu2a/. All scripts for setup and equilibration of seed structures, as well as the trimmed seed structures, are available at https://github.com/sukritsingh/kinase-seeding. Simulation setup and distributed-sampling scripts are available online at https://github.com/sukritsingh/distributed-sampling, including scripts describing simulation parameters for unbiased simulations. Scripts for the framework of adaptive sampling on Folding@home are available at https://github.com/sukritsingh/fahdaptive. The raw MD simulation trajectories for ABL, EGFR and MET comprise approximately 0.5, 1.5 and 2.3 TB of data, respectively. Owing to the size of these datasets and the lack of online storage capacity, the raw trajectories are not deposited in a public repository but are freely available upon any request.



