Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Jun 22.
Published in final edited form as: Mol Cell. 2026 Apr 30;86(10):1896–1910.e7. doi: 10.1016/j.molcel.2026.04.007

Sub-2 Å Cryo-EM Structures of Transcribing RNA Polymerase II Reveal Critical Roles of Water Molecules in Catalysis

Qingrong Li 1,, Gangshun Yi 2,3,, Yue Wu 4, Sophy Xu 1, Jenny Chong 1, Xuhui Huang 4, Peijun Zhang 2,3,5,*, Dong Wang 1,6,7,8,*
PMCID: PMC7619134  NIHMSID: NIHMS2171308  EMSID: EMS214063  PMID: 42066756

SUMMARY

RNA polymerase II (Pol II) is central to gene expression, but its catalytic mechanism remains elusive due to the absence of high-resolution structural data. The role of water molecules in Pol II catalysis is unknown. Here we present three high-resolution cryo-electron microscopy structures of active Saccharomyces cerevisiae Pol II elongation complexes in distinct catalytic states: two pre-catalysis states at 1.96 Å and 2.26 Å resolution, and a post-catalysis state at 2.33 Å resolution. Each structure reveals over 700-1,350 ordered water molecules, many located at functionally critical positions. Comparative analysis shows that these waters play essential roles in proton transfer steps during Pol II catalysis, facilitating substrate recognition and trigger loop folding during nucleotide addition. Strikingly, these waters are conserved between prokaryotic and eukaryotic transcription machineries (see accompanying paper). These findings provide unprecedented mechanistic insights into Pol II catalysis and reveal vital and evolutionarily conserved roles of water molecules in transcription.

Graphical Abstract

graphic file with name nihms-2171308-f0001.jpg

In brief

The roles of water molecules in transcription have long been overlooked due to resolution limitations. Li et al. resolve high-resolution cryo-EM structures of RNA polymerase II and visualize previously undetected water molecules. These waters play essential roles in RNA polymerase II catalysis and in mediating interactions within the transcription machinery.

INTRODUCTION

RNA polymerase II (Pol II) is a central transcription machinery in the first step of gene expression, and it synthesizes pre-mRNA, non-coding RNAs, and snoRNAs1-3. However, our current understanding of Pol II transcription catalytic mechanism remains incomplete. Previous research has revealed that a mobile, conserved motif, termed trigger loop (TL), folds into a closed state upon cognate substrate binding3-8. This pre-catalysis state (substrate bound, TL-closed state) is essential for aligning active site residues, RNA primer, substrate, and metal ions in a reaction-ready configuration5,9-11. However, despite extensive research efforts over the past 25 years, only two pre-catalysis state crystal structures of Pol II elongation complex (EC) were determined at resolutions lower than 3.0 Å resolution, largely due to the mobile nature of TL as well as technical limitations in X-ray and cryo-EM approaches5,12. Due to limited resolution and crystal packing effects, these structures left uncertainty on the configurations of active site as well as the correct substrate and metal coordination poised for reaction, and failed to resolve the complete transcription bubble5,12. The functional roles of key residues for Pol II transcription remain ambiguous.

The nucleotide addition reaction is proposed through an SN2 mechanism13. However, the molecular mechanism of proton transfer steps, the center piece of catalytic mechanism of Pol II transcription remains uncertain. The identities of proton acceptor for deprotonation of 3´-OH group of the RNA primer (before nucleophilic attack) and the proton donor for pyrophosphate (PPi) product (after nucleophilic attack) remain unclear.

Finally, our current understanding of Pol II transcription is protein-centered. While emerging evidences suggest that water acts as an essential and active participant in maintaining structure, stability, dynamics and function of biomolecules14-19, the functions of water molecules in transcription machinery remain largely unexplored and underappreciated—as the “dark matter” of transcription, mainly due to a lack of high-resolution structures of Pol II elongation complexes. In particular, the roles of water molecules in Pol II catalysis, substrate recognition, and protein-nucleic acids interactions are not understood.

A direct visualization of water densities requires a resolution better than 2.5 Å. To overcome the previous challenges in obtaining high-resolution cryo-EM structure of Pol II EC complexes, in particular preferred orientation and air-water interface denaturation20, we developed a cryo-EM affinity grid with monodispersed single-particle streptavidin (mspSA) on lipid monolayers and a tethering method to attach Pol II ECs to mspSA grids21. By employing streptavidin-affinity grids and optimized data collection strategies, we have successfully determined three high-resolution cryo-EM structures of Pol II EC in pre-catalysis (at 1.96 Å and 2.26 Å resolution, respectively) and post-catalysis states (at 2.33 Å resolution). These structures reveal a complete and detailed map of a fully assembled active site and over 700-1,350 well-resolved ordered water molecules. Our work unveils unprecedented critical roles of water molecules in substrate recognition and catalytic mechanism of transcription. Strikingly, these functional waters are highly conserved between prokaryotic and eukaryotic multi-subunit RNA polymerases (see accompanying paper). The elucidation of a complete active site of transcription machinery reveals that functional waters are evolutionarily conserved and integral components of transcription machinery with critical roles in transcription catalysis, which marks a major conceptual leap beyond the traditional 'protein-centered' paradigm of transcription.

RESULTS

Cryo-EM Pol II EC structures at pre-catalysis state unambiguously define the reaction-ready configurations of a complete active site

We determined a cryo-EM structure of a cognate substrate-bound Pol II EC at the pre-catalysis state with a global resolution of 2.26 Å (Fig. S1 and Table 1). We further improved the resolution of Pol II EC at the pre-catalysis state to 1.96 Å by including elongation factor Elf1 (with the highest local resolution reaching 1.92 Å near the active site) (Fig. 1A; Fig. S2; Table 1; Movie S1). Both pre-catalysis Pol II EC structures are captured in a substrate-bound, TL closed state (highly consistent with 0.34 Å RMSD at the active site) and reveal a fully ordered complete transcription bubble with clear upstream and downstream bubble edges (Fig. 1B; Fig S3; Movie S1). These structures define the pre-catalysis active site, including a bound substrate at the addition site (Fig. 1C and Movie S1), a bent bridge helix (BH) (Fig. 1D and Movie S1), a fully closed TL (Fig. 1E and Movie S1), the RNA-DNA hybrid (Fig. 1F and Movie S1) and the kinked template DNA (Fig. 1G and Movie S1), with all side chains and ordered water molecules unambiguously resolved.

Table 1.

Cryo-EM data collection, refinement and validation statistics.

Pre-
catalysis
(with Elf1)
EC core
Pre-
catalysis
(with Elf1)
Nucleic acid
Pre-
catalysis
(Elf1-free)
EC core
Pre-
catalysis
(Elf1-free)
Nucleic acid
Pre-
catalysis
(Elf1-free)
Rpb4/7
Pre-
catalysis
(Elf1-free)
Rpb9
Pre-
catalysis
(Elf1-free)
Rpb12/Wall
Pre-
catalysis
(Elf1-free)
Jaw/Rpb9
Pre-
catalysis
(Elf1-free)
Composite
Post-
catalysis
PDB 9SV6 N/A 9QEB N/A N/A N/A N/A N/A N/A 9RYB
EMDB 55239 55240 53053 53064 53056 53060 53063 53062 53057 54374
Data collection
 Microscope Titan Krios (OPIC) Titan Krios (OPIC) Titan Krios (eBIC) Titan Krios (eBIC) Titan Krios (eBIC) Titan Krios (eBIC) Titan Krios (eBIC) Titan Krios (eBIC) Titan Krios (eBIC) Titan Krios (eBIC)
 Detector Falcon 4 Falcon 4 Gatan K3 Gatan K3 Gatan K3 Gatan K3 Gatan K3 Gatan K3 Gatan K3 Gatan K3
 Cs (mm) 2.7 2.7 2.7 2.7 2.7 2.7 2.7 2.7 N/A 2.7
 Magnification 130k 130k 105k 105k 105k 105k 105k 105k N/A 105k
 Pixel size (Å) 0.932 0.932 0.829 0.829 0.829 0.829 0.829 0.829 N/A 0.825
 Electron dose (e-/Å2) 50 50 50 50 50 50 50 50 N/A 50
 Defocus range (μm) −0.5 to 2.5 −0.5 to 2.5 −0.5 to 2.5 −0.5 to 2.5 −0.5 to 2.5 −0.5 to 2.5 −0.5 to 2.5 −0.5 to 2.5 N/A −0.5 to 2.5
 Micrograph Number 16,482 16,482 10,599 10,599 10,599 10,599 10,599 10,599 N/A 18,913
Reconstruction
 Particles refinement 360,403 67,830 692,163 50,808 107,723 130,110 143,632 134,532 N/A 743,682
 symmetry C1 C1 C1 C1 C1 C1 C1 C1 N/A C1
 resolution (Å) 1.96 2.38 2.26 3.12 3.11 3.05 3.03 3.00 N/A 2.33
 sharpening B-factor (Å2) 46.0 40.8 73.6 57.8 86.7 150.0 117.9 100.50 N/A 71.2
Model composition
 number of atoms 35033 N/A 33759 N/A N/A N/A N/A N/A N/A 33090
 protein residues 3983 N/A 3908 N/A N/A N/A N/A N/A N/A 3931
 nucleotides 97 N/A 97 N/A N/A N/A N/A N/A N/A 54
 Ligand 1 (ATP) N/A 1 (ATP) N/A N/A N/A N/A N/A N/A 1 (PPi)
Bonds RMSD
 Bonds lengths (Å) 0.023 N/A 0.003 N/A N/A N/A N/A N/A N/A 0.012
 Bonds angles (°) 0.93 N/A 0.628 N/A N/A N/A N/A N/A N/A 0.521
Validation
 MolProbity score 1.48 N/A 1.65 N/A N/A N/A N/A N/A N/A 1.97
 Clash score 6.31 N/A 9 N/A N/A N/A N/A N/A N/A 7.28
 Rotamer outliers 0.2% N/A 1.0% N/A N/A N/A N/A N/A N/A 3.7%
 C-beta outliers 0.0% N/A 0.0% N/A N/A N/A N/A N/A N/A 0.0%
Ramachandran plot
 Favored (%) 97.3% N/A 97.0% N/A N/A N/A N/A N/A N/A 97.0%
 Allowed (%) 2.7% N/A 3.0% N/A N/A N/A N/A N/A N/A 2.9%
 Outlier (%) 0.0% N/A 0.1% N/A N/A N/A N/A N/A N/A 0.1%

Figure 1.

Figure 1.

High-resolution cryo-EM structure of Pol II elongation complex (EC) at the pre-catalysis state with substrate ATP binding.

(A) Pol II EC structure at 1.96 Å resolution. Color code and abbreviation: RNA (hot pink); template DNA (ts-DNA cyan); non-template DNA (nts DNA, lime); bridge helix (BH, light green); trigger loop (TL, purple); ATP (red). Bottom panel: local resolution map from 2.0 Å (blue, high) to 3.6 Å (red, low). (B) Cryo-EM density and scheme of the complete transcription bubble in the Pol II EC.(C–G) Cryo-EM densities of selected regions, including (C) the active site (gray) showing two Mg2+ ions (green) and ATP substrate (hot pink); (D) BH; (E) TL; (F) RNA-DNA hybrid; and (G) kinked tsDNA interacting with switch loop 1 (silver) and switch loop 2 (salmon). Cryo-EM density maps are contoured at 5-σ (contour level = 0.12). In panel (C), the 3′-OH group (not present in the structure) was modeled onto the RNA primer to illustrate the expected coordination with metal A.

These high-resolution pre-catalysis Pol II EC structures define the reaction-ready configuration of a complete active site. The 1.96 Å cryo-EM map reveals un unambiguous density of ATP substrate with a chair-like triphosphate conformation precisely coordinated with two Mg2+ ions (Fig. 1C; Movie S1). Importantly, the octahedral coordination of metal B (Mg2+) is clearly resolved with tridentate coordination with all three phosphate groups (α, β, and γ phosphate) of the ATP substrate, which is different from previously reported low-resolution structures5,12,22,23 (Fig. S4). A water molecule, W10, is found to coordinate with metal B as the sixth coordinate (Figs. 1C and 2B-II). The carboxyl groups of Asp481 (Rpb1) and Asp483 (Rpb1) bridge two Mg2+ ions (Mg2+ A and B), which also align closely with their counterparts in DNA polymerase24-26 (Figs. S4K-S4M). In addition to Asp481 and 483 residues, Mg2+ A is further coordinated with α phosphate of ATP substrate, Asp485, an ordered water molecule (W0), and presumably 3´-OH of RNA primer (Figs. 1C and 2B-I).

Figure 2.

Figure 2.

Water molecules mediate interactions between substrate ATP and the catalytic center at the pre-catalysis state.

(A) Hydrogen-bonding interactions between the substrate ATP (hot pink) and the catalytic center in the classic protein-centered view (left) and updated water-integrated view (right). Residues from different domains are color-coded as indicated. Water molecules are highlighted as red spheres. the residues that interact with ATP via water molecules are outlined with red borders (identified in this study). Water-mediated hydrogen bonds are shown as blue dashed lines; direct protein–substrate hydrogen bonds as gray dashed lines; and Mg2+-mediated coordination bonds as light purple dashed lines. A semi-transparent 3′-OH group was modeled onto the RNA primer to illustrate the expected coordination with metal A. (B) Cryo-EM structures of water-mediated interactions within the Pol II substrate recognition network. The middle panel displays the spatial distribution of water molecules (transparent red spheres) involved in substrate recognition. Hydrogen bonds and coordination bonds are indicated by blue and green dashed lines, respectively. Insets (I–VIII) present close-up views of local regions with cryo-EM density maps (shown as mesh). Contour levels are indicated in panels highlighting active-site side chains; all other close-up views are contoured 5-σ. Thick dashed lines represent hydrophobic interactions between TL residue Leu1081 and the nucleobase of the substrate ATP in the middle panel and inset VI.

These high-resolution structures enable the determination of the rotamers and conformations of all key residue side chains and provide an accurate model for understanding Pol II transcription (Movie S1). The TL is in a fully closed conformation with six turns of the proximal half (from Met1063 to Asn1082) and four turns of the distal half (from Lys1092 to Val1107), forming a fully folded alpha helix (Movie S1). We also clearly discerned the density of side chains and main chains of TL tip (from Thr1083 to Ser1091), which was poorly resolved in previous reported structures5,12 (Fig S5; Movie S1; Table S1). The cryo-EM map clearly resolves the rotamers of all critical TL residues, including Gln1078, Leu1081, Asn1082, Phe1084, and His1085 (Figs. 2A, 2B-IV, 2B-V, 2B-VI, and 2B-VII; Fig. S5). Key interactions and functional roles of these critical TL residues in the substrate recognition network are clarified (Figs. 2A and 2B; Movie S1). In addition, hydrophobic interactions at the TL tip—such as Phe1086(TL)–Val1089(TL)–Ile756(funnel) and Phe1084(TL)–Pro765(HB)–Tyr769(HB)–Leu842(BH)—are now visualized (Figs. S5G and S5H). We further reveal that the minor-groove steric gate residue Pro448 adopts a cis-conformation (Fig. 2B-VIII, inset), in contrast to the ambiguous conformations reported in earlier Pol II structures.

Cryo-EM Pol II EC structure unveils extensive water-mediated network of interactions between substrate and Pol II catalytic center

Most importantly, all water molecules come into view in the 1.96 Å Pol II EC structure, unveiling a previously unrecognized pivotal role of water molecules in mediating the substrate recognition network in Pol II (Fig. 2A). We identified 1,357 ordered waters (with a Q-score higher than 0.7 and a density value above 4-σ) and one monovalent metal ion located at protein-protein and protein-nucleic acids interfaces of the Pol II EC (Fig. S6; Supplementary Spreadsheet 1 and 2). Among these waters, thirteen are at the active site, whose mechanistic roles are described below (Fig. 2).

In contrast to the traditional protein-centered view (Fig. 2A, left), we now reveal that thirteen ordered water molecules form interaction networks that connect all moieties of the substrate to key active site residues (Fig. 2A, right). The involvement of these water molecules greatly expands current understanding of the substrate interaction network to encompass seven additional residues from distinct domains: Gln1078, Leu1081 and Asn1082 from the TL, Pro448 from the active site domain, Tyr769 and Asp837 from hybrid binding (HB) domain, and Thr827 from the bridge helix (BH) (Fig. 2B, residues with red rectangle outline). Notably, these residues have previously been examined by biochemical and mutagenesis studies, yet the structural basis for their functional effects remain unclear. Our water-integrated model provides a framework to rationalize these observations, for example for mutations such as N479S5, Q1078S27 or N1082S27, through disruption of water-mediated interactions (Fig. S7; Table S2).

Intriguingly, water-mediated interactions are important for Pol II recognition of the sugar moiety of substrate (Figs. 2A, 2B-VI and 2B-VIII). In addition to previously reported direct interaction between the 2´-hydroxyl group (2´-OH) of ATP and Arg4465,6, water-mediated interactions are made between an ordered water molecule (W2) and the 2´-OH group of ATP and the amide groups of Asn479 and Gln1078 of the TL (Fig. 2B-VIII), supporting that TL folding is a checkpoint for NTP ribo/deoxyribo discrimination from an early single-molecule fluorescence spectroscopy study28. W2 further interacts with a water chain consisting of W8, W30 and W25, which in turn interacts with Arg446, Pro448, Pro477, Ser454 and Tyr478 on the active site domain (Fig. 3B and Fig. S5I). The 3′-OH of the substrate ATP is directly recognized by Gln1078 and Asn479 (Figs. 2B-VI and 2B-VIII), clarifying previous ambiguity of these interactions12. In addition, we identified a water molecule (W3) that acts as a hub coordinating the 3′-OH and the β-phosphate group of ATP, the amide group of Asn1082, and the main chain of Leu1081 (Fig. 2B-VI).

Figure 3.

Figure 3.

Water molecules engage in trigger loop folding.

(A) Ordered waters (red spheres) surround the folded TL at the Pol II active site. The TL tip inserts into a pocket between the funnel (orange) and hybrid-binding (HB, steel blue) domains. (B) Water-mediated hydrogen-bonding network surround the folded TL at the Pol II active site. Water molecules are classified into three categories: category A (red sphere with orange rim), category B (red sphere with blue rim), and category C (red sphere with green rim). The representations of hydrogen bonds and coordination bonds are consistent with those in Figure 2A. (C) Water-mediated interactions that stabilize the TL tip and its contacts with funnel (orange, Rpb1) and HB (steel blue, Rpb2) domains. (D) Water-mediated interactions underlying the helical bundle formation between the TL and the BH. The middle panel shows the spatial distribution of water molecules involved in bundle formation, with four close-up views. All cryo-EM density maps in C and D are shown as mesh contoured at 5-σ.

In addition to sugar recognition, we also observed several waters interacting with the triphosphate moiety of ATP. The water molecule W6 bridges the α-phosphate of ATP to its nucleobase (N7) through a water network (W6–W13–W1), and additionally via W6–W13–W7 to TL residue Leu1081 and BH residue Thr827 (Fig. 2A). For β-phosphate group recognition, beyond the previously reported direct interaction with conserved His10855, here we identified two water molecules, W3 and W4, that connect to the β phosphate group of ATP, which can act as proton donors for proton transfer during nucleotide addition (see Discussion section on Pol II catalysis). These two water molecules mediate the interaction between the β phosphate group and the TL with residues Leu1081 and Asn1082 (Fig. 2B-VI). In addition to the direct interactions formed by HB residues Arg766 and Arg1020 with the γ phosphate group, a water molecule (W11) is observed bridging the α- and γ-phosphate groups and mediating interactions with HB residue Tyr769 (Fig. 2B-III).

Water molecules are important for Mg2+ ion coordination. Specifically, W0 is identified to coordinate with metal A and interact with active-site residue Asp485, as well as connecting to a water chain that includes W5 and W9 (Fig. 2B-I). Notably, W0 may also serve as a proton acceptor to facilitate deprotonation of the 3´-OH of the RNA primer during nucleotide addition (see Discussion on Pol II catalysis). Similarly, water molecule W10 is identified coordinating with metal B, completing its octahedral coordination. To our knowledge, a fully coordinated metal B in RNA Pol II has not been clearly observed in previous structures. (Fig. 2B-II).

To evaluate the kinetic stability of resolved water molecules in the cryo-EM structure, we performed all-atom molecular dynamics (MD) simulations with position restraints applied to Pol II heavy atoms and used the “residence time” of hydration sites as a straightforward metric. Key hydration sites centered on critical water molecules exhibited residence times significantly longer than bulk water (Fig. S7). Most water molecules mediating interactions between the ATP substrate and the Pol II catalytic center (Fig. 2A) demonstrated residence times ranging from 102 to 104 ps (for comparison, bulk water: 3ps). These results highlight the critical role of water molecules in stabilizing substrate interactions at the active site.

Water molecules are critical for TL folding and interactions with other catalytic center motifs

The high-resolution Pol II EC structure also reveals that water molecules play a crucial role in stabilizing the closed TL conformation as well as its interactions with other key structural motifs, including the bridge helix (BH), funnel, hybrid-binding (HB) domain, anchor, cleft, and active site domain (Figs. 3A and 3B; Fig. S5). We identified more than 50 water molecules that form extensive water-mediated hydrogen bonds that bridge residues of TL and other functional motifs in the active center (Fig. 3B). These water-mediated interactions profoundly alter the canonical protein-centered view of TL function.

We divided these waters into three categories based on their function and location: category A water (total 17 water molecules; orange-red shade) that mediated interactions between TL tip (Thr1083–Ser1091) and funnel and HB domains; category B water (total 31 water molecules; blue shade) that mediated interactions between two stable stem helices of TL (Met1063–Asn1082 and Lys1092–Val1107) with substrate and other active center motifs; category C water (seven water molecules; green shade) that mediated interactions among TL and BH helices bundle (Fig. 3B).

The category A water molecules are crucial for stabilizing the sharp-turn conformation of the TL in its fully closed state (Fig. 3C, Fig. S5). The TL tip region is flexible and can switch between a fully-folded state and disordered open states during nucleotide addition. Upon TL folding, the TL tip inserts into a pocket formed by the BH, funnel, and HB domains (Fig. 3C, Fig. S5). In the classic protein-centered view, interactions at the TL tip are primarily attributed to hydrophobic contacts, as the tip region is enriched in hydrophobic residues (Phe1084, Phe1086, Ala1087, Val1089, and Ala1090) (Figs. 3B and 3C). In this view, only two direct hydrogen bonds are observed at this interface: one between HB residue Gln763 and TL residue His1085, and another between His1085 and the β-phosphate of the incoming ATP substrate (Fig. 3B). However, our structure reveals an extensive water-mediated network bridging the TL tip with the funnel and HB domains. In total, 17 ordered water molecules are positioned at this interface, mediating interactions between the TL tip and 11 surrounding residues—two from the HB domain and nine from the funnel domain (Figs. 3B and 3C). These interactions are predominantly formed through water-mediated hydrogen bonds involving the main chain of TL tip residues, effectively expanding the interaction network beyond the limited direct contacts observed previously (Fig. 3C). In addition, one ordered water molecule (W19) coordinates three TL residues by interacting with the main-chain carbonyl groups of Phe1086 and Val1089, and the side chain of Thr1083, thereby stabilizing the fully closed conformation of the TL tip (Fig. 3C). Together, these findings highlight the critical role of water molecules in stabilizing TL folding and mediating interactions between TL-tip and other functional domains.

Category B waters (total 30 water molecules) mediate substrate recognition (also see above substrate recognition section) as well as communication with the anchor, cleft, and active site domains, and the Rpb6 subunit (Fig. 3B; Figs. S5I-S5L). Specifically, these extensive water-mediated interactions establish extensive connections between TL and nine cleft residues, three anchor residues, nine active site residues, and one Rpb6 residue, respectively (Fig. 3B; Figs. S5I-S5L). These waters are important for mediating long-range communication between TL and other functional domains within Pol II transcription machinery.

As for category C waters, we identified seven additional water molecules (W7, W27, W29, W34, W36, W48 and W55) spanning the entire helix bundle interface between the TL (Val1066–Thr1083) and the BH (Gly823 to Ile848) (Fig. 3D). These water-mediated interactions likely contribute to the stability and coupling of the TL/BH helices bundle.

Extensive water-mediated interactions at protein-nucleic acid interfaces

As a key feature conserved in all multi-subunit RNA elongation complexes, template DNA (tsDNA) crosses over the BH domain and reaches the active site, where it forms a kinked conformation at the position between i+1 and i+2, with an angle of nearly 90°. The negatively charged phosphate groups get much closer in this kinked configuration. However, how this kinked configuration is maintained is not fully understood. Here we reveal that the water molecules play an important role in maintaining the kinked conformation of tsDNA.

We found extensive water-mediated interactions with the phosphodiester groups within the kinked region (i-1, i+1, and i+2) (Fig. 4A). One water molecule (W72) was found in the core of the kinked region, bridging the phosphate groups from i+1 and i+2 tsDNA and potentially stabilizing the kinked conformation (Fig. 4A, close-up view). In addition, we found 19 waters (W39, W71–81, W83, W85, W86, W88–91) that connect the kinked DNA with bridge helix (Ala828, Tyr836, and Arg839), the switch loop 1 (Phe1402 and Glu1403), the switch loop 2 (Arg337 and Lys332), and the anchor domain (Glu1132 and Met1133) (Fig. 4A).

Figure 4.

Figure 4.

Water molecules stabilize kinked template DNA, RNA-DNA hybrid, and protein-nucleic acids interfaces.

(A) Water-mediated interactions between the kinked template DNA (cyan) and surrounding domains: BH (light green), switch loop 1 (silver), switch loop 2 (salmon), and anchor domain (olive). An inset shows the cryo-EM density of the kinked region, bent at ~90°, with the associated water molecule (density contoured at 5-σ). (B) Water molecules in the RNA-DNA hybrid region of Pol II EC at the pre-catalysis state. Water molecules are grouped into four clusters according to their interaction partners within the hybrid region: interactions with the major groove (blue), minor groove (red), phosphodiester backbone (light blue), and 2´-OH of the RNA strand (purple). (C–D) Detailed views of water-mediated interactions at the i+1 (C) and i−1 positions (D). CryoEM densities are displayed as mesh contoured at 5-σ, and hydrogen bonds are indicated with blue dashed lines. (E) Cartoon representation of water-mediated interactions for the backbone of the RNA-DNA hybrid at the pre-catalysis state. Residues from different domains are color-coded as indicated. The residues forming water-mediated contacts with the RNA-DNA backbone are outlined with red borders (identified in this study). Hydrogen bonds and coordination bonds are depicted as blue and purple dashed lines, respectively.

In addition to the kinked template DNA region, we also observed extensive water-mediated interactions in upstream RNA-DNA hybrid region (Figs. 4B to 4E; Fig. S8). A hydration shell is identified to cover the RNA-DNA hybrid region (Fig. 4B). Based on the locations and interactions, the water molecules can be categorized into four groups, including the backbone-interacting water molecules (63 waters, light blue sphere), major groove-interacting water molecules (23 waters, blue sphere), minor groove-interacting water molecules (24 waters, red sphere), and a group of water molecules involved in ribose discrimination (nine waters, purple sphere) (Fig. 4B). Water molecules were found to form extensive hydrogen bonds with the polar groups of nucleobases on both RNA and template DNA strands (Figs. 4C and 4D). These waters form “water spines” that run along the minor and major grooves of the RNA-DNA hybrid, which are important for RNA-DNA hybrid stability29,30 (Fig. 4B; Fig. S8). The water spine observed within the Pol II EC is similar to the previously reported crystal structure of a nine-mer RNA-DNA hybrid duplex (PDB: 421D)29. Notably, these ordered waters are mainly located between i+1 to i-6 base pairs, suggesting a potential role of these waters in stabilizing the RNA-DNA hybrid during transcription process (including initiation, elongation, and termination), in particular for short RNA-DNA hybrid during de novo transcription.

Intriguingly, in sharp comparison with traditional view of Pol II-nucleic acids interaction (manly via direct interactions with positive residues), we now observed extensive water-mediated interactions at the protein-nucleic acid interfaces involving a much broader range of Pol II residues, including charged (both positive and negative), polar, and nonpolar residues (Fig. 4E, residues with red rectangle outlines). We identified a total of 63 waters and 39 residues that are involved in protein-nucleic acids interactions (Fig. 4E). Among them, only eight residues are involved in direct interaction. These water molecules primarily mediate contacts between the phosphodiester backbones of RNA or DNA and Pol II residues, forming hydration shells at the protein-nucleic acids interface in a sequence-independent manner (Fig. 4E).

These water molecules at protein-nucleic acids interfaces likely act as molecular lubricants, facilitating the sliding of the nucleic acids scaffold along the binding interface and thereby promoting Pol II translocation during transcription elongation. Two mechanisms may contribute to underlie this effect. First, displacing water molecules on the DNA surface requires less energy than breaking direct protein-DNA contacts, as shown by Overhauser Dynamic Nuclear Polarization (ODNP) studies31. Second, water bridges enable a broader range of residues—including negatively charged, polar and non-polar ones—to interact with nucleic acids, beyond the typical positively charged Lys and Arg. Indeed, we identified eight acidic residues that engage the RNA-DNA hybrid via water-mediated interactions, reducing the overall positive charges of nucleic-acid binding surface and allowing smoother nucleic acid movement (Fig. 4E).

Water molecule dynamics during Pol II catalysis

To gain insight into the roles of waters during Pol II NTP addition reaction, we determined a cryo-EM structure of the Pol II EC captured at a transient post-catalysis state at 2.33 Å resolution by allowing NTP addition on cryo-EM grid (Figs. 5A and 5B; Fig. S9). Examination of the active center in this structure revealed well-defined densities corresponding to the newly formed phosphodiester bond between the RNA primer and the incoming substrate (Fig. 5B inset, red arrow), as well as pyrophosphate (PPi) product (Fig. 5B). Interestingly, the PPi migrates away from its original position shared by metal B and beta and gamma phosphate of the NTP in pre-catalysis state, suggesting a transition pathway toward release via the secondary channel. In contrast to the well-defined density of metal B in pre-catalysis state, the density corresponding to metal B in post-catalysis state is less defined, reflecting its dynamic nature. TL adopts an unfolded conformation in the post-catalysis state (Fig. 5B). Notably, Leu1081, a key residue involved in stabilizing the incoming substrate NTP in the pre-catalysis state, remains detectable in the post-catalysis structure (Fig. 5B). The nucleic acid scaffold remains the same as in the pre-translocation configuration (Fig. 5B).

Figure 5.

Figure 5.

Cryo-EM structure of Pol II EC captured at the post-catalysis state at 2.33 Å.

(A) Overview of the Pol II EC in the post-catalysis state with a cross-sectional view. The product pyrophosphate (PPi) is shown in gold. Other elements are color-coded as indicated. (B) Cryo-EM density maps and corresponding atomic models of the active center in the post-catalysis structure. A close-up view highlights the newly formed phosphodiester bond, with the density map (contoured at 4-σ, contour level = 0.17). The red arrow marks the position of the new bond. Water molecules in the post-catalysis structure are shown in sky-blue sphere. (C) Interaction network of the newly incorporated nucleotide at the post-catalysis state. Water molecules identified in the post-catalysis structure are labeled O1–O7 (sky-blue sphere). Water molecules O2–O7 are conserved from the pre-catalysis state, with their corresponding identifiers (W0, W2, W5, W6, W8, and W9) indicated. A unique water molecule, O1, is observed exclusively in the post-catalysis state, where it coordinates with Mg2+ A. By contrast, water molecules W1, W3, W4, W7, W10, W11, and W13—present in the pre-catalysis state—are absent post-catalysis and shown as transparent red spheres. Residues from different domains are color-coded as indicated. The representations of hydrogen bonds and coordination bonds are consistent with those in Figure 2A. (D) Detailed structural view of the interactions involving the newly incorporated nucleotide. A close-up view shows the density maps of Mg2+ A coordination, contoured at 4-σ. (E) Structural comparison of water molecules associated with TL folding and unfolding between the pre-catalysis (salmon) and post-catalysis (steel blue) states. Water molecules that interact with the folded TL tip in the pre-catalysis state but are absent in the post-catalysis state due to TL unfolding. (F) Superposition of pre-catalysis (salmon) and post-catalysis (steel blue) structures highlighting overlapping water molecules positioned at the interface between the TL, active site domain, and anchor domain.

In the post-catalysis state, we identified over 712 ordered water molecules. To further investigate the functional and structural roles of water molecules within the Pol II EC during catalytic reaction, we compared water distribution patterns between pre- and post-catalysis structures. While most water molecules (>92%, 660 overlapped water out of 712 total water in post-catalysis structure) occupy similar positions in both states, those located at the active site show markedly distinct distributions (Figs. 5C and 5D; Supplementary Spreadsheet 3). Based on these differences, we classified the waters into two major categories: catalytic state-specific waters (class 1) and structural waters (class 2).

Class 1 waters are specific to a particular catalytic state and are primarily localized near the active site. Accompanying nucleotide addition and TL unfolding, a total of 28 water molecules (W3-4, W7, W10-13, W18-20, W22-24, W27, W36-37, W40, W43, W45-48, W58, W60-62, W64, W66) are absent in the post-catalysis structure, underscoring their functional importance in stabilizing the folded TL and its interaction with other active center motifs in the pre-catalysis state (Figs. 5C to 5E; Movie S2). Following catalysis, we observed the appearance of a unique Mg2+ A–coordinated water molecule (O1) exclusively present in the post-catalysis structure. This water replaces the coordination previously provided by the α-phosphate oxygen of the NTP in the pre-catalysis state. As a result, two waters (O1 and W0) coordinate with Mg2+ A in the post-catalysis state (Figs. 5B to 5D).

Class 2 waters are present in both pre- and post-catalysis structures (Fig. 5F; Supplementary Spreadsheet 3). Notably, 660 out of 712 identified in the post-catalysis structure (over 92%) of water positions overlap with corresponding waters in the pre-catalysis structure within a 1.5 Å threshold, highlighting their remarkable positional consistency. This strong conservation underscores the role of structural water molecules as integral components of Pol II EC architecture. These waters contribute to both subunit-subunit interactions and protein-nucleic acid interfaces.

Remarkably, despite the nucleotide sequence variation between the two states, a persistent hydration shell—particularly clusters of ordered water molecules— is observed along the i+1 to i-6 base pairs in both structures (Figs. S8I and S8J). While individual water molecules interacting with major groove and minor groove edges may shift slightly to accommodate different base pair function groups, the overall hydration pattern remains conserved (Figs. S8I and S8J). This sequence-independent clustering suggests that water-mediated interactions in the i+1 to i-6 region are an intrinsic structural feature of Pol II’s engagement with the RNA-DNA hybrid. These interactions likely play essential roles in precise positioning of the hybrid and substrate during catalysis and may also help stabilize short RNA-DNA hybrid during de novo transcription, thereby reducing abortive transcription. Similarly, the pattern of water-mediated interactions within the kinked template DNA region is highly conserved across both states.

DISCUSSION

Water molecules are essential for proton transfer during Pol II catalysis

High-resolution structures with water molecules also offer mechanistic insights into proton transfer during Pol II catalysis. The absolutely conserved TL residue His1085 was hypothesized to function as a proton donor as it is the only pol II residue that directly contacts the β-phosphate of ATP5,8,32. However, genetic and biochemical results of this conserved residue are perplexing: Several substitutions at this position (e.g., H1085A, H1085D, H1085N, or H1085F) are lethal in yeast, whereas certain substitutions (e.g., H1085Q or H1085L) remain viable, albeit with a 2–10-fold reduction in transcription elongation rate in these mutants33,34. These results suggest the proton-donating capacity of His1085 is not strictly essential. Importantly, the viability of H1085Q and H1085L mutants suggests the presence of other proton donors for Pol II catalysis in these H1085 mutants.

Our high-resolution structures now provide a clear structural explanation for this observation. We identify two ordered water molecules (W3 and W4) that directly engage the β-phosphate of the incoming NTP and are geometrically positioned to participate in proton transfer. These observations highlight the functional importance of these waters function as proton donors in WT as well as H1085L and H1085Q mutants. In addition to these water-mediated interactions, we also unambiguously confirmed the direct interaction between His1085 and beta phosphate (Fig. 6, see below).

Figure 6.

Figure 6.

Substrate recognition and catalytic mechanism of Pol II EC in water-integrated view.

Proposed Pol II catalytic mechanism, including deprotonation of the RNA primer 3′-OH (A), SN2 attack and protonation of PPi (B), and PPi release (C). During the deprotonation step (A), the proton of 3′-OH of the RNA primer is transferred through a water chain (W0–W5–W9) that connects to the bulk solvent. In SN2 attack step (B), Water molecules W3 and W4 are positioned to function as proton donors. Pathway “1” represents the water-mediated protonation route involving W3/W4, whereas pathway “2” denotes the alternative His1085-mediated route. In the post-catalysis state (C), formation of the PPi (orange) product is coupled to TL unfolding (depicted by purple dashed lines). Electron transfer pathways are illustrated by blue arrows.

Leveraging our high-resolution cryo-EM structures and previous studies5,12,22,35,36, we now propose a water-mediated SN2 catalytic model for the Pol II EC that explicitly incorporates the critical roles of water molecules (Fig. 6; Movie S2). Prior to catalysis, water molecules mediate key interactions that support TL-folding and substrate recognition (Figs. 2A, 3B, and 6A). The base, sugar and phosphate moieties of the incoming NTP are recognized via both direct and water-mediated interactions (Figs. 2A and 6A). During catalysis, a chain of water molecules plays critical roles in deprotonating 3′-OH. We found that water molecule W0, coordinates with Mg2+ A, acts as a proton acceptor, abstracting a proton from the 3′-OH of RNA primer. This deprotonation step reduces the active energy barrier for SN2 nucleophilic attack on the α-phosphate of the NTP. W0 transfers the proton to the solvent via a water chain relay involving W5 and W9, consistent with prior computational studies showing this pathway has the lowest energy barrier (Fig. 6A) 32.

As the reaction proceeds, a pentavalent phosphate intermediate form (Fig. 6B). The electron transfer pathway is illustrated by blue arrows (Fig. 6B). Consequently, a new phosphodiester bond is formed between the 3´- RNA primer and the α-phosphate of the substrate NTP, breaking the bond between the α- and β-phosphates, extending the RNA primer by one nucleotide, and releasing pyrophosphate (PPi) (Fig. 6C). Our structures also clarify the long-standing question of how PPi is protonated. We reveal that water molecules (W3 and/or W4) and/or His1085 provide redundant protonation pathways, ensuring catalytic robustness even when His1085 is mutated. These water molecules serve as proton donors to ensure that catalysis proceeds even in the absence of His1085.

Water molecules are evolutionarily conserved among prokaryotic and eukaryotic transcription machineries

Our work revealed the essential roles of waters in Pol II transcription. This also raises an intriguing question: whether these functional important waters are evolutionally conserved across different RNA polymerases among prokaryotic and eukaryotic transcription machineries. Intriguingly, we compare water peaks identified from Pol II elongation complex and E. coli RNAP complex (Please see accompanying paper from Seth Darst’s group). Strikingly, as shown in Fig. S10 and Supplementary Spreadsheet 3, we are able to identify over 230 common waters that are shared among different structures within the cutoff 1.5 Å. Key water molecules that are involved in deprotonation and protonation steps during catalytic reaction are strictly conserved between E. coli RNAP and S. cerevisiae Pol II (Fig. S10). These waters include W0/W5/W9 water chains-proposed to deprotonate 3′-OH and W3—proposed to act as a proton donor for PPi pronation during the nucleotide incorporation reaction (Fig. S10). Another potential proton donor, water molecule W4, is absent in E. coli, which may be attributed to the substitution of the TL residue Asn1082 in S. cerevisiae Pol II with Arg933 in E. coli RNAP (Fig. S10). The majority of water molecules for substrate recognition are also highly conserved (11 out of 13 common waters, Fig. S10). These striking findings further highlight that these functional water molecules are evolutionarily conserved integral components of transcription machineries and play critical roles in the structure and function of transcription.

Conclusion

Our high-resolution Pol II EC structures at distinct stages of nucleotide addition resolve long-standing ambiguities in active-site architecture, including precise substrate and metal ion coordination within a fully assembled catalytic center. More importantly, we reveal important roles of water molecules in Pol II catalysis during nucleotide addition, substrate recognition, and TL folding. Unlike a canonical protein-centered view, we now uncover extensive and previously underappreciated water-mediated interaction networks across transcription machinery. They also serve as integral structural elements, stabilizing kinked DNA, mediating subunit interfaces, and facilitating protein-nucleic acids interactions. Critically, our findings redefine the nature of Pol II-nucleic acid contacts, revealing that water-bridges enable interactions involving a broad spectrum of residue types—including negatively charged and nonpolar residues—beyond the canonical positively charged ones. Such knowledge is highly relevant to many other nucleic acids enzymes or nucleic acids binding proteins. The concept of water as a molecular lubricant may represent a general mechanism for facilitating nucleic acid translocation across diverse nucleic acid-processing enzymes and motor proteins.

Together with the study of E. coli RNAP, we revealed that these functionally important waters in RNAP catalysis are evolutionarily conserved between prokaryotic and eukaryotic multi-subunit RNA polymerases. These observations highlight the fundamental importance of waters in transcription. Our work provides a foundational framework for future molecular dynamics and time-resolved structural analyses aimed at probing the dynamic roles of ordered and disordered water and ions during transcription.

Limitations of the study

Our sub-2 Å cryo-EM structures identify two ordered water molecules positioned toward the β-phosphate of the incoming NTP. Based on their geometry and proximity, these waters provide a structurally plausible proton donor pathway operating in parallel with His1085. It is worthy to note that previous mutation and biochemical studies as well as current structural studies cannot completely exclude a potential contribution from His1085 as a proton donor in addition to these waters in WT Pol II. At the physiological pH, His1085 can exist in both protonated (charged) and deprotonated (neutral) states, which both can form hydrogen bonds with the phosphate of substrate. The protonated state of His 1085 would be readily to serve as a proton donor during SN2 reaction. Accordingly, we propose a parallel model for protonation transfer in which water-mediated and His1085-mediated pathways may coexist. However, this study cannot quantify the relative contributions of water-mediated and His1085-mediated proton transfer in WT Pol II. Thus, the precise role and relative contribution of His1085 in the proton transfer step remains an open question that will require future biochemical and computational investigations to resolve.

STAR★METHODS

EXPERIMENTAL MODEL AND STUDY PARTICIPANT DETAILS

Plasmids and strains

Saccharomyces cerevisiae strains expressing endogenously TAP-tagged RNA polymerase II were cultured in 2×YPD medium at 28 °C.

Plasmids encoding Rpb4/7 and Elf1 were transformed into E. coli BL21 (DE3) cells and expressed recombinantly in standard LB medium.

METHOD DETAILS

Protein purification and complex assembly

Saccharomyces cerevisiae 12-subunit RNA polymerase II (Pol II) was purified following procedures described previously5,45. To assemble the complete 12-subunit Pol II, the purified 10-subunit Pol II was mixed with the recombinant Rpb4/7 subcomplex in vitro and subjected to gel filtration chromatography purification. Fractions containing the fully assembled 12-subunit Pol II were collected for further structural analysis.

The biotinylated Pol II elongation complex was assembled in vitro as previously reported21. Briefly, all RNA and DNA oligonucleotides used in this study were PAGE-purified and purchased from Integrated DNA Technologies (IDT). The sequences used to form the elongation complex were: Biotinylated template strand DNA (tsDNA): 5′-/BiotinTEG/ TTT TTT GAT ATT TTT GGA TCC CGC TCT GCT CCT TCT CCC ATC CTC TCG ATG GCT ATG AGA TCA ACT AGG AAT TC-3′; Biotinylated non-template strand DNA (ntsDNA): 5′-/BiotinTEG/ TTT TTA TGT ATT AAT GAA TTC CTA GTT GAT CTC ATA GCC CAT TCC TAC TTG GGA GAA GGA GCA GAG CGG GAT CC-3′; 3′-deoxy RNA (for pre-catalysis): 5′-AUC GAG AG/3′dG; regular RNA (for post-catalysis): 5′-AUC GAG AGG.

To assemble the nucleic acid scaffold, the RNA and template-strand DNA (tsDNA) were first mixed and annealed. For pre-catalysis samples, a 3′-deoxy RNA was used to prevent substrate incorporation, whereas for post-catalysis samples, a regular RNA was used to permit catalysis. Pol II was then added to assemble the elongation complex (EC), followed by addition of the non-template DNA (ntsDNA) strand to yield the fully assembled EC. Complex formation was carried out in elongation buffer [20 mM Tris-HCl (pH 7.5), 40 mM KCl, 10 mM MgCl2, and 5 mM DTT]. The final concentrations of components in the assembled EC were 1 μM Pol II, 1.2 μM tsDNA, 1.5 μM ntsDNA, and 1.2 μM RNA. To improve resolution of the pre-catalysis state, 5 μM Elf1 was added to one pre-catalysis sample, while a parallel sample assembled with Pol II alone (without Elf1) was prepared for comparison.

Reactions were initiated by substrate addition. For pre-catalysis samples, 5 mM ATP was added to the fully assembled EC, followed by incubation at room temperature for 10 min prior to cryo-EM grid preparation. For post-catalysis samples, the “substrate gradient grid” method was applied by adding 10 mM NTP to initiate the reaction. Detailed procedures for this method are described in the cryo-EM sample preparation section.

Cryo-EM sample preparation

The optimized protocol for preparing mspSA affinity-enriched grids is outlined as follows21: DOPC (18:1 (Δ9-Cis) PC, Avanti) and biotin-labeled lipid (16:0 Biotinyl Cap PE, Avanti) are mixed in a 9:1 ratio to reach a final concentration of 1 mg/mL, Prepare a streptavidin solution (streptavidin from Streptomyces avidinii, Sigma) at a concentration of 0.01 mg/mL in a buffer consisting of 20 mM HEPES (pH 7.5) and 150 mM NaCl46. Add 30 μL of the streptavidin solution to a Teflon well (on ice), and then carefully lay 0.7 μL of the lipid mixture over the surface of the streptavidin solution. Enclose the Teflon block in a humidity chamber containing a small amount of Milli-Q water at the bottom of the chamber. Seal the lid of the chamber with Vaseline and incubate overnight at room temperature. After 6 hours, carefully place the carbon side of holy carbon grids (Quantifoil copper 2/1, 300 mesh) onto the surface of the Teflon well (without prior glow discharge) for 2 minutes to allow for lipid monolayer transfer. Wash the grid with three drops (120 μL) of sample buffer to remove unbound streptavidin. Add 4 μL of the Pol II EC sample to affinity grids and incubate for approximately 5 minutes. After incubation, remove the excess sample solution using filter paper. Subsequently, plunge-freeze the grids using the Thermo Fisher Vitrobot IV, applying the following parameters: 3.5 seconds blotting time, −15 blotting force, at 4°C and 100% humidity with adding 3.5 μL of sample buffer. The blotting force should be optimized to achieve optimal vitrification results.

To capture the post-catalysis structure of the Pol II EC, the PoI II EC was first applied to the affinity grids and incubated for 10 minutes. Grids were washed with 50 μL fresh buffer, and excess solution was removed using filter paper. The grids were then transferred to Thermo Fisher Vitrobot IV. 3 μL fresh buffer was added to each grid and subsequently 0.2 μL of 10mM NTP substrate was introduced onto the grid at the top side immediately before blotting, generating a “Substrate Gradient Grid (SGG)”. The regions selected for data collection were marked on the grid (Fig. S9).

Data collection and processing

Initial screening to obtain integral grids with well-formed monolayers and particles was conducted using a 200 kV Thermo Scientific Glacios transmission electron microscope. Cryo-EM data were collected using a FEI Titan Krios operating at 300 kV, equipped with a Gatan K3 with GIF Quantum camera (at eBIC, Diamond) or Falcon 4 with GIF Quantum camera (OPIC, STRUBI, University of Oxford). Data acquisition was performed automatically using EPU software, with defocus values ranging from −0.5 to −2.5 μm. The imaging pixel size for data collection was 0.829 Å for pre-catalysis Pol II EC (Bio-Quantum K3), 0.825 Å for post-catalysis Pol II EC (Bio-Quantum K3) and 0.932 Å for pre-catalysis Pol II EC associated with Elf1 (Falcon 4 with SelectrisX), and the total dose applied was 50 electrons per Å2, distributed across 50 frames. A total of 10,599 images were collected for pre-catalysis Pol II EC, 18,913 images for post-catalysis Pol II EC and 16,482 images for pre-catalysis Pol II EC associated with Elf1.

All datasets were processed via CryoSPARC (v4.5.3)38. For pre-catalysis Pol II EC without Elf1, raw micrographs were imported, followed by motion correction and calculation of the contrast transfer function (CTF). The resulting micrographs were manually curated to exclude images with poor quality, (such as CTF fit resolution: <4Å; Relative ice thickness: 1< values < 1.1; Total full-frame motion distance (pixels): < 60 and astigmatism: < 5000). All micrographs were then subjected to automatic particle picking (Blob picking), with particle diameters ranging from 130 Å to 170 Å, to generate initial 2D templates. Following several rounds of 2D classification, particles with good 2D class averages were selected for Topaz training47,48 using the complete dataset. After further 2D classification, high-quality particles with distinct 2D class averages were merged, and duplicate particles were excluded. The output particles were used for ab-initio reconstruction, followed by hetero-refinement. Particles selected from good 3D classes after hetero-refinement were used for Topaz training. After multiple rounds of Topaz training on different particle sets, high-quality particles obtained from both Blob picking and Topaz training were merged to eliminate duplicates. These refined particle sets were then used for ab-initio reconstruction, followed by hetero-refinement. The best 3D class was selected for defocus refinement, global CTF refinement, and non-uniform refinement. To minimize the influence of the dynamic region of the Rbp4/7 stalk and enhance the resolution, a truncated mask without Rbp4/7 was applied for local refinement, resulting in a 2.26Å core structure of the EC complex. To obtain a complete 12-subunit Pol II EC structure, the five flexible regions were subjected to further refinement. The five flexible regions include the nucleic acid scaffold, Rpb4/7, Rpb9, Rpb12/wall, and Jaw/Rpb9. Focused masks were employed for 3D classification processing on the flexible regions, then local or homogeneous refinement jobs were applied to refine each local region. The final complete map was generated by compositing maps of core Pol II and the flexible regions. For the post-catalysis Pol II EC, a processing strategy similar to that used for the pre-catalysis complex was applied. Following motion correction, contrast transfer function (CTF) calculation, particle picking, 2D classification, ab-initio reconstruction, hetero-refinement, and homo-refinement, 3D classification with a solvent mask identified 6 out of 15 classes exhibiting TL unfolding and bond formation. Particles from these classes were selected for further non-uniform refinement, resulting in a final 2.33 Å density map. For pre-catalysis Pol II EC associated with Elf1, the same processing strategy as post-catalysis Pol II EC was used, except that 3D classification with a solvent mask identified 2 out of 15 classes exhibiting TL folding. Particles from these classes were selected for further non-uniform refinement, followed by reference-based motion correction and an additional round of non-uniform refinement, resulting in a 1.96 Å density map.

In all cases, the resolution was determined by gold-standard Fourier shell correlation (FSC). The local resolution estimation was calculated in Chimera40 based on the output maps from CryoSPARC.

Atomic model building, refinement, and validation

The initial models for the 12-subunit S. cerevisiae RNA Pol II were derived from previously published crystal structures (PDB: 2E2H5 and 8U9R12). For Elf1 in the pre-catalysis structure, the initial model was obtained from the published cryo-EM structure (PDB: 8TVY37). The resulting model was rigid-body fitted into the cryo-EM density map using ChimeraX UCSF (v1.8)39. Refinement of the fitted model was conducted in PHENIX (v1.21)41 using the phenix.real_space_refine program. The refined model was subsequently manually adjusted in COOT (v0.9.8)42. Model validation was carried out with MolProbity49. All structural figures were rendered using ChimeraX UCSF (v1.8)39.

Identification of water and ion molecules

Based on previously established methodologies, we developed a systematic workflow to accurately identify and model water and ion molecules in our cryo-EM structures. This workflow comprises three primary steps: candidate peak identification, peak validation, and molecular assignment of water and ions (Fig. S6A). Initially, we segmented the sharpened cryo-EM density maps using the Segger plugin in UCSF Chimera. Potential peaks corresponding to water or ion molecules were identified by the automated peak identification program SWIM43 and additional manual peak selection.

To ensure accuracy and remove spurious noise peaks, we applied two stringent validation thresholds: a minimum density threshold of 4-σ (contour level = 0.17), and a Q-score cutoff of 0.7. The Q-score, a quantitative metric designed to evaluate the local fit of atomic models within cryo-EM density maps, has been previously established and widely utilized in recent high-resolution cryo-EM studies to validate the accurate placement of water and ion molecules15,50,51. To further enhance the robustness of our validation, we employed cross-validation using two independent sharpened half-maps (half-A and half-B) in conjunction with the full map. Each peak was rigorously cross-checked to ensure that its Q-score exceeded 0.7 consistently across all three maps. This cross-validation strategy is conceptually analogous to anomalous signal validation in X-ray crystallography and has recently been successfully applied in cryoEM studies to robustly confirm water and ion placements15,43. For detailed Q-score value of individual water and ion molecules, please refer to Supplementary Spreadsheet 1.

Validated peaks were subsequently evaluated using our molecular assignment scoring system to differentiate between water molecules and specific ion types. This assignment scoring system calculates the probability that each peak corresponds to one of four possible molecular species S{H2O, Na+, K+, Mg2+} (Fig. S6A). Each candidate peak was ultimately assigned the molecular identity corresponding to the highest calculated score. The overall score for each molecular identity was derived from the product of three distinct scoring functions: distance-scoring function (fdistanceS), the coordination-number scoring function (fcountsS), the charge-scoring function (fchargeS), and the clash function (fclashS).

(1). Distance-score fdistanceS

For each candidate site we measure all of its n polar-contact distances (di) to surrounding polar atoms. We then score each distance with a flat-top Gaussian kernel, which gives perfect score inside an “ideal” coordination distance range (μs±hs) and smoothly decays to zero over a tail of length l (default: 0.2 Å). The μs and hs respectively represent the mean distance and half width of coordination distance range for each molecule species (S). The theoretical distance ranges for the distinct molecule coordination were input as: 2.9 ± 0.5 Å for water molecules, 2.0 ± 0.2 Å for Mg2+, 2.5 ± 0.1 Å for Na+, and 2.9 ± 0.3 Å for K+. The scoring function for each distance is defined as:

kis(di)={1,diμShSexp[12(diμShSσ)2],diμS<hS0,diμShS+l}

The Gaussian standard deviation (σ) is chosen so that the kernel falls to 0.5 at the midpoint of the tail:

σ=l22ln2

The average distance-score is:

fdistanceS=1Ni=1Nkis(di)
(2). Contact-counts score fcountsS

We utilized a linear contact-count score to quantify how the total number of polar contacts “nS” at a candidate site agrees with the expected coordination number of each species. For water molecules, the ideal maximum number of polar contacts is four. Accordingly, the scoring function for water assigns a maximum score of 1 when nPwater is less than or equal to 4. If the number of detected polar contacts exceeds four, a penalty is applied to reflect the reduced likelihood of accurate water assignment. The Δ is a soft-cutoff width that governs the rate at which the score decays (default Δ= 8). The complete coordination-scoring function for the water molecule is defined as:

fcountswater(n)={1,1nPwatercmaxwaterΔ,0,0<nPwatercmaxwatercmaxwater<nPwater6nPwater>6}

In contrast to water, ion placement requires a minimum amount of coordinating oxygen to be chemically plausible. Therefore, if the coordination number of oxygen (nOion) falls below a defined minimum threshold (cminion), the candidate site is penalized accordingly. To model this behavior, we applied a scoring function that assigns low values for insufficient coordination and gradually increases as the number of contacts approaches and exceeds cminion. The minimum coordination number cminion was set to four for all tested ion species. The maximum expected coordination number cminion was set to six for Na+ and Mg2+, and eight for K+. The full formulation of the coordination-number scoring function for ion candidates is as follows:

fcountsion(n)={nOioncminion1,0,0<nOioncminioncminion<nOioncmaxionnOion>cmaxion}
(3). Electrostatic - Charge Term fchargeS

The local electrostatic environment is another critical consideration in molecule assignment. We defined the nitrogen from the side chain of residues Arg and Lys as the basic atoms and oxygen from the side chain of residues Asp and Glu, and phosphate groups on nucleic acid backbone as the acidic atoms. The set of contact for charged atoms was denoted as Cpositive and Cnegative. For water molecules, given their ability to form polar interactions with both positively and negatively charged groups, the charge-scoring function always returns a score of “1”. In contrast, metal ions cannot form stable coordination bonds with positively charged groups. Therefore, if a positively charged group is detected within the given distance value of 4 Å, the charge-scoring function scores “0”. In addition, the coordination of Mg2+ ion requires more than one negative charge group. Taken together, the complete charge-scoring function is defined as:

kchargeS={1,S=H2O,{1,S{Na+,K+},Cpositive=00,S{Na+,K+},Cpositive>0}{1,S{Mg2+},Cpositive=0Cnegative>00,S{Mg2+},Cpositive>0Cnegative=0}}
(4). Clash Term fclashS

To avoid potential clash, we used a penalty scoring function for multiple “too-short” contacts. The polar contact within distance of (μshsI) was defined as a “too-short” contact.

nshortS=#{i:di<μShSl}fshortS=cos(min(nshortSα,π2))

Regarding the ion assignment, main-chain nitrogen contacts (atom name “N”) within the filter window are rarely compatible with metal coordination. We count “nNion” of these and apply the same half-cosine form:

fNion=cos(min(nNionα,π2))

The final distance-scoring function is:

fclashwater=fshortwaterfclashion=fshortion×fNion
(5). Combined Score and Classification

For each species (S), we compute the combined assignment score:

ScoreS=fdistancS×fcountsS×fchargeS×fcalshS

The species with the highest ScoreS was assigned to the site. To quantify the assignment’s reliability, we compared the top score (ScoreMaxS) to the second highest (Score2ndS). Specifically, we defined the confident score (CS) as:

CS=log2(ScoreMaxSScore2ndS)

Such that CS ≥ 0.6 (a ratio ≥20.6 ≈ 1.5) indicates a high-confidence assignment. All sites with CS < 0.6 or those assigned as ions were further examined by manual inspection of coordination geometry and local density (Fig. S6D). The expected geometry for water molecules is tetrahedral coordination, whereas for Mg2+, K+, and Na+ ions, we employed the octahedral coordination geometries. For the detailed score and geometry information, please refer to Supplementary Spreadsheet 2.

Identification of positional overlapped water

To evaluate the spatial position of water molecules across catalytic states of RNA polymerase II, we developed a structural alignment and distance analysis workflow. The procedure comprises the following steps:

Local structural alignment: Given the conformational differences between the pre- and post-catalysis structures, local structural alignment was performed prior to water mapping to enable accurate identification of positionally overlapped water molecules. Specifically, three independent local alignments were conducted, each focusing on a distinct subunit—Rpb1, Rpb2, and Rpb3—resulting in three corresponding pairs of locally aligned structures.

Water Molecule Extraction and Mapping: water molecules were extracted from each pair of structures. Each water molecule in the pre-catalysis structure was then spatially compared against all water molecules in the post-catalysis structure. A water pair was considered positionally overlapped if the distance between their oxygen atoms was below a defined threshold (typically 1.5 Å). In cases where a water molecule from one structure had multiple potential overlapping partners in the other, only the closest pair was retained to avoid redundancy. The results from three local-region alignments were compiled into an Excel table (Supplementary Spreadsheet 3), listing individual water molecule identifiers, coordinates, and matched pairs across the two states.

MD simulation of Pol II elongation complex

The starting model is based on our high-resolution cryo-EM structure, including resolved water molecules. Short missing loops in Rpb1, Rpb2, and Rpb8 were modeled using ChimeraX39 and Modeller52. Due to a large unresolved gap in the cryo-EM structure, Rpb4 was modeled as two separate chains. Missing heavy atoms were added using PDBFixer53, Rpb1 residue H1085 was protonated54, and Zn-coordinated cysteines were deprotonated.

For the MD simulation, the Pol II elongation complex was placed in a triclinic box with a 15-Å space between the complex and the edge of the box and solvated with water molecules. Na+ and Cl ions were added to neutralize the system and achieve a final salt concentration of 0.15 mol/L. All MD simulations were performed using the GROMACS 2022.544 simulation package, with the Amber14SB 55 protein force field, OL1556 force field for DNA and OL357 force field for RNA. ATP parameters were taken from a published force field58, and water was modeled using the TIP3P 59 model. Long-range electrostatic interactions were treated using the Particle Mesh Ewald (PME)60 method, and Van der Waals interactions were computed with a 12-Å cutoff.

The system underwent an initial energy minimization using a steepest descent algorithm for 10,000 steps. The LINCS61 algorithm was then applied to constrain bonds involving hydrogen atoms during subsequent steps. A 1-ns NVT MD simulation was performed with position restraints (force constant = 1,000 kJ mol−1 nm−2) on all heavy atoms as well as the oxygen atoms in the pre-placed water molecules from the cryo-EM structure. This was followed by a 1-ns NPT equilibration under the same position restraint settings. A velocity-rescaling62 thermostat (coupling constant = 0.1 ps) was used to maintain a temperature of 300 K, and the Berendsen barostat63 was applied during NPT equilibration, with a reference pressure of 1 bar and coupling constant of 0.5 ps. Subsequently, 25 independent production MD simulations were performed under NVT conditions at 300 K, each initiated with different velocities. During these simulations, heavy atoms (excluding water molecules) were restrained with a moderate restraint force constant of 209.2 kJ/mol·nm2 (0.5 kcal/mol·Å2), as described in a previous report64. Frames were saved at 5-ps intervals. Temperature annealing from 50 K to 300 K was applied during the first 2.5 ns. The first 10 ns of each production run were discarded, and the following 30 ns were used for the final analysis. To ensure a maximum lag time of 10 ns for MSD calculations, the final 10 ns were excluded as starting points, resulting in an effective simulation length of 20 ns. In total, 500 ns of effective simulation time was accumulated across all production runs.

Mean squared displacements (MSD) were calculated using a hydration site approach. For instance, to analyze the diffusion properties of X water, the hydration site was defined as a sphere with a 1.5 Å radius centered on the X position cryo-EM structure. During simulations, any water molecule entering this hydration site was tracked, and its subsequent displacement was included in the MSD calculation for the X water hydration site. Because positional restraints were applied to heavy atoms, structural alignment was not required. The residence time was defined as the maximum duration for which the MSD remained below 10 Å2. The maximum lag time to calculate MSD was 10 ns; any water positions with a residence time exceeding 10 ns were recorded as 10 ns. For bulk water, the MSD exceeded 10 Å2 within even the smallest saving interval (5 ps), which is consistent with the previously reported result65 showing that the diffusion constant for the TIP3P water in bulk is overestimated at approximately 5×105 cm2 s−1. To enable comparison, we set up a pure water system within a cubic box measuring 10 nm per side. Using the same MD procedure, except for adjusting the saving interval to 1 ps, we calculated the residence time for bulk water to be 3 ps.

QUANTIFICATION AND STATISTICAL ANALYSIS

For cryo-EM structure determination, the number of particles contributing to each 3D reconstruction is summarized in Table 1 and detailed in Figures S1, S2, and S9. Resolution estimation was performed using Fourier shell correlation (FSC) analysis in cryoSPARC38, following standard procedures.

To assess the reliability of candidate densities corresponding to water molecules and ions, signal-to-noise ratios were quantified using Q-scores calculated with the SWIM program43 in UCSF Chimera40. All Q-score values are provided in Supplementary Spreadsheet 1. For assignment of solvent species, we developed an assignment scoring scheme (see Method Details) to estimate the likelihood of each water or ion placement; the resulting scores are reported in Supplementary Spreadsheet 2.

To further evaluate the conservation and positional consistency of water molecules across different EC structures, we established a structural alignment and distance analysis workflow. Pairwise distance measurements for overlapping water positions are analyzed in Supplementary Spreadsheet 3.

Supplementary Material

1
2

Movie S1, high-resolution cryo-EM density of Pol II EC revealing order water molecules, related to Figure 1.

Download video file (141.2MB, mp4)
3

Movie S2, water-mediated SN2 chemistry and dynamic reorganization of water networks during Pol II catalysis, related to Figure 5 and 6.

Download video file (75.9MB, mp4)
4

Supplementary spreadsheet 1, Q-score for water and ion molecules, related to STAR Methods.

5

Supplementary spreadsheet 2, assignment score for water and ion molecules, related to STAR Methods .

6

Supplementary spreadsheet 3, positional overlap of water molecules between pre- and post-catalysis states, related to Figure 5 and STAR Methods.

• Document S1. Figures S1–S10, Tables S1–S2, and supplemental references

KEY RESOURCES TABLE

REAGENT or RESOURCE SOURCE IDENTIFIER
Bacterial and virus strains
BL21 (DE3) E. coli NEB Cat#C2527
Chemicals, peptides, and recombinant proteins
ATP Solution (100 mM) Thermo Fisher Scientific Cat#R0441
CTP Solution (100 mM) Thermo Fisher Scientific Cat#R0451
UTP Solution (100 mM) Thermo Fisher Scientific Cat#R0471
GTP Solution (100 mM) Thermo Fisher Scientific Cat#R0461
Biotin-labeled lipid (16:0 Biotinyl Cap PE) Avanti Cat#870277
DOPC (18:1 (Δ9-Cis) PC) Avanti Cat#850375
Streptavidin from Streptomyces avidinii Sigma Cat#S0677
Holy carbon grids (Quantifoil copper 2/1, 300 mesh) Quantifoil Cat#Q350CR1-2nm
Teflon well Ma et al.21 N/A
Deposited data
1.96 Å high-resolution map of the Pol II EC at pre-catalysis state (with Elf1) This study EMD-55239
Local refinement map of the nucleic acid scaffold at pre-catalysis state (with Elf1) This study EMD-55240
2.26 Å high-resolution map of the Pol II EC at pre-catalysis state (Elf1-free) This study EMD-53053
A composite map of the global Pol II EC at pre-catalysis state (Elf1-free) This study EMD-53057
Local refinement map of nucleic acid scaffold at pre-catalysis state (Elf1-free) This study EMD-53064
Local refinement map of Rpb4/7 at pre-catalysis state (Elf1-free) This study EMD-53056
Local refinement map of Rpb9 at pre-catalysis state (Elf1-free) This study EMD-53060
Local refinement map of Rpb12/Wall at pre-catalysis state (Elf1-free) This study EMD-53063
Local refinement map of Jaw/Rpb9 at pre-catalysis state (Elf1-free) This study EMD-53062
2.33 Å high-resolution map of the Pol II EC at post-catalysis state (Elf1-free) This study EMD-54374
1.96 Å high-resolution model of the Pol II EC at pre-catalysis state (with Elf1) This study PDB 9SV6
2.26 Å high-resolution model of the Pol II EC at pre-catalysis state (Elf1-free) This study PDB 9QEB
2.33 Å high-resolution model of the Pol II EC at post-catalysis state (Elf1-free) This study PDB 9RYB
Experimental models: Organisms/strains
S. cerevisiae: Strain background: BJ926 (TAP-tag on Rpb3) Wang et al.5 N/A
Oligonucleotides
Biotinylated template strand DNA (tsDNA): 5′-/BiotinTEG/ TTT TTT GAT ATT TTT GGA TCC CGC TCT GCT CCT TCT CCC ATC CTC TCG ATG GCT ATG AGA TCA ACT AGG AAT TC-3′ IDT N/A
Biotinylated non-template strand DNA (ntsDNA): 5′-/BiotinTEG/ TTT TTA TGT ATT AAT GAA TTC CTA GTT GAT CTC ATA GCC CAT TCC TAC TTG GGA GAA GGA GCA GAG CGG GAT CC-3′ IDT N/A
3′-deoxy RNA (for pre-catalysis): 5′-AUC GAG AG/3′dG IDT N/A
regular RNA (for post-catalysis): 5′-AUC GAG AGG IDT N/A
Recombinant DNA
Plasmid: pGEX-6P-1_Elf1 Sarsam et al.37 N/A
Plasmid: Rpb4/7 Wang et al.5 N/A
Software and algorithms
EPU Thermo Fisher Scientific https://www.thermofisher.com/us/en/home/electron-microscopy/products/software-em-3d-vis/epu-software.html
cryoSPARC (v4.5.3) Punjani et al.38 https://cryosparc.com/
UCSF ChimeraX (v1.8) Goddard et al.39 https://www.rbvi.ucsf.edu/chimerax/
UCSF Chimera (v1.17) Pettersen et al.40 https://www.cgl.ucsf.edu/chimera/
PHENIX (v1.21) Adams et al.41 https://phenix-online.org/
COOT (v0.9.8) Emsley et al.42 https://www2.mrc-lmb.cam.ac.uk/personal/pemsley/coot/
Segger(v2.9.1)/SWIM Zhang et al.43 https://github.com/gregdp/segger/tree/master
GROMACS (v2022.5) Abraham et al.44 https://www.gromacs.org/

Highlights.

  • High-resolution cryo-EM structures reveal functional waters in Pol II elongation.

  • Evolutionarily conserved waters play critical roles in Pol II catalysis.

  • Waters are crucial for substrate recognition and trigger loop assembly.

  • Waters mediate interactions at subunit and protein–nucleic acid interfaces.

ACKNOWLEDGMENTS

This work was supported by grants from the National Institutes of Health (R01 GM102362 and GM148476 to D.W.; R01 GM147652 to X.H. and D.W.; U54 AI170791-7522 and R21 CA280467 to P.Z.). This work was also supported by the UK Wellcome Trust Investigator Award 206422/Z/17/Z (P.Z.), the UK Wellcome Discovery Award 311427/Z/24/Z (P.Z.), and ERC AdG grant 101021133 (P.Z.). We thank Diamond Light source for access and support of the cryo-EM facilities at the UK National Electron Bio-Imaging Centre (eBIC) (proposal NT29812). Further electron microscopy provision was provided through the OPIC electron microscopy facility, a UK Instruct-ERIC Centre, which was founded by a Wellcome JIF award (060208/Z/00/Z) and is supported by a Wellcome equipment grant (093305/Z/10/Z). Computation was performed at the Oxford Biomedical Research Computing (BMRC) facility, a joint development between the Wellcome Centre for Human Genetics (Wellcome Trust Core Award 203141/Z/16/Z) and the Big Data Institute (BDI) supported by Health Data Research UK and the NIHR Oxford Biomedical Research Centre. We thank Dr. Andreas Mueller and Dr. Seth A. Darst for sharing their unpublished cryo-EM data of E. coli RNAP for comparison.

Footnotes

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

RESOURCE AVAILABILITY

Lead contact

Correspondence and requests for materials should be addressed to Dong Wang (dongwang@ucsd.edu).

Materials availability

Materials are available from Dong Wang upon request under a material transfer agreement.

Data and code availability
  • Cryo-EM density maps and corresponding atomic coordinates have been deposited in the Electron Microscopy Data Bank (EMDB) and Protein Data Bank (PDB), respectively. Pre-catalysis (Elf1-bound): 1.96 Å high-resolution map of the Pol II EC in the presence of Elf1 (EMD-55239) and corresponding atomic model (PDB 9SV6). Local refinement map of the nucleic acid scaffold has also been deposited (EMD-55240). Pre-catalysis (Elf1-free): 2.26 Å high-resolution map of the Pol II EC in the absence of Elf1 (EMD-53053) and corresponding atomic model (PDB 9QEB). A composite map of the global Pol II EC is deposited under accession code EMD-53057. Additional maps of flexible regions are available: nucleic acid scaffold (EMD-53064), Rpb4/7 (EMD-53056), Rpb9 (EMD-53060), Rpb12/Wall (EMD-53063), and Jaw/Rpb9 (EMD-53062). Post-catalysis: 2.33 Å high-resolution map of the Pol II EC (EMD-54374) and corresponding atomic model (PDB 9RYB).
  • This paper does not report original code.
  • Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.

DECLARATION OF INTERESTS

The authors declare no competing interests.

REFERENCES

  • 1.Kornberg RD (1999). Eukaryotic transcriptional control. Trends in Cell Biology 9, M46–M49. 10.1016/S0962-8924(99)01679-7. [DOI] [PubMed] [Google Scholar]
  • 2.Sims RJ, Mandal SS, and Reinberg D (2004). Recent highlights of RNA-polymerase-II-mediated transcription. Current Opinion in Cell Biology 16, 263–271. 10.1016/j.ceb.2004.04.004. [DOI] [PubMed] [Google Scholar]
  • 3.Osman S, and Cramer P (2020). Structural Biology of RNA Polymerase II Transcription: 20 Years On. Annual Review of Cell and Developmental Biology 36, 1–34. 10.1146/annurev-cellbio-042020-021954. [DOI] [PubMed] [Google Scholar]
  • 4.Gnatt AL, Cramer P, Fu J, Bushnell DA, and Kornberg RD (2001). Structural Basis of Transcription: An RNA Polymerase II Elongation Complex at 3.3 Å Resolution. Science 292, 1876–1882. 10.1126/science.1059495. [DOI] [PubMed] [Google Scholar]
  • 5.Wang D, Bushnell DA, Westover KD, Kaplan CD, and Kornberg RD (2006). Structural Basis of Transcription: Role of the Trigger Loop in Substrate Specificity and Catalysis. Cell 127, 941–954. 10.1016/j.cell.2006.11.023. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Westover KD, Bushnell DA, and Kornberg RD (2004). Structural Basis of Transcription: Nucleotide Selection by Rotation in the RNA Polymerase II Active Center. Cell 119, 481–489. 10.1016/j.cell.2004.10.016. [DOI] [PubMed] [Google Scholar]
  • 7.Kettenberger H, Armache K-J, and Cramer P (2004). Complete RNA Polymerase II Elongation Complex Structure and Its Interactions with NTP and TFIIS. Molecular Cell 16, 955–965. 10.1016/j.molcel.2004.11.040. [DOI] [PubMed] [Google Scholar]
  • 8.Unarta IC, Goonetilleke EC, Wang D, and Huang X (2023). Nucleotide addition and cleavage by RNA polymerase II: Coordination of two catalytic reactions using a single active site. Journal of Biological Chemistry 299, 102844. 10.1016/j.jbc.2022.102844. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Kaplan CD, Larsson K-M, and Kornberg RD (2008). The RNA Polymerase II Trigger Loop Functions in Substrate Selection and Is Directly Targeted by α-Amanitin. Molecular Cell 30, 547–556. 10.1016/j.molcel.2008.04.023. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Larson MH, Zhou J, Kaplan CD, Palangat M, Kornberg RD, Landick R, and Block SM (2012). Trigger loop dynamics mediate the balance between the transcriptional fidelity and speed of RNA polymerase II. Proceedings of the National Academy of Sciences 109, 6555–6560. 10.1073/pnas.1200939109. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Qiu C, Erinne OC, Dave JM, Cui P, Jin H, Muthukrishnan N, Tang LK, Babu SG, Lam KC, Vandeventer PJ, et al. (2016). High-Resolution Phenotypic Landscape of the RNA Polymerase II Trigger Loop. PLOS Genetics 12, e1006321. 10.1371/journal.pgen.1006321. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Lin G, Barnes CO, Weiss S, Dutagaci B, Qiu C, Feig M, Song J, Lyubimov A, Cohen AE, Kaplan CD, et al. (2024). Structural basis of transcription: RNA polymerase II substrate binding and metal coordination using a free-electron laser. Proceedings of the National Academy of Sciences 121, e2318527121. 10.1073/pnas.2318527121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Sosunov V, Sosunova E, Mustaev A, Bass I, Nikiforov V, and Goldfarb A (2003). Unified two-metal mechanism of RNA synthesis and degradation by RNA polymerase. The EMBO Journal 22, 2234–2244. 10.1093/emboj/cdg193. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Ball P. (2008). Water as an Active Constituent in Cell Biology. Chem. Rev 108, 74–108. 10.1021/cr068037a. [DOI] [PubMed] [Google Scholar]
  • 15.Kretsch RC, Li S, Pintilie G, Palo MZ, Case DA, Das R, Zhang K, and Chiu W (2025). Complex water networks visualized by cryogenic electron microscopy of RNA. Nature, 1–3. 10.1038/s41586-025-08855-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Perillo MA, Burgos I, Clop EM, Sanchez JM, and Nolan V (2023). The role of water in reactions catalysed by hydrolases under conditions of molecular crowding. Biophys Rev 15, 639–660. 10.1007/s12551-023-01104-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Pocker Y. (2000). Water in enzyme reactions: biophysical aspects of hydration-dehydration processes. CMLS, Cell. Mol. Life Sci 57, 1008–1017. 10.1007/PL00000741. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Levy Y, and Onuchic JN (2006). WATER MEDIATION IN PROTEIN FOLDING AND MOLECULAR RECOGNITION. Annual Review of Biophysics 35, 389–415. 10.1146/annurev.biophys.35.040405.102134. [DOI] [PubMed] [Google Scholar]
  • 19.Laage D, Elsaesser T, and Hynes JT (2017). Water Dynamics in the Hydration Shells of Biomolecules. Chem. Rev 117, 10694–10725. 10.1021/acs.chemrev.6b00765. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Lahiri I, Xu J, Han BG, Oh J, Wang D, DiMaio F, and Leschziner AE (2019). 3.1 A structure of yeast RNA polymerase II elongation complex stalled at a cyclobutane pyrimidine dimer lesion solved using streptavidin affinity grids. J Struct Biol 207, 270–278. 10.1016/j.jsb.2019.06.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Ma J, Yi G, Ye M, MacGregor-Chatwin C, Sheng Y, Lu Y, Li M, Li Q, Wang D, Gilbert RJC, et al. (2024). Open architecture of archaea MCM and dsDNA complexes resolved using monodispersed streptavidin affinity CryoEM. Nat Commun 15, 10304. 10.1038/s41467-024-53745-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Vassylyev DG, Vassylyeva MN, Zhang J, Palangat M, Artsimovitch I, and Landick R (2007). Structural basis for substrate loading in bacterial RNA polymerase. Nature 448, 163–168. 10.1038/nature05931. [DOI] [PubMed] [Google Scholar]
  • 23.Abdelkareem M, Saint-André C, Takacs M, Papai G, Crucifix C, Guo X, Ortiz J, and Weixlbaumer A (2019). Structural Basis of Transcription: RNA Polymerase Backtracking and Its Reactivation. Molecular Cell 75, 298–309.e4. 10.1016/j.molcel.2019.04.029. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Li Y, Korolev S, and Waksman G (1998). Crystal structures of open and closed forms of binary and ternary complexes of the large fragment of Thermus aquaticus DNA polymerase I: structural basis for nucleotide incorporation. The EMBO Journal 17, 7514–7525. 10.1093/emboj/17.24.7514. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Wang M, Xia S, Blaha G, Steitz TA, Konigsberg WH, and Wang J (2011). Insights into Base Selectivity from the 1.8 Å Resolution Structure of an RB69 DNA Polymerase Ternary Complex. Biochemistry 50, 581–590. 10.1021/bi101192f. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Biertümpfel C, Zhao Y, Kondo Y, Ramón-Maiques S, Gregory M, Lee JY, Masutani C, Lehmann AR, Hanaoka F, and Yang W (2010). Structure and mechanism of human DNA polymerase η. Nature 465, 1044–1048. 10.1038/nature09196. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Kaplan CD, Jin H, Zhang IL, and Belyanin A (2012). Dissection of Pol II Trigger Loop Function and Pol II Activity–Dependent Control of Start Site Selection In Vivo. PLOS Genetics 8, e1002627. 10.1371/journal.pgen.1002627. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Mazumder A, Lin M, Kapanidis AN, and Ebright RH (2020). Closing and opening of the RNA polymerase trigger loop. Proceedings of the National Academy of Sciences 117, 15642–15649. 10.1073/pnas.1920427117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Xiong Y, and Sundaralingam M (1998). Crystal structure and conformation of a DNA–RNA hybrid duplex with a polypurine RNA strand: d(TTCTTBr5CTTC)–r(GAAGAAGAA). Structure 6, 1493–1501. 10.1016/S0969-2126(98)00148-8. [DOI] [PubMed] [Google Scholar]
  • 30.Gyi JI, Lane AN, Conn GL, and Brown T (1998). The orientation and dynamics of the C2′-OH and hydration of RNA and DNA·RNA hybrids. Nucleic Acids Research 26, 3104–3110. 10.1093/nar/26.13.3104. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Franck JM, Ding Y, Stone K, Qin PZ, and Han S (2015). Anomalously Rapid Hydration Water Diffusion Dynamics Near DNA Surfaces. J. Am. Chem. Soc 137, 12013–12023. 10.1021/jacs.5b05813. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Carvalho ATP, Fernandes PA, and Ramos MJ (2011). The Catalytic Mechanism of RNA Polymerase II. J. Chem. Theory Comput 7, 1177–1188. 10.1021/ct100579w. [DOI] [PubMed] [Google Scholar]
  • 33.Palo MZ, Zhu J, Mishanina TV, and Landick R (2021). Conserved Trigger Loop Histidine of RNA Polymerase II Functions as a Positional Catalyst Primarily through Steric Effects. Biochemistry 60, 3323–3336. 10.1021/acs.biochem.1c00528. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Mishanina TV, Palo MZ, Nayak D, Mooney RA, and Landick R (2017). Trigger loop of RNA polymerase is a positional, not acid–base, catalyst for both transcription and proofreading. Proceedings of the National Academy of Sciences 114, E5103–E5112. 10.1073/pnas.1702383114. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Yang W, Lee JY, and Nowotny M (2006). Making and Breaking Nucleic Acids: Two-Mg2+-Ion Catalysis and Substrate Specificity. Molecular Cell 22, 5–13. 10.1016/j.molcel.2006.03.013. [DOI] [PubMed] [Google Scholar]
  • 36.Nakamura T, Zhao Y, Yamagata Y, Hua Y, and Yang W (2012). Watching DNA polymerase η make a phosphodiester bond. Nature 487, 196–201. 10.1038/nature11181. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Sarsam RD, Xu J, Lahiri I, Gong W, Li Q, Oh J, Zhou Z, Hou P, Chong J, Hao N, et al. (2024). Elf1 promotes Rad26’s interaction with lesion-arrested Pol II for transcription-coupled repair. Proceedings of the National Academy of Sciences 121, e2314245121. 10.1073/pnas.2314245121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Punjani A, Rubinstein JL, Fleet DJ, and Brubaker MA (2017). cryoSPARC: algorithms for rapid unsupervised cryo-EM structure determination. Nat Methods 14, 290–296. 10.1038/nmeth.4169. [DOI] [PubMed] [Google Scholar]
  • 39.Goddard TD, Huang CC, Meng EC, Pettersen EF, Couch GS, Morris JH, and Ferrin TE (2018). UCSF ChimeraX: Meeting modern challenges in visualization and analysis. Protein Science 27, 14–25. 10.1002/pro.3235. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Pettersen EF, Goddard TD, Huang CC, Couch GS, Greenblatt DM, Meng EC, and Ferrin TE (2004). UCSF Chimera—A visualization system for exploratory research and analysis. Journal of Computational Chemistry 25, 1605–1612. 10.1002/jcc.20084. [DOI] [PubMed] [Google Scholar]
  • 41.Adams PD, Afonine PV, Bunkóczi G, Chen VB, Davis IW, Echols N, Headd JJ, Hung L-W, Kapral GJ, Grosse-Kunstleve RW, et al. (2010). PHENIX: a comprehensive Python-based system for macromolecular structure solution. Acta Cryst D 66, 213–221. 10.1107/S0907444909052925. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Emsley P, Lohkamp B, Scott WG, and Cowtan K (2010). Features and development of Coot. Acta Cryst D 66, 486–501. 10.1107/S0907444910007493. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Zhang K, Pintilie GD, Li S, Schmid MF, and Chiu W (2020). Resolving individual atoms of protein complex by cryo-electron microscopy. Cell Res 30, 1136–1139. 10.1038/s41422-020-00432-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Abraham MJ, Murtola T, Schulz R, Páll S, Smith JC, Hess B, and Lindahl E (2015). GROMACS: High performance molecular simulations through multi-level parallelism from laptops to supercomputers. SoftwareX 1–2, 19–25. 10.1016/j.softx.2015.06.001. [DOI] [Google Scholar]
  • 45.Wang L, Zhou Y, Xu L, Xiao R, Lu X, Chen L, Chong J, Li H, He C, Fu X-D, et al. (2015). Molecular basis for 5-carboxycytosine recognition by RNA polymerase II elongation complex. Nature 523, 621–625. 10.1038/nature14482. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Wang F, Liu Y, Yu Z, Li S, Feng S, Cheng Y, and Agard DA (2020). General and robust covalently linked graphene oxide affinity grids for high-resolution cryo-EM. Proceedings of the National Academy of Sciences 117, 24269–24273. 10.1073/pnas.2009707117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Bepler T, Kelley K, Noble AJ, and Berger B (2020). Topaz-Denoise: general deep denoising models for cryoEM and cryoET. Nat Commun 11, 5208. 10.1038/s41467-020-18952-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Bepler T, Morin A, Rapp M, Brasch J, Shapiro L, Noble AJ, and Berger B (2019). Positive-unlabeled convolutional neural networks for particle picking in cryo-electron micrographs. Nat Methods 16, 1153–1160. 10.1038/s41592-019-0575-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Chen VB, Arendall WB, Headd JJ, Keedy DA, Immormino RM, Kapral GJ, Murray LW, Richardson JS, and Richardson DC (2010). MolProbity: all-atom structure validation for macromolecular crystallography. Acta Cryst D 66, 12–21. 10.1107/S0907444909042073. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Pintilie G, Zhang K, Su Z, Li S, Schmid MF, and Chiu W (2020). Measurement of atom resolvability in cryo-EM maps with Q-scores. Nat Methods 17, 328–334. 10.1038/s41592-020-0731-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Pintilie G, Shao C, Wang Z, Hudson BP, Flatt JW, Schmid MF, Morris KL, Burley S, and Chiu W (2025). Q-score as a reliability measure for protein, nucleic acid and small-molecule atomic coordinate models derived from 3DEM maps. Acta Cryst D 81, 410–422. 10.1107/S2059798325005923. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Sali A, and Blundell TL (1993). Comparative protein modelling by satisfaction of spatial restraints. J Mol Biol 234, 779–815. 10.1006/jmbi.1993.1626. [DOI] [PubMed] [Google Scholar]
  • 53.Eastman P, Swails J, Chodera JD, McGibbon RT, Zhao Y, Beauchamp KA, Wang L-P, Simmonett AC, Harrigan MP, Stern CD, et al. (2017). OpenMM 7: Rapid development of high performance algorithms for molecular dynamics. PLOS Computational Biology 13, e1005659. 10.1371/journal.pcbi.1005659. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Huang X, Wang D, Weiss DR, Bushnell DA, Kornberg RD, and Levitt M (2010). RNA polymerase II trigger loop residues stabilize and position the incoming nucleotide triphosphate in transcription. Proceedings of the National Academy of Sciences 107, 15745–15750. 10.1073/pnas.1009898107. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Maier JA, Martinez C, Kasavajhala K, Wickstrom L, Hauser KE, and Simmerling C (2015). ff14SB: Improving the Accuracy of Protein Side Chain and Backbone Parameters from ff99SB. J. Chem. Theory Comput 11, 3696–3713. 10.1021/acs.jctc.5b00255. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Galindo-Murillo R, Robertson JC, Zgarbová M, Šponer J, Otyepka M, Jurečka P, and Cheatham TEI (2016). Assessing the Current State of Amber Force Field Modifications for DNA. J. Chem. Theory Comput 12, 4114–4127. 10.1021/acs.jctc.6b00186. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Zgarbová M, Otyepka M, Šponer J, Mládek A, Banáš P, Cheatham TEI, and Jurečka P (2011). Refinement of the Cornell et al. Nucleic Acids Force Field Based on Reference Quantum Chemical Calculations of Glycosidic Torsion Profiles. J. Chem. Theory Comput 7, 2886–2902. 10.1021/ct200162x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Meagher KL, Redman LT, and Carlson HA (2003). Development of polyphosphate parameters for use with the AMBER force field. Journal of Computational Chemistry 24, 1016–1025. 10.1002/jcc.10262. [DOI] [PubMed] [Google Scholar]
  • 59.Mark P, and Nilsson L (2001). Structure and Dynamics of the TIP3P, SPC, and SPC/E Water Models at 298 K. J. Phys. Chem. A 105, 9954–9960. 10.1021/jp003020w. [DOI] [Google Scholar]
  • 60.Darden T, York D, and Pedersen L (1993). Particle mesh Ewald: An N·log(N) method for Ewald sums in large systems. The Journal of Chemical Physics 98, 10089–10092. 10.1063/1.464397. [DOI] [Google Scholar]
  • 61.Hess B, Bekker H, Berendsen HJC, and Fraaije JGEM (1997). LINCS: A linear constraint solver for molecular simulations. Journal of Computational Chemistry 18, 1463–1472. 10.1002/(SICI)1096-987X(199709)18:12<1463::AID-JCC4>3.0.CO;2-H. [DOI] [Google Scholar]
  • 62.Bussi G, Donadio D, and Parrinello M (2007). Canonical sampling through velocity rescaling. The Journal of Chemical Physics 126, 014101. 10.1063/1.2408420. [DOI] [PubMed] [Google Scholar]
  • 63.Berendsen HJC, Postma JPM, van Gunsteren WF, DiNola A, and Haak JR (1984). Molecular dynamics with coupling to an external bath. The Journal of Chemical Physics 81, 3684–3690. 10.1063/1.448118. [DOI] [Google Scholar]
  • 64.Wall ME, Calabró G, Bayly CI, Mobley DL, and Warren GL (2019). Biomolecular Solvation Structure Revealed by Molecular Dynamics Simulations. J. Am. Chem. Soc 141, 4711–4720. 10.1021/jacs.8b13613. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Kumar H, Dasgupta C, and Maiti PK (2014). Structure, dynamics and thermodynamics of single-file water under confinement: effects of polarizability of water molecules. RSC Adv. 5, 1893–1901. 10.1039/C4RA08730E. [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1
2

Movie S1, high-resolution cryo-EM density of Pol II EC revealing order water molecules, related to Figure 1.

Download video file (141.2MB, mp4)
3

Movie S2, water-mediated SN2 chemistry and dynamic reorganization of water networks during Pol II catalysis, related to Figure 5 and 6.

Download video file (75.9MB, mp4)
4

Supplementary spreadsheet 1, Q-score for water and ion molecules, related to STAR Methods.

5

Supplementary spreadsheet 2, assignment score for water and ion molecules, related to STAR Methods .

6

Supplementary spreadsheet 3, positional overlap of water molecules between pre- and post-catalysis states, related to Figure 5 and STAR Methods.

RESOURCES