Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2019 Mar 28.
Published in final edited form as: Q Rev Biophys. 2018 Jan;51:e8. doi: 10.1017/S0033583518000057

ANTICIPATING INNOVATIONS IN STRUCTURAL BIOLOGY

Helen M Berman 1,*, Catherine L Lawson 1, Brinda Vallat 1, Margaret J Gabanyi 1
PMCID: PMC6438187  NIHMSID: NIHMS996514  PMID: 30912485

Abstract

In this review, we describe how the interplay among science, technology and community interests contributed to the evolution of four structural biology data resources. We present the method by which data deposited by scientists are prepared for worldwide distribution, and argue that data archiving in a trusted repository must be an integral part of any scientific investigation.

1. Introduction

The structural biology community has been uniquely proactive in establishing data resources that archive the results of research and provide services to access and analyze those data. The Protein Data Bank (PDB) was established as a repository for biomacromolecular structural data more than 45 years ago (Protein Data Bank, 1971). It now contains more than 1370,000 structures determined by X-ray crystallography, Nuclear Magnetic Resonance (NMR) spectroscopy, and three-dimensional electron microscopy (3DEM). A diverse community of researchers, students, educators and the general public downloads more than 1.5 million data sets every day. In this review, we demonstrate how the synergies among science, technology and community enabled the PDB to preserve the past while constantly evolving to reflect contemporary needs. We describe how and why two other structural biology data resources were created to supplement and collaborate with the PDB. We conclude by demonstrating how the experiences of the past inform how we are meeting the current challenges presented by the more recent determination of structural models of large macromolecular machines.

2. The synergies of science, technology and community in the development of the PDB

In 1957, the structure of myoglobin was determined (Kendrew et al., 1958), followed shortly thereafter by hemoglobin (Perutz et al., 1960). Thus began the era of structural biology in which, one by one, structures of small proteins including enzymes such as lysozyme (Blake et al., 1965), ribonuclease (Kartha et al., 1967; Wyckoff et al., 1967), and carboxypeptidase (Quiocho & Lipscomb, 1971) were determined using X-ray crystallography. By the late sixties, more than a dozen structures had been determined. In those days, X-ray crystallographic methods involved the use of calculators, newly emerging computers, and manual model building relying on the Richards Box, an optical comparator that had to be housed in a large room (Richards, 1968). A single determination took years of painstaking work. The three-dimensional atomic coordinates obtained from these structure determinations contained a treasure trove of information that would eventually reveal new insights into biology, medicine, biophysics and biochemistry. Indeed, the award of the Nobel Prize to Kendrew and Perutz in 1962 (Nobelprize.org, 2017) recognized not just their achievements, but also the potential of X-ray crystallography. However, for others to help build on that knowledge, it would be necessary to have access to the three-dimensional coordinates produced by all of these new structure determinations.

The coordinate data were stored on punched cards, paper tape and magnetic tape. Because the Internet was only beginning to be established, transfer of data between laboratories involved recording the data onto appropriate media and mailing it. Starting in 1966, a small community of scientists met periodically to discuss how best to archive and distribute these structures. In 1971, a seminal meeting was held in Cold Spring Harbor (Phillips, 1972) in which the practitioners and now pioneers of structural biology described their structures to a rapt and inspired audience. Among the attendees was Walter Hamilton, an energetic and highly respected chemical crystallographer from Brookhaven National Laboratory (BNL). Walter had been collaborating with Edgar Meyer who was creating a Protein Library (Meyer, 1997). When presented with the problem of needing an archive for biomacromolecular structures, Hamilton immediately offered to house one at BNL. He contacted Olga Kennard who was then head of the Cambridge Crystallographic Data Center (CCDC) in Cambridge, UK (Allen et al., 1973) and they agreed to set up the Protein Data Bank (PDB) (Protein Data Bank, 1971) as collaboration between BNL and CCDC. After Hamilton’s death in 1973, Tom Koetzle took over the direction of the PDB. In 1979 there were 53 structures in the PDB (Figure 1), some of which are shown in Figure 2.

Figure 1.

Figure 1.

Growth chart of structures in the PDB with indicators of each decade.

a. The number of structures released per year (blue) and the cumulative number of structures (orange).

b. The same information, using a log scale.

The number of structures released at the end of each decade is shown in black.

Figure 2.

Figure 2.

Examples of structures determined in the 1970’s. Ribbon representations were generated using UCSF Chimera (Pettersen et al., 2004).

a. Myoglobin (Watson, 1969). First protein structure determined using X-ray crystallography.

b. Lysozyme (Blake et al., 1965; Kelly et al., 1979). First enzyme structure determined using X-ray crystallography.

c. Yeast phenylalanine transfer RNA (Rich & Kim, 1978; Robertus et al., 1974). First RNA structure determined using X-ray crystallography.

The 1980’s saw a steady growth of structures in the PDB in large part because of the emergence of powerful new technologies. Genetic engineering made it possible to clone and express large quantities of protein without resorting to extraction from natural biological sources. Chemical synthesis could be used to obtain purified fragments of DNA. The advent of synchrotron sources allowed the collection of data with intense X-ray beams (Harmsen et al., 1976). At the same time, development of the multiple anomalous diffraction phasing method (MAD) (Hendrickson et al., 1985) leveraged the ability to tune the X-ray wavelength using synchrotron radiation. Flash freezing (Hope, 1988) to prevent crystal decay began to be more widely used. Multi-wire detectors made it possible to collect many diffraction reflections at once (Hamlin, 1985). Computing technology continued to improve. In particular, molecular graphics made it possible to fit structural models to electron density (Jones, 1978), replacing the need for the Richards Box. During this period, NMR spectroscopy began to be used for determining the structures of small proteins (Horst et al., 2001), thus eliminating the requirement of crystallinity. During the 1980’s, the first atomic structures of viruses were determined (Erickson et al., 1985; Hopper et al., 1984) as were those of DNA (Dickerson et al., 1982) (Figure 3).

Figure 3.

Figure 3.

Examples of structures determined in the 1980’s.

a. A, B and Z DNA (Dickerson et al., 1982). This representation of the three canonical forms of DNA is taken from the Molecule of the Month.

b. Rhinovirus (Arnold & Rossmann, 1988). This was one of the early virus structures determined using X-ray crystallography. Three unique chains (grey, pink, orange surfaces) are repeated 60-fold to create a virus capsid with icosahedral symmetry.

With the potential of structural biology being realized at an increasing pace, members of the scientific community began to be concerned that valuable data would be lost if deposition of structures into the PDB were not mandatory (Barinaga, 1989). Starting in about 1982, committees were set up to determine exactly which data should be archived. Fred Richards created a petition signed by many of the leading structural biologists, urging deposition into the PDB (Hufton, 2014). In 1989, the International Union of Crystallography (IUCr) published guidelines for the deposition, archival and release of structural data (International Union of Crystallography, 1989). The National Institute of General Medical Sciences (NIGMS) then made a ruling that structure determinations funded by that institute had to be archived by the PDB. In time, virtually all journals required deposition of coordinates in the PDB as a mandatory condition of publication. Another important event in the 1980’s was the inclusion of structural biology as a focus of research by Howard Hughes investigators (Howard Hughes Medical Institute, 2017). By 1989, there were 365 structures in the PDB (Figure 1).

The rate of data deposition rapidly took off in the 1990’s as even better methods for data collection, structure determination and refinement were developed and adopted. Computer performance continued to improve dramatically and structural biologists were more than eager to embrace the new capabilities. During this period, the very first atomic structure determined by electron microscopy methods was deposited into the PDB (Henderson et al., 1990). The 1990’s saw the deposition of many protein nucleic acid complexes into the archive, including the structure of the nucleosome (Luger et al., 1997) (Figure 4). By 1999, there were 10,963 structures in the PDB (Figure 1).

Figure 4.

Figure 4

Examples of structures determined in the 1990’s.

a. The structure of a regulator of transcription called the TATA binding protein bound to DNA. The binding of beta sheets into the minor groove of DNA causes a profound bend in the DNA (Patikoglou et al., 1999).

b. Nucleosome (Luger et al., 1997). The DNA shown in orange wraps around the histone proteins shown in blue, taken from the Molecule of the Month.

When the PDB was first established, the focus was on the collection of the coordinate data as well as some other descriptive data. The PDB Format (Westbrook & Fitzgerald, 2009) was widely adopted because it was simple and ‘human’-readable. However, it was lacking in many other ways: relationships among data items were implicit and not explicit, there was no controlled vocabulary, there were limitations on the number of atoms and residues, and some of the definitions of data items were vague. In 1990, the IUCR set up a Working Group (WG) to create a Macromolecular Crystallographic Information File (mmCIF). It was originally supposed to be a variant of the Crystallographic Information File (CIF) that was already established for small molecules (Hall et al., 1991). The mmCIF WG decided to use the opportunity to not only create richer data content with precise definitions for the macromolecular crystallographic experiment and its results, but also to improve the data representation for PDB entries. A new data model was created that had data type definitions, explicit parent-child relationships among data items, enumerations for controlled vocabulary, and many other features. Workshops were held to obtain community feedback; by 1996, more than three thousand definitions were instantiated into a computer readable dictionary (Fitzgerald et al., 2005). When the PDB moved from management by BNL to the Research Collaboratory for Structural Bioinformatics (RCSB) in 1998, mmCIF became the underlying data model that allowed for the creation of a relational database. However, uptake by the community was slow and it was not until 2011 that mmCIF became the Master Format for the PDB, allowing the PDB Format to be retired. As larger structures of macromolecular assemblies started to be deposited into the PDB, the limitations of the PDB format became more apparent, leading to wider acceptance of the mmCIF format.

The 2000’s saw even more growth in the PDB. Ribosome structures, representing some of the very largest and most complex structures in the PDB, were deposited (Ban et al., 2000; Carter et al., 2000; Schluenzen et al., 2000) (Figure 5). Not surprisingly, the feat of determining these structures led to the award of a Nobel Prize in Chemistry in 2009, shared by three structural biologists. During the same period, the Protein Structure Initiative (PSI) began in which structures were determined on a genomic scale, resulting in nearly 7000 new structures in the PDB. In 2009, there were 61,812 structures in the PDB (Figure 1).

Figure 5.

Figure 5.

Ribosome subunits. The small subunit is shown on the left and the large on the right (Ban et al., 2000; Carter et al., 2000; Schluenzen et al., 2000). The protein is shown in blue and the RNA in orange and yellow. Taken from the Molecule of the Month.

When the PDB was first established, it was international in nature. Under BNL management, only one site curated the data, although there were multiple mirror or distribution sites. After RCSB was awarded the grant to manage the PDB, other sites were eager to become deposition sites. In 2003, three data centers, RCSB PDB in the US, MSD (later PDBe) at the EMBL-EBI, and PDBj in Osaka, established the worldwide Protein Data Bank (wwPDB)(Berman et al., 2003). A formal agreement was created to ensure that all structures curated by the data centers follow the same rules for data processing, and that there would be one archive with identical copies distributed by the wwPDB partners. At the time of this first agreement, compliance was difficult because there were two completely different processing pipelines. To ensure that the curated data files were in fact following the same rules, there were regular exchanges among the wwPDB partners to revalidate the data. The need for a single data processing pipeline became apparent. The project to create OneDep began in 2007; this new pipeline system was put into production in 2014 (Young et al., 2017).

By establishing an international consortium whose goal was to develop and maintain a single, high quality archive, it became possible to remediate existing data to meet more modern standards. One of the most important accomplishments was updating the PDB to use IUPAC nomenclature for standard amino acids and nucleotides (Henrick et al., 2008). Other efforts resulted in an incrementally improved corpus of data. Structures that had been represented in multiple, inconsistent ways, for example peptides and viruses, were corrected, and curation of data going forward was improved (Dutta et al., 2014; Lawson et al., 2008).

During this same era, the requirement for creating more stringent validation criteria emerged from the community. An important milestone was reached in 2008, when all crystallographic depositions were required to be accompanied by structure factors (Wlodawer et al., 2008); in 2010, chemical shifts were required for NMR structures. There was also increasing concern about the possibility that fraudulent structures had become a part of the archive. In 2008, the first of many method-specific wwPDB sponsored Validation Task Forces (VTFs) was set up. The charge to the X-ray VTF was to make recommendations to the wwPDB about validation of structures determined by that method. The X-ray VTF examined all available methods, tested them on the entire archive and reported their findings in a paper published in Structure (Read et al., 2011). Their recommendations became the basis of the wwPDB OneDep Validation module (Gore et al., 2017).

In this section, we have demonstrated how the PDB content and policies have evolved over the last 45 years and how the PDB has been agile in responding to rapid and unexpected scientific advances, technical improvements and strongly held beliefs of many stakeholders. Long before the introduction of the “FAIR” guiding principles (Wilkinson et al., 2016), the PDB archive has been making the results of structural biology investigations Findable, Accessible, Interoperable and Reusable.

3. Structural genomics and the Structural Biology Knowledgebase

PDB contains many related structures, including homologs from different organisms, biomolecular complexes with different ligands, and even systematic small mutations of proteins introduced to investigate the effect on folding and activity; for example, PDB contains 566 structures of Bacteriophage T4 lysozyme variants (Matthews, 1996) and more than 250 structures of small molecule – HIV protease complexes (Wlodawer, 2002). The Protein Structure Initiative (PSI) was launched to enable the determination of unique and diverse structures on a genomic scale (Norvell & Berg, 2007). The first phase focused on determining structures of proteins with extremely low sequence similarity to known structures, with the goal of finding new folds. The second phase focused on biology and linked the high throughput centers with projects on specific biological problems that would benefit from systematic structural approaches. For example, there were substantial gains made in determining structures of previously intractable membrane proteins (Pieper et al., 2013). New high-throughput approaches were developed that allowed for advances in every part of the structure determination pipeline, including methods for producing pure protein samples, robotic crystallization, robotic crystal mounting and positioning and automated structure determination. Counter to some earlier concerns, the quality of the structures improved and the cost per structure determination decreased significantly (Grabowski et al., 2016).

To meet the data management requirements of the PSI project, the Structural Biology Knowledgebase (SBKB) was created in 2008 (Berman et al., 2009; Gabanyi et al., 2011). The SBKB consisted of several modules that addressed the varying needs of the PSI project, described below.

TargetTrack provided information about the status of over 330,000 targets studied by the PSI Centers, including selection rationale, histories of protein production trials, and structure determination and deposition. It also collected and made public more than a thousand protocols routinely used by the centers, with variations noted on a trial-by-trial basis. Sequence-based annotations were also calculated and aggregated into each TargetTrack record. The data collected by TargetTrack were usually the first pieces of information available about a given sequence; to share it in the public domain, not only within the PSI Network, was unprecedented at that time.

A Technology Portal provided reports about the various technologies being developed to enable high-throughput protein production and structure determination (Gifford et al., 2012). Summaries of over 450 novel technologies or protocols, along with their use cases, contact information, and references were collected. Categorization by experimental step enabled researchers to find new ideas for overcoming barriers that they could translate into their own lab.

Biosync (Flippen-Andersen et al., 2010; Kuller et al., 2002) became a module of the SBKB. This data resource collects synchrotron beamline parameters and experimental capabilities, and tracks the number of structures released per facility and beamline.

The Publication Portal tracked PSI publications along with their citations and journal impact factors. To date, 80% of the 2300+ articles published by the PSI have at least 5 citations.

The PSI Materials Repository, collected 90,000+ clones and 120 novel cloning and expression vectors created by the PSI centers and distributed them to researchers all over the world (Seiler et al., 2014).

The Protein Model Portal (PMP) (Bordoli & Schwede, 2012) was created to help researchers locate homology models based on experimentally determined structures, thus further leveraging their impact. Users search the PMP by sequence or UniProt identifier, retrieving a list from among 22.8 million homology models pre-computed by Swiss-Model Repository (Kopp & Schwede, 2004), MODBASE (Pieper et al., 2009), and the modeling groups within the PSI centers, as well as experimental structures from the PDB. A graphical map indicated how much of the sequence was covered by an experimental structure or derived from a model, and quality estimates were provided regarding the reliability of a model. If no model existed, new models could be requested and calculated by 6 public modeling servers. In 2013, the PMP group, with support of the PSI and modeling community, created the Model Archive (Haas & Schwede, 2013). This new archive stores the computational model coordinates and details about assumptions, parameters and constraints applied in modeling. The Model Archive is open to all modelers and provides stable identifiers within publications as well as data storage and access in the public domain. To develop validation criteria for the modeling community, the PMP also constructed the Continuous Automated Model Evaluation (CAMEO) (Haas et al., 2013) server that continuously evaluates the accuracy of predicted models, thus fostering the development of better modelling techniques.

The SBKB website integrated the results of the PSI with over 100 publicly available sequence, structure, function, proteomics and medicine databases. A search for any given protein sequence yielded all relevant annotations or products, presenting a view of what information was known, or still to be discovered. All structures, models, targets, and clones >40% identical in sequence were returned to allow for new connections to be found within the data. If a particular sequence yielded no annotations through the SBKB, users could nominate it for structure determination through the community-nomination portal, where users would be matched to collaborate with a PSI center. As the outreach arm of the PSI project, the SBKB also partnered with the Nature Publishing Group (now Macmillan Group) to write 320 research highlights on PSI advances for the SBKB portal. David Goodsell, author of the RCSB PDB’s Molecule of the Month series (Goodsell et al., 2015), also created 90 illustrated essays of key PSI structures. PSI workshops were also archived on the SBKB.

By mid-2017, the PSI program produced 6,920 structures, contributing over 5% of the current PDB archive (Table 1). Nearly 80% of these entries were distinct from each other and had less than 30% sequence identity to any structure pre-existing in the PDB (Dessailly et al., 2009). 600 structures were motivated by community requests. During PSI:Biology (2010–2015), the 9 membrane protein centers determined 160 structures and developed ~40 novel technologies/methods for this difficult-to-determine class of proteins. Although the PSI program was terminated in 2015, the high throughput methods that enabled its productivity have endured. The SBKB is no longer operational following the end of the PSI program, but some of the modules continue to be available, including Protein Modeling Portal (Haas et al., 2013) and Biosync (Flippen-Andersen et al., 2010). The TargetTrack dataset has been archived (doi: 10.5281/zenodo.821.654).

Table 1.

Summary statistics of structures and other research products produced by the Protein Structure Initiative (PSI), 2000-2017.

Products of the Protein Structure Initiative Total number
PSI structures 6,920
 Distinct structures 5,472
 Community-nominated structures 599
 Membrane Proteins 148
Homology Models 22.8 M
Targets Selected 335,714
Technology Reports 458
Publications 2,313
 Publication with ≥5 Citations 1926
 Citations of PSI publications 117,611
Research Highlights from Nature Publishing 320
Illustrated Featured Molecules/Systems 90

4. Electron Microscopy Data Bank

Bacterial rhodopsin was the first structure determined by electron microscopy deposited into the PDB (Henderson et al., 1990). Because electron crystallography was used, it was possible for the PDB to curate the entry using a variation of the procedure for structures determined by X-ray crystallography. The determination of structures by cryo electron microscopy (3DEM) became popular in the 2000’s as software for reconstruction of 3D density maps from 2D single particle images became available, even though the level of detail produced was typically limited (Chiu et al., 2005). 3DEM scientists began to determine the overall shapes of large macromolecular complexes that could not be crystallized, opening up an important new avenue for structural biology investigations. The maps derived from 3DEM experiments could frequently be fitted with structures derived from X-ray crystallography, NMR spectroscopy or homology modeling, yielding “pseudo-atomic” models that were able to provide useful insights and leads for further research (Rossmann et al., 2005).

In 2002, a new data archive called EM Data Bank (EMDB) containing maps and metadata was established at the EMBL-EBI (Editorial, 2003; Henrick et al., 2003). Structures determined by 3DEM methods began to be deposited with maps archived in EMDB and models separately archived in PDB. An initial dictionary of data terms to describe 3DEM experiments was drafted jointly by the groups at EBI and RCSB, based on requirements provided by the 3DEM community in a series of international workshops. In 2006, the EBI and RCSB groups joined forces with Wah Chiu at the National Center for Macromolecular Imaging (NCMI) to create a “one stop shop” for deposition and retrieval of maps and models at EMDataBank.org (Lawson et al., 2011). Both groups launched “serial” map+model deposition and annotation systems that directed users first to deposit their maps to EMDB using EmDep (Henrick et al., 2003) and second to deposit their models to PDB with transfer of relevant experimental metadata, as defined in the 3DEM data dictionary. The serial systems worked remarkably well, even though the underlying coordinate deposition and processing systems at the two sites were substantially different (Section 2). Over a nine-year period (2008–2015), nearly 4000 3DEM maps and 1000 3DEM models were processed in this manner. Truly joint map+model deposition for 3DEM structures was instantiated in 2016 with the OneDep system recently implemented by the wwPDB (Young et al., 2017).

There has been substantial growth in 3DEM derived structures over the past few years (Figure 6). Major technological advances, including introduction of the direct electron detector and better data processing methods, have enabled the determination of structures derived from 2D single particle images to near-atomic resolution, making it increasingly possible to visualize amino acid sidechains and nucleotide bases (Vinothkumar & Henderson, 2016). The award of the 2017 Nobel Prize in Chemistry to 3DEM pioneers Henderson, Frank, and Dubochet recognized the potential of this rapidly evolving method to contribute to structural biology. Figure 7 provides several examples of maps deposited into EMDB just in the past year, each with a reported resolution of 4.5 Å or better.

Figure 6.

Figure 6.

Cumulative growth of 3DEM Structures. The number of structures available in EMDB for each recent year is indicated in purple (resolution better than 5.0 Å, dark purple); the number of EM-derived models available in PDB is indicated in green (resolution better than 5.0 Å, dark green).

Figure 7.

Figure 7.

Sampling of 3DEM structures recently released in EMDB: a. GroEL (Roh et al., 2017), b. DNA Protein Kinase (Sharif et al., 2017), c. Heterotrimeric Gs protein complex (Liang et al., 2017), d. Glutamate A2 receptor (Twomey et al., 2017), e. Spliceosome (Wan et al., 2016), f. Rhinovirus/Fab (Dong et al., 2017).

The deluge of high-resolution 3DEM structures has made it a priority to establish robust validation methods for 3DEM derived maps and models. With OneDep now providing the facilities for 3DEM deposition, the current focus of EMDataBank.org is on enabling development of validation methods for 3DEM.

5. The current PDB pipeline

The PDB is responsible for collecting data entries from structural biologists and distributing curated data entries to users. To accomplish this goal, it is necessary to implement a data management pipeline with components for data deposition, curation, validation, archiving and distribution. Over time, data management has changed. Next, we describe current practices in the PDB data management pipeline (Figure 8).

Figure 8.

Figure 8.

The Data Processing Pipeline, from Data Creation through Distribution. Each component of the PDB pipeline is described in Section 5.

Requirements:

In addition to the atomic coordinates, a considerable body of metadata is collected to describe how the coordinates were derived. The metadata are based on the details of each experimental method currently supported by the PDB: X-ray crystallography, nuclear magnetic resonance spectroscopy, and electron microscopy. Table 2 provides a summary of the various aspects of each method that need to be considered for data deposition.

Table 2.

Experimental metadata requirements for the methods currently supported by the PDB.

Method X-ray Crystallography NMR Spectroscopy 3D Electron Microscopy
Sample • Buffer
• Crystallization Procedure
• Buffer
• Isotope Labeling
• Buffer
• Sample Support
• Vitrification
Experiment • Sample Conditions
• X-ray Source
• Detector
• Collection Protocol
• Sample Conditions
• Spectrometer
• Acquisition Parameters
• Sample Conditions
• Electron Source
• Detector
• Imaging Parameters
Measurements/Data • Diffraction Images
• Structure Factors
• Processing Software
• Statistics (resolution, Rsym)
• Resonance Spectra
• Resonance Assignments
• Chemical Shifts
• Contraints
• Processing Protocol
• Particle Images
• Final 3D Map
• Processing Software
• Processing Protocol
• Resolution (FSC)
Structure Modeling • Structure solution method
• Refinement software, restraints
• Fit-to-data statistics
• Structure calculation method
• Refinement software
• Modeling method
• Model source
• Fitting software

Decisions about which data items must be collected are made in consultation with the community via the respective wwPDB Task Forces. Because the science, technology development and community sentiment change over time, the scope and level of granularity of the data to be collected also change over time. It is notable that protein production procedures are currently not collected. The PSI did in fact have procedures in place for collecting protein production protocols through TargetTrack (see Section 3 above). However, compliance from the community was poor, which suggests that the time was not right for collecting and archiving protein production data.

Standards:

To make the PDB archive computer searchable, it is essential that there are clear definitions for each data item collected. The PDBx/mmCIF format that is entirely computer readable is now the PDB Master Format. The data dictionary contains the definitions for all of the methods currently supported by the PDB (mmcif.wwpdb.org). The dictionary is extensible and allows for changes in existing methods and inclusion of new methods. A standing committee reviews the changing requirements and when necessary adds new definitions. In anticipation of changing needs, the dictionary also contains definitions for data items not currently in the PDB archive.

Data Curation:

All PDB entries are extensively curated. Many different aspects of the structure are carefully checked using a modular series of computational tools. For the polymer sequence, the following tasks are performed: cross checks of author-provided sample sequence and coordinate sequence versus the sequence database, cross checks of author-provided source organism versus the taxonomy database, assignments of database references and taxonomy identifiers to modeled protein polymer entities, and annotation of sequence discrepancies between sample sequence and database reference. For ligands, a search is performed to determine whether the ligand geometry is novel or equivalent to one of the ligands found in existing PDB entries. The ligand geometry is checked using a variety of 2D and 3D views. Derived data including the biological assembly are determined.

Data Validation:

Data in the PDB are validated according to recommendations made by Validation Task Forces that are convened by the wwPDB. Because X-ray crystallography is the oldest method supported by the PDB, its community has had the time and experience to develop the most extensive validation procedures (Read et al., 2011). The wwPDB has implemented the recommendations of the X-ray VTF directly into the data processing pipeline. Covalent geometry is checked against established standards. Intermolecular and intramolecular geometries of the polymer chains are checked for clashes using Molprobity (Chen et al., 2010). The geometry of ligands is checked against standards derived from small molecule structures archived in the CCDC (Bruno et al., 2004). The deposition of structure factors allows the checking of real space R factors for each residue and each ligand. A Validation Report is produced with the detailed analysis of the geometrical features of the model as well as the fit of the structure to the underlying experimental data. Graphical representation in the form of sliders gives a summary of the quality of the structure.

Validation of NMR derived structures follows the recommendation of the NMR VTF (Montelione et al., 2013). The model geometry is checked in the same way as for X-ray derived structures. Consistency checks across models are carried out for NMR structures along with examination of outliers in NMR restraints. For 3DEM-derived structures, the 3DEM Validation Task Force recommended that the validation of model geometry follow the same criteria developed for X-ray derived structures and that new methods be developed for 3DEM map validation and map-to-model fit (Henderson et al., 2012). One of the ways to achieve this goal involves engaging the community in EM Challenges (Lawson et al., 2016), where participants attempt to fit models to benchmarked maps, followed by assessment of the results. These exercises are likely to result in more robust methods for validating 3DEM structural models.

To enable efficient data deposition, curation and processing, a new tool called OneDep was developed by the wwPDB (Young et al., 2017) (Figure 9). OneDep has a Deposition and Annotation Workflow system containing the modules required for making data curation as thorough and automatic as possible. Skilled wwPDB biocurators review all of the results of data processing and work with the depositors to ensure the best possible representation of the submitted data.

Figure 9.

Figure 9.

OneDep System. Deposition is provided for X-ray, 3DEM and NMR. The annotation pipeline is made up of several modules that check the chemistry of the components, add new annotations and validate the structural model against standard geometries and the experimental data.

Archiving:

Once the data are processed, the files are put into a temporary archive until they are ready for release, usually upon publication of the structure. The released structures reside in the PDB Archive, which can be accessed using methods such as the File Transport Protocol (FTP) and rsync. The PDB Archive consists of flat files that contain several types of data, including atomic coordinates, molecular description of macromolecules and ligands, metadata describing the experimental method, and experimental data including structure factors, chemical shifts and restraints. 3DEM map data are curated by EMDB partner sites and archived under a separate, parallel branch of the archive. The PDB Archive is mirrored by all three wwPDB partners.

Data Distribution:

The PDB is distributed in several ways. Data can be downloaded via rsync or ftp protocols following the directions provided on the wwPDB website (https://www.wwpdb.org/download/downloads). In addition, each of the wwPDB data centers has websites that provide a multitude of services including downloading, searching and browsing (Berman et al., 2000; Kinjo et al., 2017; Rose et al., 2017; Ulrich et al., 2008; Velankar et al., 2016). Coordinate sets are currently downloaded from the wwPDB FTP and websites more than 550,000,000 times per year.

6. The future: Integrative hybrid (I/H) methods

Traditionally, each PDB entry contains an atomic structural model derived from a single structure determination method, including X-ray crystallography, NMR spectroscopy and 3D electron microscopy. Recently, integrative/hybrid (I/H) methods have been developed that simultaneously use data from multiple experimental techniques to compute structures of single macromolecules or macromolecular complexes (Ward et al., 2013). In some cases, data from a primary method such as NMR is combined with additional information obtained from a secondary method such as small-angle solution scattering (SAS). In other cases, information from multiple experimental sources, such as Fluorescence resonance energy transfer (FRET), SAS, chemical crosslinking (CX) and mass spectrometry (MS) are pooled together to derive a set of spatial restraints that enable computation of a structural model. Combining multiple complementary experimental methods makes it possible to determine structures of large macromolecular machines that have previously eluded traditional structure determination methods. I/H methods have led to the elucidation of structures of macromolecular assemblies such as the nuclear pore complex (Alber et al., 2007a; Alber et al., 2007b) and its sub-complexes (Figure 10, (Kim et al., 2014; Shi et al., 2014)), the type III secretion system needle (Loquet et al., 2012), the proteasomal lid complex (Politis et al., 2014), the exosome complex (Shi et al., 2015) and the mediator complex (Robinson et al., 2015). Although many important structures have been determined using I/H methods, there are no standard mechanisms to archive these structures and make them available to the public. An important distinction between structural models obtained through I/H methods and the atomistic models currently archived in the PDB is that I/H models are often coarse-grained. The existing PDB data pipeline expects fully atomistic models and hence cannot process coarse-grained I/H models.

Figure 10.

Figure 10.

I/H model of the Nup84 sub-complex from the Nuclear Pore Complex (Shi et al., 2014) available from PDB-Dev (Burley et al., 2017; Vallat et al., 2016c). Multi-scale structural model of the heptameric Nup84 sub-complex is shown (colored ribbons and spheres) along with the localization densities of the sampled structures (colored contoured surfaces). The model is obtained using the Integrative Modeling Platform (IMP) software (Russel et al., 2012) and visualized using the ChimeraX software (Goddard et al., 2018).

In 2014, thirty eight experimental and computational scientists assembled at the EMBL-EBI to discuss how best to archive the results of I/H structure determinations. The wwPDB I/H methods Task Force (I/HTF) made the following series of recommendations that would enable the wwPDB to address this problem (Sali et al., 2015): (1) a flexible model representation should be developed, allowing for multi-scale models (with atomistic and non-atomistic coarse-grained representations), multi-state models (existing in various conformations), ensembles of models, and models related by time or other order; (2) procedures for estimating the uncertainty of integrative models should be developed, validated, and adopted; (3) all relevant experimental data and metadata as well as experimental and computational protocols should be archived; (4) a Federation of model and data archives should be created; and (5) publication standards for integrative models should be established.

To address these recommendations, two subgroups of the I/HTF have been established: the Model Validation Subgroup and the Federation Subgroup. The concept of a Federation of model and data repositories would allow individual disciplines to create appropriate repositories for their experimental data based on the requirements of their communities. Mechanisms for data exchange would promote seamless interoperation among the federated repositories (Figure 11).

Figure 11.

Figure 11.

Conceptual diagram of the I/H Methods Federation. At the center are the three structural biology model repositories: the PDB archives experimentally determined structures of macromolecules (Berman et al., 2000); the Model Archive (MA), part of the Protein Model Portal (PMP), archives in silico structural models (Bordoli & Schwede, 2012; Haas et al., 2013; Haas & Schwede, 2013); and PDB-Development (PDB-Dev) is the prototype system for archiving I/H models (Burley et al., 2017; Vallat et al., 2016c). The outer circle consists of experimental data repositories that contribute to structural biology. Only a limited set of experimental data archives have been identified at present and many others may be included as the field evolves and the respective research communities build their own repositories.

Following the recommendations of the I/HTF, a preliminary dictionary has been created to address the flexible data representation required to describe I/H results (Berman et al., 2016; Vallat et al., 2016a; Vallat et al., 2016b; Vallat et al., 2017). This dictionary is a modular extension of the PDBx/mmCIF dictionary (Fitzgerald et al., 2005) used by the PDB archive and contains data definitions necessary to describe the details of I/H models, associated spatial restraints and modeling protocols. The newly developed I/H methods extension dictionary provides the fundamental data specifications required for archiving I/H models. Based on this dictionary extension, a prototype pipeline called PDB-Development (PDB-Dev; pdb-dev.wwpdb.org) has been built to enable testing and development of deposition and archiving for I/H structural models (Burley et al., 2017; Vallat et al., 2016c). Nine I/H models obtained using different modeling software such as the Integrative Modeling Platform (IMP) (Russel et al., 2012), Rosetta (Leaver-Fay et al., 2011), HADDOCK (Dominguez et al., 2003), TADbit (Serra et al., 2017) and XPLOR-NIH (Schwieters et al., 2018) have been deposited into PDB-Dev in a format compliant with the I/H methods dictionary. These include the Nup84 sub-complex of the nuclear pore complex (Shi et al., 2014), the exosome complex (Shi et al., 2015), the mediator complex (Robinson et al., 2015), lysine-linked Diubiquitin complex (Liu et al., 2018), structures of the human serum albumin domains in their native environment (Belsom et al., 2016), the chromatin model of the first 4.5Mb of chromosome 2L from Drosophila Melanogaster (Trussart et al., 2015) and the ribosomal RNA small subunit methyltransferase A complexed with 16S ribosomal RNA (van Zundert et al., 2015). These structures are now publicly available from the PDB-Dev website (Burley et al., 2017; Vallat et al., 2016c) and can be downloaded and visualized using the ChimeraX software (Goddard et al., 2018) (Figure 10).

The lessons learned from creating and maintaining the PDB archive are informing the process of developing the PDB-Dev system for archiving I/H structures. To adapt to the evolving needs of the scientific community, many important tasks have been accomplished: consulting with the community to determine requirements, carefully creating standard dictionary definitions and making sure that those dictionary standards are extensible. Once the PDB-Dev system is fully developed, it will be straightforward to include structures derived from I/H methods in the PDB archive, thus making the rich content from structures of complex macromolecular machines available to PDB users.

7. Conclusion

In this review, we describe the interplay among science, technology and community in creating data resources. The way in which the PDB developed in many ways follows the principles set forth by Elinor Ostrom for the management of natural resources (Ostrom, 1990). Those principles emphasize that bottom-up collective action can work better than top-down enforcement. Although building a community resource in this way can take much longer, the involvement of the various stakeholders in meaningful ways can better ensure its sustainability.

Domain repositories such as the PDB are key to the conduct of science and development of scientific knowledge. Preserving the data and making it freely available enables reproducibility and the ability to build on previous work to carry out new research. Structural biologists were early adopters of the concept of archiving as being an integral part of the research and publication life cycle. Not only has the availability of data helped enable further discoveries in the field, but it also has allowed computational biologists to analyze the entire corpus of data to understand the underlying principles that govern protein folding and interactions; it is impossible to imagine structural bioinformatics without the PDB. The PDB thus provides a compelling roadmap that could be applied to all of science.

Acknowledgements

We thank the wwPDB data center members, the EMDataBank group, and the SBKB partners, with special thanks to Stephen Burley and Wah Chiu for their leadership of the PDB and EM projects and John Westbrook for his vision with respect to the mmCIF effort.

This work has been supported by grants to the RCSB PDB from NSF, NIH and DOE (DBI-1338415), SBKB (U01 GM093324), EMDataBank (R01 GM079429), I/H methods (NSF EAGER award DBI-1519158), and the Enabling Data Science in Biology BD2K curriculum development project (R25 LM012286).

REFERENCES

  1. ALBER F, DOKUDOVSKAYA S, VEENHOFF L, ZHANG W, KIPPER J, DEVOS D, SUPRAPTO A, KARNI-SCHMIDT O, WILLIAMS R, CHAIT B, ROUT M & SALI A (2007a). Determining the architectures of macromolecular assemblies. Nature, 450(7170), 683–694. [DOI] [PubMed] [Google Scholar]
  2. ALBER F, DOKUDOVSKAYA S, VEENHOFF L, ZHANG W, KIPPER J, DEVOS D, SUPRAPTO A, KARNI-SCHMIDT O, WILLIAMS R, CHAIT B, SALI A & ROUT M (2007b). The molecular architecture of the nuclear pore complex. Nature, 450(7170), 695–701. [DOI] [PubMed] [Google Scholar]
  3. ALLEN FH, KENNARD O, MOTHERWELL WDS, TOWN WG & WATSON DG (1973). Cambridge Crystallographic Data Centre. II. Structural Data File. J. Chem. Doc, 13, 119–123. [Google Scholar]
  4. ARNOLD E & ROSSMANN MG (1988). The use of molecular-replacement phases for the refinement of the human rhinovirus 14 structure. Acta Crystallographica Section A, 44 (Pt 3), 270–282. [DOI] [PubMed] [Google Scholar]
  5. BAN N, NISSEN P, HANSEN J, MOORE PB & STEITZ TA (2000). The complete atomic structure of the large ribosomal subunit at a 2.4 Å resolution. Science, 289, 905–920. [DOI] [PubMed] [Google Scholar]
  6. BARINAGA M (1989). The missing crystallography data. Science, 245(4923), 1179–1181. [DOI] [PubMed] [Google Scholar]
  7. BELSOM A, SCHNEIDER M, FISCHER L, BROCK O & RAPPSILBER J (2016). Serum Albumin Domain Structures in Human Blood Serum by Mass Spectrometry and Computational Biology. Mol Cell Proteomics, 15(3), 1105–1116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. BERMAN HM, HENRICK K & NAKAMURA H (2003). Announcing the worldwide Protein Data Bank. Nat Struct Biol, 10(12), 980. [DOI] [PubMed] [Google Scholar]
  9. BERMAN HM, WESTBROOK J, FENG Z, GILLILAND G, BHAT TN, WEISSIG H, SHINDYALOV IN & BOURNE PE (2000). The Protein Data Bank. Nucleic Acids Res, 28(1), 235–242. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. BERMAN HM, WESTBROOK J, VALLAT B, WEBB B & SALI A (2016). A Data Dictionary For Archiving Integrative/Hybrid Models. In 66th Annual Meeting of the American Crystallographic Association, pp. 85–SA, Denver, CO, USA. [Google Scholar]
  11. BERMAN HM, WESTBROOK JD, GABANYI MJ, TAO W, SHAH R, KOURANOV A, SCHWEDE T, ARNOLD K, KIEFER F, BORDOLI L, KOPP J, PODVINEC M, ADAMS PD, CARTER LG, MINOR W, NAIR R & LA BAER J (2009). The protein structure initiative structural genomics knowledgebase. Nucleic Acids Res, 37(Database issue), D365–368. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. BLAKE CCF, KOENIG DF, MAIR GA, NORTH ACT, PHILLIPS DC & SARMA VR (1965). Structure of hen egg-white lysozyme. A three dimensional Fourier synthesis at 2 Å resolution. Nature, 206, 757–761. [DOI] [PubMed] [Google Scholar]
  13. BORDOLI L & SCHWEDE T (2012). Automated protein structure modeling with SWISS-MODEL Workspace and the Protein Model Portal. Methods in molecular biology, 857, 107–136. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. BRUNO IJ, COLE JC, KESSLER M, LUO J, MOTHERWELL WD, PURKIS LH, SMITH BR, TAYLOR R, COOPER RI, HARRIS SE & ORPEN AG (2004). Retrieval of crystallographically-derived molecular geometry information. J Chem Inf Comput Sci, 44(6), 2133–2144. [DOI] [PubMed] [Google Scholar]
  15. BURLEY SK, KURISU G, MARKLEY JL, NAKAMURA H, VELANKAR S, BERMAN HM, SALI A, SCHWEDE T & TREWHELLA J (2017). PDB-Dev: A prototype system for depositing integrative/hybrid structural models. Structure, 25, 1317–1318. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. CARTER AP, CLEMONS WM, BRODERSEN DE, MORGAN-WARREN RJ, WIMBERLY BT & RAMAKRISHNAN V (2000). Functional insights from the structure of the 30S ribosomal subunit and its interactions with antibiotics. Nature, 407, 340–348. [DOI] [PubMed] [Google Scholar]
  17. CHEN VB, ARENDALL WB 3RD, HEADD JJ, KEEDY DA, IMMORMINO RM, KAPRAL GJ, MURRAY LW, RICHARDSON JS & RICHARDSON DC (2010). MolProbity: all-atom structure validation for macromolecular crystallography. Acta Crystallogr D Biol Crystallogr, 66(Pt 1), 12–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. CHIU W, BAKER ML, JIANG W, DOUGHERTY M & SCHMID MF (2005). Electron cryomicroscopy of biological machines at subnanometer resolution. Structure, 13(3), 363–372. [DOI] [PubMed] [Google Scholar]
  19. DESSAILLY BH, NAIR R, JAROSZEWSKI L, FAJARDO JE, KOURANOV A, LEE D, FISER A, GODZIK A, ROST B & ORENGO C (2009). PSI-2: Structural Genomics to Cover Protein Domain Family Space. Structure, 17(6), 869–881. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. DICKERSON RE, DREW HR, CONNER BN, WING RM, FRATINI AV & KOPKA ML (1982). The Anatomy of a-DNA, B-DNA, and Z-DNA. Science, 216(4545), 475–485. [DOI] [PubMed] [Google Scholar]
  21. DOMINGUEZ C, BOELENS R & BONVIN AM (2003). HADDOCK: a protein-protein docking approach based on biochemical or biophysical information. J Am Chem Soc, 125(7), 1731–1737. [DOI] [PubMed] [Google Scholar]
  22. DONG Y, LIU Y, JIANG W, SMITH TJ, XU Z & ROSSMANN MG (2017). Antibody-induced uncoating of human rhinovirus B14. Proc Natl Acad Sci U S A, 114(30), 8017–8022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. DUTTA S, DIMITROPOULOS D, FENG Z, PERSIKOVA I, SEN S, SHAO C, WESTBROOK J, YOUNG J, ZHURAVLEVA MA, KLEYWEGT GJ & BERMAN HM (2014). Improving the representation of peptide-like inhibitor and antibiotic molecules in the Protein Data Bank. Biopolymers, 101(6), 659–668. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. EDITORIAL. (2003). A database for ‘em. Nat Struct Biol, 10(5), 313. [DOI] [PubMed] [Google Scholar]
  25. ERICKSON JW, SILVA AM, MURTHY MR, FITA I & ROSSMANN MG (1985). The structure of a T = 1 icosahedral empty particle from southern bean mosaic virus. Science, 229(4714), 625–629. [DOI] [PubMed] [Google Scholar]
  26. FITZGERALD PMD, WESTBROOK JD, BOURNE PE, MCMAHON B, WATENPAUGH KD & BERMAN HM (2005). 4.5 Macromolecular dictionary (mmCIF). In International Tables for Crystallography G Definition and exchange of crystallographic data eds. Hall SR and McMahon B), pp. 295–443. Dordrecht, The Netherlands: Springer. [Google Scholar]
  27. FLIPPEN-ANDERSEN J, GABANYI MJ, CHEN L, SALA R, WESTBROOK JD & BERMAN HM (2010). BioSync: A Structural Biologist’s Guide to High Energy Data Collection Facilities, vol. 2017. [Google Scholar]
  28. GABANYI MJ, ADAMS PD, ARNOLD K, BORDOLI L, CARTER LG, FLIPPEN-ANDERSEN J, GIFFORD L, HAAS J, KOURANOV A, MCLAUGHLIN WA, MICALLEF DI, MINOR W, SHAH R, SCHWEDE T, TAO YP, WESTBROOK JD, ZIMMERMAN M & BERMAN HM (2011). The Structural Biology Knowledgebase: a portal to protein structures, sequences, functions, and methods. J Struct Funct Genomics, 12(2), 45–54. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. GIFFORD LK, CARTER LG, GABANYI MJ, BERMAN HM & ADAMS PD (2012). The Protein Structure Initiative Structural Biology Knowledgebase Technology Portal: a structural biology web resource. Journal of Structural and Functional Genomics, 13(2), 57–62. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. GODDARD TD, HUANG CC, MENG EC, PETTERSEN EF, COUCH GS, MORRIS JH & FERRIN TE (2018). UCSF ChimeraX: Meeting modern challenges in visualization and analysis. Protein Sci, 27(1), 14–25. [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. GOODSELL DS, DUTTA S, ZARDECKI C, VOIGT M, BERMAN HM & BURLEY SK (2015). The RCSB PDB “Molecule of the Month”: Inspiring a Molecular View of Biology. PLoS Biol, 13(5), e1002140. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. GORE S, SANZ GARCIA E, HENDRICKX PMS, GUTMANAS A, WESTBROOK JD, YANG H, FENG Z, BASKARAN K, BERRISFORD JM, HUDSON BP, IKEGAWA Y, KOBAYASHI N, LAWSON CL, MADING S, MAK L, MUKHOPADHYAY A, OLDFIELD TJ, PATWARDHAN A, PEISACH E, SAHNI G, SEKHARAN MR, SEN S, SHAO C, SMART OS, ULRICH EL, YAMASHITA R, QUESADA M, YOUNG JY, NAKAMURA H, MARKLEY JL, BERMAN HM, BURLEY SK, VELANKAR S & KLEYWEGT GJ (2017). Validation of the Structures in the Protein Data Bank. Structure, 25, 1916–1927. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. GRABOWSKI M, NIEDZIALKOWSKA E, ZIMMERMAN MD & MINOR W (2016). The impact of structural genomics: the first quindecennial. J Struct Funct Genomics, 17(1), 1–16. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. HAAS J, ROTH S, ARNOLD K, KIEFER F, SCHMIDT T, BORDOLI L & SCHWEDE T (2013). The Protein Model Portal--a comprehensive resource for protein structure and model information. Database (Oxford), 2013, bat031. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. HAAS J & SCHWEDE T (2013). Model Archive, vol. 2016. [Google Scholar]
  36. HALL SR, ALLEN FH & BROWN ID (1991). The Crystallographic Information File (Cif) - a New Standard Archive File for Crystallography. Acta Crystallographica Section A, 47, 655–685. [Google Scholar]
  37. HAMLIN RC (1985). Multiwire area x-ray diffractometers. Meth. Enzymol, 114, 416–452. [DOI] [PubMed] [Google Scholar]
  38. HARMSEN A, LEBERMAN R & SCHULZ GE (1976). Comparison of protein crystal diffraction patterns and absolute intensities from synchrotron and conventional x-ray sources. J Mol Biol, 104(1), 311–314. [DOI] [PubMed] [Google Scholar]
  39. HENDERSON R, BALDWIN JM, CESKA TA, ZEMLIN F, BECKMANN E & DOWNING KH (1990). Model for the structure of bacteriorhodopsin based on high-resolution electron cryo-microscopy. J Mol Biol, 213(4), 899–929. [DOI] [PubMed] [Google Scholar]
  40. HENDERSON R, SALI A, BAKER ML, CARRAGHER B, DEVKOTA B, DOWNING KH, EGELMAN EH, FENG Z, FRANK J, GRIGORIEFF N, JIANG W, LUDTKE SJ, MEDALIA O, PENCZEK PA, ROSENTHAL PB, ROSSMANN MG, SCHMID MF, SCHRODER GF, STEVEN AC, STOKES DL, WESTBROOK JD, WRIGGERS W, YANG H, YOUNG J, BERMAN HM, CHIU W, KLEYWEGT GJ & LAWSON CL (2012). Outcome of the first electron microscopy validation task force meeting. Structure, 20(2), 205–214. [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. HENDRICKSON WA, SMITH JL & SHERIFF S (1985). Direct phase determination based on anomalous scattering. Methods Enzymol, 115, 41–55. [DOI] [PubMed] [Google Scholar]
  42. HENRICK K, FENG Z, BLUHM WF, DIMITROPOULOS D, DORELEIJERS JF, DUTTA S, FLIPPEN-ANDERSON JL, IONIDES J, KAMADA C, KRISSINEL E, LAWSON CL, MARKLEY JL, NAKAMURA H, NEWMAN R, SHIMIZU Y, SWAMINATHAN J, VELANKAR S, ORY J, ULRICH EL, VRANKEN W, WESTBROOK J, YAMASHITA R, YANG H, YOUNG J, YOUSUFUDDIN M & BERMAN HM (2008). Remediation of the protein data bank archive. Nucleic Acids Res, 36(Database issue), D426–433. [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. HENRICK K, NEWMAN R, TAGARI M & CHAGOYEN M (2003). EMDep: a web-based system for the deposition and validation of high-resolution electron microscopy macromolecular structural information. J Struct Biol, 144(1–2), 228–237. [DOI] [PubMed] [Google Scholar]
  44. HOPE H (1988). Cryocrystallography of biological macromolecules: a generally applicable method. Acta Crystallogr, B44, 22–26. [DOI] [PubMed] [Google Scholar]
  45. HOPPER P, HARRISON SC & SAUER RT (1984). Structure of tomato bushy stunt virus. V. Coat protein sequence determination and its structural implications. J Mol Biol, 177(4), 701–713. [DOI] [PubMed] [Google Scholar]
  46. HORST R, DAMBERGER F, LUGINBÜHL P, GÜNTERT P, PENG G, NIKONOVA L, LEAL WS & WÜTHRICH K (2001). NMR structure reveals intramolecular regulation mechanism for pheromone binding and release. Proc Natl Acad Sci USA, 98. [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. HOWARD HUGHES MEDICAL INSTITUTE. (2017). History, vol. 2017. [Google Scholar]
  48. HUFTON AL (2014). Sharing the Structures. In Nature Milestones: Crystallography. (1970s) Open software and crystallographic databases. Nature, Scientific Data. [Google Scholar]
  49. INTERNATIONAL UNION OF CRYSTALLOGRAPHY. (1989). Commission on Biological Macromolecules. Acta Crystallographica Section A, 45(9), 658. [Google Scholar]
  50. JONES TA (1978). FRODO: A graphic model building and refinement system for macromolecules. J. Appl. Cryst, 11, 268–272. [Google Scholar]
  51. KARTHA G, BELLO J & HARKER D (1967). Tertiary structure of ribonuclease. Nature, 213, 862–865. [DOI] [PubMed] [Google Scholar]
  52. KELLY JA, SIELECKI AR, SYKES BD, JAMES MN & PHILLIPS DC (1979). X-ray crystallography of the binding of the bacterial cell wall trisaccharide NAM-NAG-NAM to lysozyme. Nature, 282(5741), 875–878. [DOI] [PubMed] [Google Scholar]
  53. KENDREW JC, BODO G, DINTZIS HM, PARRISH RG, WYCKOFF H & PHILLIPS DC (1958). A three-dimensional model of the myoglobin molecule obtained by x-ray analysis. Nature, 181, 662–666. [DOI] [PubMed] [Google Scholar]
  54. KIM SJ, FERNANDEZ-MARTINEZ J, SAMPATHKUMAR P, MARTEL A, MATSUI T, TSURUTA H, WEISS TM, SHI Y, MARKINA-INARRAIRAEGUI A, BONANNO JB, SAUDER JM, BURLEY SK, CHAIT BT, ALMO SC, ROUT MP & SALI A (2014). Integrative structure-function mapping of the nucleoporin nup133 suggests a conserved mechanism for membrane anchoring of the nuclear pore complex. Mol Cell Proteomics, 13(11), 2911–2926. [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. KINJO AR, BEKKER GJ, SUZUKI H, TSUCHIYA Y, KAWABATA T, IKEGAWA Y & NAKAMURA H (2017). Protein Data Bank Japan (PDBj): updated user interfaces, resource description framework, analysis tools for large structures. Nucleic Acids Res, 45(D1), D282–D288. [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. KOPP J & SCHWEDE T (2004). The SWISS-MODEL Repository of annotated three-dimensional protein structure homology models. In Nucleic Acids Res, vol. 32, pp. D230–234. [DOI] [PMC free article] [PubMed] [Google Scholar]
  57. KULLER A, FLERI W, BLUHM WF, SMITH JL, WESTBROOK J & BOURNE PE (2002). A biologist’s guide to synchrotron facilities: the BioSync web resource. TIBS, 27, 213–215. [DOI] [PubMed] [Google Scholar]
  58. LAWSON CL, BAKER ML, BEST C, BI C, DOUGHERTY M, FENG P, VAN GINKEL G, DEVKOTA B, LAGERSTEDT I, LUDTKE SJ, NEWMAN RH, OLDFIELD TJ, REES I, SAHNI G, SALA R, VELANKAR S, WARREN J, WESTBROOK JD, HENRICK K, KLEYWEGT GJ, BERMAN HM & CHIU W (2011). EMDataBank.org: unified data resource for CryoEM. Nucleic Acids Res, 39(Database issue), D456–464. [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. LAWSON CL, DUTTA S, WESTBROOK JD, HENRICK K & BERMAN HM (2008). Representation of viruses in the remediated PDB archive. Acta Crystallogr D Biol Crystallogr, D64(Pt 8), 874–882. [DOI] [PMC free article] [PubMed] [Google Scholar]
  60. LAWSON CL, PATWARDHAN A, BAKER ML, HRYC C, GARCIA ES, HUDSON BP, LAGERSTEDT I, LUDTKE SJ, PINTILIE G, SALA R, WESTBROOK JD, BERMAN HM, KLEYWEGT GJ & CHIU W (2016). EMDataBank unified data resource for 3DEM. Nucleic Acids Res, 44(D1), D396–403. [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. LEAVER-FAY A, TYKA M, LEWIS SM, LANGE OF, THOMPSON J, JACAK R, KAUFMAN K, RENFREW PD, SMITH CA, SHEFFLER W, DAVIS IW, COOPER S, TREUILLE A, MANDELL DJ, RICHTER F, BAN YE, FLEISHMAN SJ, CORN JE, KIM DE, LYSKOV S, BERRONDO M, MENTZER S, POPOVIC Z, HAVRANEK JJ, KARANICOLAS J, DAS R, MEILER J, KORTEMME T, GRAY JJ, KUHLMAN B, BAKER D & BRADLEY P (2011). ROSETTA3: an object-oriented software suite for the simulation and design of macromolecules. Meth. Enzymol, 487, 545–574. [DOI] [PMC free article] [PubMed] [Google Scholar]
  62. LIANG YL, KHOSHOUEI M, RADJAINIA M, ZHANG Y, GLUKHOVA A, TARRASCH J, THAL DM, FURNESS SGB, CHRISTOPOULOS G, COUDRAT T, DANEV R, BAUMEISTER W, MILLER LJ, CHRISTOPOULOS A, KOBILKA BK, WOOTTEN D, SKINIOTIS G & SEXTON PM (2017). Phase-plate cryo-EM structure of a class B GPCR-G-protein complex. Nature, 546(7656), 118–123. [DOI] [PMC free article] [PubMed] [Google Scholar]
  63. LIU Z, GONG Z, CAO Y, DING YH, DONG MQ, LU YB, ZHANG WP & TANG C (2018). Characterizing Protein Dynamics with Integrative Use of Bulk and Single-Molecule Techniques. Biochemistry, 57(3), 305–313. [DOI] [PubMed] [Google Scholar]
  64. LOQUET A, SGOURAKIS NG, GUPTA R, GILLER K, RIEDEL D, GOOSMANN C, GRIESINGER C, KOLBE M, BAKER D, BECKER S & LANGE A (2012). Atomic model of the type III secretion system needle. Nature, 486(7402), 276–279. [DOI] [PMC free article] [PubMed] [Google Scholar]
  65. LUGER K, MADER AW, RICHMOND RK, SARGENT DF & RICHMOND TJ (1997). Crystal structure of the nucleosome core particle at 2.8 A resolution. Nature, 389(6648), 251–260. [DOI] [PubMed] [Google Scholar]
  66. MATTHEWS BW (1996). Structural and genetic analysis of the folding and function of T4 lysozyme. FASEB J, 10(1), 35–41. [DOI] [PubMed] [Google Scholar]
  67. MEYER EF (1997). The first years of the Protein Data Bank. Protein Sci, 6(7), 1591–1597. [DOI] [PMC free article] [PubMed] [Google Scholar]
  68. MONTELIONE GT, NILGES M, BAX A, GUNTERT P, HERRMANN T, RICHARDSON JS, SCHWIETERS CD, VRANKEN WF, VUISTER GW, WISHART DS, BERMAN HM, KLEYWEGT GJ & MARKLEY JL (2013). Recommendations of the wwPDB NMR Validation Task Force. Structure, 21(9), 1563–1570. [DOI] [PMC free article] [PubMed] [Google Scholar]
  69. NOBELPRIZE.ORG. (2017). The Nobel Prize in Chemistry 1962. [Google Scholar]
  70. NORVELL JC & BERG JM (2007). Update on the protein structure initiative. Structure, 15(12), 1519–1522. [DOI] [PubMed] [Google Scholar]
  71. OSTROM E (1990). Governing the Commons: The Evolution of Institutions for Collective Action.: Cambridge University Press. [Google Scholar]
  72. PATIKOGLOU GA, KIM JL, SUN L, YANG SH, KODADEK T & BURLEY SK (1999). TATA element recognition by the TATA box-binding protein has been conserved throughout evolution. Genes Dev, 13, 3217–3230. [DOI] [PMC free article] [PubMed] [Google Scholar]
  73. PERUTZ MF, ROSSMANN MG, CULLIS AF, MUIRHEAD H, WILL G & NORTH ACT (1960). Structure of haemoglobin: a three-dimensional Fourier synthesis at 5.5 Å resolution, obtained by X-ray analysis. Nature, 185, 416–422. [DOI] [PubMed] [Google Scholar]
  74. PETTERSEN EF, GODDARD TD, HUANG CC, COUCH GS, GREENBLATT DM, MENG EC & FERRIN TE (2004). UCSF Chimera--a visualization system for exploratory research and analysis. Journal of Computational Chemistry, 25(13), 1605–1612. [DOI] [PubMed] [Google Scholar]
  75. PHILLIPS DC (1972). Protein crystallography 1971: Coming of age In Cold Spring Harbor Symposia on Quantitative Biology, vol. 36 Cold Spring Harbor: Cold Spring Harbor Laboratory Press. [DOI] [PubMed] [Google Scholar]
  76. PIEPER U, ESWAR N, WEBB BM, ERAMIAN D, KELLY L, BARKAN DT, CARTER H, MANKOO P, KARCHIN R, MARTI-RENOM MA, DAVIS FP & SALI A (2009). MODBASE, a database of annotated comparative protein structure models and associated resources. Nucleic Acids Res, 37(Database issue), D347–354. [DOI] [PMC free article] [PubMed] [Google Scholar]
  77. PIEPER U, SCHLESSINGER A, KLOPPMANN E, CHANG GA, CHOU JJ, DUMONT ME, FOX BG, FROMME P, HENDRICKSON WA, MALKOWSKI MG, REES DC, STOKES DL, STOWELL MH, WIENER MC, ROST B, STROUD RM, STEVENS RC & SALI A (2013). Coordinating the impact of structural genomics on the human alpha-helical transmembrane proteome. Nat Struct Mol Biol, 20(2), 135–138. [DOI] [PMC free article] [PubMed] [Google Scholar]
  78. POLITIS A, STENGEL F, HALL Z, HERNANDEZ H, LEITNER A, WALZTHOENI T, ROBINSON CV & AEBERSOLD R (2014). A mass spectrometry-based hybrid method for structural modeling of protein complexes. Nat. Methods, 11(4), 403–406. [DOI] [PMC free article] [PubMed] [Google Scholar]
  79. PROTEIN DATA BANK. (1971). Crystallography: Protein Data Bank. Nature New Biol, 233(42), 223–223.20480989 [Google Scholar]
  80. QUIOCHO FA & LIPSCOMB WN (1971). Carboxypeptidase A: a protein and an enzyme. Adv Protein Chem, 25, 1–78. [DOI] [PubMed] [Google Scholar]
  81. READ RJ, ADAMS PD, ARENDALL WB 3RD, BRUNGER AT, EMSLEY P, JOOSTEN RP, KLEYWEGT GJ, KRISSINEL EB, LUTTEKE T, OTWINOWSKI Z, PERRAKIS A, RICHARDSON JS, SHEFFLER WH, SMITH JL, TICKLE IJ, VRIEND G & ZWART PH (2011). A new generation of crystallographic validation tools for the protein data bank. Structure, 19(10), 1395–1412. [DOI] [PMC free article] [PubMed] [Google Scholar]
  82. RICH A & KIM S-H (1978). The three-dimensional structure of transfer RNA. Sci. Am, 238, 52–62. [DOI] [PubMed] [Google Scholar]
  83. RICHARDS FM (1968). The matching of physical models to three-dimensional electron-density maps: a simple optical device. J. Mol. Biol, 37, 225–230. [DOI] [PubMed] [Google Scholar]
  84. ROBERTUS JD, LADNER JE, FINCH JT, RHODES D, BROWN RS, CLARK BFC & KLUG A (1974). Structure of yeast phenylalanine tRNA at 3 Å resolution. Nature, 250, 546–551. [DOI] [PubMed] [Google Scholar]
  85. ROBINSON PJ, TRNKA MJ, PELLARIN R, GREENBERG CH, BUSHNELL DA, DAVIS R, BURLINGAME AL, SALI A & KORNBERG RD (2015). Molecular architecture of the yeast Mediator complex. Elife, 4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  86. ROH SH, HRYC CF, JEONG HH, FEI X, JAKANA J, LORIMER GH & CHIU W (2017). Subunit conformational variation within individual GroEL oligomers resolved by Cryo-EM. Proc Natl Acad Sci U S A, 114(31), 8259–8264. [DOI] [PMC free article] [PubMed] [Google Scholar]
  87. ROSE PW, PRLIC A, ALTUNKAYA A, BI C, BRADLEY AR, CHRISTIE CH, COSTANZO LD, DUARTE JM, DUTTA S, FENG Z, GREEN RK, GOODSELL DS, HUDSON B, KALRO T, LOWE R, PEISACH E, RANDLE C, ROSE AS, SHAO C, TAO YP, VALASATAVA Y, VOIGT M, WESTBROOK JD, WOO J, YANG H, YOUNG JY, ZARDECKI C, BERMAN HM & BURLEY SK (2017). The RCSB protein data bank: integrative view of protein, gene and 3D structural information. Nucleic Acids Res, 45(D1), D271–D281. [DOI] [PMC free article] [PubMed] [Google Scholar]
  88. ROSSMANN MG, MORAIS MC, LEIMAN PG & ZHANG W (2005). Combining X-ray crystallography and electron microscopy. Structure, 13(3), 355–362. [DOI] [PMC free article] [PubMed] [Google Scholar]
  89. RUSSEL D, LASKER K, WEBB B, VELAZQUEZ-MURIEL J, TJIOE E, SCHNEIDMAN-DUHOVNY D, PETERSON B & SALI A (2012). Putting the pieces together: integrative modeling platform software for structure determination of macromolecular assemblies. PLoS Biol, 10(1), e1001244. [DOI] [PMC free article] [PubMed] [Google Scholar]
  90. SALI A, BERMAN HM, SCHWEDE T, TREWHELLA J, KLEYWEGT G, BURLEY SK, MARKLEY J, NAKAMURA H, ADAMS P, BONVIN AM, CHIU W, PERARO MD, DI MAIO F, FERRIN TE, GRUNEWALD K, GUTMANAS A, HENDERSON R, HUMMER G, IWASAKI K, JOHNSON G, LAWSON CL, MEILER J, MARTI-RENOM MA, MONTELIONE GT, NILGES M, NUSSINOV R, PATWARDHAN A, RAPPSILBER J, READ RJ, SAIBIL H, SCHRODER GF, SCHWIETERS CD, SEIDEL CA, SVERGUN D, TOPF M, ULRICH EL, VELANKAR S & WESTBROOK JD (2015). Outcome of the First wwPDB Hybrid/Integrative Methods Task Force Workshop. Structure, 23(7), 1156–1167. [DOI] [PMC free article] [PubMed] [Google Scholar]
  91. SCHLUENZEN F, TOCILJ A, ZARIVACH R, HARMS J, GLUEHMANN M, JANELL D, BASHAN A, BARTELS H, AGMON I, FRANCESCHI F & YONATH A (2000). Structure of functionally activated small ribosomal subunit at 3.3 Å resolution. Cell, 102, 615–623. [DOI] [PubMed] [Google Scholar]
  92. SCHWIETERS CD, BERMEJO GA & CLORE GM (2018). Xplor-NIH for molecular structure determination from NMR and other data sources. Protein Sci, 27(1), 26–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
  93. SEILER CY, PARK JG, SHARMA A, HUNTER P, SURAPANENI P, SEDILLO C, FIELD J, ALGAR R, PRICE A, STEEL J, THROOP A, FIACCO M & LABAER J (2014). DNASU plasmid and PSI:Biology-Materials repositories: resources to accelerate biological research. Nucleic Acids Res, 42(Database issue), D1253–1260. [DOI] [PMC free article] [PubMed] [Google Scholar]
  94. SERRA F, BAU D, GOODSTADT M, CASTILLO D, FILION GJ & MARTI-RENOM MA (2017). Automatic analysis and 3D-modelling of Hi-C data using TADbit reveals structural features of the fly chromatin colors. PLoS Comput Biol, 13(7), e1005665. [DOI] [PMC free article] [PubMed] [Google Scholar]
  95. SHARIF H, LI Y, DONG Y, DONG L, WANG WL, MAO Y & WU H (2017). Cryo-EM structure of the DNA-PK holoenzyme. Proc Natl Acad Sci U S A, 114(28), 7367–7372. [DOI] [PMC free article] [PubMed] [Google Scholar]
  96. SHI Y, FERNANDEZ-MARTINEZ J, TJIOE E, PELLARIN R, KIM SJ, WILLIAMS R, SCHNEIDMAN-DUHOVNY D, SALI A, ROUT MP & CHAIT BT (2014). Structural characterization by cross-linking reveals the detailed architecture of a coatomer-related heptameric module from the nuclear pore complex. Mol Cell Proteomics, 13(11), 2927–2943. [DOI] [PMC free article] [PubMed] [Google Scholar]
  97. SHI Y, PELLARIN R, FRIDY PC, FERNANDEZ-MARTINEZ J, THOMPSON MK, LI Y, WANG QJ, SALI A, ROUT MP & CHAIT BT (2015). A strategy for dissecting the architectures of native macromolecular assemblies. Nat Methods, 12(12), 1135–1138. [DOI] [PMC free article] [PubMed] [Google Scholar]
  98. TRUSSART M, SERRA F, BAU D, JUNIER I, SERRANO L & MARTI-RENOM MA (2015). Assessing the limits of restraint-based 3D modeling of genomes and genomic domains. Nucleic Acids Res, 43(7), 3465–3477. [DOI] [PMC free article] [PubMed] [Google Scholar]
  99. TWOMEY EC, YELSHANSKAYA MV, GRASSUCCI RA, FRANK J & SOBOLEVSKY AI (2017). Channel opening and gating mechanism in AMPA-subtype glutamate receptors. Nature. [DOI] [PMC free article] [PubMed] [Google Scholar]
  100. ULRICH EL, AKUTSU H, DORELEIJERS JF, HARANO Y, IOANNIDIS YE, LIN J, LIVNY M, MADING S, MAZIUK D, MILLER Z, NAKATANI E, SCHULTE CF, TOLMIE DE, KENT WENGER R, YAO H & MARKLEY JL (2008). BioMagResBank. Nucleic Acids Res, 36(Database issue), D402–408. [DOI] [PMC free article] [PubMed] [Google Scholar]
  101. VALLAT B, WEBB B, WESTBROOK J, SALI A & BERMAN H (2016a). Integrative/Hybrid Methods PDBx/mmCIF dictionary extension, vol. 2016. [Google Scholar]
  102. VALLAT B, WEBB B, WESTBROOK J, SALI A & BERMAN H (2016b). Integrative/Hybrid Methods PDBx/mmCIF dictionary extension documentation, vol. 2016. [Google Scholar]
  103. VALLAT B, WEBB B, WESTBROOK J, SALI A & BERMAN HM (2016c). The PDB-Dev prototype deposition and archiving system, vol. 2016. [Google Scholar]
  104. VALLAT B, WEBB B, WESTBROOK J, SALI A & BERMAN HM (2017). A Data Dictionary For Archiving Integrative/Hybrid Models. In 24th IUCr Congress and General Assembly International Union of Crystallography, Hyderabad, India. [Google Scholar]
  105. VAN ZUNDERT GCP, MELQUIOND ASJ & BONVIN A (2015). Integrative Modeling of Biomolecular Complexes: HADDOCKing with Cryo-Electron Microscopy Data. Structure, 23(5), 949–960. [DOI] [PubMed] [Google Scholar]
  106. VELANKAR S, VAN GINKEL G, ALHROUB Y, BATTLE GM, BERRISFORD JM, CONROY MJ, DANA JM, GORE SP, GUTMANAS A, HASLAM P, HENDRICKX PM, LAGERSTEDT I, MIR S, FERNANDEZ MONTECELO MA, MUKHOPADHYAY A, OLDFIELD TJ, PATWARDHAN A, SANZ-GARCIA E, SEN S, SLOWLEY RA, WAINWRIGHT ME, DESHPANDE MS, IUDIN A, SAHNI G, SALAVERT TORRES J, HIRSHBERG M, MAK L, NADZIRIN N, ARMSTRONG DR, CLARK AR, SMART OS, KORIR PK & KLEYWEGT GJ (2016). PDBe: improved accessibility of macromolecular structure data from PDB and EMDB. Nucleic Acids Res, 44(D1), D385–395. [DOI] [PMC free article] [PubMed] [Google Scholar]
  107. VINOTHKUMAR KR & HENDERSON R (2016). Single particle electron cryomicroscopy: trends, issues and future perspective. Q Rev Biophys, 49, e13. [DOI] [PubMed] [Google Scholar]
  108. WAN R, YAN C, BAI R, HUANG G & SHI Y (2016). Structure of a yeast catalytic step I spliceosome at 3.4 A resolution. Science, 353(6302), 895–904. [DOI] [PubMed] [Google Scholar]
  109. WARD AB, SALI A & WILSON IA (2013). Biochemistry. Integrative structural biology. Science, 339(6122), 913–915. [DOI] [PMC free article] [PubMed] [Google Scholar]
  110. WATSON HC (1969). The stereochemistry of the protein myoglobin. Prog. Stereochem, 4, 299. [Google Scholar]
  111. WESTBROOK JD & FITZGERALD PMD (2009). Chapter 10 The PDB format, mmCIF formats, and other data formats In Bioinformatics Structural, Second Edition eds. P. E. Bourne and J. Gu), pp. 271–291. Hoboken, NJ: John Wiley & Sons, Inc. [Google Scholar]
  112. WILKINSON MD, DUMONTIER M, AALBERSBERG IJ, APPLETON G, AXTON M, BAAK A, BLOMBERG N, BOITEN JW, DA SILVA SANTOS LB, BOURNE PE, BOUWMAN J, BROOKES AJ, CLARK T, CROSAS M, DILLO I, DUMON O, EDMUNDS S, EVELO CT, FINKERS R, GONZALEZ-BELTRAN A, GRAY AJ, GROTH P, GOBLE C, GRETHE JS, HERINGA J, T HOEN PA, HOOFT R, KUHN T, KOK R, KOK J, LUSHER SJ, MARTONE ME, MONS A, PACKER AL, PERSSON B, ROCCA-SERRA P, ROOS M, VAN SCHAIK R, SANSONE SA, SCHULTES E, SENGSTAG T, SLATER T, STRAWN G, SWERTZ MA, THOMPSON M, VAN DER LEI J, VAN MULLIGEN E, VELTEROP J, WAAGMEESTER A, WITTENBURG P, WOLSTENCROFT K, ZHAO J & MONS B (2016). The FAIR Guiding Principles for scientific data management and stewardship. Sci Data, 3, 160018. [DOI] [PMC free article] [PubMed] [Google Scholar]
  113. WLODAWER A (2002). Rational approach to AIDS drug design through structural biology. Annu Rev Med, 53, 595–614. [DOI] [PubMed] [Google Scholar]
  114. WLODAWER A, MINOR W, DAUTER Z & JASKOLSKI M (2008). Protein crystallography for non-crystallographers, or how to get the best (but not more) from published macromolecular structures. FEBS J, 275(1), 1–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  115. WYCKOFF HW, HARDMAN KD, ALLEWELL NM, INAGAMI T, TSERNOGLOU D, JOHNSON LN & RICHARDS FM (1967). The structure of ribonuclease-S at 6 Å resolution. J. Biol. Chem, 242, 3749–3753. [PubMed] [Google Scholar]
  116. YOUNG JY, WESTBROOK JD, FENG Z, SALA R, PEISACH E, OLDFIELD TJ, SEN S, GUTMANAS A, ARMSTRONG DR, BERRISFORD JM, CHEN L, CHEN M, DI COSTANZO L, DIMITROPOULOS D, GAO G, GHOSH S, GORE S, GURANOVIC V, HENDRICKX PM, HUDSON BP, IGARASHI R, IKEGAWA Y, KOBAYASHI N, LAWSON CL, LIANG Y, MADING S, MAK L, MIR MS, MUKHOPADHYAY A, PATWARDHAN A, PERSIKOVA I, RINALDI L, SANZ-GARCIA E, SEKHARAN MR, SHAO C, SWAMINATHAN GJ, TAN L, ULRICH EL, VAN GINKEL G, YAMASHITA R, YANG H, ZHURAVLEVA MA, QUESADA M, KLEYWEGT GJ, BERMAN HM, MARKLEY JL, NAKAMURA H, VELANKAR S & BURLEY SK (2017). OneDep: Unified wwPDB System for Deposition, Biocuration, and Validation of Macromolecular Structures in the PDB Archive. Structure, 25(3), 536–545. [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES