Skip to main content
Data in Brief logoLink to Data in Brief
. 2015 Aug 19;5:23–27. doi: 10.1016/j.dib.2015.08.003

Data from a proteomic baseline study of Assemblage A in Giardia duodenalis

Samantha J Emery a, Ernest Lacey b, Paul A Haynes a,
PMCID: PMC4556777  PMID: 26380841

Abstract

Eight Assemblage A strains from the protozoan parasite Giardia duodenalis were analysed using label-free quantitative shotgun proteomics, to evaluate inter- and intra-assemblage variation and complement available genetic and transcriptomic data. Isolates were grown in biological triplicate in axenic culture, and protein extracts were subjected to in-solution digest and online fractionation using Gas Phase Fractionation (GPF). Recent reclassification of genome databases for subassemblages was evaluated for database-dependent loss of information, and proteome composition of different isolates was analysed for biologically relevant assemblage-independent variation. The data from this study are related to the research article “Quantitative proteomics analysis of Giardia duodenalis Assemblage A – a baseline for host, assemblage and isolate variation” published in Proteomics (Emery et al., 2015 [1]).

Keywords: Assemblage A, Giardia duodenalis, Label-free quantitative shotgun proteomics, Variant Surface Protein, Variable genome, Parasite proteomics


Specifications table

Subject area Biology
More specific subject area Quantitative proteomic data of 8 Giardia duodenalis Assemblage A isolates using gas phase fractionation and normalised spectral abundance factors (NSAF).
Type of data Table, Figure, Supplementary Tables
How data was acquired Protein extracts from biological triplicates were digested in solution, and fractionated online using GPF with mass range fraction optimised for the G. duodenalis A1 subassemblage genome. Data was acquired on a LTQ-XL Linear Ion Trap (Thermo).
Data format Raw data, reproducibly identified proteins.
Experimental factors 8 G. duodenalis strains grown in Axenic culture from animal and human hosts, covering both subassemblage A1 and A2 to analyse isolate variation. Data was searched against both A1 subassemblage genome database and recently released A2 subassemblage database to compare database-specific losses.
Experimental features Sample triplicates were combined to produce reproducibly identified proteins and spectral counts of each protein were used to calculate NSAF values for each protein.
Data source location Sydney, NSW, Australia
Data accessibility Data is available from http://www.ebi.ac.uk/pride/archive/projects/PXD001272 and will also be made available through the giardiadb.org website later in 2015.
Value of the data
  • First proteomic baseline for taxonomy and isolate variation in Assemblage A strains.

  • Provides proteome coverage of isolates from animal and human hosts, both A1 and A2 subassemblages, with an emphasis on Australian isolates.

  • Evaluates database-dependent losses based on new genome reclassifications and releases in Assemblage A.

  • Identifies sources of inter- and intra-assemblage A isolate variation and its impacts.

1. Experimental design, materials and methods

1.1. Isolate selection, axenic culture, protein extraction and digestion

Eight Assemblage A strains [1], including the A1 genome strain, were assembled from animal and human infections, previously characterised in the literature according to karotype [2,3], subassemblage [4], virulence [2], geographic variation [5,6] and drug resistance [7]. The full description of strains can be seen in Table 1.

Table 1.

Classification information for the eight G. duodenalis strains used in this study including subassemblage, geographic origin, and the host species the strain was isolated from. Strain identification coincides with those previously published in the literature.

Strain Assemblage Origin Host source
BRIS/83/HEPU 106 A1 Brisbane, Australia Human
BRIS87/HEPU/713 A1 Brisbane, Australia Human
OAS1 A1 Canada Sheep (Ovis aries)
Bac2 A1 Australia Cat (Felis catus)
BRIS/95/HEPU/2041 A1 Victoria, Australia Cockatoo (Cacatua galerita)
BRIS/89/HEPU/1065 A1 Brisbane, Australia Human
WB* A1 Afghanistan Human
BRIS/89/HEPU/1003 A2 Brisbane, Australia Human

Assemblage A1 genome strain (ATCC 50803).

G. duodenalis strains were cultured in triplicate axenically in TYI-S33 media supplemented with 10% newborn calf serum and 1% bile as previously described [8] and harvested from confluent cultures in late log-phase. Trophozoites were harvested by centrifugation, washed twice in ice-cold PBS to remove media traces [9] and pellets of 108 trophozoites were extracted into 1 mL ice-cold SDS sample buffer containing 1 mM EDTA and 5% beta-mercaptoethanol, then disulphides were reduced at 75 °C for 10 min. Trophozoite protein extracts were centrifuged at 0 °C at 13,000×g for 10 min to remove debris, and protein concentration was measured by BCA assay (Pierce). A 500 µg protein pellet was extracted using methanol–chloroform precipitation [10] and in-solution digestion was performed using a modified filter aided sample preparation (FASP) [11]. After peptide extraction all samples were dried using a vacuum centrifuge and reconstituted to 60 µL with 2% formic acid, 2% 2,2,2-trifluorethanol (TFE).

1.2. Nanoflow LC-MS/MS using gas phase fractionation

Optimised gas phase fractionation (GPF) mass ranges were calculated using the 2.5 release of the G. duodenalis WB genome for Assemblage A from giardiaDB.org [12]. Charge states +2 and +3 were considered as well as carbamidomethyl as a cysteine modification, and 4 mass ranges were calculated over 400–2000 amu. The mass ranges were as following: the low mass range was 400–518 amu, the low-medium mass range was 518–691 amu, the medium-high mass range was 691–988 amu and the high mass range was 988–2000 amu. Each FASP protein digest for the triplicates of each strain were analysed by nanoLC-MS/MS on an LTQ-XL linear ion trap mass spectrometer (Thermo, San Jose, CA). Peptides were separated on a 150×0.2 mm I.D fused-silica column packed with Magic C18AQ (200 Å, 5 µm diameter, Michrom Bioresources, California) connected to an Advance CaptiveSpray Source (Michrom Bioresources, California). Each FASP protein digest was analysed as 4 repeat injections, with the mass spectrometer scanning for 180 min runs for each of the four calculated mass ranges. Samples were injected onto the column using a Surveyor autosampler, followed by an initial wash step with buffer A (0.1% v/v formic acid, 1 mM ammonium formate, 0.2% v/v methanol) for 4 min followed by 150 µL/min for 2 min. Peptides were eluted from the column with 0–80% buffer B (100% v/v ACN, 0.1% v/v formic acid) at 150 µL/min for 167 min finished by a wash step with buffer A for 6 min at 150 µL/min. Spectra in the positive ion mode were scanned over the respective GPF ranges and, using Xcalibur software (Version 2.06, Thermo), automated peak recognition, dynamic exclusion and MS/MS of the top six most-intense ions at 35% normalisation collision energy were performed.

1.3. Database searching for protein/peptide information

The LTQ-XL raw output files were converted into mzXML files and searched against the Giardiadb.org 4.0 release of G. duodenalis strain Assemblage A1 and A2 genome using the global proteome machine (GPM) software (version 2.1.1) and the X!Tandem algorithm. The 4 fractions for the GPF of each replicate were processed sequentially with output files generated for each individual fraction, and a merged, non-redundant output file for protein identifications with log(e) values<−1. Peptide identification was determined using MS and MS/MS tolerances of +2 Da and +0.4 Da. Carbamidomethyl was considered a complete modification, and partial modifications considered included oxidation of methionine and tryptophan.

1.4. Data processing and quantitation

The output from the GPM software (version 2.1.1) [13,14] constituted low stringency protein and peptide identifications, and was used to assess experimental consistency. These data were further processed using the Scrappy software package [15], which combines biological triplicates into a single list of reproducibly identified proteins, which we define in this study as those proteins present reproducibly in all three replicates of at least one strain, with a total spectral count (SpC) of ≥5 [15]. Reversed database searching was used for calculating peptide and protein false discovery rates (FDRs) as previously described [15]. Complete protein and peptide data for replicates, including database-dependent losses are shown in Supplementary data 1, Table 1 and in Giardia specific gene-families in Supplementary data, Table 2. Protein abundance was calculated using NSAF values [16]. Distribution of reproducibly identified proteins by strain can be viewed in Fig. 1. The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium [17] via the PRIDE partner repository with the dataset identifier PXD001272.

Fig. 1.

Fig. 1

Distribution of shared and unique proteins in the A1 subassemblage between the 1197 non-redundant proteins identified within the seven isolates analysed. The 1197 proteins were reproducibly identified in at least one isolate, with 149 (12.4%) of these proteins identified within only one isolate, and therefore considered to be uniquely expressed. Part A (left) shows the distribution of these 149 uniquely expressed proteins by isolate in the seven A1 isolates analysed in this study. Part B (right) shows the distribution of the shared proteins between the seven subassemblage A1 isolates. A total of 503 (42%) proteins were identified in all seven isolates examined in this study, and are considered common between isolates of the A1 subassemblage. The remaining segments indicates proteins common within decreasing numbers of isolates, while the final elevated segment indicates the 149 isolate-unique proteins.

2. Direct link to deposited data

Data is available through the PRIDE proteomics database through the following link http://www.ebi.ac.uk/pride/archive/projects/PXD001272 and will also be made available through the giardiadb.org website later in 2015.

3. Conflict of interest

The authors declare that there is no conflict of interest on any work published in this paper.

Acknowledgements

SJE acknowledges funding from the Australian Government in the form of an APA scholarship, as well as financial support from Macquarie University. SJE wishes to thank Dr Jacqui Upcroft for supplying the Giardia samples used and for the ongoing support received from colleagues at Microbial Screening Technologies. PAH wishes to thank Justin Lane for continued support and encouragement.

Footnotes

Appendix A

Supplementary data associated with this article can be found in the online version at doi:10.1016/j.dib.2015.08.003.

Appendix A. Supporting information

Supplementary data

mmc1.docx (19.1KB, docx)

Supplementary data

mmc2.docx (13.5KB, docx)

References

  • 1.Emery S.J., Lacey E., Haynes P.A. Quantitative proteomics analysis of Giardia duodenalis Assemblage A – a baseline for host, assemblage and isolate variation. Proteomics. 2015;15:2281–2285. doi: 10.1002/pmic.201400434. [DOI] [PubMed] [Google Scholar]
  • 2.Williamson A.L., O׳Donoghue P.J., Upcroft J.A., Upcroft P. Immune and pathophysiological responses to different strains of Giardia duodenalis in neonatal mice. Int. J. Parasitol. 2000;30:129–136. doi: 10.1016/s0020-7519(99)00181-2. [DOI] [PubMed] [Google Scholar]
  • 3.Upcroft J.A., Chen N., Upcroft P. Mapping variation in chromosome homologues of different Giardia strains. Mol. Biochem. Parasitol. 1996;76:135–143. doi: 10.1016/0166-6851(95)02554-5. [DOI] [PubMed] [Google Scholar]
  • 4.Nolan M.J., Jex A.R., Upcroft J.A., Upcroft P., Gasser R.B. Barcoding of Giardia duodenalis isolates and derived lines from an established cryobank by a mutation scanning-based approach. Electrophoresis. 2011;32:2075–2090. doi: 10.1002/elps.201100283. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Upcroft J.A., Boreham P.F., Campbell R.W., Shepherd R.W., Upcroft P. Biological and genetic analysis of a longitudinal collection of Giardia samples derived from humans. Acta Trop. 1995;60:35–46. doi: 10.1016/0001-706x(95)00100-s. [DOI] [PubMed] [Google Scholar]
  • 6.Upcroft J.A., Boreham P.F., Upcroft P. Geographic variation in Giardia karyotypes. Int. J. Parasitol. 1989;19:519–527. doi: 10.1016/0020-7519(89)90082-9. [DOI] [PubMed] [Google Scholar]
  • 7.Upcroft J.A., Healey A., Murray D.G., Boreham P.F., Upcroft P. A gene associated with cell division and drug resistance in Giardia duodenalis. Parasitology. 1992;104(Pt 3):397–405. doi: 10.1017/s0031182000063642. [DOI] [PubMed] [Google Scholar]
  • 8.Keister D.B. Axenic culture of Giardia lamblia in TYI-S-33 medium supplemented with bile. Trans. R. Soc. Trop. Med. Hyg. 1983;77:487–488. doi: 10.1016/0035-9203(83)90120-7. [DOI] [PubMed] [Google Scholar]
  • 9.Dunn L.A., Upcroft J.A., Fowler E.V., Matthews B.S., Upcroft P. Orally administered Giardia duodenalis extracts enhance an antigen-specific antibody response. Infect. Immun. 2001;69:6503–6510. doi: 10.1128/IAI.69.10.6503-6510.2001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Wessel D., Flugge U.I. A method for the quantitative recovery of protein in dilute solution in the presence of detergents and lipids. Anal. Biochem. 1984;138:141–143. doi: 10.1016/0003-2697(84)90782-6. [DOI] [PubMed] [Google Scholar]
  • 11.Chapman B., Castellana N., Apffel A., Ghan R. Plant proteogenomics: from protein extraction to improved gene predictions. Methods Mol. Biol. 2013;1002:267–294. doi: 10.1007/978-1-62703-360-2_21. [DOI] [PubMed] [Google Scholar]
  • 12.Scherl A., Shaffer S.A., Taylor G.K., Kulasekara H.D. Genome-specific gas-phase fractionation strategy for improved shotgun proteomic profiling of proteotypic peptides. Anal. Chem. 2008;80:1182–1191. doi: 10.1021/ac701680f. [DOI] [PubMed] [Google Scholar]
  • 13.Fenyo D., Beavis R.C. A method for assessing the statistical significance of mass spectrometry-based protein identifications using general scoring schemes. Anal. Chem. 2003;75:768–774. doi: 10.1021/ac0258709. [DOI] [PubMed] [Google Scholar]
  • 14.Craig R., Beavis R.C. TANDEM: matching proteins with tandem mass spectra. Bioinformatics. 2004;20:1466–1467. doi: 10.1093/bioinformatics/bth092. [DOI] [PubMed] [Google Scholar]
  • 15.Neilson K.A., George I.S., Emery S.J., Muralidharan S. Analysis of rice proteins using SDS-PAGE shotgun proteomics. Methods Mol. Biol. 2014;1072:289–302. doi: 10.1007/978-1-62703-631-3_21. [DOI] [PubMed] [Google Scholar]
  • 16.Zybailov B., Mosley A.L., Sardiu M.E., Coleman M.K. Statistical analysis of membrane proteome expression changes in Saccharomyces cerevisiae. J. Proteome Res. 2006;5:2339–2347. doi: 10.1021/pr060161n. [DOI] [PubMed] [Google Scholar]
  • 17.Vizcaino J.A., Cote R.G., Csordas A., Dianes J.A. The PRoteomics IDEntifications (PRIDE) database and associated tools: status in 2013. Nucleic Acids Res. 2013;41:D1063–1069. doi: 10.1093/nar/gks1262. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary data

mmc1.docx (19.1KB, docx)

Supplementary data

mmc2.docx (13.5KB, docx)

Articles from Data in Brief are provided here courtesy of Elsevier

RESOURCES