Skip to main content
. 2020 Nov 4;11:575377. doi: 10.3389/fmicb.2020.575377

TABLE 2.

Participants, their background and applied data processing and important comments from their summaries.

Participants Participants’ sector Data processing workflow Participants comments
P1 Food – Quality assessment w/FastQC (Wingett and Andrews, 2018) – Quality trimming w/fastp (Chen et al., 2018) – Taxonomic classification w/Kraken (custom database and MiniKraken DB; Wood and Salzberg, 2014) – Additionally pathoLive analysis for classification of viral reads (Tausch et al., 2018) – FastQC revealed bases of bad quality at the beginning of the reads. Therefore the reads were trimmed – kraken analysis with custom database resulted in many false-positive results; therefore, results were confirmed with BLASTn (Boratyn et al., 2012).
P2 Human – Quality assessment w/FastQC (Wingett and Andrews, 2018) – Quality trimming w/Trimmomatic (Bolger et al., 2014) – Taxonomic classification based on mapping and assembly w/Pikavirus (in-house in-development tool at https://github.com/BU-ISCIII/PikaVirus) – Taxonomic classification based on mapping and assembly w/oases (Schulz et al., 2012) – Taxonomic classification based on rRNA clustering w/MeTRS (Cottier et al., 2018) – Taxonomic classification based on protein identity analysis w/Kaiju (Menzel et al., 2016) – trimming parameters: nucleotides at 3′ with phred quality <10 or average quality ≤15 (window size 4), removal of reads shorter 50 bp – trimming dropped 8137323 sequences (81.71%) – unusual bad quality 5′ end was observed at the 25 firsts bases
P3 Human – Taxonomic classification w/Kraken (Wood and Salzberg, 2014) as implemented on Galaxy public server – Norovirus GV (murine norovirus, not a human pathogen) Unlikely to be on food sample – Hepatitis C virus (human pathogen, but route of transmission is via blood), highly unlikely to be found on food sample, and contamination with human blood?
P4 Food – Quality assessment w/FastQC (Wingett and Andrews, 2018) – Quality/Adapter trimming w/BBduk v. 36.49 (https://sourceforge.net/projects/bbmap/) – Taxonomic classification w/MGmapper (Petersen et al., 2017) – mapped against a phiX174 reference sequence to remove potential control library reads – Certain contaminations can be difficult to identify without control samples. For example, certain microorganisms may be part of the natural microbiome of fish or could have been introduced during sample handling and processing.
P5 Veterinary – Quality assessment w/FastQC (Wingett and Andrews, 2018) – Quality/Adapter trimming w/Trimmomatic (Bolger et al., 2014) – Host sequence removal w/BWA-MEM (Li, 2013) – Taxonomic classification w/Kraken (MiniKraken database; ref. Wood and Salzberg, 2014) – Filtered for minimum read count (threshold 500 reads) – species for which there were less than 1,000 reads would need further confirmation before release of the information – PhiX carry over from the sequencing lab
P6 Human – Quality assessment w/FastQC (Wingett and Andrews, 2018) – Quality trimming w/Trimmomatic (Bolger et al., 2014) – Host sequence removal w/BBmap (id threshold 0.65; https://sourceforge.net/projects/bbmap/) – Taxonomic classification w/Kraken (database version 13/10/2017; Wood and Salzberg, 2014) – FastQC result: No adapters detected
P7 Human – Quality/Adapter trimming w/Trimmomatic (Bolger et al., 2014) – Host sequence removal w/Bowtie2 (Langmead and Salzberg, 2012) – Taxonomic classification w/MALT (Herbig et al., 2016), DIAMOND (Buchfink et al., 2015), MEGAN (Huson et al., 2016), custom database with refseq viruses, bacterial, fungi, and protists – Pathogen of importance is Norovirus GV – Abundance of a cloning vector could be an artifact of sequencing reagents and preparation
P8 Veterinary – Quality trimming w/RIEMS (Scheuch et al., 2015) – Taxonomic classification w/RIEMS (Scheuch et al., 2015); ncbi nt – The Calicivirdae/Norwalk virus reads indicate the presence of noroviruses in the sample. This is the most important entero-pathogenic virus in the analyzed sample
P9 Food – Species-level classification w/Kraken (Wood and Salzberg, 2014) – Multi-locus sequence types (MLSTs) reconstructed w/MetaMLST (Zolfo et al., 2017) – Strain-level identification w/PanPhlAn (Scholz et al., 2016) None
P10 Food – Quality assessment w/FastQC (Wingett and Andrews, 2018) – Quality trimming w/cutadapt (part of MGmapper processing; (Martin, 2011) – Taxonomic classification w/MGmapper (Petersen et al., 2017) – no hits w/default settings, re-analysis w/adjusted parameters (max mismatch ratio = 0.15, min read count = 20) – sequence GC content measured by FastQC is reported as failure
P11 Veterinary – Host sequences were removed by blasting (BLASTn) against a database created from the Oncorhynchus mykiss isolate Swanson WGS data (NCBI acc. MSJN00000000.1) using an E-value cutoff 1E-100 – Taxonomic classification w/carried out by blasting (BLASTn) the remaining reads against an NCBI nt database using TimeLogic® DeCypher® server (Active Motif Inc., Carlsbad, CA, United States) with an E-value cutoff of 1E-5. The assignment of sequences to species were carried out by an in-house Python script using the nucl_gb.accessions2taxid (accession to taxid) and names.dmp (taxid to scientific names) files available from the resources at NCBI – We included Murine norovirus in the table despite it is not human pathogen
P12 Veterinary/Food – Quality assessment w/FastQC (Wingett and Andrews, 2018) – Quality/Adapter trimming w/cutadapt (part of MGmapper processing) (Martin, 2011) Taxonomic classification w/MGmapper (Databases: Bacteria, Bacteria_draft, Human Microbiome, Virus, Fungi, Protozoa, and MetaHitAssembly) (Petersen et al., 2017) None