Summary
Bam-readcount is a utility for generating low-level information about sequencing data at specific nucleotide positions. Originally designed to help filter genomic mutation calls, the metrics it outputs are useful as input for variant detection tools and for resolving ambiguity between variant callers (Koboldt et al., 2013a; Kothen-Hill et al., 2018). In addition, it has found broad applicability in diverse fields including tumor evolution, single-cell genomics, climate change ecology, and tracking community spread of SARS-CoV-2 (Miller et al., 2018; Müller et al., 2018; Paiva et al., 2020; Sun et al., 2020).
Statement of need
Bam-readcount is designed to meet two related needs related to genomic sequence analysis. The first is rapid genotyping of specific locations from a bam file, reporting not just the dominant bases, but counts of all bases. One context in which this is important is residual disease monitoring, where base changes with frequency below the sensitivity of standard genomic variant callers may still be informative. The second is reporting 15 key metrics for each reported base, including summarized mapping and base qualities, strandedness information, mismatch counts, and position within the reads. This information can be useful in a large number of contexts, with one frequent application being variant filtering, to remove false-positive calls, either with straightforward application of heuristic cutoffs or with semi-automated machine-learning approaches (Ainscough et al., 2018; Koboldt et al., 2013b). Another common use case is in ensemble variant calling situations where there is disagreement about base counts or key metrics at particular sites. Bam-readcount can be used to produce consistent, tool-agnostic metrics that are helpful in resolving such ambiguity (Anzar et al., 2019; Kockan et al., 2017; Kothen-Hill et al., 2018).
Implementation and results
The ongoing adoption of compressed data formats has necessitated additions to the code, and the version 1.0 release that we report on here utilizes an updated version of HTSlib to support rapid CRAM file access (Bonfield et al., 2021). This has also improved performance, and Bam-readcount can report on 100,000 randomly selected sites from a 30x whole-genome sequencing (WGS) BAM in around 5 minutes (Griffith, Miller, et al., 2015). Its performance scales nearly linearly with the number of genomic sites queried and average sequencing depth (Figure 1). Querying the same 100,000 sites from a BAM with 300x WGS takes 48 minutes, roughly 10x as long.
Figure 1:

Performance of Bam-readcount when querying randomly selected genomic positions from BAMs (left) or corresponding CRAMs (right) of varying sequencing depth. Colors correspond to average sequencing depth of the downsampled BAM/CRAM file.
Memory usage likewise is dependent on depth of sequencing, but still requires less than 1 GB of RAM for a 300x WGS BAM. Processing small CRAM files is somewhat slower than BAMs with comparable amounts of data, due to the increased CPU usage for decompression, but as depth increases, retrieval from disk becomes the bottleneck and operations on CRAMs exceed the speed of BAM. In our testing, on a fast SSD tier of networked disks, this transition occurs at a depth of about 180x. The problem is also embarrassingly parallel, so assuming adequate disk I/O, a roughly linear increase in speed can be achieved with a scatter/gather approach.
To lower barriers to adoption, we provide docker images for containerized workflows, and have developed a python wrapper that annotates a VCF file with read counts produced from this tool, available as part of the VAtools package (http://vatools.org).
Conclusions
Bam-readcount provides fast and accurate genomic readcounts and associated metrics, which allow it to fill a key niche in many genomic workflows. It has been adopted as a lightweight variant caller, finding known mutations in pre-leukemic phenotypes and used for detecting therapy-altering mutations from cell-free DNA (Wyatt et al., 2016; Xie et al., 2014). Viral researchers have tracked nucleotide changes across samples to understand diversity in Varicella Zoster Virus Encephalitis and to perform epidemiological surveillance in wastewater of SARS-CoV-2 (Depledge et al., 2018; Mondal et al., 2021). Those with RNA-sequencing data have found it useful for identifying allele-specific expression in cancer, or for enabling copy-number detection in single-cell RNA sequencing by retrieving allele frequencies (Cancer Genome Atlas Research Network et al., 2013; Müller et al., 2018). Its feature-rich output has also enabled deep learning approaches to variant calling and filtering (Ainscough et al., 2018; Anzar et al., 2019). In these roles, and other related ones, Bam-readcount has served as key infrastructure that supports groups of all sizes, from exploratory analyses to core facility pipelines to large multi-institution workflows (Griffith, Griffith, et al., 2015; Jensen et al., 2017; Sandmann et al., 2018). In the NCI’s Genomic Data Commons pipelines alone, its use in variant filtering means that it has been run on tens of thousands of cancer genomes.
Looking forward, we anticipate that as machine learning makes deeper inroads into genomics, the ability to extract highly informative features from large cohorts in a rapid manner will continue to make Bam-readcount useful for the next generation of genomics research.
The Bam-readcount tool is available at https://github.com/genome/bam-readcount and is shared under a MIT license to enable broad re-use.
Acknowledgements
This work was supported by the National Cancer Institute [R50CA211782 to CAM, P01CA101937 to TJL, K22CA188163 to OLG, 1U01CA209936 to OLG, U24CA237719 to OLG], the Edward P. Evans Foundation (to MJW), and the National Human Genome Research Institute [R00 HG007940 to MG]
Data availability
The WGS data used for benchmarking is available through dbGaP study phs000159, under sample id 452198/AML31. The summary data and scripts used to generate the figure are available at https://github.com/genome/bam-readcount/tree/joss-paper/figures. An archived snapshot of this 1.0 release is available at https://doi.org/10.5281/zenodo.5142454
Software
References
- Ainscough BJ, Barnell EK, Ronning P, Campbell KM, Wagner AH, Fehniger TA, Dunn GP, Uppaluri R, Govindan R, Rohan TE, Griffith M, Mardis ER, Swamidass SJ, & Griffith OL (2018). A deep learning approach to automate refinement of somatic variant calling from cancer sequencing data. Nat. Genet, 50(12), 1735–1743. 10.1038/s41588-018-0257-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- Anzar I, Sverchkova A, Stratford R, & Clancy T (2019). NeoMutate: An ensemble machine learning framework for the prediction of somatic mutations in cancer. BMC Med. Genomics, 12(1), 63. 10.1186/s12920-019-0508-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bonfield JK, Marshall J, Danecek P, Li H, Ohan V, Whitwham A, Keane T, & Davies RM (2021). HTSlib: C library for reading/writing high-throughput sequencing data. Gigascience, 10(2). 10.1093/gigascience/giab007 [DOI] [Google Scholar]
- Cancer Genome Atlas Research Network, Ley TJ, Miller C, Ding L, Raphael BJ, Mungall AJ, Robertson AG, Hoadley K, Triche TJ Jr, Laird PW, Baty JD, Fulton LL, Fulton R, Heath SE, Kalicki-Veizer J, Kandoth C, Klco JM, Koboldt DC, Kanchi K-L, … Eley G (2013). Genomic and epigenomic landscapes of adult de novo acute myeloid leukemia. N. Engl. J. Med, 368(22), 2059–2074. 10.1056/NEJMoa1301689 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Depledge DP, Cudini J, Kundu S, Atkinson C, Brown JR, Haque T, Houldcroft CJ, Koay ES, McGill F, Milne R, Whitfield T, Tang JW, Underhill G, Bergstrom T, Norberg P, Goldstein R, Solomon T, & Breuer J (2018). High viral diversity and mixed infections in cerebral spinal fluid from cases of varicella zoster virus encephalitis. J. Infect. Dis, 218(10), 1592–1601. 10.1093/infdis/jiy358 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Griffith M, Griffith OL, Smith SM, Ramu A, Callaway MB, Brummett AM, Kiwala MJ, Coffman AC, Regier AA, Oberkfell BJ, Sanderson GE, Mooney TP, Nutter NG, Belter EA, Du F, Long RL, Abbott TE, Ferguson IT, Morton DL, … Wilson RK (2015). Genome modeling system: A knowledge management platform for genomics. PLoS Comput. Biol, 11(7), e1004274. 10.1371/journal.pcbi.1004274 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Griffith M, Miller CA, Griffith OL, Krysiak K, Skidmore ZL, Ramu A, Walker JR, Dang HX, Trani L, Larson DE, Demeter RT, Wendl MC, McMichael JF, Austin RE, Magrini V, McGrath SD, Ly A, Kulkarni S, Cordes MG, … Wilson RK (2015). Optimizing cancer genome sequencing and analysis. Cell Syst, 1(3), 210–223. 10.1016/j.cels.2015.08.015 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jensen MA, Ferretti V, Grossman RL, & Staudt LM (2017). The NCI genomic data commons as an engine for precision medicine. Blood, 130(4), 453–459. 10.1182/blood-2017-03-735654 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Koboldt DC, Larson DE, & Wilson RK (2013a). Using VarScan 2 for germline variant calling and somatic mutation detection. Curr. Protoc. Bioinformatics, 44, 15.4.1–17. 10.1002/0471250953.bi1504s44 [DOI] [Google Scholar]
- Koboldt DC, Larson DE, & Wilson RK (2013b). Using VarScan 2 for germline variant calling and somatic mutation detection. Current Protocols in Bioinformatics, 44(1), 15–14. 10.1093/bioinformatics/bty316 [DOI] [Google Scholar]
- Kockan C, Hach F, Sarrafi I, Bell RH, McConeghy B, Beja K, Haegert A, Wyatt AW, Volik SV, Chi KN, Collins CC, & Sahinalp SC (2017). SiNVICT: Ultra-sensitive detection of single nucleotide variants and indels in circulating tumour DNA. Bioinformatics, 33(1), 26–34. 10.1093/bioinformatics/btw536 [DOI] [PubMed] [Google Scholar]
- Kothen-Hill ST, Zviran A, Schulman RC, Deochand S, Gaiti F, Maloney D, Huang KY, Liao W, Robine N, Omans ND, & Landau DA (2018, February). Deep learning mutation prediction enables early stage lung cancer detection in liquid biopsy.
- Miller CA, Dahiya S, Li T, Fulton RS, Smyth MD, Dunn GP, Rubin JB, & Mardis ER (2018). Resistance-promoting effects of ependymoma treatment revealed through genomic analysis of multiple recurrences in a single patient. Cold Spring Harb Mol Case Stud, 4(2). 10.1101/mcs.a002444 [DOI] [Google Scholar]
- Mondal S, Feirer N, Brockman M, Preston MA, Teter SJ, Ma D, Goueli SA, Moorji S, Saul B, & Cali JJ (2021). A direct capture method for purification and detection of viral nucleic acid enables epidemiological surveillance of SARS-CoV-2. Sci. Total Environ, 795, 148834. 10.1101/2021.05.06.21256753 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Müller S, Cho A, Liu SJ, Lim DA, & Diaz A (2018). CONICS integrates scRNA-seq with DNA sequencing to map gene expression to tumor sub-clones. Bioinformatics, 34(18), 3217–3219. 10.1093/bioinformatics/bty316 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Paiva MHS, Guedes DRD, Docena C, Bezerra MF, Dezordi FZ, Machado LC, Krokovsky L, Helvecio E, da Silva AF, Vasconcelos LRS, Rezende AM, Silva S. J. R. da, Sales K. G. S. da, Sá B. S. L. F. de, da Cruz DL, Cavalcanti CE, Neto A. de M., Silva C. T. A. da, Mendes RPG, … Wallau GL (2020). Multiple introductions followed by ongoing community spread of SARS-CoV-2 at one of the largest metropolitan areas of northeast brazil. Viruses, 12(12). 10.3390/v12121414 [DOI] [Google Scholar]
- Sandmann S, Karimi M, Graaf A. O. de, Rohde C, Göllner S, Varghese J, Ernsting J, Walldin G, van der Reijden BA, Müller-Tidow C, Malcovati L, Hellström-Lindberg E, Jansen JH, & Dugas M (2018). appreci8: A pipeline for precise variant calling integrating 8 tools. Bioinformatics, 34(24), 4205–4212. 10.1093/bioinformatics/bty518 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sun Y, Bossdorf O, Grados RD, Liao Z, & Müller-Schärer H (2020). Rapid genomic and phenotypic change in response to climate warming in a widespread plant invader. Glob. Chang. Biol, 26(11), 6511–6522. 10.1111/gcb.15291 [DOI] [PubMed] [Google Scholar]
- Wyatt AW, Azad AA, Volik SV, Annala M, Beja K, McConeghy B, Haegert A, Warner EW, Mo F, Brahmbhatt S, Shukin R, Le Bihan S, Gleave ME, Nykter M, Collins CC, & Chi KN (2016). Genomic alterations in Cell-Free DNA and enzalutamide resistance in Castration-Resistant prostate cancer. JAMA Oncol, 2(12), 1598–1606. 10.1001/jamaoncol.2016.0494 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Xie M, Lu C, Wang J, McLellan MD, Johnson KJ, Wendl MC, McMichael JF, Schmidt HK, Yellapantula V, Miller CA, Ozenberger BA, Welch JS, Link DC, Walter MJ, Mardis ER, Dipersio JF, Chen F, Wilson RK, Ley TJ, & Ding L (2014). Age-related mutations associated with clonal hematopoietic expansion and malignancies. Nat. Med, 20(12), 1472–1478. 10.1038/nm.3733 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The WGS data used for benchmarking is available through dbGaP study phs000159, under sample id 452198/AML31. The summary data and scripts used to generate the figure are available at https://github.com/genome/bam-readcount/tree/joss-paper/figures. An archived snapshot of this 1.0 release is available at https://doi.org/10.5281/zenodo.5142454
