Skip to main content
Molecular Therapy. Methods & Clinical Development logoLink to Molecular Therapy. Methods & Clinical Development
. 2020 Mar 30;17:752–757. doi: 10.1016/j.omtm.2020.03.024

VSeq-Toolkit: Comprehensive Computational Analysis of Viral Vectors in Gene Therapy

Saira Afzal 1,, Raffaele Fronza 2, Manfred Schmidt 1,2
PMCID: PMC7177155  PMID: 32346552

Abstract

Viral vector characterization and analysis are important components for the development of safe gene therapeutic products, elucidating the potential genotoxic and immunogenic effects of vectors and establishing their safety profiles. Here, we present VSeq-Toolkit, which offers varying analysis modes for viral gene therapy data. The first mode determines the undesirable known contaminants and their frequency in viral preparations or other sequencing data. The second mode is designed for the analysis of intra-vector fusion breakpoints and the third mode for unraveling the viral-host fusion events distribution. Analysis modes of our toolkit can be executed independently or together and allow the analysis of multiple viral vectors concurrently. It has been designed and evaluated for the analysis of short read high-throughput sequencing data, including whole-genome or targeted sequencing. VSeq-Toolkit is developed in Perl and Bash programming languages and is available at https://github.com/CompMeth/VSeq-Toolkit.

Graphical Abstract

graphic file with name fx1.jpg


In-depth analysis of viral vectors is a key component in assessment of safety and efficacy of gene therapy. Afzal et al. describes a computational platform to deepen the understanding of viral vectors in pre-clinical and clinical gene therapy data, as well as dissecting the contaminants in viral preparations.

Introduction

Advancements in gene therapy and approval of products for retinal dystrophy and lipoprotein lipase deficiency reinforce the promises to treat challenging disorders ranging from hereditary, infectious, metabolic, cardiovascular, and ophthalmologic to various cancer types.1, 2, 3, 4, 5, 6 The use of viruses as carriers of genetic products requires extensive and in-depth understanding of viral vectors starting from viral preparation until clinical employment.7, 8, 9, 10 The viral preparations must be strictly quality assessed11,12 for exclusion of any viral or cellular contaminants and impurities to exclude any immunogenic effects. Additionally, during the pre-clinical and clinical settings, viral gene therapy should be closely monitored for any insertional mutagenesis effects.13,14 Analyzing and tracking viral integration events within the host genome helps to estimate the likelihood of viral vector safety in relevance to deregulation of any tumor suppressor or oncogene.15, 16, 17 Certain vectors, for example adeno-associated viruses (AAVs), are susceptible to vector-vector rearranged junction formation.18, 19, 20 Therefore, elucidating vector-vector breakpoint fusions is important in enhancing the understanding of viral vectors.

Here, we present a comprehensive toolkit for the analysis of viral gene therapy data starting from the analysis of viral vector preparations or any contaminant estimations to the pre-clinical and clinical monitoring of vector-vector or vector-host fusions. Various tools are available that deal with the vector-host fusions or integration site (IS) analysis, in the context of viral cancers mainly21, 22, 23, 24, 25 and in gene therapy focusing on PCR-based methods as linear-amplification-mediated (LAM) PCR26, 27, 28, 29, 30, 31 and targeted sequencing.28 However, here, our aim is to provide an easy-to-use tool suite that can provide multiple analyses, including contaminant distribution within the data, intra-viral vector fusion events that were not previously addressed by available methods, and viral-genome fusion events in whole-genome sequencing (WGS) or targeted sequencing data. In addition, we aimed for the analysis of multiple viral vectors simultaneously within a single sample. The analysis modes of VSeq-Toolkit can be used independently or jointly for reliable and precise characterization and estimation of contaminants and their frequency, vector-vector, and vector-genome fusion distributions. Our method is designed for Illumina short read paired-end (PE) data and can be reliably used for the pre-clinical assessment to clinical monitoring of viral vectors risk and safety profiles.

Results

To evaluate the reliability and accuracy of VSeq-Toolkit (Figure 1), in silico datasets generated for each respective analysis mode were analyzed. The 250 bp PE dataset D1 comprising of 6,000 reads was analyzed with contaminant analysis mode, which was able to detect all respective sequences correctly, including 1,550 and 800 contaminant one and two sequences, respectively, 1,750 vector sequences, and 1,900 reads from human genome assembly (hg38) (Figure 2A). Similarly, D2 and D3 in silico datasets were analyzed with vector-vector and vector-host fusion analysis modes, respectively. The vector-vector breakpoints of 250 bp PE reads were accurately identified for each vector reference sequence. Vector-host fusion analysis mode also identified all fusion event reads in 250 bp PE data (Figure 2A; Data S1). The datasets comprising of contaminant, vector-vector, and vector-host 150 bp PE reads were also evaluated that showed similarly high precision and recall values. Additionally, D1A, D2A, and D3A datasets with 0.25% error rates were analyzed in a similar manner and showed reliability of toolkit modes (Figure 2A). We estimated the percentages of read pairs (contaminant analysis) or fusion events (vector-vector and vector-host) within ± 3 bp of the expected position. In case of contaminant analysis mode (D1), all read-pair positions were detected at the expected position, whereas in case of vector-vector (D2) and vector-host (D3) modes, approximately >76% fusion breakpoint positions were detected at expected positions, 19% within ± 1 bp difference and 3.5%–4.5% in ± 2 bp of expected positions (Figure 2B).

Figure 1.

Figure 1

Schematics of Basic VSeq-Toolkit Workflow

The VSeq-Toolkit is comprised of five modules with three main analysis modes. The input module takes FASTQ paired-end reads, and the pre-processing module performs the quality filtering and trimming. The contaminant analysis mode allows simultaneous detection of contaminants or vector in the sample along with their respective distribution and fragment size statistics. The vector-vector fusion mode provides the analysis of vector rearranged breakpoint events. The vector-genome fusion mode analyzes the distribution of vector integration sites within the host genome.

Figure 2.

Figure 2

Performance and Reliability Evaluation of VSeq-Toolkit Modes on In Silico Datasets

(A) Statistical measures, recall and precision estimation on in silico datasets are depicted here. Each mode of VSeq-Toolkit; contaminant analysis, vector-vector fusion, and vector-host fusion is evaluated with respective in silico datasets without errors D1, D2, and D3 and with 0.25% error rate D1A, D2A, and D3A. (B) The histogram represents the percentages of detected read pair positions (in contaminant analysis mode on D1) or fusion positions (in vector-vector analysis mode on D2 and vector-host analysis mode on D3) within a ± 3 bp range of the expected positions. For each dataset, both the positions were evaluated for each read pair (in contaminant analysis mode) or each fusion event region (in vector-vector and vector-host analysis modes).

We had additionally analyzed experimental datasets from previously published studies. The S1 experimental dataset,28 i.e., a control lentiviral vector (LV) sample with three known vector-genome ISs, was evaluated with vector-host fusion mode. Similarly, AAV-based publicly available samples32 S2 and S3 were analyzed to mainly show the performance of vector-vector fusion analysis mode. In the case of S1 sample of about 37 million reads, the three expected vector-genome integration events were accurately identified with significantly high sequence count numbers (Data S2). In the case of LV sample, as expected, no significant vector-vector fusions were detected, and only one breakpoint was detected, which could be most likely an experimental artifact sequence. On the other hand, the AAV-based samples S2 and S3, comprising of 5,561,416 and 2,540,471 reads, respectively, were analyzed to show the performance of vector-vector fusion analysis mode. The significant numbers of vector-vector breakpoints were detected in both samples, mainly at the inverted terminal repeat (ITR) regions, as shown in Figure 3. Here, the breakpoints were evaluated by considering zero and one as a cutoff value for sequence count.

Figure 3.

Figure 3

Distribution of Vector-Vector Breakpoint Junctions

The vector-vector fusion breakpoints analyzed in the two experimental datasets S2 and S3 are presented here. The x axis represents the vector reference, and the y axis represents the frequency of breakpoint positions in a logarithmic scale. The dashed line represents the cutoff value of one for fusion sequence count. The breakpoints are accumulated mainly in the ITR regions.

We have also evaluated the computational time efficiency of each mode of VSeq-Toolkit on in silico datasets DS1.1, DS2.1, and DS3.1 for contaminant, vector-vector, and vector-host fusion analysis modes, respectively (Figure 4A). Each of these datasets consists of 100k reads (250 bp PE). Additionally, we have also evaluated the time consumption for analyzing experimental datasets S1 (for vector-host fusions), and S2 and S3 (for vector-vector fusions). The S1 dataset of about 37 million reads was analyzed for vector-host fusions in 71 min. The S2 and S3 (about 5.5 and 2.5 million reads) were analyzed for vector-vector fusions within less than 12 and 6 min, respectively (Figure 4B). The S1 dataset was 100 bp PE, whereas S2 and S3 were 250 bp PE.

Figure 4.

Figure 4

Computational Performance of VSeq-Toolkit Modes

(A) Representation of run time for contaminant analysis, vector-vector, and vector-host fusion analysis modes of VSeq-Toolkit on respective in silico datasets D1.1, D2.1, and D3.1, each comprising of 100k reads. (B) Time consumption for the S1 experimental dataset, comprising of approximately 37 million reads analyzed by vector-host fusion (mode 3), and for the S2 and S3 experimental datasets, consisting of about 5.5 and 2.5 million reads, respectively, with vector-vector fusion analysis (mode 2). (C) Time taken by different processing steps of vector-host analysis mode for the S1 dataset (VH1, quality control; VH2, vector-genome mapping; VH3, extraction of vector-genome candidates; VH4, selection and processing; VH5, parameter based estimation and filtering; VH6, clustering, annotation, and results generation). (D) Time consumed per each main module of vector-vector fusion analysis mode for S2 and S3 datasets (VV1, quality control; VV2, vector mapping; VH3, extraction of potential vector-vector reads; VV4, selection, filtering, and processing; VV5, parameter estimation and feature designation; and VV6, results generation).

Additionally, we measured the time consumption per each main module of the toolkit for vector-host (S1 dataset) and vector-vector (S2 and S3 datasets) analysis modes. In case of vector-host analysis for S1, the highest time is consumed by mapping step, followed by quality check (Figure 4C). In vector-vector analysis the major time consumption was in quality check module followed by mapping (Figure 4D).

Discussion

We have described VSeq-Toolkit, which combines the functionalities to cover a range of viral gene therapy data analysis requirements at one platform in a comprehensive and computationally efficient manner. It provides specific modes for analyzing contaminants distribution, vector-vector rearranged junctions, and vector-genome breakpoint distribution in high-throughput sequencing data. The toolkit allows characterization of the contaminants in viral preparations along with determination of their respective frequencies and fragment sizes. The vector-vector fusion analysis mode allows characterization of vector breakpoint events. Currently, to the best of our knowledge, no tools are available so far that are specifically designed and evaluated for vector breakpoint distribution profiling for gene therapy data.

The viral-genome fusion mode of the toolkit helps to unravel the distribution of ISs of vectors within the respective genomes. Although a wide range of tools are available for IS analysis specifically suitable for LAM-PCR26, 27, 28, 29, 30, 31 and targeted or WGS with focus on viral cancers21, 22, 23, 24, 25 and gene therapy.28 Here, we provide the analysis of different aspects of gene therapy data with a single toolkit for WGS or targeted sequencing data. VSeq-Toolkit provides an added advantage that both vector-vector fusions and vector-genome integration events can be investigated together. At the first stage, vector-vector breakpoint reads are analyzed, followed by vector-genome fusion events analysis. The exclusion of vector-vector fusions at the first step increases the specificity of subsequent vector-genome distribution analysis, along with providing a useful and detailed vector breakpoint profile. In contrast to the available viral IS tools, the vector-fusion mode of our toolkit investigates in depth not only the genomic part of the fusion read, but also the vector region that provides a more transparent view of vector positions and their distribution. Furthermore, as compared to the other available methods,21, 22, 23, 24, 25,28 the toolkit modes allow analysis of multiple viral references within a sample concurrently. This feature has broader implications for analysis of viral cancers as well.

We have evaluated each analysis mode of toolkit with in silico datasets and have shown the reliability and accuracy of our method. Additionally, we depicted the performance of analysis modes on LV and AAV experimental datasets. The toolkit is easy to use and allows multiple adjustable parameters that can be tailored according to analysis requirements. Additionally, the specificity and sensitivity levels for selection of reads are user adjustable. A range of other configurable parameters are also provided, including minimum vector or genome region length, maximum number of non-mapped or overlapped bases between the fusion events, the minimum identity percentage of each fusion event region, etc. In the case of vector-host IS analysis, the clustering of reads based on genomic positions can be performed within the required range. The results are reported in an easy-to-read format for the end-user with detailed information by each respective mode. The toolkit can be employed with any host genome or viral vector sequence as reference.

In-depth analysis of various dimensions of gene therapy data is an important step to assess safety and efficacy of viral-vector based gene therapy. VSeq-Toolkit provides a compact workflow to analyze contaminant distribution, viral vector breakpoint profiles, and viral-genome ISs in a reliable and highly time-efficient way. It additionally has broader implications in analyzing next-generation sequencing data for the presence of viral and non-viral contaminants. Moreover, it allows investigation of data from insertional mutagenesis screens, viral cancers, and infectious diseases.

Materials and Methods

Toolkit

The input module of VSeq-Toolkit accepts FASTQ PE data that is processed initially with the quality-control module for quality filtering of reads and trimming of sequencing adapters with skewer.33 In the next step, data is processed, based on user requirements, through one or all of the main analysis modes: contaminant analysis, vector-vector fusion analysis, and vector-host fusion analysis (Figure 1). The contaminant analysis mode characterizes the undesired known contaminants and estimates their abundance. The filtered and trimmed dataset is first mapped with the reference genome by the Burrows-Wheeler Aligner (BWA) MEM aligner.34 Multiple potential contaminants can be provided as a concatenated single reference file, if desired, along with vector or host genome sequences. In the next steps, the sequence alignment/map (SAM)35 file is processed for duplicates removal with Samtools, and correctly mapped read pairs are selected for further processing. Subsequently, the distribution of each contaminant sequence, vector, and host is analyzed, and frequency values are calculated. The fragment size distribution for the contaminants and vector or host genome references is estimated from the SAM file for the final processed subset of reads. The read number are estimated within six categories of fragment sizes, ranging from less than 50 to greater than 1,000. At the final step, the contaminant sequences are removed from the actual raw data files, and cleaned FASTQ files are generated for further investigation either by external tools or VSeq-Toolkit.

The second mode of vector-vector fusion analysis is designed to detect the rearranged vector sequence breakpoints. The first step of this analysis mode performs mapping of reads with single or multiple combined vector references. In the subsequent steps, the soft-clip reads are extracted and potential vector-vector fusion reads within the same vector reference are selected. These are filtered further based on the user configurable stringency levels. The candidate reads are then processed in the next steps to determine individual vector-vector breakpoint positions that undergo a next round of final selection based on user-adjustable parameters for minimum length, identity, and maximum overlap or distance among fused regions. Finally, the detailed information about each fusion breakpoint is generated in the result files.

In the vector-host fusion analysis mode, the first module performs mapping of data with the combined host and vector reference sequences. The soft-clipped read pairs with alternative mapping with either vector or genome are extracted, and genome-genome rearranged reads are excluded at this stage. This subset of reads is subsequently processed for selection of correctly mapped read pairs, followed by a number of steps leading to the determination of the exact fusion junction positions for vector and genome regions. Similar to vector-vector fusion mode, the vector-genome final fusion reads are also filtered based on user-customizable parameters. Each of the viral vector-genome files is then processed for position-based clustering followed by annotation with custom scripts and bedtools.36 Finally, the resulted IS output per each vector type is concatenated and provided as a result file.

In sensitive mode for vector-vector and vector-host fusion analysis, the sequencing reads that do not show alternative mapping are processed with a more sensitive remapping module to rescue candidate fusions. In all three modes, each sequence is tagged with a unique and non-unique label, based on either a read pair (first mode) or the regions of a read that contribute to afusion event (second and third modes) is mapped uniquely or non-uniquely in the respective reference.

Datasets

We generated in silico datasets for analyzing the performance of all three modes of VSeq-Toolkit. For contaminant analysis mode, the in silico dataset was designed by extracting random sequences from two reference sequences regarded as contaminants, a vector reference and hg38. A random region of 500 bp length was extracted from each of the aforementioned references. The 250 bp PE dataset comprising of a total of 6,000 reads (D1) was designed that includes 1,550 and 800 read pairs for first and second contaminant sequences, respectively. In addition, the dataset contains 1,750 vector reads and 1,900 reads from hg38. For vector-vector fusion analysis, the in silico dataset was created by randomly extracting the sequences of 500 bp from the two different vectors to create 250 bp PE end data, then we excluded 50 bp region from these reads and inserted a known vector region to simulate vector-vector fusion reads. This dataset is comprised of 215 reads (D2) from each of the first and second vectors. For vector-host fusion analysis, 19,550 random sequences were extracted from the hg38 genome and 250 PE reads were created. For each vector type, different regions of the vector were selected and introduced within the reads, resulting in final dataset of 19,550 reads (D3). Furthermore, these datasets were simulated with 0.25% errors to generate D1A, D2A, and D3A datasets, respectively, for each mode. The vector-vector and vector-host datasets were generated with 150 bp PE read lengths as well. In addition, we also generated in a similar manner DS1.1 (comprising reads from two contaminants and one vector sequence, excluding hg38), DS2.1, and DS3.1 for contaminant analysis, vector-vector, and vector-host fusion modes, respectively. Each of these datasets was comprised of 100k reads to evaluate the time-efficiency for each VSeq-Toolkit analysis mode.

The analysis was performed on in silico datasets with respective modules, and basic statistical measures, including precision and recall, were estimated. An event is considered true positive if it lies within ± 3 bp of the expected position, false negative if any event is not detected by the respective module of the toolkit, and false positive if other than the expected positions are reported. In the case of vector-vector and vector-genome fusion analysis, the statistical-measures estimation takes into account both positions of the fusion event, i.e., the position of both vector regions that makes one fusion event in the case of vector-vector analysis, and in the case of the vector-genome, the fusion positions of vector as well as genomic regions are taken into account.

Additionally, we have analyzed the experimental datasets to depict the reliability and performance of vector-vector and vector-host fusion modes of the toolkit. We have analyzed a control sample of LV-transduced HeLa cells with three integration events comprising of approximately 37 million PE reads reported in a previous study.28 In this study, we have regarded this sample as S1. Furthermore, we have analyzed two publicly available datasets (SRR7085814 and SRR7085812)32 with AAV generated by targeted enrichment sequencing (https://www.ncbi.nlm.nih.gov/Traces/study/?acc=SRP144114). Here, in this work, we refer to these samples as S2 and S3, respectively.

Author Contributions

S.A. and M.S. conceived the project. S.A. designed and developed the methods, performed the analysis, and wrote the manuscript. M.S. and R.F. revised the manuscript. M.S. provided the resources.

Conflicts of Interest

M.S. is co-founder and CEO of GeneWerk GmbH.

Acknowledgments

We are grateful to our colleagues at the Translational Oncology group for helpful feedback and discussions.

Footnotes

Supplemental Information can be found online at https://doi.org/10.1016/j.omtm.2020.03.024.

Supplemental Information

Data S1. Detail of In Silico Dataset (D1, D2, and D3) Analysis Output
mmc1.xlsx (2MB, xlsx)
Data S2. Detail of S1 Data Analysis Output
mmc2.xlsx (10.2KB, xlsx)

References

  • 1.Biffi A., Montini E., Lorioli L., Cesani M., Fumagalli F., Plati T., Baldoli C., Martino S., Calabria A., Canale S. Lentiviral hematopoietic stem cell gene therapy benefits metachromatic leukodystrophy. Science. 2013;341:1233158. doi: 10.1126/science.1233158. [DOI] [PubMed] [Google Scholar]
  • 2.Linette G.P., Stadtmauer E.A., Maus M.V., Rapoport A.P., Levine B.L., Emery L., Litzky L., Bagg A., Carreno B.M., Cimino P.J. Cardiovascular toxicity and titin cross-reactivity of affinity-enhanced T cells in myeloma and melanoma. Blood. 2013;122:863–871. doi: 10.1182/blood-2013-03-490565. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.MacLaren R.E., Groppe M., Barnard A.R., Cottriall C.L., Tolmachova T., Seymour L., Clark K.R., During M.J., Cremers F.P., Black G.C. Retinal gene therapy in patients with choroideremia: initial findings from a phase 1/2 clinical trial. Lancet. 2014;383:1129–1137. doi: 10.1016/S0140-6736(13)62117-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Naldini L. Gene therapy returns to centre stage. Nature. 2015;526:351–360. doi: 10.1038/nature15818. [DOI] [PubMed] [Google Scholar]
  • 5.Kassner U., Hollstein T., Grenkowitz T., Wühle-Demuth M., Salewsky B., Demuth I., Dippel M., Steinhagen-Thiessen E. Gene Therapy in Lipoprotein Lipase Deficiency: Case Report on the First Patient Treated with Alipogene Tiparvovec Under Daily Practice Conditions. Hum. Gene Ther. 2018;29:520–527. doi: 10.1089/hum.2018.007. [DOI] [PubMed] [Google Scholar]
  • 6.Ludwig P.E., Freeman S.C., Janot A.C. Novel stem cell and gene therapy in diabetic retinopathy, age related macular degeneration, and retinitis pigmentosa. Int. J. Retina Vitreous. 2019;5:7. doi: 10.1186/s40942-019-0158-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Connolly J.B. Lentiviruses in gene therapy clinical research. Gene Ther. 2002;9:1730–1734. doi: 10.1038/sj.gt.3301893. [DOI] [PubMed] [Google Scholar]
  • 8.Zhang X., Godbey W.T. Viral vectors for gene delivery in tissue engineering. Adv. Drug Deliv. Rev. 2006;58:515–534. doi: 10.1016/j.addr.2006.03.006. [DOI] [PubMed] [Google Scholar]
  • 9.Lukashev A.N., Zamyatnin A.A., Jr. Viral vectors for gene therapy: Current state and clinical perspectives. Biochemistry (Mosc.) 2016;81:700–708. doi: 10.1134/S0006297916070063. [DOI] [PubMed] [Google Scholar]
  • 10.Finer M., Glorioso J. A brief account of viral vectors and their promise for gene therapy. Gene Ther. 2017;24:1–2. doi: 10.1038/gt.2016.71. [DOI] [PubMed] [Google Scholar]
  • 11.Merten O.-W., Wright J.F. Towards routine manufacturing of gene therapy drugs. Mol. Ther. Methods Clin. Dev. 2016;3:16021. doi: 10.1038/mtm.2016.21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.van der Loo J.C.M., Wright J.F. Progress and challenges in viral vector manufacturing. Hum. Mol. Genet. 2016;25(R1):R42–R52. doi: 10.1093/hmg/ddv451. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Howe S.J., Mansour M.R., Schwarzwaelder K., Bartholomae C., Hubank M., Kempski H., Brugman M.H., Pike-Overzet K., Chatters S.J., de Ridder D. Insertional mutagenesis combined with acquired somatic mutations causes leukemogenesis following gene therapy of SCID-X1 patients. J. Clin. Invest. 2008;118:3143–3150. doi: 10.1172/JCI35798. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Wu C., Dunbar C.E. Stem cell gene therapy: the risks of insertional mutagenesis and approaches to minimize genotoxicity. Front. Med. 2011;5:356–371. doi: 10.1007/s11684-011-0159-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Schmidt M., Schwarzwaelder K., Bartholomae C., Zaoui K., Ball C., Pilz I., Braun S., Glimm H., von Kalle C. High-resolution insertion-site analysis by linear amplification-mediated PCR (LAM-PCR) Nat. Methods. 2007;4:1051–1057. doi: 10.1038/nmeth1103. [DOI] [PubMed] [Google Scholar]
  • 16.Hacein-Bey-Abina S., Garrigue A., Wang G.P., Soulier J., Lim A., Morillon E., Clappier E., Caccavelli L., Delabesse E., Beldjord K. Insertional oncogenesis in 4 patients after retrovirus-mediated gene therapy of SCID-X1. J. Clin. Invest. 2008;118:3132–3142. doi: 10.1172/JCI35700. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Negre O., Bartholomae C., Beuzard Y., Cavazzana M., Christiansen L., Courne C., Deichmann A., Denaro M., de Dreuzy E., Finer M. Preclinical evaluation of efficacy and safety of an improved lentiviral vector for the treatment of β-thalassemia and sickle cell disease. Curr. Gene Ther. 2015;15:64–81. doi: 10.2174/1566523214666141127095336. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Nowrouzi A., Penaud-Budloo M., Kaeppel C., Appelt U., Le Guiner C., Moullier P., von Kalle C., Snyder R.O., Schmidt M. Integration frequency and intermolecular recombination of rAAV vectors in non-human primate skeletal muscle and liver. Mol. Ther. 2012;20:1177–1186. doi: 10.1038/mt.2012.47. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Colella P., Ronzitti G., Mingozzi F. Emerging Issues in AAV-Mediated In Vivo Gene Therapy. Mol. Ther. Methods Clin. Dev. 2017;8:87–104. doi: 10.1016/j.omtm.2017.11.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Hanlon K.S., Kleinstiver B.P., Garcia S.P., Zaborowski M.P., Volak A., Spirig S.E., Muller A., Sousa A.A., Tsai S.Q., Bengtsson N.E. High levels of AAV vector integration into CRISPR-induced DNA breaks. Nat. Commun. 2019;10:4439. doi: 10.1038/s41467-019-12449-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Chen Y., Yao H., Thompson E.J., Tannir N.M., Weinstein J.N., Su X. VirusSeq: software to identify viruses and their integration sites using next-generation sequencing of human cancer tissue. Bioinformatics. 2013;29:266–267. doi: 10.1093/bioinformatics/bts665. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Wang Q., Jia P., Zhao Z. VirusFinder: software for efficient and accurate detection of viruses and their integration sites in host genomes through next generation sequencing data. PLoS ONE. 2013;8:e64465. doi: 10.1371/journal.pone.0064465. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Ho D.W.H., Sze K.M.F., Ng I.O.L. Virus-Clip: a fast and memory-efficient viral integration site detection tool at single-base resolution with annotation capability. Oncotarget. 2015;6:20959–20963. doi: 10.18632/oncotarget.4187. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Forster M., Szymczak S., Ellinghaus D., Hemmrich G., Rühlemann M., Kraemer L., Mucha S., Wienbrandt L., Stanulla M., Franke A., UFO Sequencing Consortium within I-BFM Study Group Vy-PER: eliminating false positive detection of virus integration events in next generation sequencing data. Sci. Rep. 2015;5:11534. doi: 10.1038/srep11534. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Shieh F.S., Jongeneel P., Steffen J.D., Lin S., Jain S., Song W., Su Y.H. ChimericSeq: An open-source, user-friendly interface for analyzing NGS data to identify and characterize viral-host chimeric sequences. PLoS ONE. 2017;12:e0182843. doi: 10.1371/journal.pone.0182843. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Arens A., Appelt J.-U., Bartholomae C.C., Gabriel R., Paruzynski A., Gustafson D., Cartier N., Aubourg P., Deichmann A., Glimm H. Bioinformatic clonality analysis of next-generation sequencing-derived viral vector integration sites. Hum. Gene Ther. Methods. 2012;23:111–118. doi: 10.1089/hgtb.2011.219. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Hocum J.D., Battrell L.R., Maynard R., Adair J.E., Beard B.C., Rawlings D.J., Kiem H.P., Miller D.G., Trobridge G.D. VISA--Vector Integration Site Analysis server: a web-based server to rapidly identify retroviral integration sites from next-generation sequencing. BMC Bioinformatics. 2015;16:212. doi: 10.1186/s12859-015-0653-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Afzal S., Wilkening S., von Kalle C., Schmidt M., Fronza R. GENE-IS: Time-Efficient and Accurate Analysis of Viral Integration Events in Large-Scale Gene Therapy Data. Mol. Ther. Nucleic Acids. 2017;6:133–139. doi: 10.1016/j.omtn.2016.12.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Spinozzi G., Calabria A., Brasca S., Beretta S., Merelli I., Milanesi L., Montini E. VISPA2: a scalable pipeline for high-throughput identification and annotation of vector integration sites. BMC Bioinformatics. 2017;18:520. doi: 10.1186/s12859-017-1937-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Berry C.C., Nobles C., Six E., Wu Y., Malani N., Sherman E., Dryga A., Everett J.K., Male F., Bailey A. INSPIIRED: Quantification and Visualization Tools for Analyzing Integration Site Distributions. Mol. Ther. Methods Clin. Dev. 2016;4:17–26. doi: 10.1016/j.omtm.2016.11.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Calabria A., Beretta S., Merelli I., Spinozzi G., Brasca S., Pirola Y., Benedicenti F., Tenderini E., Bonizzoni P., Milanesi L., Montini E. γ-TRIS: a graph-algorithm for comprehensive identification of vector genomic insertion sites. Bioinformatics. 2020;36:1622–1624. doi: 10.1093/bioinformatics/btz747. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Senís E., Mosteiro L., Wilkening S., Wiedtke E., Nowrouzi A., Afzal S., Fronza R., Landerer H., Abad M., Niopek D. AAVvector-mediated in vivo reprogramming into pluripotency. Nat. Commun. 2018;9:2651. doi: 10.1038/s41467-018-05059-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Jiang H., Lei R., Ding S.-W., Zhu S. Skewer: a fast and accurate adapter trimmer for next-generation sequencing paired-end reads. BMC Bioinformatics. 2014;15:182. doi: 10.1186/1471-2105-15-182. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Li H., Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics. 2009;25:1754–1760. doi: 10.1093/bioinformatics/btp324. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Li H., Handsaker B., Wysoker A., Fennell T., Ruan J., Homer N., Marth G., Abecasis G., Durbin R., 1000 Genome Project Data Processing Subgroup The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25:2078–2079. doi: 10.1093/bioinformatics/btp352. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Quinlan A.R., Hall I.M. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 2010;26:841–842. doi: 10.1093/bioinformatics/btq033. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data S1. Detail of In Silico Dataset (D1, D2, and D3) Analysis Output
mmc1.xlsx (2MB, xlsx)
Data S2. Detail of S1 Data Analysis Output
mmc2.xlsx (10.2KB, xlsx)

Articles from Molecular Therapy. Methods & Clinical Development are provided here courtesy of American Society of Gene & Cell Therapy

RESOURCES