Skip to main content

This is a preprint.

It has not yet been peer reviewed by a journal.

The National Library of Medicine is running a pilot to include preprints that result from research funded by NIH in PMC and PubMed.

medRxiv logoLink to medRxiv
[Preprint]. 2024 May 16:2024.03.15.24304216. Originally published 2024 Mar 18. [Version 2] doi: 10.1101/2024.03.15.24304216

Mapping structural variants to rare disease genes using long-read whole genome sequencing and trait-relevant polygenic scores

Cas LeMaster 1, Carl Schwendinger-Schreck 1, Bing Ge 2, Warren A Cheung 1, Rebecca McLennan 1, Jeffrey J Johnston 1, Tomi Pastinen 1, Craig Smail 1
PMCID: PMC10984062  PMID: 38562793

Abstract

Recent studies have revealed the pervasive landscape of rare structural variants (rSVs) present in human genomes. rSVs can have extreme effects on the expression of proximal genes and, in a rare disease context, have been implicated in patient cases where no diagnostic single nucleotide variant (SNV) was found. Approaches for integrating rSVs to date have focused on targeted approaches in known Mendelian rare disease genes. This approach is intractable for rare diseases with many causal loci or patients with complex, multi-phenotype syndromes. We hypothesized that integrating trait-relevant polygenic scores (PGS) would provide a substantial reduction in the number of candidate disease genes in which to assess rSV effects. We further implemented a method for ranking PGS genes to define a set of core/key genes where a rSV has the potential to exert relatively larger effects on disease risk. Among a subset of patients enrolled in the Genomic Answers for Kids (GA4K) rare disease program (N=497), we used PacBio HiFi long-read whole genome sequencing (lrWGS) to identify rSVs intersecting genes in trait-relevant PGSs. Illustrating our approach in Autism (N=54 cases), we identified 22, 019 deletions, 2,041 duplications, 87,826 insertions, and 214 inversions overlapping putative core/key PGS genes. Additionally, by integrating genomic constraint annotations from gnomAD, we observed that rare duplications overlapping putative core/key PGS genes were frequently in higher constraint regions compared to controls (P = 1×10−03). This difference was not observed in the lowest-ranked gene set (P = 0.15). Overall, our study provides a framework for the annotation of long-read rSVs from lrWGS data and prioritization of disease-linked genomic regions for downstream functional validation of rSV impacts. To enable reuse by other researchers, we have made SV allele frequencies and gene associations freely available.

INTRODUCTION

Structural variants (SVs) are a significant source of genetic diversity and have been increasingly recognized as contributors to rare and complex disease (Groza et al. 2024; Kainer et al. 2023; Merker et al. 2018). Typical SVs are genetic alterations longer than 50 bp in length and classified as deletions, duplications, insertions, and inversions. Considering such lengths, the analysis of SVs has been constrained by the limitations of short-read sequencing, which often fails to accurately resolve complex and repetitive genomic rearrangements (Merker et al. 2018; Marx 2023). In recent years, new long-read high fidelity (HiFi) sequencing technologies have enabled the detection and characterization of SVs with greater resolution and accuracy (Hon et al. 2020; Kucuk et al. 2023).

Despite improved SV detection, in a rare disease context it remains highly challenging to identify candidate pathogenic SVs. One solution to this challenge is to first understand the landscape of disease-associated genes for a given rare disorder. One such approach is in integrating trait-relevant polygenic scores (PGS), which quantify inherited genetic predisposition for a given phenotype by estimating disease effects across multiple genetic loci (Simona et al. 2023). In a rare disease context, increased polygenic liability has been observed for many disease cohorts, including in schizophrenia (Davies et al. 2020), severe neurodevelopmental disorders (Niemi et al. 2018), and autism spectrum disorder (Gaugler et al. 2014). Through integrating PGS, we can gain enhanced resolution into genes with impacts on rare diseases.

In this study, we leveraged PacBio long-read HiFi genomes available in the Genomic Answers for Kids (GA4K) study at Children’s Mercy Research Institute to investigate the role of rare structural variants (rSVs) in autism spectrum disorder. The GA4K study is a large-scale, phenotypically diverse pediatric rare disease cohort of over 13,000 individuals (Cohen et al. 2022). At the time of our study, there were 497 rare disease probands with HiFi long read whole genome sequences (lrWGS) completed. Further integrating open-source PGS available from PGS Catalog (Lambert et al. 2021), we computed individual PGS liability for autism spectrum disorder and additional traits and phenotypes. Our study focused on rare deletions, duplications, insertions, and inversions overlapping trait-relevant PGS genes. We further conducted a genomic constraint analysis of rSVs, providing insights on their evolutionary tolerance and potential pathogenicity.

Overall, we found an enrichment of deletions, duplications and insertions at trait-relevant genes and a significantly higher constraint of duplications in individuals with autism. Furthermore, we trace SV-overlapped genes back to autism and implicate anterograde trans-synaptic signaling using ontology enrichment. Our study utilizes the intersection of rSVs, trait-relevant PGS, and genomic constraint to explore the structural variant landscape in a large-scale rare disease cohort, with a focus on genes associated with autism spectrum disorder.

RESULTS

Structural variant type and frequency show prevalence of deletions and insertions

We analyzed 497 probands with HiFi long read whole genome sequences (lrWGS) in the Children’s Mercy GA4K database (N = 254 male; N = 243 female). We assessed ancestry using Somalier, which predicted 437 European, 58 Admixed American, and 2 Asian individuals. Structural variants (SVs) were called using PacBio SV (PBSV) after alignment with human genome GrCh38 (Pedersen et al. 2020). All individuals were assessed for total number of SVs (Supplemental Table 1, Supplemental Figure 1). We developed a python application called GA4K-SV-FINDER where researchers can query these SVs, including allelic frequencies and overlapping genes, by inputting user-specified genomic regions or gene names (Data Availability).

To identify potential large-effect rare structural variants (rSVs) impacting rare disease phenotypes we removed common SVs from the dataset (>0.5% MAF; Methods) (Supplemental Figure 1). This reduced a total of 10,631,806 SVs to 1,017,180 rSVs (Figure 1A). The mean number of SVs per individual was 21,521, and 2,047 for rSVs (Figure 1A).

Figure 1: SV type frequency and average length in study population.

Figure 1:

(A) Two dot-bar graphs representing SV type (Deletions (DEL), duplications (DUP), insertions (INS), inversions (INV)) counts before and after filtering for rare variants. (B) Two horizontal bar graphs with outlier whiskers displaying the average lengths for SVs before and after filtering.

Representing the lowest frequency (0.5%) of rSV type, inversions (INVs) had the largest mean length (25,022 bp) compared to deletions (DEL, 20.3%, 879 bp), duplications (DUP, 2.4%, 8,000 bp), and insertions (INS, 76.8%, 885 bp) (Figure 1). When comparing SV type across sexes, frequencies were similar and within 5% variability between males and females for length and count. The frequency and length of rSVs in our study closely correspond with population frequencies seen in other studies (Kosugi et al. 2019; Guo et al. 2021).

Cohort phenotypes aligned with polygenic scores in five GA4K phenotypes

We calculated PGS for all probands across various rare phenotypes (short stature/height; global developmental delay/fluid intelligence; autism/autism; hypotonia/grip strength; and seizure/epilepsy). Probands were then grouped by their phenotype to construct case/control cohorts. Probands not possessing a given phenotype were considered controls in each mapping. We chose to analyze the top-five most frequent phenotypes in GA4K that had accompanying lrWGS. At the time of assessment, the largest phenotype was Global Developmental Delay at 138 cases, followed by Seizures (86), Hypotonia (62), Autism (54), and Short Stature (45) (Figure 2A). We observed significant differences in Autism and Short Stature PGS, with scores higher for Autism and lower for height, respectively (Figure 2A).

Figure 2: Box plots show significant mappings for Autism and Short Stature.

Figure 2:

(A) Box and whisker plots showing differences in selected PGS mapped to five rare disease phenotypes in GA4K for cases and controls. Red indicates case individuals. Blue indicates control individuals. PGS Z-score is indicated on the Y axis (N cohort = 497, N GDD cases = 138, N Seizure = 86, N Hypotonia = 62, N Autism = 54, N Short Stature = 45). P-values were computed using a Mann-Whitney test. (B) Box plots showing the count of rSV type (DEL, DUP, INS, INV) (left) and rSV type length (right) across individuals for cases (N = 54) and controls (N = 443) in the Autism cohort.

Focusing on the larger case number of the two significant phenotypes, we then compared rSVs in the Autism group (N cases = 54) group with controls (N controls = 443). We observed that probands with Autism had slightly higher median counts of DUPs and INVs, and longer mean lengths, except for INVs, when compared with controls (Figure 2B).

High weight PGS genes reveal rSV enrichment in Autism

Due to the complex genetic nature of Autistic disorders, we decided to further focus on individuals in our cohort with Autism. We evaluated genes harboring variants included in a PGS for Autism (PGS Catalog ID: PGS000327). Each variant is weighted according to its relative association with a phenotype (Methods). We intersected the PGS variant coordinates with the coordinates of genes. This revealed 5,830 genes intersected by at least one overlapping PGS variant. We then intersected cohort rSVs with these PGS-linked genes and found at least one overlap in 4,271 genes (Supplemental Table 2). Among this gene set, Ras Suppressor Protein 1 (RSU1) had the highest PGS weighting (gene weight = 0.139). The lowest weighted gene was Importin 7 (IPO7) (gene weight = 0.023) (Supplemental Table 2; Figure 3). Additionally, there are 262 genes associated with the Human Phenotype Ontology (HPO) ID for autism spectrum disorder (HP:0000717) (Kohler et al. 2021); of those genes, 84 were found to have overlap with rSVs (average gene weight = 0.05).

Figure 3: Top quartile PGS genes show a wide range of weights and a high count of insertions.

Figure 3:

Dot plot of genes within their quartiles based on the intersecting PGS gene weights. The highest weighted genes are in Q4 on the left. The lowest weighted genes are in Q1, to the right. The y-axis for gene weights displays the range of gene weights. The inset bar plot shows the rSV count for each quartile. A color legend is displayed above the y-axis.

We then organized these 4,271 genes into 4 quartiles according to their PGS gene weight (Figure 3). Each quartile contained 1063–1071 genes. All 4,271 genes were evaluated for frequency of rSVs both within the gene body and 200kb up- and down-stream of the gene body (Supplemental Table 2). Focusing on the top quartile, we observed 112,100 rSV overlaps. INSs were the most abundant, both in the gene body (78.3%) and in the flanking regions (upstream: 77%, downstream: 79.3%). This was followed by DELs (body: 19.6%, up: 20.5%, down: 18.1%), DUPs (body: 1.8%, up: 2.1%, down: 2.4%), and INVs (body: 0.2%, up: 0.3%, down: 0.2%). INS and DEL frequency in the gene body paralleled a descending gene weight across all quartiles, where lower gene weights correlated with a lower number of the SV (Figure 3). Duplications followed a similar trend, but the bottom quartile had more than the third quartile. Inversions were more abundant in the bottom quartile, however 431 of 514 INVs were found overlapping the gene Chymotrypsinogen (CTRB1/CTRB2) – a gene region harboring many INVs as previously described (Rosendahl et al. 2018). When evaluating the up- and down-stream 200kb flanking regions for each gene, higher up- and down-stream rSV frequencies associated with a lower gene weight quartile, where the lower gene weight had higher flanking rSV counts.

We then quantified genes in the top quartile of PGS effect with rSV frequency significantly different (>5% diff, Chi-Square test, p-value < 0.05) between Autism and control groups (Table 1; Methods). In doing this we move from entire cohort frequencies to an Autism-control enrichment. While we define significant frequency differences between groups where genes have >5% enrichment over the other group, our aim here is to resolve higher impact rSV-gene enrichments between Autism and controls (>10%) in the top quartile for in-depth downstream analysis (Table 1). Following this approach, we observed that DELs were enriched in GRNA Binding Motif Single Stranded Interacting Protein 3 (RBMS3), Contactin-associated Protein 2 (CNTNAP2), Human Immunodeficiency Virus Type I Enhancer Binding Protein 3 (HIVEP3), Ribosomal Protein S6 Kinase A5 (RPS6KA5), Myomesin 2 (MYOM2), and Astrotactin-2 (ASTN2) (Table 1). Only one gene, Repeat-Containing Protein 7 (WDR7), had >10% enrichment of DUPs (Table 1). The most frequent rSV type, INSs, was enriched in Catenin Alpha-3 (CTNNA3), Neuronal Guanine Nucleotide Exchange Factor (NGEF), Calmodulin-lysine N-methyltransferase (CAMKMT), Coiled-Coil Domain Containing 171 (CCDC171), IQ Motif Containing GTPase Activating Protein 2 (IQGAP2), Myosin ID (MYO1D), Death-associated Protein Kinase 1 (DAPK1), and Neuregulin 3 (NRG3) (Table 1). No INVs were found to be > 10% differentially enriched (Table 1). The genes RBMS3, CTNAP2, HIVEP3, ASTN2, CTNNA3, and DAPK1 have been previously reported in connection to autism spectrum disorders (Pereanu et al. 2018).

Table 1:

SV enriched genes strongly associated with Autism.

Gene Body
Gene %>Control SV Type Locus
CTNNA3 * 20% INS chr10:66007727–67358623
WDR7 * 19% DUP chr18:56826057–57024955
NGEF * 17% INS chr2:232890042–232996364
RBMS3 * 16% DEL chr3:28628062–29007547
CNTNAP2 * 16% DEL chr7:146151646–148366160
CAMKMT * 15% INS chr2:44410220–44589793
CCDC171 * 15% INS chr9:15568002–15883218
HIVEP3 * 14% DEL chr1:41531282–41606928
RPS6KA5 * 14% DEL chr14:91044310–91044392
MYOM2 * 13% DEL chr8:2081050–2130306
IQGAP2 * 12% INS chr5:76591340–76699893
MYO1D * 11% INS chr17:32566830–32862227
DAPK1 * 10% INS chr9:87575220–87652891
NRG3 * 10% INS chr10:82132483–82534600
ASTN2 * 10% DEL chr9:116439185–117282935
NIPAL1 9% DEL chr4:47976159–47976340
AUTS2 9% DEL chr7:70003736–70742818
PDE11A 8% DEL chr2:177946020–177947205
RYR2 7% DEL chr1:237179371–237591190
AMPH 7% DEL chr7:38610296–38610398
KCNQ3 6% INS chr8:132128593–132187729
ATP8A2 6% DEL chr13:25512904–25997038
PTPRT 5% DUP chr20:42622625–42623766
SLC35F4 5% DEL chr14:57652619–57904836
EYS 5% INV chr6:65180037–65180262
LEPR 5% DEL chr1:65575256–65578219
TCAIM 5% DUP chr3:44378928–44385564
FHOD3 5% DEL chr18:36763451–36763532
TSC22D2 5% INS chr3:150448184–150448513
Upstream
Gene %>Control SV Type Locus
CDH2 * 23% INS chr18:27855601–27856513
SFSWAP * 19% DEL chr12:131560091–131701808
F13A1 * 18% INS chr6:6045738–6046653
TPO * 18% DEL chr2:1221441–1346236
SH2D4A * 17% INS chr8:19172662–19172667
RSPO4 * 16% INS chr20:764442–949516
CACNA1A * 15% INS chr19:13185003–13185154
BAZ1A * 15% INS chr14:34557496–34704350
TNS3 * 14% DEL chr7:47111621–47112597
SLC12A7 * 14% INS chr5:853602–1045207
C14orf159 * 14% DEL chr14:91044310–91044392
BCKDHB * 10% INS chr6:80054612–80058252
DHX30 9% INS chr3:47784344–47784344
KATNAL1 9% INS chr13:30003095–30030025
FAM19A5 9% DEL chr22:48306054–48415152
VRK2 9% DEL chr2:57755085–57755160
FBXO31 8% DEL chr16:87162696–87219126
BCAS1 8% DUP chr20:53868937–53869001
CUL1 8% INS chr7:148595522–148617124
ARMC4 8% DEL chr10:27707038–27753250
GALNT2 8% INS chr1:229870399–229870918
SPTBN5 7% DUP chr15:41656319–41656377
MSRB3 7% INS chr12:65083822–65194817
VWA8 7% DEL chr13:41399627–41400559
JARID2 7% DEL chr6:15052128–15194297
DPP10 7% DEL chr2:114371011–114372499
TNKS 7% DUP chr8:9486153–9496734
ARPP21 7% DEL chr3:35564178–35564416
RAB3C 7% INS chr5:58420384–58422140
MLIP 7% INS chr6:53744882–53801696
AKR1C1 6% DEL chr10:4815750–4815964
ITGA7 6% DEL chr12:55661570–55661646
GPR141 6% INS chr7:37498387–37527837
BCO1 6% INS chr16:81144808–81235200
DNAJC12 6% INS chr10:67755364–67755364
CNTN4 6% DEL chr3:1917906–1987559
ESR2 6% INS chr14:63948722–64038340
LRRTM3 6% INS chr10:66791760–66867117
GLIS1 5% DEL chr1:53408988–53468247
Downstream
Gene %>Control SV Type Locus
BANP * 17% DEL chr16:88123428–88237677
SOX5 * 16% INS chr12:24007152–24007683
PPP2R2D * 16% INS chr10:132149554–132159794
AUTS2 * 16% INS chr7:70845576–70872295
TENM3 * 16% INS chr4:182831825–182834588
CAMTA1 * 16% INS chr1:7803433–7953145
U2AF2 * 16% INS chr19:55835451–55835451
SCCPDH * 16% INS chr1:246783075–246785899
EFHB * 15% INS chr3:19969952–19969952
C2orf76 * 12% DEL chr2:119444380–119552830
CUL1 * 12% INS chr7:148839053–148944552
MTUS2 * 12% DEL chr13:29558245–29559472
ITGA7 * 11% INS chr12:55871380–55905934
PEMT * 10% INS chr17:17619242–17656596
CORIN 9% DEL chr4:47976159–47976340
TNS3 9% INV chr7:47614271–47615264
SYNE2 9% INS chr14:64354319–64354319
RB1CC1 8% INS chr8:52878480–52928679
UGDH 7% DEL chr4:39619159–39712200
IFT140 7% DEL chr16:1651054–1761705
ASIC2 7% INS chr17:34302244–34302244
XIRP2 7% INV chr2:167316081–167424243
SV2B 7% DEL chr15:91304256–91304539
SOX30 7% INS chr5:157778905–157778905
CTNND2 6% DEL chr5:11930715–12057490
KLB 6% DEL chr4:39619159–39619751
FBN2 6% INS chr5:128818627–128818639
MYH15 5% DUP chr3:108705618–108705704
MCTP1 5% INS chr5:95299561–95409472
ATG10 5% DEL chr5:82466234–82466779
DNAJC6 5% DEL chr1:65575256–65578219
ALMS1 5% DEL chr2:73624051–73686145
PRKD1 5% INS chr14:30385461–30385461
CFAP77 5% DEL chr9:132716657–132751514
SEPT14 5% INS chr7:55887600–55887855

Genes with significantly (Chi-Square, p-value < 0.05) higher frequency rSV enrichment (>5%) in Autism individual gene bodies (top), upstream (middle) and downstream (bottom) regions. The group frequency differential and SV type is displayed to the right of the gene.

*

Indicates high impact (>10%) rSV-gene differential enrichment in Autism over controls.

We also assessed enrichment in the 200kb flanking regions of these high weight PGS-linked genes, revealing significant enrichment of DELs down-stream of BTG3 Associated Nuclear Protein (BANP), MOB Kinase Activator 1A (C2orf6), and Microtubule Associated Scaffold Protein 2 (MTUS2). However, most down-stream enrichments were INSs proximal to SRY-box Transcription Factor 5 (SOX5), Protein Phosphatase 2 Regulatory Subunit B Delta (PPP2R2D), Autism Susceptibility Candidate 2 (AUTS2), Saccharopine Dehydrogenase (SCCPDH), Teneurin Transmembrane Protein 3 (TENM3), Calmodulin Binding Transcription Activator 1 (CAMTA1), U2 Small Nuclear RNA Auxiliary Factor 2 (U2AF2), EF-Hand Domain Family Member B (EFHB), Cullin 1 (CUL1), Integrin Subunit Alpha 7 (ITGA7), and Phosphatidylethanolamine N-Methyltransferase (PEMT) (Table 1). Furthermore, we observed an enrichment of DELs up-stream of Splicing Factor SWAP (SFSWAP), Thyroid Peroxidase (TPO), Tensin 3 (TNS3), and D-Glutamate Cyclase (C14orf159), as well as INSs enriched up-stream of genes Cadherin 2 (CDH2), Coagulation Factor XIII A Chain (F13A1), SH2 Domain Containing 4A (SH2D4A), R-Spondin 4 (RSPO4), Calcium Voltage-Gated Channel Subunit Alpha1 A (CACNA1A), Bromodomain Adjacent To Zinc Finger Domain 1A (BAZ1A), Solute Carrier Family 12 Member 7 (SLC12A7), and Branched Chain Keto Acid Dehydrogenase E1 Subunit Beta (BCKDHB). We did not observe an enrichment for DUPs or INVs within the PGS-linked gene set.

While we focused on differential enrichment between groups, it is also important to note that the two genes with the highest frequency of rSVs in Autism individuals were not found in SFARI or AutDB Autism databases (MYOM2, 100% [state what this percentage means, just to remind the reader]; NGEF, 98%) (Table 1; Figure 4). Moreover, the top 3 most frequent genes associated with Autism – CTNNA3 (43%), CNTNAP2 (35%), and RBMS3 (26%), while differentially enriched in Autism individuals, were less frequent than those genes with previous associations to the disease (Table 1; Figure 4) (Pereanu et al. 2018). The only gene with enrichments in the gene body as well as a flanking region (downstream) was AUTS2, where DELs were found in the gene body in 13% of individuals and 42% had INSs in the downstream region (Figure 4). Furthermore, genes TNS3 (Upstream: 22% DEL; Downstream: 13% INV), CUL1 (Upstream: 11% INS; Downstream: 26% INS), and ITGA7 (Upstream: 7% DEL; Downstream: 24% INS) had enrichments in both up- and down-stream regions (Figure 4). Three of four genes with differential enrichments in more than one region of a gene were found on chromosome 7 (Figure 4).

Figure 4: Genes with differential rSV enrichment tend have close chromosomal proximity to one another whereas rSV enriched genes with high frequency tend to have more distance between.

Figure 4:

PhenoGram showing the frequency of individuals with rSVs at the cut-off genes by chromosomal position (color scale from 0–100% frequency of case individuals). Cytogenic Giemsa bands show GC- (blue) and AT-rich (gray) regions. Symbols indicate whether the enrichment occurs in the gene body, the upstream, or the downstream regions of the listed gene. The color gradient of symbols indicates the degree of enrichment in the case cohort, with higher enrichment indicated in orange and red.

Insertions were the most common type of enrichment (49/103), followed by DELs (44/103), DUPs (7/103), and INVs (3/103). Two of the 19 chromosomes with rSV enrichments at high weight PGS-linked genes were exclusive to INSs. We also found that while DUPs were relatively infrequent, differentially enriching the gene body of 3 of 62 total genes, DUPs at WDR7 had the second highest differential percentage of an enriched gene body (19% > Controls, p value = 0.0001) (Table 1). Though WDR7 was not found in AutDB or in HPO gene associations, it is positioned down-stream of CDH2 on chromosome 18, the most frequently up-stream enriched gene (76%) associated with Autism (Abrahams et al. 2013). Additionally, PTPRT and BCAS1, two more of the seven genes enriched by DUPs, are both positioned on chromosome 20 and associated with Autism. Moreover, enriching DUPs were at genes with a mean gene weight in the top 3% of PGS-linked genes, whereas DELs, INSs, and INVs were in the top 14%. Overall, we identified genes with differential SV enrichments in Autism over control individuals, specific to the rSV type and locus, finding a high frequency enrichment of DELs and INSs overlapping PGS-linked genes and their flanking regions.

SV enriched genes in the top quartile of PGS effect size significantly associate with autism and potentially associated pathways

After deriving a list of genes with high enrichment of rSVs in Autism (>5%; N genes = 103) (Table1), we sought to determine if these genes specifically enrich the Autism phenotype or candidate biological pathways. We used the DisGeNET database with Enrichr and found a significant enrichment (adj. P = 5×10−06) with Autistic Disorder using our gene list (Table 1). Furthermore, using the GO Cellular Component database, we found that rSVs were enriched in cases for the genes Formin Homology 2 Domain Containing 3 (FHOD3), MYO1D, Katanin Catalytic Subunit A1 Like 1 (KATNAL1), AUTS2, DAPK1, MTUS2, Amphiphysin (AMPH), IQGAP2, CORIN, and Spectrin Repeat Containing Nuclear Envelope Protein 2 (SYNE2) all of which are involved in the cytoskeleton (GO:0015629; GO:0005856) (adj. p-value = 0.02) (Table 1). Our results here provide a framework for linking increased rSV burden in Autism to potential downstream impacts on gene pathways relevant to Autism pathogenesis.

Duplications are enriched at regions of high constraint in top PGS genes

We next used 1kb gnomAD constraint windows to evaluate potential pathogenicity amongst rSVs (Methods). In short, higher constraint indicates an evolutionary intolerance to variation within that locus. Though constraint scores were not available across every 1kb region of the genome, we were able to intersect constraint windows for 7.9% of rSVs in the gene body and 6.7% in the flanking regions. We graphed the density of rSV constraint scores across different SV types (Figure 5). This analysis revealed that rare DELs and DUPs in the Autism group tended to overlap higher constraint windows compared with controls (DEL: p < 0.0001, Autism z mean = 0.14, control z mean = −0.08; DUP: p < 0.0001, Autism z mean = 0.48, control z mean = 0.03). Significant differences were not observed for INSs or INVs (INS: p = 0.66, Autism z mean = 0.2, control z mean = 0.18; INV: p = 0.14, Autism z mean = −2.77, control z mean = −2.6). Furthermore, Autism DUPs tended to overlap higher constraint windows within 200kb up-stream of PGS-linked genes (DUP: p < 0.0001, Autism z mean = 0.66, control z mean = 0.14). We observed a similar pattern for DUPs in Autism individuals for rSVs within the down-stream region (DUP: p < 0.0001, Autism z mean = 0.42, control z mean = 0.11). Thus, DUPs were the only SV type in Autism with significantly higher constraint than controls in all three regions.

Figure 5: Duplication constraints are significantly different between Autism cases and controls inside and outside of the gene body.

Figure 5:

Four density graphs representing two SV types with significantly different constraint scores across Autism-control groups and gene regions. The top two graphs show significantly different constraints across groups for the gene body. The bottom two show up-stream and down-stream differences.

We then focused on DUPs in the top quartile of PGS gene weight and observed a significant enrichment in rSVs overlapping higher constraint regions in Autism compared to controls (p = 0.001, Autism z mean = 0.40, control z mean = −0.66) (Figure 6). This trend was not observed for genes in the bottom quartile of PGS effect (p = 0.14, case z mean = 1.0, control z mean = 0.54) (Figure 6) or across other SV types. In summary, DUPs were found to overlap higher constraint windows more frequently in the Autism group compared to controls for SVs overlapping, or proximal to, large-effect autism PGS genes.

Figure 6: Constraint scores for duplications in top PGS quartile genes are significantly different across Autism cases and controls.

Figure 6:

Violin plot of Autism and control duplication constraint scores within top and bottom quartile genes.

DISCUSSION

Our study highlights the significance of rare structural variants (rSVs) in complex rare diseases such as Autism. By utilizing long-read PacBio HiFi sequencing, we have enhanced the detection and characterization of rSVs in a rare disease cohort, particularly deletions (DELs), duplications (DUPs), and insertions (INSs) observed within or proximal to autism genes. Our analysis indicates a higher incidence of INSs and DELs in genes linked to Autism, with DUPs occurring frequently in large-effect PGS genes with high genomic constraint, suggesting a potential pathogenic influence. The integration of PGS has provided further insight into the genetic predisposition of various phenotypes within the cohort by using correlations between PGS and the clinical presentation of Autism.

Large-scale lrWGS assays in a rare disease cohort such as GA4K enables insights into structural variation not previously explored. This analysis identified specific genes with significant rSV intersections in the context of Autism, such as INSs in the gene body of CTNNA3, the upstream region of CDH2, and downstream of SOX5. Interestingly, all three genes were not only associated with the PGS for Autism, but they were also found in autism gene databases (Banerjee-Basu and Packer 2010; Pereanu et al. 2018). Their regional frequencies were complex, where the flanking enrichments of CDH2 and SOX5 were seen in over 63% of the Autism group, and only 42% of the group was enriched in the body of CTNNA3. This suggests that flanking enrichments of previously correlated Autism genes may provide important insights into the pathogenic influence of these genes on associated disease phenotypes. Additionally evaluating rSVs in genes with investigative association to Autism, we further resolved rSV enriched genes in Autism individuals as autistic-disorder- and cytoskeleton-specific through gene ontology enrichment. Furthermore, the online database of Autism spectrum disorder genes by the Simons Foundation Autism Research Initiative (SFARI) lists 1,176 genes related to the disorder through SNVs (Abrahams et al. 2013). In our study, we implicated over 5,000 genes using PGS associations and resolved 4,271 of those genes with rSV overlaps in our cohort. Of those 1,176 SFARI genes, 503 were in our list of intersected PGS-genes. That leaves 3,768 genes absent from SFARI but with PGS and rSV Autism overlaps identified in our cohort. Because long-read mapping-based methods for SV calling are more accurate than short-read approaches, due to their ability to provide more accurate alignments by capturing entire SV alleles in a single fragment – additionally considering that HiFi lrWGS is relatively new – our study underscores the importance in adopting lrWGS in rare disease where no single nucleotide variant is implicated.

In future studies, there should be consideration into the complexities of rare disease phenotypes and any potential comorbidities. Controlling for phenotype, age, gender, and family transmission might strengthen gene-rSV associations. Our findings underscore the complexity of the genetic factors impacting Autism risk and reinforce the importance of comprehensive analysis required in uncovering the nuances of rare disease variants. These insights lay the groundwork for more precise genetic diagnostics and the development of targeted interventions involving rSVs.

METHODS

Study Cohort

The study cohort described includes 497 affected probands across 419 families with 49% females and 51% males. Subjects were enrolled in the Genomics Answers for Kids (GA4K) program. Probands at age of enrollment ranged from 0 to 57 years (median = 8) (older individuals were typically ascertained as follow-up from an affected family member). Subjects were eligible for the study if they had a suspected genetic diagnosis based on clinical presentation and/or an existing molecular or cytogenic findings. The study complies with all relevant ethical regulations as approved by the Children’s Mercy Institutional Review Board (IRB) (Study #11120514). Informed or written consent was obtained from all participants prior to study inclusion. Participants were not compensated for participation.

SV Calling

PacBio HiFi Sequencing results from GA4K proband subjects were processed through PacBio’s structural variant (PBSV, Version: 2.6.2) calling workflows after alignment to human reference GRCh38. Variants achieved a PASS score if they had at least 1 read per strand, were at least 1000 bp from the end of a contig, and away (< 1000 bp) from gaps (run of >= 50 Ns) in the reference assembly. Additionally, for PBSV, we used the following parameters: -ccs (Circular Consensus Sequence/HiFi), -m 20 (SVs larger than 20 bp were used), -A 3 (3 supporting reads were required to call a variant as homozygous alternate), -O 3 (3 supporting reads were required to call a variant as heterozygous), -P 20 (minimum phred) were used for structural variant calling. Due to known false-positives, CNVs were not included in variant analysis. Additionally, only SVs greater than 50 bp in length were used for each variant type. rSVs were defined as having a minor allele frequency (MAF) < 0.5% or allele count =< 4 in the study cohort. Somalier was used to predict ancestry (> 0.5 probability).

GA4K-SV-FINDER

We created a python-based application for querying allele frequencies and genes associated with SVs in our cohort. This application can be found, along with a readme and the raw data file, in a GitHub repository available at https://github.com/smail-lab-cmh/ga4k-sv-finder.

Polygenic Risk Scores

Polygenic risk scores were calculated for each specified polygenic score using open-source PGS available in the PGS catalog (www.pgscatalog.org). Associated PGS variant coordinates were downloaded from their respective catalog entries. PGS variants were converted to hg38 coordinates using dbSNP26 (version 155) where necessary. PGS variants were restricted to autosomes only and variants mapping to the HLA region were removed. Individual-level PGS were calculated using PLINK (version 1.9) (Purcell et al. 2007) (“sum” flag) on all variants available in the GA4K imputed genotype callset with R2>=0.8 (“exclude-if-info” flag). PGS scores were converted to Z-scores within each PGS in the full GA4K EUR ancestry cohort (proband and other available family members, N = 7,436). Individuals with an extreme outlier PGS Z-score (abs(PGS Z-score)>=10) in one or more PGS were removed from the final PGS dataset.

Intersect

To create compatible chromosomal coordinate files for structural and PGS variants, pseudo-BED files were first created for each PGS SNP coordinate file, containing the chromosome, coordinates, base change, and associated weight within the PGS as a z-score. This file was first intersected, using the bedtools intersect command (v2.31.0), with a gene position BED file (gencode v26 hg38) to determine the gene the SNP was associated with. After determining the genes associated with the PGS, they were then intersected with pseudo-BED files generated from each pbsv VCF from the cohort. This created intersect files that provide the chromosomal coordinates for the SV and the genes, which included sv type, length, weight, and overlap length.

Flanking Regions

Both up- and down-stream intersects were performed by first using BEDtools Slop command on gene start and end coordinates separately. For example, start coordinates are transformed into a 200kb upstream value, expanding the gene start coordinate within the boundaries of its respective chromosome. When intersecting these expanded genes with SVs, any SV found to enrich the gene body that is also matching an SV found in the new flanking region is removed from the flanking region analysis, as it was already intersecting the gene body. In doing this, we separated SVs found up- or down-stream of the gene from the SVs overlapping the gene body itself.

Gene Weighting

Gene coordinates were retrieved from the gencode v26 hg38 basic annotation file (gencodegenes.org). Gene coordinates were intersected by weighted PGS variant coordinates. For each gene, PGS variant weights were transformed to absolute values before taking the maximum value across all intersecting variants. In turn, this provided an absolute maximum weight per gene. Following gene weighting, gene coordinates were intersected by rSVs.

Constraint Scoring

Each rSV was intersected with the gnomAD v3 QCed genomic constraint by 1kb regions file (file name: constraint_z_genome_1kb.qc.download.txt.gz). This file scores 1kb regions throughout the genome with a constraint z-score value. Each of the proband rSVs were intersected with these 1kb constraint windows. For window density assessment of each rSV type, all intersected constraint windows are considered. For top and bottom quartile assessment, the scores are averaged across windows for each SV, excluding non-coding regions or any undefined region lacking a constraint value. This produces a single score for each SV intersecting a gene.

Statistical and Enrichment Analysis

Graphing and statistics were generated using Graph Pad Prism 9 and Python (SciPy-Stats & Pandas). For enrichment, SV frequencies between Autism and control groups were determined and Chi-Square tests were used to determine significant differences. When SV counts were low for a gene (<5), a Fisher’s Exact test was used. These tests were performed for body, upstream, and downstream of genes separately.

Gene Analysis

Gene analysis was conducted using the DISEASES database, AutDB, GeneCards Version 5.19, and Enrichr (Grissa et al. 2022; Stelzer et al. 2016; Chen et al. 2013; Pereanu et al. 2018).

Supplementary Material

Supplement 1
media-1.pdf (290.4KB, pdf)
Supplement 2
media-2.csv (33.2KB, csv)
Supplement 3
media-3.csv (641.6KB, csv)

ACKNOWLEDGEMENTS

C.S. is supported by NIH grant R35GM146966. We thank all the individuals that participated in making the GA4K study possible. This work was funded through internal institutional funds from Children’s Mercy Research Institute and Children’s Mercy Kansas City.

Footnotes

COMPETING INTERESTS

The authors declare no competing interests.

DATA AVAILABILITY

GA4K study data can be found at the ANVIL host at https://anvilproject.org/data/studies/phs002206/workspaces. Code used, as well as the application for querying SVs (GA4K-SV-FINDER) can be found in a git-hub repository at https://github.com/smail-lab-cmh/ga4k-sv-finder.

REFERENCES

  1. Abrahams B. S., Arking D. E., Campbell D. B., Mefford H. C., Morrow E. M., Weiss L. A., Menashe I., Wadkins T., Banerjee-Basu S., and Packer A.. 2013. ‘SFARI Gene 2.0: a community-driven knowledgebase for the autism spectrum disorders (ASDs)’, Mol Autism, 4: 36. [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Banerjee-Basu S., and Packer A.. 2010. ‘SFARI Gene: an evolving database for the autism research community’, Dis Model Mech, 3: 133–5. [DOI] [PubMed] [Google Scholar]
  3. Chen E. Y., Tan C. M., Kou Y., Duan Q., Wang Z., Meirelles G. V., Clark N. R., and Ma’ayan A.. 2013. ‘Enrichr: interactive and collaborative HTML5 gene list enrichment analysis tool’, BMC Bioinformatics, 14: 128. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Cohen A. S. A., Farrow E. G., Abdelmoity A. T., Alaimo J. T., Amudhavalli S. M., Anderson J. T., Bansal L., Bartik L., Baybayan P., Belden B., Berrios C. D., Biswell R. L., Buczkowicz P., Buske O., Chakraborty S., Cheung W. A., Coffman K. A., Cooper A. M., Cross L. A., Curran T., Dang T. T. T., Elfrink M. M., Engleman K. L., Fecske E. D., Fieser C., Fitzgerald K., Fleming E. A., Gadea R. N., Gannon J. L., Gelineau-Morel R. N., Gibson M., Goldstein J., Grundberg E., Halpin K., Harvey B. S., Heese B. A., Hein W., Herd S. M., Hughes S. S., Ilyas M., Jacobson J., Jenkins J. L., Jiang S., Johnston J. J., Keeler K., Korlach J., Kussmann J., Lambert C., Lawson C., Le Pichon J. B., Leeder J. S., Little V. C., Louiselle D. A., Lypka M., McDonald B. D., Miller N., Modrcin A., Nair A., Neal S. H., Oermann C. M., Pacicca D. M., Pawar K., Posey N. L., Price N., Puckett L. M. B., Quezada J. F., Raje N., Rowell W. J., Rush E. T., Sampath V., Saunders C. J., Schwager C., Schwend R. M., Shaffer E., Smail C., Soden S., Strenk M. E., Sullivan B. R., Sweeney B. R., Tam-Williams J. B., Walter A. M., Welsh H., Wenger A. M., Willig L. K., Yan Y., Younger S. T., Zhou D., Zion T. N., Thiffault I., and Pastinen T.. 2022. ‘Genomic answers for children: Dynamic analyses of >1000 pediatric rare disease genomes’, Genet Med, 24: 1336–48. [DOI] [PubMed] [Google Scholar]
  5. Davies R. W., Fiksinski A. M., Breetvelt E. J., Williams N. M., Hooper S. R., Monfeuga T., Bassett A. S., Owen M. J., Gur R. E., Morrow B. E., McDonald-McGinn D. M., Swillen A., Chow E. W. C., van den Bree M., Emanuel B. S., Vermeesch J. R., van Amelsvoort T., Arango C., Armando M., Campbell L. E., Cubells J. F., Eliez S., Garcia-Minaur S., Gothelf D., Kates W. R., Murphy K. C., Murphy C. M., Murphy D. G., Philip N., Repetto G. M., Shashi V., Simon T. J., Suner D. H., Vicari S., Scherer S. W., Brain International 22q, Consortium Behavior, Bearden C. E., and Vorstman J. A. S.. 2020. ‘Using common genetic variation to examine phenotypic expression and risk prediction in 22q11.2 deletion syndrome’, Nat Med, 26: 1912–18. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Gaugler T., Klei L., Sanders S. J., Bodea C. A., Goldberg A. P., Lee A. B., Mahajan M., Manaa D., Pawitan Y., Reichert J., Ripke S., Sandin S., Sklar P., Svantesson O., Reichenberg A., Hultman C. M., Devlin B., Roeder K., and Buxbaum J. D.. 2014. ‘Most genetic risk for autism resides with common variation’, Nat Genet, 46: 881–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Grissa D., Junge A., Oprea T. I., and Jensen L. J.. 2022. ‘Diseases 2.0: a weekly updated database of disease-gene associations from text mining and data integration’, Database (Oxford), 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Groza C., Schwendinger-Schreck C., Cheung W. A., Farrow E. G., Thiffault I., Lake J., Rizzo W. B., Evrony G., Curran T., Bourque G., and Pastinen T.. 2024. ‘Pangenome graphs improve the analysis of structural variants in rare genetic diseases’, Nat Commun, 15: 657. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Guo M., Li S., Zhou Y., Li M., and Wen Z.. 2021. ‘Comparative Analysis for the Performance of Long-Read-Based Structural Variation Detection Pipelines in Tandem Repeat Regions’, Front Pharmacol, 12: 658072. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Hon T., Mars K., Young G., Tsai Y. C., Karalius J. W., Landolin J. M., Maurer N., Kudrna D., Hardigan M. A., Steiner C. C., Knapp S. J., Ware D., Shapiro B., Peluso P., and Rank D. R.. 2020. ‘Highly accurate long-read HiFi sequencing data for five complex genomes’, Sci Data, 7: 399. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Kainer D., Templeton A. R., Prates E. T., Jacboson D., Allan E. R. O., Climer S., and Garvin M. R.. 2023. ‘Structural variants identified using non-Mendelian inheritance patterns advance the mechanistic understanding of autism spectrum disorder’, HGG Adv, 4: 100150. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Kohler S., Gargano M., Matentzoglu N., Carmody L. C., Lewis-Smith D., Vasilevsky N. A., Danis D., Balagura G., Baynam G., Brower A. M., Callahan T. J., Chute C. G., Est J. L., Galer P. D., Ganesan S., Griese M., Haimel M., Pazmandi J., Hanauer M., Harris N. L., Hartnett M. J., Hastreiter M., Hauck F., He Y., Jeske T., Kearney H., Kindle G., Klein C., Knoflach K., Krause R., Lagorce D., McMurry J. A., Miller J. A., Munoz-Torres M. C., Peters R. L., Rapp C. K., Rath A. M., Rind S. A., Rosenberg A. Z., Segal M. M., Seidel M. G., Smedley D., Talmy T., Thomas Y., Wiafe S. A., Xian J., Yuksel Z., Helbig I., Mungall C. J., Haendel M. A., and Robinson P. N.. 2021. ‘The Human Phenotype Ontology in 2021’, Nucleic Acids Res, 49: D1207–D17. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Kosugi S., Momozawa Y., Liu X., Terao C., Kubo M., and Kamatani Y.. 2019. ‘Comprehensive evaluation of structural variation detection algorithms for whole genome sequencing’, Genome Biol, 20: 117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Kucuk E., Bpgh van der Sanden, O’Gorman L., Kwint M., Derks R., Wenger A. M., Lambert C., Chakraborty S., Baybayan P., Rowell W. J., Brunner H. G., Vissers Lelm, Hoischen A., and Gilissen C.. 2023. ‘Comprehensive de novo mutation discovery with HiFi long-read sequencing’, Genome Med, 15: 34. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Lambert S. A., Gil L., Jupp S., Ritchie S. C., Xu Y., Buniello A., McMahon A., Abraham G., Chapman M., Parkinson H., Danesh J., MacArthur J. A. L., and Inouye M.. 2021. ‘The Polygenic Score Catalog as an open database for reproducibility and systematic evaluation’, Nat Genet, 53: 420–25. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Marx V. 2023. ‘Method of the year: long-read sequencing’, Nat Methods, 20: 6–11. [DOI] [PubMed] [Google Scholar]
  17. Merker J. D., Wenger A. M., Sneddon T., Grove M., Zappala Z., Fresard L., Waggott D., Utiramerur S., Hou Y., Smith K. S., Montgomery S. B., Wheeler M., Buchan J. G., Lambert C. C., Eng K. S., Hickey L., Korlach J., Ford J., and Ashley E. A.. 2018. ‘Long-read genome sequencing identifies causal structural variation in a Mendelian disease’, Genet Med, 20: 159–63. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Niemi M. E. K., Martin H. C., Rice D. L., Gallone G., Gordon S., Kelemen M., McAloney K., McRae J., Radford E. J., Yu S., Gecz J., Martin N. G., Wright C. F., Fitzpatrick D. R., Firth H. V., Hurles M. E., and Barrett J. C.. 2018. ‘Common genetic variants contribute to risk of rare severe neurodevelopmental disorders’, Nature, 562: 268–71. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Pedersen B. S., Bhetariya P. J., Brown J., Kravitz S. N., Marth G., Jensen R. L., Bronner M. P., Underhill H. R., and Quinlan A. R.. 2020. ‘Somalier: rapid relatedness estimation for cancer and germline studies using efficient genome sketches’, Genome Med, 12: 62. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Pereanu W., Larsen E. C., Das I., Estevez M. A., Sarkar A. A., Spring-Pearson S., Kollu R., Basu S. N., and Banerjee-Basu S.. 2018. ‘AutDB: a platform to decode the genetic architecture of autism’, Nucleic Acids Res, 46: D1049–D54. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Purcell S., Neale B., Todd-Brown K., Thomas L., Ferreira M. A., Bender D., Maller J., Sklar P., de Bakker P. I., Daly M. J., and Sham P. C.. 2007. ‘PLINK: a tool set for whole-genome association and population-based linkage analyses’, Am J Hum Genet, 81: 559–75. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Rosendahl J., Kirsten H., Hegyi E., Kovacs P., Weiss F. U., Laumen H., Lichtner P., Ruffert C., Chen J. M., Masson E., Beer S., Zimmer C., Seltsam K., Algul H., Buhler F., Bruno M. J., Bugert P., Burkhardt R., Cavestro G. M., Cichoz-Lach H., Farre A., Frank J., Gambaro G., Gimpfl S., Grallert H., Griesmann H., Grutzmann R., Hellerbrand C., Hegyi P., Hollenbach M., Iordache S., Jurkowska G., Keim V., Kiefer F., Krug S., Landt O., Leo M. D., Lerch M. M., Levy P., Loffler M., Lohr M., Ludwig M., Macek M., Malats N., Malecka-Panas E., Malerba G., Mann K., Mayerle J., Mohr S., Te Morsche R. H. M., Motyka M., Mueller S., Muller T., Nothen M. M., Pedrazzoli S., Pereira S. P., Peters A., Pfutzer R., Real F. X., Rebours V., Ridinger M., Rietschel M., Rosmann E., Saftoiu A., Schneider A., Schulz H. U., Soranzo N., Soyka M., Simon P., Skipworth J., Stickel F., Strauch K., Stumvoll M., Testoni P. A., Tonjes A., Werner L., Werner J., Wodarz N., Ziegler M., Masamune A., Mossner J., Ferec C., Michl P., Drenth J P. H., Witt H., Scholz M., Sahin-Toth M., and A. C. P. all members of the PanEuropean Working group on. 2018. ‘Genome-wide association study identifies inversion in the CTRB1-CTRB2 locus to modify risk for alcoholic and non-alcoholic chronic pancreatitis’, Gut, 67: 1855–63. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Simona A., Song W., Bates D. W., and Samer C. F.. 2023. ‘Polygenic risk scores in pharmacogenomics: opportunities and challenges-a mini review’, Front Genet, 14: 1217049. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Stelzer G., Rosen N., Plaschkes I., Zimmerman S., Twik M., Fishilevich S., Stein T. I., Nudel R., Lieder I., Mazor Y., Kaplan S., Dahary D., Warshawsky D., Guan-Golan Y., Kohn A., Rappaport N., Safran M., and Lancet D.. 2016. ‘The GeneCards Suite: From Gene Data Mining to Disease Genome Sequence Analyses’, Curr Protoc Bioinformatics, 54: 1 30 1–1 30 33. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplement 1
media-1.pdf (290.4KB, pdf)
Supplement 2
media-2.csv (33.2KB, csv)
Supplement 3
media-3.csv (641.6KB, csv)

Data Availability Statement

GA4K study data can be found at the ANVIL host at https://anvilproject.org/data/studies/phs002206/workspaces. Code used, as well as the application for querying SVs (GA4K-SV-FINDER) can be found in a git-hub repository at https://github.com/smail-lab-cmh/ga4k-sv-finder.


Articles from medRxiv are provided here courtesy of Cold Spring Harbor Laboratory Preprints

RESOURCES