Abstract
Decades of genetic association testing in human cohorts have provided important insights into the genetic architecture and biological underpinnings of complex traits and diseases. However, for certain traits, genome-wide association studies (GWAS) for common single nucleotide polymorphisms (SNPs) are approaching signal saturation, which underscores the need to explore other types of genetic variation to further understand the genetic basis of traits and diseases. Copy number variation (CNV) genome-wide is an important source of heritability that is well-known to functionally affect human traits. Recent technological and computational advances are enabling the large-scale, genome-wide evaluation of CNVs which is likely to impact several downstream applications. Here, we review the current state of the art for GWAS beyond SNPs; highlight areas of opportunity; discuss current limitations in resource infrastructure that need to be overcome to enable the wider uptake of CNV-GWAS results; and suggest guidelines and standards for future GWAS for variation beyond SNPs at scale.
Introduction
Genome-wide association studies (GWAS) have shown remarkable potential for modelling genotype to phenotype relationships — typically single nucleotide polymorphisms (SNPs) — using a simple additive assumption against both continuous and binary human phenotypes1,2. Early findings led to the rapid development of high-throughput genotyping assays (SNP arrays) and novel analytical techniques for genetic studies into common, complex diseases3. The resulting ‘explosion’ in data generation saw a sharp increase in the number of cohorts genotyped for SNPs, allowing extensive genotype-to-phenotype modelling across a diverse set of human traits4. Dedicated repositories such as the NHGRI-EBI GWAS Catalog5, which now contains over half a million lead associations from more than 90,000 GWAS, provide expert curated results and underlying summary statistics to the scientific community in an open, consistent and reliable way following the FAIR recommendations6. Standardization of results and metadata has been key to enable productive downstream use, for example, the generation of polygenic scores (PGS; also known as polygenic risk score or PRS) for risk assessment7, causal inference through Mendelian randomization or the incorporation of data into drug discovery pipelines, such as the Open Targets Platform8.
Understanding the full extent of genetic contributions to disease risk or trait distributions is crucial for understanding biological mechanisms, creating better models for early disease detection and tailoring therapeutic strategies. However, of the thousands genetic variants that have been associated with human traits or diseases using GWAS, most are common SNPs with low effect sizes, that is, they confer fairly small changes in overall disease risk. Moreover, despite ever increasing sample sizes, GWAS are approaching signal saturation for common SNPs for certain traits9,10. The largest GWAS to date used a combined cohort of over 5 million individuals across multiple ancestry groups to investigate the genetic contributions to height9. Researchers produced a saturated map of over 12,000 common SNPs, which are estimated to explain 40% of the observed phenotypic variance in populations of European ancestries. Furthermore, a sample size of 5 million was sufficient to map more than 90% of the genetic variance explained by common SNPs in European ancestries. Nevertheless, a substantial amount of phenotypic variation cannot be accounted for by common SNPs, as causal variants are not well tagged by common SNPs11, which has prompted researchers to explore alternative approaches, including rare-variant analyses, and other types of genetic variation in the context of human trait distributions12. There remains many open questions around how much of the observed “missing heritability” (phenotypic variance) can be explained by different factors including variation other than SNPs, rare variants, cryptic relatedness, epistasis and environmental factors13.
Large-scale sequencing studies have shown that copy number variants (CNVs) — a class of structural variation (SV) displaying a genomic imbalance in their number of copies14 and often involving large segments of DNA that are duplicated or deleted — have a substantial impact on complex phenotypes; owing to their size range from kilobases to megabases, they have great potential to affect genome structure and the regulatory landscape, both in close proximity and more distally to their locations15. One of the first large studies looking at CNV associations for common disease described significant associations between specific CNVs and four major common diseases (Crohn’s disease, rheumatoid arthritis, and type 1 and type 2 diabetes mellitus)16. At the time, it was thought that large-scale and systematic GWAS tests for CNV would result in many novel associations that could have a strong relevance to human health. However, most large cohorts were set up almost exclusively using SNP genotyping arrays, which provide low resolution for the detection of biallelic CNVs and limited representation of multi-allelic CNVs17, resulting in substantial under-representation of CNV studies and associations in the literature and in the key data resources compared to SNP-based associations (Figure 1). In contrast to SNP arrays, high-throughput DNA sequencing (also known as next-generation sequencing (NGS)) enables the detection of CNVs at single base-pair resolution and is now available in multiple population-scale human cohorts18. As a result, CNV associations are starting to be incorporated more widely into genome-wide tests, although challenges remain to scale these efforts19.
Figure 1. The cumulative total of association studies added to the GWAS Catalog between 2021 – 2023 for SNP- and CNV-based tests.
The cumulative number of new association studies added to the GWAS catalog per month with SNP GWAS shown in blue and CNV GWAS shown in orange. The y-axis has two different scale for SNP GWAS (left in blue) and CNV GWAS (right in orange) showing maximums of 90,000 and 1,000 for new SNP and CNV GWAS studies respectively.
In this Review, we describe the current state of GWAS beyond SNPs and provide recommendations for the next phase of large-scale application in human cohorts. We highlight methodologies, software and infrastructure that are currently missing, and require further consideration before CNV-GWAS results can be fully integrated with existing GWAS resources.
Biological and clinical relevance of CNVs
Copy number variation (CNV) defines any region of the genome that varies in copy between two or more individuals. Normally, CNVs are assessed with respect to a reference genome whereby sequencing reads are aligned to a reference backbone and copy-number-variable locations are identified (or ‘called’) across samples. Often this process involves defining the normal (or ‘dynamic average’) copy number for all locations present in the reference genome and then detecting differences in copy number estimates compared to the reference baseline20,21. Genetic variation comes in multiple different varieties ranging from single base pair differences (that is, SNPs) to small insertions and deletions (INDELs) all the way up to larger SV and copy number changes (that is, CNVs). CNVs can occur via different mechanisms, including deletion, insertion, duplication either tandem (directly adjacent) or interspersed, or translocation between chromosomes (Figure 2a).
Figure 2. CNV types, differential frequencies in world populations, loss of function modes of action and relevance to pharmacogenetics.
a) Depiction of different classes of genetic variants and highlighting the different types of structural and copy number variants that can occur in a genome. b) Effect a single copy deletion of the UGT2B17 gene on the metabolism of testosterone with reported higher risks of prostate cancer and the the reported differential frequency of the single copy loss across world populations. c) Illustration of different some modes of action by which CNV events can result in loss of function for different alleles. d) Effect of both a whole-gene duplication and deletion of the CYP2D6 gene on the metabolism of codeine, showing the resulting increased and decreased formation of morphine respectively resulting in either over sensitivity or a lack of efficiency of Codeine of individuals with these different CNVs.
Of the datasets shared in the GWAS Catalog, ~16,000 significant (P<10−5) CNV–trait associations have been reported in total, compared to >500 million SNP lead associations. CNVs occur at different frequencies within a population, and, like SNPs, most are rare. However, there are many genomic locations that harbour commonly variable CNV locations (MAF >0.05)22–25; some of these CNV alleles will be in linkage disequilibrium (LD) with surrounding SNPs and so could have been captured by previous GWAS if appropriate tagging SNPs were present, but the actual causal variation would still be unknown. Importantly, recurrent CNVs are far more common across the genome than recurrent SNPs and account for a larger fraction of the total base-pair differences between genomes26,27.
Population genetics
Large reference population databases such as gnomAD28 have improved the ability to investigate genetic differences across world populations and are an essential component of variant interpretation pipelines29. CNVs are a major source of genetic diversity, with geographical patterns of variation similar to SNPs30 that reflect human origins in Africa, followed by population movements during evolutionary history. Furthermore, the role of CNV in producing phenotypic differences has been long observed and known to be important during genotype to phenotype mapping in agricultural and livestock species with multiple studies investigating CNV effects against numerous important traits in a variety of species31,32.
Similar to SNP haplotypes, variation in patterns of LD across the globe present challenges for imputation, confounded by the high mutation rate of some CNVs33, which further erodes LD.
CNV reference sets from multiple populations provide well-calibrated frequency distributions in unaffected controls and are useful for the exclusion of likely benign variants within a clinical diagnostic setting34 and for the assessments of frequency differences between populations35. It is worth noting the critical point that correct frequency estimation for CNV relies on accurate CNV breakpoint estimation which can be somewhat overcome by relaxing the definition of CNV matching using a fuzzier definition of start and end locations for the event or an iterative reciprocal overlap rule36, however all approaches can result in the incorrect merging of CNV events which can in turn result in inconsistent frequency estimates. There are several classic examples of CNV frequency skew that have, in some cases, been attributed to historical, cultural or even behavioural differences37 (Figure 2b). Perhaps the best studied is the salivary amylase gene AMY1, which has copy numbers ranging from 2 to >10 copies38. Increased copy number of AMY1 was the first reported example of positive selection on CNVs across world populations, with individuals from populations with historically high-starch diets having higher copy numbers of the AMY1 gene39. More recent investigations have implicated CNVs at AMY1 and its complex ancestral haplotypes (SNPs and CNVs) in glucose metabolism impairment and obesity with carriers of lower AMY1 copy number potentially having a higher risk of developing insulin resistance38,40.
Population-level estimates of CNV frequencies can help derive measures of dosage sensitivity. These measures can predict the likely pathogenicity of variants in a genomic region by considering how rare these events are within a population41. Here it is important to distinguish between regions of the genome that are more likely to harbour structural rearrangements due to mechanisms such as non-allelic homologous recombination (NAHR) and their deleteriousness42. Certain genomic contexts such as high levels of repeats and segmental duplications shown higher variant loads and increased rates of somatic mutations (mutational hotspots). Variants that are likely to have arisen due to somatic mutation (e.g. in whole blood) can potentially be detected and corrected for using the level mosaicism that they display43 and further discriminated from background artefacts using novel machine learning based approaches44. The absence or aggregate rarity of loss-of-function single-nucleotide variants (SNVs) in protein-coding genes, such as those reported by the gnomAD loss-of-function observed/expected (LOEUF) score45, has shown value for the discovery of gene-disrupting variation and the inference of ‘human knockouts’. Genome-wide dosage sensitivity maps can be useful for making new genetic discoveries, interpreting previously known associations, and developing diagnostic guidelines within clinical genetics settings41,46–48. With larger cohort sizes and higher genome resolution across diverse sets of population groups49, there is a strong need to aggregate information across all variant class types, to allow more comprehensive association testing, improved dosage sensitivity maps and more accurate clinical interpretation.
Human disease and clinical diagnosis
Recent years have made it increasingly clear that CNVs are involved in the susceptibility to complex traits and conditions50, although their association with rare Mendelian disorders has been known since the early days of CNV detection51,52. Numerous CNV–disease associations have been proposed through the analysis of single affected families using techniques such as array-based comparative genomic hybridization (aCGH)53,54. Furthermore, several cohort-based studies have assessed the contribution of CNV to overall disease burden12,55–57, although these have often been limited in resolution, relying on either targeted sequencing or array-based assays58–60. Recurrent de novo CNVs have been implicated in a range of genomic disorders often in relation to disorders with neurodevelopmental delay, including Prader-Wlli, DeGeroge syndromes and many others61.
Neurodevelopmental and neuropsychiatric disorders represent a notable set of conditions in which incorporation of CNVs into analytical pipelines has led to considerable insights into disease aetiology, with a significant proportion of disease risk of schizophrenia, bipolar disorder and autism spectrum disorders explained by rare variants62. Today, this class of variants is widely accepted as an important susceptibility determinant, in some cases with effect sizes close to those typically observed in Mendelian disease. A key example is the chromosomal region 22q11.2, which has been associated with a high odds ratio for schizophrenia if deleted (OR=67.7) and protective effect if duplicated (OR=0.15)63,64. The availability of large-scale disease cohorts has enabled multiple informative CNV studies, including a recent effort showing how somatic CNVs explain a small but notable proportion of schizophrenia65 and another recent study into the genetic architecture of autism using extensive whole genome sequence annotation for rare and common variations including CNVs66.
To date, hundreds of CNV locations throughout the genome have been prioritized for diagnostic testing within a clinical genetics settings, with the National Genomic Test Directory being developed with the guidance of NHS England Rare and Inherited Disease working group and recommendations for the interpretation and reporting of CNVs from the American College of Medical Genetics and the Clinical Genome Resource (ClinGen)67. CNVs have different loss of function modes of actions, with impact on either one or both copies of a gene (or genomic element), and can interact with different genetic variants on opposite alleles, resulting in complete loss of function (or compound heterozygous variation). These different modes of action are important to consider in variant prioritization and are often accounted for in clinical diagnostic tests (Figure 2c). Most diagnostic tests have been developed using CNV classifications from either exon sequencing or multiplex ligation-dependent probe amplification (MPLA) and are evaluated by clinical geneticists, specialist clinicians, scientists and specialized patient groups (Supplementary File 1).
Pharmacogenetics
Pharmacogenetics is the study of how genetic variation influences our response to medication. Genes that encode the proteins involved in adsorption, distribution, metabolism and excretion (ADME) can influence the pharmacokinetics of a drug, and genes that encode drug targets or proteins that interact with drug targets, or their pathways can influence the pharmacodynamics of a drug. Genetic variation, including CNVs, in these genes can affect the efficacy of drug treatment, dosage required by a patient and the risk of adverse effects. Pharmacogenetic testing, either pre-emptively or prior to prescribing, is now clinically established in multiple healthcare systems for key gene–drug pairs with sufficient evidence of clinical utility68.
One well-established example is copy number variation at CYP2D6, for which there are published genotype-based therapeutic guidelines from several organizations69 owing to its role in the metabolism of commonly prescribed drugs such as painkillers and anti-depressants. CYP2D6 lies in a region of chromosome 22 with two pseudogenes (CYP2D7 and CYP2D8). It is highly polymorphic and structurally complex, with over 130 haplotypes (termed ‘star alleles’), including those containing SNVs, entire gene deletions (*5), identical and non-identical gene duplications and multiplications, and conversions or hybrid genes structures with CYP2D7. As more diverse populations are studied, new alleles are being reported70, and multiple methods have been developed for calling the CYP2D6 alleles from NGS data71–74.
The Clinical Pharmacogenetic Implementation Consortium provides evidence-based recommendations for prescribing through translation of a patient’s diplotype to the implication on therapeutic outcome (Figure 2d). Ultrarapid CYP2D6 metabolizers have an activity score range of >2.25; example diplotypes include whole-gene duplication(s) of the wild type allele (*1/*1xN). When treated with codeine or tramadol, these patients have increased formation of morphine and should avoid use of these drugs due to a higher risk for serious drug-induced toxicity69. Conversely, CYP2D6 poor metabolizers have an activity score of 0 (for example the whole-gene deletion diplotype *5/*5). These patients have greatly reduced morphine formation leading to diminished analgesia and therefore codeine use is not recommended due to the lack of drug efficacy69. This is a key example of the clinical impact of CNVs and may have implications for better opioid use management75. Warnings issued by regulatory bodies regarding CYP2D6 metabolizer status now appear on drug labelling information, including for codeine, and for several drugs, testing of metabolizer status is required prior to prescription (https://www.pharmgkb.org/gene/PA128/labelAnnotation). Further examples of complex loci involving differences in drug metabolism and efficacy include the major histocompatibility complex (MHC) and immunoglobulin-like receptors KIR76.
Conducting a CNV-GWAS
Obtaining CNV-relevant information appropriate for association testing requires a set of steps to go from either sequence or array-based data to relative coverage information or CNV calls as a proxy for copy number genome-wide. Multiple tools have been developed to perform these steps and are reported in detail elsewhere77–79. Once these steps have been performed and following both sample and probe level quality control (normally both prior to CNV calling and post CNV calling) there is a standard set of tasks required to perform before any CNV association test can be reliably carried out. Like SNP-GWAS tests, it is sensible to perform principal component analysis (PCA) on the derived copy number estimates (or CNV call sets) to obtain a set of potential confounding factors that can be included within an association testing model80. This should normally be done by using the copy number estimates across samples for rare and common events separately (common being defined as, for example, events seen an an estimated population frequency of greater than 1%), resulting in a set of PCs for each variant class. A standard approach would then be to include the first 20 PCs from each of the rare and common PCAs within the association testing model. The PCs from a CNV based PCA are likely to include measures of noise and other confounding factors that are difficult to model during copy number estimation as well as broad genetic ancestry groups and cryptic relatedness that should be accounted for during a genome wide test81.
For association testing, models can be applied using a variety of different assumptions21,50,57,82 but the most straightforward and the most similar to standard SNP-GWAS models is the use of a copy number estimate on the probe level as a linear dosage variable within a linear or logistic regression test. In this mode, the copy number estimates are used as a proxy for the underlying copy number events, which additionally allows CNVs with different boundaries around a region to be combined within the same test without the need to impose a set of overlap rules related to how individual CNVs should be combined or not. Furthermore, it negates some potential problems due to differential CNV calling sensitivities between samples for certain calling algorithms83. Like SNP-GWAS, it is sensible to include further covariates within the association model, some standard covariates would include, sex, age, height, weight, the precise definition of which should be considered with regard to the type of trait being tested84. Once the inputs and covariates to the model are decided, the most common approach would be to test each genomic region independently using a linear or logistic regression against the trait of interest, resulting in a P value and effect size estimate for all probe locations. Again, as for SNP-GWAS, a frequency cut-off and certain additional quality control measures should be used to exclude probe locations from the reported results. A standard cut-off of 1% or 5% population frequency for CNVs should be observed as a starting point, however care should be given to how accurate breakpoint estimation has been in the calling set and exactly how the reference frequency has been calculated, which could be defined using either external population cohorts, or an internal frequency estimate or both85,86. In practise it is possible that inclusion of CNV events or probe locations with much lower frequency estimates than the 1–5% range may still give reliable association results. Given reliable genome wide testing for CNV using a standard CNV GWAS model an important question arises into how an acceptable significance threshold should be set genome wide. A conservative approach is to use a Bonferroni correction87,88 to account for the number of tests performed and, just like which SNP based GWAS, the CNVs or probe location can be selected based on their observed frequency across a population. Other approaches could be based on FDR or permutation and there remains an open question into what a standard genome wide threshold for CNVs and could be based on factors such as how quickly LD between CNV regions breaks down genome wide89.
Sequencing-based CNV callers
Since the development of short-read NGS technology, a plethora of tools for calling CNVs and other variants from sequencing data have been developed83. However, these tend to be optimized for rare events, and very few can be scaled to biobank-sized data. To address this shortcoming, CNest, a read-depth-based caller, was developed and applied to 200,000 whole-exome sequences (WES) from UK Biobank, identifying 862 CNV associations with 78 different traits21. More recently, a haplotype-informed method was employed on 468,570 WES to identify CNVs associated with 41 quantitative traits along with hypertension and type 2 diabetes mellitus90. Towards the end of 2021, UK Biobank released whole-genome sequencing (WGS) data on 200,000 individuals, making it one of the largest WGS repositories in the world91. As part of this release, deCODE performed structural variant calling in 150,119 UK Biobank participants using Manta92, an Illumina software that uses evidence from mapped paired-end and split reads to identify CNVs and copy number neutral events, such as inversions, translocations and insertions. Calls were combined with those from a long-read sequencing study93 and genotyped across all individuals using GraphTyper94. Very similar methods were employed on the final set of half a million WGS from UK Biobank, except that Manta was replaced with the DRAGEN SV caller91. DRAGEN SV caller (Illumina) is also being applied at Genomics England, the All of Us Research Program49 and PRECISE (Singapore). These new methodologies show promise for expanding the genetic investigation of disease to variation previously untestable at scale.
Genome-wide association testing with CNVs
While association testing on single variants can link phenotypic consequences to individual events, it can be underpowered especially for rare events (both SNVs and CNVs). Many previous studies into rare SNV testing genome-wide have shown the benefit of including different levels of variant effect predictions, such as protein loss of function (pLOF) and rare deleterious missense variant within the burden style tests to boost statistical power for trait association testing95,96. To boost statistical power, discovery studies on CNVs have been performed at group-level by collapsing rare CNVs based on the functional regions they overlap (for example, genes) and CNV categories (that is, duplication or deletion). The significance of a set of CNVs associated with traits could be tested using the same approaches and implementations applied to sequence variants through burden tests (Fisher’s exact test for binary traits or linear regression for quantitative traits)82,90,97,98 and/or kernel-based tests (SKAT)99,100. While burden-style analysis interrogates the impact of CNVs that disrupt protein sequences for causing disease by using the presence of CNVs as an indicator variable, the models of CNVs that cause disease by modulating gene dosage can also be tested by using copy number estimates as continuous variables in a dosage-sensitive approach. CNV kernel association tests have also been proposed to account for the heterogeneity of CNVs of similar effects101, potentially improving discovery power.
For rare collapsing gene-based burden tests involving CNVs there is a need to classify individual events based on both positional and frequency information. Similar to SNVs, predictions into how damaging different CNVs are likely to be on gene function should be included, for example the difference between a whole gene duplication and a gene truncating duplication can be quite different with the latter being much more likely to result in a loss of function102. Further cases such as single exon of partial gene deletions could be assigned differential loss of function likelihoods based on the genomic content deleted and its likely impact on transcription and / or protein structure103. Other cases involving CNVs encompassing multiple genes can be difficult to model in a gene collapsing test with the simplest solution being to consider the CNV separately for each gene and only being included if it has the relevant characteristics for the gene in question (e.g. predicted loss of function). Some previous studies have shown the validity of moving away from the need to algorithmically call individual CNVs, followed by regional or gene-level burden style tests, and opt for a probe-level copy number estimate21,104 (Figure 3). Here, similar to SNV-GWAS approaches, rather than trying to understand the precise genome structure at any given genomic location, one performs a well-defined standard discovery process and only attempts to fine-map the underlying variants for statistically significant regions of interest105,106.
Figure 3. Differences between common SNP and CNV variant detection and example data flow for CNV GWAS results.
a) Differences in genome-wide resolution between array and exome sequence based GWAS tests, highlighting that most common SNPs are either intronic or intergenic whereas CNVs from exome sequencing alter coding elements. b) Data flow from population cohorts to data deposition of CNV-GWAS summary statistics and some examples of downstream applications.
Applications of these approaches in deep-phenotyped biobank cohorts have enabled the discovery of numerous novel CNVs modulating gene–trait associations21,50,57,82,90,98. We highlight a recent phenotype association study of coding CNVs using WES on 57 heritable quantitative traits in the UK Biobank90. Collapsing rare coding CNVs into protein loss-of-function variant counts with SNVs in gene burden testing discovered 100 new protein loss-of-function gene–trait associations (20% more compared with using SNVs/indels only), which also included some novel gene–trait associations. Application of dosage models by testing the regional common coding CNV burden with traits have also shown 99 loci that are not well-tagged by nearby SNPs after removal of all associations potentially explainable by linkage disequilibrium (LD) with imputed SNPs and indels within 3Mb highlighting the need to incorporate CNVs in GWAS and understanding the interaction between CNVs and SNVs with a reported 20% increase in the number of associations that can be made above a SNP/Indel only model90.
Limitations
Some technical challenges need to be overcome before CNV-GWAS can become more routine and of greater utility to downstream applications. These limitations can be grouped into three areas: data sources and availability, methods and models, and infrastructure and standards (Figure 3).
Data sources and availability
Most large human cohorts have varying levels of CNV-relevant data that often display marked differences in resolution and sensitivity. SNP genotyping data is the most commonly available but with limited resolution; short-read sequencing data is becoming more widespread with higher resolution; and long-read sequencing data has limited availability but the highest capacity to resolve complex rearrangements. Classical and more recent cytogenomic methods for CNV detection and interpretation include Fluorescence in situ hybridization (FISH)107,108 and optical mapping109,110 respectively and provide a different view and orthogonal technique for validating the underlying structure of complex loci111. All these data types have benefits and limitations that are important to consider when applying or integrating GWAS summary statistics. Furthermore, as better CNV imputation methods are developed, there is a need for accurate benchmarking of relevant imputation reference panels. CNV imputation methods could largely operate in a similar way to SNV imputation however haplotype reference panels will need to include high resolution phased CNV calls112. It is worth noting that the CNV information must be provided as event level genotypes along with SNV variation to allow linkage of CNV locations and copy number state to informative surrounding SNV genotypes.
A number of large human cohorts (for example, UK Biobank, Finngen and AllofUs) have generated CNV calls using a variety of CNV callers21,92,113,114, allowing large-scale CNV-GWAS to be further explored. Agreement on a set of standard GWAS models for CNVs will have several key advantages and perhaps the most tractable approach is the use of a probe-level CNV dosage variable in a standard linear GWAS test with an additive assumption84. Since CNVs are a major source of genetic diversity, addressing the over-representation of individuals clustering with the 1000 Genomes EUR superpopulation is a key goal to expand the utility of CNV-GWAS. Extensive African CNV diversity remains under-ascertained. Continued efforts to increase cohort sizes across diverse sets of population groups are urgently required to address the issue.
Methods and models
Many methods are available for calling CNVs from NGS data all of which use different modelling choices83. There is also a lack of agreement in the community over which calling software is most appropriate for different applications and a pervasive concern over large numbers of false positives115, plus no gold standard or community agreed CNV-GWAS tools and guidelines. An increasing number of publications describe CNV modelling applied to trait association testing (Supplementary File 2), yet the majority still use SNP genotyping arrays or targeted sequencing, which suffer from low resolution and limited dose response116,117. Furthermore, most CNV-GWAS models involve defining a set of overlap and merging rules followed by burden-style tests98,118, and there is no standard software choice for discovery-mode CNV-GWAS, the equivalent to packages such as REGENIE80 or PLINK119 for SNP-GWAS studies. There are currently only a few studies that have used NGS-derived CNV data in a CNV-GWAS setting21,90,104,120 and no standard file formats to contain CNV dosages genome-wide.
Although methods using co-inheritance patterns between SNPs and CNVs (that is, CNV tagging SNPs) have been described121–123, there are no established community-maintained resources to allow these methods to be accurately benchmarked. Having community-led guidelines and agreed approaches for CNV modelling within a GWAS setting opens many opportunities for further methods development and extended applications. However, one key resource that is needed is a set of well-maintained and reliable LD maps between SNPs and CNVs, allowing linkage of GWAS signals and inheritance patterns between variant classes and providing sufficient coverage of human genetic diversity to ensure that results are equitable for all. Information into when and how SNP-CNV LD breaks down genome-wide will be valuable for joint modelling of SNPs and CNVs, variant to gene mapping and allow assessments into how much CNV tagging is possible using common SNPs, and to what degree trait to CNV associations may have been modelled by previous SNP-GWAS.
Standards and infrastructure
Standards in CNV association reporting are lacking, which corresponds to difficulties in including data in repositories. For example, the GWAS Catalog accepts submissions of CNV associations, but the variant representation differs widely between datasets, including chromosome and base-pair location with deletion or duplication indicated either as the effect allele or in the statistical model, or genomic location with start and end of the window. Different association models may be included in the same or separate files. Key data and metadata fields have not been defined, with varying levels of detail provided about the models used, making it difficult to compare datasets without referring to the original publications. Consequently, CNV data shared in the GWAS Catalog are only available as downloadable files and not integrated into the searchable dataset, meaning cross-study queries are not possible, limiting accessibility. Using data encompassing CNV-GWAS summary statistics submission, literature search and GWAS Catalog CNV pilot data, we found 116 manuscripts that described some level of CNV association testing, only 4 of which were obtained via data submission to the Catalog (Supplementary File 2). The remaining 114 potential datasets contain a diverse set of sample sizes and different modelling choices, which will require expert curation prior to potential inclusion in public resources; of note, most have not made their genome-wide (or regional) summary statistics available following publication.
Conversely, standards have been integral to the utility of SNP-GWAS data. The data model defined by the GWAS Catalog at inception, and subsequently developed standards for sample metadata124 and genome-wide summary statistics submissions125, have enabled GWAS data to be reused in a wide range of contexts, for example, analysed by software packages (gwasrapidd126, pandasGWAS127), converted to formats suitable for Mendelian randomization (OpenGWAS128), integrated with other omics data for drug discovery (Open Targets8, Knowledge Portal network129) and linked to PGS generated from the data (PGS Catalog). Recent implementation of the community-led standard for GWAS summary statistics increased the number of shared datasets containing all key data fields required for PGS and Mendelian randomization from 50% to ~100%125. Availability of a submission portal for SNP-GWAS endorsed by journals, and associated tools for data preparation, have resulted in full genome-wide datasets being shared for more than 30% of all new publications. Equivalent standards for sharing CNV-GWAS data are urgently required. Whilst the growth of SNP-GWAS was initially slow (~800 GWAS published in 2008–2010 vs ~5,000 in 2018–2020) allowing standards to evolve over time, reduced genotyping cost, availability of deeply phenotyped datasets, and increased speed of computation mean that CNV-GWAS are likely to proliferate rapidly, resulting in a huge number of non-interoperable datasets becoming available.
Standardization of CNVs and sharing of full datasets in a standardized repository such as the GWAS Catalog has potential to hugely increase the downstream utility of CNV-GWAS, including interpretation of SNPs and GWAS in context. Integration of SNP and CNV data in the same resource and particularly in the context of LD will facilitate the assessment of whether a given CNV strengthens the evidence for a known locus or represents a novel finding and allow joint modelling with SNPs in effector gene/drug discovery pipelines, in the generation of PGSs and in Mendelian randomization analyses.
Emerging opportunities
Long reads and graph genomes
Long-read sequencing data in theory give access to very precise breakpoint positions for CNVs and other SV owing to the ability to provide single reads spanning entire structural variant boundaries130 (Figure 4). Several SV callers have been developed specifically for long-read sequencing131–134. However, the availability of long-read data in very large human cohorts remains limited, and computing requirements for storing such datasets and applying SV detection methods are a substantial challenge19,135,136. Studies such as the 1000 Genomes ONT sequencing consortium137 have generated high-depth long-read data in thousands of individuals from diverse populations138 and are applying new techniques to represent a more complete view of genomic variants using methods such as pan-genome graphs139,140. The use of long-read-derived SV representations allows imputation of SVs from genotyping or short-read sequencing data141, enabling an improved ability to call, characterize and interpret SVs across large groups of individuals.
Figure 4. Discerning the structural variation landscape of CNV associations by using genome graphs.
1) The result of a standard probe level genome wide test for CNVs discovering several significant associations. 2) Identification of individuals carrying CNV changes of an effect allele for a single association locus. 3) Alignment of short read data in relevant individuals to a genome graph to help discern the precise structure and breakpoints of the underlying genomic variants responsible for the association observed.
Polygenic scores
GWAS summary statistics and/or individual-level data can be used to estimate an individual’s genetic propensity for a specific trait or disease. The effects of all relevant genetic variants can be aggregated into a single number proportional to the genetic propensity, commonly referred to as a PGS. PGSs are of intense interest in human genetics as they have been shown to be significantly predictive for many traits and diseases, frequently beyond conventional (risk) factors142,143. PGSs have been shown to have potential clinical utility, and there have been recent efforts to develop common standards for reporting, clinical implementation and risk communication7,144,145.
PGSs have historically been constructed using SNP data, primarily because SNPs have been the dominant variant type in GWAS, and it is more straightforward to model one variant type. However, studies have begun to combine SNP-based PGSs together with CNVs146 to assess their relative and joint contribution to neuropsychiatric and cognitive traits, for example schizophrenia146,147 and psychopathology148,149. Recently, these analyses have been extended to a comparative analysis of schizophrenia, cognitive and socioeconomic phenotypes150, emphasizing the importance of analysing CNVs in the context of broader (genome-wide) polygenic risk. Importantly, PGSs frequently harbour limitations when it comes to generalizability, particularly with respect to transferring scores across genetic ancestries151. Relatively few studies have assessed the transferability of CNV signals across genetic ancestries and even fewer seek to combine this with polygenic risk information.
PGS ideally comprise the totality of genome variation that can be both functionally linked as causal of a given phenotype or which maximally tags the true but unverifiable causal variant(s). In this way, we expect the inclusion of CNVs either as more efficient tags or as the true causal variants themselves to improve the performance of PGSs in some cases (i.e. large effect sizes and difficult to tag CNVs) and therefore their potential utility in aetiological studies and clinical risk prediction. Diversifying the types of genetic variation beyond SNPs to CNVs may in principle improve the potential for approaches which disrupt polygenic risk. Furthermore, for risk prediction, the identification of causal alleles and corresponding effects both improve the performance of PGS within ancestry and their transferability between ancestries152,153. The inclusion of CNVs in fine mapping is anticipated to improve the identification of causal alleles and thus the downstream transferability of PGSs as they are equitably integrated into healthcare systems.
Rare disease and common CNV modifiers
Genetic variants disrupting the same gene can occur at the same phase (disrupting one gene copy) or different phases (disrupting two gene copies) and thus could have different effects. Genes with recessive inheritance are usually identified in large human cohorts with low effective population size (such as Finnish154) or high consanguinity (such as South Asian populations155). Software packages such as Beagle156, Eagle2157, and SHAPEIT158 use population haplotype reference panels and frequency information to accurately impute variation. However, these methods rely on the reference set containing all variants of interest and so phasing rare variation can be a challenge depending on the size and completeness of the haplotype reference panel. Recent development of accurate statistical phasing methods scalable to biobank cohorts, such the SHAPEIT5 algorithm159 and including NGS data genome wide are providing novel opportunities to test associations between phase-aware allelic series and phenotypes for both common and rare variation159,160. Their application to UK Biobank data enabled the discovery of compound heterozygous events in the general population and their contribution to complex common diseases159,161. However, the phenotypic impact of other complex allelic series, such as considering CNV phase with or without SNVs, has not been studied. Further discovery studies using phase information would not only help to discover new disease-associated genes but also update our knowledge on essential human genes, thus providing genetic support for drug target development.
Drug target identification
As well as understanding the potential impact of genetic variation on how we respond to medications (pharmacogenetics), CNV-GWAS can provide new insights for the identification of novel drug targets or help in the prioritisation of targets for drug discovery. Systematic CNV-GWAS across large biobank data and diverse populations will provide novel trait associations, which may provide new potential targets. They may also provide further evidence to strengthen a gene–disease association where previously only signals from non-coding SNVs have been identified and mapped to a likely causal gene through predictive methods such as Locus2Gene162. CNV–trait associations importantly provide evidence for the effect of modulation that a drug may have, and thus inform on what modulation and dosage level may be suitable to produce a particular therapeutic outcome. For example, whether inhibition of the target (gene deletion) versus activation (gene duplication) would be therapeutically beneficial. Conversely, this can also inform on whether modulation of the target would be detrimental and likely result in side effects. These can help prioritise targets for your mechanism of action of choice, and de-prioritize targets that may have severe unwanted consequences if modulated.
Conclusions
Many SNP-GWAS have collectively contributed to an extensive and publicly available genetic association knowledge base that has had an impact on multiple domains including clinical genetics, therapeutic target identification, complex traits and population genetics1,163–165. Nonetheless, the mapping of common SNP-GWAS loci to causal variants or genes remains challenging, and the contribution of rare SNPs and other classes of genetic variation, such as CNVs, remains under-represented in human studies. The next phase of GWAS is likely to be a combination of increasingly wider applications of high-throughput sequencing, allowing a better representation of rare variants and extensive modelling of other variant classes, such as CNV. Advances in these areas could then be combined to improve discovery potential, facilitate further association methods (such as joint SNV–CNV models), enhance prediction of causal variants166 and give rise to a new generation of resources aimed at improving the interpretability of GWAS signals in the context of other genetic variants, genetic backgrounds and environmental factors.
The future of GWAS will include some CNV-specific components, both individually and jointly with other variant classes (for example, SNPs). There is a need to standardize model choices to allow better integration with existing resources and the large body of knowledge available from SNP-GWAS repositories that serve multiple downstream applications. We recommend that the community adopts a standard set of CNV-GWAS models facilitating the integration of CNV-GWAS results in public repositories and maximising their utility to downstream applications. A probe-level test using linear or logistic regression models with an additive assumption could form the basis of a standard CNV-GWAS test, with different modelling choices, such as non-linear models or burden style tests being extensions to the standard test. The use of such a standard model not only gives better linkage between summary statistics across different cohorts but also improves the ability to interpret GWAS signals between SNPs and CNVs in context, due to similar modelling assumptions for genome-wide discovery. The ability to interpret CNV-GWAS and SNP-GWAS signals in context holds great promise for improving our understanding of genome to trait associations and has applications in causal variant mapping, joint modelling and downstream applications such as drug discovery and genetic disease risk modelling.
CNV-GWAS datasets will benefit from the annotation and standardization that has already been applied to SNP-GWAS, for example ontology annotation of phenotype definitions and controlled vocabulary for population descriptors facilitating combination and comparison of datasets. For CNV-GWAS data to adhere to the FAIR principles, data must be described with appropriate metadata, accessioned with a persistent identifier, indexed and searchable in a recognized resource such as the GWAS Catalog (Findable, Accessible). The metadata must be described with formal language that is shared with other data types (Interoperable) and contain minimum elements for reuse of the data according to agreed community standards (Reusable). For example, key metadata for CNVs may include caller and model, with controlled vocabulary for each. There is a need for community engagement to define these elements, before resources such as the GWAS Catalog can fully integrate CNV-GWAS datasets and serve data for the maximum range of downstream uses.
Methods related to linkage between SNPs and CNVs need further development and improvement using population-scale assessments into variant tagging and shared inheritance patterns. The resulting SNP-CNV resources are a critical component that is currently missing but acutely needed for the interpretation of genome wide association results in context. These resources also drive the development of better genome imputation models at the population level allowing a more complete picture of genetic variation and the relationship between variant classes. Full integration and accurate representation of GWAS signals in context within public resources such as the GWAS Catalog will serve as an important new resource to the community.
Supplementary Material
Footnotes
Open Targets Platform https://platform.opentargets.org/
gnomAD https://gnomad.broadinstitute.org/
GWAS Catalog https://www.ebi.ac.uk/gwas/
References
- [1].Wellcome Trust Case Control Consortium. Genome-wide association study of 14,000 cases of seven common diseases and 3,000 shared controls. Nature. 2007;447(7145):661–678. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [2].Barrett JC, Cardon LR. Evaluating coverage of genome-wide association studies. Nat Genet. 2006;38(6):659–662. [DOI] [PubMed] [Google Scholar]
- [3].LaFramboise T. Single nucleotide polymorphism arrays: a decade of biological, computational and technological advances. Nucleic Acids Res. 2009;37(13):4181–4193. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [4].Hofker MH, Fu J, Wijmenga C. The genome revolution and its role in understanding complex diseases. Biochim Biophys Acta. 2014;1842(10):1889–1895. [DOI] [PubMed] [Google Scholar]
- [5].Sollis E, Mosaku A, Abid A, Buniello A, Cerezo M, Gil L, et al. The NHGRI-EBI GWAS Catalog: knowledgebase and deposition resource. Nucleic Acids Res. 2023;51(D1):D977–D985. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [6].Wilkinson MD, Dumontier M, Aalbersberg IJJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [7].Wand H, Lambert SA, Tamburro C, Iacocca MA, O’Sullivan JW, Sillari C, et al. Improving reporting standards for polygenic scores in risk prediction studies. Nature. 2021;591(7849):211–219. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [8].Ochoa D, Hercules A, Carmona M, Suveges D, Baker J, Malangone C, et al. The next-generation Open Targets Platform: reimagined, redesigned, rebuilt. Nucleic Acids Res. 2023;51(D1):D1353–D1359. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [9].Yengo L, Vedantam S, Marouli E, Sidorenko J, Bartell E, Sakaue S, et al. A saturated map of common genetic variants associated with human height. Nature. 2022;610(7933):704–712. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [10].Yengo L, Sidorenko J, Kemper KE, Zheng Z, Wood AR, Weedon MN, et al. Meta-analysis of genome-wide association studies for height and body mass index in ∼700000 individuals of European ancestry. Hum Mol Genet. 2018;27(20):3641–3649. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [11].Zhu H, Zhou X. Statistical methods for SNP heritability estimation and partition: A review. Comput Struct Biotechnol J. 2020;18:1557–1568. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [12].Deciphering Developmental Disorders Study. Large-scale discovery of novel genetic causes of developmental disorders. Nature. 2015;519(7542):223–228. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [13].Manolio TA, Collins FS, Cox NJ, Goldstein DB, Hindorff LA, Hunter DJ, et al. Finding the missing heritability of complex diseases. Nature. 2009;461(7265):747–753. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [14].Yang L. A practical guide for structural variation detection in the human genome. Curr Protoc Hum Genet. 2020;107(1). doi: 10.1002/cphg.103 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [15].Montavon T, Thevenet L, Duboule D. Impact of copy number variations (CNVs) on long-range gene regulation at the HoxD locus. Proc Natl Acad Sci U S A. 2012;109(50):20204–20211. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [16].Wellcome Trust Case Control Consortium, Craddock N, Hurles ME, Cardin N, Pearson RD, Plagnol V, et al. Genome-wide association study of CNVs in 16,000 cases of eight common diseases and 3,000 shared controls. Nature. 2010;464(7289):713–720. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [17].Verlouw JAM, Clemens E, de Vries JH, Zolk O, Verkerk AJMH, Am Zehnhoff-Dinnesen A, et al. A comparison of genotyping arrays. Eur J Hum Genet. 2021;29(11):1611–1624. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [18].Rapti M, Zouaghi Y, Meylan J, Ranza E, Antonarakis SE, Santoni FA. CoverageMaster: comprehensive CNV detection and visualization from NGS short reads for genetic medicine applications. Brief Bioinform. 2022;23(2). doi: 10.1093/bib/bbac049 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [19].Tanjo T, Kawai Y, Tokunaga K, Ogasawara O, Nagasaki M. Practical guide for managing large-scale human genome data in research. J Hum Genet. 2021;66(1):39–52. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [20].Vacic V, McCarthy S, Malhotra D, Murray F, Chou HH, Peoples A, et al. Duplications of the neuropeptide receptor gene VIPR2 confer significant risk for schizophrenia. Nature. 2011;471(7339):499–503. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [21].Fitzgerald T, Birney E. CNest: A novel copy number association discovery method uncovers 862 new associations from 200,629 whole-exome sequence datasets in the UK Biobank. Cell Genom. 2022;2(8):100167. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [22].Conrad DF, Hurles ME. The population genetics of structural variation. Nat Genet. 2007;39(7 Suppl):S30–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [23].Conrad DF, Pinto D, Redon R, Feuk L, Gokcumen O, Zhang Y, et al. Origins and functional impact of copy number variation in the human genome. Nature. 2010;464(7289):704–712. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [24].Lee C, Scherer SW. The clinical context of copy number variation in the human genome. Expert Rev Mol Med. 2010;12:e8. [DOI] [PubMed] [Google Scholar]
- [25].Lupski JR. Genomic rearrangements and sporadic disease. Nat Genet. 2007;39(7 Suppl):S43–7. [DOI] [PubMed] [Google Scholar]
- [26].Campbell CD, Eichler EE. Properties and rates of germline mutations in humans. Trends Genet. 2013;29(10):575–584. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [27].Belyeu JR, Brand H, Wang H, Zhao X, Pedersen BS, Feusier J, et al. De novo structural mutation rates and gamete-of-origin biases revealed through genome sequencing of 2,396 families. Am J Hum Genet. 2021;108(4):597–607. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [28].Gudmundsson S, Singer-Berk M, Watts NA, Phu W, Goodrich JK, Solomonson M, et al. Variant interpretation using population databases: Lessons from gnomAD. Hum Mutat. 2022;43(8):1012–1030. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [29].Chen S, Francioli LC, Goodrich JK, Collins RL, Kanai M, Wang Q, et al. A genomic mutational constraint map using variation in 76,156 human genomes. Nature. 2024;625(7993):92–100. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [30].Sudmant PH, Rausch T, Gardner EJ, Handsaker RE, Abyzov A, Huddleston J, et al. An integrated map of structural variation in 2,504 human genomes. Nature. 2015;526(7571):75–81. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [31].Taghizadeh S, Gholizadeh M, rahimi-Mianji G, Moradi MH, Costilla R, Moore S, et al. Genome-wide identification of copy number variation and association with fat deposition in thin and fat-tailed sheep breeds. Sci Rep. 2022;12(1). doi: 10.1038/s41598-022-12778-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [32].Delledonne A, Punturiero C, Ferrari C, Bernini F, Milanesi R, Bagnato A, et al. Copy number variant scan in more than four thousand Holstein cows bred in Lombardy, Italy. PLoS One. 2024;19(5):e0303044. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [33].Zhang F, Gu W, Hurles ME, Lupski JR. Copy number variation in human health, disease, and evolution. Annu Rev Genomics Hum Genet. 2009;10:451–481. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [34].Wright CF, Fitzgerald TW, Jones WD, Clayton S, McRae JF, van Kogelenberg M, et al. Genetic diagnosis of developmental disorders in the DDD study: a scalable analysis of genome-wide research data. Lancet. 2015;385(9975):1305–1314. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [35].Redon R, Ishikawa S, Fitch KR, Feuk L, Perry GH, Andrews TD, et al. Global variation in copy number in the human genome. Nature. 2006;444(7118):444–454. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [36].Coutelier M, Holtgrewe M, Jäger M, Flöttman R, Mensah MA, Spielmann M, et al. Combining callers improves the detection of copy number variants from whole-genome sequencing. Eur J Hum Genet. 2022;30(2):178–186. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [37].Hollox EJ, Zuccherato LW, Tucci S. Genome structural variation in human evolution. Trends Genet. 2022;38(1):45–58. [DOI] [PubMed] [Google Scholar]
- [38].Rossi N, Aliyev E, Visconti A, Akil ASA, Syed N, Aamer W, et al. Ethnic-specific association of amylase gene copy number with adiposity traits in a large Middle Eastern biobank. NPJ Genom Med. 2021;6(1):8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [39].Perry GH, Dominy NJ, Claw KG, Lee AS, Fiegler H, Redon R, et al. Diet and the evolution of human amylase gene copy number variation. Nat Genet. 2007;39(10):1256–1260. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [40].Higuchi R, Iwane T, Iida A, Nakajima K. Copy Number Variation of the Salivary Amylase Gene and Glucose Metabolism in Healthy Young Japanese Women. J Clin Med Res. 2020;12(3):184–189. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [41].Collins RL, Glessner JT, Porcu E, Lepamets M, Brandon R, Lauricella C, et al. A cross-disorder dosage sensitivity map of the human genome. Cell. 2022;185(16):3041–3055.e25. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [42].Barra V, Fachinetti D. The dark side of centromeres: types, causes and consequences of structural abnormalities implicating centromeric DNA. Nat Commun. 2018;9(1). doi: 10.1038/s41467-018-06545-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- [43].Cook CB, Armstrong L, Boerkoel CF, Clarke LA, du Souich C, Demos MK, et al. Somatic mosaicism detected by genome-wide sequencing in 500 parent–child trios with suspected genetic disease: clinical and genetic counseling implications. Cold Spring Harb Mol Case Stud. 2021;7(6):a006125. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [44].Elrick H, Sauer CM, Espejo Valle-Inclan J, Trevers K, Tanguy M, Zumalave S, et al. SAVANA: reliable analysis of somatic structural variants and copy number aberrations in clinical samples using long-read sequencing. bioRxiv. Published online July 25, 2024. doi: 10.1101/2024.07.25.604944 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [45].Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q, et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature. 2020;581(7809):434–443. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [46].Thaxton C, Good ME, DiStefano MT, Luo X, Andersen EF, Thorland E, et al. Utilizing ClinGen gene-disease validity and dosage sensitivity curations to inform variant classification. Hum Mutat. 2022;43(8):1031–1040. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [47].Huang N, Lee I, Marcotte EM, Hurles ME. Characterising and predicting haploinsufficiency in the human genome. PLoS Genet. 2010;6(10):e1001154. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [48].Rice AM, McLysaght A. Dosage-sensitive genes in evolution and disease. BMC Biol. 2017;15(1):78. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [49].All of Us Research Program Genomics Investigators. Genomic data in the All of Us Research Program. Nature. Published online February 19, 2024. doi: 10.1038/s41586-023-06957-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- [50].Auwerx C, Jõeloo M, Sadler MC, Tesio N, Ojavee S, Clark CJ, et al. Rare copy-number variants as modulators of common disease susceptibility. Genome Med. 2024;16(1):5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [51].Kirschner R, Rosenberg T, Schultz-Heienbrok R, Lenzner S, Feil S, Roepman R, et al. RPGR transcription studies in mouse and human tissues reveal a retina-specific isoform that is disrupted in a patient with X-linked retinitis pigmentosa. Hum Mol Genet. 1999;8(8):1571–1578. [DOI] [PubMed] [Google Scholar]
- [52].Shaikh TH. Copy Number Variation Disorders. Curr Genet Med Rep. 2017;5(4):183–190. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [53].Xu HH, Zhang Y, He ZH, Di XH, Pan FY, Shi WW. Familial 5.29 Mb deletion in chromosome Xq22.1-q22.3 with a normal phenotype: a rare pedigree and literature review. BMC Med Genomics. 2023;16(1):111. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [54].Naseer MI, Chaudhary AG, Rasool M, Kalamegam G, Ashgan FT, Assidi M, et al. Copy number variations in Saudi family with intellectual disability and epilepsy. BMC Genomics. 2016;17(Suppl 9):757. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [55].Wolstencroft J, Wicks F, Srinivasan R, Wynn S, Ford T, Baker K, et al. Neuropsychiatric risk in children with intellectual disability of genetic origin: IMAGINE, a UK national cohort study. Lancet Psychiatry. 2022;9(9):715–724. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [56].Zarrei M, Burton CL, Engchuan W, Higginbotham EJ, Wei J, Shaikh S, et al. Gene copy number variation and pediatric mental health/neurodevelopment in a general population. Hum Mol Genet. 2023;32(15):2411–2421. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [57].Auwerx C, Lepamets M, Sadler MC, Patxot M, Stojanov M, Baud D, et al. The individual and global impact of copy-number variants on complex human traits. Am J Hum Genet. Published online February 25, 2022. doi: 10.1016/j.ajhg.2022.02.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [58].Ceyhan-Birsoy O, Pugh TJ, Bowser MJ, Hynes E, Frisella AL, Mahanta LM, et al. Next generation sequencing-based copy number analysis reveals low prevalence of deletions and duplications in 46 genes associated with genetic cardiomyopathies. Mol Genet Genomic Med. 2016;4(2):143–151. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [59].Singer ES, Ross SB, Skinner JR, Weintraub RG, Ingles J, Semsarian C, et al. Characterization of clinically relevant copy-number variants from exomes of patients with inherited heart disease and unexplained sudden cardiac death. Genet Med. 2021;23(1):86–93. [DOI] [PubMed] [Google Scholar]
- [60].Nfonsam L, Huang L, Carson N, McGowan-Jordan J, Beaulieu Bergeron M, Goobie S, et al. ALU transposition induces familial hypertrophic cardiomyopathy. Mol Genet Genomic Med. 2020;8(1):e951. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [61].Wilfert AB, Sulovari A, Turner TN, Coe BP, Eichler EE. Recurrent de novo mutations in neurodevelopmental disorders: properties and clinical implications. Genome Med. 2017;9(1). doi: 10.1186/s13073-017-0498-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- [62].Malhotra D, Sebat J. CNVs: harbingers of a rare variant revolution in psychiatric genetics. Cell. 2012;148(6):1223–1241. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [63].Marshall CR, Howrigan DP, Merico D, Thiruvahindrapuram B, Wu W, Greer DS, et al. Contribution of copy number variants to schizophrenia from a genome-wide study of 41,321 subjects. Nat Genet. 2017;49(1):27–35. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [64].Davies RW, International 22q11.2 Brain and Behavior Consortium, Fiksinski AM, Breetvelt EJ, Williams NM, Hooper SR, et al. Using common genetic variation to examine phenotypic expression and risk prediction in 22q11.2 deletion syndrome. Nat Med. 2020;26(12):1912–1918. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [65].Maury EA, Sherman MA, Genovese G, Gilgenast TG, Kamath T, Burris SJ, et al. Schizophrenia-associated somatic copy-number variants from 12,834 cases reveal recurrent NRXN1 and ABCB11 disruptions. Cell Genom. 2023;3(8):100356. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [66].Trost B, Thiruvahindrapuram B, Chan AJS, Engchuan W, Higginbotham EJ, Howe JL, et al. Genomic architecture of autism from comprehensive whole-genome sequence annotation. Cell. 2022;185(23):4409–4427.e18. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [67].Riggs ER, Andersen EF, Cherry AM, Kantarci S, Kearney H, Patel A, et al. Technical standards for the interpretation and reporting of constitutional copy-number variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics (ACMG) and the Clinical Genome Resource (ClinGen). Genet Med. 2020;22(2):245–257. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [68].Hippman C, Nislow C. Pharmacogenomic Testing: Clinical Evidence and Implementation Challenges. J Pers Med. 2019;9(3). doi: 10.3390/jpm9030040 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [69].Crews KR, Monte AA, Huddart R, Caudle KE, Kharasch ED, Gaedigk A, et al. Clinical Pharmacogenetics Implementation Consortium Guideline for CYP2D6, OPRM1, and COMT Genotypes and Select Opioid Therapy. Clin Pharmacol Ther. 2021;110(4):888–896. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [70].Twesigomwe D, Drögemöller BI, Wright GEB, Adebamowo C, Agongo G, Boua PR, et al. Characterization of CYP2D6 Pharmacogenetic Variation in Sub-Saharan African Populations. Clin Pharmacol Ther. 2023;113(3):643–659. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [71].Twist GP, Gaedigk A, Miller NA, Farrow EG, Willig LK, Dinwiddie DL, et al. Constellation: a tool for rapid, automated phenotype assignment of a highly polymorphic pharmacogene, CYP2D6, from whole-genome sequences. NPJ Genom Med. 2016;1:15007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [72].Lee SB, Wheeler MM, Patterson K, McGee S, Dalton R, Woodahl EL, et al. Stargazer: a software tool for calling star alleles from next-generation sequencing data using CYP2D6 as a model. Genet Med. 2019;21(2):361–372. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [73].Chen X, Shen F, Gonzaludo N, Malhotra A, Rogert C, Taft RJ, et al. Cyrius: accurate CYP2D6 genotyping using whole-genome sequencing data. Pharmacogenomics J. 2021;21(2):251–261. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [74].Twesigomwe D, Drögemöller BI, Wright GEB, Siddiqui A, da Rocha J, Lombard Z, et al. StellarPGx: A Nextflow Pipeline for Calling Star Alleles in Cytochrome P450 Genes. Clin Pharmacol Ther. 2021;110(3):741–749. [DOI] [PubMed] [Google Scholar]
- [75].Cavallari LH, Johnson JA. A case for genotype-guided pain management. Pharmacogenomics. 2019;20(10):705–708. [DOI] [PubMed] [Google Scholar]
- [76].Tayeh MK, Gaedigk A, Goetz MP, Klein TE, Lyon E, McMillin GA, et al. Clinical pharmacogenomic testing and reporting: A technical standard of the American College of Medical Genetics and Genomics (ACMG). Genet Med. 2022;24(4):759–768. [DOI] [PubMed] [Google Scholar]
- [77].Singh AK, Olsen MF, Lavik LAS, Vold T, Drabløs F, Sjursen W. Detecting copy number variation in next generation sequencing data from diagnostic gene panels. BMC Med Genomics. 2021;14(1):214. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [78].Wang K, Li M, Hadley D, Liu R, Glessner J, Grant SFA, et al. PennCNV: an integrated hidden Markov model designed for high-resolution copy number variation detection in whole-genome SNP genotyping data. Genome Res. 2007;17(11):1665–1674. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [79].Behera S, Catreux S, Rossi M, Truong S, Huang Z, Ruehle M, et al. Comprehensive and accurate genome analysis at scale using DRAGEN accelerated algorithms. bioRxiv. Published online January 6, 2024. doi: 10.1101/2024.01.02.573821 [DOI] [Google Scholar]
- [80].Mbatchou J, Barnard L, Backman J, Marcketta A, Kosmicki JA, Ziyatdinov A, et al. Computationally efficient whole-genome regression for quantitative and binary traits. Nat Genet. Published online May 20, 2021. doi: 10.1038/s41588-021-00870-7 [DOI] [PubMed] [Google Scholar]
- [81].Romdhane L, Kefi S, Mezzi N, Abassi N, Jmel H, Romdhane S, et al. Ethnic and functional differentiation of copy number polymorphisms in Tunisian and HapMap population unveils insights on genome organizational plasticity. Sci Rep. 2024;14(1). doi: 10.1038/s41598-024-54749-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [82].Hujoel MLA, Sherman MA, Barton AR, Mukamel RE, Sankaran VG, Terao C, et al. Influences of rare copy-number variation on human complex traits. Cell. 2022;185(22):4233–4248.e27. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [83].Gabrielaite M, Torp MH, Rasmussen MS, Andreu-Sánchez S, Vieira FG, Pedersen CB, et al. A Comparison of Tools for Copy-Number Variation Detection in Germline Whole Exome and Whole Genome Sequencing Data. Cancers. 2021;13(24). doi: 10.3390/cancers13246283 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [84].Uffelmann E, Huang QQ, Munung NS, de Vries J, Okada Y, Martin AR, et al. Genome-wide association studies. Nature Reviews Methods Primers. 2021;1(1):1–21. [Google Scholar]
- [85].Collins RL, Brand H, Karczewski KJ, Zhao X, Alföldi J, Francioli LC, et al. A structural variation reference for medical and population genetics. Nature. 2020;581(7809):444–451. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [86].Gross AM, Ajay SS, Rajan V, Brown C, Bluske K, Burns NJ, et al. Copy-number variants in clinical genome sequencing: deployment and interpretation for rare and undiagnosed disease. Genet Med. 2019;21(5):1121–1130. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [87].Fadista J, Manning AK, Florez JC, Groop L. The (in)famous GWAS P-value threshold revisited and updated for low-frequency variants. Eur J Hum Genet. 2016;24(8):1202–1205. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [88].Kaler AS, Purcell LC. Estimation of a significance threshold for genome-wide association studies. BMC Genomics. 2019;20(1). doi: 10.1186/s12864-019-5992-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [89].Null M, Yilmaz F, Astling D, Yu HC, Cole JB, Hallgrímsson B, et al. Genome-wide analysis of copy number variants and normal facial variation in a large cohort of Bantu Africans. HGG Adv. 2022;3(1):100082. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [90].Hujoel MLA, Handsaker RE, Sherman MA, Kamitaki N, Barton AR, Mukamel RE, et al. Hidden protein-altering variants influence diverse human phenotypes. bioRxiv. Published online June 9, 2023. doi: 10.1101/2023.06.07.544066 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [91].Li S, Carss KJ, Halldorsson BV, Cortes A, UK Biobank Whole-Genome Sequencing Consortium. Whole-genome sequencing of half-a-million UK Biobank participants. bioRxiv. Published online December 8, 2023. doi: 10.1101/2023.12.06.23299426 [DOI] [Google Scholar]
- [92].Halldorsson BV, Eggertsson HP, Moore KHS, Hauswedell H, Eiriksson O, Ulfarsson MO, et al. The sequences of 150,119 genomes in the UK Biobank. Nature. 2022;607(7920):732–740. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [93].Beyter D, Ingimundardottir H, Oddsson A, Eggertsson HP, Bjornsson E, Jonsson H, et al. Long-read sequencing of 3,622 Icelanders provides insight into the role of structural variants in human diseases and other traits. Nat Genet. 2021;53(6):779–786. [DOI] [PubMed] [Google Scholar]
- [94].Eggertsson HP, Kristmundsdottir S, Beyter D, Jonsson H, Skuladottir A, Hardarson MT, et al. GraphTyper2 enables population-scale genotyping of structural variation using pangenome graphs. Nat Commun. 2019;10(1):5402. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [95].Backman JD, Li AH, Marcketta A, Sun D, Mbatchou J, Kessler MD, et al. Exome sequencing and analysis of 454,787 UK Biobank participants. Nature. 2021;599(7886):628–634. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [96].Wang Q, Dhindsa RS, Carss K, Harper AR, Nag A, Tachmazidou I, et al. Rare variant contribution to human disease in 281,104 UK Biobank exomes. Nature. 2021;597(7877):527–532. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [97].Li YR, Glessner JT, Coe BP, Li J, Mohebnasab M, Chang X, et al. Rare copy number variants in over 100,000 European ancestry subjects reveal multiple disease associations. Nat Commun. 2020;11(1):255. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [98].Aguirre M, Rivas MA, Priest J. Phenome-wide Burden of Copy-Number Variation in the UK Biobank. Am J Hum Genet. 2019;105(2):373–383. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [99].Babadi M, Fu JM, Lee SK, Smirnov AN, Gauthier LD, Walker M, et al. GATK-gCNV enables the discovery of rare copy number variants from exome sequencing data. Nat Genet. 2023;55(9):1589–1597. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [100].Wu MC, Lee S, Cai T, Li Y, Boehnke M, Lin X. Rare-variant association testing for sequencing data with the sequence kernel association test. Am J Hum Genet. 2011;89(1):82–93. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [101].Zhan X, Girirajan S, Zhao N, Wu MC, Ghosh D. A novel copy number variants kernel association test with application to autism spectrum disorders studies. Bioinformatics. 2016;32(23):3603–3610. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [102].Dougherty ML, Underwood JG, Nelson BJ, Tseng E, Munson KM, Penn O, et al. Transcriptional fates of human-specific segmental duplications in brain. Genome Res. 2018;28(10):1566–1576. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [103].Egorova TV, Galkin II, Velyaev OA, Vassilieva SG, Savchenko IM, Loginov VA, et al. In-frame deletion of dystrophin exons 8–50 results in DMD phenotype. Int J Mol Sci. 2023;24(11):9117. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [104].Schmitz D, Li Z, Lo Faro V, Rask-Andersen M, Ameur A, Rafati N, et al. Copy number variations and their effect on the plasma proteome. Genetics. 2023;225(4). doi: 10.1093/genetics/iyad179 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [105].de Los Campos G, Grueneberg A, Funkhouser S, Pérez-Rodríguez P, Samaddar A. Fine mapping and accurate prediction of complex traits using Bayesian Variable Selection models applied to biobank-size data. Eur J Hum Genet. 2023;31(3):313–320. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [106].Broekema RV, Bakker OB, Jonkers IH. A practical view of fine-mapping and gene prioritization in the post-genome-wide association era. Open Biol. 2020;10(1):190221. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [107].Zhang C, Cerveira E, Rens W, Yang F, Lee C. Multicolor fluorescence in situ hybridization (FISH) approaches for simultaneous analysis of the entire human genome. Curr Protoc Hum Genet. 2018;99(1). doi: 10.1002/cphg.70 [DOI] [PubMed] [Google Scholar]
- [108].Gribble SM, Ng BL, Prigmore E, Fitzgerald T, Carter NP. Array painting: a protocol for the rapid analysis of aberrant chromosomes using DNA microarrays. Nat Protoc. 2009;4(12):1722–1736. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [109].Mantere T, Neveling K, Pebrel-Richard C, Benoist M, van der Zande G, Kater-Baats E, et al. Optical genome mapping enables constitutional chromosomal aberration detection. Am J Hum Genet. 2021;108(8):1409–1422. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [110].Schrauwen I, Rajendran Y, Acharya A, Öhman S, Arvio M, Paetau R, et al. Optical genome mapping unveils hidden structural variants in neurodevelopmental disorders. Sci Rep. 2024;14(1). doi: 10.1038/s41598-024-62009-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- [111].Louzada S, Yang F. High-resolution FISH analysis using DNA fibers generated by molecular combing. In: Cancer Cytogenetics and Cytogenomics. Methods in molecular biology (Clifton NJ). Springer US; 2024:185–203. [DOI] [PubMed] [Google Scholar]
- [112].Choi J, Kim S, Kim J, Son HY, Yoo SK, Kim CU, et al. A whole-genome reference panel of 14,393 individuals for East Asian populations accelerates discovery of rare functional variants. Sci Adv. 2023;9(32). doi: 10.1126/sciadv.adg6319 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [113].Lepamets M, Auwerx C, Nõukas M, Claringbould A, Porcu E, Kals M, et al. Omics-informed CNV calls reduce false-positive rates and improve power for CNV-trait associations. HGG Adv. 2022;3(4):100133. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [114].Hujoel MLA, Handsaker RE, Sherman MA, Kamitaki N, Barton AR, Mukamel RE, et al. Protein-altering variants at copy number-variable regions influence diverse human phenotypes. Nat Genet. 2024;56(4):569–578. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [115].Gordeeva V, Sharova E, Babalyan K, Sultanov R, Govorun VM, Arapidi G. Benchmarking germline CNV calling tools from exome sequencing data. Sci Rep. 2021;11(1):14416. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [116].Zhou Z, Wang W, Wang LS, Zhang NR. Integrative DNA copy number detection and genotyping from sequencing and array-based platforms. Bioinformatics. 2018;34(14):2349–2355. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [117].Montanucci L, Lewis-Smith D, Collins RL, Niestroj LM, Parthasarathy S, Xian J, et al. Genome-wide identification and phenotypic characterization of seizure-associated copy number variations in 741,075 individuals. Nat Commun. 2023;14(1):4392. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [118].Owen D, Bracher-Smith M, Kendall KM, Rees E, Einon M, Escott-Price V, et al. Effects of pathogenic CNVs on physical traits in participants of the UK Biobank. BMC Genomics. 2018;19(1):867. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [119].Purcell S, Neale B, Todd-Brown K, Thomas L, Ferreira MAR, Bender D, et al. PLINK: a tool set for whole-genome association and population-based linkage analyses. Am J Hum Genet. 2007;81(3):559–575. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [120].Fawcett KA, Demidov G, Shrine N, Paynton ML, Ossowski S, Sayers I, et al. Exome-wide analysis of copy number variation shows association of the human leukocyte antigen region with asthma in UK Biobank. BMC Med Genomics. 2022;15(1):119. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [121].Liu J, Zhou Y, Liu S, Song X, Yang XZ, Fan Y, et al. The coexistence of copy number variations (CNVs) and single nucleotide polymorphisms (SNPs) at a locus can result in distorted calculations of the significance in associating SNPs to disease. Hum Genet. 2018;137(6–7):553–567. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [122].Wineinger NE, Pajewski NM, Tiwari HK. A Method to Assess Linkage Disequilibrium between CNVs and SNPs Inside Copy Number Variable Regions. Front Genet. 2011;2:17. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [123].Estivill X, Armengol L. Copy number variants and common disorders: filling the gaps and exploring complexity in genome-wide association studies. PLoS Genet. 2007;3(10):1787–1799. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [124].Morales J, Welter D, Bowler EH, Cerezo M, Harris LW, McMahon AC, et al. A standardized framework for representation of ancestry data in genomics studies, with application to the NHGRI-EBI GWAS Catalog. Genome Biol. 2018;19(1):21. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [125].Hayhurst J, Buniello A, Harris L, Mosaku A, Chang C, Gignoux CR, et al. A community driven GWAS summary statistics standard. bioRxiv. Published online July 18, 2022. doi: 10.1101/2022.07.15.500230 [DOI] [Google Scholar]
- [126].Magno R, Maia AT. gwasrapidd: an R package to query, download and wrangle GWAS catalog data. Bioinformatics. 2020;36(2):649–650. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [127].Cao T, Li A, Huang Y. pandasGWAS: a Python package for easy retrieval of GWAS catalog data. BMC Genomics. 2023;24(1):238. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [128].Elsworth B, Lyon M, Alexander T, Liu Y, Matthews P, Hallett J, et al. The MRC IEU OpenGWAS data infrastructure. bioRxiv. Published online August 10, 2020:2020.08.10.244293. doi: 10.1101/2020.08.10.244293 [DOI] [Google Scholar]
- [129].Costanzo MC, Roselli C, Brandes M, Duby M, Hoang Q, Jang D, et al. Cardiovascular Disease Knowledge Portal: A Community Resource for Cardiovascular Disease Research. Circ Genom Precis Med. 2023;16(6):e004181. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [130].Chen Y, Wang AY, Barkley CA, Zhang Y, Zhao X, Gao M, et al. Deciphering the exact breakpoints of structural variations using long sequencing reads with DeBreak. Nat Commun. 2023;14(1):283. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [131].Smolka M, Paulin LF, Grochowski CM, Horner DW, Mahmoud M, Behera S, et al. Detection of mosaic and population-level structural variants with Sniffles2. Nat Biotechnol. Published online January 2, 2024. doi: 10.1038/s41587-023-02024-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- [132].Sedlazeck FJ, Rescheneder P, Smolka M, Fang H, Nattestad M, von Haeseler A, et al. Accurate detection of complex structural variations using single-molecule sequencing. Nat Methods. 2018;15(6):461–468. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [133].Dierckxsens N, Li T, Vermeesch JR, Xie Z. A benchmark of structural variation detection by long reads through a realistic simulated model. Genome Biol. 2021;22(1):342. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [134].Jiang T, Liu Y, Jiang Y, Li J, Gao Y, Cui Z, et al. Long-read-based human genomic structural variation detection with cuteSV. Genome Biol. 2020;21(1):189. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [135].Amarasinghe SL, Su S, Dong X, Zappia L, Ritchie ME, Gouil Q. Opportunities and challenges in long-read sequencing data analysis. Genome Biol. 2020;21(1):30. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [136].De Coster W, Weissensteiner MH, Sedlazeck FJ. Towards population-scale long-read sequencing. Nat Rev Genet. 2021;22(9):572–587. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [137].Gustafson JA, Gibson SB, Damaraju N, Zalusky MP, Hoekzema K, Twesigomwe D, et al. Nanopore sequencing of 1000 Genomes Project samples to build a comprehensive catalog of human genetic variation. medRxiv. Published online March 7, 2024. doi: 10.1101/2024.03.05.24303792 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [138].Schloissnig S, Pani S, Rodriguez-Martin B, Ebler J, Hain C, Tsapalou V, et al. Long-read sequencing and structural variant characterization in 1,019 samples from the 1000 Genomes Project. bioRxiv. Published online April 20, 2024. doi: 10.1101/2024.04.18.590093 [DOI] [Google Scholar]
- [139].Groza C, Schwendinger-Schreck C, Cheung WA, Farrow EG, Thiffault I, Lake J, et al. Pangenome graphs improve the analysis of structural variants in rare genetic diseases. Nat Commun. 2024;15(1):657. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [140].Ebler J, Ebert P, Clarke WE, Rausch T, Audano PA, Houwaart T, et al. Pangenome-based genome inference allows efficient and accurate genotyping across a wide spectrum of variant classes. Nat Genet. 2022;54(4):518–525. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [141].Noyvert B, Erzurumluoglu AM, Drichel D, Omland S, Andlauer TFM, Mueller S, et al. Imputation of structural variants using a multi-ancestry long-read sequencing panel enables identification of disease associations. bioRxiv. Published online December 22, 2023. doi: 10.1101/2023.12.20.23300308 [DOI] [Google Scholar]
- [142].Lambert SA, Gil L, Jupp S, Ritchie SC, Xu Y, Buniello A, et al. The Polygenic Score Catalog as an open database for reproducibility and systematic evaluation. Nat Genet. 2021;53(4):420–425. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [143].Xiang R, Kelemen M, Xu Y, Harris LW, Parkinson H, Inouye M, et al. Recent advances in polygenic scores: translation, equitability, methods and FAIR tools. Genome Med. 2024;16(1):33. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [144].Hao L, Kraft P, Berriz GF, Hynes ED, Koch C, Korategere V Kumar P, et al. Development of a clinical polygenic risk score assay and reporting workflow. Nat Med. 2022;28(5):1006–1013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [145].Lennon NJ, Kottyan LC, Kachulis C, Abul-Husn NS, Arias J, Belbin G, et al. Selection, optimization and validation of ten chronic disease polygenic risk scores for clinical implementation in diverse US populations. Nat Med. 2024;30(2):480–487. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [146].Bergen SE, Ploner A, Howrigan D, CNV Analysis Group and the Schizophrenia Working Group of the Psychiatric Genomics Consortium, O’Donovan MC, Smoller JW, et al. Joint Contributions of Rare Copy Number Variants and Common SNPs to Risk for Schizophrenia. Am J Psychiatry. 2019;176(1):29–35. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [147].Taniguchi S, Ninomiya K, Kushima I, Saito T, Shimasaki A, Sakusabe T, et al. Polygenic risk scores in schizophrenia with clinically significant copy number variants. Psychiatry Clin Neurosci. 2020;74(1):35–39. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [148].Mollon J, Schultz LM, Huguet G, Knowles EEM, Mathias SR, Rodrigue A, et al. Impact of Copy Number Variants and Polygenic Risk Scores on Psychopathology in the UK Biobank. Biol Psychiatry. 2023;94(7):591–600. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [149].Alexander-Bloch A, Huguet G, Schultz LM, Huffnagle N, Jacquemont S, Seidlitz J, et al. Copy Number Variant Risk Scores Associated With Cognition, Psychopathology, and Brain Structure in Youths in the Philadelphia Neurodevelopmental Cohort. JAMA Psychiatry. 2022;79(7):699–709. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [150].Saarentaus EC, Havulinna AS, Mars N, Ahola-Olli A, Kiiskinen TTJ, Partanen J, et al. Polygenic burden has broader impact on health, cognition, and socioeconomic outcomes than most rare and high-risk copy number variants. Mol Psychiatry. 2021;26(9):4884–4895. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [151].Kachuri L, Chatterjee N, Hirbo J, Schaid DJ, Martin I, Kullo IJ, et al. Principles and methods for transferring polygenic risk scores across global populations. Nat Rev Genet. 2024;25(1):8–25. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [152].Hu S, Ferreira LAF, Shi S, Hellenthal G, Marchini J, Lawson DJ, et al. Leveraging fine-scale population structure reveals conservation in genetic effect sizes between human populations across a range of human phenotypes. bioRxiv. Published online August 9, 2023:2023.08.08.552281. doi: 10.1101/2023.08.08.552281 [DOI] [Google Scholar]
- [153].Hou K, Ding Y, Xu Z, Wu Y, Bhattacharya A, Mester R, et al. Causal effects on complex traits are similar for common variants across segments of different continental ancestries within admixed individuals. Nat Genet. 2023;55(4):549–558. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [154].Heyne HO, Karjalainen J, Karczewski KJ, Lemmelä SM, Zhou W, FinnGen, et al. Mono- and biallelic variant effects on disease at biobank scale. Nature. 2023;613(7944):519–525. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [155].Song P, Gupta A, Goon IY, Hasan M, Mahmood S, Pradeepa R, et al. Data Resource Profile: Understanding the patterns and determinants of health in South Asians-the South Asia Biobank. Int J Epidemiol. 2021;50(3):717–718e. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [156].Browning SR, Browning BL. Rapid and accurate haplotype phasing and missing-data inference for whole-genome association studies by use of localized haplotype clustering. Am J Hum Genet. 2007;81(5):1084–1097. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [157].Loh PR, Danecek P, Palamara PF, Fuchsberger C, A Reshef Y, K Finucane H, et al. Reference-based phasing using the Haplotype Reference Consortium panel. Nat Genet. 2016;48(11):1443–1448. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [158].Delaneau O, Zagury JF, Robinson MR, Marchini JL, Dermitzakis ET. Accurate, scalable and integrative haplotype estimation. Nat Commun. 2019;10(1). doi: 10.1038/s41467-019-13225-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- [159].Hofmeister RJ, Ribeiro DM, Rubinacci S, Delaneau O. Accurate rare variant phasing of whole-genome and whole-exome sequencing data in the UK Biobank. Nat Genet. 2023;55(7):1243–1249. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [160].Browning BL, Browning SR. Statistical phasing of 150,119 sequenced genomes in the UK Biobank. Am J Hum Genet. 2023;110(1):161–165. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [161].Lassen FH, Venkatesh SS, Baya N, Zhou W, Bloemendal A, Neale BM, et al. Exome-wide evidence of compound heterozygous effects across common phenotypes in the UK Biobank. medRxiv. Published online July 3, 2023. doi: 10.1101/2023.06.29.23291992 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [162].Mountjoy E, Schmidt EM, Carmona M, Schwartzentruber J, Peat G, Miranda A, et al. An open approach to systematically prioritize causal variants and genes at all published human GWAS trait-associated loci. Nat Genet. 2021;53(11):1527–1533. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [163].Abdellaoui A, Yengo L, Verweij KJH, Visscher PM. 15 years of GWAS discovery: Realizing the promise. Am J Hum Genet. 2023;110(2):179–194. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [164].Namba S, Konuma T, Wu KH, Zhou W, Global Biobank Meta-analysis Initiative, Okada Y. A practical guideline of genomics-driven drug discovery in the era of global biobank meta-analysis. Cell Genom. 2022;2(10):100190. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [165].Klein RJ, Zeiss C, Chew EY, Tsai JY, Sackler RS, Haynes C, et al. Complement factor H polymorphism in age-related macular degeneration. Science. 2005;308(5720):385–389. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [166].Arruda AL, Morris AP, Zeggini E. Advancing equity in human genomics through tissue-specific multi-ancestry molecular data. Cell Genom. 2024;4(2):100485. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.





