Skip to main content

This is a preprint.

It has not yet been peer reviewed by a journal.

The National Library of Medicine is running a pilot to include preprints that result from research funded by NIH in PMC and PubMed.

Research Square logoLink to Research Square
[Preprint]. 2025 Oct 31:rs.3.rs-6238738. [Version 1] doi: 10.21203/rs.3.rs-6238738/v1

Genetics of Major Depressive Disorder in a Homogeneous Population with Uniform Phenotyping

Floris Huider 1,2,3, Yuri Milaneschi 2,4,5, René Pool 1,2, Bernardo de APC Maciel 3, Scott D Gordon 6, M Liset Rietman 7, Almar AL Kok 2,4,8, Tessel E Galesloot 9, Brittany L Mitchell 6, Leen M ‘t Hart 2,8,10,11, Femke Rutters 2,8, Marieke T Blom 2,12, Didi Rhebergen 2,4,5, Marjolein Visser 2,13, Ingeborg A Brouwer 2,13, Edith Feskens 14, Catharina A Hartman 15, Albertine J Oldehinkel 15, Mariska Bot 2,4,5, Eco JC de Geus 2,3, Lambertus A Kiemeney 9,16, Martijn Huisman 2,8,17, H Susan J Picavet 7, WM Monique Verschuren 7,18, Nicholas G Martin 6, Conor Dolan 2,3, Hanna M van Loo 15, Brenda WJH Penninx 2,4,5, Jouke-Jan Hottenga 2,3,*, Dorret I Boomsma 1,2,*
PMCID: PMC12636711  PMID: 41282258

Abstract

Harmonized phenotyping and diverse population-specific studies are crucial for advancing gene discovery in psychiatric genetics. We conducted a genome-wide association study (GWAS) of DSM-defined lifetime Major Depressive Disorder (MDD) in 64,941 participants (25.7% cases) from the Dutch BIObanks Netherlands Internet Collaboration (BIONIC) consortium. SNP-based heritability was estimated at 13.4%, exceeding recent global meta-analyses, with a high genetic correlation (r = 0.89) to the latest major depression GWAS by the Psychiatric Genetics Consortium (PGC-MD). We identified a novel genome-wide significant locus in PALMD (P = 3.26 × 10−8), that was confirmed by GWAS-by-subtraction. Polygenic scores (PGSs) based on BIONIC predicted MDD in UKBiobank, and PGSs from PGC-MD predicted into BIONIC, with within-family analyses indicating minimal confounding. Genetic causal inference revealed associations with over 30 phenotypes. Twin concordance for MDD increased with polygenic burden, reinforcing its genetic architecture. This study emphasizes the power of harmonized phenotyping and regional biobanks in uncovering the genetic architecture of MDD, highlighting the value of population-specific studies for improving risk prediction and advancing psychiatric genetics.

Keywords: GWAS, MDD, uniform population, DSM criteria, PGS


Major depressive disorder (MDD) is among the most common psychiatric disorders with a lifetime prevalence that can reach up to 19%, and is a leading cause of disability worldwide1,2. Episodes are frequently characterized by severe dysfunction and pervasiveness3,4. MDD is a genetically complex trait with a twin-based heritability of 37%5 and a polygenic architecture, characterized by the additive effect of many variants with small effects. In addition, MDD’s intrinsic heterogeneity, heterogeneity in depression diagnoses and assessment, and the relatively modest heritability compared to other psychiatric conditions, have all been suggested as contributing to low statistical power for genome-wide association studies6,7. One successful approach to combating several of these problems is through large international collaborative efforts that bring together dozens of cohorts and millions of individuals8–11.

At the same time, a discussion has emerged regarding a trade-off in sample size and phenotypic definition in genetic studies. Larger sample sizes have led to finding large numbers of genetic variants associated with broad concepts of depression. Growing sample sizes often come at the expense of uniform phenotyping; combining samples across multiple studies tends to yield a lower adherence to clinical diagnostic criteria and phenotypic uniformity. The use of self-reported diagnosis or treatment for depression, or depressed mood derived from a cut-off applied to a self-report symptom questionnaire in research, coined ‘minimal phenotyping’, may not align with MDD according to diagnostic criteria or represent associated but different genetic vulnerabilities12–15. In a comparison of depression phenotypes derived from the UK Biobank, Cai et al. (2020) found that minimal phenotyping definitions are epidemiologically distinct from strictly defined MDD, have lower SNP-based heritability (SNP-h2) estimates (11% for self-reported depression vs. 26% for MDD following diagnostic criteria), and that GWAS hits from minimal phenotyping may not be specific to MDD, and instead apply to a broader range of psychopathology and personality measures15. Consequently, follow-up characterization of GWAS loci from minimal phenotyping studies may not inform on biology specific to MDD.

The inclusion of data based on heterogeneous phenotypic assessment associates with a decay of SNP-h2 estimates16,17. This has led to a call for approaches with uniform phenotyping, for example by adherence to clinical diagnostic criteria. In addition to phenotypic heterogeneity, meta-analyses can introduce genetic heterogeneity. Nearly all current studies tend to limit participation to participants of European ancestry, but even here there can be considerable population stratification and admixture affecting heterogeneity and GWAS results18, even after adjustments through ancestry-informative principal components. Tropf et al.19 find that heterogeneity across populations is one of the contributing factors to missing heritability, the common difference between GWAS-derived SNP-h2 estimates and the upper limits set by twin-based approaches.

To address these challenges, we initiated the BIObanks Netherlands Internet Collaboration (BIONIC)20. This nation-wide collaborative project combines genotype data from 16 Dutch cohorts, and focuses on a phenotype definition of MDD according to DSM-5 criteria. The Netherlands has a population of 18 million people, with ~15.5 million of Dutch ancestry, and a population density of ~535 people per square kilometer (1,390 people/sq mi). In such single populations, population stratification is smaller, although not entirely absent21,22. We enriched the existing population cohorts and clinical cohorts in the Netherlands with an infrastructure for nation-wide depression phenotyping so that all MDD cases were defined based on DSM-5 (Diagnostic and Statistical Manual of Mental Disorders)-criteria3,23. All cohorts agreed to share genotype and phenotype data in a central location for mega-analysis, where QC and genotype imputation were conducted simultaneously and harmonized across all arrays and participants.

With the BIONIC dataset we ask the question if we can find evidence for a contribution of single nucleotide polymorphisms (SNP) to MDD as evidenced by genome-wide significant associations, significant SNP heritability and polygenic prediction. We report results from the mega-analysis of lifetime MDD case-control status and genome-wide SNP data in 64,941 individuals (Ncase = 16,655, Ncontrol = 48,286). We leverage clinical diagnostic criteria in a relatively homogeneous sample to identify genetic variant associations and consider the benefits of such an approach for capturing genetic signal, comparing our results to the largest GWAS of major depression to date (by the Psychiatric Genetics Consortium MDD workgroup)10. We benchmark these results with a genome-wide association mega-analysis of the well-characterized trait of height, for which a saturated map of common genetic variants has been generated24. We apply GWAS-by-subtraction, in which we searched for a uniquely Dutch MDD signal by ‘subtracting’ the largest global major depression GWAS effort from our GWAS results. We extend the GWA mega-analysis of MDD with a series of complementary analyses. These include fine-mapping approaches, such as gene-based tests and tissue expression analyses, as well as broader investigations into phenome-wide genetic correlations and causal relationships. We evaluate polygenic risk scores (PGS) through in- and out-of-sample predictions, assess potential confounding via within- and between-family tests, and examine twin concordance across PGS deciles.

Results

Genome-wide association mega-analysis of lifetime MDD

We ran a GWAS mega-analysis of lifetime MDD including 64,941 individuals (16,655 lifetime MDD cases and 48,286 screened controls). The majority of the sample was female (60.8% in the full sample) and average age was 50.7 (SD 16.0)(eTable 1). Figure 1 shows the Manhattan Plot of genetic variants and the significance of their association with lifetime MDD (QQplot in Supplementary Figure 2). The rs3818852 SNP on chromosome 1, an intronic variant of the PALMD gene, reached significance at the genome-wide-adjusted alpha level (OR = 0.93, p = 3.26 × 10−8)(regional association plot in Supplementary Figure 3). PALMD encodes the palmdelphin protein involved in cellular processes including cytoskeletal organization and cell shape modulation. There have been no previous GWAS hits for rs3818852 registered in the GWAS catalogue, indicating a potentially novel cross-domain phenotypic association. A second signal, rs4131791 on chromosome 18, showed suggestive significance (OR = 1.07, p = 2.17 × 10−7), and maps to the LINC03035 gene. Table 1 lists the 10 most significant independent SNPs, with the summary statistics for the top 1000 loci available in the Supplementary Material.

Figure 1. Manhattan plot of the GWAS mega-analysis of major depressive disorder.

Figure 1.

Manhattan Plot of the genome-wide association mega-analysis of lifetime MDD in BIONIC. The x-axis indicates chromosomal position. Y-axis denotes −log10 P-value and indicates the strength of association between the SNP and outcome phenotype. The red line indicates Bonferroni-corrected genome-wide significance; P < 5×10−8. The blue line indicates suggestive singificance; P < 1×10−5.

Table 1.

Top 10 independent SNPs in the genome-wide association mega-analysis of lifetime MDD. Positions and rsID are in build GRCh17.

Chr SNP BP A1/A2 F(A1) OR s.e. P Nearest Gene
1 rs3818852 100160177 A/T 0.38 0.93 0.01 3.26e-8 PALMD
18 rs4131791 52747871 C/T 0.58 1.07 0.02 2.17e-7 LINC3035
18 rs12327102 50787839 G/C 0.46 0.93 0.01 6.96e-6 DCC
16 rs62037142 57477652 C/T 0.98 0.76 0.04 6.97e-6 CIAPIN1
1 rs150207136 119607219 G/A 0.98 0.79 0.04 7.53e-6 WARS2
1 rs1970168 232761116 G/A 0.51 1.07 0.02 9.13e-6 SIPA1L2
4 rs10000918 186268576 A/G 0.74 1.08 0.02 1.04e-6 SNX25
5 rs33935044 21145325 A/T 0.98 1.32 0.08 1.30e-6 L0C102723561
4 rs138362425 75121348 C/A 0.98 1.35 0.09 1.37e-6 MTHFD2L
13 rs7491779 82485534 A/G 0.68 0.93 0.01 2.21e-6 None

Chromosome (chr), single-nucleotide polymorphism (SNP), base pair (BP), alleles (A1/A2; effect allele/non effect allele), effect allele frequency (F(A1)), effect allele odds ratio (OR), standard error (s.e.).

To study the feasibility of our analysis of MDD in this relatively small sample, we conducted a genome-wide association mega-analysis of height in 52,893 individuals with information on this phenotype from the BIONIC dataset (Figure 2, Supplementary Figure 4). Despite the modest sample size, we were able to identify 94 independent significant SNPs associated with height at the genome-wide significance level.

Figure 2. Manhattan Plot of GWAS mega-analysis of height.

Figure 2.

Manhattan Plot of the genome-wide association mega-analysis of height in BIONIC. The x-axis indicates chromosomal position. Y-axis denotes −log10 P-value and indicates the strength of association between the SNP and outcome phenotype. The red line indicates Bonferroni-corrected genome-wide significance; P < 5 × 10−8. The blue line indicates suggestive singificance; P < 5 × 10−5.

SNP-based heritability and genetic correlation with PGC-MD

We ran Linkage Disequilibrium score regression (LDSC)25 on the GWA mega-analysis results of lifetime MDD, hereafter called the BIONIC MDD GWAS, to compute SNP-based heritability (SNP-h2). SNP-h2 for lifetime MDD in the effective sample size (Neff = 63,742.64; Ncases = 16,655) was estimated to be 0.134 (se = 0.021) on the liability scale, suggesting 13.4% of phenotypic variance is jointly captured by the measured SNPs. The LD intercept estimate was 1.018 (se = 0.007), suggesting negligible inflation in GWAS test statistics due to sources other than polygenicity. We found a higher LDSC SNP-h2 point estimate than that in the largest GWAS to date, the Psychiatric Genetic Consortium meta-analysis of major depression (PGC-MD; SNP-h2 = 8.4%, se = 0.07%)10. This is consistent with evidence from previous studies showing that uniform phenotyping and clinical diagnostic criteria in a relatively homogeneous population benefit the capture of genetic signal of MDD15,17. We computed the genetic correlation between the BIONIC MDD GWAS and PGC-MD (European samples including 23andMe, Neff = 1,523,738) in LDSC, finding a genetic correlation of rG = 0.869 (se = 0.048).

Height LDSC SNP-h2 was estimated to be 0.379 (se = 0.026; Neff = 49,036), or 37.9% of phenotypic variance. The genetic correlation with the largest GWAS of height to date, Yengo et al.24, was rG = 0.980 (se = 0.016).

GWAS-by-subtraction

We investigated if a uniquely Dutch MDD signal might be identified through GWAS-by-subtraction26 (Figure 3A). The BIONIC MDD GWAS and PGC-MD GWAS were regressed on a latent factor representing genetic variance in depression aspects that are shared across the two studies, followed by the BIONIC MDD GWAS was further regressed on a latent factor representing genetic variance unique to the BIONIC MDD GWAS. This served as input for a GWAS of uniquely Dutch MDD (Figure 3B). In this GWAS, no SNPs reached significance at the genome-wide adjusted alpha level. However, the rs3818852 SNP on chromosome 1 that was found in the BIONIC MDD GWAS remained among the top signals (OR = 0.83, P-value = 1.38 × 10−6). LDSC SNP-h2 for this model was 0.036 (se = 0.014).

Figure 3. GWAS results of uniquely Dutch MDD (GWAS-by-subtraction).

Figure 3.

A) GWAS-by-subtraction model (Demange et al. (2021)). B) Manhattan Plot of the genome-wide association analysis of uniquely Dutch MDD, the latent factor resulting from subtracting the PGC-MD GWAS from the BIONIC MDD GWAS. The x-axis indicates chromosomal position. Y-axis denotes −log10 P-value and indicates the strength of association between the single-nucleotide polymorphism and the outcome phenotype. The red line indicates Bonferroni-corrected genome-wide significance; P < 5 × 10−8. The blue line indicates suggestive singificance; P < 5 × 10−5.

Polygenic score prediction

We tested predictive utility of the major depression polygenic score (MD PGS) based on PGC-MD in the BIONIC MDD GWAS sample (unrelated N = 44,144; 27% cases). The MD PGS was significantly associated with lifetime MDD status in BIONIC (OR = 1.054, p < 0.001), and explained 3.94% of variance on the liability scale. Figure 4A displays the PGS divided into deciles and their effect size.

Figure 4. MD PGS prediction of lifetime MDD and twins by MDD PGS decile.

Figure 4.

A) Predictive PGS performance of MD PGS based on PGC-MD of lifetime MDD in BIONIC (unrelated N = 44,144). The x-axis denotes the MD PGS decile. The y-axis denotes the odds ratio of each decile from the generalized estimating equations model where lifetime MDD status in the BIONIC sample was predicted by PGS decile. Odds ratios were estimated relative to decile 1. 95% confidence intervals are given. B) Number of twins concordant for lifetime MDD status per MD PGS decile (N = 266 twins).

We then tested the performance of a PGS based on the BIONIC MDD GWAS for two phenotypic definitions of depression in the UKBiobank: strict (ICD code for depression) and broad (‘minimal phenotypic’). The BIONIC MDD PGS was significantly associated with both strict (N = 23,755, OR = 1.163, Z = 17.1) and broad definitions of depression (N = 132,122, OR = 1.107, Z = 29.5), explaining 0.39% and 0.29% of phenotypic variance on the liability scale, respectively. These results highlight the transferability of the BIONIC MDD PGS, showing robust out-of-sample prediction for different phenotypic definitions.

Within- and between-family PGS prediction

Genotype-phenotype associations and the PGSs derived from them may arise through confounding effects, including population stratification, assortative mating and gene-by-environment correlation27. To evaluate whether PGS prediction of lifetime MDD may be subject to confounding, we evaluated MD PGS performance based on PGC-MD in a subset of dizygotic twins from the BIONIC MDD GWAS sample. In an unmatched binary logistic regression model of 2282 dizygotic twins (1141 complete pairs), the major depression PGS significantly predicted lifetime MDD (OR = 1.434, p = 9.52 × 10−9). We then applied a random effects logistic model to predict lifetime MDD with between- and within-family MD PGS, and family as a random variable. Both the between- and within-family MD PGS were significantly associated with lifetime MDD (OR = 1.409, p = 9.04 × 10−5; OR = 1.948, p = 1.09 × 10−6, respectively). A chi-square test of the difference in between- and within-family MD PGS effects gave χ2 (1, N = 2282) = 3.983, p = 0.046. The between- and within-family PGS effects were significantly different, albeit barely so, suggesting that if there is confounding, it is minimal. Sensitivity analyses showed that the confounding may in part be due to sex-specific effects, as the difference for the between- and within-family MD PGS effects was strongest in sex-discordant twins (p = 0.025), but absent in sex-concordant twin pairs (p = 0.435)(Supplementary Information).

Twin Concordance and PGS deciles

We computed MD PGS deciles in N = 2963 twin pairs in BIONIC. We distinguished three subsamples based on lifetime MDD concordance in the twin pairs: concordantly affected (N = 133 pairs), concordantly unaffected (N = 2257), and discordantly affected (N = 573). We expected concordantly affected twin pairs to be overrepresented in the high PGS deciles, and concordantly unaffected twin pairs to be overrepresented in the low PGS deciles. Figure 4B shows the distribution of concordantly affected twins across MD PGS deciles (concordantly unaffected and discordantly affected are shown in Supplementary Figure 5). Indeed, we found the number of concordantly affected twin pairs to increase in the higher MD PGS deciles, and the number of concordantly unaffected twin pairs to decrease in the higher MD PGS deciles. A similar pattern was observed in an independent replication sample from the Australian Genetics of Depression Study (Supplementary Information).

We ran an ordinal logistic regression model with twin MDD concordance as outcome to test whether the mean MD PGS predicted twin MDD concordance status in the N = 2963 twins in BIONIC. We found each increase of one standard deviation in mean MD PGS to be significantly associated with a 1.3-fold greater risk for concordant MDD status in twins (OR = 1.289; 95% CI = 1.174–1.404).

Genetic causal associations

We computed genetic correlations between BIONIC lifetime MDD and 1461 complex traits and diseases, finding 388 significant genetic correlations (False Discovery Rate (FDR) < 5%) with well-established correlates including self-harm, substance use, anxiety, loneliness, neuroticism, irritable bowel syndrome, and chronic pain (Supplementary Material). With the latent causal variable (LCV) method28,29, we identified 33 traits with a putative causal genetic association with MDD at FDR < 5% (|GCP| > 0.60), and 4 traits with limited partial genetic causality (|GCP| ≤ 0.60)(Table 2; Figure 5). We identified 34 traits as putative risk factors for MDD, and 3 traits as putative outcomes of MDD. The putative causal genetic associations with MDD include a diverse range of traits, including well-known risk factors such as hypersomnia, stroke, and trauma. Several significant putative causal traits were related to one’s work environment, including working with paints and chemicals and workplace with very cold temperature, which might also reflect differences in socio-economic status. Others included lifestyle traits such as vitamin D levels and vegetable consumption. Finally, cardiometabolic traits were found, including stroke, ECG load, and HDL cholesterol.

Table 2.

Traits with an inferred putative causal relationship with MDD.

Phenotypes that cause MDD GCP GCP se GCP P rG rG se rG P
Sleeping too much −0.88 0.11 8.64e-16 0.65 0.16 4.01e-05
Workplace often full of chemical or other fumes −0.86 0.13 2.32e-11 0.73 0.19 7.56e-05
Mouth ulcers −0.83 0.13 5.31e-11 0.28 0.07 4.74e-05
ECG load −0.85 0.13 2.03e-10 −0.68 0.18 1.68e-04
Had other major operations −0.94 0.15 6.11e-10 0.41 0.10 7.53e-05
Triglycerides −0.77 0.14 6.77e-08 0.20 0.07 6.87e-03
Self-reported cervical spondylosis −0.80 0.15 1.86e-07 0.44 0.15 2.38e-03
Sometimes worked with paints, thinners or glues −0.80 0.16 3.31e-07 0.54 0.12 1.18e-05
Mouth-teeth dental problems: Mouth ulcers −0.49 0.10 1.10e-06 0.28 0.08 2.49e-04
Treatment/medication code: cod liver oil capsule −0.79 0.17 4.97e-06 −0.66 0.22 3.08e-03
Ever had period of mania / excitability −0.79 0.18 6.68e-06 0.60 0.12 5.69e-07
Vitamin and mineral supplements: Vitamin B −0.76 0.18 1.56e-05 0.36 0.13 6.57e-03
Treatment/medication code: rosiglitazone −0.90 0.21 1.89e-05 0.35 0.14 1.20e-02
Vegetable consumption −0.60 0.14 2.21e-05 −0.43 0.17 1.30e-02
Vitamin D −0.75 0.18 4.35e-05 −0.18 0.06 1.97e-03
Workplace often very cold −0.75 0.18 4.45e-05 0.51 0.15 4.61e-04
HDL cholesterol −0.67 0.18 1.73e-04 −0.12 0.04 3.22e-03
Workplace often very hot −0.55 0.15 2.43e-04 0.50 0.15 5.84e-04
Did your sleep change? −0.75 0.21 3.44e-04 0.69 0.27 1.13e-02
Difficulty concentrating during worst depression −0.75 0.21 3.47e-04 0.75 0.14 1.60e-07
Illnesses of mother: None of the above (group 2) −0.72 0.20 3.75e-04 −0.41 0.12 1.12e-03
Workplace had a lot of cigarette smoke from other people smoking: Often −0.74 0.21 4.57e-04 0.65 0.16 6.64e-05
More creative or having more ideas than usual −0.59 0.17 5.10e-04 0.51 0.17 2.24e-03
Probable Recurrent major depression (moderate) −0.75 0.22 6.30e-04 0.76 0.14 2.40e-08
Diabetes diagnosed by doctor −0.69 0.20 7.56e-04 0.24 0.07 3.16e-04
Self-reported stroke −0.72 0.22 8.71e-04 0.64 0.20 1.35e-03
Treatment/medication code: ranitidine −0.70 0.21 9.20e-04 0.46 0.14 8.28e-04
Treatment/medication code: diazepam −0.72 0.22 1.03e-03 0.63 0.24 1.00e-02
Self-reported diabetes −0.68 0.21 1.23e-03 0.23 0.07 1.38e-03
Stroke (European biobanks) −0.84 0.26 1.47e-03 0.30 0.11 8.61e-03
Diaphragmatic hernia −0.68 0.22 2.30e-03 0.49 0.14 3.52e-04
Age at menopause (last menstrual period) −0.65 0.22 2.99e-03 −0.22 0.06 7.67e-04
Illnesses of mother: Chronic bronchitis/emphysema −0.67 0.23 3.42e-03 0.41 0.11 2.23e-04
Stroke (global biobanks) −0.64 0.23 4.74e-03 0.27 0.09 2.94e-03
Phenotypes caused by MDD
Breathing problems improved/stopped away from workplace or on holiday 0.85 0.12 3.55e-13 0.69 0.17 4.71e-05
Been in a serious accident believed to be life-threatening 0.78 0.16 1.08e-06 0.68 0.17 4.94e-05
Back pain for 3+ months 0.53 0.12 5.32e-06 0.31 0.11 6.39e-03

This table displays 37 traits with a significant (FDR < 5%) genetic causal proportion with MDD. P-values are before FDR correction, P-values after FDR correction are listed in Supplementary Material. GCP: genetic causal proportion; se: standard error; rG: genetic correlation.

Gene-level analyses & tissue enrichment

We ran a MAGMA30 gene-based analysis on the BIONIC MDD GWAS results, mapping input SNPs to 19,210 protein coding genes. Two genes were significantly associated with MDD at the genome-wide significance level: the PALMD gene on chromosome 1 (Z-score = 4.725, P = 1.15 × 10−6) and the PIACIN1 gene on chromosome 16 (Z-score = 4.627, P = 1.85 × 10−6)(Figure 5). We ran MAGMA gene-set analysis for pathway-specific enrichment for gene sets C2 and C5 from MsigDB v7.0. We found no significant gene-sets at the Bonferroni-corrected significance threshold (3.23 × 10−6). The most significant gene sets were involved in regulation of T-cell migration (Beta = 1.61, P = 6.75 × 10−6) and vitamin binding (Beta = 0.31, P = 7.75 × 10−6). The full summary statistics of the gene-based test and gene set test are available in the Supplementary Material. In the MAGMA tissue expression analysis, in which we compared SNP associations from the BIONIC MDD GWAS with gene expression levels from the GTEx v8 database for 53 tissue types, none of the investigated tissues showed significant enrichment after multiple testing correction (Supplementary Figure 8).

Gene prioritization

We applied FLAMES31 to find the most likely effector gene for the genome-wide significant locus in the BIONIC MDD GWAS (lead variant: rs3818852 on chromosome 1). There were 23 SNPs with an LD > 0.60 implicated as potential alternative causal SNPs, which were used for credible set generation in FINEMAP. The PALMD gene was assigned the highest FLAMES score (0.436), suggesting it is the most likely effector gene for the genetic association between rs3818852 and MDD.

Sensitivity analyses

As an alternative to the GWAS mega-analysis approach, we conducted a genome-wide association fixed effect meta-analysis combining the seven array group GWAS results of lifetime MDD case-control status. Results were similar to those of the main model (rG = 1, Supplementary Information). Further, we ran a GWAS mega-analysis of MDD in which the control group included individuals with evidence for psychopathology other than MDD to explore whether screening of controls led to biased SNP-h2 and rG estimates. Total sample size increased to N = 67,440 (16,655 cases). The less stringent screening of controls had very little effect on the outcome (rG = 0.99 with main model, Supplementary Information). SNP-h2 in this model was 0.118 (0.015), which is comparable to the first model (SNP-h2 = 0.134). This suggests the main model SNP-h2 estimate was not inflated by the screening of controls.

Discussion

We ran a genome-wide association (GWAS) mega-analysis of lifetime major depressive disorder (MDD) in the Dutch BIONIC dataset, including MDD and genotype data on 64,941 participants. We created this resource to help unravel the complex genetic architecture of MDD by focusing on uniform phenotyping based on DSM-5 criteria in a genetically homogenous sample.

We found that common SNPs explained 13.4% of variance in lifetime MDD on the liability scale. This is a notable estimate compared to recent large-scale meta-analyses using a combination of broad and diagnostic definitions (8.4%)9–11. Though the proportion of variance explained by common SNPs can be subject to variation due to many aspects of study design, a notable contributor is heterogeneity in population, study design, and phenotype. Tropf et al.19 found that heterogeneity across populations was one of the sources for the difference between twin-based heritability and SNP-based heritability, the so-called missing heritability, and Cai et al.15 found that restricting samples to those meeting clinical criteria benefitted explained genetic variance. Our findings corroborate empirical evidence that both sample uniformity and phenotypic precision can enhance genetic signal detection17,19, joining other recent efforts that focus on this32,33. The success of our efforts reinforces the valuable contributions regional biobanks can have and suggests potential applications in studying other complex traits.

We report one genome-wide significant finding for MDD: rs3818852, an intronic variant of the PALMD gene on chromosome 1. To our knowledge, PALMD has not been implicated in MDD before. While its exact biological function is largely unknown, PALMD encodes the palmdelphin protein, which may be involved in cellular processes including cytoskeletal organization and cell shape modulation and is expressed in dendritic spines, membrane, and cytoplasm (NCBI, checked Feb 2025). Though rs3818852 has no previous associations, PALMD has been associated with complex traits from different phenotypic domains such as aortic valve calcification and the vaginal microbiome34–36. The A allele of the rs3818852 variant on chromosome 1 (PALMD) has a distinctly lower minor allele frequency (MAF) in BIONIC (0.379) and in the Genome of the Netherlands reference panel (0.374)22 compared to the global population (0.486, TOPMED, https://www.ncbi.nlm.nih.gov/, checked Feb 2025). Difference in allele frequencies has been proposed as the leading reason for the success of the CONVERGE study in detecting two of the first GWAS loci for MDD37, neither of which was significantly associated in GWAS on European samples9,11. The PALMD signal was further reiterated by the GWAS-by-subtraction analysis, in which the signal remained after subtracting the global depression signal captured by the international PGC-MD effort. The gene-based analysis identified two significantly associated genes with MDD, corroborating the PALMD and finding the CIAPIN1 gene. and found a significant association with the CIAPIN1 gene. CIAPIN1 is a protein coding cytokine-induced inhibitor of apoptosis involved in mitochondrial function, methyltransferase activity and iron-sulfur cluster binding. It has been found to inhibit hippocampal neuronal cell damage through MAPK and apoptotic signaling pathways38. To our knowledge, a role in MDD has not been established before.

Polygenic scores (PGSs) based on the BIONIC MDD GWAS significantly predicted distinct definitions of lifetime MDD in UKbiobank. The explained variance of the BIONIC GWAS PGS was higher for ICD10 diagnosis-based criteria, suggesting PGS based on strict phenotyping approaches may have higher transferability to other settings that use DSM criteria, such as the clinic. These results highlight that our PGS generalizes to other populations and phenotypic definitions. We also leveraged some of the unique features of the BIONIC dataset to explore polygenic score prediction of MDD in twin pairs. PGSs are increasingly becoming predictors and phenotypic proxies for phenotypes in genetic epidemiological designs39. GWAS and PGSs derived from these are naïve to the way in which genotype-phenotype associations arise, meaning they can reflect biological mechanisms but also may capture environmental factors, and thus may be subject to confounding27,40. By leveraging features of twin- and within-family designs, we examined confounding in PGS prediction of MDD, and found that if there was confounding due to population stratification, assortative mating and gene-environment correlation, it was minimal, corroborating findings from previous work where estimates of assortative mating in MDD were low20,41,42. We also explored the association between twin concordance and polygenic burden following a recent example in schizophrenia and bipolar disorder43. We show that twin concordance for MDD is partly a function of polygenic burden, with concordantly affected twins having a higher PGS than pairs where one twin was affected, who in turn had a higher PGS than concordantly unaffected twins.

Estimates of genetic correlations and genetic causal associations with other traits identified several putative risk factors and outcomes of MDD. We found hypersomnia to increase the risk of MDD, together with workplace-related traits such as a cold workplace, working with fumes and paint thinners, having had major operations, and cardiovascular indicators like stroke and triglyceride levels. We found other traits to reduce the likelihood of MDD, including vitamin D levels, HDL cholesterol levels, vegetable consumption and cod liver oil capsule consumption. MDD predicted that breathing problems improved when away from work, suggesting MDD may predict that breathing problems were caused by dissatisfaction or stress at work rather than an innate condition. MDD also predicted back pain to last over 3 months, and having been in a serious accident. In an earlier study of genetic causal associations with broad depression29, quite similar results were found, such as putative causal effects of hypersomnia, vitamin D levels, triglyceride levels, and various antidepressants. Conversely, the study found no putative outcomes of broad depression. It has been shown that minimally-defined depression includes nonspecific liability to psychopathology and that this affects genetic correlation estimates15. It is possible that this also affects genetic causal estimation estimates, and that considering clinically-defined MDD may reveal unique putative causal targets and outcomes.

There were several strengths and limitations. The sample size of this study was small compared to the standard of GWAS being published nowadays. Further, the results in this study were achieved by considering two factors, phenotypic homogeneity and population homogeneity. We cannot distinguish to which to attribute the increase in genetic signal, as indicated by the increase in SNP-h2. Finally, it cannot go without mention that the lack of diversity in the sample might hinder generalization of the findings to culturally-dependent phenotypic definitions and other ancestry-diverse populations. This study has been successful in recruiting cohorts that have not before been included in a GWAS of MDD, which we approached because they already had genotyping in the Dutch population available. We are not the first to show that a homogeneous and strict phenotyping definition benefits MDD genetic variant identification37. It is clear that such a strategy still holds promise for finding genome-wide significant associations with MDD in populations.

With this effort we add to the growing evidence base on the genetic architecture of MDD. We show that an effort with uniform phenotyping in homogeneous populations can aid genetic signal detection for MDD. As the genetic signal captured by broad phenotyping strategies plateaus, it may be time to invest in clinically-relevant phenotyping approaches, for which we provide a proof of concept.

Online Methods

Sample and phenotype descriptions

This genome-wide association (GWAS) mega-analysis combines major depressive disorder (MDD) and genotype data from sixteen Dutch cohorts in the BIONIC (BIObanks Netherlands Internet Collaboration) project. Descriptions of the sixteen cohorts can be found in the Supplementary Information. The majority of phenotype data were collected through the Lifetime Depression Assessment of Survey (LIDAS)23, developed to assess DSM-5 MDD. Other data were derived from clinical interviews and questionnaires (see eMethods for a detailed description). MDD cases were uniformly defined according to DSM-5 diagnostic criteria as ever having had a period of at least two weeks with five or more depression symptoms causing dysfunction in life, of which at least one a cardinal symptom44. MDD controls were defined as having no such period, and – when information on other psychopathology was available – having no other psychopathology (eMethods). The final analysis set included 64,941 European-ancestry individuals with non-missing MDD, covariate (sex, age) and genotype data, 16,655 lifetime MDD cases and 48,286 screened controls. Height data were available in 52,893 individuals.

Genotyping, quality control, and imputation

Details on the quality control (QC) procedure have been described previously20. In brief, the sixteen cohorts had collected genotype data on seven different arrays across multiple batches, resulting in 19 datasets. This included some smaller datasets that could not be analyzed independently because they were too small or had an overrepresentation of cases. To include the largest number of individuals, SNP data genotyped on the same arrays were combined after basic sample and SNP quality control (details in eMethods). This resulted in seven array groups with raw genotype data from Affymetrix 6, Axiom Finngen, Axiom-NL, Illumina CytoSNP, Global Screening Array, Human Core Exome, and Omniexpress chip. These array groups were the units for further QC and imputation. Genotype data were imputed against the Human Reference Consortium (HRC) panel (v1.1)45. After imputation, the seven array groups had identical genome coverage and were merged into a single set for mega-analysis. All arrays included genotype data on the 22 autosomes. The availability of X-chromosome data varied (Supplementary Figure 1), with the first pseudo-autosomal region (PAR), non-PAR, and second PAR available in 55,141, 55,863, and 41,441 individuals, respectively.

We conducted principal component analysis (PCA) in PLINK (v1.9)46 on the imputed data based on the three superpopulations (African, Asian, European) from the 1000 Genomes Project reference panel (phase 3v5)47,48. The BIONIC genotype data were SNP and LD (linkage disequilibrium) pruned and projected onto the 1000 Genomes Project PCA space to compute principal components (PCs). Ancestry outliers in the BIONIC data were excluded based on a 4 SD distance from the mean of 0 for the first 6 standardized PCs. Additional ancestry outliers were removed based on visual inspection of PC clustering.

Genome-wide association mega-analysis

We conducted a GWAS mega-analysis of lifetime MDD and height for SNPs on the merged seven array HRC-imputed dataset. Analyses were conducted in fastGWA49 from the Genome-wide Complex Trait Analysis (GCTA) software50, with a generalized linear mixed model for lifetime MDD and a mixed linear model for height. We computed a genetic relatedness matrix (GRM) in GCTA to address relatedness in the sample (sparse-GRM threshold 0.05). SNPs with a minor allele frequency below 0.01 and imputation quality below 0.40 were excluded from analysis. Analyses were corrected for sex, age, 10 ancestry-informative principal components and genotype array. Independent GWAS signals were identified by conditional and joint (COJO) analysis in GCTA51, where the BIONIC GWAS MDD sample served as the LD reference sample, restricted to one individual per related pair (relatedness ≥ 0.125) as calculated by the KING software (V2.2.6)52.

SNP-based heritability and genetic correlation

The SNP-based heritability (SNP-h2) was estimated by Linkage Disequilibrium score regression (v1.0.1; LDSC)25. Effective N was calculated following previous literature53. For lifetime MDD, SNP-h2 was estimated on the liability scale, specifying 18% as population prevalence3 and sample prevalence as observed in the data (27% for the main GWAS model).

GWAS-by-subtraction

To explore if there was evidence for specific Dutch MDD genetic signal, we carried out GWAS-by-subtraction as implemented in the genomic structural equation modeling (Genomic SEM)54 software with the GWAS summary statistics from the largest major depression (MD) GWAS to date10 (PGC-MD) and the BIONIC MDD GWAS as input. When applied to two traits, GWAS-by-subtraction performs a GWAS of the unique genetic variation, i.e. the variance not shared, between the two traits26. We defined unique Dutch MDD as the genetic variation in the BIONIC MDD GWAS that was not explained by the PGC-MD European GWAS (excluding the Netherlands), running a GWAS on the residual unique Dutch MDD genetic variation in the BIONIC MDD GWAS after the PGC-MD European GWAS was regressed out.

Polygenic score prediction

We explored genetic risk prediction of lifetime MDD through in-sample and out-of-sample polygenic score (PGS) prediction. For in-sample prediction we generated MD PGSs in BIONIC based on the PGC-MD GWAS summary statistics (European samples and including 23andMe), excluding all datasets from the Netherlands10. PGSs were generated in the LDpred software (v0.9.1)55, a Bayesian method that derives posterior mean causal effect sizes from GWAS summary statistics by assuming a prior for the genetic architecture and linkage disequilibrium (LD) information. LD block information for West-European ancestry was derived from the UKBiobank56. The weights from the infinitesimal model were then used to perform allele scoring on the target sample with PLINK (v1.9)46. The association between the MD PGS and lifetime MDD case-control status in BIONIC was estimated by logistic regression in unrelated individuals with age at assessment, sex, and 10 ancestry informative PCs as covariates. The proportion of variance explained by the PGS on the liability scale for MDD case-control status was estimated according to Lee at al.57.

We explored out-of-sample MD PGS prediction based on the BIONIC MDD GWAS summary statistics in the UK Biobank, for a strict definition of MDD (clinical criteria) and broad depression (clinical criteria, sumscore cut-offs or self-report single item measures). Details on data field codes and genotype QC are in eMethods. A total of 23,755 and 132,122 individuals met the criteria for strict and broad phenotype definitions, respectively. PGS were generated in similar fashion to the in-sample prediction in LDpred (v1.0.10) and PLINK (v1.9) assuming an infinitesimal model. PGS prediction was estimated by logistic regression in unrelated individuals with age at measurement, sex, genotype array and 10 PCs as covariates.

Within-family designs are able to account for many sources of passive genotype-environment correlation, with dizygotic (DZ) twins having the additional benefit that all shared environmental factors are time-invariant among twins27. We sought to evaluate MD PGS prediction in a within-family design, following the approach of Selzam et al.27 with dizygotic twins from the Netherlands Twin Register (NTR; a collaborating cohort in BIONIC). We identified 1141 dizygotic twin pairs with non-missing lifetime MDD case-control status and MD PGSs. First, PGS prediction of MDD was assessed in the entire DZ twin sample through a logistic generalized linear mixed-effects model, correcting for sex and genotype array. Next, between-family PGS effects were defined as the change in lifetime MDD risk with a change in mean family PGS values, and within-family PGS effects as the change in lifetime MDD risk with a change in the difference between individual PGS and the family average PGS. The two were fit as predictors of MDD in the same model, together with a random effect for family, so that individual estimates are adjusted for and independent of the effect of the other estimate. The statistical difference between the between- and within-family PGS effects were then empirically tested through a chi-square test of the difference between their coefficients divided by the standard deviations of the sampling distribution of the estimate differences58,59.

Following a recent example in bipolar disorder and schizophrenia43, we computed MD PGS deciles in N = 2963 complete monozygotic and dizygotic twin pairs from the NTR cohort who took part in BIONIC. We then distinguished 3 groups: concordant affected (133) pairs, concordant unaffected (2257) pairs and discordant (573) pairs. We hypothesized that concordant affected twins would be overrepresented in high MD PGS deciles. We repeated this approach in an independent sample for replication, the Australian Genetics of Depression Study60 (Supplementary Information). We formally tested the association between the mean MD PGS in NTR twin pairs and MDD concordance in an ordinal logistic regression model, correcting for two ancestry-informative PCs43.

Genetic causal associations

We computed genetic correlations between the BIONIC lifetime MDD GWA mega-analysis and 1461 disease, personality and lifestyle traits using bivariate LDSC in the Complex-Traits Genetics Virtual Lab (CTG-VL) analysis pipeline (https://vl.genoma.io/). Details on their derivation and acquisition are given in the eMethods. We applied the bivariate latent causal variable (LCV) model to the 388 traits from the CTG-VL pipeline with a significant genetic correlation with the BIONIC MDD GWAS. The LCV method identifies potentially causal relationships among heritable traits28, quantifying the degree to which a genetic correlation between two traits may be explained by (partial) genetic causality as opposed to full horizontal pleiotropy through the genetic causal proportion (GCP) metric. The LCV method assumes a genetic correlation between trait A and B is mediated by a latent variable (L) that represents the causal component between the two traits. GCP is based on the correlations between trait A and L, and trait B and L, and quantifies the proportion to which a genetic correlation can be explained by potential causal effects. Assuming no bi-directional causality, a GCP value of 0 suggests the genetic correlation is purely due to horizontal pleiotropic effects, whereas a |GCP| of 1 indicates complete genetic causality. Intermediate GCP values denote partial genetic causality, where some, but not all, of the genetic correlation can be explained by causal effects. A |GCP| > 0.60 is considered to be robust61–63. A negative GCP would suggest that a given trait is likely to be a risk or protective factor for MDD, whereas a positive GCP would suggest that MDD causes the trait. Benjamini-Hochberg’s False Discovery Rate < 5% was applied to correct for multiple testing at both the genetic correlation and LCV steps.

Gene-level analyses & tissue enrichment

We performed a gene-based test and gene-set analysis of MDD in the MAGMA (Multi-marker Analysis of GenoMic Annotation; v1.08)30 tool. MAGMA employs a multiple linear principal components regression, and F test to obtain P values for genes30. In the gene-based test, input SNPs from the BIONIC MDD GWAS were mapped to 19,210 protein coding genes (symmetric gene window of 10kb), and tested for association with MDD in a SNP-wise top model. Pathway-specific enrichment of genetic signal for MDD was explored in competitive gene-set analysis with the gene-based analysis of MDD as input. Association between gene-set membership and gene Z-scores was computed using MAGMA. Gene sets were derived from categories C2 and C5 from the publicly-accessible MsigDB v7.0. Significance was set at the Bonferroni-corrected threshold of P = 0.05 / 15,488 = 3.23 × 10−6. We further performed tissue expression analysis in MAGMA with the gene-based test results as input. The MAGMA gene-property test compares average gene-expression for 53 tissue categories conditioning on average expression across all categories (one-sided) to test for tissue-specific enrichment. Gene expression datasets were obtained from GTEx v8 RNAseq. Significance was defined as P = 0.05 / 18,062 = 2.77 × 10−6.

Gene prioritization

We fine-mapped the genetic association results in the Fine-mapped Locus Assessment Model of Effector geneS (FLAMES)31 software. FLAMES integrates machine learning predictions linking SNPs to genes with GWAS-wide convergence of gene interactions, to predict the most likely effector gene in a locus. Significant loci in the BIONIC MDD GWAS were subjected to FLAMES to determine the effector genes. Gene Z-scores and tissue enrichment estimates were derived from the MAGMA gene-based test and tissue expression analyses, respectively, and served as input for Polygenic Priority Score (PoPS; v0.2)64. An LD matrix was generated based on the BIONIC MDD GWAS sample in LDstore2 (v2)65 and posterior probabilities for 95% credible sets of causal SNPs were derived from the MDD GWAS summary statistics using FINEMAP (v1.4)66. Information from the MAGMA, PoPS and FINEMAP analyses were combined in FLAMES to generate the FLAMES score, an estimate between 0 and 1 assigned to genes near the associated locus where the highest estimate indicates the most likely effector gene. A cut-off FLAMES score of 0.05 was used to determine the putative effector gene in each locus.

Supplementary Material

Supplementary Files

This is a list of supplementary les associated with this preprint. Click to download.

Figure 5. Causal architecture plots for MDD.

Figure 5.

Figure 5.

Causal architecture plots depicting the LCV phenome-wide analysis results for BIONIC MDD. Dots represent traits with a significant genetic correlation with MDD. The x-axis reflects the proportion of genetic causality (GCP). The y-axis shows the GCP absolute Z-score (statistical significance). The red-dashed line represents the statistical significance threshold (FDR < 5%), while the gray-dashed line represents the division between traits causally influencing MDD (left) and traits causally being influenced by MDD (right). Dot size reflects estimate accuracy (se). Annotation is provided for traits with a |GCP| > 0.60 that are significant at False Discovery Rate (FDR) < 5%, and that have a positive (plot A) or negative (plot B) genetic correlation with MDD.

Figure 5. Gene-based analysis of MDD.

Figure 5.

Manhattan Plot of the gene-based analysis of lifetime MDD in BIONIC. The x-axis indicates chromosomal position. The y-axis denotes −log10 P-value and indicates the strength of association between the gene and the outcome phenotype. The red line indicates Bonferroni-corrected genome-wide significance; P < 2.50 × 10−6 (based on 19,994 tested protein-coding genes).

Acknowledgements & Funding

We are very grateful to everyone who participated in this research or worked on this project and its contributing studies.

Funding for the BIONIC project was awarded to Dorret Boomsma and Brenda Penninx by the Biobanking and Biomolecular Resources Research Infrastructure (BBMRI-NL: 184.021.007;184.033.111). Below are cohort-specific funding declarations and acknowledgements.

We would like to thank the research participants and employees of 23andMe for making this work possible. The genome-wide summary statistics for the Adams et al. analysis of 23andMe, Inc., data were obtained under a data transfer agreement with the Vrije Universiteit Amsterdam.

Lifelines

The Lifelines initiative has been made possible by subsidy from the Dutch Ministry of Health, Welfare and Sport, the Dutch Ministry of Economic Affairs, the University Medical Center Groningen (UMCG), Groningen University and the Provinces in the North of the Netherlands (Drenthe, Friesland, Groningen). NARSAD Young Investigator Grant from the Brain & Behavior Research Foundation. VENI grant from the Talent Program of the Netherlands Organisation for Scientific Research (NWO-ZonMW 09150161810021). We thank Trynke de Jong for the contribution to Lifelines data collection. We thank Martje Bos and Victória Trindade Pons for their help in preparing the Lifelines phenotype data.

MooDFOOD

European Union FP7 funding for MooDFOOD Project ‘Multi-country cOllaborative project on the rOle of Diet, FOod-related behaviour, and Obesity in the prevention of Depression’ (grant agreement no. 613598).

TRAILS

Participating centers of the TRacking Adolescents’ Individual Lives Survey (TRAILS) include the University Medical Center and University of Groningen, the University of Utrecht, the Radboud Medical Center Nijmegen, and the Parnassia Group, all in the Netherlands. TRAILS has been financially supported by various grants from the Netherlands Organization for Scientific Research NWO (Medical Research Council program grant GB-MW 940-38-011; ZonMW Brainpower grant 100-001-004; ZonMw Risk Behavior and Dependence grant 60-60600-97-118; ZonMw Culture and Health grant 261-98-710; Social Sciences Council medium-sized investment grants GB-MaGW 480-01-006 and GB-MaGW 48007-001; Social Sciences Council project grants GB-MaGW 452-04-314 and GB-MaGW 452-06-004; ZonMw Longitudinal Cohort Research on Early Detection and Treatment in Mental Health Care grant 636340002; NWO large-sized investment grant 175.010.2003.005; NWO Longitudinal Survey and Panel Funding 481-08-013 and 481-11-001; NWO Vici 016.130.002, 453-16-007/2735, and Vi.C.191.021; NWO Gravitation 024.001.003), the Dutch Ministry of Justice (WODC), the European Science Foundation (EuroSTRESS project FP-006), the European Research Council (ERC-2017-STG-757364 and ERC-CoG-2015-681466), Biobanking and Biomolecular Resources Research Infrastructure BBMRI-NL (CP 32), the Gratama foundation, the Jan Dekker foundation, the participating universities, and Accare. Statistical analyses are carried out on the Genetic Cluster Computer (http://www.geneticcluster.org), which is financially supported by the Netherlands Scientific Organization (NWO 480-05-003) along with a supplement from the Dutch Brain Foundation.

LASA

The Longitudinal Aging Study Amsterdam is largely supported by grants from the Netherlands Ministry of Health, Welfare and Sport, Directorate of Long-Term Care.

NQplus

NQplus was core funded by ZonMw (ZonMw, Grant 91110030); add-on funding was provided by ZonMW Gezonde Voeding (ZonMw, Grant 115100007), BBMRI (Grant BBMRI-NL RP9 and CP2011-38) and Wageningen University and Research.

MOTAR

The MOTAR study was funded by NWO VICI grant number 91811602 of B.W.J.H. Penninx. NWO had no role in the design of the study, the collection, analysis and interpretation of the data, or in the preparation, review, or approval of the manuscript.

The Hoorn Studies

The GWAS in the Hoorn studies was supported by the Amsterdam University Medical Center, a grant from the Foundation for the National Institutes of Health through the Accelerating Medicines Partnership (no. HART17AMP) and the Dutch String of Pearls Initiative. We appreciate the corporation of the participants and research assistants who have been involved in the Hoorn Study and New Hoorn Study. We would like to thank Tootje Hoovers and Jolanda Bosman as well as all the researchers previously involved for the organization of both studies.

Netherlands Twin Register

NTR acknowledges funding from the Netherlands Organization for Scientific Research (NWO): Biobanking and Biomolecular Research Infrastructure (BBMRI–NL, 184.033.111) and the BBMRI-NL funded BIOS Consortium (NWO184.021.007); The Netherlands Twin Register is supported by grant NWO 480-15-001/674: Netherlands Twin Registry Repository: researching the interplay between genome and environment, the Avera Institute for Human Genetics and by multiple grants from the Netherlands Organization for Scientific Research (NWO). Genotyping was made possible by grants from NWO/SPI 56-464-14192, Genetic Association Information Network (GAIN) of the Foundation for the National Institutes of Health, Rutgers University Cell and DNA Repository (NIMH U24 MH 068457-06), the Avera Institute, Sioux Falls (USA) and the National Institutes of Health (NIH R01 HD042157-01A1, MH081802, Grand Opportunity grants 1RC2 MH089951 and 1RC2 MH089995) and European Research Council (ERC-230374). DIB acknowledges the Royal Netherlands Academy of Science Professor Award (PAH/6635).

Nijmegen Biomedical Study

The Nijmegen Biomedical Study is a population-based survey conducted at the Department for Health Evidence and the Department of Laboratory Medicine of the Radboud university medical center. Principal investigators of the Nijmegen Biomedical Study are L.A.L.M. Kiemeney, A.L.M. Verbeek, D.W. Swinkels en B. Franke.

Doetinchem Cohort Study

The Doetinchem Cohort Study is supported by the Dutch Ministry of Health, Welfare and Sport and the National Institute for Public Health and the Environment. We thank the respondents, epidemiologists and fieldworkers of the Municipal Health Service in Doetinchem for their contribution to the data collection for this study. The authors want to acknowledge the logistic management which was provided by P Vissink, and the data managers J van der Laan, R J de Kleine, I Toxopeus. Further, we thank all (senior) researchers who contributed to the data for collection, in particular in (alphabetical order): J M A de Boer, H B Bueno de Mesquita, P Engelfriet, G C Herber-Gast, G Hulsegge, D Kromhout, L Launer, A C J Nooyens, M C Ocke, S H van Oostrom, K Proper, J C Seidell, H A Smit, W G C Wendel-Vos.

NESDA & NESDAsib

The infrastructure for the NESDA study (www.nesda.nl) is funded through the Geestkracht program of the Netherlands Organisation for Health Research and Development (ZonMw, grant number 100001002) and financial contributions by participating universities and mental health care organizations (VU University Medical Center, GGZ inGeest, Leiden University Medical Center, Leiden University, GGZ Rivierduinen, University Medical Center Groningen, University of Groningen, Lentis, GGZ Friesland, GGZ Drenthe, Rob Giel Onderzoekscentrum).

NESDO

The infrastructure for NESDO is funded through the Fonds NutsOhra, Stichting tot Steun VCVGZ, NARSAD The Brain and Behaviour Research Fund, and the participating universities and mental health care organizations (VU University Medical Center, Leiden University Medical Center, University Medical Center Groningen, Radboud University Nijmegen Medical Center, and GGZ inGeest, GGNet, GGZ Nijmegen, GGZ Rivierduinen, Lentis, and Parnassia).

Funding Statement

We are very grateful to everyone who participated in this research or worked on this project and its contributing studies.

Funding for the BIONIC project was awarded to Dorret Boomsma and Brenda Penninx by the Biobanking and Biomolecular Resources Research Infrastructure (BBMRI-NL: 184.021.007;184.033.111). Below are cohort-specific funding declarations and acknowledgements.

We would like to thank the research participants and employees of 23andMe for making this work possible. The genome-wide summary statistics for the Adams et al. analysis of 23andMe, Inc., data were obtained under a data transfer agreement with the Vrije Universiteit Amsterdam.

Footnotes

Competing Interests None.

Ethics Statement

All relevant ethical regulations for working with human participants were followed in the conduct of the study, and written informed consent was obtained from all participants.

Data Availability

Data are available upon reasonable request from the contributing cohorts. Full summary statistics from the BIONIC GWA mega-analysis of MDD can be found here: https://tweelingenregister.vu.nl/bionic-mdd-2025

References

  • 1.Kessler RC, Bromet EJ. The Epidemiology of Depression Across Cultures. Annu Rev Public Health. 2013;34(1):119–138. doi: 10.1146/annurev-publhealth-031912-114409 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Vos T, Abajobir AA, Abate KH, et al. Global, regional, and national incidence, prevalence, and years lived with disability for 328 diseases and injuries for 195 countries, 1990–2016: a systematic analysis for the Global Burden of Disease Study 2016. The Lancet. 2017;390(10100):1211–1259. doi: 10.1016/S0140-6736(17)32154-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Fedko IO, Hottenga JJ, Helmer Q, et al. Measurement and genetic architecture of lifetime depression in the Netherlands as assessed by LIDAS (Lifetime Depression Assessment Self-report). Psychol Med. Published online 2020:1–10. doi: 10.1017/S0033291720000100 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Malhi GS, Mann JJ. Depression. The Lancet. 2018;392(10161):2299–2312. doi: 10.1016/S0140-6736(18)31948-2 [DOI] [PubMed] [Google Scholar]
  • 5.Sullivan PF, Neale MC, Kendler KS. Genetic Epidemiology of Major Depression: Review and Meta-Analysis. Am J Psychiatry. 2000;157(10):1552–1562. doi: 10.1176/appi.ajp.157.10.1552 [DOI] [PubMed] [Google Scholar]
  • 6.Flint J. The genetic basis of major depressive disorder. Mol Psychiatry. Published online January 26, 2023:1–12. doi: 10.1038/s41380-023-01957-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Levinson DF, Mostafavi S, Milaneschi Y, et al. Genetic studies of major depressive disorder: why are there no genome-wide association study findings and what can we do about it? Biol Psychiatry. 2014;76(7):510–512. doi: 10.1016/j.biopsych.2014.07.029 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Als TD, Kurki MI, Grove J, et al. Depression pathophysiology, risk prediction of recurrence and comorbid psychiatric disorders using genome-wide analyses. Nat Med. 2023;29(7):1832–1844. doi: 10.1038/s41591-023-02352-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Howard DM, Adams MJ, Shirali M, et al. Genome-wide association study of depression phenotypes in UK Biobank identifies variants in excitatory synaptic pathways. Nat Commun. 2018;9(1):1470. doi: 10.1038/s41467-018-03819-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Adams MJ, Streit F, Meng X, et al. Trans-ancestry genome-wide study of depression identifies 697 associations implicating cell types and pharmacotherapies. Cell. 2025;0(0). doi: 10.1016/j.cell.2024.12.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Wray NR, Ripke S, Mattheisen M, et al. Genome-wide association analyses identify 44 risk variants and refine the genetic architecture of major depression. Nat Genet. 2018;50(5):668–681. doi: 10.1038/s41588-018-0090-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Boyd JH, Weissman MM, Thompson WD, Myers JK. Screening for Depression in a Community Sample: Understanding the Discrepancies Between Depression Symptom and Diagnostic Scales. Arch Gen Psychiatry. 1982;39(10):1195–1200. doi: 10.1001/archpsyc.1982.04290100059010 [DOI] [PubMed] [Google Scholar]
  • 13.McIntosh AM, Sullivan PF, Lewis CM. Uncovering the Genetic Architecture of Major Depression. Neuron. 2019;102(1):91–103. doi: 10.1016/j.neuron.2019.03.022 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Mitchell AJ, Vaze A, Rao S. Clinical diagnosis of depression in primary care: a meta-analysis. The Lancet. 2009;374(9690):609–619. doi: 10.1016/S0140-6736(09)60879-5 [DOI] [PubMed] [Google Scholar]
  • 15.Cai N, Revez JA, Adams MJ, et al. Minimal phenotyping yields genome-wide association signals of low specificity for major depression. Nat Genet. 2020;52(4):437–447. doi: 10.1038/s41588-0200594-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Meng X, Navoly G, Giannakopoulou O, et al. Multi-ancestry genome-wide association study of major depression aids locus discovery, fine mapping, gene prioritization and causal inference. Nat Genet. Published online January 4, 2024. doi: 10.1038/s41588-023-01596-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Wang X, Walker A, Revez JA, et al. Polygenic risk prediction: why and when out-of-sample prediction R2 can exceed SNP-based heritability. Am J Hum Genet. 2023;110(7):1207–1215. doi: 10.1016/j.ajhg.2023.06.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Gouveia MH, Bentley AR, Leal TP, et al. Unappreciated subcontinental admixture in Europeans and European Americans and implications for genetic epidemiology studies. Nat Commun. 2023;14(1):6802. doi: 10.1038/s41467-023-42491-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Tropf FC, Lee SH, Verweij RM, et al. Hidden heritability due to heterogeneity across seven populations. Nat Hum Behav. 2017;1(10):757–765. doi: 10.1038/s41562-017-0195-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Huider F, Milaneschi Y, Hottenga JJ, et al. Genomics Research of Lifetime Depression in the Netherlands: The BIObanks Netherlands Internet Collaboration (BIONIC) Project. Twin Res Hum Genet. Published online March 18, 2024:1–11. doi: 10.1017/thg.2024.4 [DOI] [PubMed] [Google Scholar]
  • 21.Abdellaoui A, Hottenga JJ, de Knijff P, et al. Population structure, migration, and diversifying selection in the Netherlands. Eur J Hum Genet. 2013;21(11):1277–1285. doi: 10.1038/ejhg.2013.48 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Francioli LC, Menelaou A, Pulit SL, et al. Whole-genome sequence variation, population structure and demographic history of the Dutch population. Nat Genet. 2014;46(8):818–825. doi: 10.1038/ng.3021 [DOI] [PubMed] [Google Scholar]
  • 23.Bot M, Middeldorp CM, de Geus EJC, et al. Validity of LIDAS (LIfetime Depression Assessment Self-report): a self-report online assessment of lifetime major depressive disorder. Psychol Med. 2017;47(2):279–289. doi: 10.1017/S0033291716002312 [DOI] [PubMed] [Google Scholar]
  • 24.Yengo L, Vedantam S, Marouli E, et al. A saturated map of common genetic variants associated with human height. Nature. 2022;610(7933):704–712. doi: 10.1038/s41586-022-05275-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Bulik-Sullivan BK, Loh PR, Finucane HK, et al. LD Score regression distinguishes confounding from polygenicity in genome-wide association studies. Nat Genet. 2015;47(3):291–295. doi: 10.1038/ng.3211 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Demange PA, Malanchini M, Mallard TT, et al. Investigating the genetic architecture of noncognitive skills using GWAS-by-subtraction. Nat Genet. 2021;53(1):35–44. doi: 10.1038/s41588-020-00754-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Selzam S, Ritchie SJ, Pingault JB, Reynolds CA, O’Reilly PF, Plomin R. Comparing Within- and Between-Family Polygenic Score Prediction. Am J Hum Genet. 2019;105(2):351–363. doi: 10.1016/j.ajhg.2019.06.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.O’Connor LJ, Price AL. Distinguishing genetic correlation from causation across 52 diseases and complex traits. Nat Genet. 2018;50(12):1728–1734. doi: 10.1038/s41588-018-0255-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Aman AM, García-Marín LM, Thorp JG, et al. Phenome-wide screening of the putative causal determinants of depression using genetic data. Hum Mol Genet. 2022;31(17):2887–2898. doi: 10.1093/hmg/ddac081 [DOI] [PubMed] [Google Scholar]
  • 30.de Leeuw CA, Mooij JM, Heskes T, Posthuma D. MAGMA: Generalized Gene-Set Analysis of GWAS Data. PLoS Comput Biol. 2015;11(4):e1004219. doi: 10.1371/journal.pcbi.1004219 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Schipper M, de Leeuw CA, Maciel BAPC, et al. Prioritizing effector genes at trait-associated loci using multimodal evidence. Nat Genet. 2025;57(2):323–333. doi: 10.1038/s41588-025-02084-7 [DOI] [PubMed] [Google Scholar]
  • 32.Mitchell BL, Campos AI, Whiteman DC, et al. The Australian Genetics of Depression Study: New Risk Loci and Dissecting Heterogeneity Between Subtypes. Biol Psychiatry. Published online November 2, 2021. doi: 10.1016/j.biopsych.2021.10.021 [DOI] [PubMed] [Google Scholar]
  • 33.Davies MR, Kalsi G, Armour C, et al. The Genetic Links to Anxiety and Depression (GLAD) Study: Online recruitment into the largest recontactable study of depression and anxiety. Behav Res Ther. 2019;123:103503. doi: 10.1016/j.brat.2019.103503 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Miyazawa K, Ito K, Ito M, et al. Cross-ancestry genome-wide analysis of atrial fibrillation unveils disease biology and enables cardioembolic risk prediction. Nat Genet. 2023;55(2):187–197. doi: 10.1038/s41588-022-01284-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Fan W, Kan H, Liu HY, et al. Association between Human Genetic Variants and the Vaginal Bacteriome of Pregnant Women. mSystems. 2021;6(4):e00158. doi: 10.1128/mSystems.00158-21 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Surendran P, Drenos F, Young R, et al. Trans-ancestry meta-analyses identify rare and common variants associated with blood pressure and hypertension. Nat Genet. 2016;48(10):1151–1161. doi: 10.1038/ng.3654 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Cai N, Bigdeli TB, Kretzschmar W, et al. Sparse whole-genome sequencing identifies two loci for major depressive disorder. Nature. 2015;523(7562):588–591. doi: 10.1038/nature14659 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Yeo HJ, Shin MJ, Yeo EJ, et al. Tat-CIAPIN1 inhibits hippocampal neuronal cell damage through the MAPK and apoptotic signaling pathways. Free Radic Biol Med. 2019;135:68–78. doi: 10.1016/j.freeradbiomed.2019.02.028 [DOI] [PubMed] [Google Scholar]
  • 39.Plomin R. Blueprint: How DNA Makes Us Who We Are. MIT Press; 2019. [Google Scholar]
  • 40.Belsky DW, Harden KP. Phenotypic Annotation: Using Polygenic Scores to Translate Discoveries From Genome-Wide Association Studies From the Top Down. Curr Dir Psychol Sci. 2019;28(1):82–90. doi: 10.1177/0963721418807729 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Howe LJ, Nivard MG, Morris TT, et al. Within-sibship genome-wide association analyses decrease bias in estimates of direct genetic effects. Nat Genet. Published online May 9, 2022:1–12. doi: 10.1038/s41588-022-01062-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Horwitz TB, Balbona JV, Paulich KN, Keller MC. Evidence of correlations between human partners based on systematic reviews and meta-analyses of 22 traits and UK Biobank analysis of 133 traits. Nat Hum Behav. 2023;7(9):1568–1583. doi: 10.1038/s41562-023-01672-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Song J, Pasman JA, Johansson V, et al. Polygenic Risk Scores and Twin Concordance for Schizophrenia and Bipolar Disorder. JAMA Psychiatry. Published online August 28, 2024. doi: 10.1001/jamapsychiatry.2024.2406 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.American Psychiatric Association. Diagnostic and Statistical Manual of Mental Disorders (DSM5®). 5th ed. American Psychiatric Association; 2013. [Google Scholar]
  • 45.McCarthy S, Das S, Kretzschmar W, et al. A reference panel of 64,976 haplotypes for genotype imputation. Nat Genet. 2016;48(10):1279–1283. doi: 10.1038/ng.3643 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Purcell S, Neale B, Todd-Brown K, et al. PLINK: A Tool Set for Whole-Genome Association and Population-Based Linkage Analyses. Am J Hum Genet. 2007;81(3):559–575. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Auton A, Abecasis GR, Altshuler DM, et al. A global reference for human genetic variation. Nature. 2015;526(7571):68–74. doi: 10.1038/nature15393 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Novembre J, Stephens M. Interpreting principal component analyses of spatial population genetic variation. Nat Genet. 2008;40(5):646–649. doi: 10.1038/ng.139 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Jiang L, Zheng Z, Qi T, et al. A resource-efficient tool for mixed model association analysis of large-scale data. Nat Genet. 2019;51(12):1749–1755. doi: 10.1038/s41588-019-0530-8 [DOI] [PubMed] [Google Scholar]
  • 50.Yang J, Lee SH, Goddard ME, Visscher PM. GCTA: A Tool for Genome-wide Complex Trait Analysis. Am J Hum Genet. 2011;88(1):76–82. doi: 10.1016/j.ajhg.2010.11.011 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Yang J, Ferreira T, Morris AP, et al. Conditional and joint multiple-SNP analysis of GWAS summary statistics identifies additional variants influencing complex traits. Nat Genet. 2012;44(4):369–S3. doi: 10.1038/ng.2213 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Manichaikul A, Mychaleckyj JC, Rich SS, Daly K, Sale M, Chen WM. Robust relationship inference in genome-wide association studies. Bioinformatics. 2010;26(22):2867–2873. doi: 10.1093/bioinformatics/btq559 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Ip HF, van der Laan CM, Krapohl EML, et al. Genetic association study of childhood aggression across raters, instruments, and age. Transl Psychiatry. 2021;11(1):1–9. doi: 10.1038/s41398-02101480-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Grotzinger AD, Rhemtulla M, de Vlaming R, et al. Genomic structural equation modelling provides insights into the multivariate genetic architecture of complex traits. Nat Hum Behav. 2019;3(5):513–525. doi: 10.1038/s41562-019-0566-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Vilhjálmsson BJ, Yang J, Finucane HK, et al. Modeling Linkage Disequilibrium Increases Accuracy of Polygenic Risk Scores. Am J Hum Genet. 2015;97(4):576–592. doi: 10.1016/j.ajhg.2015.09.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Bycroft C, Freeman C, Petkova D, et al. The UK Biobank resource with deep phenotyping and genomic data. Nature. 2018;562(7726):203–209. doi: 10.1038/s41586-018-0579-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Lee SH, Goddard ME, Wray NR, Visscher PM. A Better Coefficient of Determination for Genetic Profile Analysis. Genet Epidemiol. 2012;36(3):214–224. doi: 10.1002/gepi.21614 [DOI] [PubMed] [Google Scholar]
  • 58.Clogg CC, Petkova E, Haritou A. Statistical Methods for Comparing Regression Coefficients Between Models. Am J Sociol. 1995;100(5):1261–1293. doi: 10.1086/230638 [DOI] [Google Scholar]
  • 59.Paternoster R, Brame R, Mazerolle P, Piquero A. Using the Correct Statistical Test for the Equality of Regression Coefficients. Criminology. 1998;36(4):859–866. doi: 10.1111/j.1745-9125.1998.tb01268.x [DOI] [Google Scholar]
  • 60.Byrne EM, Kirk KM, Medland SE, et al. Cohort profile: the Australian genetics of depression study. BMJ Open. 2020;10(5):e032580. doi: 10.1136/bmjopen-2019-032580 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.García-Marín LM, Campos AI, Cuéllar-Partida G, Medland SE, Kollins SH, Rentería ME. Large-scale genetic investigation reveals genetic liability to multiple complex traits influencing a higher risk of ADHD. Sci Rep. 2021;11(1):22628. doi: 10.1038/s41598-021-01517-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.García-Marín LM, Campos AI, Martin NG, Cuéllar-Partida G, Rentería ME. Inference of causal relationships between sleep-related traits and 1,527 phenotypes using genetic data. Sleep. 2021;44(1):zsaa154. doi: 10.1093/sleep/zsaa154 [DOI] [PubMed] [Google Scholar]
  • 63.Haworth S, Kho PF, Holgerson PL, et al. Assessment and visualization of phenome-wide causal relationships using genetic data: an application to dental caries and periodontitis. Eur J Hum Genet. 2021;29(2):300–308. doi: 10.1038/s41431-020-00734-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Weeks EM, Ulirsch JC, Cheng NY, et al. Leveraging polygenic enrichments of gene features to predict genes underlying complex traits and diseases. Nat Genet. 2023;55(8):1267–1276. doi: 10.1038/s41588-023-01443-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Benner C, Havulinna AS, Järvelin MR, Salomaa V, Ripatti S, Pirinen M. Prospects of Fine-Mapping Trait-Associated Genomic Regions by Using Summary Statistics from Genome-wide Association Studies. Am J Hum Genet. 2017;101(4):539–551. doi: 10.1016/j.ajhg.2017.08.012 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Benner C, Spencer CCA, Havulinna AS, Salomaa V, Ripatti S, Pirinen M. FINEMAP: efficient variable selection using summary data from genome-wide association studies. Bioinformatics. 2016;32(10):1493–1501. doi: 10.1093/bioinformatics/btw018 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Data are available upon reasonable request from the contributing cohorts. Full summary statistics from the BIONIC GWA mega-analysis of MDD can be found here: https://tweelingenregister.vu.nl/bionic-mdd-2025


Articles from Research Square are provided here courtesy of American Journal Experts

RESOURCES