Abstract
Normative models in neuroimaging learn the brain patterns of healthy population distribution and estimate how disease subjects like Alzheimer’s Disease (AD) deviate from the norm. Existing variational autoencoder (VAE)-based normative models using multimodal neuroimaging data aggregate information from multiple modalities by estimating product or averaging of unimodal latent posteriors. This can often lead to uninformative joint latent distributions which affects the estimation of subject-level deviations. In this work, we addressed the prior limitations by adopting the Mixture-of-Product-of-Experts (MoPoE) technique which allows better modelling of the joint latent posterior. Our model labelled subjects as outliers by calculating deviations from the multimodal latent space. Further, we identified which latent dimensions and brain regions were associated with abnormal deviations due to AD pathology.
Index Terms—: normative modelling, multimodal, variational autoencoders, mixture-of-product-of-experts
1. INTRODUCTION
Normative modelling for neurodegenerative disorders like Alzheimer’s Disease (AD) learn the brain patterns of cognitively unimpaired (healthy control) subjects. The trained model can be used to detect deviations at subject-level related to AD pathology [1, 2]. Recent works on VAE-based normative models have a unimodal structure with a single encoder and decoder that can handle only a single imaging modality [3–6]. However, brain disorders like AD are multifactorial, showing deviations from the norm in features across multiple imaging modalities. Since each modality is sensitive to disease effects to a varying degree, it is important to develop normative models that can handle multiple imaging modalities (e.g. Magnetic Resonance Imaging (MRI) and Positron Emission Tomography (PET)).
Existing multimodal VAE frameworks for normative modelling have modality-specific encoders and decoders and aggregate the encoding distributions to learn a joint latent representation [4]. Each modality can be treated as an ‘expert’ and the aggregated latent distributions are estimated by either taking the product (Product-of-Experts (PoE) [7]) or summation across the unimodal expert’s densities (Mixture-of-Experts (MoE) [8]). However, both PoE and MoE have their own sets of limitations. PoE joint distribution may be biased towards overconfident but miscalibrated experts, ignoring other modalities and this leads to a sub-optimal joint representation [8]. In MoE, the aggregating experts do not result in a distribution that is sharper than the other experts. Even if we increase the number of experts, the shared latent representation does not become more informative as in PoE, which means that no proper aggregated inference is possible in the latent space.
To address the above limitations of MoE and PoE, we adopted the Mixture-of-Product-of-Experts (MoPoE) approach [9] for aggregating information from multiple modalities in the latent space. MoPoE combines the strengths of both MoE and PoE, leading to more informative joint latent distribution. Previous works have shown that normative deviations calculated in the latent space outperform feature-space deviations [4, 10]. Our model quantified the deviation of AD patients based on the joint latent space and label subjects with abnormal (statistically significant) deviations as outliers. We showed that normative modelling with MoPoE better captures outliers compared to the baseline methods. We clinically validated the latent deviations to check if they were sensitive to the different AD stages and if they were significantly associated with cognition. Finally, we identified latent dimensions with abnormal deviations, mapped them to feature-space deviations and analyzed regional brain deviations associated with AD pathology.
2. METHODS
2.1. Approximate the joint posterior
Let be a set of conditionally independent N modalities. The backbone of our model is a multi-modal variational autoencoder (mVAE), a generative model of the form , where z is the latent variable and is the prior. mVAE optimizes the ELBO (Evidence Lower Bound) which is a combination of modality-specific likelihood distributions and the KL divergence between the approximate joint posterior and prior .
In mVAE, each modality can be treated as an “expert” and the approximate joint posterior can be estimated either by taking the product of the unimodal posteriors (Product-of-Experts (PoE)) [7] or by the Mixture-of-Experts (MoE) approach, where the joint inference distribution is represented by the sum of the unimodal inferences [8]. In PoE, if the inference for a particular modality is very sharp, the joint inference will be heavily dominated by it. Hence, the optimization of unimodal inference with low precision might be greatly degraded. On the other hand, MoE spreads its density over all individual experts and hence the joint latent posterior is not sharper than the unimodal posteriors [11].
2.2. Mixture-of-Product-of-Experts
To mitigate the challenges of both PoE and MoE, we adopted a generalization of both MoE and PoE, called Mixture-of-Product-of-Experts (MoPoE) proposed by [9] (Equation 2) where represents a random subset of N modalities and represents the power-set of all N modalities. MoPoE can be considered as a hierarchical distribution. First, the the unimodal posterior approximations of a subset are combined by PoE (Equation 1). Next, the subset approximations are combined by MoE (Equation 2). This combines the strengths of both MoE and PoE while addressing their weaknesses. The corresponding ELBO for MoPoE is represented by Equation 3.
2.3. Multi-modal normative modelling
Our framework (Figure 1) has separate encoders to encode the 2 modalities into their corresponding latent parameters (mean and variance). The unimodal latents were aggregated through the MoPoE approach to estimate the shared latent parameters (joint latent distribution). The shared latents were passed through the modality-specific decoders to reconstruct each modality. The model was first trained to characterize the healthy population cohort. We assumed disease abnormality can be quantified by measuring how AD subjects deviate from the joint space (latent deviations) [4, 10] or from the reconstruction errors of healthy controls (feature-space deviations) [3, 6]. At test time, the trained model was applied on the AD cohort to estimate both latent and feature deviations.
Fig. 1.

Proposed MoPoE normative modelling framework
| (1) |
| (2) |
| (3) |
2.4. Multi-modal latent and feature deviations
Mahalanobis distance:
To quantify how much AD each subject deviates from the latent distribution of healthy controls, we measured the Mahalanobis distance [4], which accounts for correlations between latent vectors.
where is a sample from the joint posterior distribution for subject and represent the mean and covariance of the healthy cohort latent position. We additionally derived a multivariate feature-space deviation index based on the Mahalonobis distance that quantifies how the reconstruction error of an AD subject deviate from the reconstruction errors of controls.
where is the mean squared reconstruction error between original and reconstructed input for subject j and brain region . and are the mean and covariance of the healthy cohort reconstruction error respectively.
Z-score latent deviations:
To analyse which latent dimensions and brain regions are associated with abnormal deviations due to AD pathology, we calculated both latent space Z-scores (for each latent dimension) and feature-space Z-scores and represent the latent values and reconstruction error of test subject j for i-th position respectively. and are the mean and standard deviations of healthy cohort latent values (reconstruction error) respectively.
3. EXPERIMENTS AND RESULTS
3.1. Data and feature processing
For training, we selected 248 cognitively unimpaired (healthy control) subjects from the Alzheimer’s Disease Neuroimaging (ADNI) dataset [12] with Clinical Dementia Rating (CDR) = 0 and no amyloid pathology. We used regional brain volumes extracted from T1-weighted MRI scans and regional Standardized Uptake Value Ratio (SUVR) values extracted from AV45 Amyloid PET scans as the two input modalities for our model (Figure 1). Both brain volumes and SUVR values were extracted from 66 cortical (Desikan-Killiany atlas) and 24 subcortical regions. At test time, we used 48 healthy controls for a separate holdout cohort and a disease cohort of 726 individuals across the following AD stages: preclinical stage with no symptoms (CDR = 0, A+) (N = 305), (b) CDR = 0.5 (N = 236) and (c) . Each brain ROI was normalised by removing the mean and dividing by the standard deviation of the healthy control cohort brain regions. We conditioned our model on the age and sex of patients, represented as one-hot encoding vectors, to remove the effects of covariates from the MRI and PET features.
3.2. Baselines and implementation details
We compared our proposed MoPoE framework to the following classes of baselines: (i) aggregation strategies in normative modelling: MoE and generalized PoE [4], PoE [3], (ii) unimodal baselines: MRI only, amyloid only and concatenated MRI + amyloid and (iii) state-of-the-art multimodal VAE models: mmJSD [13], JMVAE [14] and MVTCAE [15]. All baseline models except the unimodal ones were implemented using the open-source Multi-view-AE python package [16]. All models were trained using Adam optimizer with hyperparameters as follows: epochs = 500, learning rate = 10−5, batch size = 64 and latent dimensions in the range [5,10,15,20]. The encoder and decoder networks have 2 fully-connected layers of sizes 64, 32 and 32, 64 respectively.
3.3. Evaluating outlier detection performance
For each model, we calculated and for the healthy holdout cohort and disease cohort from ADNI. For both deviation metrics, we identified subjects with statistically significant deviations [17]. A good normative model is supposed to correctly identify disease subjects as outliers and healthy individuals lying within the normative distribution. Following aguila et al., [16], we used the positive likelihood ratio for assessing how well our MoPoE normative model can detect outliers.
Latent deviations achieved greater likelihood ratios compared to feature-space deviations (Table 1). All models performed similarly when using . Our proposed model with MoPoE achieved the best overall performance across different latent dimensions with the highest likelihood ratio for d = 5 and d = 10. Also, the multimodal models performed better than than the unimodal ones, indicating that better normative models can be learnt by modelling the joint distribution between modalities.
Table 1.
Likelihood ratio calculated for (ADNI)
| Latent dimensions | Latent Mahalnobis | Feature Mahalnobis | ||||||
|---|---|---|---|---|---|---|---|---|
| d = 5 | d = 10 | d = 15 | d = 20 | d = 5 | d = 10 | d = 15 | d = 20 | |
| MOPOE (proposed) | 6.73 | 8.1 | 5.72 | 4.56 | 1.85 | 2.52 | 2.16 | 1.76 |
| POE | 6.21 | 7.95 | 5.68 | 4.89 | 1.67 | 1.92 | 2.22 | 1.32 |
| MOE | 6.26 | 7.57 | 5.45 | 5.68 | 1.37 | 1.53 | 1.77 | 1.68 |
| gPOE | 6.64 | 7.8 | 5.75 | 4.3 | 2.11 | 1.65 | 1.86 | 1.55 |
| mri only | 4.92 | 4.65 | 4.25 | 4.67 | - | - | - | - |
| amyloid only | 5.25 | 5.12 | 4.41 | 4.83 | - | - | - | - |
| mri_amyloid_concat | 4.72 | 5.95 | 4.43 | 4.2 | 1.42 | 1.67 | 1.33 | 1.2 |
| mmJSD | 6.31 | 7.21 | 4.66 | 5.25 | 2.21 | 1.81 | 1.4 | 1.25 |
| JMVAE | 5.68 | 7.85 | 5.1 | 4.51 | 1.68 | 2.33 | 2.1 | 1.75 |
| MVTCAE | 6.34 | 6.83 | 5.56 | 3.82 | 1.35 | 1.68 | 1.52 | 1.21 |
3.4. Clinical validation of latent deviations
We demonstrated the clinical validation of our proposed multimodal latent deviations generated by MoPoE in two ways: (i) sensitivity with respect to disease staging and (ii) correlation with subject cognition. For the ADNI dataset, showed a monotonous increasing pattern across the AD stages and differed between groups overall (Figure 2 left). Thus, latent deviations generated by our proposed model can capture the neuroanatomical alterations in the brain due to the progressive stages of AD. We also measured the association between and Alzheimer’s Disease Assessment Scale (ADAS) cognition [18]. ADAS scores range from 0 to 70 with high scores indicating worse cognition. Age and sex-adjusted linear regressions were fitted between the and cognition scores, with Pearson Correlation Coefficient measuring the correlation between them. Our results showed that latent deviations were significantly associated with ADAS scores () with a correlation coefficient of 0.62 (Figure 2 right).
Fig. 2.

Left: Box plot showing the latent deviations across cognitively unimpaired (CU) subjects and the AD groups (in order of severity). Statistical annotations: ns: not significant, . Right: Association between and cognition scores (ADAS). Each point in the plot represents a subject and the red line denotes the linear regression fit of the points, adjusted by age and sex.
3.5. Interpretability analysis
3.5.1. Mapping from latent to feature deviations
Ideally, all latent dimensions can be used to reconstruct the input data and quantify feature-space deviations. The latent dimensions with statistically significant mean absolute Z-scores indicate the latent dimensions which show deviation between control and disease cohorts. This can provide an interpretation of how latent space deviations can be mapped to deviations in the feature-space. We passed these selected latent vectors through the decoders setting the remaining latent dimensions and covariates to be 0 such that the reconstructions and feature deviations reflect only the information encoded in the selected latent vectors. From the ADNI dataset, we identified 3 out of 10 latent dimensions (4,5 and 7) whose mean absolute and used them for generating the feature-space deviations (Figure 3 left).
Fig. 3.

Left: Latent dimensions (4,5 and 7) with statistically significant deviations (mean absolute or ). The dotted red line indicates . Latent dimensions above the dotted line were used for mapping to feature-space deviations. Right: Effect size maps showing the region-level pairwise group differences in between control subjects and each of the AD stages for both the modalities. The color bar represents the Cohen’s d statistic effect size (0.5 is considered a small effect, 1.5 a medium effect and 2.5 a large effect). Gray regions represent that no participants have statistically significant deviations after False Discovery Rate (FDR) correction.
3.5.2. Brain regions associated with AD abnormality
We visualized pairwise group comparisons in for each modality at each of 90 regions using the Cohen’s d-statistic effect size maps after False Discovery Rate (FDR) correction, adjusted by age and sex. High effect size corresponding to a region indicates significantly more grey matter volume atrophy (neurodegeneration) or amyloid deposition (pathological abnormalities) in the brain compared to healthy control subjects. Region-level group differences in MRI atrophy were most evident within the temporal, parietal and hippocampal regions which is consistent with the observations in existing literature [19, 20] (Figure 2B). Higher group differences in amyloid loading were mostly observed in the accumbens, precuneus, frontal and temporal regions, which are sensitive to amyloid pathology accumulation [21] (Figure 3 right).
4. CONCLUSION
We built on recent studies [3, 4] and introduced a multimodal VAE normative model which provides an alternative method of learning the joint normative distribution between multiple modalities to address the limitations of existing approaches. We adopted the MoPoE technique of aggregating information from modalities which results in more informative latent space and better capture of outliers as evidenced by the better significance ratio. The latent deviations were found to be sensitive towards the AD stages and significantly associated with subject cognition. We also identified latent dimensions with statistically Z-score deviations, mapped only those selected latent vectors to feature-space deviations and analyzed brain regions associated with AD abnormality.
ACKNOWLEDGEMENT
The preparation of this report was supported by the Centene Corporation contract (P19-00559) for the Washington University-Centene ARCH Personalized Medicine Initiative and the National Institutes of Health (NIH) (R01-AG067103). Computations were performed using the facilities of the Washington University Research Computing and Informatics Facility, which were partially funded by NIH grants S10OD025200, 1S10RR022984-01A1 and 1S10OD018091-01. Additional support is provided The McDonnell Center for Systems Neuroscience.
Footnotes
COMPLIANCE WITH ETHICAL STANDARDS
Ethical approval was not required as confirmed by the license attached with the open access data. The study was approved by the Institutional Review Board at the Washington University in St. Louis.
REFERENCES
- [1].Dong et al. , “Heterogeneity of neuroanatomical patterns in prodromal alzheimer’s disease: links to cognition, progression and biomarkers,” Brain, vol. 140, no. 3, pp. 735–747, 2017. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [2].Marquand et al. , “Conceptualizing mental disorders as deviations from normative functioning,” Molecular psychiatry, vol. 24, no. 10, pp. 1415–1424, 2019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [3].Kumar et al. , “Normative modeling using multimodal variational autoencoders to identify abnormal brain volume deviations in alzheimer’s disease,” in Medical Imaging 2023: Computer-Aided Diagnosis. SPIE, 2023, vol. 12465, p. 1246503. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [4].Lawry et al. , “Multi-modal variational autoencoders for normative modelling across multiple imaging modalities,” in International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2023, pp. 425–434. [Google Scholar]
- [5].Pinaya et al. , “Using deep autoencoders to identify abnormal brain structural patterns in neuropsychiatric disorders: A large-scale multi-sample study,” Human brain mapping, vol. 40, no. 3, pp. 944–954, 2019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [6].Pinaya et al. , “Using normative modelling to detect disease progression in mild cognitive impairment and alzheimer’s disease in a cross-sectional multi-cohort study,” Scientific reports, vol. 11, no. 1, pp. 1–13, 2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [7].Wu others, “Multimodal generative models for scalable weakly-supervised learning,” Advances in Neural Information Processing Systems, vol. 31, 2018. [Google Scholar]
- [8].Shi et al. , “Variational mixture-of-experts autoencoders for multi-modal deep generative models,” Advances in neural information processing systems, vol. 32, 2019. [Google Scholar]
- [9].Sutter et al. , “Generalized multimodal elbo,” arXiv preprint arXiv:2105.02470, 2021. [Google Scholar]
- [10].Lawry et al. , “Conditional vaes for confound removal and normative modelling of neurodegenerative diseases,” in International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2022, pp. 430–440. [Google Scholar]
- [11].Daunhawer et al. , “On the limitations of multimodal vaes,” arXiv preprint arXiv:2110.04121, 2021. [Google Scholar]
- [12].Mueller et al. , “Ways toward an early diagnosis in alzheimer’s disease: the alzheimer’s disease neuroimaging initiative (adni),” Alzheimer’s & Dementia, vol. 1, no. 1, pp. 55–66, 2005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [13].Sutter et al. , “Multimodal generative learning utilizing jensen-shannon-divergence,” Advances in neural information processing systems, vol. 33, pp. 6100–6110, 2020. [Google Scholar]
- [14].Suzuki et al. , “Joint multimodal learning with deep generative models,” arXiv preprint arXiv:1611.01891, 2016. [Google Scholar]
- [15].Hwang et al. , “Multi-view representation learning via total correlation objective,” Advances in Neural Information Processing Systems, vol. 34, pp. 12194–12207, 2021. [Google Scholar]
- [16].Aguila et al. , “Multi-view-ae: A python package for multi-view autoencoder models,” Journal of Open Source Software, vol. 8, no. 85, pp. 5093–5093, 2023. [Google Scholar]
- [17].Tabachnick et al. , Using multivariate statistics, vol. 6, pearson; Boston, MA, 2013. [Google Scholar]
- [18].Rosen et al. , “A new rating scale for alzheimer’s disease.,” The American journal of psychiatry, vol. 141, no. 11, pp. 1356–1364, 1984. [DOI] [PubMed] [Google Scholar]
- [19].Leech et al. , “The role of the posterior cingulate cortex in cognition and disease,” Brain, vol. 137, no. 1, pp. 12–32, 2014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [20].Young et al. , “Anatomical heterogeneity of alzheimer disease: based on cortical thickness on mris,” Neurology, vol. 83, no. 21, pp. 1936–1944, 2014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [21].Levitis et al. , “Differentiating amyloid beta spread in autosomal dominant and sporadic alzheimer’s disease,” Brain Communications, vol. 4, no. 3, pp. fcac085, 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
