Abstract
Accurate and early diagnosis of Alzheimer’s disease (AD) is critical for effective intervention and requires integrating complementary information from multimodal neuroimaging data. However, conventional fusion approaches often rely on simple concatenation of features, which cannot adaptively balance the contributions of biomarkers such as amyloid PET and MRI across brain regions. In this work, we propose MREF-AD, a Multimodal Regional Expert Fusion model for AD diagnosis. It is a Mixture-of-Experts (MoE) framework that models mesoscopic brain regions within each modality as independent experts and employs a gating network to learn subject-specific fusion weights. Utilizing tabular neuroimaging and demographic information from the Alzheimer’s Disease Neuroimaging Initiative (ADNI), MREF-AD achieves competitive performance over strong classic and deep baselines while providing interpretable, modality- and region-level insight into how structural and molecular imaging jointly contribute to AD diagnosis. The source code is available at https://github.com/PennShenLab/mref-ad.
Index Terms—: Alzheimer’s disease, multimodal imaging, mixture of experts, amyloid PET, magnetic resonance imaging (MRI)
I. Introduction
Alzheimer’s disease is a progressive neurodegenerative disorder, leading to cognitive decline and dementia [1], [2]. Despite the importance of early and accurate diagnosis for effective intervention, it remains challenging because clinical symptoms may overlap with healthy aging and other dementias. As such, in clinical practice, diagnosis often requires integrating multiple sources of evidence [3], [4]. This motivates the development of computational models that can combine complementary modalities for more robust and early AD detection.
Two commonly used imaging modalities in AD diagnosis are amyloid positron emission tomography (PET) and magnetic resonance imaging (MRI), which provide molecular and structural information of the brain, respectively. These modalities offer complementary information of two distinct yet related biological processes. In particular, amyloid PET captures the burden of amyloid- plaques in the brain, while MRI captures brain volume and is used to measure structural atrophy across cortical and subcortical regions. Both amyloid- plaques and brain volume are established markers that are widely associated with AD [2], [4]–[6]. In clinical and large-scale studies such as ADNI [7], neuroimaging scans are commonly also summarized as their quantitative measurements in tabular format for specific regions of interests (ROIs), providing fine-grained, high-dimensional anatomical details. Measurements are quantified in standardized uptake value ratios (SUVR) of amyloid- plaques for PET and regional brain volumes for MRI [8], [9].
Despite their utility, current multimodal models often rely on simple feature concatenation when handling structured tabular data, where features across modalities are analyzed as a single input and treated as equally informative [10], [11] (Fig. 1). Hence, this design ignores how the predictive value of different brain regions across different modalities might vary from patient to patient. Furthermore, interpreting these tabular models is difficult. Standard methods offer interpretability at the level of individual columns, or specific ROIs. For clinicians, the granularity of per-ROI analysis that treats localized anatomical points in isolation is often too fragmented to be insightful. Consequently, there is a lack of models that can provide mesoscopic brain regional interpretations, where the individual point ROI measurements are grouped into broader, clinically relevant brain regions such as the frontal or temporal lobes, that clinicians could use to evaluate disease progression and pathology [8], [12].
Fig. 1.

A comparison of strategies for multimodal Alzheimer’s study. A. Conventional methods usually adopt early-fusion by concatenating different imaging (e.g., PET, MRI) as well as other (e.g., demographics) modalities to form a single feature matrix to feed the model, resulting in conflated and non-specific analysis. As an illustration here, the PET and MRI data are both in the tabular form consisting of a hierarchy of regional summary statistics, while the demographic variates like age, sex, etc. are also stored in a table. SUVR stands for Standardized Uptake Value Ratio. B. MREF-AD introduces multimodal, regional experts design through MoE, ensuring the subsequent importance analysis are at brain region/group level rather than per feature column. More technical details can be seen later in Fig. 2. C. As a result, MREF-AD enables brain region-aligned, subject-specific explainability. Here, two subjects with early AD (high risk) and stable MCI (lower risk) are shown as examples.
Furthermore, in real-world clinical settings, multimodal data are frequently incomplete for some or all patients due to variability in imaging resources, institutional constraints, and patient-specific factors. While structural MRI is currently widely accessible, other more costly and specialized technologies, such as PET, remain limited in availability [13]. Moreover, certain patients may be unable to undergo specific imaging procedures due to contraindications, including biomedical implants, pre-existing comorbidities, or physical and behavioral limitations that prevent the acquisition of a complete multimodal dataset [14], [15]. Yet, standard multimodal frameworks typically deteriorate in performance in the event of missing modality [16], [17]. As such, they often require complete data, which excludes a significant portion of the broader patient population from model-assisted diagnosis [18]. Thus, ensuring robust model performance despite missing modality at inference time is crucial for clinical deployment [19]. This avoids the impractical need of maintaining separate models for different data or modality combinations to serve all patients when new cohorts of data become available.
To address these limitations, we propose MREF-AD, a Multimodal Regional Expert Fusion framework for AD diagnosis built upon an MoE design. MREF-AD decomposes each tabular neuroimaging modality—amyloid PET and structural MRI—into mesoscopic regional experts (e.g., frontal, temporal, parietal lobes), where each expert analyzes corresponding local imaging features (e.g., cortical thickness or regional SUVR values; Fig. 2). The gating mechanism enables MREF-AD to capture both cross-modality and within-modality (regional) heterogeneity, facilitating subject-specific and region-level interpretability. Each expert is implemented as a lightweight three-layer multilayer perceptron (MLP), ensuring that observed improvements arise from adaptive fusion rather than increased capacity. Furthermore, the gating network learns subject-specific fusion weights, dynamically emphasizing modality- and region-specific information most relevant to each individual.
Fig. 2. MREF-AD framework.

(A) Brain regions (n = 14) in the neuroimaging data and their corresponding FreeSurfer measurements. (B) In MREF-AD, 28 region-level experts from amyloid PET and MRI, together with one demographic expert, model modality–region features. Gating network assigns expert weights for Alzheimer’s disease diagnosis, enabling brain region-level interpretability.
We evaluated MREF-AD on data derived from the ADNI database, where the task is to classify subjects into cognitively normal (CN), mild cognitive impairment (MCI), and AD groups. The experimental results show that MREF-AD outperforms other baselines trained on concatenated imaging and demographic features, achieving improved accuracy, F1 score, and AUROC while offering interpretable insights into how brain region-specific structural and molecular imaging jointly contribute to disease classification. These results demonstrate the promise of our proposed adaptive and interpretable multi-modal fusion for Alzheimer’s disease diagnosis.
Our contributions can be summarized as follows:
We introduce MREF-AD, an MoE architecture that decomposes each modality into regional experts and learns subject-specific modality–region fusion weights for adaptive, interpretable multimodal imaging.
We show that this regional expert fusion improves three-way AD classification on the ADNI cohort compared to a set of representative baselines, especially in the event of missing modality.
We provide region-level interpretability analyses that reveal how structural MRI and amyloid PET features are differentially prioritized across brain regions, yielding a data-driven atlas of regional biomarker relevance.
II. Related Work
Early approaches to AD diagnosis classification often rely on straightforward multimodal fusion such as feature concatenation or kernel-based combinations [6], [10], [11], [20]–[22]. Many machine learning and deep learning-based studies have adopted early fusion strategies, combining modality-specific features into a single, flattened representation, followed by some downstream classifiers [6], [10], [11]. Other studies explored late fusion, for example, through ensemble learning or hierarchical fusion of single-modality predictors [20]–[22]. Among the various methods, ensemble-based Random Forest (RF) and XGBoost have demonstrated robust performance on AD prediction by capturing complex, non-linear interactions between multimodal features [22]–[24], while logistic regression (LR) remains a clinical standard due to its simplicity and clarity [25]. However, simple early fusion assumes all modalities are present and equally relevant and late fusion typically ignores the complex interactions across modalities. They provide limited insight into which modality or brain region drives a given prediction and often generate suboptimal results.
Modern approaches pursue more adaptive and interpretable multimodal fusion, leveraging techniques such as attention mechanisms, graph neural networks, contrastive learning, etc. [26]–[37] In particular, Transformer-derived attention mechanisms have emerged as the dominant strategy for integrating heterogeneous modalities. They enable learned feature weighting across modalities, improving interaction between multimodal signals [32]–[34]. As an example, FT-Transformer [37] represents one of the most significant works adapting the Transformer architecture for structured tabular data. It tokenizes each feature into an embedding and feeds all embeddings to Transformer layers with self-attention to capture high-order, non-linear feature interactions. However, FT-Transformer still takes multimodal input as a flat concatenation of features and can not explicitly encode modality/brain region contributions. Moreover, attention-based methods in general are sensitive to missing modality and require large training data with heavy computations due to the model complexity.
In addition, interpreting these methods for AD prediction remains challenging. Conventional classifiers such as logistic regression, random forests, XGBoost, and MLPs usually provide feature importance either natively (from the coefficients or split-based importance in tree ensembles) or through post-hoc explainability analyses (most notably SHAP [38]). Although these approaches can effectively rank individual features such as neuroimaging traits and methods like SHAP have been increasingly adopted to explain complex black-box models for AD [22], [39], they generally operate at the granularity of individual features and provide per-column feature importance of the input data matrix. Aggregating these dispersed feature attributions into clinically meaningful insights (e.g., temporal lobe regional contribution) requires manual post-processing and may not faithfully reflect how the models utilize multi-modal (e.g., MRI and PET neuroimaging), grouping signals (e.g., brain structures/regions) during prediction.
Recently, methods based on mixture-of-experts strategies have emerged to achieve adaptive fusion and emphasize intrinsic interpretability by the flexible experts design for several biomedical tasks including AD [40]–[45]. MoE architectures use a gating network to select among multiple expert subnetworks, allowing the model to tailor the feature processing to each input subject. Yun et al. developed Flex-MoE for AD classification, which implemented generalized and specialized routers through a sparse MoE structure to deal with missing modality within imaging, clinical, biospecimen, and genetic multimodal data [40]. In another work, Jiang et al. proposed a multi-task multi-gate MoE framework, M3AD, for simultaneous AD diagnosis and cognitive decline prediction from structural MRI images [41]. It utilizes specialized experts for diagnosis-specific pathological patterns and shared experts for common structural features across cognitive decline progression. Despite their strong performance across heterogeneous datasets, current MoE models for AD only operate at the whole-modality level and are not able to detect region-specific contributions within each modality.
III. Methods
Data used in the preparation of this article were obtained from the ADNI database1. The ADNI was launched in 2003 as a public-private partnership led by Principal Investigator Michael W. Weiner, MD. The primary goal of ADNI has been to test whether serial MRI, PET, other biological markers, and clinical and neuropsychological assessment can be combined to measure the progression of mild cognitive impairment (MCI) and early AD. All participants provided written informed consent, and study protocols were approved by each participating site’s Institutional Review Board (IRB). Up-to-date information about the ADNI is available at www.adni-info.org.
This study focuses on the amyloid and MRI subset of the ADNI [46] dataset, which integrates multimodal neuroimaging data, including amyloid PET, structural MRI, as well as demographic information (age, sex, education, race, and ethnicity). Amyloid PET and MRI were selected as representative molecular and structural biomarkers, providing complementary information on amyloid- deposition and neurodegenerative atrophy, respectively.
After preprocessing and quality control, the final dataset included 1,530 unique participants, each contributing one imaging visit corresponding to their last available amyloid–MRI session (Table I). Each visit was labeled with the participant’s clinical diagnosis of cognitively normal (CN), mild cognitive impairment (MCI), or Alzheimer’s disease (AD) from the ADNI database.
TABLE I.
Participant demographics at the last available visit in the ADNI dataset. Values are reported as mean ± standard deviation unless otherwise noted. Participant counts refer to unique individuals, with diagnoses assigned based on their last recorded imaging session.
| Characteristic | CN | MCI | AD | Total |
|---|---|---|---|---|
| Participants (n) | 637 | 557 | 336 | 1530 |
| Age (years) | 74.8 ± 8.0 | 75.2 ± 8.0 | 76.6 ± 8.2 | 75.3 ± 8.1 |
| Sex (M/F) | 268/369 | 304/253 | 192/144 | 764/766 |
| Education (years) | 16.7 ± 2.4 | 16.0 ± 2.7 | 15.9 ± 2.6 | 16.3 ± 2.6 |
Amyloid PET features were aligned with MRI features by matching both subject ID and visit code to ensure same-session correspondence between modalities. For each modality, region-level measurements were defined using the FreeSurfer cortical and subcortical parcellation atlas [47], [48]. By mapping individual brain scans into a standardized set of ROI measures, FreeSurfer derives morphometric features (68 in total from left and right hemispheres), which are widely utilized as a testbed for evaluating computational models within AD research and serve as a robust and reproducible benchmark. For amyloid PET, we extracted standardized uptake value ratios (SUVRs) indicating regional amyloid- plaque burden, and for MRI, regional volumes (mm3) of the brain.
A. Model Architecture
We formulate the three-way Alzheimer’s disease classification task as a multimodal learning problem. Let represent input features from modality-region pairs. Our objective is to learn a robust mapping that predicts the diagnostic label even when specific features are missing.
MREF-AD (Fig. 2) is an MoE model of modality–region experts and a learned gating network that adaptively fuses their outputs. In our implementation, each neuroimaging modality (amyloid PET and MRI) is partitioned into 14 mesoscopic brain regions, yielding 28 imaging experts. Furthermore, we include demographic variables and treat them as an additional expert within our multimodal fusion framework. Thus, this results in a total of experts (28 imaging experts + 1 demographic).
Expert Networks.
Each expert is implemented as a three-layer multilayer perceptron (MLP) with architecture , where is the input dimension for expert (varying by modality-region pair), H1 and H2 are hidden dimensions with value 145, and C=3 is the number of output classes. Each expert maps region-specific input features to class logits for .
Gating Network.
To capture within- and cross-modality structure, MREF-AD employs a gating scheme that produces subject-specific fusion weights. The gating network first computes modality-level logits for modalities (MRI, PET, demographics), which are converted to modality weights via:
| (1) |
For each modality , the gating network then computes region-level logits for the regions within that modality, which are normalized to regional weights:
| (2) |
where denotes the set of regional experts within modality . The final expert weight combines these two levels:
| (3) |
where denotes the modality associated with expert . Note that by construction. While MREF-AD is designed to support this hierarchical gating scheme, the framework also allows for a flat gating configuration where a single gating network assigns weights across all 29 modality-region experts simultaneously. For the primary results reported in this study, the flat gating configuration was utilized as it was empirically found to offer superior diagnostic performance and more flexible weight allocation.
To regularize the gating mechanism, we added a sparsity penalty based on the entropy of the gating weights and a diversity penalty encouraging decorrelated expert outputs. The total loss function is:
| (4) |
where is the cross-entropy loss. The sparsity penalty minimizes the entropy of gating weights to encourage focused expert selection:
| (5) |
where is the batch size. The diversity penalty encourages decorrelated expert predictions by minimizing pairwise cosine similarities:
| (6) |
Optimization used the AdamW optimizer (weight decay 1 × 10−4) for up to 40 epochs, with early stopping based on validation loss (patience = 10). Hyperparameters were tuned using Optuna [49], searching across learning rates from 1 × 10−5 to 1 × 10−2, dropout rates from 0.0 to 0.5, sparsity penalties from 0.0 to 0.2, and diversity penalties from 0.0 to 0.1. The best performing configuration utilized an optimal learning rate of 5.3 × 10−4, dropout rate of 0.22, , and .
Final Prediction.
The model’s output combines all expert predictions weighted by their gating scores:
| (7) |
where . Missing modalities are handled by masking their corresponding experts (setting for missing modality ) before applying the softmax normalization, which naturally excludes them from the weighted sum.
Although we focus on amyloid PET, MRI, and demographics in this study, the formulation is modality-agnostic and can be directly applied to other multimodal imaging settings.
B. Handling Missing Modality
In real-world clinical practice, multimodal data are often incomplete and modalities might be missing at inference time. To evaluate model robustness, we simulated clinical scenarios where a modality is unavailable by withholding all features for a specific modality (e.g., amyloid PET or MRI). Specifically, all models are trained on complete data, while at inference time, a modality is nulled by removing the numerical values within the test sets.
MREF-AD is explicitly designed to handle missing modalities by (i) masking and excluding missing modalities or experts from the gating distribution and (ii) normalizes the remaining gate weights such that predictions are based only from observed inputs.
To ensure that every expert in MREF-AD receives a valid numerical input, even when specific modalities are unavailable, we first impute missing values with the training set median. Binary indicators are associated with each expert to indicate input data availability and are propagated to the gating networks. They mask expert participation at the gating stage, such that experts with missing data do not contribute to the mixture. Hence, missingness is handled structurally, where missing modalities are internally ignored by the gating mechanism through the masking of their gating logits prior to normalization. As a result, missing experts are guaranteed to receive zero mixture weight, and the remaining weights are renormalized to sum to one. More formally, given and are the modality- and region-level gate weights, then
| (8) |
This guarantees final predictions to be computed only from existing modalities.
As such, MREF-AD is able to perform inference on missing modality through modality and expert masking, without the need for retraining or modality-specific model variants. Furthermore, the learned expert weights provide a sample-specific explanation of expert contribution to each prediction, conditional on availability.
Baseline models (detailed below) that cannot natively process missing inputs similarly apply median imputation, or mean statistics following their default implementation.
C. Baselines
To provide fair comparisons, we implemented two multimodal fusion baselines using the same three-layer MLP building block and training hyperparameters as MREF-AD. In the concatenation baseline, all modality- and region-level features were flattened into a single feature vector and passed through a three-layer MLP classifier. In the late-fusion baseline, separate MLPs were trained for each modality, and their softmax probabilities were averaged to obtain the final prediction. Furthermore, we compared our approach against state-of-the-art architectures for tabular and multimodal data. First, we compared with Flex-MoE [40], a sparse mixture-of-experts model specifically designed for AD classification with data across imaging and clinical modalities. We also compared against the sophisticated tabular-attention model, FT-Transformer [37], and trained traditional machine learning classifiers (Random Forest, Logistic Regression, XGBoost) on the same concatenated feature representation.
All neural models were trained with cross-entropy loss and class-balanced weights to address label imbalance.
D. Training and evaluation
We evaluated all models by subsampling across 10 random seeds, partitioning into training, validation, and test sets with a 80/10/10 ratio, respectively. For each seed, we performed hyperparameter optimization for each model with Optuna [49] on the training and validation subsets. To ensure sufficient search space exploration of hyperparameters, we ran 100 Optuna trials for the traditional machine learning baselines and 200 trials for neural network-based architectures, including MREF-AD. We then obtained the best-performing hyperparameters for each model, which were used to train the final model on the combined training and validation data. The model’s performance was evaluated on the held-out test set. Finally, we report the accuracy and macro-F1 mean and standard deviation across all 10 seeds to provide a robust assessment of model performance and stability. All features were z-score normalized within each training fold and the corresponding statistics were applied to the validation data.
IV. Results
Table II summarizes performance on the three-way diagnostic task (CN vs MCI vs AD). Under optimal conditions with complete data modality, MREF-AD consistently outperforms the baseline and state-of-the-art models. While traditional ML models seemingly provide strong baselines, MREF-AD achieves significant improvements in F1 score over all baselines (paired one-sided Wilcoxon tests with Holm correction, p < 0.01855), and similar significant gains in accuracy over Logistic Regression, MLP, XGBoost, FT-Transformer, and Flex-MoE. More importantly, while traditional and neural models provide competitive performance in these optimal settings, they lack the architectural flexibility to handle the data irregularities common in clinical practice.
TABLE II.
Diagnostic performance comparison between traditional and deep multimodal fusion methods, including our proposed MREF-AD framework, across three classes of AD diagnosis (CN, MCI, AD). Results are reported as mean ± standard deviation across ten seeds.
| Model | Accuracy | F1-score |
|---|---|---|
| Logistic Regression | 0.636 ± 0.019 | 0.628 ± 0.019 |
| Random Forest | 0.638 ± 0.030 | 0.635 ± 0.032 |
| XGBoost | 0.626 ± 0.032 | 0.619 ± 0.033 |
| MLP | 0.619 ± 0.020 | 0.608 ± 0.024 |
| FT-Transformer | 0.611 ± 0.027 | 0.598 ± 0.027 |
| Flex-MoE | 0.597 ± 0.033 | 0.584 ± 0.037 |
| MREF-AD (ours) | 0.657 ± 0.024 | 0.664 ± 0.028 |
A. Robustness to Missing Modalities
Next, we evaluate the robustness of MREF-AD under realistic clinical settings where one modality is missing at inference time.
Taking models trained on the complete dataset and removing one of the neuroimaging modalities from the test data at inference time, MREF-AD significantly outperforms FT-Transformer, MLP, and logistic regression in both missing PET and MRI modality conditions (Fig. 3). In both missing modality conditions, MREF-AD maintains an F1-score of around 0.50. The next best-performing baseline is FT-Transformer, with an F1-score that drops to approximately 0.38, but similar to other baselines, experience a great decrease when MRI is missing. On the other hand, MLP and logistic regression yield results near 0.20 in both missing modality conditions.
Fig. 3. Model robustness to missing modalities at inference time.

Comparison of F1-scores for MREF-AD, Flex-MoE, FT-Transformer, MLP, and logistic regression baseline models under missing amyloid PET or MRI data at inference time. Error bars indicate standard deviation across 10 seeds and data splits.
Moreover, although logistic regression achieves competitive performance under full-modality conditions (Table II), notably, it shows the most significant performance decline in the event of missing modality. In contrast, MREF-AD demonstrates robustness to missing modality inputs. This shows that the handling of missing modality within its architecture helped the model maintain the highest diagnostic performance.
B. Clinical Interpretability
The gating network in MREF-AD produces sample-specific softmax weights over experts, determining their relative influence on the final prediction. This enables model interpretability by aggregating these weights across subjects to obtain expert contributions that quantify the average importance of each regional expert in diagnostic decision-making.
Fig. 4 summarizes these contributions across test sets. Both MRI and amyloid PET experts receive substantial weights, with MRI experts showing slightly higher aggregate contributions. Among MRI experts, Temporal, Subcortical Temporal and Ventricular regions rank highly, suggesting that structural alterations in these areas offer complementary information for disease classification. On the amyloid side, regions such as Brainstem, Striatum/Basal Ganglia, and Corpus Callosum White Matter exhibit prominent contributions, reflecting strong molecular cues. Several regions—including the Subcortical Temporal area—show high contributions in both modalities, indicating convergent molecular and structural relevance, whereas regions such as the Brainstem exhibit dominant importance within a single modality, demonstrating the model’s capacity to disentangle modality-specific signals. The Demographic expert ranks among the highest-contributing experts, underscoring the consistent influence of subject-level variables.
Fig. 4. Average modality–region expert contributions.

Top: mean gating weights across subjects for each regional expert, grouped by modality (MRI, PET, demographics). Middle and bottom: MRI and amyloid PET expert contributions projected onto cortical and subcortical regions; Color intensity (darker shades) indicates greater model reliance on that regional expert.
The averaged pattern shows the model’s overall behavior, however MREF-AD also offers interpretability for individual clinical profiles. To illustrate this, Fig. 4 shows the model’s subject-specific gating weights for representative CN, MCI, and AD cases, showing how the model shifts its emphasis across the disease stages. For a CN individual, the temporal lobe in PET receives the highest weight. In MCI, the model distributes more weight on the corpus callosum and frontal regions in PET. In an AD patient, the weighting shifts toward subcortical temporal structures in MRI. Together, these examples show that MREF-AD offers personalized interpretability into each patient’s diagnosis and adapts its regional and modality-level focus in a manner aligned with known clinical progression of Alzheimer’s disease.
Overall, these patterns indicate that MREF-AD adaptively balances shared and modality-specific regional information, offering biologically interpretable insights into Alzheimer’s disease mechanisms. The elevated contributions of temporal and subcortical/ventricular regions are consistent with established imaging findings in AD [2], [4], [5], while the learned modality–region expert weights reveal how structural MRI, amyloid PET, and demographics are differentially prioritized across regions in a way that is not prescribed a priori by the model, yielding a data-driven map of regional biomarker relevance.
C. Ablations
Next, we conducted comprehensive ablation experiments to evaluate the robustness and design choices of MREF-AD (Table III). Under leave-one-modality-out retraining settings, performance decreased modestly relative to the full model, indicating that MREF-AD effectively leverages complementary information across modalities. Among modalities, removing amyloid PET led to the largest performance decline while MRI removal produced smaller changes.
TABLE III.
Ablation and sparsity analysis of MREF-AD. Values are mean ± SD over 10 seeds.
| Configuration | Accuracy | F1-score |
|---|---|---|
| MREF-AD | 0.657 ± 0.024 | 0.664 ± 0.028 |
| (A) Modality ablation | ||
| w/o Amyloid | 0.592 ± 0.023 | 0.596 ± 0.022 |
| w/o MRI | 0.610 ± 0.021 | 0.606 ± 0.019 |
| w/o Demographic | 0.638 ± 0.026 | 0.644 ± 0.030 |
| Amyloid only | 0.624 ± 0.016 | 0.615 ± 0.023 |
| MRI only | 0.568 ± 0.030 | 0.570 ± 0.034 |
| (B) Expert sparsity (top-k gating) | ||
| Top-5 experts | 0.620 ± 0.025 | 0.609 ± 0.039 |
| Top-3 experts | 0.642 ± 0.022 | 0.643 ± 0.030 |
| Top-1 expert | 0.589 ± 0.039 | 0.562 ± 0.039 |
| (C) Gating architecture ablation | ||
| Modality-only gate | 0.652 ± 0.007 | 0.654 ± 0.009 |
| Region and Modality gates | 0.656 ± 0.029 | 0.665 ± 0.035 |
We further examined the effect of expert sparsity and gating design. Top- gating, specifically Top-3 sparse MREF-AD, yielded performance comparable to the full model, demonstrating that MREF-AD can achieve competitive performance with significantly fewer active experts, thereby enhancing efficiency. While a hierarchical modality-region gating scheme was initially designed to capture structured dependencies, empirical results indicated that a flat gating architecture—directly assigning weights across all 29 experts—offered better diagnostic performance.
D. Model Complexity
While performance in AD diagnosis, robustness to missing modality scenarios, and clinical interpretability are crucial, beyond that, the practicality of deep learning models in clinical settings is also influenced by the computational cost and ease of deployment [50], [51]. As such, we evaluate the complexity of MREF-AD by comparing its total model parameters, the number of active parameters actively used during inference, and active Floating Point Operations (FLOPs) against the neural network-based baselines (Table IV).
TABLE IV.
Model Complexity and Computational Efficiency Comparison
| Method | Active Params | Active FLOPs | Total Params |
|---|---|---|---|
| MLP (Baseline) | 161,923 | 364,044 | 161,923 |
| FT-Transformer | 941,379 | 2,666,760 | 941,379 |
| Flex-MoE | 629,891 | 1,424,550 | 1,948,291 |
| MREF-AD (Top-1) | 111,544 | 221,820 | 456,092 |
| MREF-AD (Top-3) | 135,036 | 268,364 | |
| MREF-AD (Top-5) | 158,528 | 314,908 | |
| MREF-AD (Top-10) | 217,258 | 431,268 | |
| MREF-AD (Dense) | 440,432 | 873,436 |
Note: Active parameters and FLOPs refer to the computational cost per single inference. Lower values indicate more efficient prediction. For MREF-AD, Top- denotes the number of experts activated by the gating network.
Our results show that MREF-AD achieves parameter-efficient learning through its sparse MoE design. It has a total model capacity of 456,092 parameters, smaller than the 941,379 parameters of the FT-Transformer and 1,948,291 for Flex-MoE. More importantly, MREF-AD uses only a subset of these parameters during inference through its sparse top-k expert activation mechanism, whereas FT-Transformer and MLP activate all available parameters. In the Top-1 configuration, MREF-AD activates just 111,544 parameters per prediction, which significantly reduces the model’s effective complexity at inference time. This translates to just 221,820 FLOPs per pass, which is roughly a 12-fold reduction compared to FT-Transformer (2.67M FLOPs) and makes MREF-AD even more efficient compared to the simpler MLP baseline (364K FLOPs). Similarly, in Top-3 setting, the active parameters and FLOPs of MREF-AD remain lower than MLP while still maintaining a higher predictive performance (Tables II and III). In the dense setting, when all experts are activated, MREF-AD’s complexity remains lower than FT-Transformer and Flex-MoE, fewer active parameters (440k vs. 941k and 629,891, respectively) and fewer floating-point operations (873k vs. 2.67M and 1.42M FLOPs). Ultimately, MREF-AD offers a flexible, tunable configuration for clinical deployment through the top- expert parameter. Hence, allowing the model to be adapted to different clinical environments with varying computational resources.
V. Discussion
Our study proposes MREF-AD, which enables adaptive, interpretable fusion of Alzheimer’s multimodal data by explicitly treating mesoscopic brain regions as distinct experts. Evaluated on ADNI benchmark data, the model demonstrates its strength not only in improved classification performance but also in handling missing modality and promoting clinical explainability into the disease process. Indeed, one significant advantage of MREF-AD is its intrinsic interpretability. Unlike black-box deep learning models or post-hoc analysis approaches [22], MREF-AD’s gating weights provide a direct, probabilistic measure of feature importance. The gating mechanism allows dynamic weights adjustment for different subjects, facilitating personalized interpretability.
Another key finding from our experiments is MREF-AD’s stability under missing-MRI and -PET conditions, especially in contrast to logistic regression’s significantly larger drop in performance. Interestingly, logistic regression has competitive performance under full-modality. This behavior has also been observed in previous works, where linear models often perform well in healthcare settings under complete data assumptions [52] even when compared with neural networks like MLP [53]. Hence, the performance degradation under missing-modality scenarios highlights a key limitation that such models lack the architectural ability to handle missing data and rely on the quality of the imputation. While various modality imputation strategies have been explored, recent work increasingly emphasizes architecture-level approaches that inherently accommodate missing modalities [17]. Thus, our results further validate the importance of considering architectural robustness to incomplete modalities for clinical deployment, where fully observed multimodal data are rarely guaranteed, and the need to move beyond complete data assumptions.
Beyond robustness to missing data, our results suggest that the modality-region expert decomposition of MREF-AD offers meaningful insights into AD neuroimaging. Although MREF-AD and the MLP baseline share similar neural network building blocks, MREF-AD consistently improves F1 and accuracy, indicating that the gains are not simply driven by model size. Instead, the gating mechanism captures heterogeneity in how different brain regions and modalities contribute to diagnosis across individuals. The aggregated expert contributions from MRI and amyloid PET experts align with the clinical patterns of neurodegenerative atrophy and molecular pathology. The modality ablations on how the model uses complementary information confirm the best performance arises from joint molecular and structural evidence. The way MREF-AD approaches these patterns at a mesoscopic level also match how clinicians often summarize AD-related abnormalities in practice.
Despite its promise, this study has several limitations. First, while we utilized ROI-level tabular data to ensure lightweight efficiency (e.g., as shown in Table IV), it discards voxel-level information and spatial patterns that raw image-based models could leverage. Future work could replace the MLP experts in the MoE with 3D-CNN or visual transformer [54] blocks to learn directly from the images. Moreover, here we use only one imaging session per subject (last available PET/MRI visit), whereas future work could utilize longitudinal data for progression trajectories modeling or early detection, which is more important for clinical intervention of AD. Finally, although assessed on MRI, Amyloid PET and demographic data as a proof-of-concept here, MREF-AD could be extended further to incorporate additional modalities such as tau PET, FDG PET, cognitive scores, CSF markers, and genetic data, to allow a more comprehensive view of AD etiology.
VI. Conclusions
We presented MREF-AD, an adaptive Mixture-of-Experts framework for multimodal neuroimaging-based Alzheimer’s disease diagnosis. By modeling amyloid PET, MRI, and demographic features as independent experts and using a gating network for subject-specific fusion, MREF-AD achieves robust and interpretable predictions. Furthermore, our results show that the model maintains AD diagnostic utility even when a specific imaging modality is missing at inference. Compared with baseline multilayer perceptron models and traditional classifiers, MREF-AD achieved superior diagnostic performance across accuracy and F1-score, demonstrating the benefit of adaptive fusion. Beyond performance gains, the interpretability analysis revealed biologically meaningful expert weighting patterns, including modality–region combinations that align with known biomarker progression and highlight temporal and subcortical involvement in AD. These results highlight MREF-AD’s potential as both a predictive and explanatory framework for multimodal neuroimaging. While we illustrate its utility on Alzheimer’s disease diagnosis, the architecture is general and can be extended to other neurological disorders, additional imaging or molecular modalities, and longitudinal or prognostic modeling tasks.
Fig. 5. Subject-specific expert contributions.

Comparison of gating weights for CN (left), MCI (middle), and AD (right) patients. The color intensity is shown on the same scale as in Fig. 4
Acknowledgments
This work was supported in part by the NIH grants U01 AG066833, U01 AG068057, P30 AG073105, R01 AG068191, R01 AG07147, U19 AG074879, and R01 EB037101. The ADNI data sets were obtained from the Alzheimer’s Disease Neuroimaging Initiative (https://adni.loni.usc.edu), funded by NIH grant U01 AG024904. The authors declare no competing interests.
Footnotes
Data used in preparation of this article were obtained from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database (adni.loni.usc.edu). As such, the investigators within the ADNI contributed to the design and implementation of ADNI and/or provided data but did not participate in analysis or writing of this report. A complete listing of ADNI investigators can be found at: http://adni.loni.usc.edu/wp-content/uploads/how_to_apply/ADNI_Acknowledgement_List.pdf
References
- [1].Selkoe DJ, “Alzheimer’s disease: genes, proteins, and therapy,” Physiological Reviews, 2001. [DOI] [PubMed] [Google Scholar]
- [2].Jack CR, Knopman DS, Jagust WJ, Shaw LM, Aisen PS, Weiner MW et al. , “Hypothetical model of dynamic biomarkers of the alzheimer’s pathological cascade,” The Lancet Neurology, vol. 9, no. 1, pp. 119–128, 2010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [3].Jack C Jr, Bennett D, Blennow K, Carrillo M, Dunn B, Haeberlein S et al. , “Nia-aa research framework: Toward a biological definition of alzheimer’s disease,” Alzheimer’s & Dementia, vol. 14, no. 4, pp. 535–562, 2018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [4].Frisoni GB, Fox NC, Jack CR Jr, Scheltens P, and Thompson PM, “The clinical use of structural mri in alzheimer disease,” Nature Reviews Neurology, vol. 6, no. 2, pp. 67–77, 2010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [5].Jagust W, “Imaging the evolution and pathophysiology of alzheimer disease,” Nature Reviews Neuroscience, vol. 19, no. 11, pp. 687–700, 2018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [6].Zhang D, Wang Y, Zhou L, Yuan H, Shen D, A. D. N. Initiative et al. , “Multimodal classification of alzheimer’s disease and mild cognitive impairment,” NeuroImage, vol. 55, no. 3, pp. 856–867, 2011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [7].Petersen RC, Aisen PS, Beckett LA, Donohue MC, Gamst AC et al. , “Alzheimer’s disease neuroimaging initiative (adni) clinical characterization,” Neurology, vol. 74, no. 3, pp. 201–209, 2010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [8].Jagust WJ, Koeppe RA, Rabinovici GD, Villemagne VL, Harrison TM et al. , “The adni pet core at 20,” Alzheimer’s & Dementia, vol. 20, no. 10, pp. 7340–7349, 2024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [9].Jack CR Jr, Arani A, Borowski BJ, Cash DM, Crawford K et al. , “Overview of adni mri,” Alzheimer’s & Dementia, vol. 20, no. 10, pp. 7350–7360, 2024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [10].iu S, Liu S, Cai W, Che H, Pujol S, Kikinis R et al. , “Multimodal neuroimaging feature learning for multiclass diagnosis of alzheimer’s disease,” IEEE Transactions on Biomedical Engineering, vol. 62, no. 4, pp. 1132–1140, 2014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [11].Suk H-I, Lee S-W, Shen D, and A. D. N. Initiative, “Latent feature representation with stacked auto-encoder for ad/mci diagnosis,” Brain Structure and Function, vol. 220, no. 2, pp. 841–859, 2015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [12].Alsaleh A, “An interpretable multimodal deep learning framework for alzheimer’s disease diagnosis,” Digital Health, vol. 11, p. 20552076251390281, 2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [13].Rabinovici GD, Knopman DS, Arbizu J, Benzinger TL, Donohoe KJ et al. , “Updated appropriate use criteria for amyloid and tau pet: a report from the alzheimer’s association and society for nuclear medicine and molecular imaging workgroup,” Journal of Nuclear Medicine, vol. 66, no. Supplement 2, pp. S5–S31, 2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [14].Dill T, “Contraindications to magnetic resonance imaging,” Heart, vol. 94, no. 7, pp. 943–948, 2008. [DOI] [PubMed] [Google Scholar]
- [15].Johnson KA, Minoshima S, Bohnen NI, Donohoe KJ, Foster NL et al. , “Appropriate use criteria for amyloid pet: a report of the amyloid imaging task force, the society of nuclear medicine and molecular imaging, and the alzheimer’s association,” Alzheimer’s & Dementia, vol. 9, no. 1, pp. E1–E16, 2013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [16].Ma M, Ren J, Zhao L, Testuggine D, and Peng X, “Are multi-modal transformers robust to missing modality?” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 18 177–18 186. [Google Scholar]
- [17].Wu R, Wang H, Chen H-T, and Carneiro G, “Deep multimodal learning with missing modality: A survey,” arXiv preprint arXiv:2409.07825, 2024. [Google Scholar]
- [18].Abdelaziz M, Wang T, and Elazab A, “Alzheimer’s disease diagnosis framework from incomplete multimodal data using convolutional neural networks,” Journal of biomedical informatics, vol. 121, p. 103863, 2021. [DOI] [PubMed] [Google Scholar]
- [19].Stahlschmidt SR, Ulfenborg B, and Synnergren J, “Multimodal deep learning for biomedical data fusion: a review,” Briefings in bioinformatics, vol. 23, no. 2, p. bbab569, 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [20].Suk H-I, Lee S-W, Shen D, A. D. N. Initiative et al. , “Hierarchical feature representation and multimodal fusion with deep learning for ad/mci diagnosis,” NeuroImage, vol. 101, pp. 569–582, 2014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [21].Venugopalan J, Tong L, Hassanzadeh HR, and Wang MD, “Multimodal deep learning models for early detection of alzheimer’s disease stage,” Scientific reports, vol. 11, no. 1, p. 3254, 2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [22].Yi F, Yang H, Chen D, Qin Y, Han H et al. , “Xgboost-shap-based interpretable diagnostic framework for alzheimer’s disease,” BMC medical informatics and decision making, vol. 23, no. 1, p. 137, 2023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [23].Bao Y-W, Wang Z-J, Shea Y-F, Chiu PK-C, Kwan JS et al. , “Combined quantitative amyloid-β pet and structural mri features improve alzheimer’s disease classification in random forest model-a multicenter study,” Academic Radiology, vol. 31, no. 12, pp. 5154–5163, 2024. [DOI] [PubMed] [Google Scholar]
- [24].Guo AY, Laporte JP, Singh K, Bae J, Bergeron K et al. , “Machine learning diagnosis of mild cognitive impairment using advanced diffusion mri and csf biomarkers,” Alzheimer’s & Dementia: Diagnosis, Assessment & Disease Monitoring, vol. 17, no. 3, p. e70182, 2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [25].Li Q, Cui L, Guan Y, Li Y, Xie F, and Guo Q, “Prediction model and nomogram for amyloid positivity using clinical and mri features in individuals with subjective cognitive decline,” Human Brain Mapping, vol. 46, no. 8, p. e70238, 2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [26].Liu M, Cheng D, Wang K, Wang Y, and A. D. N. Initiative, “Multi-modality cascaded convolutional neural networks for alzheimer’s disease diagnosis,” Neuroinformatics, vol. 16, no. 3, pp. 295–308, 2018. [DOI] [PubMed] [Google Scholar]
- [27].Meng L and Zhang Q, “Research on early diagnosis of alzheimer’s disease based on dual fusion cluster graph convolutional network,” Biomedical Signal Processing and Control, vol. 86, p. 105212, 2023. [Google Scholar]
- [28].Klepl D, He F, Wu M, Blackburn DJ, and Sarrigiannis P, “Eeg-based graph neural network classification of alzheimer’s disease: An empirical evaluation of functional connectivity methods,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol. 30, pp. 2651–2660, 2022. [DOI] [PubMed] [Google Scholar]
- [29].Zhou H, He L, Zhang Y, Shen L, and Chen B, “Interpretable graph convolutional network of multi-modality brain imaging for alzheimer’s disease diagnosis,” in 2022 IEEE 19th International Symposium on Biomedical Imaging (ISBI). IEEE, 2022, pp. 1–5. [Google Scholar]
- [30].Zhou H, He L, Chen BY, Shen L, and Zhang Y, “Multi-modal diagnosis of alzheimer’s disease using interpretable graph convolutional networks,” IEEE Transactions on Medical Imaging, 2024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [31].Li J, Xu H, Yu H, Jiang Z, and Zhu L, “Multi-modal feature selection with anchor graph for alzheimer’s disease,” Frontiers in Neuroscience, vol. 16, p. 1036244, 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [32].Bahdanau D, “Neural machine translation by jointly learning to align and translate,” arXiv preprint arXiv:1409.0473, 2014. [Google Scholar]
- [33].Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN et al. , “Attention is all you need,” Advances in Neural Information Processing Systems, vol. 30, 2017. [Google Scholar]
- [34].Tsai Y-HH, Bai S, Liang PP, Kolter JZ, Morency L-P, and Salakhutdinov R, “Multimodal transformer for unaligned multimodal language sequences,” in Proceedings of the Conference of the Association for Computational Linguistics, vol. 2019, 2019, p. 6558. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [35].Hollmann N, Müller S, Eggensperger K, and Hutter F, “Tabpfn: A transformer that solves small tabular classification problems in a second,” arXiv preprint arXiv:2207.01848, 2022. [Google Scholar]
- [36].Hollmann N, Müller S, Purucker L, Krishnakumar A, Körfer M et al. , “Accurate predictions on small data with a tabular foundation model,” Nature, vol. 637, no. 8045, pp. 319–326, 2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [37].Gorishniy Y, Rubachev I, Khrulkov V, and Babenko A, “Revisiting deep learning models for tabular data,” Advances in neural information processing systems, vol. 34, pp. 18 932–18 943, 2021. [Google Scholar]
- [38].Lundberg SM and Lee S-I, “A unified approach to interpreting model predictions,” Advances in neural information processing systems, vol. 30, 2017. [Google Scholar]
- [39].Vahid A, Zamani J, and Hadi Hosseini S, “Predicting future amyloid conversion: the role of baseline amyloid distribution, genetic and cognitive resilience,” npj Dementia, vol. 1, no. 1, p. 40, 2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [40].Yun S, Choi I, Peng J, Wu Y, Bao J et al. , “Flex-moe: Modeling arbitrary modality combination via the flexible mixture-of-experts,” Advances in Neural Information Processing Systems, vol. 37, pp. 98 782–98 805, 2024. [Google Scholar]
- [41].Jiang Y, Ding H, Chen H, Lan J, Teng X et al. , “M ˆ3 ad: Multi-task multi-gate mixture of experts for alzheimer’s disease diagnosis with conversion pattern modeling,” arXiv preprint arXiv:2508.01819, 2025. [Google Scholar]
- [42].Xin J, Yun S, Peng J, Choi I, Ballard JL et al. , “I2moe: Interpretable multimodal interaction-aware mixture-of-experts,” arXiv preprint arXiv:2505.19190, 2025. [PMC free article] [PubMed] [Google Scholar]
- [43].Ding R, Lu H, and Liu M, “Denseformer-moe: A dense transformer foundation model with mixture of experts for multi-task brain image analysis,” IEEE Transactions on Medical Imaging, 2025. [DOI] [PubMed] [Google Scholar]
- [44].Jiang Y and Shen Y, “M4oe: A foundation model for medical multimodal image segmentation with mixture of experts,” in international conference on medical image computing and computer-assisted intervention. Springer, 2024, pp. 621–631. [Google Scholar]
- [45].Li J, Zhang Y, Shu W, Feng X, Wang Y et al. , “M4: Multi-proxy multi-gate mixture of experts network for multiple instance learning in histopathology image analysis,” Medical Image Analysis, vol. 103, p. 103561, 2025. [DOI] [PubMed] [Google Scholar]
- [46].Weiner MW, Veitch DP, Aisen PS, Beckett LA, Cairns NJ et al. , “Recent publications from the alzheimer’s disease neuroimaging initiative: Reviewing progress toward improved ad clinical trials,” Alzheimer’s & Dementia, vol. 13, no. 4, pp. e1–e85, 2017. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [47].Fischl B, “Freesurfer,” NeuroImage, vol. 62, no. 2, pp. 774–781, 2012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [48].Desikan RS, Ségonne F, Fischl B, Quinn BT, Dickerson BC et al. , “An automated labeling system for subdividing the human cerebral cortex on mri scans into gyral based regions of interest,” NeuroImage, vol. 31, no. 3, pp. 968–980, 2006. [DOI] [PubMed] [Google Scholar]
- [49].Akiba T, Sano S, Yanase T, Ohta T, and Koyama M, “Optuna: A next-generation hyperparameter optimization framework,” in Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining, 2019, pp. 2623–2631. [Google Scholar]
- [50].Abulibdeh R, Tu K, and Sejdić E, “Balancing model complexity and clinical deployability in deep learning for sociodemographic information extraction,” Journal of Primary Care & Community Health, vol. 16, p. 21501319251404193, 2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [51].Sheikh F, Marouf AA, Rokne JG, and Alhajj R, “Lightweight deep learning models with explainable ai for early alzheimer’s detection from standard mri scans,” Diagnostics, vol. 15, no. 21, p. 2709, 2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [52].Cabanillas-Carbonell M and Zapata-Paulini J, “Evaluation of machine learning models for the prediction of alzheimer’s: In search of the best performance,” Brain, Behavior, & Immunity-Health, vol. 44, p. 100957, 2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [53].Cary MP Jr, Zhuang F, Draelos RL, Pan W, Amarasekara S et al. , “Machine learning algorithms to predict mortality and allocate palliative care for older patients with hip fracture,” Journal of the American Medical Directors Association, vol. 22, no. 2, pp. 291–296, 2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [54].Liu Y, Zhang Y, Wang Y, Hou F, Yuan J et al. , “A survey of visual transformers,” IEEE transactions on neural networks and learning systems, vol. 35, no. 6, pp. 7478–7498, 2023. [DOI] [PubMed] [Google Scholar]
