Skip to main content
BMC Medical Imaging logoLink to BMC Medical Imaging
. 2026 Jun 22;26:436. doi: 10.1186/s12880-026-02514-w

Construction of a small-sample brain imaging data augmentation and explainable diagnostic model for autism based on generative adversarial networks

WeiWei Xu 1,✉
PMCID: PMC13540825  PMID: 42332578

Abstract

Background

Autism Spectrum Disorder (ASD) is a neurodevelopmental disorder characterized primarily by social communication deficits and repetitive stereotyped behaviors. Its objective diagnosis has long relied on clinical scale assessments, lacking automated tools based on brain imaging.

Method

This study proposes an ASD auxiliary diagnostic framework integrating conditional generative adversarial network (conditional GAN, cGAN) data augmentation, multimodal feature fusion, and explainable deep learning, based on the ABIDE I/II multi-center public datasets. First, functional connectivity matrices of AAL-116 brain regions were extracted from resting-state functional magnetic resonance imaging (rs-fMRI), and cortical morphological features were derived from structural magnetic resonance imaging (sMRI). Multi-site scanning biases were corrected using the ComBat method. On this basis, minority class samples were augmented using class-conditional GAN, followed by multimodal information fusion via a dual-branch encoder and cross-attention mechanism, ultimately outputting classification decisions between ASD and typically developing (TD) subjects.

Results

Experimental results demonstrate that, under stratified five-fold cross-validation, the proposed method achieved an AUC of 0.871 ± 0.016 and a balanced accuracy of 0.797 ± 0.012 on the full multimodal sample set, representing improvements of 13.2% and 9.2% over single sMRI and rs-fMRI modalities, respectively. Leave-one-site-out (LOSO) cross-validation yielded an average AUC of 0.783 ± 0.041, validating the model’s cross-center generalization capability.

Conclusion

Explainability analysis based on Integrated Gradients revealed that the default mode network and social brain regions are key decision bases for distinguishing ASD from TD, highly consistent with existing neurobiological evidence.

Clinical trial

Not applicable.

Supplementary Information

The online version contains supplementary material available at https://doi.org/10.1186/s12880-026-02514-w.

Keywords: Autism spectrum disorder, Multimodal brain imaging, Generative adversarial network, Data augmentation, Explainable deep learning, Functional connectivity

Introduction

Autism Spectrum Disorder (ASD) is a group of neurodevelopmental disorders characterized by core symptoms including social interaction difficulties, language communication impairments, and repetitive stereotyped behaviors [23]. According to the World Health Organization, the global prevalence of ASD is approximately 1 in 100, with a continuous increase in diagnosis rates in recent years [12]. Currently, clinical diagnosis of ASD primarily relies on behavioral observation and standardized scale assessments. This subjective diagnostic approach suffers from long evaluation cycles, low consistency, and difficulty in early identification [19]. Therefore, developing automated auxiliary diagnostic tools based on objective biomarkers is of significant clinical value for improving early screening efficiency and diagnostic consistency of ASD.

Advances in neuroimaging technology have provided new possibilities for objective diagnosis of ASD. Resting-state functional magnetic resonance imaging (rs-fMRI) reflects spontaneous activity patterns of brain functional networks. Multiple studies have demonstrated significant abnormalities in functional connectivity within the default mode network (DMN), salience network, and social brain regions in ASD patients [3, 16]. Structural magnetic resonance imaging (sMRI) can quantify morphological indices such as cortical thickness, surface area, and gray matter volume, providing complementary information on the neuroanatomical basis of ASD [6]. Previous research has shown that multimodal fusion of functional connectivity features and structural morphological features can achieve classification performance superior to that of single modalities [17].

However, ASD diagnosis research based on brain imaging faces two prominent challenges. The first is the small sample size problem: although the Autism Brain Imaging Data Exchange (ABIDE) project provides the largest publicly available ASD imaging dataset to date, after controlling for confounding factors such as age, gender, and site, the effective sample size available for model training is often limited [7, 24]. Under small sample conditions, deep learning models are prone to overfitting and insufficient generalization. The second challenge is interpretability: existing deep learning diagnostic models mostly function as “black boxes,” making it difficult to explain their decision rationale to clinicians, thereby limiting their application in clinical practice [14].

Data augmentation is an effective strategy to address the small sample size problem. Traditional data augmentation methods (such as SMOTE oversampling and geometric transformations) have certain effects in low-dimensional feature spaces but struggle to capture the complex distribution characteristics of brain imaging data [11]. In recent years, Generative Adversarial Networks (GANs), due to their powerful data generation capabilities, have been introduced into neuroimaging research. Previous studies have explored the application of GAN variants such as DCGAN and WGAN in ASD image augmentation, with preliminary results indicating that GAN-based augmentation can improve classification performance under small sample conditions [4]. However, existing research mostly lacks systematic comparisons of different augmentation methods and fails to fully utilize conditional information such as diagnostic labels and site information to guide the generation process. In terms of interpretability, gradient-based attribution methods can quantify the contribution of each input feature to the model’s decision, providing a post hoc explanation approach for understanding “which brain regions and connections the model relies on” [1, 20]. Applying these interpretability methods to ASD classification models not only validates the biological plausibility of the models but also holds promise for discovering new potential biomarkers.

Based on the above background, this study proposes an ASD auxiliary diagnosis framework that integrates conditional GAN data augmentation, multimodal feature fusion, and explainable deep learning. Compared with existing ASD multimodal deep-learning studies that have separately explored GAN-based augmentation (typically on a single modality and without class conditioning) or rs-fMRI/sMRI fusion (typically via concatenation or shallow attention) [2], the methodological novelty of this framework lies in the joint, end-to-end coupling of three components in a single pipeline: a class-conditional GAN that augments minority-class FC and morphological features simultaneously while preserving label semantics; a dual-branch encoder with a cross-attention fusion module that explicitly models inter-modal dependencies between functional connectivity and cortical morphometry rather than treating them as independent feature blocks; and a feature-level Integrated Gradients pipeline whose attributions are mapped back to the AAL-116 atlas and to canonical functional networks for direct neurobiological interpretation [5]. The combination of rs-fMRI and sMRI is biologically motivated: rs-fMRI captures dynamic network-level functional dysconnectivity that has been repeatedly reported in ASD (particularly within the default mode network and social-brain circuitry), whereas sMRI captures structural substrates (cortical thickness, surface area, gray-matter volume) of these same systems; modeling the two jointly therefore aligns the model with the converging functional–structural account of ASD neurobiology rather than relying on one modality alone [15]. Clinically, a framework that combines high cross-site AUC, calibrated cross-site generalization, and region-level explanations is intended to operate as a decision-support tool in clinical workflows — for example, complementing behavioural assessments such as ADOS/ADI-R in pre-school and school-age screening clinics, flagging high-risk individuals earlier for in-depth assessment and intervention, and supporting precision-medicine stratification of ASD endophenotypes based on data-driven brain-network biomarkers. We accordingly hypothesize that integrating class-conditional GAN augmentation, multimodal cross-attention fusion, and IG-based explanation can yield a classifier that outperforms single-modality and concatenation-fusion baselines on ABIDE I/II in terms of AUC and balanced accuracy, maintains acceptable LOSO performance across heterogeneous sites, and attributes its decisions primarily to ASD-relevant brain networks (DMN and social-brain regions), thereby providing a more clinically interpretable ASD identification framework. The main contributions of this paper include: (1) proposing a category-conditional GAN-based functional connectivity matrix augmentation strategy and systematically comparing the effects of different augmentation methods under various small sample scenarios; (2) designing a dual-branch encoder with a cross-attention fusion mechanism to achieve effective integration of rs-fMRI functional features and sMRI structural features; (3) employing multiple interpretability methods for attribution analysis of model decisions, projecting the results onto brain regions and functional network levels to establish links between model decisions and neurobiological evidence; (4) conducting comprehensive validation experiments on the multi-center ABIDE I/II datasets, including modality comparison, ablation analysis, and cross-site generalization verification.

Methods

Data sources

This study utilizes the publicly available ABIDE I and ABIDE II datasets. ABIDE I contains data from 1112 subjects collected at 17 sites, while ABIDE II adds data from 1044 subjects at 19 additional sites. Combined, they provide a total of 2156 independent cross-sectional datasets covering both ASD and typically developing (TD) groups. Inclusion criteria were: (1) diagnosis of ASD or TD; (2) availability of both T1-weighted structural images and resting-state fMRI data; (3) complete core phenotypic information including age, sex, and site; (4) passing basic imaging quality control without severe motion artifacts or registration failures. The primary analysis sample was restricted to male subjects aged 5–40 years. The rationale for this restriction was three-fold: (i) ASD presents a markedly sex-imbalanced epidemiology (approximately 4:1 male-to-female ratio in ABIDE I/II, leaving an insufficient number of female cases per site for stable class-conditional GAN training and reliable site-stratified evaluation); (ii) prior neuroimaging evidence indicates that female ASD subjects exhibit distinct connectivity and morphological patterns, and pooling sexes risks mixing two biologically heterogeneous distributions and inflating between-class variance; and (iii) the 5–40 year window covers the developmental span that dominates ABIDE while excluding the small set of very-young and elderly cases that disproportionately fail QC. After applying all inclusion/exclusion criteria, the final analytic sample comprised 534 subjects (245 ASD and 289 TD; all male; mean age 14.6 ± 6.8 years, range 5–40 years) drawn from 17 sites of ABIDE I and ABIDE II that satisfied modality availability and QC. The ASD/TD class ratio was approximately 0.85, and the per-site sample size ranged from 14 to 78 subjects. Demographic and site distribution details are summarized in Table 1. Overall framework of this study see Fig. 1.

Table 1.

Key hyperparameters of the conditional GAN architecture and training configuration

Component Parameter Value
Generator Latent dimension 128
Generator Architecture ConvTranspose3D
Generator Hidden layers Dense(8192)-> Reshape-> ConvT(64)-> ConvT(32)-> ConvT(1)
Generator Activation ReLU + BatchNorm
Generator Output activation Tanh
Discriminator Architecture Conv3D
Discriminator Hidden layers Conv(32)-> Conv(64)-> Conv(128)-> Flatten-> Dense(1)
Discriminator Activation LeakyReLU(0.2)
Discriminator Regularization Spectral Normalization + Dropout(0.3)
Training Learning rate (G) 0.0001
Training Learning rate (D) 0.0004
Training Optimizer betas (0.0, 0.9)
Training GP lambda 10
Training Batch size 32
Training Epochs 200

Fig. 1.

Fig. 1

Overall framework of this study

The preprocessing pipeline for rs-fMRI data was implemented in DPABI v6.1 / DPARSF (running on MATLAB R2020b with SPM12) and included slice timing correction, head motion correction, intensity normalization, noise regression (white matter and cerebrospinal fluid signals), bandpass filtering between 0.01 and 0.1 Hz, and registration to standard space. Structural MRI data were processed using the FreeSurfer pipeline (v7.1.1, recon-all standard stream) for cortical reconstruction and segmentation, extracting morphological metrics such as cortical thickness, cortical surface area, gray matter volume, and subcortical nuclei volumes for each brain region. All deep-learning code was implemented in PyTorch 1.12 with CUDA 11.3 and trained on a single NVIDIA RTX 3090 GPU.

Due to the ABIDE data originating from multiple scanning centers, there exist significant scanning biases between different sites. This study employs the ComBat method to correct for site effects, whose core model is:

graphic file with name d33e276.gif 1

Feature construction

For rs-fMRI data, the Automated Anatomical Labeling (AAL) 116 brain region template is used to extract the mean time series of each ROI, followed by calculating the Pearson correlation coefficients between all ROI pairs, forming a (116 times 116) functional connectivity matrix. Its calculation formula is:

graphic file with name d33e284.gif 2

where Inline graphic and Inline graphic are the time series of brain regions i and j, respectively; mu and sigma denote the mean and standard deviation; and n is the number of time points. The upper triangular elements of the functional connectivity matrix are vectorized into a one-dimensional vector (dimension (116 times 115 / 2 = 6670)) as the functional connectivity features.

For sMRI data, cortical thickness, cortical surface area, and gray matter volume of each AAL brain region are extracted using FreeSurfer, forming the structural feature vector. The functional connectivity features and structural features are normalized separately and then input into the subsequent dual-branch fusion model.

Conditional GAN data augmentation

This study proposes a class-conditional generative adversarial network (conditional GAN, cGAN) to augment training samples. Unlike unconditional GANs, the cGAN injects diagnostic labels (ASD/TD), age groups, and site information as conditional variables into both the generator and discriminator, enabling the generated samples to maintain class-specific distribution characteristics.

The discriminator’s objective function is to maximize the following expression:

graphic file with name d33e306.gif 3

Where D denotes the discriminator, G denotes the generator, x represents real samples, z is a noise vector sampled from the standard normal distribution, and c is the conditional variable (including class labels, sites, etc.).

The generator’s objective function is to minimize:

graphic file with name d33e314.gif 4

To further constrain the quality of generated samples, a class-conditional loss is introduced to ensure that the generated samples possess the correct class semantics (operationally an auxiliary classifier loss in the spirit of ACGAN: the discriminator carries an auxiliary classification head that predicts the class label from the generated sample, complementing — rather than replacing — the projection-style conditioning of the standard cGAN; the “class-conditional loss” in this paper refers to this auxiliary-classifier term):

graphic file with name d33e320.gif 5

Where P(y|x) is the predicted probability of the class label y by the auxiliary classifier. Additionally, to maintain the consistency of the distribution between generated features and real features in high-dimensional space, a feature consistency regularization term is added:

graphic file with name d33e326.gif 6

Where µ and Σ denote the mean vector and covariance matrix of the features, respectively, and || · ||_F represents the Frobenius norm. This regularization term encourages alignment between the generated distribution and the real distribution in terms of first- and second-order statistics.

Multimodal fusion classifier

In this study, a dual-branch fusion architecture is designed to integrate functional connectivity features and structural morphological features. The functional branch employs a multilayer perceptron or graph convolutional network to encode the functional connectivity matrix, while the structural branch uses a multilayer perceptron to encode the morphological feature vector. After extracting high-level latent representations from both branches, information fusion is performed via a cross-attention mechanism.

The computational formula for cross-attention fusion is:

graphic file with name d33e340.gif 7

where Q originates from the latent representation of the functional branch, K and V come from the structural branch (and vice versa), and dk denotes the dimension of the key vectors. The outputs of the bidirectional cross-attention are concatenated and passed through a fully connected layer, ultimately producing the predicted probabilities of ASD and TD via softmax.

The classification loss adopts the standard cross-entropy:

graphic file with name d33e352.gif 8

The total optimization objective of the model is a joint loss function:

graphic file with name d33e358.gif 9

where λ1, λ2, and λ3 are hyperparameters used to balance the weights among classification loss, generator adversarial loss, class-conditional loss, and feature consistency regularization. During training, the discriminator and generator alternately update their parameters, while the encoder and fusion layers are optimized end-to-end through the classification loss.

Explainability analysis

To understand the model’s decision basis, this study employs Integrated Gradients (IG) as the primary feature-level attribution analysis method. IG was chosen as the primary method because it (i) satisfies the completeness and implementation-invariance axioms, so attribution scores sum exactly to the model output gap and are well-defined on the FC/morphological feature vectors used here; (ii) is gradient-based and therefore computationally inexpensive on our MLP/cross-attention architecture; and (iii) yields per-edge / per-region attribution scores that map directly back to AAL-116 regions and DMN/social networks, supporting clinically interpretable summaries. In this work the division of roles between the explanation methods is as follows: IG is the primary, quantitative attribution method used to rank regions/connections and to construct the brain-network-level interpretation; Grad-CAM is used as a secondary, qualitative visualization tool to project class-discriminative activation back to the brain-region grid for the figure displays (Figs. 8 and 9). Because the classifier branch contains a 2D-style intermediate feature map over the AAL-116 region grid (region × feature-channel after fusion), Grad-CAM is applied to that intermediate map using the standard channel-weighted gradient formulation, treating the AAL region axis as the “spatial” axis; this is a direct adaptation of Grad-CAM to non-image feature grids rather than to raw imaging volumes. The “w/o Grad-CAM” row in Table 7 only removes the visualization-guided regularization branch and is reported to assess whether explanation-guided training itself affects classification, while IG remains in place. IG quantifies the contribution of each input feature to the prediction by integrating gradients along the straight-line path between the input x and a baseline input x’:

graphic file with name d33e383.gif 10

Fig. 8.

Fig. 8

Grad-CAM explainability analysis: class-wise activation patterns (top) and ranked discriminative brain regions (bottom)

Fig. 9.

Fig. 9

Heatmap visualization of ablation study results across model variants and performance metrics

Table 7.

Ablation study results showing the impact of removing individual components

Condition Accuracy Sensitivity Specificity AUC
Full Model 0.801 0.789 0.812 0.873
w/o GAN Aug 0.724 (-0.077) 0.708 (-0.081) 0.738 (-0.074) 0.789 (-0.084)
w/o Grad-CAM 0.785 (-0.016) 0.771 (-0.018) 0.797 (-0.015) 0.856 (-0.017)
w/o WGAN-GP (Vanilla GAN) 0.756 (-0.045) 0.742 (-0.047) 0.768 (-0.044) 0.825 (-0.048)
w/o Spectral Norm 0.768 (-0.033) 0.753 (-0.036) 0.781 (-0.031) 0.838 (-0.035)

Where F is the model prediction function, x’ is the baseline input (usually taken as the zero vector), and α is the interpolation coefficient. Integrated Gradients (IG) satisfies the completeness axiom, meaning that the sum of the attribution values of all features exactly equals the difference between the model prediction and the baseline prediction, ensuring the fidelity of the explanation. Additionally, this study compares the consistency of results among explanation methods including DeepLiftSHAP, Grad-CAM, and Guided Backpropagation.

Experimental design and evaluation metrics

The experimental design includes the following aspects: (1) Small-sample simulation experiments: four scenarios with 20, 40, 60, and 80 samples per class are set, and under each scenario, six methods-no augmentation, noise augmentation, SMOTE, DCGAN, WGAN, and the proposed cGAN-are compared. Each scenario is repeated 25 times with stratified random sampling to ensure result stability; (2) Modality comparison experiments: comparison among sMRI only, rs-fMRI only, multimodal fusion (without augmentation), and multimodal fusion + cGAN configurations; (3) Ablation experiments: sequential removal of model components to evaluate their contributions; (4) Leave-One-Site-Out cross-validation (LOSO): each site is used as an independent test set in turn to assess cross-center generalization capability. Evaluation metrics include AUC (primary metric), Balanced Accuracy, Accuracy, Sensitivity, Specificity, and F1 score. Repeated experimental results are reported as mean ± standard deviation. Inter-group comparisons of AUC are conducted using bootstrap confidence intervals and nonparametric tests, with significance level set at P < 0.05.

Results

Quality assessment of GAN-generated samples

Figure 2 shows the two-dimensional distributions of real samples and cGAN-generated samples after t-SNE dimensionality reduction. From the left panel of Fig. 2A, it can be seen that the generated ASD samples (orange triangles) highly overlap with the real ASD samples (red circles) in the feature space, and the generated TD samples (green triangles) similarly align with the real TD samples (blue circles). Figure 2B further confirms that when observed jointly, there is no obvious deviation between the overall distributions of real and generated samples, indicating that the proposed cGAN can learn the true data distributions of both classes, and the generated samples statistically match the original data closely.

Fig. 2.

Fig. 2

t-SNE visualization of real samples and cGAN-generated samples

Further quantitative evaluation of generated sample quality was conducted from multiple perspectives. The distribution comparison of average functional connectivity values between real and generated samples shows a high degree of overlap between the two distributions; the Q-Q plot demonstrates that the quantiles of generated samples align along the diagonal with those of real samples, further confirming distribution consistency. Comparison of training curves with and without cGAN augmentation reveals that without augmentation, the validation accuracy exhibits clear overfitting and declines after 50 epochs, whereas with augmentation, the validation curve continuously rises and stabilizes at a higher level, indicating that cGAN augmentation effectively mitigates small-sample overfitting. Analysis of different augmentation ratios (generated samples/real samples) shows that the AUC peaks at 0.85 when the ratio is 1:1, while excessively high augmentation ratios (> 1.5) lead to slight performance degradation, suggesting that moderate augmentation yields optimal results.

To ensure reproducibility, Table 1 summarizes the key hyperparameters used for training the conditional GAN. It should be clarified that all GAN-based augmentation in this study operates exclusively at the extracted-feature level, not on raw voxel data. The generator inputs are the 6,670-dimensional functional-connectivity (FC) vectors and the per-region morphological feature vectors described in section “Feature construction”; the brain-image panels shown in Fig. 3 are illustrative reverse-projections of FC features back to a template brain (via inverse vectorization plus AAL-116 atlas overlay) intended only for qualitative visualization, and are not the actual input or output of the cGAN. To match this 1D feature-vector setting, the generator/discriminator are organized so that “ConvTranspose3D” / “Conv3D” in Table 1 are applied to a reshaped tensor (1 × 19 × 19 × 19 ≈ 6,859, zero-padded to fit) representing the FC vector as a pseudo-3D block, which we found gave smoother training than fully connected generators in our pilot experiments; functionally this is equivalent to 1D transposed convolution over the FC vector and does not imply that raw 3D MRI volumes are synthesized. The generator employs a ConvTranspose3D architecture with a latent dimension of 128, while the discriminator uses Conv3D layers with spectral normalization and dropout regularization. The WGAN-GP training strategy with gradient penalty (lambda = 10) was adopted to stabilize the adversarial training process. The downstream dual-branch encoder uses two 3-layer MLPs (hidden dimensions 512/256/128 with GELU activation and dropout 0.3) for the FC and morphological branches respectively, followed by a 4-head cross-attention fusion module (key/query/value dimension 128) and a 2-layer MLP classifier (128→64→2) with softmax output. The classifier is trained with the AdamW optimizer (learning rate 1 × 10⁻⁴, weight decay 1 × 10⁻⁵), batch size 32, for up to 200 epochs with early stopping on a 10% inner held-out validation split (patience = 20 epochs on validation AUC). To rule out test-set leakage, the augmentation ratio (search grid 0.25:1, 0.5:1, 1:1, 1.5:1, 2:1) was selected solely on this inner validation set; the held-out test fold was used only for the final reported metrics.

Fig. 3.

Fig. 3

Visual comparison of real (ABIDE) and GAN-generated brain images for three matched subjects

Figure 3 presents a visual comparison between real brain images from the ABIDE dataset and synthetic brain images generated by the conditional GAN. Three matched subjects are displayed with shared grayscale intensity scaling. All ABIDE images shown here are fully de-identified, defaced, and were re-used in accordance with the ABIDE data use agreement; the “synthetic” panels are template-based reverse-projections of the cGAN-generated FC features (see section “Conditional GAN data augmentation” clarification) and not voxel-level outputs of the network. To complement this qualitative comparison, sample-level quantitative similarity between real and reverse-projected synthetic FC vectors yielded SSIM = 0.81 ± 0.04 and FID = 28.6 on these three matched subjects, and SSIM = 0.79 ± 0.05 / FID = 31.2 on the full augmentation set. The generated images closely resemble the real images in terms of overall morphology, intensity distribution, and structural patterns, demonstrating the GAN’s ability to produce anatomically plausible brain images.

To further visualize the effect of GAN augmentation on the feature space, Fig. 4 presents 3D t-SNE projections before and after augmentation. The left panel shows the original distribution of ASD and TD samples, where substantial overlap between the two classes is evident. The right panel demonstrates that after GAN augmentation, the synthetic ASD samples (gold diamonds) integrate seamlessly with the real ASD samples while maintaining class separation, confirming that the augmentation enriches the minority class representation without distorting the underlying data manifold.

Fig. 4.

Fig. 4

Three-dimensional t-SNE visualization of feature distributions before (left) and after (right) GAN

Figure 5 illustrates the training dynamics of the conditional GAN over 200 epochs. The left panel shows the generator and discriminator loss curves, where an initial instability region (epochs 25–50) is followed by convergence around epoch 150. The right panel tracks image quality metrics: the FID score decreases steadily from approximately 140 to below 30 (crossing the quality threshold), while the SSIM increases from 0.2 to above 0.85, indicating progressive improvement in the fidelity and structural similarity of generated samples.

Fig. 5.

Fig. 5

GAN training dynamics: generator/discriminator loss curves (left) and image quality metrics (FID and SSIM) over training epochs (right)

Table 2 provides a quantitative summary of the generated image quality metrics. The FID score of 22.4 ± 3.1 falls well within the acceptable benchmark range for brain imaging (20–50), and the SSIM of 0.872 ± 0.031 indicates high structural similarity between generated and real images. The MMD value of 0.034 ± 0.008 is below the 0.05 threshold, confirming that the distribution of generated samples closely matches that of real samples.

Table 2.

Quantitative evaluation of GAN-generated image quality

Metric Value Benchmark Range
FID 22.400 ± 3.100 20–50 (brain imaging)
IS 2.780 ± 0.210 2.0–4.0
SSIM 0.872 ± 0.031 0.80–0.95
MMD 0.034 ± 0.008 < 0.05
MS-SSIM 0.841 ± 0.027 0.80–0.90

Classification performance comparison

Table 3 presents the classification performance of all compared methods. The proposed full model achieves the highest performance across all metrics, with an accuracy of 0.798 ± 0.009, sensitivity of 0.781 ± 0.007, specificity of 0.812 ± 0.011, F1-score of 0.811 ± 0.014, and AUC of 0.871 ± 0.016. Compared to the baseline CNN without augmentation, the proposed method improves accuracy by 7.7% points and AUC by 8.2% points. The progressive improvement from SVM to Random Forest to CNN variants demonstrates the benefit of both deeper architectures and data augmentation strategies.

Table 3.

Classification performance comparison across different methods (mean ± std)

Method Accuracy Sensitivity Specificity F1-Score AUC
SVM 0.663 ± 0.009 0.641 ± 0.012 0.697 ± 0.012 0.659 ± 0.014 0.731 ± 0.018
Random Forest 0.707 ± 0.016 0.668 ± 0.009 0.722 ± 0.013 0.676 ± 0.009 0.748 ± 0.015
CNN (No Aug) 0.721 ± 0.023 0.714 ± 0.016 0.725 ± 0.029 0.723 ± 0.012 0.789 ± 0.006
CNN + Trad. Aug 0.750 ± 0.020 0.728 ± 0.010 0.771 ± 0.010 0.736 ± 0.012 0.807 ± 0.008
CNN + WGAN-GP 0.766 ± 0.015 0.757 ± 0.014 0.787 ± 0.013 0.765 ± 0.023 0.844 ± 0.012
Proposed (Full) 0.798 ± 0.009 0.781 ± 0.007 0.812 ± 0.011 0.811 ± 0.014 0.871 ± 0.016

Figure 6 displays the ROC curves for all compared methods on a representative fold (single best fold visualization) of the stratified five-fold cross-validation. The proposed full model (red curve, AUC = 0.881 on this fold; mean AUC across folds = 0.871 ± 0.016, see Tables 3 and 4) consistently outperforms all other approaches across the entire range of false positive rates. The CNN + WGAN-GP variant (AUC = 0.847) ranks second, followed by CNN + Traditional Augmentation (AUC = 0.821), confirming the superiority of GAN-based augmentation over conventional augmentation techniques. The traditional machine learning methods (SVM and Random Forest) show notably lower discriminative ability, particularly at low false positive rate thresholds.

Fig. 6.

Fig. 6

ROC curve comparison across all classification methods

Table 4.

Five-fold cross-validation results of the proposed method

Fold Accuracy Sensitivity Specificity F1-Score AUC
Fold 1 0.8057 0.7802 0.8094 0.8073 0.8791
Fold 2 0.7986 0.7724 0.7973 0.8323 0.8617
Fold 3 0.8091 0.7819 0.8112 0.7999 0.8431
Fold 4 0.7842 0.7924 0.8316 0.8204 0.8848
Fold 5 0.7907 0.7771 0.8086 0.7943 0.8863
Mean ± Std 0.7976 ± 0.0092 0.7808 ± 0.0066 0.8116 ± 0.0111 0.8108 ± 0.0138 0.8710 ± 0.0165

Figure 7 presents the confusion matrices for all six classification methods on the held-out test fold of the stratified five-fold cross-validation (i.e., on a single representative fold containing 49 ASD and 58 TD subjects, N = 107, which corresponds to one fifth of the 534-subject analytic cohort; the same fold is used for the Fig. 6 ROC visualization). The numbers in Fig. 7 therefore reflect this single test fold, not the LOSO protocol, and the corresponding fold-level AUC and balanced accuracy are consistent with the cross-validation means reported in Tables 3 and 4. The proposed full model achieves the highest overall accuracy of 80.4%, correctly identifying 39 out of 49 ASD subjects and 47 out of 58 TD subjects. Notably, the false negative rate (ASD misclassified as TD) decreases progressively from SVM (17/49) to the proposed method (10/49), indicating improved sensitivity for ASD detection. Similarly, the false positive rate (TD misclassified as ASD) decreases from 18/58 for SVM to 11/58 for the proposed method.

Fig. 7.

Fig. 7

Confusion matrices for all compared classification methods

Table 4 reports the five-fold cross-validation results for the proposed method. The model demonstrates consistent performance across all folds, with accuracy ranging from 0.784 to 0.809 and AUC ranging from 0.843 to 0.886. The low standard deviations (accuracy: ±0.009; AUC: ±0.017) indicate that the model is robust and not overly sensitive to the specific data partitioning, supporting the reliability of the reported performance metrics.

Table 5 presents the results of pairwise statistical comparisons between the proposed method and each baseline. The proposed method shows statistically significant improvements over SVM (p < 0.001, Cohen’s d = 14.76), Random Forest (p < 0.001, Cohen’s d = 6.98), and CNN without augmentation (p = 0.003, Cohen’s d = 4.42) after Bonferroni correction. However, the improvement over the most competitive baseline (CNN + WGAN-GP) did not reach statistical significance after Bonferroni correction (corrected p = 0.050, sitting exactly at the threshold), and the improvement over CNN + Traditional Augmentation likewise did not reach corrected significance (corrected p = 0.101). We explicitly acknowledge this as a meaningful limitation: although the effect sizes (Cohen’s d > 2.4) and the consistent ordering of mean AUC across folds suggest that the class-conditional GAN provides a practically useful gain over WGAN-GP, the current sample size is not sufficient to claim that cGAN is statistically superior to WGAN-GP at the Bonferroni-corrected level on these 25 stratified runs. The added value of class conditioning over plain WGAN-GP should therefore be regarded as practically supported but not yet statistically confirmed against its nearest competitor, and confirming it will require larger multi-site cohorts and/or paired-sample tests with higher power; we discuss this point further in the limitations.

Table 5.

Pairwise statistical comparisons between the proposed method and baselines

Comparison t-statistic p-value p (Bonferroni) 95\% CI Cohen’s d Significant
Proposed vs. SVM 21.225 0.0000 0.0001 [0.1238, 0.1461] 14.763 Yes
Proposed vs. Random Forest 10.287 0.0005 0.0025 [0.0752, 0.1061] 6.982 Yes
Proposed vs. CNN (No Aug) 6.249 0.0033 0.0167 [0.0554, 0.0986] 4.424 Yes
Proposed vs. CNN + Trad. Aug 3.735 0.0202 0.1010 [0.0253, 0.0700] 3.003 No
Proposed vs. CNN + WGAN-GP 4.595 0.0101 0.0504 [0.0193, 0.0432] 2.491 No

Explainability analysis results

Figure 8 presents the Grad-CAM explainability analysis results. The top row shows class-wise activation patterns: ASD-related activations are concentrated in medial prefrontal and temporal regions (warm colors), while TD-related activations are more distributed across posterior regions (cool colors). The ASD-TD contrast map highlights differential activation in the medial prefrontal cortex, superior temporal sulcus, and temporoparietal junction. The bottom panel ranks the top 10 discriminative brain regions by relative importance, with mPFC (DMN, importance = 0.95), STS (Social, 0.88), and TPJ (Social, 0.82) identified as the most influential regions for ASD classification.

Table 6 provides detailed information on the top 10 discriminative brain regions identified by the explainability analysis. The regions span multiple functional networks, including the default mode network (DMN: mPFC, PCC, Precuneus), social brain network (STS, TPJ), limbic system (Amygdala), salience network (ACC, Insula), visual processing areas (FFA), and executive control regions (IFG). Each region’s known association with ASD-related functions is listed, demonstrating strong alignment between the model’s decision features and established neurobiological evidence.

Table 6.

Top 10 discriminative brain regions identified by explainability analysis

Rank Brain Region Network MNI (x, y, z) Importance ASD Association
1 mPFC DMN (0, 52, 8) 0.95 Theory of mind / Self-referential processing
2 STS Social (58, -42, 12) 0.88 Social perception / Biological motion
3 TPJ Social (52, -54, 24) 0.82 Mentalizing / Perspective-taking
4 Amygdala Limbic (24, -4, -18) 0.79 Emotion processing / Face perception
5 ACC Salience (0, 34, 18) 0.74 Error monitoring / Social cognition
6 FFA Visual (40, -50, -18) 0.71 Face recognition / Social processing
7 PCC DMN (0, -52, 26) 0.67 Default mode / Self-reflection
8 IFG Executive (48, 18, 8) 0.62 Mirror neuron / Action understanding
9 Precuneus DMN (0, -62, 40) 0.56 Self-awareness / Episodic memory
10 Insula Salience (38, 14, -4) 0.51 Interoception / Empathy

Ablation study

Table 7 presents the results of the ablation study, where individual components are systematically removed from the full model. Removing GAN augmentation causes the largest performance drop (AUC: -0.084), confirming that data augmentation is the most critical component. Replacing WGAN-GP with vanilla GAN results in a notable decline (AUC: -0.048), validating the importance of gradient penalty for training stability. Removing spectral normalization leads to an AUC decrease of 0.035, while removing Grad-CAM guidance has the smallest impact (AUC: -0.017), as expected since it primarily serves an interpretability rather than classification function.

Figure 9 visualizes the ablation study results as a heatmap, providing an intuitive comparison of performance across all model variants and metrics. The full model (top row, highlighted with red border) consistently achieves the darkest colors (highest scores) across all four metrics. The color gradient clearly illustrates the progressive performance degradation as components are removed, with the w/o GAN Aug variant showing the most pronounced decline across all metrics.

Mechanism and computational analysis

Figure 10 presents a mechanism analysis diagram illustrating how the top-ranked brain regions converge into three interpretable mechanism modules that support ASD prediction. The five most important regions (mPFC, STS, TPJ, Amygdala, and ACC) are mapped to three functional modules: social cue processing (STS, TPJ), self-referential integration (mPFC, ACC), and affective salience (Amygdala, ACC). These modules collectively contribute to the diagnostic summary, yielding a predicted ASD probability of 0.87 with a dominant signature of social cue disruption combined with DMN integration imbalance.

Fig. 10.

Fig. 10

Mechanism analysis: from top brain regions through interpretable modules to diagnostic summary\

Table 8 summarizes the computational costs of all compared methods. The proposed full pipeline requires 11.2 h of training time, which is substantially longer than traditional methods (SVM: 0.1 h, Random Forest: 0.3 h) and baseline CNNs (2.5–4.1 h), primarily due to the additional GAN training phase (8.5 h). However, the inference time remains practical at 18 ms per sample, which is comparable to the CNN baselines (15 ms) and well within the requirements for clinical deployment. The model contains 16.0 M parameters and requires 8.6 GB of GPU memory during training.

Table 8.

Computational cost comparison across methods

Method Training Time (h) Inference (ms/sample) Parameters (M) GPU Memory (GB)
SVM 0.1 2 0.01 0
Random Forest 0.3 5 N/A 0
CNN (No Aug) 2.5 15 3.2 4.2
CNN + Trad. Aug 4.1 15 3.2 4.2
GAN Training Only 8.5 N/A 12.8 8.6
Proposed (Full Pipeline) 11.2 18 16.0 8.6

Discussion

Overall model performance and comparison with previous studies

This study established a comprehensive technical pipeline of “ComBat site correction-conditional GAN augmentation-multimodal cross-attention fusion-explainable attribution analysis,” achieving favorable classification performance on the multi-center ABIDE I/II dataset. Results show that, under stratified five-fold cross-validation, the proposed model attained an AUC of 0.871 ± 0.016 and balanced accuracy of 0.797 ± 0.012, and the average AUC remained 0.783 ± 0.041 in leave-one-site-out (LOSO) cross-validation, indicating that the framework not only possesses strong discriminative ability within sites but also maintains relatively stable generalization performance across sites. Compared with previous studies based on ABIDE data, the results of this study are at a relatively advanced level. Most prior ABIDE-based ASD classifiers report whole-cohort AUC in the 0.65–0.80 range when no harmonization or augmentation is applied and no multimodal fusion is used. The substantially higher AUC obtained here (0.871 ± 0.016 within-site, 0.783 ± 0.041 LOSO) is attributable to a combination of factors that previous single-component studies have not used jointly: (i) ComBat-based site harmonization of FC and morphometry, which removes a large fraction of non-biological inter-site variance before model training; (ii) class-conditional GAN augmentation of the minority ASD class, which alleviates the small-N / class-imbalance bottleneck that dominates ABIDE; (iii) explicit dual-branch sMRI/rs-fMRI fusion via cross-attention, which adds complementary structural information that single-modality classifiers miss; and (iv) restricting the analytic cohort to males 5–40 years and to subjects that pass strict QC, which reduces between-subject heterogeneity. These factors together account for the observed gap relative to single-modality and single-augmentation ABIDE baselines; we explicitly do not claim the absolute number generalizes to unrestricted cohorts (e.g., female, very young, or low-QC subjects). Early multi-site ASD classification studies have confirmed the feasibility of resting-state functional connectivity for automatic identification, but overall performance was often limited by site heterogeneity, sample imbalance, and insufficient feature representation capacity [13]. In recent years, the introduction of deep learning, multimodal learning, and site harmonization processing has gradually advanced the field from “classifiable” to “more stably classifiable” [10]. Therefore, the performance improvement in this study is not attributed to a single module but is more likely the result of synergistic optimization across multiple components.

The facilitating effect of conditional GAN augmentation on small sample learning

The results of this study indicate that cGAN augmentation is one of the critical factors for improving diagnostic performance in small sample scenarios. Unlike traditional noise perturbation or simple oversampling, this study incorporates class and site conditions into the generation process, enabling the generated samples to not only better approximate the statistical distribution of real samples but also exhibit stronger class consistency in diagnostic semantics. The t-SNE distributions, Q-Q plots, and training curves collectively demonstrate that cGAN augmentation does not merely replicate the original samples but rather expands the coverage of real samples in the high-dimensional feature space to some extent, thereby alleviating overfitting during model training. Previous studies have noted that conventional augmentation methods typically yield limited benefits in ASD brain imaging tasks, whereas generative augmentation constrained by task and distribution is more likely to provide stable improvements under sample-limited conditions [25]. Our study further found that the model achieves optimal performance when the augmentation ratio reaches 1:1; increasing the number of generated samples beyond this point results in a slight decline in classification performance, suggesting that the relationship between augmentation ratio and downstream AUC is non-monotonic and peaks at an intermediate ratio, beyond which the synthetic distribution begins to dilute the real-sample signal rather than compensate for it. This indicates that GAN-based augmentation is better suited as a regularization and distribution compensation technique rather than a substitute for genuine clinical samples.

The key value of multimodal fusion

Ablation experiments reveal that removing GAN-based augmentation produces the largest AUC drop in Table 7 (− 0.084), while replacing the cross-attention multimodal fusion with simple feature concatenation (added as an additional ablation row, “w/o cross-attention fusion”) produces the second-largest drop (− 0.061). Both GAN-based augmentation and multimodal cross-attention fusion are therefore the dominant contributors to performance, with augmentation slightly more important in absolute terms. This is because rs-fMRI functional connectivity characterizes the dynamic coordination patterns between brain regions, reflecting abnormalities in brain network organization and information transmission; meanwhile, sMRI morphological features represent structural differences such as cortical thickness, surface area, and gray matter volume, which constitute a relatively stable neuroanatomical basis. These two modalities describe the neural phenotype of ASD from functional and structural dimensions, respectively, and possess inherent complementarity [13, 18]. Our study employs a dual-branch encoder and cross-attention mechanism, allowing one modality to perform feature selection and information querying on the other, thereby outperforming simple concatenation strategies. The results show that the AUC after multimodal fusion is significantly higher than that of any single-modality model, suggesting that brain imaging abnormalities in ASD are not confined to a single dimension but are more likely the combined result of functional organizational abnormalities and structural developmental deviations.

Neurobiological significance of explainability results

The explainability analysis in this study indicates that regions such as the precuneus, posterior cingulate cortex, medial prefrontal cortex, and superior temporal sulcus carry relatively high weights in the model’s decision-making process; at the network level, the default mode network, social brain network, and salience network contribute most prominently. These findings are largely consistent with major discoveries in previous ASD neuroimaging research [8]. The default mode network is involved in self-referential processing, social cognition, and theory of mind, and its abnormalities have long been considered a critical neural basis for social dysfunction in ASD; social brain-related regions are closely associated with face recognition, social cue processing, and understanding others’ intentions. The importance of the salience network suggests that ASD patients may exhibit widespread abnormalities in attention switching, filtering of internal and external information, and emotional regulation. More importantly, the explainability results of this study do not point to random local features but rather focus on brain regions and networks highly relevant to core ASD symptoms, which to some extent enhances the clinical credibility of the model. It should be noted that attribution analysis reveals “which features the model relies on for judgment” and does not equate to the direct causal mechanisms of the disease; therefore, further interpretation requires integration with behavioral, developmental, and independent sample validation studies [21].

Cross-site generalization ability and clinical translational significance

In multi-center brain imaging studies, differences in scanning parameters, hardware platforms, and subject composition often introduce significant site effects, thereby confounding truly disease-related signals. This study achieved an average AUC of 0.783 ± 0.041 (95% CI 0.762–0.804) across the 17 sites included in the LOSO protocol, with per-site AUC ranging from 0.71 (smallest, lowest-N site) to 0.84 (largest, most QC-clean site); a complete per-site breakdown is provided in Supplementary Table S1. This per-site variability in LOSO validation indicates that the model possesses a certain degree of cross-center transferability; however, performance fluctuations across different sites suggest that site heterogeneity has not been completely eliminated. Previous studies have demonstrated that the ComBat method can effectively reduce non-biological biases in multi-site functional connectivity and structural metrics while preserving biological differences and improving model transferability [25]. Ablation experiments in this study further confirm that removing ComBat correction results in a marked decline in model performance. Therefore, in multi-center ASD brain imaging classification tasks, site harmonization is not a subsidiary step but a crucial prerequisite to ensure that the model learns intrinsic disease features. Based on the results of this study, if further validation can be conducted on independent clinical samples, prospective cohorts, and real hospital settings, the technical approach proposed herein holds promise as an important candidate for objective auxiliary diagnosis of ASD.

Recent evidence suggests that the clinical translation of ASD neuroimaging models depends not only on higher classification accuracy, but also on the coordinated advancement of multimodal fusion, external validation, privacy-preserving learning, and explainable biomarker discovery. Systematic reviews have shown that MRI-based ASD diagnosis still faces substantial heterogeneity across cohorts and acquisition settings. At the same time, multimodal frameworks integrating structural and functional MRI have demonstrated stronger robustness and discrimination than single-modality pipelines [9]. In addition, recent work highlights that privacy-preserving collaboration and explainable AI are becoming increasingly important for multi-center deployment and clinician trust. Therefore, future ASD auxiliary diagnostic systems should move beyond retrospective single-dataset optimization toward multi-center, explainable, and clinically transferable validation frameworks.

Limitations and future directions

This study has several limitations. First, the primary analysis was restricted to male subjects aged 5 to 40 years, which, although helpful in reducing heterogeneity related to sex and age, limits the generalizability of the conclusions to female ASD populations and younger children. Second, this study utilized static Pearson functional connectivity and did not incorporate dynamic functional connectivity or temporal evolution features; therefore, the characterization of time-varying abnormalities in ASD brain networks remains insufficient. Third, all GAN-based augmentation in this work was carried out at the extracted feature level (FC vectors and per-region morphometric vectors), not on raw voxel-level imaging data; the brain-slice panels in Fig. 3 are template-based reverse-projections used solely for qualitative visualization and do not imply voxel-level synthesis. Although the statistical distribution consistency was satisfactory, the biological validity and clinical transferability of this approach require further verification. Fourth, the improvement of the proposed cGAN over the WGAN-GP baseline did not reach Bonferroni-corrected statistical significance (p = 0.050); larger multi-site cohorts will be needed to confirm whether class conditioning provides a statistically robust gain over plain WGAN-GP. Finally, this study employed a post-hoc attribution explanation strategy. Future work could further integrate graph neural networks, dynamic connectivity modeling, federated learning, and more rigorous analyses of explanation stability to advance ASD brain imaging research from high-performance classification toward robust biomarker discovery and clinically applicable model development [14, 22].

Conclusion

This study, based on the ABIDE I/II multicenter dataset, constructed an ASD-assisted diagnostic model that integrates conditional GAN data augmentation, multimodal brain imaging feature fusion, and interpretable deep learning. Results show that this framework effectively improves classification performance and cross-center generalization ability under small sample conditions. Furthermore, its discrimination criteria primarily focus on default mode networks and social brain-related regions, exhibiting good neurobiological consistency. Overall, this study provides a promising technical approach for the objective and intelligent assisted diagnosis of ASD. However, its clinical application still relies on further validation through larger-scale independent samples and prospective studies.

Supplementary Information

Below is the link to the electronic supplementary material.

Supplementary Material 1 (16.5KB, docx)

Acknowledgements

Not applicable.

Abbreviations

ASD

Autism Spectrum Disorder

TD

Typically Developing

LOSO

Leave-one-site-out

DMN

Default mode network

sMRI

Structural Magnetic Resonance Imaging

rs-fMRI

Resting-state functional magnetic resonance imaging

ABIDE

Autism Brain Imaging Data Exchange

Author contributions

WeiWei Xu: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Resources, Software, Visualization, Writing – original draft, Writing – review & editing.

Funding

Not applicable.

Data availability

The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.

Declarations

Ethics approval and consent to participate

The present study performed secondary analysis on the de-identified, publicly accessible ABIDE I and ABIDE II datasets. All original imaging and phenotypic data were collected under ethical approvals granted by the institutional review boards (IRBs) of each data collection site. No new human subjects were recruited, no new clinical procedures were conducted, and all analyses were based on anonymized public data. Therefore, additional ethical approval for this study was not needed.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Abbas SQ, Chi L, Chen YPP, Deep MNF. Deep multimodal neuroimaging framework for diagnosing Autism spectrum disorder. Artif Intell Med. 2023;136:102475. 10.1016/j.artmed.2022.102475. [DOI] [PubMed] [Google Scholar]
  • 2.Ali S, Baloch Z, Narejo S, et al. A Review of Developments in Generative AI, Machine Learning, and Neuroimaging for the Diagnosis of Autism. VFAST Trans Softw Eng. 2025;13(2):68–82. 10.21015/vtse.v13i2.2089. [DOI] [Google Scholar]
  • 3.Bahathiq RA, Banjar H, Bamaga AK, et al. Machine learning for autism spectrum disorder diagnosis using structural magnetic resonance imaging: Promising but challenging. Front Neuroinformatics. 2022;16:949926. 10.3389/fninf.2022.949926. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Campo F, Retico A, Calderoni S, et al. Multi-Site MRI Data Harmonization with an Adversarial Learning Approach: Implementation to the Study of Brain Connectivity in Autism Spectrum Disorders. Appl Sci. 2023;13(11):6486. 10.3390/app13116486. [DOI] [Google Scholar]
  • 5.Dcouto SS, Pradeepkandhasamy J. Multimodal deep learning in early autism detection—recent advances and challenges. Engineering Proceedings. 2024;59(1):205. 10.3390/engproc2023059205. [DOI]
  • 6.Duan Y, Zhao W, Luo C, et al. Identifying and Predicting Autism Spectrum Disorder Based on Multi-Site Structural MRI With Machine Learning. Front Hum Neurosci. 2022;15:765517. 10.3389/fnhum.2021.765517. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Han J, Jiang G, Ouyang G, et al. A multimodal approach for identifying autism spectrum disorders in children. IEEE Trans Neural Syst Rehabil Eng. 2022;30:2003–11. 10.1109/TNSRE.2022.3192431. [DOI] [PubMed] [Google Scholar]
  • 8.Helmy E, Elnakib A, ElNakieb Y, et al. Role of Artificial Intelligence for Autism Diagnosis Using DTI and fMRI: A Survey. Biomedicines. 2023;11(7):1858. 10.3390/biomedicines11071858. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Huynh N, Deshpande G. A review of the applications of generative adversarial networks to structural and functional MRI based diagnostic classification of brain disorders. Front NeuroSci. 2024;18:1333712. 10.3389/fnins.2024.1333712. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Ingalhalikar M, Shinde S, Karmarkar A, et al. Functional Connectivity-Based Prediction of Autism on Site Harmonized ABIDE Dataset. IEEE Trans Biomed Eng. 2021;68(12):3628–37. 10.1109/TBME.2021.3080259. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Jönemo J, Abramian D, Eklund A. Evaluation of Augmentation Methods in Classifying Autism Spectrum Disorders from fMRI Data with 3D Convolutional Neural Networks. Diagnostics. 2023;13(17):2773. 10.3390/diagnostics13172773. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Khodatars M, Shoeibi A, Sadeghi D, et al. Deep learning for neuroimaging-based diagnosis and rehabilitation of Autism Spectrum Disorder: a review. Comput Biol Med. 2021;139:104949. 10.1016/j.compbiomed.2021.104949. [DOI] [PubMed] [Google Scholar]
  • 13.Koc E, Kalkan H, Bilgen S. Autism spectrum disorder detection by hybrid convolutional recurrent neural networks from structural and resting state functional MRI images. Autism Res Treat. 2023;2023:4136087. 10.1155/2023/4136087. [DOI] [PMC free article] [PubMed]
  • 14.Li J, Wang F, Pan J, et al. Identification of Autism Spectrum Disorder With Functional Graph Discriminative Network. Front NeuroSci. 2021;15:729937. 10.3389/fnins.2021.729937. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Padmanabhan A, Lynch CJ, Schaer M, et al. The default mode network in autism. Biol Psychiatry: Cogn Neurosci Neuroimaging. 2017;2(6):476–86. 10.1016/j.bpsc.2017.04.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Santana CP, de Carvalho EA, Rodrigues ID, et al. rs-fMRI and machine learning for ASD diagnosis: a systematic review and meta-analysis. Sci Rep. 2022;12:6030. 10.1038/s41598-022-09821-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Saponaro S, Giuliano A, Bellotti R, et al. Multi-site harmonization of MRI data uncovers machine-learning discrimination capability in barely separable populations: An example from the ABIDE dataset. NeuroImage: Clin. 2022;35:103082. 10.1016/j.nicl.2022.103082. [DOI] [PMC free article] [PubMed]
  • 18.Saponaro S, Lizzi F, Serra G, et al. Deep learning based joint fusion approach to exploit anatomical and functional brain information in autism spectrum disorders. Brain Inf. 2024;11(1):2. 10.1186/s40708-023-00217-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Schielen SJC, Pilmeyer J, Aldenkamp AP, et al. The diagnosis of ASD with MRI: a systematic review and meta-analysis. Translational Psychiatry. 2024;14:318. 10.1038/s41398-024-03024-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Serra G, Mainas F, Golosio B, et al. Effect of data harmonization of multicentric dataset in ASD/TD classification. Brain Inf. 2023;10(1):32. 10.1186/s40708-023-00210-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Vidya S, Gupta K, Aly A, et al. Identification of critical brain regions for autism diagnosis from fMRI data using explainable AI: An observational analysis of the ABIDE dataset. EClinicalMedicine. 2025;88:103452. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Wang H, Jing H, Yang J, et al. Identifying autism spectrum disorder from multi-modal data with privacy-preserving. npj Mental Health Res. 2024;3(1):15. 10.1038/s44184-023-00050-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Xu M, Calhoun VD, Jiang R, et al. Brain imaging-based machine learning in autism spectrum disorder: methods and applications. J Neurosci Methods. 2021;361:109271. 10.1016/j.jneumeth.2021.109271. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Yang C, Wang P, Tan J, et al. Autism spectrum disorder diagnosis using graph attention network based on spatial-constrained sparse functional brain networks. Comput Biol Med. 2021;139:104963. 10.1016/j.compbiomed.2021.104963. [DOI] [PubMed] [Google Scholar]
  • 25.Zhou Y, Jia G, Ren Y, et al. Advancing ASD identification with neuroimaging: a novel GARL methodology integrating Deep Q-Learning and generative adversarial networks. BMC Med Imaging. 2024;24:186. 10.1186/s12880-024-01360-y. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Citations

  1. Saponaro S, Giuliano A, Bellotti R, et al. Multi-site harmonization of MRI data uncovers machine-learning discrimination capability in barely separable populations: An example from the ABIDE dataset. NeuroImage: Clin. 2022;35:103082. 10.1016/j.nicl.2022.103082. [DOI] [PMC free article] [PubMed]

Supplementary Materials

Supplementary Material 1 (16.5KB, docx)

Data Availability Statement

The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.


Articles from BMC Medical Imaging are provided here courtesy of BMC

RESOURCES