Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 Apr 30;16:20075. doi: 10.1038/s41598-026-50461-x

MACAFNet transformer-based multi-atlas fusion framework for autism spectrum disorder classification using functional connectivity

D Aarthi 1,✉, S Kannimuthu 1
PMCID: PMC13323724  PMID: 42062424

Abstract

Autism Spectrum Disorder (ASD) is a neurodevelopmental condition characterized by impairments in communication, social interaction and behavior. ASD individuals develop symptoms such as recurrent actions, atypical facial expressions and challenges in social engagement. This study proposes a Multi-Atlas Context-Aware Fusion Network (MACAFNet) for ASD classification using functional connectivity (FC) features derived from resting-state functional MRI (rs-fMRI) to enable reliable ASD classification. This framework integrates information from multiple brain atlases such as AAL, CC200, Dosenbach160, EZ, HO and TT to capture complementary neurofunctional features across diverse brain parcellations. Each atlas-specific connectivity representation is projected into a shared embedding space and fused using a Transformer-based attention mechanism that explicitly models inter-atlas contextual dependencies. Unlike traditional static fusion systems, this allows for dynamic and adaptive feature integration. The fused representation is then processed by a neural classifier for ASD classification. Experiments conducted on the ABIDE-I dataset demonstrate that the proposed approach achieves an accuracy of 88.46% and an AUC of 0.9546, indicating strong discriminative capability. The results highlight the effectiveness of context-aware multi-atlas fusion in capturing complex brain connectivity patterns for research-oriented ASD classification.

Keywords: Autism spectrum disorder (ASD), Resting-state fMRI (rs-fMRI), Functional connectivity (FC), Multi-atlas fusion, Transformer attention, Cross-atlas learning

Subject terms: Computational biology and bioinformatics, Diseases, Health care, Mathematics and computing, Neuroscience

Introduction

ASD is a complex neurodevelopmental syndrome characterized by difficulties in social interaction and communication flexibility. The Diagnostic and Statistical Manual of Mental Disorders has categorized ASD into three levels in reference to the support required for day-to-day functioning1. Progressive outcomes can be achieved using early identification of ASD-related patterns. Children getting early interventions show better behavioral and social flexibility2. Furthermore, cerebellar anomalies have also been connected to repetitive behaviors and decreased environmental exploration in children with ASD3. Thus, early identification and clinical assessment of neurological differences play a crucial role in improving treatment outcomes4. Few works highlight that the timeline for ASD assessment in children can be minimized when comprehensive pre-assessment information is available in advance. In contrast, ASD in adults requires more diagnostic interactions5. ASD classification is a multidisciplinary approach that includes the coordination of neurologists, psychiatrists, developmental pediatricians, and therapists6. This multidisciplinary approach provides a strong platform for integrating modern computational tools with classification and treatment.

Abnormal neurological or behavioral indicators are automatically identified using computational method, which help doctors customize treatment plans7. AI-based medical imaging assists neurologists in measuring changes in the structural and functional regions of the brain. This enables more reliable data-driven clinical decision-making8. MRI data is a promising modality for investigating the neurological basis of ASD. Structural MRI (sMRI) and functional MRI (fMRI) have aided in the detection of various abnormalities in ASD individuals. These abnormalities include irregular brain folding, excessive reduction of neural connections and poor connectivity between brain regions9. The examination of fMRI data highlights that ASD is associated with decreased inter-regional connectivity and higher segregation across functional networks10. Conventional models fail to capture neuroanatomical and functional anomalies that are better identified by modern approaches11. Enhanced volumetric and connectivity analyses using rs-fMRI provide greater support to clinicians in investigating brain developmental activities. Further, they help in examining the efficiency of therapeutic interventions over time12. Strong neurobiological biomarkers can be developed through the integration of AI with multimodal neuroimaging data13. Hence, neuroimaging-based approaches have become the basis for identifying brain-level signatures of ASD. Few studies focus on gamified mobile health applications to support children with disabilities, particularly in communication and self-management14. Mobile health and wearable technologies enable real-time behavioral monitoring and social skill training. Whereas, AI-based applications capture facial, gesture and speech patterns. However, limitations such as data confidentiality, the need for expert supervision and long-term assessment challenges remain15.

Video-based motion analysis using DL has become a viable technique for early ASD identification. It monitors head, trunk and hand movements to identify behavioral features16. The limitations of such models include data availability and generalization issues. The key features of ASD are atypical prosody, pragmatic language complications and speech impairments. Approaches such as Natural Language Processing (NLP) and speech analytics provide insights into these aspects; however, findings remain conflicting and require further investigation17. These techniques demonstrate the capability of AI in ASD diagnosis using facial, motion and speech-based datasets. But several challenges remain in clinical application. Variations exist across datasets obtained from different clinical settings. These may be due to differences in imaging procedures, feature extraction strategies and demographic composition. Many studies have developed Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN)-based pipelines for ASD classification. Still, most existing approaches either rely on voxel-level analysis or direct processing of connectivity matrices without explicitly modeling inter-atlas relationships. As a result, most existing approaches either rely on voxel-level analysis or direct processing of connectivity matrices without explicitly modeling inter-atlas relationships, limiting their ability to capture complementary information across diverse brain parcellations. Moreover, prior hybrid models typically operate as black-box classifiers. It offers limited interpretability on the contributing brain regions.

In recent years, rs-fMRI has become an important modality for analyzing alterations in FC associated with ASD. Instead of voxel-level analysis, the brain is commonly partitioned into regions of interest (ROIs) using atlases such as Automated Anatomical Labeling (AAL), Craddock 200 Functional (CC200), Dosenbach 160 Functional Regions of Interest (Dosenbach160), Eickhoff–Zilles (EZ), Harvard–Oxford (HO) and Talairach–Tournoux (TT) atlases. These atlas-based representations reduce dimensionality and improve interpretability while capturing meaningful connectivity patterns. Prior studies have shown that ASD-related abnormalities can be effectively identified through disrupted functional connectivity among distributed brain regions18,19. In parallel, single-atlas approaches such as EAG-RS leverage explainability-guided ROI selection to identify informative brain regions20. Recent works have explored multi-atlas integration to enhance ASD classification performance. Models such as MADE-for-ASD and M3ASD demonstrate that combining multiple atlases provides complementary connectivity information and improves robustness. However, effectively capturing and integrating complementary information across atlas representations remains underexplored. This results in the requirement for innovative frameworks to effectively capture complex inter-atlas relationships.

To address these limitations, this work proposes MACAFNet, a novel Multi-Atlas Context-Aware Fusion Network for ASD classification using FC features derived from rs-fMRI. Existing multi-atlas approaches depend on static or shallow fusion strategies, whereas the proposed framework models inter-atlas contextual dependencies using a lightweight Transformer Encoder. The proposed model enables dynamic weighting and interaction modeling across atlas-specific representations. This representation allows the model to effectively capture complementary neurofunctional characteristics. By integrating six widely used brain atlases, The proposed approach combines six brain atlases and presents a unified representation encompassing anatomical, functional and cytoarchitectonic perspectives. Evaluation across 63 atlas combinations demonstrates the effectiveness of the proposed fusion strategy. It achieves a strong performance on the ABIDE-I dataset.

The main contributions of this study are outlined as follows:

  • A novel MACAFNet framework for ASD classification using multi-atlas FC features from rs-fMRI.

  • A Transformer-based attention mechanism for context-aware inter-atlas feature fusion.

  • A multi-atlas representation integrating six brain atlases to capture complementary neurofunctional information.

  • An ablation study (63 atlas combinations) validating the effectiveness of multi-atlas integration.

  • Competitive performance on ABIDE-I (88.46% accuracy and 0.9546 AUC).

  • A mathematically grounded formulation ensuring reproducibility and clarity.

The remainder of this paper is organized as follows. Section  2 reviews the literature work related to ASD prediction. Section  3 presents the architecture of the proposed model and Sect.  4 describes the experimental setup. Section  5 describes the results and Sect.  6 shows the discussion. The last section presents the conclusion.

Related work

Recent studies employ deep learning (DL) techniques on neuroimaging, behavioral and linguistic data in ASD screening research. MRI-based models aid in identifying structural and functional abnormalities associated with ASD. Eye-tracking data, facial expression analysis and motion-based features support automated screening systems. Speech and language variances are analyzed using NLP. The gamified mobile health applications increase accessibility and engagement for ASD children. Few works on ensemble and multimodal frameworks have proved enhanced predictive performance by combining diverse data sources. Despite these advancements, challenges remain in terms of dataset diversity, model generalization and clinical adoption. This section reviews various approaches for ASD detection, their strengths and limitations to provide a clear understanding of recent developments in this field. Table 1 presents an overview of models ranging from conventional machine learning (ML) approaches to advanced DL and multimodal frameworks used for ASD detection.

Table 1.

Comprehensive summary of existing studies that utilize ML, DL and multimodal data integration techniques to support ASD classification and screening.

Author Model/method Data Modality Dataset Pros Cons
Santana et al.12 SVM, ANN, other classifiers (meta-analysis) rs-fMRI 55 rs-fMRI studies SVM robust; ANN scales with larger datasets; multimodal approaches improved sensitivity to 84.7% Accuracy decreased with larger samples; poor methodological quality; limited clinical applicability
Ning Qiang et al.21 HRVAE + LASSO rs-fMRI 871 subjects from ABIDE Captures hierarchical brain networks; 82.1% accuracy; promising for clinical assessment support Complexity; potential overfitting; requires large datasets
Hazlett et al.22 DL on cortical surface hyper-expansion sMRI High-risk infants Early prediction at 24 months; 88% sensitivity; non-invasive biomarkers Small cohort; longitudinal follow-up required
Bengs et al.23 4D spatio-temporal DL (CNN + Conv-Recurrent) fMRI fMRI data Captures spatial and temporal features; F1-score 0.71 Moderate performance; high computational cost
Sachdeva et al.24 MLP on connectivity + graph topology features rs-fMRI + T1-weighted MRI rs-fMRI + T1-weighted images 83.57% accuracy; 0.978 AUC; identifies key ROIs Feature engineering needed; limited interpretability
Heinsfeld et al.25 DNN on functional connectivity matrices rs-fMRI 505 ASD individuals and 530 TC from ABIDE I One of the earliest DL models for ASD prediction using ABIDE; achieved 70% accuracy; demonstrated feasibility of DL on connectivity features Performance limited; sensitive to site variability
Parisot et al.26 Graph Convolutional Network (GCN) using population graph rs-fMRI 871 subjects from ABIDE, 731 subjects from ADNI Demonstrates the relationships between subjects; improves ASD classification; integrates imaging and non-imaging data Requires graph construction; performance depends on population graph design
Liu et al.27 A Stacked Sparse Denoising Autoencoder (SSDAE) and Multi-Layer Perceptron (MLP)

rs-fMRI

(Multi-atlas)

505 ASD individuals and 530 TC from ABIDE I Higher accuracy than prior ABIDE studies; strong ASD detection; excellent performance on curated subset Lower specificity; generalization variability; subset bias risk; requires external validation
Yang et al.28 Multi-Atlas Multi-View Learning rs-fMRI 407 ASD and 468 TC, 6 to 64 years Integrates connectivity patterns from multiple atlases; improves representation learning Limited modeling of inter-atlas contextual relationships
Jung et al.29 Explainability-guided ROI selection (EAG-RS) framework rs-fMRI 539 ASD and 573 TD, aged 7 to 64 years. Explainability-guided ROI selection; captures high-order inter-regional relationships; provides neuroscientific insights into ASD subtypes MLP-based architecture with high parameter complexity; relies only on FC (single modality); limited multimodal integration.
Federica Cilia et al.30 CNN on scanpath visual maps Eye-tracking (gaze scan-paths) 59 school-age participants Non-invasive; targets social attention; 90% accuracy Small sample; variability in stimuli; limited external validity
Kanhirakadavath et al.31 DNN, SVM, RF on ETSP images Eye-tracking spatial plots (ETSP) Child eye-tracking dataset DNN achieved 97% AUC (augmented); captures gaze dynamics; non-invasive Small cohort; task/stimuli dependent; dataset variability
Farhat et al.32 Ensemble CNN (VGG16 + Xception) Facial images Kaggle ASD Face Image High accuracy (97–99%); strong generalization Limited dataset diversity; may not generalize clinically
Anderson et al.33 Markerless video-based gait analysis (OpenPose) Video (gait and movement) Children 1–5 years with/without NDD Non-invasive; accurate gait metrics; early-age feasibility Step length affected by body size; methodological factors
Jin et al.34 ML on home-based video Video (home recordings) 9,959 subjects; 19 studies High sensitivity/specificity; user-friendly; early ASD detection Variable video quality; feature selection may affect results
Dia et al.35 Transformer-based affect recognition Video (facial and motion cues) SSBD dataset; YouTube videos Continuous affect recognition; ASD-specific dataset; highlights stimming & face analysis Data quality/diversity issues; model underperforms if trained on neurotypical data only
Ganjigunte et al.36 BAR pipeline Video (behavioral actions) 400 ASD + 125 ODD children; SSBD Untrimmed video analysis; 78–81% accuracy; clinically relevant actions Requires labeled actions; dataset-specific; computationally intensive
Abbas et al.37 Multi-modular ML assessment (parent + video + clinician) Multimodal (video, clinical, parental) 375 children, 18 to 72 months Outperforms conventional screeners; AUC improvement; time-efficient Needs validation in general populations; limited primary care testing
Ma et al.38 ML on natural speech prosody Speech/audio Meta-analysis Pitch-related features distinguish ASD; 75–80% sensitivity/specificity Temporal features less informative; age-dependent effects
Chi et al.39 RF and CNN Speech/audio Self-recorded child speech (mobile game) CNN achieved 79% accuracy; works in home environment; scalable Audio quality variable; requires preprocessing
Mahmoudi et al.14 Scoping review on gamified apps Mobile health (mHealth) applications 38 studies covering 32 apps Increases engagement/motivation; supports communication & education Mixed effectiveness; some disability groups underexplored; long-term impact unknown

The reviewed literature works demonstrate ML and DL approaches for ASD detection across multiple modalities. Studies based on rs-fMRI neuroimaging have shown promising results in identifying FC alterations and reliable biomarkers. Furthermore, non-invasive approaches such as eye-tracking, facial analysis, video-based behavior recognition and speech processing provide scalable solutions for early screening. Many neuroimaging studies rely on single-atlas representations or simple connectivity features. This limits the ability of the models to capture the complex and heterogeneous nature of brain connectivity. Recent multi-atlas and explainability-guided methods also do not fully exploit contextual relationships across atlas representations. Also, the challenges such as dataset heterogeneity and limited generalization remain significant concerns.

Methods

The methods section describes the proposed Multi-Atlas Context-Aware Fusion Network (MACAFNet) developed for ASD classification using rs-fMRI connectivity features. The framework utilizes FC representations derived from multiple brain atlases to capture complementary neurofunctional information across different brain parcellations. This context-aware fusion strategy allows the model to effectively capture inter-atlas dependencies and enhances predictive robustness and generalization for rs-fMRI-based ASD classification.

Proposed architecture

The proposed architecture MACAFNet, is designed to classify ASD-related patterns using FC features derived from rs-fMRI data. The framework utilizes atlas-based brain parcellations to represent neural interactions among predefined ROIs, instead of directly processing MRI slices. Six brain atlases such as AAL, CC200, Dosenbach160, EZ, HO and TT are employed to generate complementary FC representations of the brain. In this study, each subject’s ROI time series is normalized using z-score normalization. It is then followed by FC computation using Pearson correlation. The upper triangular elements of the FC matrix are extracted as feature vectors. These connectivity features are stored and used as input vectors to the model. These features are transformed into a shared embedding space. This enables consistent representation across different atlas configurations. A Transformer-based attention module is employed to model contextual relationships among the atlas representations and learn the relative importance of each atlas in the classification process.

The attention mechanism facilitates context-aware fusion of multi-atlas connectivity features. This allows the network to capture complementary information from different brain parcellation schemes. Each atlas-specific FC vector is independently projected into a shared embedding space using fully connected layers. The resultant embeddings are stacked and passed through a Transformer encoder to model inter-atlas dependencies. A lightweight Transformer configuration with one layer, four heads and embedding size of 64 is adopted to balance model complexity and generalization. This is configured to adapt to the limited sample size of neuroimaging datasets. Finally, mean pooling is applied across atlas embeddings to obtain a unified representation for classification. This architecture enables the model to effectively capture complex inter-regional connectivity patterns and improves the robustness and generalization capability of rs-fMRI-based ASD prediction. Figure 1 illustrates the overall workflow of the proposed MACAFNet framework for ASD classification.

Fig. 1.

Fig. 1

Architecture of the proposed framework used for ASD classification. Multi-atlas rs-fMRI connectivity features are embedded into a shared space and fused using a Transformer-based attention module and fully connected layers for robust ASD classification.

Dataset description and experimental data splitting strategy

The proposed framework is evaluated on the publicly available Autism Brain Imaging Data Exchange I (ABIDE-I) dataset40. It consists of rs-fMRI data collected from 17 different international sites. The sites are NYU Langone Medical Center (NYU), University of Michigan (UM), University of Utah School of Medicine (USM), University of California Los Angeles (UCLA), Kennedy Krieger Institute (KKI), Oregon Health & Science University (OHSU), Stanford University (SU), University of Pittsburgh School of Medicine (PITT), Yale Child Study Center (YALE), Trinity College Dublin (TRINITY), University of Leuven (LEUVEN), University of Cambridge (UC), California Institute of Technology (CALTECH), Social Brain Lab BCN NIC UMC Groningen (SBL), San Diego State University (SDSU), University of Miami (UM), and Max Planck Institute for Human Cognitive and Brain Sciences (MAX_MUN). The original dataset contains 1112 subjects. This work employed 1035 subjects with complete phenotypic and imaging information. The final dataset consists of 505 ASD and 530 typically developing (TD) controls. The preprocessed functional data were obtained from the Preprocessed Connectomes Project (PCP) repository41,42, specifically using derivatives generated through the Configurable Pipeline for the Analysis of Connectomes (CPAC) with the filt_global strategy (global signal regression). ROI-based time-series data corresponding to multiple atlases (CC200, AAL, Dosenbach160, EZ, TT, and HO) were directly utilized from the provided .1D derivative files.

Pairwise Pearson correlation was applied to standardized ROI time-series signals to construct functional connectivity (FC) matrices. The upper triangular elements of these matrices were extracted and vectorized to form atlas-specific FC feature vectors. These atlases provide complementary perspectives of brain organization such as anatomical, functional and cytoarchitectonic representations. This helps in enabling a richer characterization of brain connectivity patterns. Diagnostic labels were obtained from the phenotypic metadata. It consists of a column named “DX_GROUP” which defines the ground truth (1: ASD, 2: TD). These labels were converted into a binary format, with ASD as ‘0’ and TD as ‘1’ for carrying out the training process. These labels were associated with the corresponding connectivity features to construct input–target pairs. Subjects with missing diagnostic labels were excluded from the analysis.

Let Inline graphic denote the number of ROIs in atlas a. The FC matrix for each atlas is represented as a symmetric matrixInline graphic, where each element corresponds to the Pearson correlation coefficient between ROI time series. To eliminate redundancy, only the upper triangular elements of the FC matrix are retained. This resulted in a feature vector of size Inline graphic for each atlas. For each atlas combination experiment, the dataset is randomly partitioned into training (70%), validation (15%), and testing (15%) subsets using a stratified splitting strategy to preserve class distribution, thereby preventing site-specific data leakage. A fixed random seed (42) is used to ensure reproducibility across all experiments.

Multi-atlas feature representation

The study incorporates multiple brain atlases such as AAL, CC200, Dosenbach160, EZ, HO and TT to capture complementary aspects of brain organization such as anatomical structure, functional connectivity and cytoarchitectonic properties. Each atlas provides a unique perspective on brain parcellation. It helps in the extraction of diverse and informative features. This multi-atlas strategy enhances the representation of complex neural patterns and improves the ability to detect subtle abnormalities associated with ASD. It leads to more robust and reliable classification performance compared to conventional single-atlas approaches. The number of ROIs and corresponding feature dimensions vary across atlases. This enables the model to capture brain connectivity patterns at different spatial resolutions as outlined in Table 2.

Table 2.

Summary of the brain atlases used in this study.

Atlas ROIs Feature dimension Representation type Regions captured Relevance to ASD analysis
AAL 116 6670 Anatomical Divides the brain into anatomically defined regions based on structural boundaries Helps analyze structural-functional disruptions commonly observed in ASD
CC200 200 19900 Functional Clusters brain regions based on functional similarity from rs-fMRI signals Captures fine-grained functional connectivity patterns altered in ASD
Dosenbach160 161 12880 Functional Defines ROIs based on task-control and functional network regions Useful for studying attention, control networks, and cognitive impairments in ASD
EZ 116 6670 Cytoarchitectonic Based on cellular structure and cortical microarchitecture Enables analysis at a finer biological level relevant to neurodevelopmental disorders
HO 111 6105 Anatomical Probabilistic atlas derived from structural MRI data Provides robust anatomical reference for cross-subject consistency
TT 97 4656 Anatomical Classical coordinate-based brain atlas using standardized space Facilitates spatial normalization and comparison across studies

It outlines the number of ROIs, feature dimensionality, representation type, the type of brain regions captured by each atlas and their relevance to ASD analysis.

Each atlas captures brain organization at a different spatial scale. Thus, the integration of multi-atlases enables multi-scale feature representation learning. This method combines fine-grained local connectivity patterns with coarse-grained global brain network structures. The atlas-specific FC feature vectors are independently projected into a shared embedding space and subsequently fused using a Transformer-based attention mechanism. This results in improved discriminative capability for ASD classification.

Functional connectivity (FC) extraction

Functional connectivity represents statistical relationships between spatially separated brain regions. This is extensively used to depict neural communication patterns in rs-fMRI data. In this study, FC vectors are computed from ROI time series obtained through the PCP41 using Pearson correlation. For each subject, the brain is parcellated into ROIs using six brain atlases such as AAL, CC200, Dosenbach160, EZ, HO and TT. Each atlas captures the neural activity at different levels of granularity. These atlases provide a unique spatial partitioning of the brain. The corresponding FC representations are derived from ROI-level time-series using standard preprocessing pipelines. These are provided as vectorized connectivity features. For each atlas, a symmetric FC matrix is constructed. But only the upper triangular elements are extracted to form the feature vector. This eliminates the redundant information. These connectivity features represent pairwise statistical dependencies between brain regions and capture large-scale brain network interactions associated with ASD. The resulting connectivity feature vectors vary in dimensionality as each atlas defines a different number of ROIs. These atlas-specific FC feature vectors are independently used as inputs to the embedding layers of the proposed MACAFNet framework.

Multi-atlas context-aware fusion network (MACAFNet) module

MACAFNet is designed to combine connectivity representations obtained from multiple brain atlases to capture complementary functional patterns associated with ASD. The atlas-specific connectivity features are transformed into compact embeddings using linear projection layers with ReLU activation and dropout regularization. These embeddings capture high-level representations of FC patterns while reducing redundancy and computational complexity. The embeddings from all atlases are then stacked and processed using a Transformer-based attention mechanism.

This module enables the model to capture relationships between different brain atlases by leveraging an attention-based weighting mechanism. It adaptively emphasizes informative connectivity patterns while suppressing less relevant features. This process helps in enhancing the feature representation. A Transformer encoder is employed to model inter-atlas dependencies it consists of a single layer with four attention heads, an embedding dimension of 64 and a dropout rate of 0.4. This configuration provides a balance between model capacity and generalization. The outputs from the Transformer are subsequently aggregated using mean pooling to obtain a unified feature representation for final classification. This final representation consists of individual atlas information and their relationships. This gives a clear understanding of brain connectivity. A dropout layer is employed to reduce overfitting. ReLU activation is used in the embedding layers to help the model learn more complex relationships from the input features. Figure 2 outlines the workflow of MACAFNet fusion.

Fig. 2.

Fig. 2

MACAFNet architecture showing atlas-wise embedding, Transformer-based cross-atlas attention, feature fusion via pooling, and fully connected layers for ASD classification.

Classification and model training

Once the multi-atlas feature fusion is completed, the integrated representation is passed through fully connected layers to classify each subject into ASD or typically developing (TD) categories for research analysis. Each fused feature vector is processed through a ReLU-activated dense layer. It is followed by a dropout layer with a rate of 0.4 to reduce overfitting and improve generalization. The final output layer produces logits for ASD classification. The model is optimized using binary cross-entropy with logits loss (BCEWithLogitsLoss) with class weighting to address class imbalance. It ensures numerical stability during training. To address class imbalance in the dataset, a class weighting strategy is applied based on the distribution of ASD and TD samples in the training set. Adam optimizer is employed to provide efficient adaptive learning rate updates for stable convergence. The initial learning rate is set to 0.001.

Training is performed using mini-batch gradient descent with a batch size of 32. This batch size is chosen to balance computational efficiency and stable gradient estimation under limited dataset conditions. The model is trained for a maximum of 50 epochs to ensure sufficient convergence. To prevent overfitting, early stopping with a patience of 7 epochs is employed. Moreover, a ReduceLROnPlateau scheduler reduces the learning rate by a factor of 0.5 when validation performance does not improve for three consecutive epochs. A stratified train–validation–test split is performed while ensuring that subjects from the same acquisition site are not distributed across different splits, thereby preventing site-specific data leakage. The model achieving the highest validation F1-score is selected as the best model. The held-out test set is used only for final unbiased evaluation. This is done to ensure no data leakage.

For each experiment, the dataset is split into training (70%), validation (15%), and testing (15%) subsets using a stratified sampling strategy to preserve class distribution. Experiments are conducted across all 63 possible combinations of the six brain atlases, ranging from single-atlas models to full multi-atlas fusion. During training, the network learns discriminative functional connectivity (FC) patterns from each atlas while MACAFNet explicitly models inter-atlas contextual dependencies using a lightweight Transformer encoder, enabling dynamic weighting of atlas-specific connectivity patterns rather than static fusion. This integrated design enables MACAFNet to capture complex brain connectivity patterns. This results in an enhanced ASD classification performance. To address class imbalance and enhance classification robustness, the decision threshold is not fixed at 0.5. Instead, the optimal threshold is determined by maximizing the F1-score on the validation set. During inference, the output logits are passed through a sigmoid activation function. The optimized threshold is applied to obtain final ASD/TD predictions on the test set. All experiments are conducted with a fixed random seed of 42 to ensure reproducibility of results.

Mathematical modeling of the proposed architecture

The overall pipeline is mathematically formulated to systematically transform raw rs-fMRI signals into discriminative multi-atlas representations, followed by attention-based fusion using the proposed MACAFNet architecture. Let the dataset consist of N subjects. For each subject, functional connectivity matrices are extracted from six brain atlases.

Data preparation

Raw rs-fMRI signals are preprocessed and organized into labeled samples. Each subject is associated with multiple atlas-based representations. Preprocessing includes denoising, filtering and motion correction. The complete dataset D after preprocessing and organizing is given in Eq. (1).

graphic file with name d33e927.gif 1

where, Inline graphic represents total number of subjects, Inline graphicis rs-fMRI data of subject s and Inline graphic denotes the label (0 = Control, 1 = ASD).

Dataset splitting

The dataset is evaluated using stratified split to ensure reliable and unbiased performance estimation as represented in Eq. (2). This split helps to preserve class distribution across all subsets. Further, it ensures balanced and unbiased model training and evaluation.

graphic file with name d33e952.gif 2

ROI time-series extraction

For each subject and atlas, brain activity is represented as time-series signals extracted from predefined ROIs, capturing temporal neural dynamics. Brain signals are extracted from regions of interest across time points using Eq. (3).

graphic file with name d33e963.gif 3

where, Inline graphic denotes ROI time-series matrix for subject s, Inline graphicrepresents the number of ROIs (depending on the atlas) and T denotes the number of time frames.

Feature vector construction

The upper triangular elements of the symmetric connectivity matrix are extracted and vectorized to form a compact feature representation by removing redundant and self-correlated values, which is denoted in Eq. (4).

graphic file with name d33e987.gif 4

where, Inline graphic denotes the feature vector for subject s, Inline graphic is the extraction of upper triangle and Inline graphic represents the vectorization operator.

Feature normalization

Feature normalization helps to ensure consistent feature scaling across subjects. It is represented in Eq. (5).

graphic file with name d33e1015.gif 5

Multi-atlas feature representation

Multiple atlas-specific feature vectors are generated for each subject. Features from multiple atlases are combined into a unified representation as represented in Eq. (6).

graphic file with name d33e1026.gif 6

where, M represents the number of atlases, Inline graphic is the feature vector from atlas ‘i’, i ∈ {1, 2, …, M} are the atlas index and Inline graphic is the fused feature vector.

Feature embedding

Each atlas-specific feature vector is transformed into a common latent space using a linear projection followed by a non-linear activation function. This ensures that features from different atlases have a consistent representation. This enables effective learning of inter-atlas relationships. Equation (7) represents the embedded feature.

graphic file with name d33e1050.gif 7

where, Inline graphic​represents the embedded feature, Inline graphic is the non-linear Rectified Linear Unit activation function, Inline graphic​ is the learnable weight matrix, Inline graphic​ represents the input feature vector from atlas ‘i’ and Inline graphic​ is the bias term.

Sequence formation

The embedded feature vectors from all atlases are stacked to form an ordered sequence, which serves as input to the attention mechanism. This sequence representation given in Eq. (8) allows the model to learn relationships and dependencies across different atlases.

graphic file with name d33e1083.gif 8

where, Inline graphic is the embedding matrix and Inline graphicis the embedded vector.

Query, key, value projection

The embedding matrix is linearly transformed into Query, Key, and Value representations using learnable weight matrices. These projections enable the model to compute attention scores and capture relationships between atlas features. The formula for transformation is given in Eqs. (9), (10) and (11) respectively.

graphic file with name d33e1110.gif 9
graphic file with name d33e1114.gif 10
graphic file with name d33e1118.gif 11

where, Inline graphic denotes the query matrix (represents “what to attend”), Inline graphicdenotes the key matrix (represents “what is available”), Inline graphicdenotes the value matrix (represents “actual information”), Inline graphicdenotes the input embedding matrix (stacked atlas features), Inline graphic are the learnable weight matrix for Query, Key and Values respectively,

Attention computation

Attention weights are computed by measuring the similarity between Query and Key representations. It is followed by normalization using the SoftMax function. These weights determine the relative importance of each atlas feature. These computed weights are then applied to the Value representations to generate context-aware features that capture inter-atlas relationships. It is given in Eqs. (12) and (13).

graphic file with name d33e1155.gif 12
graphic file with name d33e1159.gif 13

where, Inline graphicrepresents the attention weight matrix, Inline graphic, Similarity matrix between Query and Key, Inline graphic ​denotes the scaling factor to stabilize gradients, Inline graphic: Attention output (context-aware representation) and Inline graphic denotes the value matrix.

Multi-head attention

Multi-head attention applies multiple parallel attention mechanisms to the same input. It allows the model to learn diverse and complementary inter-atlas relationships. Each attention head focuses on different aspects of the feature interactions. Finally, their outputs are combined to form a richer and more informative representation as given by Eq. (14).

graphic file with name d33e1192.gif 14

where, Inline graphic denotes the final multi-head attention output, Inline graphic ​is the output of the attention heads, Inline graphic is the output projection matrix, H is the number of attention heads and Inline graphic is the Concatenated output.

Context-aware fusion

The attention outputs from all atlases are aggregated to form a single unified feature vector that captures both individual atlas characteristics and their interrelationships. This fusion step represented in Eq. (15) integrates context-aware information into a compact representation for effective classification.

graphic file with name d33e1221.gif 15

where, Inline graphic: Final fused feature vector.

Classification layer

The fused feature vector Inline graphic is transformed through a fully connected layer with ReLU activation to obtain a high-level hidden representation Inline graphic, which is given in Eq. (16).

graphic file with name d33e1246.gif 16

where, Inline graphicis the weight matrix, Inline graphic is Bias vector of the first fully connected layer.

The hidden representation Inline graphic is projected through a linear layer to produce logits Inline graphic​ for binary ASD classification, which is given in Eq. (17).

graphic file with name d33e1273.gif 17

where, Inline graphic is the weight matrix of the output layer, Inline graphic is the bias vector of the output layer.

Loss function

The model is optimized using a weighted binary cross-entropy with logits loss, which combines sigmoid activation and cross-entropy in a numerically stable formulation. Class weights are incorporated to address imbalance between ASD and TD samples, ensuring balanced learning and improved sensitivity to minority class predictions. The loss function Inline graphic is represented in Eq. (18).

graphic file with name d33e1298.gif 18

where, Inline graphic is the contribution when true label is ASD,

Inline graphic is the contribution when true label is Control and Inline graphic represents the positive class weight.

Optimization process (Adam optimizer)

Adaptive moment estimation (Adam) is an efficient optimization procedure. It combines momentum and adaptive learning rates. Momentum supports smooth gradients via a moving average. Adaptive learning rates adjust the learning rate for each parameter on the basis of its history. The first momentum called the mean Inline graphic, defines the exponential moving average of the gradients. It reduces noise updates and allows faster convergence. It is given by Eq. (19).

graphic file with name d33e1328.gif 19

where Inline graphic refers to the decay rate (i.e., 0.9 for the momentum term), Inline graphicis the exponential moving average of past momentum, and Inline graphic refers to the gradient of the loss at time step t.

The second moment called the variance Inline graphic tracks of the magnitude of the gradient over time. It is shown in Eq. (20).

graphic file with name d33e1358.gif 20

where Inline graphic represents the decay rate (i.e.) 0.999 for variance term and Inline graphic is the exponential moving average of the squared gradients.

The bias-corrected first moment Inline graphic helps in removing the initialization bias in the early stage of the training steps. It is initialized to zero when the training starts. The correction bias is adjusted via Inline graphic. It is given in Eq. (21).

graphic file with name d33e1385.gif 21

The bias-corrected second moment Inline graphic helps in removing the initial bias toward zero. Inline graphic helps to correct bias which is defined via Eq. (22).

graphic file with name d33e1402.gif 22

With respect to the parameter update Inline graphic, the model weights are updated. It integrates the momentum and adaptive learning rate to produce a balanced and fast converging optimization step as defined in Eq. (23).

graphic file with name d33e1415.gif 23
graphic file with name d33e1419.gif

Model selection

The optimal model is selected based on best F1-score validation performance. It is denoted in Eq. (24) is identified by maximizing the validation F1-score, ensuring a balanced trade-off between precision and recall for ASD classification.

graphic file with name d33e1429.gif 24

Proposed methodology

This section presents the proposed methodology for ASD classification using multi-atlas FC features derived from rs-fMRI data. The framework consists of a structured pipeline where FC matrices are extracted from six different brain atlases. It is followed by atlas-specific feature embedding, contextual feature learning through a Transformer encoder, multi-atlas fusion and final classification. Brain atlases were employed to obtain complementary FC representations of brain networks. For each atlas, ROI-wise time-series signals were extracted and pairwise Pearson correlation coefficients were computed to generate connectivity matrices. These matrices were transformed into feature vectors by extracting the upper triangular elements of the symmetric FC matrices to remove redundancy and projected into a shared embedding space. The proposed MACAFNet framework integrates atlas-specific features using a Transformer-based attention mechanism with a single encoder layer and four attention heads that captures contextual dependencies across different atlas representations. This enables the model to learn complementary connectivity patterns associated with ASD. The fused representation is then passed through a fully connected classification layer to predict whether a subject belongs to the ASD or typically developing class. The final output consists of logits, which are converted to probabilities using a sigmoid activation function during inference. The detailed workflow of the proposed approach is summarized in Algorithm 1.

Algorithm 1.

Algorithm 1

MACAFNet multi-atlas ASD classification model

Figure generation and visualization tools

All figures in this study were generated using a combination of tools. The architectural diagrams were created using draw.io (diagrams.net) for clear schematic representation of the proposed framework. Performance plots, including accuracy curves, loss curves, and ROC curves, were generated using Python-based visualization libraries, specifically Matplotlib and Scikit-learn visualization utilities. Data processing and visualization were performed in a Python 3.8 environment.

Results

This section presents the experimental evaluation of the proposed MACAFNet for ASD classification using multi-atlas FC features derived from rs-fMRI data. All experiments were conducted on an NVIDIA GPU server (NVIDIA-SMI 535.183.01) equipped with four NVIDIA RTX A6000 GPUs (48 GB VRAM each), running CUDA 12.2.

Comparison with existing studies

The performance of the proposed MACAFNet framework was compared with several prior studies that utilized ABIDE I dataset consisting of 1,035 subjects for ASD classification. Table 3 outlines the performance measure of the existing models and the proposed MACAFNet framework in terms of test accuracy. AE + DNN25, CNN-based model43 and AE + SLP44 produced an accuracy of 70% approximately, while SAE + MLP45 produced 70.8% and MADE-for-ASD27 produced 75.2%. The proposed MACAFNet achieves an accuracy of 88.46% for the atlas combination CC200 + EZ + TT. This improvement demonstrates the effectiveness of leveraging complementary connectivity information from multiple brain parcellations. It also emphasizes the benefit of context-aware multi-atlas fusion in capturing discriminative neural patterns associated with ASD. Furthermore, MACAFNet dynamically models inter-atlas dependencies leading to improved generalization. The results clearly indicate that MACAFNet outperforms existing approaches by providing a more reliable and robust framework for ASD classification model. The reported performance corresponds to the best-performing model selected based on the highest validation F1-score (0.9252) during training with early stopping, and evaluated on a held-out test set.

Table 3.

Comparation of performance of the proposed MACAFNet framework with the existing ASD classification methods on the ABIDE I CPAC dataset.

Study Model Number of subjects Accuracy (%)
Heinfield et al. 25 AE + DNN 1035 70
Sherkatghanad et al.43 CNN 1035 70.22
Eslami et al.44 AE + SLP 1035 70.3
Almuqhim et al.45 SAE + MLP 1035 70.8
Liu et al. 27 MADE-for-ASD 1035 75.2
Proposed MACAFNet MACAFNet + Multi-Atlas Context-Aware Fusion 1035 88.46

Ablation study: effect of multi-atlas combinations

An ablation study was conducted using all possible 63 combinations of six brain atlases to evaluate the contribution of individual atlases and their combinations. The combinations range from single-atlas models to multi-atlas fusion. The results demonstrate that multi-atlas fusion consistently outperforms single-atlas models across all metrics. Among all combinations, CC200 + EZ + TT achieved the highest performance, with a test accuracy of 88.46%, F1-score of 0.8448, and AUC of 0.9546, while maintaining balanced precision (0.8448) and recall (0.8448). High sensitivity (0.8448) and specificity (0.9082) indicate effective detection of both ASD and TD subjects. The performance gain is not strictly monotonic with the number of atlases, indicating that the quality and complementarity of selected atlases are more important than simply increasing the number of inputs. The performance metrics of the top 10 multi-atlas combinations are shown in Table 4. All results reported in this study correspond to models trained using a stratified 70%–15%–15% split (training–validation–test), with class distribution preserved across splits. The best model for each atlas combination is selected based on validation F1-score and evaluated on the held-out test set to ensure fair and unbiased comparison.

Table 4.

Performance metrics of the top 10 multi-atlas combinations used for ablation study.

Atlas combinations Test accuracy Training accuracy Validation accuracy Precision F1-score AUC Sensitivity Specificity
CC200 + EZ + TT 0.8846 0.9876 0.9290 0.8448 0.8448 0.9546 0.8448 0.9082
AAL + CC200 + HO 0.8782 0.9848 0.9097 0.8679 0.8288 0.9229 0.7931 0.9286
AAL + CC200 + Dosenbach160 + HO + TT 0.8718 0.9834 0.9226 0.7879 0.8387 0.9372 0.8466 0.8571
AAL + CC200 + Dosenbach160 + EZ + TT 0.8654 0.9820 0.9032 0.7937 0.8264 0.9367 0.8621 0.8673
AAL + CC200 0.8590 0.9751 0.9032 0.8000 0.8136 0.9263 0.8276 0.8776
AAL + CC200 + EZ + HO + TT 0.8526 0.9738 0.9161 0.7869 0.8067 0.9340 0.8276 0.8673
AAL + Dosenbach160 + EZ 0.8462 0.9682 0.9484 0.7361 0.8154 0.9305 0.8138 0.8061
AAL + EZ + HO 0.8397 0.9862 0.9032 0.7143 0.8148 0.9335 0.8483 0.7755
Dosenbach160 + HO 0.8269 0.9751 0.9161 0.7123 0.7939 0.9353 0.8448 0.7857

Significant values are in bold.

The ablation study highlights several key findings:

  • Single atlases such as CC200 provide strong baseline performance (accuracy 87.18%, F1-score 0.8438), confirming their informative connectivity representation.

  • Adding complementary atlases improves recall and sensitivity, showing enhanced detection of ASD subjects.

  • Precision-recall trade-offs are observed across different atlas combinations, emphasizing the value of context-aware fusion to balance false positives and false negatives.

  • The proposed attention-based fusion mechanism effectively learns optimal weighting across atlas features, avoiding redundancy and enhancing discriminative feature learning.

Additionally, to handle class imbalance and improve robustness, the classification threshold is optimized by maximizing the F1-score on predicted probabilities rather than using a fixed threshold of 0.5. This strategy contributes to improved balance between sensitivity and specificity across atlas combinations.

The training and validation accuracy and loss curves for the CC200 + EZ + TT atlas combination outlined in Fig. 3 demonstrates smooth convergence behavior. A strong generalization capability is obtained with minimal divergence between training and validation accuracy. The training loss and validation loss gradually decreased highlighting effective optimization, model convergence, reduced overfitting and improved generalization. The ROC curve depicted in Fig. 4 shows an AUC of 0.9546, confirming excellent class separability. This high AUC indicates strong discriminative ability and robust ranking performance across varying decision thresholds.

Fig. 3.

Fig. 3

Training and validation accuracy (a) and loss (b) curves of the MACAFNet model for the CC200 + EZ + TT atlas combination. The curves show smooth convergence and minimal overfitting.

Fig. 4.

Fig. 4

ROC curve of the MACAFNet model (CC200 + EZ + TT), with AUC = 0.9546, demonstrating strong discriminative ability for ASD classification.

Hyperparameter tuning performed using various ranges of batch size and epoch revealed that there was a slight decrease in the performance of the model for different batch sizes of {4, 8, 16, 64} and epochs between (40, 100). Specifically, smaller batch sizes introduced noisy gradient updates, while larger batch sizes and increased epochs led to overfitting. The model produced peak performance at a batch size of 32 and epoch of 50. This indicates an optimal balance between convergence stability and generalization.

Error analysis and model performance

The performance of the top MACAFNet model (CC200 + EZ + TT) was further analyzed using the confusion matrix to understand misclassifications between ASD and TD subjects, which is depicted in Fig. 5. The confusion matrix corresponds to the best-performing model (selected using validation F1-score) evaluated on the held-out test set.

Fig. 5.

Fig. 5

Confusion matrix of the MACAFNet model (CC200 + EZ + TT), showing the distribution of true and predicted classes for ASD and TD subjects.

The performance of the top MACAFNet model (CC200 + EZ + TT) was further analyzed using the confusion matrix to understand misclassifications between ASD and TD subjects. From the results it is evident that the model correctly classified 89 TD and 49 ASD subjects. It misclassified 9 TD subjects as ASD and 9 ASD subjects as TD. This indicates that the model maintains a balanced performance across both classes. It effectively detects ASD subjects while minimizing false positives among TD subjects. The error analysis highlights that the proposed MACAFNet achieved high sensitivity (0.8448) and specificity (0.9082). Despite strong performance, a small number of misclassifications may arise due to inter-subject variability, site-specific differences in data acquisition and overlapping connectivity patterns between ASD and TD groups. Further, training and validation curves exhibit smooth convergence with minimal divergence, indicating stable learning and good generalization. The top-performing atlas combination achieved a test accuracy of 88.46%, F1-score of 0.8448, and AUC of 0.9546. This indicates consistent and robust predictive capability.

Discussion

The experimental results demonstrate that the proposed MACAFNet effectively leverages complementary FC information from multiple brain atlases to improve ASD classification. Multi-atlas fusion consistently outperforms single-atlas models across all metrics. The CC200 + EZ + TT combination achieved the highest accuracy of 88.46%, with a recall of 0.8448 and an AUC of 0.9546, indicating strong discriminative capability and generalization. A key observation is that different atlases capture distinct aspects of brain organization. The superior performance of CC200 + EZ + TT is due to their complementarity, where CC200 captures fine-grained functional connectivity, EZ provides cytoarchitectonic information, and TT contributes stable anatomical structure. This combination enables the model to learn multi-level representations of brain connectivity. Importantly, increasing the number of atlases does not always improve performance; instead, selecting complementary atlases is more effective. The Transformer-based attention mechanism further enhances performance by learning inter-atlas dependencies and emphasizing the most informative features while suppressing noise. The model’s robustness is supported by effective training strategies such as dropout, early stopping, and learning rate scheduling, which improve convergence and reduce overfitting. From a clinical perspective, the achieved sensitivity and specificity indicate strong potential for reliable ASD classification. Although the proposed model achieves strong performance, variability across acquisition sites in ABIDE remains a challenge. Future work will explore domain adaptation and site-invariant learning strategies.

Conclusion

This study presented a MACAFNet for ASD classification using functional connectivity features derived from rs-fMRI data. The framework effectively integrates connectivity representations from multiple brain atlases and leverages a transformer-based fusion mechanism to capture contextual dependencies among atlas-specific features. Experimental evaluation across 63 atlas combinations, using validation-guided model selection and held-out test evaluation, demonstrated that multi-atlas fusion consistently outperforms single-atlas models. Notably, the CC200 + EZ + TT combination achieved the highest performance, with an accuracy of 88.46% and an F1-score of 0.8448. The results confirm that combining complementary atlas representations enables richer and more discriminative feature learning, while the attention-based fusion mechanism enhances the modeling of inter-atlas relationships. Overall, MACAFNet provides a robust and scalable approach for ASD detection using rs-fMRI data. Experimental evaluation across 63 atlas combinations, using validation-guided model selection and held-out test evaluation, demonstrated that multi-atlas fusion consistently outperforms single-atlas models. Despite the promising results, the study is limited by dataset heterogeneity and sample size constraints inherent to ABIDE-I. Future work will focus on integrating multi-modal neuroimaging data, exploring graph-based connectivity representations, and validating the framework on larger and more diverse datasets. These advancements are expected to further improve model generalizability and clinical applicability.

Acknowledgements

The authors would like to acknowledge the support and resources provided by Karpagam College of Engineering, which were invaluable for the completion of this research.

Author contributions

D.A. conceived the experiment(s). D.A. and S.K. developed the methodology and conducted the experiment(s). S.K. performed the formal analysis and visualized the results. D.A. and S.K. carried out the investigation and curated the data. D.A. and S.K. were responsible for software implementation, validation, and resource management. D.A. prepared the original draft, and S.K. reviewed and edited the manuscript. S.K. supervised the project and handled project administration. D.A. and S.K. acquired funding. All authors reviewed and agreed to the final version of the manuscript.

Data availability

The preprocessed functional data were obtained from the Preprocessed Connectomes Project (PCP) repository (http://preprocessed-connectomes-project.org/abide/), specifically from the ABIDE I release using the Configurable Pipeline for the Analysis of Connectomes (CPAC) with the filt_global preprocessing strategy. The dataset is publicly available for research purposes under the data usage terms of the PCP initiative, and appropriate citations have been followed. The complete code and analysis scripts used to process the data and reproduce the results are publicly available via Zenodo at: https://doi.org/10.5281/zenodo.19563123.

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Lord, C. et al. Autism spectrum disorder. Nat. Rev. Dis. Primers. 6, 5. 10.1038/s41572-019-0138-4 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Zwaigenbaum, L. et al. Early intervention for children with autism spectrum disorder under 3 years of age: recommendations for practice and research. Pediatrics136 (Suppl 1), S60–S81. 10.1542/peds.2014-3667E (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Pierce, K. & Courchesne, J. Evidence for a cerebellar role in reduced exploration and stereotyped behavior in autism. Biol. Psychiatry. 49, 655–664. 10.1016/S0006-3223(00)01008-8 (2001). [DOI] [PubMed] [Google Scholar]
  • 4.Aarthi, D. & Kannimuthu, S. A comprehensive analysis of autism spectrum disorder using machine learning algorithms: survey. In: International Conference on Power Engineering and Intelligent Systems, Delhi, India. 10.1007/978-981-99-7216-6_20 (2024).
  • 5.McKenzie, K. et al. Factors influencing waiting times for diagnosis of autism spectrum disorder in children and adults. Res. Dev. Disabil.45–46, 300–306. 10.1016/j.ridd.2015.07.033 (2015). [DOI] [PubMed] [Google Scholar]
  • 6.Schielen, S. J. C., Pilmeyer, J., Aldenkamp, A. P. & Zinger, S. The diagnosis of ASD with MRI: a systematic review and meta-analysis. Transl. Psychiatry. 14, 318. 10.1038/s41398-024-03024-5 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Zhang, S. AI-assisted early screening, diagnosis, and intervention for autism in young children. Front. Psychiatry. 16, 1513809. 10.3389/fpsyt.2025.1513809 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Wang, M., Xu, D., Zhang, L. & Jiang, H. Application of multimodal MRI in the early diagnosis of autism spectrum disorders: a review. Diagnostics13, 3027. 10.3390/diagnostics13193027 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Anagnostou, E. & Taylor, M. J. Review of neuroimaging in autism spectrum disorders: what have we learned and where we go from here. Mol. Autism. 2, 4. 10.1186/2040-2392-2-4 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Alves, C. L. et al. Diagnosis of autism spectrum disorder based on functional brain networks and machine learning. Sci. Rep.13, 8072. 10.1038/s41598-023-34650-6 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Gkintoni, E. et al. Leveraging AI-driven neuroimaging biomarkers for early detection and social function prediction in autism spectrum disorders: a systematic review. Healthcare13, 1776. 10.3390/healthcare13151776 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Santana, C. P. et al. rs-fMRI and machine learning for ASD diagnosis: a systematic review and meta-analysis. Sci. Rep.12, 6030. 10.1038/s41598-022-09821-6 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Varshney, R. K., Katiyar, A. & Johri, P. Hybrid CNN–RNN models for multimodal analysis of autism spectrum disorder neuroimaging. In: 2025 International Conference on Automation and Computation (AUTOCOM), Dehradun, India, pp 155–160. 10.1109/AUTOCOM64127.2025.10956945 (2025).
  • 14.Mahmoudi, E. et al. Gamification in mobile apps for children with disabilities: scoping review. JMIR Serious Games. 12, e49029. 10.2196/49029 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Shahini, A. et al. A systematic review for artificial intelligence-driven assistive technologies to support children with neurodevelopmental disorders. Inf. Fusion. 12410.1016/j.inffus.2025.103441 (2025).
  • 16.Yang, Z., Zhang, Y., Ning, J., Wang, X. & Wu, Z. Early diagnosis of autism: a review of video-based motion analysis and deep learning techniques. IEEE Access.13, 2903–2928. 10.1109/ACCESS.2024.3523872 (2025). [Google Scholar]
  • 17.Trayvick, J. et al. Speech and language patterns in autism: towards natural language processing as a research and clinical tool. Psychiatry Res.340, 116109. 10.1016/j.psychres.2024.116109 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Abraham, A. et al. Deriving reproducible biomarkers from multi-site resting-state data: An Autism-based example. NeuroImage147, 736–745. 10.1016/j.neuroimage.2016.10.045 (2017). [DOI] [PubMed] [Google Scholar]
  • 19.Bassett, D. S. & Sporns, O. Network neuroscience. Nat. Neurosci.20, 353–364. 10.1038/nn.4502 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Jung, W., Jeon, E., Kang, E. & Suk, H. I. EAG-RS: A novel explainability-guided ROI-selection framework for ASD diagnosis via inter-regional relation learning. IEEE Trans. Med. Imaging. 43 (4), 1400–1411. 10.1109/TMI.2023.3337362 (2024). [DOI] [PubMed] [Google Scholar]
  • 21.Ning, Q. et al. A deep learning method for autism spectrum disorder identification based on interactions of hierarchical brain networks. Behav. Brain Res.452, 114603. 10.1016/j.bbr.2023.114603 (2023). [DOI] [PubMed] [Google Scholar]
  • 22.Hazlett, H. C. et al. Early brain development in infants at high risk for autism spectrum disorder. Nature542, 348–351. 10.1038/nature21369 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Bengs, M., Gessert, N. & Schlaefer, A. 4D spatio-temporal deep learning with 4D fMRI data for autism spectrum disorder classification. arXiv preprint. 10.48550/arXiv.2004.10165 (2020). [Google Scholar]
  • 24.Sachdeva, J., Mittal, R., Mehta, J., Jain, R. & Ranjan, A. Resolving autism spectrum disorder (ASD) through brain topologies using fMRI dataset with multi-layer perceptron (MLP). Psychiatry Res. Neuroimag.. 343, 111858. 10.1016/j.pscychresns.2024.111858 (2024). [DOI] [PubMed] [Google Scholar]
  • 25.Heinsfeld, A. S., Franco, A. R., Craddock, R. C., Buchweitz, A. & Meneguzzi, F. Identification of autism spectrum disorder using deep learning and the ABIDE dataset. NeuroImage Clin.17, 16–23. 10.1016/j.nicl.2017.08.017 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Parisot, S. et al. Disease prediction using graph convolutional networks: Application to Autism Spectrum Disorder and Alzheimer’s disease. Med. Image Anal.48, 117–130. 10.1016/j.media.2018.06.001 (2018). [DOI] [PubMed] [Google Scholar]
  • 27.Liu, X., Hasan, M. R., Gedeon, T. & Hossain, M. Z. MADE-for-ASD: A multi-atlas deep ensemble network for diagnosing Autism Spectrum Disorder. Comput. Biol. Med.182, 109083. 10.1016/j.compbiomed.2024.109083 (2024). [DOI] [PubMed] [Google Scholar]
  • 28.Yang, S. et al. M3ASD: Integrating multi-atlas and multi-center data via multi-view low-rank graph structure learning for Autism Spectrum Disorder diagnosis. Brain Sci.15 (11), 1136. 10.3390/brainsci15111136 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Jung, W., Jeon, E., Kang, E. & Suk, H. I. EAG-RS: A novel explainability-guided ROI-selection framework for ASD diagnosis via inter-regional relation learning. IEEE Trans. Med. Imaging. 43 (4), 1400–1411. 10.1109/TMI.2023.3337362 (2024). [DOI] [PubMed] [Google Scholar]
  • 30.Cilia, F. et al. Computer-aided screening of autism spectrum disorder: eye-tracking study using data visualization and deep learning. JMIR Hum. Factors. 8, e27706. 10.2196/27706 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Kanhirakadavath, M. R. & Chandran, M. S. M. Investigation of eye-tracking scan path as a biomarker for autism screening using machine learning algorithms. Diagnostics12, 518. 10.3390/diagnostics12020518 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Farhat, T. et al. A deep learning-based ensemble for autism spectrum disorder diagnosis using facial images. PLoS One. 20, e0321697. 10.1371/journal.pone.0321697 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar] [Retracted]
  • 33.Anderson, J. T., Stenum, J., Roemmich, R. T. & Wilson, R. B. Validation of markerless video-based gait analysis using pose estimation in toddlers with and without neurodevelopmental disorders. Front. Digit. Health. 7, 1542012. 10.3389/fdgth.2025.1542012 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Jin, L., Cui, H., Zhang, P. & Cai, C. Early diagnostic value of home video-based machine learning in autism spectrum disorder: a meta-analysis. Eur. J. Pediatr.184, 37. 10.1007/s00431-024-05837-4 (2024). [DOI] [PubMed] [Google Scholar]
  • 35.Dia, M., Khodabandelou, G., Md Sabri, A. Q. & Othmani, A. Video-based continuous affect recognition of children with autism spectrum disorder using deep learning. Biomed. Signal. Process. Control. 89, 105712. 10.1016/j.bspc.2023.105712 (2024). [Google Scholar]
  • 36.Ganjigunte Prakash, V. et al. Video-based real-time assessment and diagnosis of autism spectrum disorder using deep neural networks. Expert Syst.10.1111/exsy.13253 (2023). [Google Scholar]
  • 37.Abbas, H. et al. Multi-modular AI approach to streamline autism diagnosis in young children. Sci. Rep.10, 5014. 10.1038/s41598-020-61213-w (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Ma, W., Xu, L., Zhang, H. & Zhang, S. Can natural speech prosody distinguish autism spectrum disorders? a meta-analysis. Behav. Sci.14, 90. 10.3390/bs14020090 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Chi, N. A. et al. Classifying autism from crowdsourced semistructured speech recordings: machine learning model comparison study. JMIR Pediatr. Parent.5, e35406. 10.2196/35406 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.ABIDE I (Autism Brain Imaging Data Exchange). ABIDE I database access. https://fcon_1000.projects.nitrc.org/indi/abide/abide_I.html, (2012).
  • 41.Preprocessed Connectomes Project. Preprocessed Connectomes Project – ABIDE Preprocessed Initiative. http://preprocessed-connectomes-project.org/abide/. (2015).
  • 42.Cameron Craddock, Y. et al. Yan, Pierre Bellec The Neuro Bureau Preprocessing Initiative: open sharing of preprocessed neuroimaging data and derivatives. Neuroinformatics 2013, Stockholm, Sweden. (2013).
  • 43.Sherkatghanad, Z. et al. Automated detection of Autism Spectrum Disorder using a convolutional neural network. Front. Neurosci.13, 1325. 10.3389/fnins.2019.01325 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Eslami, T., Mirjalili, V., Fong, A., Laird, A. R. & Saeed, F. ASD-DiagNet: A hybrid learning approach for detection of Autism Spectrum Disorder using fMRI data. Front. Neuroinform. 13, 70. 10.3389/fninf.2019.00070 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Almuqhim, F. & Saeed, F. ASD-SAENet: A sparse autoencoder and deep-neural network model for detecting Autism Spectrum Disorder (ASD) using fMRI data. Front. Comput. Neurosci.15, 654315. 10.3389/fncom.2021.654315 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The preprocessed functional data were obtained from the Preprocessed Connectomes Project (PCP) repository (http://preprocessed-connectomes-project.org/abide/), specifically from the ABIDE I release using the Configurable Pipeline for the Analysis of Connectomes (CPAC) with the filt_global preprocessing strategy. The dataset is publicly available for research purposes under the data usage terms of the PCP initiative, and appropriate citations have been followed. The complete code and analysis scripts used to process the data and reproduce the results are publicly available via Zenodo at: https://doi.org/10.5281/zenodo.19563123.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES